Hiroshi Saruwatari

dblp:87/629 · DBLP profile ↗
← Back
249ranked-venue papers
15as first author
70since 2021 · last 2026
0000-0003-0876-5617ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 204 · 11 first-author · 56 since 2021Artificial intelligence and machine learning · 134 · 10 first-author · 41 since 2021Systems, architecture and hardware · 8 · 1 first-author
YearPublicationVenuePosition
2026 J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
abstract
Spoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However, existing datasets are often limited in size, spontaneity, or linguistic coherence. To address these limitations, we introduce J-CHAT, a 76,000-hour open-source Japanese spoken dialogue corpus. Constructed using an automated, language-independent methodology, J-CHAT ensures acoustic cleanliness, diversity, and natural spontaneity. The corpus is built from YouTube and podcast data, with extensive filtering and denoising to enhance quality. Experimental results with generative spoken dialogue language models trained on J-CHAT demonstrate its effectiveness for SDS development. By providing a robust foundation for training advanced dialogue models, we anticipate that J-CHAT will drive progress in human-AI dialogue research and applications.
Wataru Nakata, Kentaro Seki, Hitomi Yanaka, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
LREC6
2026 DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
abstract
Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference.
Wataru Nakata, Yuki Saito 0001, Kazuki Yamauchi, Emiru Tsunoo, Hiroshi Saruwatari
SIGDIAL5
2026 Stride conversion algorithms for convolutional layers and its application to sampling-frequency-independent deep neural networks
abstract
We propose interpolation-based algorithms that enable convolutional and transposed convolutional layers to operate with arbitrary (including non-integer) strides. A primary motivation for the proposed algorithms is to maintain a consistent temporal resolution when adapting deep neural networks (DNNs) to different sampling frequencies (SFs). To handle untrained SFs, we previously introduced SF-independent (SFI) convolutional layers, which adjust kernel weights in accordance with the target SF. However, achieving full consistency across SFs also requires the proportional adjustment of the stride, which results in non-integer values in many practical cases. Conventional algorithms for convolutional layers cannot handle such strides directly, and commonly used approaches (e.g., stride rounding or signal resampling) lead to performance degradation. To solve this problem, we propose a feature-domain interpolation framework that constructs continuous-time representations of intermediate features. This enables sampling at arbitrary stride intervals without modifying the network architecture. Through music source separation experiments, we show that the proposed algorithms maintain a strong performance across a range of SFs, including those where the stride becomes non-integer. Our analysis reveals that the proposed algorithms are robust to the choice of interpolation method and are especially effective for sources containing pitched sounds.
Kanami Imamura, Tomohiko Nakamura, Norihiro Takamune, Kohei Yatabe, Hiroshi Saruwatari
Signal Process.5
2026 Speaker-conditioned phrase break prediction for text-to-speech with phoneme-level pre-trained language model
abstract
This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We further demonstrate that these speaker embeddings can capture speaker-related characteristics solely from the phrasing task. Besides, we explore the potential of pre-trained speaker embeddings for unseen speakers through a few-shot adaptation method. Furthermore, we pioneer the application of phoneme-level pre-trained language models to this TTS front-end task, which significantly boosts the accuracy of the phrasing model. Our methods are rigorously assessed through both objective and subjective evaluations, demonstrating their effectiveness. • Speaker-conditioned phrasing model improves accuracy in multi-speaker phrasing tasks. • We explore various speaker embeddings in phrasing models. • We apply phoneme-level pre-trained language models to enhance phrasing accuracy. • We propose a speaker adaptation method for few-shot phrasing tasks. • We verify that speaker embeddings learn human-aligned features via phrasing tasks.
Yuki Saito 0001, Takaaki Saeki, Tomoki Koriyama, Wataru Nakata, Detai Xin, Hiroshi Saruwatari
Speech Commun.7
2025 CAVIARES: Corpus for Audio-Visual Expressive Voice Agent
abstract
High-quality audio-visual corpora are essential for building voice agents capable of natural human-machine communication, but existing corpora commonly contain a limited amount of data per speaker, making personalized modeling difficult. We present CAVIARES, a new audio-visual corpus comprising 9.5 hours of expressive speech recorded by a single professional Japanese female speaker. CAVIARES consists of two subsets: acted dialogue and expressive reading, providing a diverse range of speaking styles for speech-to-facial motion modeling and multimodal learning tasks. In this paper, we describe the construction process of CAVIARES and the results of corpus analysis. CAVIARES will be released for research purposes only.
Jinsheng Chen, Yuki Saito 0001, Naoko Tanji, Hironori Doi, Byeongseon Park, Yuma Shirahata, Kentaro Tachibana, Hiroshi Saruwatari
ASRU9
2025 Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
abstract
We introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer into a self-supervised learning (SSL) model to achieve SFI speech feature extraction for MOS prediction. We present some strategies to improve the MOS prediction performance of our model: distilling knowledge from a pretrained non-SFI-SSL model and pretraining with a large-scale MOS dataset. Our submission to the AMC 2025 Track 3 ranked the first in one evaluation metric and the fourth in the final ranking. We also report the results of our ablation study to investigate essential factors of our model.
Go Nishikawa, Wataru Nakata, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari, Tomohiko Nakamura
ASRU5
2025 Analysing the Language of Neural Audio Codecs
abstract
This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC models, examining their adherence to linguistic statistical laws such as Zipf’s law and Heaps’ law, as well as their entropy and redundancy. To assess how these token-level properties relate to semantic and acoustic preservation in synthesized speech, we evaluate intelligibility using error rates of automatic speech recognition, and quality using the UTMOS score. Our results reveal that NAC tokens, particularly 3-grams, exhibit language-like statistical patterns. Moreover, these properties, together with measures of information content, are found to correlate with improved performances in speech recognition and resynthesis tasks. These findings offer insights into the structure of NAC token sequences and inform the design of more effective generative speech models.
Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando, Yuki Saito 0001, Hiroshi Saruwatari
ASRU6
2025 Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
abstract
Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features of self-supervised learning (SSL) models. This work is the first to incorporate SSL features with causality into an SE model. The causal SSL features are encoded and combined with spectrogram features using feature-wise linear modulation to estimate a mask for enhancing the noisy input speech. Simultaneously, we quantize the causal SSL features using vector quantization to represent phonetic characteristics as semantic tokens. The model not only encodes SSL features but also predicts the future semantic tokens in multi-task learning (MTL). The experimental results using VoiceBank + DEMAND dataset show that our proposed method achieves 2.88 in PESQ, especially with semantic prediction MTL, in which we confirm that the semantic prediction played an important role in causal SE.
Emiru Tsunoo, Yuki Saito 0001, Wataru Nakata, Hiroshi Saruwatari
ICASSP4
2025 RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Yusuke Kanamori, Yuki Okamoto, Taisei Takano, Shinnosuke Takamichi, Yuki Saito 0001, Hiroshi Saruwatari
INTERSPEECH6
2025 Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
abstract
We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike conventional FM modules, which use the coarse representations from the weak generator as conditions, SFM constructs intermediate states along the FM paths from these representations. During training, we introduce an orthogonal projection method to adaptively determine the temporal position of these states, and apply a principled construction strategy based on a single-segment piecewise flow. The SFM inference starts from the intermediate state rather than pure noise, thereby focusing computation on the latter stages of the FM paths. We integrate SFM into multiple TTS models with a lightweight SFM head. Experiments demonstrate that SFM yields consistent gains in speech naturalness across both objective and subjective evaluations, and significantly accelerates inference when using adaptive-step ODE solvers. Demo and codes are available at https://ydqmkkx.github.io/SFMDemo/.
Yiyi Cai, Yuki Saito 0001, Lixu Wang, Hiroshi Saruwatari
NeurIPS5
2024 Diversity-Based Core-Set Selection for Text-to-Speech with Linguistic and Acoustic Features
abstract
This paper proposes a method for extracting a lightweight subset from a text-to-speech (TTS) corpus ensuring synthetic speech quality. In recent years, methods have been proposed for constructing large-scale TTS corpora by collecting diverse data from massive sources such as audiobooks and YouTube. Although these methods have gained significant attention for enhancing the expressive capabilities of TTS systems, they often prioritize collecting vast amounts of data without considering practical constraints like storage capacity and computation time in training, which limits the available data quantity. Consequently, the need arises to efficiently collect data within these volume constraints. To address this, we propose a method for selecting the core subset (known as core-set) from a TTS corpus on the basis of a diversity metric, which measures the degree to which a subset encompasses a wide range. Experimental results demonstrate that our proposed method performs significantly better than the baseline phoneme-balanced data selection across language and corpus size.
Kentaro Seki, Shinnosuke Takamichi, Takaaki Saeki, Hiroshi Saruwatari
ICASSP4
2024 Do Learned Speech Symbols Follow Zipf's Law?
abstract
In this study, we investigate whether speech symbols, learned through deep learning, follow Zipf’s law, akin to natural language symbols. Zipf’s law is an empirical law that delineates the frequency distribution of words, forming fundamentals for statistical analysis in natural language processing. Natural language symbols, which are invented by humans to symbolize speech content, are recognized to comply with this law. On the other hand, recent breakthroughs in spoken language processing have given rise to the development of learned speech symbols; these are data-driven symbolizations of speech content. Our objective is to ascertain whether these datadriven speech symbols follow Zipf’s law, as the same as natural language symbols. Through our investigation, we aim to forge new ways for the statistical analysis of spoken language processing.
Shinnosuke Takamichi, Hiroki Maeda, Joonyong Park, Daisuke Saito, Hiroshi Saruwatari
ICASSP5
2024 Localizing Acoustic Energy in Sound Field Synthesis by Directionally Weighted Exterior Radiation Suppression
abstract
A method for synthesizing the desired sound field while suppressing the exterior radiation power with directional weighting is proposed. The exterior radiation from the loudspeakers in sound field synthesis systems can be problematic in practical situations. Although several methods to suppress the exterior radiation have been proposed, suppression in all outward directions is generally difficult, especially when the number of loudspeakers is not sufficiently large. We propose the directionally weighted exterior radiation representation to prioritize the suppression directions by incorporating it into the optimization problem of sound field synthesis. By using the proposed representation, the exterior radiation in the prioritized directions can be significantly reduced while maintaining high interior synthesis accuracy, owing to the relaxed constraint on the exterior radiation. Its performance is evaluated with the application of the proposed representation to amplitude matching in numerical experiments.
Yoshihide Tomita, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2024 Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
Takuto Igarashi, Yuki Saito 0001, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH7
2024 SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
Takaaki Saeki, Soumi Maiti, Shinnosuke Takamichi, Shinji Watanabe 0001, Hiroshi Saruwatari
INTERSPEECH5
2024 SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
Yuki Saito 0001, Takuto Igarashi, Kentaro Seki, Shinnosuke Takamichi, Ryuichi Yamamoto, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH7
2024 Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
Kentaro Seki, Shinnosuke Takamichi, Norihiro Takamune, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari
INTERSPEECH6
2024 SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
Osamu Take, Shinnosuke Takamichi, Kentaro Seki, Yoshiaki Bando, Hiroshi Saruwatari
INTERSPEECH5
2024 The T05 System for the voicemos challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
abstract
We present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturalness mean opinion score (MOS) for high-quality synthetic speech. In addition to a pretrained self-supervised learning (SSL)-based speech feature extractor, our system incorporates a pretrained image feature extractor to capture the difference of synthetic speech observed in speech spectrograms. We first separately train two MOS predictors that use either of an SSL-based or spectrogram-based feature. Then, we fine-tune the two predictors for better MOS prediction using the fusion of two extracted features. In the VMC 2024 Track 1, our T05 system achieved first place in 7 out of 16 evaluation metrics and second place in the remaining 9 metrics, with a significant difference compared to those ranked third and below. We also report the results of our ablation study to investigate essential factors of our system.
Kaito Baba, Wataru Nakata, Yuki Saito 0001, Hiroshi Saruwatari
SLT4
2024 DNN-Based Ensemble Singing Voice Synthesis With Interactions Between Singers
abstract
We propose a singing voice synthesis (SVS) method for a more unified ensemble singing voice by modeling interactions between singers. Most existing SVS methods aim to synthesize a solo voice, and do not consider interactions between singers, i.e., adjusting one’s own voice to the others’ voices. Since the production of ensemble voices from solo singing voices ignores the interactions, it can degrade the unity of the vocal ensemble. Therefore, we propose a SVS that reproduces the interactions. It is based on an architecture that uses musical scores of multiple voice parts, and loss functions that simulate the interactions’ effect to acoustic features. Experimental results show that our methods improve the unity of the vocal ensemble.
Hiroaki Hyodo, Shinnosuke Takamichi, Tomohiko Nakamura, Junya Koguchi, Hiroshi Saruwatari
SLT5
2024 Cross-Dialect Text-to-Speech In Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level Bert
abstract
We explore cross-dialect text-to-speech(CD-TTS),a task to synthesize learned speakers’voices in non-native dialects,especially in pitch-accent languages.CD-TTS is important for developing voice agents that naturally communicate with people across regions.We present a novel TTS model comprising three sub-modules to perform competitively at this task.We first train a backbone TTS model to synthesize dialect speech from a text conditioned on phoneme-level accent latent variables(ALVs)extracted from speech by a reference encoder. Then,we train an ALV predictor to predict ALVs tailored to a target dialect from input text leveraging our novel multi-dialect phoneme-level BERT.We conduct multi-dialect TTS experiments and evaluate the effectiveness of our model by comparing it with a baseline derived from conventional dialect TTS methods.The results show that our model improves the dialectal naturalness of synthetic speech in CD-TTS.
Kazuki Yamauchi, Yuki Saito 0001, Hiroshi Saruwatari
SLT3
2024 JNV corpus: A corpus of Japanese nonverbal vocalizations with diverse phrases and emotions
abstract
We present JNV (Japanese Nonverbal Vocalizations) corpus, a corpus of Japanese nonverbal vocalizations (NVs) with diverse phrases and emotions. Existing Japanese NV corpora either lack phrase diversity or focus on a small number of emotions, which makes it difficult to analyze the characteristics of Japanese NVs and support downstream tasks like emotion recognition. We first propose a corpus-design method that contains two phases: (1) collecting NVs phrases based on crowd-sourcing; (2) recording NVs by stimulating speakers with emotional scenarios. We then collect 420 audio clips from 4 speakers that cover 6 emotions based on the proposed method. Results of comprehensive objective and subjective experiments demonstrate that (1) the emotions of the collected NVs can be recognized with high accuracy by both human evaluators and statistical models; (2) the collected NVs have a high authenticity comparable to previous corpora of English NVs. Additionally, we analyze the distributions of vowel types in Japanese and conduct feature importance analysis to show discriminative acoustic features between emotion categories in Japanese NVs. We publicate JNV to advance further development in this field.
Detai Xin, Shinnosuke Takamichi, Hiroshi Saruwatari
Speech Commun.3
2024 Sound Field Estimation Based on Physics-Constrained Kernel Interpolation Adapted to Environment
abstract
A sound field estimation method based on kernel interpolation with an adaptive kernel function is proposed. The kernel-interpolation-based sound field estimation methods enable physics-constrained interpolation from pressure measurements of distributed microphones with a linear estimator, which constrains interpolation functions to satisfy the Helmholtz equation. However, a fixed kernel function would not be capable of adapting to the acoustic environment in which the measurement is performed, limiting their applicability. To make the kernel function adaptive, we represent it with a sum of directed and residual trainable kernel functions. The directed kernel is defined by a weight function composed of a superposition of exponential functions to capture highly directional components. The weight function for the residual kernel is represented by neural networks to capture unpredictable spatial patterns of the residual components. Experimental results using simulated and real data indicate that the proposed method outperforms the current kernel-interpolation-based methods and a method based on physics-informed neural networks.
Juliano G. C. Ribeiro, Shoichi Koyama, Ryosuke Horiuchi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Text-Inductive Graphone-Based Language Adaptation for Low-Resource Speech Synthesis
abstract
Neural text-to-speech (TTS) systems have made significant progress in generating natural synthetic speech. However, neural TTS requires large amounts of paired training data, which limits its applicability to a small number of resource-rich languages. Previous work on low-resource TTS has addressed the data hungriness based on transfer learning from a multilingual model to low-resource languages, but it still relies heavily on the availability of paired data for the target languages. In this paper, we propose a text-inductive language adaptation framework for low-resource TTS to address the cost of collecting the paired data for low-resource languages. To inject textual knowledge during transfer learning, our framework employs a two-stage adaptation scheme that utilizes both text-only and supervised data for the target language. In the text-based adaptation stage, we update the language-aware embedding layer with a masked language model objective using text-only data for the target language. In the supervised adaptation stage, the entire TTS model is updated using paired data for the target language. We also propose a graphone-based multilingual training method that jointly uses graphemes and International Phonetic Alphabet symbols (referred to as graphones) for resource-rich languages, while using only graphemes for low-resource languages. This approach facilitates the transfer of pronunciation knowledge from resource-rich to low-resource languages. Through extensive evaluations, we demonstrate that 1) our framework with text-based adaptation outperforms the previous supervised transfer learning approach, 2) the proposed graphone-based training method further improves the performance of both multilingual TTS and low-resource language adaptation. With only 5 minutes of paired data for fine-tuning, our method achieved highly intelligible synthetic speech with the character error rates of around 6 % for a target language.
Takaaki Saeki, Soumi Maiti, Shinji Watanabe 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 COCO-NUT: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-Based Control
abstract
In text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as text-to-image, free-form text instruction should be useful for intuitive and complicated control of voice characteristics. A sufficiently large corpus of high-quality and diverse voice samples with corresponding free-form descriptions can advance such control research. However, neither an open corpus nor a scalable method is currently available. To this end, we develop Coco-Nut, a new corpus including diverse Japanese utterances, along with text transcriptions and free-form voice characteristics descriptions. Our methodology to construct this corpus consists of 1) automatic collection of voice-related audio data from the Internet, 2) quality assurance, and 3) manual annotation using crowdsourcing. Additionally, we benchmark our corpus on the prompt embedding model trained by contrastive speech-text learning.
Aya Watanabe, Shinnosuke Takamichi, Yuki Saito 0001, Wataru Nakata, Detai Xin, Hiroshi Saruwatari
ASRU6
2023 Spatial Active Noise Control Method Based on Sound Field Interpolation from Reference Microphone Signals
abstract
A spatial active noise control (ANC) method based on the interpolation of a sound field from reference microphone signals is proposed. In most current spatial ANC methods, a sufficient number of error microphones are required to reduce noise over the target region because the sound field is estimated from error microphone signals. However, in practical applications, it is preferable that the number of error microphones is as small as possible to keep a space in the target region for ANC users. We propose to interpolate the sound field using reference microphones, which are normally placed outside the target region, instead of the error microphones. We derive a fixed filter for spatial noise reduction on the basis of the kernel ridge regression for sound field interpolation. Furthermore, to compen-sate for estimation errors, we combine the proposed fixed filter with multichannel ANC based on a transition of the control filter using the error microphone signals. Numerical experimental results indicate that regional noise can be sufficiently reduced by the proposed methods even when the number of error microphones is particularly small.
Kazuyuki Arikawa, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2023 jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
abstract
We construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese children’s songs and have six voice parts (lead vocal, soprano, alto, tenor, bass, and vocal percussion). They are divided into seven subsets, each of which features typical characteristics of a music genre such as jazz and enka. The variety in genre and voice part match vocal ensembles recently widespread in social media services such as YouTube, although the main targets of conventional vocal ensemble datasets are choral singing made up of soprano, alto, tenor, and bass. Experimental evaluation demonstrates that our corpus is a challenging resource for vocal ensemble separation. Our corpus is available on our project page.
Tomohiko Nakamura, Shinnosuke Takamichi, Naoko Tanji, Satoru Fukayama, Hiroshi Saruwatari
ICASSP5
2023 Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source Images
abstract
We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental sounds from the desired onomatopoeia texts. Onomatopoeias have another representation: visual-text representations of sounds in comics, advertisements, and virtual reality. A visual onomatopoeia (visual text of onomatopoeia) contains rich information that is not present in the text, such as a long-short duration of the image, so the use of this representation is expected to synthesize diverse sounds. Therefore, we propose visual onoma-to-wave for environmental sound synthesis from visual onomatopoeia. The method can transfer visual concepts of the visual text and sound-source image to the synthesized sound. We also propose a data augmentation method focusing on the repetition of onomatopoeias to enhance the performance of our method. An experimental evaluation shows that the methods can synthesize diverse environmental sounds from visual text and sound-source images.
Hien Ohnaka, Shinnosuke Takamichi, Keisuke Imoto, Yuki Okamoto, Kazuki Fujii, Hiroshi Saruwatari
ICASSP6
2023 Kernel Interpolation of Acoustic Transfer Functions with Adaptive Kernel for Directed and Residual Reverberations
abstract
An interpolation method for region-to-region acoustic transfer functions (ATFs) based on kernel ridge regression with an adaptive kernel is proposed. Most current ATF interpolation methods do not incorporate the acoustic properties for which measurements are performed. Our proposed method is based on a separate adaptation of directional weighting functions to directed and residual reverberations, which are used for adapting kernel functions. Thus, the proposed method can not only impose constraints on fundamental acoustic properties, but can also adapt to the acoustic environment. Numerical experimental results indicated that our proposed method outperforms the current methods in terms of interpolation accuracy, especially at high frequencies.
Juliano G. C. Ribeiro, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2023 MID-Attribute Speaker Generation Using Optimal-Transport-Based Interpolation of Gaussian Mixture Models
abstract
In this paper, we propose a method for intermediating multiple speakers’ attributes and diversifying their voice characteristics in “speaker generation,” an emerging task that aims to synthesize a nonexistent speaker’s naturally sounding voice. The conventional TacoSpawn-based speaker generation method represents the distributions of speaker embeddings by Gaussian mixture models (GMMs) conditioned with speaker attributes. Although this method enables the sampling of various speakers from the speaker-attribute-aware GMMs, it is not yet clear whether the learned distributions can represent speakers with an intermediate attribute (i.e., mid-attribute). To this end, we propose an optimal-transport-based method that interpolates the learned GMMs to generate nonexistent speakers with mid-attribute (e.g., gender-neutral) voices. We empirically validate our method and evaluate the naturalness of synthetic speech and the controllability of two speaker attributes: gender and language fluency. The evaluation results show that our method can control the generated speakers’ attributes by a continuous scalar value without statistically significant degradation of speech naturalness.
Aya Watanabe, Shinnosuke Takamichi, Yuki Saito 0001, Detai Xin, Hiroshi Saruwatari
ICASSP5
2023 Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts
abstract
We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or single-modality context, which does not fully represent the context information. The proposed method uses an acoustic context encoder and a textual context encoder to aggregate context information and feeds it to the TTS model, which enables the model to predict context-dependent prosody. We conducted comprehensive objective and subjective evaluations on a multi-speaker Japanese audiobook dataset. Experimental results demonstrate that the proposed method significantly outperforms two previous works. Additionally, we present insights about the different choices of context - modalities, lateral information and length - for audiobook TTS that have never been discussed in the literature before.
Detai Xin, Sharath Adavanne, Federico Ang, Ashish Kulkarni, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP6
2023 Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech
abstract
Pause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers’ different styles of inserting silent pauses, which can degrade the performance of the model trained on a multi-speaker speech corpus. To this end, we propose more powerful pause insertion frameworks based on a pre-trained language model. Our approach uses bidirectional encoder representations from transformers (BERT) pre-trained on a large-scale text corpus, injecting speaker embeddings to capture various speaker characteristics. We also leverage duration-aware pause insertion for more natural multi-speaker TTS. We develop and evaluate two types of models. The first improves conventional phrasing models on the position prediction of respiratory pauses (RPs), i.e., silent pauses at word transitions without punctuation. It performs speaker-conditioned RP prediction considering contextual information and is used to demonstrate the effect of speaker information on the prediction. The second model is further designed for phoneme-based TTS models and performs duration-aware pause insertion, predicting both RPs and punctuation-indicated pauses (PIPs) that are categorized by duration. The evaluation results show that our models improve the precision and recall of pause insertion and the rhythm of synthetic speech.
Tomoki Koriyama, Yuki Saito 0001, Takaaki Saeki, Detai Xin, Hiroshi Saruwatari
ICASSP6
2023 Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining
abstract
While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method for zero-shot multilingual TTS using text-only data for the target language. The use of text-only data allows the development of TTS systems for low-resource languages for which only textual resources are available, making TTS accessible to thousands of languages. Inspired by the strong cross-lingual transferability of multilingual language models, our framework first performs masked language model pretraining with multilingual text-only data. Then we train this model with a paired data in a supervised manner, while freezing a language-aware embedding layer. This allows inference even for languages not included in the paired data but present in the text-only data. Evaluation results demonstrate highly intelligible zero-shot TTS with a character error rate of less than 12% for an unseen language.
Takaaki Saeki, Soumi Maiti, Shinji Watanabe 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IJCAI6
2023 How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics
Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura, Kentaro Seki, Detai Xin, Hiroshi Saruwatari
INTERSPEECH6
2023 CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
Yuki Saito 0001, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH5
2023 ChatGPT-EDSS: Empathetic Dialogue Speech Synthesis Trained from ChatGPT-derived Context Word Embeddings
Yuki Saito 0001, Shinnosuke Takamichi, Eiji Iimori, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH5
2023 HumanDiffusion: diffusion model using perceptual gradients
Yota Ueda, Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Hiroshi Saruwatari
INTERSPEECH5
2023 Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus
Detai Xin, Shinnosuke Takamichi, Ai Morimatsu, Hiroshi Saruwatari
INTERSPEECH4
2023 Amplitude Matching for Multizone Sound Field Control
abstract
A multizone sound field control method, called amplitude matching, is proposed. The objective of amplitude matching is to synthesize a desired amplitude (or magnitude) distribution over a target region with multiple loudspeakers, whereas the phase distribution is arbitrary. Most of the current multizone sound field control methods are intended to synthesize a specific sound field including phase or to control acoustic potential energy inside the target region. In amplitude matching, a specific desired amplitude distribution can be set, ignoring sound propagation directions. Although the optimization problem of amplitude matching does not have a closed-form solution, our proposed algorithm based on the alternating direction method of multipliers (ADMM) allows us to accurately and efficiently synthesize the desired amplitude distribution. We also introduce the differential-norm penalty for a time-domain filter design with a small filter length. The experimental results indicated that the proposed method outperforms current multizone sound field control methods in terms of accuracy of the synthesized amplitude distribution.
Takumi Abe, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 PoP-IDLMA: Product-of-Prior Independent Deeply Learned Matrix Analysis for Multichannel Music Source Separation
abstract
Independent deeply learned matrix analysis (IDLMA) is a state-of-the-art determined audio source separation method based on pretrained deep neural networks (DNNs). Owing to the excellent expression power of DNNs, IDLMA can handle a wider range of sources than conventional source models such as nonegative matrix factorization (NMF). However, owing to its supervised nature, the separation performance of IDLMA often degrades in the presence of timbral mismatches between the training data and the to-be-separated data. In this paper, we propose two source models that encompass the NMF- and DNN-based source models by constructing a prior distribution of the source power spectrogram (product of priors: PoP) on the basis of the product-of-expert concept. Since the NMF-based source model works well for a fully blind situation, the proposed models can handle the timbral mismatch without losing the expression power of DNNs. By introducing the PoP-based source models into IDLMA, we propose IDLMA extensions (PoP-IDLMAs) and derive their efficient parameter estimation algorithms on the basis of the majorization–minimization algorithm. Experimental results demonstrated the effectiveness of the proposed PoP-IDLMAs and that the proposed models greatly improve the source power estimation in frequency bands above 500 Hz.
Takuya Hasumi, Tomohiko Nakamura, Norihiro Takamune, Hiroshi Saruwatari, Daichi Kitamura, Yu Takahashi, Kazunobu Kondo
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Spatial Active Noise Control Based on Individual Kernel Interpolation of Primary and Secondary Sound Fields
abstract
A spatial active noise control (ANC) method based on the individual kernel interpolation of primary and secondary sound fields is proposed. Spatial ANC is aimed at cancelling unwanted primary noise within a continuous region by using multiple secondary sources and microphones. A method based on the kernel interpolation of a sound field makes it possible to attenuate noise over the target region with flexible array geometry. Furthermore, by using the kernel function with directional weighting, prior information on primary noise source directions can be taken into consideration. However, whereas the sound field to be interpolated is a superposition of primary and secondary sound fields, the directional weight for the primary noise source was applied to the total sound field in previous work; therefore, the performance improvement was limited. We propose a method of individually interpolating the primary and secondary sound fields and formulate a normalized least-mean-square algorithm based on this interpolation method. Experimental results indicate that the proposed method outperforms the method based on total kernel interpolation.
Kazuyuki Arikawa, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2022 Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic Sounds
abstract
A differentiable digital signal processing (DDSP) autoencoder is a musical sound synthesizer that combines a deep neural network (DNN) and spectral modeling synthesis. It allows us to flexibly edit sounds by changing the fundamental frequency, timbre feature, and loudness (synthesis parameters) extracted from an input sound. However, it is designed for a monophonic harmonic sound and cannot handle mixtures of harmonic sounds. In this paper, we propose a model (DDSP mixture model) that represents a mixture as the sum of the outputs of multiple pretrained DDSP autoencoders. By fitting the output of the proposed model to the observed mixture, we can directly estimate the synthesis parameters of each source. Through synthesis parameter extraction experiments, we show that the proposed method has high and stable performance compared with a straightforward method that applies the DDSP autoencoder to the signals separated by an audio source separation method.
Masaya Kawamura, Tomohiko Nakamura, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
ICASSP4
2022 Region-to-Region Kernel Interpolation of Acoustic Transfer Function with Directional Weighting
abstract
A method of interpolating the acoustic transfer function (ATF) between regions that takes into account both the physical properties of the ATF and the directionality of region configurations is proposed. Most spatial ATF interpolation methods are limited to estimation in the region of receivers. A kernel method for region-to-region ATF interpolation makes it possible to estimate the ATFs for both source and receiver regions from a discrete set of ATF measurements. We newly formulate the reproducing kernel Hilbert space and associated kernel function incorporating directional weight to enhance the interpolation accuracy. We also investigate hyperparameter optimization methods for this kernel function. Numerical experiments indicate that the proposed method outperforms the method without the use of directional weighting.
Juliano G. C. Ribeiro, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2022 Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis
Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito 0001, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
INTERSPEECH7
2022 Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History
abstract
We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history.Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and empathetic DSS is a technology to implement this act in spoken dialogue systems.Our model is conditioned by the history of linguistic and prosody features for predicting appropriate dialogue context.As such, it can be regarded as an extension of the conventional linguistic-feature-based dialogue history modeling.To train the empathetic DSS model effectively, we investigate 1) a self-supervised learning model pretrained with large speech corpora, 2) a style-guided training using a prosody embedding of the current utterance to be predicted by the dialogue context embedding, 3) a cross-modal attention to combine text and speech modalities, and 4) a sentence-wise embedding to achieve fine-grained prosody modeling rather than utterancewise modeling.The evaluation results demonstrate that 1) simply considering prosodic contexts of the dialogue history does not improve the quality of speech in empathetic DSS and 2) introducing style-guided training and sentence-wise embedding modeling achieves higher speech quality than that by the conventional method.
Yuto Nishimura, Yuki Saito 0001, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH5
2022 SelfRemaster: Self-Supervised Speech Restoration with Analysis-by-Synthesis Approach Using Channel Modeling
abstract
We present a self-supervised speech restoration method without paired speech corpora.Because the previous general speech restoration method uses artificial paired data created by applying various distortions to high-quality speech corpora, it cannot sufficiently represent acoustic distortions of real data, limiting the applicability.Our model consists of analysis, synthesis, and channel modules that simulate the recording process of degraded speech and is trained with real degraded speech data in a self-supervised manner.The analysis module extracts distortionless speech features and distortion features from degraded speech, while the synthesis module synthesizes the restored speech waveform, and the channel module adds distortions to the speech waveform.Our model also enables audio effect transfer, in which only acoustic distortions are extracted from degraded speech and added to arbitrary high-quality audio.Experimental evaluations with both simulated and real data show that our method achieves significantly higher-quality speech restoration than the previous supervised method, suggesting its applicability to real degraded speech materials.
Takaaki Saeki, Shinnosuke Takamichi, Tomohiko Nakamura, Naoko Tanji, Hiroshi Saruwatari
INTERSPEECH5
2022 UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
abstract
We present the UTokyo-SaruLab mean opinion score (MOS) prediction system submitted to VoiceMOS Challenge 2022.The challenge is to predict the MOS values of speech samples collected from previous Blizzard Challenges and Voice Conversion Challenges for two tracks: a main track for in-domain prediction and an out-of-domain (OOD) track for which there is less labeled data from different listening tests.Our system is based on ensemble learning of strong and weak learners.Strong learners incorporate several improvements to the previous finetuning models of self-supervised learning (SSL) models, while weak learners use basic machine-learning methods to predict scores from SSL features.In the Challenge, our system had the highest score on several metrics for both the main and OOD tracks.In addition, we conducted ablation studies to investigate the effectiveness of our proposed methods.
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH6
2022 STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent
abstract
We present STUDIES, a new speech corpus for developing a voice agent that can speak in a friendly manner.Humans naturally control their speech prosody to empathize with each other.By incorporating this "empathetic dialogue" behavior into a spoken dialogue system, we can develop a voice agent that can respond to a user more naturally.We designed the STUDIES corpus to include a speaker who speaks with empathy for the interlocutor's emotion explicitly.We describe our methodology to construct an empathetic dialogue speech corpus and report the analysis results of the STUDIES corpus.We conducted a text-to-speech experiment to initially investigate how we can develop more natural voice agent that can tune its speaking style corresponding to the interlocutor's emotion.The results show that the use of interlocutor's emotion label and conversational context embedding can produce speech with the same degree of naturalness as that synthesized by using the agent's emotion label.
Yuki Saito 0001, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari
INTERSPEECH5
2022 J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis
abstract
In this paper, we construct a Japanese audiobook speech corpus called "J-MAC" for speech synthesis research.With the success of reading-style speech synthesis, the research target is shifting to tasks that use complicated contexts.Audiobook speech synthesis is a good example that requires cross-sentence, expressiveness, etc.Unlike reading-style speech, speaker-specific expressiveness in audiobook speech also becomes the context.To enhance this research, we propose a method of constructing a corpus from audiobooks read by professional speakers.From many audiobooks and their texts, our method can automatically extract and refine the data without any language dependency.Specifically, we use vocal-instrumental separation to extract clean data, connectionist temporal classification to roughly align text and audio, and voice activity detection to refine the alignment.J-MAC is open-sourced in our project page.We also conduct audiobook speech synthesis evaluations, and the results give insights into audiobook speech synthesis.
Shinnosuke Takamichi, Wataru Nakata, Naoko Tanji, Hiroshi Saruwatari
INTERSPEECH4
2022 Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTS
abstract
This paper proposes a human-in-the-loop speaker-adaptation method for multi-speaker text-to-speech.With a conventional speaker-adaptation method, a target speaker's embedding vector is extracted from his/her reference speech using a speaker encoder trained on a speaker-discriminative task.However, this method cannot obtain an embedding vector for the target speaker when the reference speech is unavailable.Our method is based on a human-in-the-loop optimization framework, which incorporates a user to explore the speakerembedding space to find the target speaker's embedding.The proposed method uses a sequential line search algorithm that repeatedly asks a user to select a point on a line segment in the embedding space.To efficiently choose the best speech sample from multiple stimuli, we also developed a system in which a user can switch between multiple speakers' voices for each phoneme while looping an utterance.Experimental results indicate that the proposed method can achieve comparable performance to the conventional one in objective and subjective evaluations even if reference speech is not used as the input of a speaker encoder directly.
Kenta Udagawa, Yuki Saito 0001, Hiroshi Saruwatari
INTERSPEECH3
2022 Personalized Filled-pause Generation with Group-wise Prediction Models
abstract
In this paper, we propose a method to generate personalized filled pauses (FPs) with group-wise prediction models. Compared with fluent text generation, disfluent text generation has not been widely explored. To generate more human-like texts, we addressed disfluent text generation. The usage of disfluency, such as FPs, rephrases, and word fragments, differs from speaker to speaker, and thus, the generation of personalized FPs is required. However, it is difficult to predict them because of the sparsity of position and the frequency difference between more and less frequently used FPs. Moreover, it is sometimes difficult to adapt FP prediction models to each speaker because of the large variation of the tendency within each speaker. To address these issues, we propose a method to build group-dependent prediction models by grouping speakers on the basis of their tendency to use FPs. This method does not require a large amount of data and time to train each speaker model. We further introduce a loss function and a word embedding model suitable for FP prediction. Our experimental results demonstrate that group-dependent models can predict FPs with higher scores than a non-personalized one and the introduced loss function and word embedding model improve the prediction performance.
Yuta Matsunaga, Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
LREC4
2022 VTTS: Visual-Text To Speech
abstract
This paper proposes a visual-text to speech (vTTS) method, a method for synthesizing speech directly from visual text (i.e., text as an image). vTTS can use visual features in visual text that should be important for speech synthesis such as emphasis and radicals (components in Chinese characters), but they are not available in conventional TTS using discrete symbols as the input. The proposed vTTS method extracts visual features with a convolutional neural network and then generates acoustic features with a non-autoregressive model in an end-to-end manner. Experimental results show that 1) the vTTS method is capable of generating speech with naturalness comparable to or better than a conventional TTS, 2) it can transfer emphasis and emotion attributes in visual text to speech without additional labels and architectures, and 3) it can synthesize more natural and intelligible speech from unseen and rare characters than conventional TTS.
Yoshifumi Nakano, Takaaki Saeki, Shinnosuke Takamichi, Katsuhito Sudoh, Hiroshi Saruwatari
SLT5
2022 Region-to-Region Kernel Interpolation of Acoustic Transfer Functions Constrained by Physical Properties
abstract
A method to interpolate the acoustic transfer function (ATF) between regions using kernel ridge regression (KRR) is proposed. Conventionally, the ATF interpolation problem is strongly restricted and situational, depending on knowledge of environmental conditions while not accounting for source position variation. We derive our interpolation function as the solution of an optimization problem defined on a function space where every element holds the acoustic properties of the ATF. By making the space a reproducing kernel Hilbert space (RKHS), we can guarantee that our problem has a known and unique optimizer. The generality of the formulation of this method enables region-to-region estimations, with variable source and receiver within the assigned bounds. The definition of a RKHS also allows for the use of kernel principal component analysis, thereby efficiently providing greater noise robustness to our interpolation function. Our proposed method is compared with a previously established region-to-region interpolation method in numerical simulations where the advantages of the KRR approach are confirmed, showing lower error and greater stability for higher frequencies.
Juliano G. C. Ribeiro, Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Sampling-Frequency-Independent Convolutional Layer and its Application to Audio Source Separation
abstract
Audio source separation is often used for the preprocessing of various tasks, and one of its ultimate goals is to construct a single versatile preprocessor that can handle every variety of audio signal. One of the most important varieties of the discrete-time audio signal is sampling frequency. Since it is usually task-specific, the versatile preprocessor must handle all the sampling frequencies required by the possible downstream tasks. However, conventional models based on deep neural networks (DNNs) are not designed for handling a variety of sampling frequencies. Thus, for unseen sampling frequencies, they may not work appropriately. In this paper, we propose sampling-frequency-independent (SFI) convolutional layers capable of handling various sampling frequencies. The core idea of the proposed layers comes from our finding that a convolutional layer can be viewed as a collection of digital filters and inherently depends on sampling frequency. To overcome this dependency, we propose an SFI structure that features analog filters and generates weights of a convolutional layer from the analog filters. By utilizing time- and frequency-domain analog-to-digital filter conversion techniques, we can adapt the convolutional layer for various sampling frequencies. As an example application, we construct an SFI version of a conventional source separation network. Through music source separation experiments, we show that the proposed layers enable separation networks to consistently work well for unseen sampling frequencies in objective and perceptual separation qualities. We also demonstrate that the proposed method outperforms a conventional method based on signal resampling when the sampling frequencies of input signals are significantly lower than the trained sampling frequency.
Koichi Saito, Tomohiko Nakamura, Kohei Yatabe, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network
abstract
Incremental text-to-speech (TTS) synthesis generates utterances in small linguistic units for the sake of real-time and low-latency applications. We previously proposed an incremental TTS method that leverages a large pre-trained language model to take unobserved future context into account without waiting for the subsequent segment. Although this method achieves comparable speech quality to that of a method that waits for the future context, it entails a huge amount of processing for sampling from the language model at each time step. In this paper, we propose an incremental TTS method that directly predicts the unobserved future context with a lightweight model, instead of sampling words from the large-scale language model. We perform knowledge distillation from a GPT2-based context prediction network into a simple recurrent model by minimizing a teacher-student loss defined between the context embedding vectors of those models. Experimental results show that the proposed method requires about ten times less inference time to achieve comparable synthetic speech quality to that of our previous method, and it can perform incremental synthesis much faster than the average speaking speed of human English speakers, demonstrating the availability of our method to real-time applications.
Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
ASRU3
2021 Deficient Basis Estimation of Noise Spatial Covariance Matrix for Rank-Constrained Spatial Covariance Matrix Estimation Method in Blind Speech Extraction
abstract
Rank-constrained spatial covariance matrix estimation (RCSCME) is a state-of-the-art blind speech extraction method applied to cases where one directional target speech and diffuse noise are mixed. In this paper, we proposed a new algorithmic extension of RCSCME. RCSCME complements a deficient one rank of the diffuse noise spatial covariance matrix, which cannot be estimated via preprocessing such as independent low-rank matrix analysis, and estimates the source model parameters simultaneously. In the conventional RC- SCME, a direction of the deficient basis is fixed in advance and only the scale is estimated; however, the candidate of this deficient basis is not unique in general. In the proposed RCSCM model, the deficient basis itself can be accurately estimated as a vector variable by solving a vector optimization problem. Also, we derive new update rules based on the EM algorithm. We confirm that the proposed method outperforms conventional methods under several noise conditions.
Yuto Kondo, Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari
ICASSP5
2021 Amplitude Matching: Majorization-Minimization Algorithm for Sound Field Control Only with Amplitude Constraint
abstract
A sound field control method for synthesizing a desired amplitude distribution inside a target region, amplitude matching, is proposed. In the conventional pressure matching, a desired sound field is set as a pressure distribution including amplitude and phase. In personal audio applications, it is sometimes not necessary to synthesize a specific phase distribution, but a certain acoustic power level should be controlled inside the target region. Since the optimization problem to achieve amplitude matching becomes nonlinear, there is no closed-form solution and numerical optimization algorithms are generally applied. We derive an efficient algorithm for the amplitude matching based on the majorization–minimization algorithm. Numerical experiments indicated that high control accuracy over the target region can be achieved with low computational cost by using the proposed algorithm.
Shoichi Koyama, Takashi Amakasu, Natsuki Ueno, Hiroshi Saruwatari
ICASSP4
2021 Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme Perception
abstract
We propose a conditional generative adversarial network (GAN) incorporating humans’ perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a real-data distribution accurately but can never represent a human-acceptable distribution, which are ranges of data in which humans accept the naturalness regardless of whether the data are real or not. A Human-GAN was proposed to model the human-acceptable distribution. A DNN-based generator is trained using a human-based discriminator, i.e., humans’ perceptual evaluations, instead of the GAN’s DNN-based discriminator. However, the HumanGAN cannot represent conditional distributions. This paper proposes the HumanACGAN, a theoretical extension of the HumanGAN, to deal with conditional human-acceptable distributions. Our HumanACGAN trains a DNN-based conditional generator by regarding humans as not only a discriminator but also an auxiliary classifier. The generator is trained by deceiving the human-based discriminator that scores the unconditioned naturalness and the human-based classifier that scores the class-conditioned perceptual acceptability. The training can be executed using the backpropagation algorithm involving humans’ perceptual evaluations. Our experimental results in phoneme perception demonstrate that our HumanACGAN can successfully train this conditional generator.
Yota Ueda, Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari
ICASSP6
2021 Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTS
abstract
We propose a method for obtaining disentangled speaker and language representations via mutual information minimization and domain adaptation for cross-lingual text-to-speech (TTS) synthesis. The proposed method extracts speaker and language embeddings from acoustic features by a speaker encoder and a language encoder. Then the proposed method applies domain adaptation on the two embeddings to obtain language-invariant speaker embedding and speaker-invariant language embedding. To get more disentangled representations, the proposed method further uses mutual information minimization between the two embeddings to remove entangled information within each embedding. Disentangled representations of speaker and language are critical for cross-lingual TTS synthesis since entangled representations make it difficult to maintain speaker identity information when changing the language representation and consequently causes performance degradation. We evaluate the proposed method using English and Japanese multi-speaker datasets with a total of 207 speakers. Experimental result demonstrates that the proposed method significantly improves the naturalness and speaker similarity of both intra-lingual and cross-lingual TTS synthesis. Furthermore, we show that the proposed method has a good capability of maintaining the speaker identity between languages.
Detai Xin, Tatsuya Komatsu, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP4
2021 Harmonic WaveGAN: GAN-Based Speech Waveform Generation Model with Harmonic Structure Discriminator
Kazuki Mizuta, Tomoki Koriyama, Hiroshi Saruwatari
Interspeech3
2021 Sequence-to-Sequence Learning for Deep Gaussian Process Based Speech Synthesis Using Self-Attention GP Layer
Taiki Nakamura, Tomoki Koriyama, Hiroshi Saruwatari
Interspeech3
2021 Cross-Lingual Speaker Adaptation Using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
Interspeech5
2021 Joint-diagonalizability-constrained multichannel nonnegative matrix factorization based on time-variant multivariate complex sub-Gaussian distribution
abstract
Multichannel nonnegative matrix factorization (MNMF) is a common blind source separation technique that employs full-rank spatial covariance matrices (SCMs). The full-rank SCMs can simulate reverberant mixing systems where the sources are spatially spread. In conventional MNMF, spectrograms of observed signals are modeled by some types of distribution, e.g., the Gaussian distribution and Student’s t distribution. However, MNMF based on the sub-Gaussian distribution has not been proposed because its cost function is difficult to minimize. In this paper, we address the statistical model extension of MNMF to the sub-Gaussian distribution to improve the source separation accuracy. In the proposed method, the generalized Gaussian distribution is utilized as the sub-Gaussian model. Moreover, to design an auxiliary function for the proposed cost function, we introduce the joint-diagonalizability constraint to SCMs similarly to FastMNMF. Two types of update rule for the proposed MNMF are derived on the basis of the majorization-minimization (MM) and majorization-equalization (ME) algorithms. Since the optimization speed of each parameter affects the source separation performance, we experimentally analyze the best combination of MM- and ME-algorithm-based update rules in the proposed method. Experiments of blind source separation reveal that the proposed MNMF based on the sub-Gaussian model can outperform conventional methods.
Keigo Kamo, Yoshiki Mitsui, Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
Signal Process.6
2021 Independent deeply learned matrix analysis with automatic selection of stable microphone-wise update and fast sourcewise update of demixing matrix
abstract
Independent deeply learned matrix analysis (IDLMA) is a fast and high-performance method for multichannel audio source separation. IDLMA utilizes the deep neural network inference of source models and the blind estimation of demixing filters based on source independence. In conventional IDLMA, iterative projection (IP) is exploited to estimate the demixing filters. Although IP is a fast algorithm, it sometimes fails to estimate an appropriate solution. This is because IP updates the demixing filters in a sourcewise manner, where only one source model is used for each update, and the update sometimes becomes unstable owing to the specific low-quality source models. In this paper, we first derive a new numerically stable microphone-wise update algorithm that exploits all source model information simultaneously. The microphone-wise update problem cannot be solved by IP; instead, a new type of vectorwise coordinate descent algorithm is introduced. Next, comparison analysis of the proposed microphone-wise update and IP reveals the tradeoff w.r.t. convergence speed and numerical stability. To resolve this tradeoff problem, we propose the automatic selection of update rules on the basis of the likelihood function of observed signals. Finally, experimental results show the efficacy of the proposed IDLMA with the automatic selection of update rules.
Naoki Makishima, Yoshiki Mitsui, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
Signal Process.5
2021 Deep Gaussian process based multi-speaker speech synthesis with latent speaker representation
abstract
This paper proposes deep Gaussian process (DGP)-based frameworks for multi-speaker speech synthesis and speaker representation learning. A DGP has a deep architecture of Bayesian kernel regression, and it has been reported that DGP-based single speaker speech synthesis outperforms deep neural network (DNN)-based ones in the framework of statistical parametric speech synthesis. By extending this method to multiple speakers, it is expected that higher speech quality can be achieved with a smaller number of training utterances from each speaker. To apply DGPs to multi-speaker speech synthesis, we propose two methods: one using DGP with one-hot speaker codes, and the other using a deep Gaussian process latent variable model (DGPLVM). The DGP with one-hot speaker codes uses additional GP layers to transform speaker codes into latent speaker representations. The DGPLVM directly models the distribution of latent speaker representations and learns it jointly with acoustic model parameters. In this method, acoustic speaker similarity is expressed in terms of the similarity of the speaker representations, and thus, the voices of similar speakers are efficiently modeled. We experimentally evaluated the performance of the proposed methods in comparison with those of conventional DNN and variational autoencoder (VAE)-based frameworks, in terms of acoustic feature distortion and subjective speech quality. The experimental results demonstrate that (1) the proposed DGP-based and DGPLVM-based methods improve subjective speech quality compared with a feed-forward DNN-based method, and (2) even when the amount of training data for target speakers is limited, the DGPLVM-based method outperforms other methods, including the VAE-based one. Additionally, (3) by using a speaker representation randomly sampled from the learned speaker space, the DGPLVM-based method can generate voices of non-existent speakers.
Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari
Speech Commun.3
2021 Incremental Text-to-Speech Synthesis Using Pseudo Lookahead With Large Pretrained Language Model
abstract
This letter presents an incremental text-to-speech (TTS) method that performs synthesis in small linguistic units while maintaining the naturalness of output speech. Incremental TTS is generally subject to a trade-off between latency and synthetic speech quality. It is challenging to produce high-quality speech with a low-latency setup that does not make much use of an unobserved future sentence (hereafter, “lookahead”). To resolve this issue, we propose an incremental TTS method that uses a pseudo lookahead generated with a language model to take the future contextual information into account without increasing latency. Our method can be regarded as imitating a human's incremental reading and uses pretrained GPT2, which accounts for the large-scale linguistic knowledge, for the lookahead generation. Evaluation results show that our method 1) achieves higher speech quality than the method taking only observed information into account and 2) achieves a speech quality equivalent to waiting for the future context observation.
Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE Signal Process. Lett.3
2021 Spatial Active Noise Control Based on Kernel Interpolation of Sound Field
abstract
An active noise control (ANC) method to reduce noise over a region in space based on kernel interpolation of sound field is proposed. Current methods of spatial ANC are largely based on spherical or circular harmonic expansion of the sound field, where the geometry of the error microphone array is restricted to a simple one such as a sphere or circle. We instead apply the kernel interpolation method, which allows for the estimation of a sound field in a continuous region with flexible array configurations. The interpolation scheme is used to derive adaptive filtering algorithms for minimizing the acoustic potential energy inside a target region. A practical time-domain algorithm is also developed together with its computationally efficient block-based equivalent. We conduct experiments to investigate the achievable level of noise reduction in a two-dimensional free space, as well as adaptive broadband noise control in a three-dimensional reverberant space. The experimental results indicated that the proposed method outperforms the multipoint-pressure-control-based method in terms of regional noise reduction.
Shoichi Koyama, Jesper Brunnström, Hayato Ito, Natsuki Ueno, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Multichannel Blind Source Separation Based on Evanescent-Region-Aware Non-Negative Tensor Factorization in Spherical Harmonic Domain
abstract
There is growing interest in new audio formats in the context of virtual reality (VR), and higher-order ambisonics (HOA) is preferred for VR systems to transmit recorded scenes owing to its transmission efficiency and its flexibility to work with different loudspeaker setups. However, the conversion between another well-known format, i.e., object format, and the HOA format is not fully addressed in the literature. To address this issue, blind source separation in a spherical harmonic (SH) domain can be considered as the best way to extract objects in terms of efficiency, i.e., decoding HOA signals for separation can be omitted. A few authors attempted to extract objects from encoded HOA signals directly by using multichannel non-negative matrix factorization (MNMF), but these approaches either assume only far-field sources or do not take array characteristics into account, which make these methods difficult to use for VR in practical situations where singers or speakers often perform close to microphones. Furthermore, MNMF generally requires a huge computational cost, although dimensional reduction to the SH domain is performed. In this work, we also model near-field sources by estimating the model parameters of non-negative tensor factorization (NTF) in the SH domain assuming that microphone signals can be obtained with a rigid spherical array. We propose a masking scheme to exclude noisy evanescent regions in the SH domain from the NTF cost function. Evaluations show that our method outperforms existing methods devised for the HOA format and that our masking approach is effective in improving the separation quality.
Yuki Mitsufuji, Norihiro Takamune, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Time-Domain Audio Source Separation With Neural Networks Based on Multiresolution Analysis
abstract
We propose a time-domain audio source separation method based on multiresolution analysis, which we call multiresolution deep layered analysis (MRDLA). The MRDLA model is based on one of the state-of-the-art time-domain deep neural networks (DNNs), Wave-U-Net, which successively down-samples features and up-samples them to have the original time resolution. From the signal processing viewpoint, we found that the down-sampling (DS) layers of Wave-U-Net cause aliasing and may discard information useful for source separation because they are implemented with decimation. These two problems are due to the decimation; thus, to achieve a more reliable source separation method, we should design DS layers capable of simultaneously overcoming these problems. With this motivation, focusing on the fact that the successive DS architecture of Wave-U-Net resembles that of multiresolution analysis, we develop DS layers based on discrete wavelet transforms (DWTs), which we call the DWT layers, because the DWTs have anti-aliasing filters and the perfect reconstruction property. We further extend the DWT layers such that their wavelet basis functions can be trained together with the other DNN components while maintaining the perfect reconstruction property. Since a straightforward trainable extension of the DWT layers does not guarantee the existence of anti-aliasing filters, we derive constraints for this guarantee in addition to the perfect reconstruction property. Through music source separation experiments including subjective evaluations, we show the efficacy of the proposed methods and the importance of simultaneously considering both the anti-aliasing filters and the perfect reconstruction property.
Tomohiko Nakamura, Shihori Kozuka, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Perceptual-Similarity-Aware Deep Speaker Representation Learning for Multi-Speaker Generative Modeling
abstract
We propose novel deep speaker representation learning that considers perceptual similarity among speakers for multi-speaker generative modeling. Following its success in accurate discriminative modeling of speaker individuality, knowledge of deep speaker representation learning (i.e., speaker representation learning using deep neural networks) has been introduced to multi-speaker generative modeling. However, the conventional discriminative algorithm does not necessarily learn speaker embeddings suitable for such generative modeling, which may result in lower quality and less controllability of synthetic speech. We propose three representation learning algorithms that utilize a perceptual speaker similarity matrix obtained by large-scale perceptual scoring of speaker-pair similarity. The algorithms train a speaker encoder to learn speaker embeddings with three different representations of the matrix: a set of vectors, the Gram matrix, and a graph. Furthermore, we propose an active learning algorithm that iterates the perceptual scoring and speaker encoder training. To obtain accurate embeddings while reducing costs of scoring and training, the algorithm selects unscored speaker-pairs to be scored next on the basis of the sequentially-trained speaker encoder's similarity prediction results. Experimental evaluation results show that 1) the proposed representation learning algorithms learn speaker embeddings strongly correlated with perceptual speaker-pair similarity, 2) the embeddings improve synthetic speech quality in speech autoencoding tasks better than conventional d-vectors learned by discriminative modeling, 3) the proposed active learning algorithm achieves higher synthetic speech quality while reducing costs of scoring and training, and 4) among the proposed similarity {vector, matrix, graph} embedding algorithms, the first achieves the best speaker similarity for synthetic speech and the third gives the most improvement in the synthetic speech naturalness.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Mutual-Information-Based Sensor Placement for Spatial Sound Field Recording
abstract
A sensor (microphone) placement method based on mutual information for spatial sound field recording is proposed. The sound field recording methods using distributed sensors enable the estimation of the sound field inside a target region of arbitrary shape; however, it is a difficult task to find the best placement of sensors. We focus on the mutual-information-based sensor placement method in which spatial phenomena are modeled as a Gaussian process (GP). We propose the use of the sound-field-interpolation kernel for the covariance of measurements in a GP model to obtain the sensor placement suitable for sound field recording. We also extend the method to treat broadband signals and derive an efficient algorithm based on block matrix inversion. Numerical simulation results indicated that the proposed method achieves accurate sound field estimation compared with a method using the generally used Gaussian kernel.
Kentaro Ariga, Tomoya Nishida, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP5
2020 Humangan: Generative Adversarial Network With Human-Based Discriminator And Its Evaluation In Speech Perception Modeling
abstract
We propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribution by fooling the discriminator that distinguishes real and generated data. Therefore, the basic GAN cannot represent the outside of a real-data distribution. In the case of speech perception, humans can recognize not only human voices but also processed (i.e., a non-existent human) voices as human voice. Such a human-acceptable distribution is typically wider than a real-data one and cannot be modeled by the basic GAN. To model the human-acceptable distribution, we formulate a backpropagation-based generator training algorithm by regarding human perception as a black-boxed discriminator. The training efficiently iterates generator training by using a computer and discrimination by human. We evaluate our HumanGAN in speech naturalness modeling and demonstrate that it can represent a human-acceptable distribution that is wider than a real-data distribution.
Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari
ICASSP5
2020 Spatial Active Noise Control Based on Kernel Interpolation with Directional Weighting
abstract
A spatial active noise control (ANC) method taking prior information on the approximate direction of primary noise sources into consideration is proposed. ANC aims to cancel incoming primary noise using secondary loudspeakers. Conventional multipoint ANC does not guarantee the reduction of noise between multiple discrete control points; therefore, several attempts have been made to reduce the noise over an entire target region, i.e., by spatial ANC. We have recently proposed a spatial ANC method based on kernel ridge regression for sound field interpolation using distributed microphones and loudspeakers, where the cost function is formulated on the basis of the regional power obtained by kernel interpolation. In this study, we incorporate prior knowledge on the noise source direction into spatial ANC based on the kernel interpolation with directional weighting. Numerical simulation results indicate that the proposed method can achieve larger regional noise reduction than the methods without the information on noise source direction.
Hayato Ito, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP4
2020 Regularized Fast Multichannel Nonnegative Matrix Factorization with ILRMA-Based Prior Distribution of Joint-Diagonalization Process
abstract
In this paper, we address a convolutive blind source separation (BSS) problem and propose a new extended framework of FastMNMF by introducing prior information for joint diagonalization of the spatial covariance matrix model. Recently, FastMNMF has been proposed as a fast version of multichannel nonnegative matrix factorization under the assumption that the spatial covariance matrices of multiple sources can be jointly diagonalized. However, its source-separation performance was not improved and the physical meaning of the joint-diagonalization process was unclear. To resolve these problems, we first reveal a close relationship between the joint-diagonalization process and the demixing system used in independent low-rank matrix analysis (ILRMA). Next, motivated by this fact, we propose a new regularized FastMNMF supported by ILRMA and derive convergence-guaranteed parameter update rules. From BSS experiments, we show that the proposed method outperforms the conventional FastMNMF in source-separation accuracy with almost the same computation time.
Keigo Kamo, Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
ICASSP5
2020 Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T Distribution
abstract
In this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a state-of-the-art BSS method that enables us to take interfrequency correlations into account, but the generative model is limited within the multivariate Gaussian distribution and its parameter optimization algorithm does not guarantee stable convergence. To resolve these problems, first, we propose to extend the generative model to a parametric multivariate Student’s t distribution that can deal with various types of signal. Secondly, we derive a new parameter optimization algorithm that guarantees the monotonic nonincrease in the cost function, providing stable convergence. Experimental results reveal that the cost function in the conventional IPSDTA does not display monotonically nonincreasing properties. On the other hand, the proposed method guarantees the monotonic nonincrease in the cost function and outperforms the conventional ILRMA and IPSDTA in the source-separation performance.
Tatsuki Kondo, Kanta Fukushige, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Rintaro Ikeshita, Tomohiro Nakatani
ICASSP5
2020 Utterance-Level Sequential Modeling for Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit
abstract
This paper presents a deep Gaussian process (DGP) model with a recurrent architecture for speech sequence modeling. DGP is a Bayesian deep model that can be trained effectively with the consideration of model complexity and is a kernel regression model that can have high expressibility. In the previous studies, it was shown that the DGP-based speech synthesis outperformed neural network-based one, in which both models used a feed-forward architecture. To improve the naturalness of synthetic speech, in this paper, we show that DGP can be applied to utterance-level modeling using recurrent architecture models. We adopt a simple recurrent unit (SRU) for the proposed model to achieve a recurrent architecture, in which we can execute fast speech parameter generation by using the high parallelization nature of SRU. The objective and subjective evaluation results show that the proposed SRU-DGP-based speech synthesis outperforms not only feed-forward DGP but also automatically tuned SRU- and long short-term memory (LSTM)-based neural networks.
Tomoki Koriyama, Hiroshi Saruwatari
ICASSP2
2020 Time-Domain Audio Source Separation Based on Wave-U-Net Combined with Discrete Wavelet Transform
abstract
We propose a time-domain audio source separation method using down-sampling (DS) and up-sampling (US) layers based on a discrete wavelet transform (DWT). The proposed method is based on one of the state-of-the-art deep neural networks, Wave-U-Net, which successively down-samples and up-samples feature maps. We find that this architecture resembles that of multiresolution analysis, and reveal that the DS layers of Wave-U-Net cause aliasing and may discard information useful for the separation. Although the effects of these problems may be reduced by training, to achieve a more reliable source separation method, we should design DS layers capable of overcoming the problems. With this belief, focusing on the fact that the DWT has an anti-aliasing filter and the perfect reconstruction property, we design the proposed layers. Experiments on music source separation show the efficacy of the proposed method and the importance of simultaneously considering the anti-aliasing filters and the perfect reconstruction property.
Tomohiko Nakamura, Hiroshi Saruwatari
ICASSP2
2020 Lifter Training and Sub-Band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials
abstract
In this paper, we propose computationally efficient and high-quality methods for statistical voice conversion (VC) with direct waveform modification based on spectral differentials. The conventional method with a minimum-phase filter achieves high-quality conversion but requires heavy computation in filtering. This is because the minimum phase using a fixed lifter of the Hilbert transform often results in a long-tap filter. One of our methods is a data-driven method for lifter training. Since this method takes filter truncation into account in training, it can shorten the tap length of the filter while preserving conversion accuracy. Our other method is sub-band processing for extending the conventional method from narrow-band (16 kHz) to full-band (48 kHz) VC, which can convert a full-band waveform with higher converted-speech quality. Experimental results indicate that 1) the proposed lifter-training method for narrow-band VC can shorten the tap length to 1/16 without degrading the converted-speech quality and 2) the proposed sub-band-processing method for full-band VC can improve the converted-speech quality than the conventional method.
Takaaki Saeki, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP4
2020 End-to-End Text-to-Speech Synthesis with Unaligned Multiple Language Units Based on Attention
Masashi Aso, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH3
2020 Multi-Speaker Text-to-Speech Synthesis Using Deep Gaussian Processes
abstract
Multi-speaker speech synthesis is a technique for modeling multiple speakers' voices with a single model.Although many approaches using deep neural networks (DNNs) have been proposed, DNNs are prone to overfitting when the amount of training data is limited.We propose a framework for multi-speaker speech synthesis using deep Gaussian processes (DGPs); a DGP is a deep architecture of Bayesian kernel regressions and thus robust to overfitting.In this framework, speaker information is fed to duration/acoustic models using speaker codes.We also examine the use of deep Gaussian process latent variable models (DGPLVMs).In this approach, the representation of each speaker is learned simultaneously with other model parameters, and therefore the similarity or dissimilarity of speakers is considered efficiently.We experimentally evaluated two situations to investigate the effectiveness of the proposed methods.In one situation, the amount of data from each speaker is balanced (speaker-balanced), and in the other, the data from certain speakers are limited (speaker-imbalanced). Subjective and objective evaluation results showed that both the DGP and DG-PLVM synthesize multi-speaker speech more effective than a DNN in the speaker-balanced situation.We also found that the DGPLVM outperforms the DGP significantly in the speakerimbalanced situation.
Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari
INTERSPEECH3
2020 Real-Time, Full-Band, Online DNN-Based Voice Conversion System Using a Single CPU
Takaaki Saeki, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH4
2020 Harmonic Lowering for Accelerating Harmonic Convolution for Audio Signals
Hirotoshi Takeuchi, Kunio Kashino, Yasunori Ohishi, Hiroshi Saruwatari
INTERSPEECH4
2020 Cross-Lingual Text-To-Speech Synthesis via Domain Adaptation and Perceptual Similarity Regression in Speaker Space
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
INTERSPEECH5
2020 Investigating Effective Additional Contextual Factors in DNN-Based Spontaneous Speech Synthesis
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
INTERSPEECH7
2020 SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on Gameplay
abstract
Developing a spontaneous speech corpus would be beneficial for spoken language processing and understanding. We present a speech corpus named the SMASH corpus, which includes spontaneous speech of two Japanese male commentators that made third-person audio commentaries during the gameplay of a fighting game. Each commentator ad-libbed while watching the gameplay with various topics covering not only explanations of each moment to convey the information on the fight but also comments to entertain listeners. We made transcriptions and topic tags as annotations on the recorded commentaries with our two-step method. We first made automatic and manual transcriptions of the commentaries and then manually annotated the topic tags. This paper describes how we constructed the SMASH corpus and reports some results of the annotations.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
LREC3
2020 DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech Corpus
abstract
In this paper, we investigate the effectiveness of using rich annotations in deep neural network (DNN)-based statistical speech synthesis. DNN-based frameworks typically use linguistic information as input features called context instead of directly using text. In such frameworks, we can synthesize not only reading-style speech but also speech with paralinguistic and nonlinguistic features by adding such information to the context. However, it is not clear what kind of information is crucial for reproducing paralinguistic and nonlinguistic features. Therefore, we investigate the effectiveness of rich tags in DNN-based speech synthesis according to the Corpus of Spontaneous Japanese (CSJ), which has a large amount of annotations on paralinguistic features such as prosody, disfluency, and morphological features. Experimental evaluation results shows that the reproducibility of paralinguistic features of synthetic speech was enhanced by adding such information as context.
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
LREC7
2020 Binaural Rendering From Distributed Microphone Signals Considering Loudspeaker Distance in Measurements
abstract
A method of binaural rendering from distributed microphone recordings that takes loudspeaker distance for measuring head-related transfer function (HRTF) into consideration is proposed. In general, to reproduce the binaural signals from the signals captured by multiple microphones in the recording area, the captured sound field is represented by plane-wave decomposition. Thus, HRTF is approximated as a transfer function from a plane-wave source in binaural rendering. To incorporate the distance in HRTF measurements, we propose a method based on the spherical-wave decomposition of a sound field, in which the HRTF is assumed to be measured from a point source. Result of experiments using HRTFs calculated by the boundary element method indicated that the accuracy of binaural signal reproduction by the proposed method based on the spherical-wave decomposition was higher than that by the plane-wave-decomposition-based method. We also evaluate the performance of signal conversion from distributed microphone measurements into binaural signals.
Naoto Iijima, Shoichi Koyama, Hiroshi Saruwatari
MMSP3
2020 Phase reconstruction from amplitude spectrograms based on directional-statistics deep neural networks
abstract
This paper presents a deep neural network (DNN)-based phase reconstruction method from amplitude spectrograms. In speech processing, an amplitude spectrogram is often used for processing, and the corresponding phases are reconstructed from the amplitude spectrogram by using the Griffin-Lim method. However, the Griffin-Lim method causes unnatural artifacts in synthetic speech. To solve this problem, we propose the directional-statistics DNNs for predicting phases from the amplitude spectrograms. We first propose the von Mises distribution DNN, which is a generative model having the von Mises distribution and models histograms of a periodic variable. We extend it for modeling group delay that has a stronger connection to the amplitude spectrograms. Furthermore, we generalize the group-delay modeling and propose another DNN called the sine-skewed generalized cardioid distribution DNN for modeling asymmetric histograms such as a group delay. Results from objective and subjective evaluations indicate that (1) our von Mises distribution DNN can predict group delay more accurately than predicting phases, (2) our DNN works as better initialization of the Griffin-Lim method, (3) the phase reconstruction methods based on our von Mises distribution DNN achieve better speech quality than the conventional Griffin-Lim method, and (4) our sine-skewed generalized cardioid distribution DNN models the group delay more accurately than our von Mises distribution DNN.
Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari
Signal Process.5
2020 Reciprocity gap functional in spherical harmonic domain for gridless sound field decomposition
abstract
A sound field decomposition method based on the reciprocity gap functional (RGF) in the spherical harmonic domain is proposed. To estimate and reconstruct a continuous sound field including sources by using multiple microphones, an intuitive and powerful strategy is to decompose the sound field into Green’s functions. Sparse-representation algorithms have been applied to this decomposition problem; however, it requires the discretization of the target region into grid points to construct a dictionary matrix. Discretization-based methods lead to decomposition errors of off-grid sources and high computational cost of sparse representation. We apply the RGF to sparse sound field decomposition, which makes it possible to decompose the sound field as a closed-form solution without discretization. In addition, the formulation in the spherical harmonic domain enables the flexible arrangement of microphones under the assumption of the spherical target region. Numerical simulation results indicated that high decomposition and reconstruction accuracies can be achieved by the proposed method, especially at low frequencies, with a low computational cost.
Yuhta Takida, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
Signal Process.4
2020 Acoustic model-based subword tokenization and prosodic-context extraction without language knowledge for text-to-speech synthesis
abstract
This paper presents text tokenization and context extraction without using language knowledge for text-to-speech (TTS) synthesis. To generate prosody, statistical parametric TTS synthesis typically requires the professional knowledge of the target language. Therefore, languages suitable for TTS synthesis are limited to only rich-resource languages. To achieve TTS synthesis without using language knowledge, we propose acoustic model-based subword tokenization and unsupervised extraction of prosodic contexts. The subword tokenization can determine language units suitable for prosody generation. The context extraction can retrieve contexts from pairs of subwords and prosody. The proposed methods function without language knowledge and can improve F0 prediction accuracy. Experimental evaluation demonstrates that 1) the training of proposed subword tokenization, which uses the expectation-maximization algorithm and deep neural networks, is empirically stable, 2) the proposed subword tokenization tokenizes text into subwords that are close to language-specific units, and 3) the proposed methods outperform the conventional methods using language model-based tokenization in terms of synthetic speech quality.
Masashi Aso, Shinnosuke Takamichi, Norihiro Takamune, Hiroshi Saruwatari
Speech Commun.4
2020 Blind Speech Extraction Based on Rank-Constrained Spatial Covariance Matrix Estimation With Multivariate Generalized Gaussian Distribution
abstract
In this article, we propose a new blind speech extraction (BSE) method that robustly extracts a directional speech from background diffuse noise by combining independent low-rank matrix analysis (ILRMA) and efficient rank-constrained spatial covariance matrix (SCM) estimation. To achieve more accurate BSE than ILRMA, which assumes each source to be a point source (rank-1 spatial model), the proposed method restores the lost spatial basis for the full-rank SCM of diffuse noise. We adopt the multivariate complex generalized Gaussian distribution (GGD) as the statistical generative model to express various types of observed signal. To estimate the model parameters for an arbitrary shape parameter of the multivariate GGD, we derive a new inequality for rank-constrained SCMs. Also, we propose new acceleration methods to accomplish much faster extraction than conventional blind source separation methods. In BSE experiments using simulated and real recorded data, we confirm that the proposed method achieves more accurate and faster speech extraction than conventional methods.
Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Multichannel Non-Negative Matrix Factorization Using Banded Spatial Covariance Matrices in Wavenumber Domain
abstract
Blind source separation exploiting multichannel information has long been a popular topic, and recently proposed methods based on the local Gaussian model have shown promising results despite its high computational cost for the case of many microphone signals. The low updating speed for such a model is mainly due to the inversion of a spatial covariance matrix, for which the complexity increases with the number of microphones, M, and is generally of order O(M3). Several projection-based approaches that attempt to concentrate energy on the diagonal part of the spatial covariance matrix have been introduced to circumvent the matrix inversion, which can reduce the complexity to O(M). In this article, we focus on the fast Fourier transform as a projection method because the energy concentration on the diagonal can be efficiently achieved compared with other projection-based methods. For the case where the diagonalization is imperfect, for example, owing to discontinuities at the edge of a linear array, we also developed a more robust algorithm approximating the tri-diagonal part of the spatial covariance matrix, which requires a complexity of O(M2) for the inversion by applying the Thomas algorithm. To remove the ad-hoc integration of post clustering after the decomposition, we also examine a self-clustering algorithm. Our evaluation shows better results than other previously proposed methods in terms of the separation quality under reverberant conditions as well as higher efficiency than multichannel non-negative matrix factorization.
Yuki Mitsufuji, Stefan Uhlich, Norihiro Takamune, Daichi Kitamura, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 Independent Low-Rank Matrix Analysis Based on Time-Variant Sub-Gaussian Source Model for Determined Blind Source Separation
abstract
Independent low-rank matrix analysis (ILRMA) is a fast and stable method of blind audio source separation. Conventional ILRMAs assume time-variant (super-)Gaussian source models, which can only represent signals that follow a super-Gaussian distribution. In this article, we focus on ILRMA based on a generalized Gaussian distribution (GGD-ILRMA) and propose a new type of GGD-ILRMA that adopts a time-variant sub-Gaussian distribution for the source model. We propose a new update scheme called generalized iterative projection for homogeneous source models (GIP-HSM) and obtain a convergence-guaranteed update rule for demixing spatial parameters by combining the GIP-HSM scheme and the majorization-minimization (MM) algorithm. Furthermore, a new extension of the MM algorithm is proposed for the convergence acceleration by applying the majorization-equalization algorithm to a multivariate case. In the experimental evaluation, we show the versatility of the proposed method, i.e., the proposed time-variant sub-Gaussian source model can be applied to various types of source signal.
Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo, Nobutaka Ono
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Feedforward Spatial Active Noise Control Based on Kernel Interpolation of Sound Field
abstract
A method for feedforward active noise control (ANC) over a spatial region is proposed. Conventional multipoint ANC aims to reduce the noise at multiple discrete positions; therefore, the noise reduction in the region between these points cannot be guaranteed. Recent studies revealed the possibility of spatial ANC, i.e., noise control in a continuous target region. These methods are essentially based on spherical/circular harmonic decomposition of the sound field by using spherical/circular arrays and have mainly been investigated for feedback control under the assumption of periodicity of the noise. We apply a sound field interpolation method based on kernel ridge regression to feedforward spatial ANC to control spatial nonstationary noise using distributed arrays. Numerical simulation results indicated that a large regional noise reduction is achieved by the proposed method compared with feedforward multipoint ANC.
Hayato Ito, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP4
2019 Robust Gridless Sound Field Decomposition Based on Structured Reciprocity Gap Functional in Spherical Harmonic Domain
abstract
A sound field reconstruction method for a region including sources is proposed. Under the assumption of spatial sparsity of the sources, this reconstruction problem has been solved by using sparse decomposition algorithms with the discretization of the target region. Since this discretization leads to the off-grid problem, we previously proposed a gridless sound field decomposition method based on the reciprocity gap functional in the spherical harmonic domain. Even though this method allows efficient estimation with a closed-form solution while avoiding the off-grid problem, the estimation using a single time-frequency bin can be greatly affected by measurement errors. We formulate an optimization problem using the identical structure of source locations in multiple time-frequency bins and derive an algorithm based on an annihilating filter. Numerical simulation results indicated that robustness against noise can be improved by the proposed method.
Yuhta Takida, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP4
2019 Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking
abstract
This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing voice leads to a richer musical experience and is used in double-tracking, a recording method in which two performances of the same phrase are recorded and mixed to create a richer, layered sound. However, singing voices synthesized using conventional DNN-based methods never vary because the synthesis process is deterministic and only one waveform is synthesized from one musical score. To address this problem, we use a GMMN to model the variation of the modulation spectrum of the pitch contour of natural singing voices and add a randomized inter-utterance variation to the pitch contour generated by conventional DNN-based singing voice synthesis. Experimental evaluations suggest that 1) our approach can provide perceptible inter-utterance pitch variation while preserving speech quality. We extend our approach to double-tracking, and the evaluation demonstrates that 2) GMMN-based neural double-tracking is perceptually closer to natural double-tracking than conventional signal processing-based artificial double-tracking is.
Hiroki Tamaru, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
ICASSP5
2019 Vocoder-free text-to-speech synthesis incorporating generative adversarial networks using low-/multi-frequency STFT amplitude spectra
abstract
This paper proposes novel training algorithms for vocoder-free text-to-speech (TTS) synthesis based on generative adversarial networks (GANs) that compensate for short-term Fourier transform (STFT) amplitude spectra in low/multi frequency resolution. Vocoder-free TTS using STFT amplitude spectra can avoid degradation of synthetic speech quality caused by the vocoder-based parameterization used in conventional TTS. Our previous work for the vocoder-based TTS proposed a method for incorporating the GAN-based distribution compensation into acoustic model training to improve synthetic speech quality. This paper extends the algorithm to the vocoder-free TTS and propose a GAN-based training algorithm using low-frequency-resolution amplitude spectra to overcome the difficulty in modeling complicated distribution of the high-dimensional spectra. In the proposed algorithm, amplitude spectra are transformed into low-frequency-resolution amplitude spectra by applying an average pooling function along with a frequency axis; then the GAN-based distribution compensation is performed in the low-frequency-resolution domain. Because the low-frequency-resolution amplitude spectra approximately emulate filter banks, the proposed algorithm is expected to improve synthetic speech quality by reducing differences in spectral envelopes of natural and synthetic speech. Furthermore, various frequency scales that are related to human speech perception (e.g., mel and inverse mel frequency scales) can be introduced to the proposed training algorithm by applying an frequency warping function to amplitude spectra. This paper also proposes a GAN-based training algorithm using multi-frequency-resolution amplitude spectra that uses both low- and original-frequency-resolution amplitude spectra to reduce the differences in not only spectral envelopes but also fine structures. Experimental results demonstrate that (1) GANs using low-frequency-resolution amplitude spectra improve speech quality and work robustly against the settings of the frequency resolution and hyperparameters, (2) in comparison among low-, original-, and multi-frequency-resolution amplitude spectra, the use of low-frequency-resolution ones work best improve the synthetic speech quality, and (3) the use of the inverse mel frequency scale for obtaining low-frequency-resolution amplitude spectra further improves synthetic speech quality.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
Comput. Speech Lang.3
2019 Bilevel Optimization Using Stationary Point of Lower-Level Objective Function for Discriminative Basis Learning in Nonnegative Matrix Factorization
abstract
In this letter, we address an audio signal separation problem and propose a new effective algorithm for solving a bilevel optimization in discriminative nonnegative matrix factorization (NMF). Recently, discriminative training of NMF bases has been developed for better signal separation in supervised NMF (SNMF), which exploits a priori training of given sample signals. The optimization in this method consists of a simultaneous minimization of two objective functions, resulting in a bilevel optimization problem with SNMF (BiSNMF), where conventional methods approximately solve this optimization. To strictly solve BiSNMF, we introduce a new algorithm with the following two features: (a) conversion of the optimization constraint into a penalty term and (b) optimization of the reformulated problem on the basis of a multiplicative steepest descent, ensuring the nonnegativity of variables. Experiments on music signal separation show the efficacy of the proposed algorithm.
Hiroaki Nakajima, Daichi Kitamura, Norihiro Takamune, Hiroshi Saruwatari, Nobutaka Ono
IEEE Signal Process. Lett.4
2019 Independent Deeply Learned Matrix Analysis for Determined Audio Source Separation
abstract
In this paper, we propose a new framework called independent deeply learned matrix analysis (IDLMA), which unifies a deep neural network (DNN) and independence-based multichannel audio source separation. IDLMA utilizes both pretrained DNN source models and statistical independence between sources for the separation, where the time-frequency structures of each source are iteratively optimized by a DNN while enhancing the estimation accuracy of the spatial demixing filters. As the source generative model, we introduce a complex heavy-tailed distribution to improve the separation performance. In addition, we address a semi-supervised situation; namely, a solo-recorded audio dataset can be prepared for only one source in the mixture signal. To solve the limited-data problem, we propose an appropriate data augmentation method to adapt the DNN source models to the observed signal, which enables IDLMA to work even in the semi-supervised situation. Experiments are conducted using music signals with a training dataset in both supervised and semi-supervised situations. The results show the validity of the proposed method in terms of the separation accuracy.
Naoki Makishima, Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hayato Sumino, Shinnosuke Takamichi, Hiroshi Saruwatari, Nobutaka Ono
IEEE ACM Trans. Audio Speech Lang. Process.7
2019 Three-Dimensional Sound Field Reproduction Based on Weighted Mode-Matching Method
abstract
A sound field reproduction method based on the spherical wavefunction expansion of sound fields is proposed, which can be flexibly applied to various array geometries and directivities. First, we formulate sound field synthesis as a minimization problem of some norm on the difference between the desired and synthesized sound fields, and then the optimal driving signals are derived by using the spherical wavefunction expansion of the sound fields. This formulation is closely related to the mode-matching method; a major advantage of the proposed method is the optimal weight on the mode determined according to the norm to be minimized instead of the empirical truncation in the mode-matching method. We also provide some examples of norms and their corresponding weights in analytical forms. Both interior and exterior sound field reproduction are considered in the proposed method, and some applications, such as multizone reproduction and interior reproduction with exterior cancellation, are also discussed. Numerical simulation results indicated that higher reproduction accuracy can be achieved by the proposed method than by the current pressure-matching and mode-matching methods.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Vectorwise Coordinate Descent Algorithm for Spatially Regularized Independent Low-Rank Matrix Analysis
abstract
Audio source separation is an important problem for many audio applications. Independent low-rank matrix analysis (ILRMA) is a recently proposed algorithm that employs the statistical independence between sources and the low-rankness of the time-frequency structure in each source. As reported in this paper, we have developed a new framework that enables us to introduce a spatial regularization of the demixing matrix in ILRMA. Since the conventional optimization cannot be applied to this regularized ILRMA, we derive a novel approach based on vectorwise coordinate descent, which does not require a step-size parameter and guarantees convergence. In experiments, ILRMA with beamforming-based regularization is evaluated as an application of the proposed framework.
Yoshiki Mitsui, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
ICASSP4
2018 Text-to-Speech Synthesis Using STFT Spectra Based on Low-/Multi-Resolution Generative Adversarial Networks
abstract
This paper proposes novel training algorithms for vocoder-free statistical parametric speech synthesis (SPSS) using short-term Fourier transform (STFT) spectra. Recently, text-to-speech synthesis using STFT spectra has been investigated since it can avoid quality degradation caused by the vocoder-based parameterization in conventional SPSS using a vocoder. In conventional SPSS using a vocoder, we previously proposed a training algorithm for integrating generative adversarial network (GAN)-based distribution compensation. To extend the algorithm to vocoder-free SPSS, we propose low- and multi-resolution GAN-based training algorithms for vocoder-free SPSS. In our algorithm that uses the low-resolution GAN, acoustic models are trained to minimize the weighted sum of the mean squared error between natural and generated spectra in the original resolution and adversarial loss to deceive discriminative models in the lower resolution. Since the low-resolution spectra are close to filter banks and their distribution becomes simpler, GAN-based distribution compensation works well. Furthermore, we propose an algorithm using multi-resolution GANs, which uses both the low-resolution GAN and original-resolution GAN. Experimental results demonstrate that 1) the low-resolution GAN works robustly to the setting of its frequency resolution and hyperparameter, and 2) compared the low-, original-, and multi-resolution GANs, the low-resolution GAN works the best to improve synthetic speech quality.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP3
2018 Sound Field Reproduction with Exterior Cancellation Using Analytical Weighting of Harmonic Coefficients
abstract
A method for sound field reproduction with the suppression of exterior radiation is proposed, which makes it possible to synthesize a desired sound field in a reverberant environment without prior knowledge of the transfer functions of the multiple loudspeakers. The objective function used to achieve this is formulated as the weighted sum of the interior reproduction error and exterior radiation power. The optimal driving signals are derived by harmonic expansion of both the interior and exterior sound fields. In contrast to the empirical coefficient truncation in the state of the art, in the proposed method, an optimal weighting of the harmonic coefficients is derived analytically. Numerical simulation results indicated that high interior reproduction accuracy and exterior power suppression can be achieved by the proposed method compared with the mode-matching method using harmonic-order truncation owing to the optimal weighting.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2018 CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects
Shinnosuke Takamichi, Hiroshi Saruwatari
LREC2
2018 Sound Field Recording Using Distributed Microphones Based on Harmonic Analysis of Infinite Order
abstract
A sound field recording method based on spherical or circular harmonic analysis for arbitrary array geometry and directivity of microphones is proposed. In current methods based on harmonic analysis, a sound field is decomposed into harmonic functions with a center given in advance, which is called a global origin, and their coefficients are obtained up to a certain truncation order using microphone measurements. However, the accuracy of the reconstructed sound field depends on the predefined position of the global origin and the truncation order, which makes it difficult to apply this technique to an asymmetric array since the criterion to determine the position of the global origin and the truncation order is not obvious. We formulate an estimate of the harmonic coefficients on the basis of infinite-order analysis. This formulation enables us to estimate the harmonic coefficients at an arbitrary desired position independently of the position of the global origin without truncation errors. Numerical simulation results indicated that the proposed method makes it possible to avoid performance degradation caused by inappropriate setting of the global origin.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
IEEE Signal Process. Lett.3
2018 Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks
abstract
A method for statistical parametric speech synthesis incorporating generative adversarial networks (GANs) is proposed. Although powerful deep neural networks techniques can be applied to artificially synthesize speech waveform, the synthetic speech quality is low compared with that of natural speech. One of the issues causing the quality degradation is an oversmoothing effect often observed in the generated speech parameters. A GAN introduced in this paper consists of two neural networks: a discriminator to distinguish natural and generated samples, and a generator to deceive the discriminator. In the proposed framework incorporating the GANs, the discriminator is trained to distinguish natural and generated speech parameters, while the acoustic models are trained to minimize the weighted sum of the conventional minimum generation loss and an adversarial loss for deceiving the discriminator. Since the objective of the GANs is to minimize the divergence (i.e., distribution difference) between the natural and generated speech parameters, the proposed method effectively alleviates the oversmoothing effect on the generated speech parameters. We evaluated the effectiveness for text-to-speech and voice conversion, and found that the proposed method can generate more natural spectral parameters and F0than conventional minimum generation error training algorithm regardless of its hyperparameter settings. Furthermore, we investigated the effect of the divergence of various GANs, and found that a Wasserstein GAN minimizing the Earth-Mover's distance works the best in terms of improving the synthetic speech quality.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Blind source separation based on independent low-rank matrix analysis with sparse regularization for time-series activity
abstract
In this paper, we propose a new blind source separation (BSS) method based on independent low-rank matrix analysis (ILRMA) with novel sparse regularization. ILRMA is a recently proposed BSS algorithm that simultaneously estimates a demixing matrix and source spectrogram models based on nonnegative matrix factorization (NMF). To improve the separation accuracy and stability, an additional constraint such as sparseness is needed but there have been no studies on this so far. In this study, we introduce an a priori statistical model for time-series amplitudes of source spectrograms, employing a new frequency-wise sparse regularization using estimates from the Bayesian postfilter to enhance the modeling accuracy. This regularization results in a bilevel optimization problem that consists of the estimation of a sparsity-emphasized source model using NMF and the separation of sources by ILRMA. In this paper, we present two approximated optimization schemes and their combination for performing regularized ILRMA. The efficacy of the proposed method is confirmed in a BSS experiment.
Yoshiki Mitsui, Daichi Kitamura, Shinnosuke Takamichi, Nobutaka Ono, Hiroshi Saruwatari
ICASSP5
2017 Spatio-temporal sparse sound field decomposition considering acoustic source signal characteristics
abstract
We propose a sound field decomposition method that takes into consideration spatio-temporal sparsity. It has been proved that sparse representation of a sound field is effective in reducing errors originating from spatial aliasing artifacts compared with conventional plane wave decomposition. In most current methods of sparse sound field decomposition, the spatial sparsity of the sound source distribution is only assumed. However, it is known that the temporal structure of the source signal to be decomposed can also be sparse in the time-frequency domain. We formulate an objective function for sparse sound field decomposition by using the ℓp,q-norm to simultaneously induce sparsity in the space and time domains. An optimization algorithm on the auxiliary function method is derived to solve it. Numerical simulations of acoustic holography indicate that the reconstruction accuracy can be improved by controlling the parameter of temporal sparsity. We also demonstrate that a statistical measure of the source signals can be used as an indicator to determine a nearly optimal parameter.
Naoki Murata, Shoichi Koyama, Norihiro Takamune, Hiroshi Saruwatari
ICASSP4
2017 Training algorithm to deceive Anti-Spoofing Verification for DNN-based speech synthesis
abstract
This paper proposes a novel training algorithm for high-quality Deep Neural Network (DNN)-based speech synthesis. The parameters of synthetic speech tend to be over-smoothed, and this causes significant quality degradation in synthetic speech. The proposed algorithm takes into account an Anti-Spoofing Verification (ASV) as an additional constraint in the acoustic model training. The ASV is a discriminator trained to distinguish natural and synthetic speech. Since acoustic models for speech synthesis are trained so that the ASV recognizes the synthetic speech parameters as natural speech, the synthetic speech parameters are distributed in the same manner as natural speech parameters. Additionally, we find that the algorithm compensates not only the parameter distributions, but also the global variance and the correlations of synthetic speech parameters. The experimental results demonstrate that 1) the algorithm outperforms the conventional training algorithm in terms of speech quality, and 2) it is robust against the hyper-parameter settings.
Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
ICASSP3
2017 Listening-area-informed sound field reproduction based on circular harmonic expansion
abstract
A sound field reproduction method that exploits prior information on listening areas is proposed. Most current methods are aimed at reproducing the sound field over the entire space or around listener locations. We formulate the objective function for this problem as the expectation minimization of the spatial squared error of the sound pressure inside the listening areas. The optimal driving signals are obtained by circular harmonic expansion. Comparing the proposed method with the mode-matching method, the advantage of the proposed method appears in the optimal weighting matrix of the circular harmonics, which depends on the location and range of the listening area. Numerical simulations indicated that high reproduction accuracy compared with other current methods can be maintained in the listening areas at high frequencies by using the proposed method.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2017 Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior Probabilities
abstract
Voice conversion (VC) using sequence-to-sequence learning of context posterior probabilities is proposed.Conventional VC using shared context posterior probabilities predicts target speech parameters from the context posterior probabilities estimated from the source speech parameters.Although conventional VC can be built from non-parallel data, it is difficult to convert speaker individuality such as phonetic property and speaking rate contained in the posterior probabilities because the source posterior probabilities are directly used for predicting target speech parameters.In this work, we assume that the training data partly include parallel speech data and propose sequence-to-sequence learning between the source and target posterior probabilities.The conversion models perform non-linear and variable-length transformation from the source probability sequence to the target one.Further, we propose a joint training algorithm for the modules.In contrast to conventional VC, which separately trains the speech recognition that estimates posterior probabilities and the speech synthesis that predicts target speech parameters, our proposed method jointly trains these modules along with the proposed probability conversion modules.Experimental results demonstrate that our approach outperforms the conventional VC.
Hiroyuki Miyoshi, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari
INTERSPEECH4
2017 Sampling-Based Speech Parameter Generation Using Moment-Matching Networks
abstract
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis.Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis produces completely the same speech, i.e., there is no inter-utterance variation in synthetic speech.To give synthetic speech natural inter-utterance variation, this paper builds DNN acoustic models that make it possible to randomly sample speech parameters.The DNNs are trained so that they make the moments of generated speech parameters close to those of natural speech parameters.Since the variation of speech parameters is compressed into a low-dimensional simple prior noise vector, our algorithm has lower computation cost than direct sampling of speech parameters.As the first step towards generating synthetic speech that has natural inter-utterance variation, this paper investigates whether or not the proposed sampling-based generation deteriorates synthetic speech quality.In evaluation, we compare speech quality of conventional maximum likelihood-based generation and proposed sampling-based generation.The result demonstrates the proposed generation causes no degradation in speech quality.
Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
INTERSPEECH3
2016 Sound field decomposition in reverberant environment using sparse and low-rank signal models
abstract
A sound field decomposition method for a reverberant environment is proposed. Sound field decomposition is the foundation of various acoustic signal processing applications and enables the estimation of the entire sound field from pressure measurements. Although spatial Fourier analysis of the sound field has been widely used, sparse decomposition of the sound field has recently been proved to be effective in several applications. However, in current methods, no constraints are imposed on ambiance components, whereas source components are assumed to be sparsely distributed in the space. This results in inaccurate decomposition in a reverberant environment. The proposed method is based on sparse and low-rank signal models, which are used for simultaneous decomposition of the observed signals into source and ambiance components. Numerical simulation results indicated that the decomposition accuracy is superior to that of current methods.
Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2016 Multichannel blind source separation based on non-negative tensor factorization in wavenumber domain
abstract
Multichannel non-negative matrix factorization based on a spatial covariance model is one of the most promising techniques for blind source separation. However, this approach is not tractable for a large number of microphones, M, because the computational cost is of order O(M3) per time-frequency bin. To circumvent this drawback, we propose non-negative tensor factorization in the wavenumber domain, which reduces the cost to the order O(M). It transforms microphone signals into the spatial frequency domain, a technique that is commonly used for soundfield reconstruction. The proposed method is compared to several blind source separation (BSS) methods in terms of separation quality and computational cost.
Yuki Mitsufuji, Shoichi Koyama, Hiroshi Saruwatari
ICASSP3
2016 Sparse sound field decomposition with multichannel extension of complex NMF
abstract
A sparse sound field decomposition method using prior information on source signals in the time-frequency domain is proposed. Sparse sound field decomposition has been proved to be effective for various acoustic signal processing applications. Current methods for sparse decomposition are based only on the spatial sparsity of the source distribution. However, it can be assumed that possible source signals to be decomposed are approximately known in advance. To exploit this prior information, we incorporated the complex nonnegative factorization model into sparse sound field decomposition. Since the magnitude spectrum of the possible source signals can be trained in advance, accuracy of the sparse decomposition can be improved even when the source signals are highly correlated and the sources are in a highly noisy environment. In addition, the proposed decomposition algorithm is derived using the auxiliary function method. Numerical experiments indicated that the sparse decomposition performance was significantly improved using the proposed method.
Naoki Murata, Shoichi Koyama, Hirokazu Kameoka, Norihiro Takamune, Hiroshi Saruwatari
ICASSP5
2016 Semi-Supervised Joint Enhancement of Spectral and Cepstral Sequences of Noisy Speech
Li Li 0063, Hirokazu Kameoka, Takuya Higuchi, Hiroshi Saruwatari
INTERSPEECH4
2016 Determined Blind Source Separation Unifying Independent Vector Analysis and Nonnegative Matrix Factorization
abstract
This paper addresses the determined blind source separation problem and proposes a new effective method unifying independent vector analysis (IVA) and nonnegative matrix factorization (NMF). IVA is a state-of-the-art technique that utilizes the statistical independence between sources in a mixture signal, and an efficient optimization scheme has been proposed for IVA. However, since the source model in IVA is based on a spherical multivariate distribution, IVA cannot utilize specific spectral structures such as the harmonic structures of pitched instrumental sounds. To solve this problem, we introduce NMF decomposition as the source model in IVA to capture the spectral structures. The formulation of the proposed method is derived from conventional multichannel NMF (MNMF), which reveals the relationship between MNMF and IVA. The proposed method can be optimized by the update rules of IVA and single-channel NMF. Experimental results show the efficacy of the proposed method compared with IVA and MNMF in terms of separation accuracy and convergence speed.
Daichi Kitamura, Nobutaka Ono, Hiroshi Sawada, Hirokazu Kameoka, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 Efficient multichannel nonnegative matrix factorization exploiting rank-1 spatial model
abstract
This paper proposes a new efficient multichannel nonnegative matrix factorization (NMF) method. Recently, multichannel NMF (MNMF) has been proposed as a means of solving the blind source separation problem. This method estimates a mixing system of sources and attempts to separate them in a blind fashion. However, this method is strongly dependent on its initial values because there are no constraints in the spatial models. To solve this problem, we introduce a rank-1 spatial model into MNMF. The proposed method estimates a demixing matrix while representing sources using NMF bases and can be optimized by the update rules of independent vector analysis and conventional single-channel NMF. Experimental results show the efficacy of the proposed method in terms of robustness and convergence speed.
Daichi Kitamura, Nobutaka Ono, Hiroshi Sawada, Hirokazu Kameoka, Hiroshi Saruwatari
ICASSP5
2015 Structured sparse signal models and decomposition algorithm for super-resolution in sound field recording and reproduction
abstract
A method for achieving super-resolution of sound field recording and reproduction is proposed. To obtain driving signals of loudspeakers for reproduction from received signals of microphones, sparse signal decomposition makes it possible to reduce spatial aliasing artifacts when the number of microphones is less than that of loudspeakers. For more accurate and robust signal decomposition, we propose three types of group sparse signal model based on the physical properties of a sound field. In addition, a decomposition algorithm is derived to address these signal models as an extension of M-FOCUSS. In the simulation experiments, the accuracy of the sparse decomposition was significantly improved compared with that of M-FOCUSS. Furthermore, the accuracy of sound field reproduction using our proposed method was higher than that using current methods, especially at frequencies above the spatial Nyquist frequency.
Shoichi Koyama, Naoki Murata, Hiroshi Saruwatari
ICASSP3
2015 Statistical modeling of binaural signal and its application to binaural source separation
abstract
This paper addresses a new statistical model of binaural signals and its application to efficient binaural source separation. Binaural source separation is always required to retain a spatial cue of the separated sound, such as a head-related transfer function (HRTF). However, the direct use of an HRTF is not realistic because this information is normally not known in advance. To cope with this problem, first, we focus on the difference between signal probability density functions at both ears, which can be blindly estimated by using our previous work on higher-order statistics. Next, we derive a sound-localization-preserved generalized minimum mean-square error short-time spectral amplitude estimator. Objective and subjective experiments show the efficacy of the proposed method in terms of spatial quality.
Yuki Murota, Daichi Kitamura, Shoichi Koyama, Hiroshi Saruwatari, Satoshi Nakamura 0001
ICASSP4
2015 Multichannel Signal Separation Combining Directional Clustering and Nonnegative Matrix Factorization with Spectrogram Restoration
abstract
In this paper, to address problems in multichannel music signal separation, we propose a new hybrid method that combines directional clustering and advanced nonnegative matrix factorization (NMF). The aims of multichannel music signal separation technology is to extract a specific target signal from observed multichannel signals that contain multiple instrumental sounds. In previous studies, various methods using NMF have been proposed, but many problems remain including poor separation accuracy and lack of robustness. To solve these problems, we propose a new supervised NMF (SNMF) with spectrogram restoration and a hybrid method that concatenates the proposed SNMF after directional clustering. Via the extrapolation of supervised spectral bases, the proposed SNMF attempts both target signal separation and reconstruction of the lost target components, which are generated by preceding directional clustering. In addition, we experimentally reveal the trade-off between separation and extrapolation abilities and propose a new scheme for adaptive divergence, where the optimal divergence can be automatically changed in each time frame according to the local spatial conditions. The results of an evaluation experiment show that our proposed hybrid method outperforms the conventional music signal separation methods.
Daichi Kitamura, Hiroshi Saruwatari, Hirokazu Kameoka, Yu Takahashi, Kazunobu Kondo, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Music signal separation based on Bayesian spectral amplitude estimator with automatic target prior adaptation
abstract
In this paper, we propose a new approach for addressing music signal separation based on the generalized Bayesian estimator with automatic prior adaptation. This method consists of three parts, namely, the generalized MMSE-STSA estimator with a flexible target signal prior, the NMF-based dynamic interference spectrogram estimator, and closed-form parameter estimation for the statistical model of the target signal based on higher-order statistics. The statistical model parameter of the hidden target signal can be detected automatically for optimal Bayesian estimation with online target-signal prior adaptation. Our experimental evaluation can show the efficacy of the proposed method.
Yuki Murota, Daichi Kitamura, Shunsuke Nakai, Hiroshi Saruwatari, Satoshi Nakamura 0001, Yu Takahashi, Kazunobu Kondo
ICASSP4
2014 Musical-noise-free blind speech extraction integrating microphone array and iterative spectral subtraction
abstract
In this paper, we propose a musical-noise-free blind speech extraction method using a microphone array for application to nonstationary noise. In our previous study, it was found that optimized iterative spectral subtraction (SS) results in speech enhancement with almost no musical noise generation, but this method is valid only for stationary noise. The proposed method consists of iterative blind dynamic noise estimation by, e.g., independent component analysis (ICA) or multichannel Wiener filtering, and musical-noise-free speech extraction by modified iterative SS, where multiple iterative SS is applied to each channel while maintaining the multichannel property reused for the dynamic noise estimators. Also, in relation to the proposed method, we discuss the justification of applying ICA to signals nonlinearly distorted by SS. From objective and subjective evaluations simulating a real-world hands-free speech communication system, we reveal that the proposed method outperforms the conventional methods.
Ryoichi Miyazaki, Hiroshi Saruwatari, Satoshi Nakamura 0001, Kiyohiro Shikano, Kazunobu Kondo, Jonathan Blanchette, Martin Bouchard 0001
Signal Process.2
2014 Alaryngeal Speech Enhancement Based on One-to-Many Eigenvoice Conversion
abstract
In this paper, we present novel speaking-aid systems based on one-to-many eigenvoice conversion (EVC) to enhance three types of alaryngeal speech: esophageal speech, electrolaryngeal speech, and body-conducted silent electrolaryngeal speech. Although alaryngeal speech allows laryngectomees to utter speech sounds, it suffers from the lack of speech quality and speaker individuality. To improve the speech quality of alaryngeal speech, alaryngeal-speech-to-speech (AL-to-Speech) methods based on statistical voice conversion have been proposed. In this paper, one-to-many EVC capable of flexibly controlling the converted voice quality by adapting the conversion model to given target natural voices is further implemented for the AL-to-Speech methods to effectively recover speaker individuality of each type of alaryngeal speech. These proposed systems are compared with each other from various perspectives. The experimental results demonstrate that our proposed systems are capable of effectively addressing the issues of alaryngeal speech, e.g., yielding significant improvements in speech quality of each type of alaryngeal speech.
Hironori Doi, Tomoki Toda, Keigo Nakamura, Hiroshi Saruwatari, Kiyohiro Shikano
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Musical noise analysis for Bayesian minimum mean-square error speech amplitude estimators based on higher-order statistics
Hiroshi Saruwatari, Suzumi Kanehara, Ryoichi Miyazaki, Kiyohiro Shikano, Kazunobu Kondo
INTERSPEECH1
2013 Design of multichannel frequency domain statistical-based enhancement systems preserving spatial cues via spectral distances minimization
Frédéric Mustière, Martin Bouchard 0001, Hossein Najaf-Zadeh, Ramin Pichevar, Louis Thibault, Hiroshi Saruwatari
Signal Process.6
2012 Musical-noise-free speech enhancement: Theory and evaluation
abstract
In this paper, we propose a new theory of nonlinear noise reduction with a perfectly musical-noise-free property, where no musical noise is generated even for a high signal-to-noise ratio. To achieve high-quality noise reduction with low musical noise, an iterative spectral subtraction method, i.e., recursively applied weak nonlinear signal processing, has been proposed. Although evaluation experiments indicated the existence of an appropriate parameter setting that gives a musical-noise-free state, no theoretical studies have been carried out. Therefore, in this paper, we theoretically derive pairs of internal parameters that satisfy the musical-noise-free condition by analysis based on higher-order statistics. It is clarified that finding a fixed point in the kurtosis of noise spectra enables the reproduction of the musical-noise-free state, and comparative experiments with commonly used noise reduction methods show the efficacy of the proposed method.
Ryoichi Miyazaki, Hiroshi Saruwatari, Takayuki Inoue, Kiyohiro Shikano, Kazunobu Kondo
ICASSP2
2012 Speech kurtosis estimation from observed noisy signal based on generalized Gaussian distribution prior and additivity of cumulants
abstract
In this paper, we propose a new method for stable estimation of the kurtosis of a speech power spectrum. Speech kurtosis can be used for the prediction of speech recognition accuracy as reported in recent studies. However, the conventional estimation method is very unstable owing to the high sensitivity of higher-order statistics. To overcome this problem, we introduce the generalized Gaussian distribution prior in order to avoid the calculation of higher-order statistics, and construct a kurtosis table that directly represents the relationship among the kurtosis of speech, noise, and their mixture in the power spectrum domain. Speech kurtosis can be estimated stably from observable data by looking up values in the table. An experimental evaluation confirms the efficacy of the proposed method.
Ryo Wakisaka, Hiroshi Saruwatari, Kiyohiro Shikano, Tomoya Takatani
ICASSP2
2012 Statistical approach to voice quality control in esophageal speech enhancement
abstract
This paper describes a voice quality control method in statistical esophageal speech enhancement. Esophageal speech is produced by one of the alternative speaking methods for laryngectomees. Its naturalness and intelligibility are much lower than those of natural voices and its voice quality sounds similar even if uttered by different laryngectomees. These issues are alleviated by a statistical voice conversion method from esophageal speech into normal speech (ES-to-Speech) based on eigenvoices. This method is capable of determining converted voice quality using a few target voice samples. In this paper, we propose ES-to-Speech using regression techniques to make it possible to manually control the converted voice quality by manipulating a few intuitively controllable parameters even if no target voice sample is available. The effectiveness of the proposed method is confirmed by experimental evaluations.
Kenzo Yamamoto, Tomoki Toda, Hironori Doi, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP4
2012 Evaluation of Many-to-Many Alignment Algorithm by Automatic Pronunciation Annotation Using Web Text Mining
abstract
The need for robust pronunciation annotation over out-of-vocabulary (OOV) words has been increasing with the development of an application that deals with proper nouns and brand-new words, such as Voice Search. In robust pronunciation annotation over OOV words, the alignment between graphemes and phonemes is vital data. For a many-to-many alignment algorithm between graphemes and phonemes, we describe its problems and methods to overcome them. An evaluation experiment of a many-to-many alignment by automatic pronunciation annotation using Web text mining is also performed. That experimental result shows that the proposed many-to-many alignment produces an alignment that has the high generalization ability for OOV words while avoiding degradation of the accuracy of the pronunciation annotation compared with the conventional approach.
Keigo Kubo, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2012 Spoken Inquiry Discrimination Using Bag-of-Words for Speech-Oriented Guidance System
abstract
INTERSPEECH 2012: The 13th Annual Conference of the International Speech Communication Association, September 9-13, 2012, Portland, Oregon, USA.
Haruka Majima, Rafael Torres 0001, Yoko Fujita, Hiromichi Kawanami, Tomoko Matsui, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH6
2012 Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech
Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
Speech Commun.3
2012 Musical-Noise-Free Speech Enhancement Based on Optimized Iterative Spectral Subtraction
abstract
In this paper, we provide a theoretical analysis of the amount of musical noise in iterative spectral subtraction, and its optimization method for the least musical noise generation. To achieve high-quality noise reduction with low musical noise, iterative spectral subtraction, i.e., iteratively applied weak nonlinear signal processing, has been proposed. Although the effectiveness of the method has been reported experimentally, there have been no theoretical studies. Therefore, in this paper, we formulate the generation process of musical noise by tracing the change in kurtosis of noise spectra, and conduct a comparison of the amount of musical noise for different parameter settings but the same achieved level of noise attenuation. Furthermore, we theoretically derive the optimal internal parameters that generate no musical noise. It is clarified that to find a fixed point in kurtosis yields the no-musical-noise property. Comparative experiments with commonly used noise reduction methods show the proposed method's efficacy.
Ryoichi Miyazaki, Hiroshi Saruwatari, Takayuki Inoue, Yu Takahashi, Kiyohiro Shikano, Kazunobu Kondo
IEEE Trans. Speech Audio Process.2
2011 Blind noise suppression for Non-Audible Murmur recognition with stereo signal processing
abstract
In this paper, we propose a blind noise suppression method for Non-Audible Murmur (NAM) recognition. NAM is a very soft whispered voice detected with NAM microphone, which is one of the body-conductive microphones. Due to its recording mechanism, the detected signal suffers from noise caused by speaker's movements. In the proposed method using a stereo signal detected with two NAM microphones, the noise is estimated with blind source separation, and then, spectral subtraction is performed in each channel to reduce the noise. Moreover, channel selection is performed frame by frame to generate less distorted monaural NAM signal. Experimental results show that 1) word accuracy in large vocabulary continuous NAM recognition is degraded from 69.2% to 53.6% by the noise and 2) it is significantly recovered to 63.3% in a simulated situation and 58.6% in a real situation with the proposed method.
Shunta Ishii, Tomoki Toda, Hiroshi Saruwatari, Sakriani Sakti, Satoshi Nakamura 0001
ASRU3
2011 Acoustic model training for non-audible murmur recognition using transformed normal speech data
abstract
In this paper we present a novel approach to acoustic model training for non-audible murmur (NAM) recognition using normal speech data transformed into NAM data. NAM is extremely soft murmur, that is so quiet that people around the speaker can hardly hear it. It is detected directly through the soft tissue of the head using a special body-conductive microphone, NAM microphone, placed on the neck below the ear. NAM recognition is one of the promising silent speech interfaces for man-machine speech communication. We have previously shown the effectiveness of speaker adaptive training (SAT) based on constrained maximum likelihood linear regression (CMLLR) in NAM acoustic model training. However, since the amount of available NAM data is still small, the effect of SAT is limited. In this paper we propose modified SAT methods capable of using a larger amount of normal speech data by transforming them into NAM data. The experimental results demonstrate that the pro posed methods yield an absolute increase of approximately 2% in word accuracy compared with the conventional method.
Denis Babani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP3
2011 An evaluation of alaryngeal speech enhancement methods based on voice conversion techniques
abstract
In this study, we evaluate our proposed methods for enhancing alaryngeal speech based on statistical voice conversion techniques. Voice conversion based on a Gaussian mixture model has been applied to the conversion of alaryngeal speech into normal speech (AL-to-Speech). Moreover, one-to-many eigenvoice conversion (EVC) has also been applied to AL-to-Speech to enable the recovery of the original voice quality of laryngectomees even if only one arbitrary utterance of the original voice is available. VC/EVC-based AL-to-Speech systems have been developed for several types of alaryngeal speech, such as esophageal speech (ES), electrolaryngeal speech (EL), and body-conducted silent electrolaryngeal speech (silent EL). These proposed systems are compared with each other from various perspectives. The experimental results demonstrate that our proposed systems yield significant enhancement effects on each type of alaryngeal speech.
Hironori Doi, Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP4
2011 Theoretical analysis of musical noise in Wiener filtering family via higher-order statistics
abstract
Recently, one of the authors has reported that the amount of generated musical noise is strongly correlated with higher-order statistics of the power spectra. On the basis of this finding, in this paper, we provide a new theoretical analysis of the amount of musical noise generated via the Wiener filtering family. Our theoretical analysis allows the universal performance description from the viewpoint of the amount of musical noise generation and that of noise reduction, enabling reasonable sound quality comparison under the same noise reduction performance. From a mathematical analysis and evaluation experiments, we also clarify which parameter settings result in less musical noise generated in the Wiener filtering family.
Takayuki Inoue, Hiroshi Saruwatari, Kiyohiro Shikano, Kazunobu Kondo
ICASSP2
2011 Robust sound field reproduction integrating multi-point sound field control and wave field synthesis
abstract
For a reproduced sound field, the competing goals between the listening area and reproduction accuracy in an actual environment is one of the most important problems in sound field reproduction using loudspeakers. In this paper, we propose a new method of balancing these goals with absolute accuracy using an inverse filter of the room acoustics: the null space of a generalized inverse matrix given by a compensation filter of the wave field outside the control points. To develop an expression for the compensation filter, we use the loudspeaker driving function of wave field synthesis (WFS) in stead of the filter used in conventional studies. By using WFS, the proposed method overcomes the compensation limitation of auditory distance and azimuth perception outside the control points. The results of computer simulations revealed that the proposed method balances the competing goals and has wide applicability in a spatial domain with high accuracy of reproduction both under free-field conditions and in a simulation model with room reflection.
Noriyoshi Kamado, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP2
2011 Automatic musical thumbnailing based on audio object localization and its evaluation
abstract
In this paper, to automatically generate musical thumbnails that con tain the main part of the original tune, we propose a new estimation method for identifying structure changes in stereo tunes based on localization information. The proposed method can estimate the main parts of a musical tune by analyzing the specific timing when localization information changes under the assumption that the changing time of the localization approximately corresponds to the timing of the musical structure change. We evaluate the effectiveness of the proposed method by objective and subjective assessments. The experimental results show that the proposed method is effective in automating musical structure analysis for generating musical thumb nails.
Hiroyuki Nawata, Noriyoshi Kamado, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP3
2011 Speaker-Adaptive Speech Synthesis Based on Eigenvoice Conversion and Language-Dependent Prosodic Conversion in Speech-to-Speech Translation
abstract
This paper describes a novel approach based on voice conversion (VC) to speaker-adaptive speech synthesis for speech-tospeech translation. Voice quality of translated speech in an output language is usually different from that of an input speaker of the translation system since a text-to-speech system is developed with another speaker’s voices in the output language. To render the input speaker’s voice quality in the translated speech, we propose a voice quality control method based on one-tomany eigenvoice conversion (EVC) and language-dependent prosodic conversion. Spectral parameters of the translated speech are effectively converted by one-to-many EVC enabling unsupervised speaker adaptation. Moreover, prosodic parameters are modified considering their global differences between the input and output languages. The effectiveness of the proposed method is confirmed by experimental evaluations on cross-lingual VC among Japanese, English, and Chinese. Index Terms: speech-to-speech translation, speech synthesis,
Nobuhiko Hattori, Tomoki Toda, Hisashi Kawai, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2011 Theoretical Analysis of Musical Noise and Speech Distortion in Structure-Generalized Parametric Blind Spatial Subtraction Array
abstract
INTERSPEECH 2011: 12th Annual Conference of the International Speech Communication Association, 28-31 August, 2011, Florence, Italy.
Ryoichi Miyazaki, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH2
2011 Blind Speech Prior Estimation for Generalized Minimum Mean-Square Error Short-Time Spectral Amplitude Estimator
abstract
In this paper, to achieve high-quality speech enhancement, we introduce the generalized minimum mean-square error short-time spectral amplitude estimator with a new blind prior estimation of the speech probability density function (p.d.f.). To deal with various types of speech signals with different p.d.f., we propose an algorithm of speech kurtosis estimation based on moment-cumulant transformation for blind adaptation to the shape parameter of speech p.d.f. From the objective and subjective evaluation experiments, we show the improved noise reduction performance of the proposed method.
Ryo Wakisaka, Hiroshi Saruwatari, Kiyohiro Shikano, Tomoya Takatani
INTERSPEECH2
2011 Theoretical Analysis of Musical Noise in Generalized Spectral Subtraction Based on Higher Order Statistics
abstract
In this paper, we provide a new theoretical analysis of the amount of musical noise generated via generalized spectral subtraction based on higher order statistics. Power spectral subtraction is the most commonly used spectral subtraction method, and in our previous study a musical noise assessment theory limited to the power spectral domain was proposed. In this paper, we propose a generalization of our previous theory on spectral subtraction for arbitrary exponent parameters. We can thus compare the amount of musical noise between any exponent domains from the results of our analysis. We also clarify that less musical noise is generated when we choose a lower exponent spectral domain; this implies that there is no theoretical justification for using power/amplitude spectral subtraction.
Takayuki Inoue, Hiroshi Saruwatari, Yu Takahashi, Kiyohiro Shikano, Kazunobu Kondo
IEEE Trans. Speech Audio Process.2
2011 Musical Noise Controllable Algorithm of Channelwise Spectral Subtraction and Adaptive Beamforming Based on Higher Order Statistics
abstract
In this paper, we propose a musical-noise-controllable algorithm for array signal processing with the aim for high-performance and high-quality noise reduction. Recently, many methods of integrating linear microphone array signal processing and nonlinear signal processing for noise reduction have been studied, but these methods often suffer from the problem of musical noise. In the proposed algorithm, channelwise spectral subtraction is applied before adaptive array signal processing. We also introduce a new automatic control algorithm to obtain the subtraction strength parameter used in the spectral subtraction, which depends on the amount of generated musical noise, measured by higher order statistics. We confirm the effectiveness of the proposed algorithm via objective and subjective evaluations.
Hiroshi Saruwatari, Yohei Ishikawa, Yu Takahashi, Takayuki Inoue, Kiyohiro Shikano, Kazunobu Kondo
IEEE Trans. Speech Audio Process.1
2010 Statistical approach to enhancing esophageal speech based on Gaussian mixture models
abstract
This paper presents a novel method of enhancing esophageal speech using statistical voice conversion. Esophageal speech is one of the alternative speaking methods for laryngectomees. Although it doesn't require any external devices, generated voices sound unnatural. To improve the intelligibility and naturalness of esophageal speech, we propose a voice conversion method from esophageal speech into normal speech. A spectral parameter and excitation parameters of target normal speech are separately estimated from a spectral parameter of the esophageal speech based on Gaussian mixture models. The experimental results demonstrate that the proposed method yields significant improvements in intelligibility and naturalness. We also apply one-to-many eigenvoice conversion to esophageal speech enhancement for flexibly controlling enhanced voice quality.
Hironori Doi, Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP4
2010 Complex Newton algorithm for blind signal extraction of speech in diffuse noise
abstract
Several recent methods for speech enhancement in presence of diffuse background noise use frequency domain blind signal separation to estimate the diffuse noise and a nonlinear post filter to suppress this estimated noise. This paper presents a frequency domain blind signal extraction method for estimating the diffuse noise in place of the frequency domain blind signal separation. The method is based on the minimization by means of a complex Newton algorithm of a cost function depending of the modulus of the extracted component. The proposed complex Newton method is compared to the gradient descent on the same cost function and to the blind signal separation approach.
Jani Even, Hiroshi Saruwatari, Kiyohiro Shikano, Tomoya Takatani
ICASSP2
2010 Speech enhancement in presence of diffuse background noise: Why using blind signal extraction?
abstract
This paper study the blind estimation of the diffuse background noise for the hands-free speech interface. Some recent papers showed that it is possible to use blind signal separation (BSS) to estimate the diffuse background noise by suppressing the speech component after all the components were separated. In particular, the scale indeterminacy of BSS is avoided by using the projection back method. In this paper, we study an alternative to the projection back for the noise estimation and justify the use of blind signal extraction BSE rather than BSS.
Jani Even, Hiroshi Saruwatari, Kiyohiro Shikano, Tomoya Takatani
ICASSP2
2010 Non-parallel training for many-to-many eigenvoice conversion
abstract
This paper presents a novel training method of an eigenvoice Gaussian mixture model (EV-GMM) effectively using non-parallel data sets for many-to-many eigenvoice conversion, which is a technique for converting an arbitrary source speaker's voice into an arbitrary target speaker's voice. In the proposed method, an initial EV-GMM is trained with the conventional method using parallel data sets consisting of a single reference speaker and multiple pre-stored speakers. Then, the initial EV-GMM is further refined using non-parallel data sets including a larger number of pre-stored speakers while considering the reference speaker's voices as hidden variables. The experimental results demonstrate that the proposed method yields significant quality improvements in converted speech by enabling us to use data of a larger number of pre-stored speakers.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP3
2010 MMSE STSA estimator with nonstationary noise estimation based on ICA for high-quality speech enhancement
abstract
In this paper, we propose a new blind speech extraction method consisting of a minimum mean-square error short-time spectral amplitude (MMSE STSA) estimator and noise estimation based on independent component analysis (ICA). First, we perform a computer simulation using the artificial noise whose stationarity could be controlled parametrically, and the obtained results indicate that the proposed method is superior to conventional methods, such as blind spatial subtraction array (BSSA) and the original MMSE STSA estimator under the non-point-source and nonstationary noise condition. Finally, we conduct an experiment in an actual railway-station environment, and objective and subjective evaluations to confirm the advantage of the proposed method in the real world.
Ryoi Okamoto, Yu Takahashi, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP3
2010 Theoretical musical-noise analysis and its generalization for methods of integrating beamforming and spectral subtraction based on higher-order statistics
abstract
In this paper, we conduct a theoretical analysis of the amount of musical noise generated via methods of integrating beamforming and spectral subtraction (SS) based on higher-order statistics under the same noise reduction performance condition. In our previous analysis, we did not consider the effect of flooring technique in SS and the fact that the noise reduction performances of the integration methods are not equivalent. Then, in this study, we analyze the amount of generated musical noise with consideration of such problems. As a result of the analysis, it is clarified that an appropriate structure depends on both the parameters of SS and the statistical characteristics of the input signal. Moreover, it is also revealed that a specific structure is proper to reduce the musical noise for almost all cases.
Yu Takahashi, Hiroshi Saruwatari, Hiroshi Shikano, Kazunobu Kondo
ICASSP2
2010 Close speaker cancellation for suppression of non-stationary background noise for hands-free speech interface
abstract
This paper presents a noise cancellation method based on the ability to efficiently cancel a close target speaker contribution from the signals observed at a microphone array. The proposed method exploits this specificity in the case of the hands-free speech interface. This method is in particular able to deal with non-stationary noise. The method can be divided in three steps. First, the steering vector pointing at the target user is estimated from the covariance of the observed signals. Then the noise estimate is obtained by cancelling the user's contribution. During this step the speech pauses are also estimated. Finally a post-filter is used to suppress this estimated noise from the observed signals. The post-filter strength is controlled by using the estimated noise during the speech pauses as reference. A 20k-words dictation task in presence of non-stationary diffuse background noise at different SNR levels illustrates the effectiveness of the proposed method.
Jani Even, Carlos Toshinori Ishi, Hiroshi Saruwatari, Norihiro Hagita
INTERSPEECH3
2010 The use of air-pressure sensor in electrolaryngeal speech enhancement based on statistical voice conversion
abstract
In our previous work, we proposed a speaking-aid system converting electrolaryngeal speech (EL speech) to normal speech using a statistical voice conversion technique. The main weakness of our system is the difficulty of estimating natural contours of the fundamental frequency (F0) from EL speech including only built-in F0 contours. This paper proposes another speaking-aid system with an air-pressure sensor to enable laryngectomees to control F0 contours of the EL speech using their breathing air. The experimental result demonstrates that 1) the correlation coefficient of F0 contours between the converted and the target speech is improved from 0.58 to 0.78 by the use of the air-pressure sensor and 2) the synthetic speech converted by the proposed system sounds more natural and is more preferred to that by our conventional aid system.
Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2010 Adaptive voice-quality control based on one-to-many eigenvoice conversion
abstract
This paper presents adaptive voice-quality control methods based on one-to-many eigenvoice conversion. To intuitively control the converted voice quality by manipulating a small number of control parameters, a multiple regression Gaussian mixture model (MR-GMM) has been proposed. The MR-GMM also allows us to estimate the optimum control parameters if target speech samples are available. However, its adaptation performance is limited because the number of control parameters is too small to widely model voice quality of various target speakers. To improve the adaptation performance while keeping capability of voice-quality control, this paper proposes an extended MR-GMM (EMR-GMM) with additional adaptive parameters to extend a subspace modeling target voice quality. Experimental results demonstrate that the EMR-GMM yields significant improvements of the adaptation performance while allowing us to intuitively control the converted voice quality.
Kumi Ohta, Tomoki Toda, Yamato Ohtani, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2010 Comparison of methods for topic classification in a speech-oriented guidance system
abstract
INTERSPEECH2010: 11th Annual Conference of the International Speech Communication Association, September 26-30, 2010, Chiba, Japan.
Rafael Torres 0001, Shota Takeuchi, Hiromichi Kawanami, Tomoko Matsui, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH5
2010 Improvement of speech recognition performance for spoken-oriented robot dialog system using end-fire array
abstract
In this paper, we propose a microphone array structure for a spoken-oriented robot dialog system that is designed to discriminate the direction of arrival (DOA) of the target speech and that of the robot internal noise. First, we investigate the performance of the noise estimation conducted by semi-blind source separation (SBSS) in presence of both the diffuse background noise and the robot internal noise. The result indicates that the noise estimation of the SBSS is not good. Next, we analyze the DOA of the robot internal noise in order to determine the reason of the above result; we find out that the internal noise is always in-phase at the microphone array and overlap spacial with the target speech. Based on this fact, we propose to change the microphone array structure from the broadside array to the end-fire array in order to discriminate the DOAs of the target speech and the internal noise. Finally, we evaluate the word accuracy in a dictation task in presence of both diffuse background noise and robot internal noise to confirm the advantage of the proposed structure. Simulation results shows that the proposed microphone array structure results in approximately 10% improvement of the speech recognition performance.
Hiroshi Sawada, Jani Even, Hiroshi Saruwatari, Kiyohiro Shikano, Tomoya Takatani
IROS3
2009 Multiple ICA-based real-time blind source extraction applied to handy size microphone
abstract
A new blind source extraction method in widespread noise conditions is proposed, which is based on multiple frequency-domain independent component analysis (FDICA) combining projection back and spectral subtraction. In addition, We implement the proposed method to digital signal processor (DSP) for a more realistic real-time operation, and develop a new blind source extraction (BSE) microphone which can extract a target sound in real-time. In this paper, we illustrate and evaluate the proposed method and BSE microphone. And experimental results reveal that the extraction performance of the proposed method are superior to that of conventional methods, and we show the efficacy of microphone.
Takashi Hiekata, Takashi Morita 0003, Youhei Ikeda, Hiroshi Hashimoto, Yu Takahashi, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP7
2009 Kernel-based nonlinear independent component analysis for underdetermined blind source separation
abstract
In this paper we propose a new unsupervised training method for nonlinear spatial filter using a new independent component analysis based on kernel infomax. The nonlinearity of the spatial filter used in this paper is equivalent to the integration of beamforming and spectral subtraction, and the whole structure is optimized by independent component analysis in the reproducing kernel Hilbert space. The optimized filter is shown to be capable of achieving better quality output than the conventional method based on time-frequency binary masking.
Shigeki Miyabe, Biing-Hwang Juang, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP3
2009 Acoustic compensation methods for body transmitted speech conversion
abstract
Statistical voice conversion is very effective for enhancing body transmitted speech recorded with Non-Audible Murmur (NAM) microphone. In this method, a probabilistic model to convert body transmitted speech into natural speech is trained previously. Because acoustic characteristics of body transmitted speech is sensitive to recording conditions such as a location of NAM microphone, significant degradation of the conversion performance is often caused in practical situations by acoustic mismatches between training and conversion processes. To alleviate this problem, we propose unsupervised acoustic compensation methods for body transmitted voice conversion. Experimental results demonstrate that the proposed methods significantly reduce the quality degradation of converted speech caused by the acoustic mismatches.
Daisuke Miyamoto, Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP4
2009 Hands-free speech recognition challenge for real-world speech dialogue systems
abstract
In this paper, we describe and review our recent development of hands-free speech dialogue system which is used for railway station guidance. In the application at the real railway station, robustness against reverberation and noise is the most essential issue for the dialogue system. To address the problem, we introduce two key techniques in our proposed hands-free system; (a) speech dialogue system construction with real speech database collection and language/acoustic model improvement, and (b) microphone array preprocessing using blind spatial subtraction array which can solve the reverberation-naiveness problem inherent in conventional microphone arrays. The experimental assessment of the proposed dialogue system reveals that our system can provide the recognition accuracy of more than 80% under realistic railway-station conditions.
Hiroshi Saruwatari, Hiromichi Kawanami, Shota Takeuchi, Yu Takahashi, Tobias Cincarek, Kiyohiro Shikano
ICASSP1
2009 Source adaptive blind signal extraction using closed-form ICA for hands-free robot spoken dialogue system
abstract
In this paper, we propose a new ICA-based BSS algorithm including estimation of sources' probability density functions (PDFs) to adapt the nonlinear activation function to various noise conditions. In the proposed method, closed-form second-order ICA is introduced as a computational-cost-efficient preprocessing to extract sources' PDFs, which is beneficial for real-time application. Compared with various type of conventional ICAs, e.g., fixed activation-function type and ML-based type, our proposed algorithm can give a faster and higher convergence. Based on the proposed source-adaptive ICA, we show a real-time noise reduction results under diffuse noise environment. Also we can demonstrate our recently developed hands-free robot spoken dialogue system via real-time ICA.
Yu Takahashi, Hiroshi Saruwatari, Yuki Fujihara, Kentaro Tachibana, Yoshimitsu Mori, Shigeki Miyabe, Kiyohiro Shikano, Akira Tanaka
ICASSP2
2009 Musical noise analysis based on higher order statistics for microphone array and nonlinear signal processing
abstract
In this paper, we conduct an analysis for reduction of musical noise in integration method of microphone array signal processing and nonlinear signal processing. In these days, for better noise reduction, integration methods of microphone array signal processing and nonlinear signal processing have been researched. However, non-linear signal processing causes musical noise. Since such musical noise make users uncomfortable, it is desired that musical noise is mitigated. Moreover, in these days, it is reported that higher-order statistics is strongly related with the amount of generated musical noise. Thus, we analyze the integrated method of microphone array signal processing and nonlinear signal processing, based on higher-order statistics. Also, we propose an architecture for reducing musical noise based on the analysis. The effectiveness of the proposed architecture and analysis correctness are shown via a computer simulation and a subjective evaluation.
Yu Takahashi, Yoshihisa Uemura, Hiroshi Saruwatari, Kiyohiro Shikano, Kazunobu Kondo
ICASSP3
2009 Musical noise generation analysis for noise reduction methods based on spectral subtraction and MMSE STSA estimation
abstract
In this paper, we reveal new findings about the generated musical noise in minimum mean-square error short-time spectral amplitude (MMSE STSA) processing. Recently we have proposed a objective metric of musical noise based on kurtosis change ratio on spectral subtraction (SS). Also we found an interesting relationship among the degree of generated musical noise, the shapes of signal-s probability density function, the strength parameter of SS processing. This paper is aimed to automatically evaluate the sound quality of various types of noise reduction methods using kurtosis change ratio. We give a mathematical analysis based on higher-order statistics viewpoint, and lead to a valuable relation in that MMSE STSA has a weakness in speech period distortion rather than noise period, and vice versa in SS.
Yoshihisa Uemura, Yu Takahashi, Hiroshi Saruwatari, Kiyohiro Shikano, Kazunobu Kondo
ICASSP3
2009 Electrolaryngeal speech enhancement based on statistical voice conversion
abstract
This paper proposes a speaking-aid system for laryngectomees using GMM-based voice conversion that converts electrolaryngeal speech (EL speech) to normal speech. Because valid F0 information cannot be obtained from the EL speech, we have so far converted the EL speech to whispering. This paper conducts the EL speech conversion to normal speech using F0 counters estimated from the spectral information of the EL speech. In this paper, we experimentally evaluate these two types of output speech of our speaking-aid system from several points of view. The experimental results demonstrate that the converted normal speech is preferred to the converted whisper.
Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2009 Many-to-many eigenvoice conversion with reference voice
abstract
In this paper, we propose many-to-many voice conversion (VC) techniques to convert an arbitrary source speaker's voice into an arbitrary target speaker's voice. We have proposed one-to-many eigenvoice conversion (EVC) and many-to-one EVC. In the EVC, an eigenvoice Gaussian mixture model (EV-GMM) is trained in advance using multiple parallel data sets of a reference speaker and many pre-stored speakers. The EV-GMM is flexibly adapted to an arbitrary speaker using a small amount of adaptation data without any linguistic constraints. In this paper, we achieve many-to-many VC by sequentially performing many-to-one EVC and one-to-many EVC through the reference speaker using the same EV-GMM. Experimental results demonstrate the effectiveness of the proposed many-to-many VC.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2009 Semi-blind suppression of internal noise for hands-free robot spoken dialog system
abstract
The speech enhancement architecture presented in this paper is specifically developed for hands-free robot spoken dialog systems. It is designed to take advantage of additional sensors installed inside the robot to record the internal noises. First a modified frequency domain blind signal separation (FD-BSS) gives estimates of the noises generated outside and inside of the robot. Then these noises are canceled from the acquired speech by a multichannel Wiener post-filter. Some experimental results show the recognition improvement for a dictation task in presence of both diffuse background noise and internal noises.
Jani Even, Hiroshi Sawada, Hiroshi Saruwatari, Kiyohiro Shikano, Tomoya Takatani
IROS3
2009 Techniques in rapid unsupervised speaker adaptation based on HMM-Sufficient Statistics
Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
Speech Commun.3
2009 Blind Spatial Subtraction Array for Speech Enhancement in Noisy Environment
abstract
We propose a new blind spatial subtraction array (BSSA) consisting of a noise estimator based on independent component analysis (ICA) for efficient speech enhancement. In this paper, first, we theoretically and experimentally point out that ICA is proficient in noise estimation under a non-point-source noise condition rather than in speech estimation. Therefore, we propose BSSA that utilizes ICA as a noise estimator. In BSSA, speech extraction is achieved by subtracting the power spectrum of noise signals estimated using ICA from the power spectrum of the partly enhanced target speech signal with a delay-and-sum beamformer. This ldquopower-spectrum-domain subtractionrdquo procedure enables better noise reduction than the conventional ICA with estimation-error robustness. Another benefit of BSSA architecture is ldquopermutation robustness". Although the ICA part in BSSA suffers from a source permutation problem, the BSSA architecture can reduce the negative affection when permutation arises. The results of various speech enhancement test reveal that the noise reduction and speech recognition performance of the proposed BSSA are superior to those of conventional methods.
Yu Takahashi, Tomoya Takatani, Keiichi Osako, Hiroshi Saruwatari, Kiyohiro Shikano
IEEE Trans. Speech Audio Process.4
2008 Frequency domain semi-blind signal separation: application to the rejection of internal noises
abstract
Recently, methods using blind signal separation were proposed to separate the signals received by a microphone array. In this paper, we propose a new frequency domain semi-blind source separation method for replacing the blind source separation method when it is possible to obtain additional information on some of the signals. This is of particular interest in situations like in hands-free speech recognition where the blind separation has to work on limited amount of data in a challenging environment. The proposed method incorporates references to some of the signals that are obtained by additional sensors. Some experimental results shows that the proposed method is able to incorporate the additional information efficiently and that the performances are improved in term of SNR and word accuracy in a speech recognition task.
Jani Even, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP2
2008 Distant talking robust speech recognition using late reflection components of room impulse response
abstract
We propose a robust and fast dereverberation technique for real-time speech recognition application. First, we effectively identify the late reflection components of the room impulse response. We use this information together with the concept of Spectral Subtraction (SS) to remove the late reflection components of the reverberant signal. In the absence of the clean speech in actual scenario, approximation is carried out in estimating the late reflection where the estimation error is corrected through multi-band SS. The multi-band coefficients are optimized during offiine training and used in the actual online dereverberation. The proposed method performs better and faster than the relevant approach using Multi-LPC and reverberant matched model. Moreover the proposed method is robust to speaker and microphone locations.
Randy Gomez, Jani Even, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP3
2008 Source-oriented localization control of stereo audio signals based on blind source separation
abstract
We propose methods to analyze and control source localization of stereo audio signals using blind source separation (BSS) based on independent component analysis (ICA). Although an inverse system of separation compensates distortion caused by ICA as reconstruction of stereo spatial characteristics, this technique is insufficient to analyze localization because it achieves compensation of distortion and reconstruction of spatial characteristics simultaneously. Thus we analyze spatial characteristics effectively by dividing the compensation into two steps: monaural-output compensation of distortion and its reconstruction of spatial characteristics. Additionally, we control the localization of each source by modifying the analyzed spatial characteristics. It is shown that the proposed method can be applied to stereo signals consisting of more than two sources.
Yuuki Haraguchi, Shigeki Miyabe, Hiroshi Saruwatari, Kiyohiro Shikano, Toshiyuki Nomura
ICASSP3
2008 Hybrid structure of inverse filtering and DOA-parameterized wavefront synthesis
abstract
Weakness against a user's position shifting is one of the most important problems of binaural reproduction using loudspeakers. In this paper we propose a new method of inverse filtering of room acoustics with high robustness against a user's position shifting by presenting a wavefront estimated from the binaural recording. To analyze and synthesize the wavefront, we introduce a new physical modelling of the superimposition of the wavefronts weighted by multiple sound sources. By following the fluctuation of the wavefront with a time-varying filter, the analysis and synthesis overcome the limitation of the number of sound sources. Utilizing arbitrary components of the generalized inverse matrix, wavefront approximation does not degrade the accuracy of reproduction at the controlled area of the inverse filter.
Yuuta Yuyama, Shigeki Miyabe, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP3
2008 Low-delay voice conversion based on maximum likelihood estimation of spectral parameter trajectory
abstract
As typical voice conversion methods, two spectral conversion processes have been proposed: 1) the frame-based conversion that converts spectral parameters frame by frame and 2) the trajectory-based conversion that converts all spectral parameters over an utterance simultaneously. The former process is capable of real-time conversion but it sometimes causes inappropriate spectral movements. On the other hand, the latter process provides the converted spectral parameters exhibiting proper dynamic characteristics but a batch process is inevitable. To achieve the real-time conversion process considering spectral dynamic characteristics, we propose a time-recursive conversion algorithm based on maximum likelihood estimation of spectral parameter trajectory. Experimental results show that the proposed method achieves the low-delay conversion process, e.g., only one frame delay, while keeping the conversion performance comparably high to that of the conventional trajectory-based conversion. Index Terms: speech synthesis, voice conversion, Gaussian mixture model, maximum likelihood estimation, time-recursive algorithm. 1.
Takashi Muramatsu, Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2008 Evaluation of speaking-aid system with voice conversion for laryngectomees toward its use in practical environments
abstract
In this paper, we evaluate our previously proposed speaking-aid system with voice conversion for laryngectomees. The proposed system employs a sound source unit generating extremely small signals to keep them from annoying other persons, and then it statistically converts articulated signals captured with a bodyattached microphone into natural speech. We have so far shown the effectiveness of the proposed system using speech data imitated by a non-laryngectomee, which have recorded in a sound proof room. In this paper, we further investigate 1) whether such small sound source signals cause the lack of auditory feedback under noisy environments and 2) whether the proposed system is effective for real laryngectomees. Experimental results demonstrate that 1) an explicit auditory feedback is useful to keep the speaker's articulation stable and 2) the voice conversion dramatically improves the naturalness of the laryngectomee's speech but it slightly degrades its intelligibility.
Keigo Nakamura, Tomoki Toda, Yoshitaka Nakajima, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2008 An improved one-to-many eigenvoice conversion system
abstract
We have previously developed a one-to-many eigenvoice conversion (EVC) system enabling the conversion from a specific source speaker's voice into an arbitrary target speaker's voice. In this system, eigenvoice Gaussian mixture model (EV-GMM) is trained in advance with multiple parallel data sets composed of utterance pairs of the source and many pre-stored target speakers. The EV-GMM is effectively adapted to an arbitrary target speaker using a small amount of adaptation data. Although this system achieves the very flexible training of the conversion model, the quality of the converted speech is still not high enough. In order to alleviate this problem, we simultaneously apply the following promising techniques to the one-to-many EVC system: 1) STRAIGHT mixed excitation, 2) the conversion algorithm considering global variance, and 3) speaker adaptive training of the EV-GMM. Experimental results demonstrate that the proposed system causes remarkable improvements in the performance of EVC.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2008 Speaker verification with non-audible murmur segments by combining global alignment kernel and penalized logistic regression machine
abstract
We investigate a novel method for speaker verification with nonaudible murmur (NAM) segments. NAM is recorded using a special microphone placed on the neck and is hard for other people to hear. We have already reported a method based on a support vector machine (SVM) using NAM segments to use a keyword phrase effectively. To further exploit keyword-specific features, we introduce a global alignment (GA) kernel and penalized logistic regression machine (PLRM). In the experiments using NAM from 55 speakers, our method achieved an error reduction rate of roughly 60% compared with the SVM-based method using a polynomial kernel.
Hideki Okamoto, Tomoko Matsui, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2008 Development and evaluation of hands-free spoken dialogue system for railway station guidance
abstract
In this paper, we describe development and evaluation of handsfree spoken dialogue system which is used for railway station guidance. In the application at the railway station, noise robustness is the most essential issue for the dialogue system. To address the problem, we introduce two key techniques in our proposed hands-free system; (a) blind spatial subtraction array (BSSA) as a preprocessing, which can efficiently reduce nonstationary and diffuse noises in real-time, and (b) robust voice activity detection (VAD) based on speech decoding for further improvement of speech recognition accuracy. The experimental assessment of the proposed dialogue system reveals that the combination of real-time BSSA and robust VAD can provide the recognition accuracy of more than 80% under adverse railway-station noise conditions.
Hiroshi Saruwatari, Yu Takahashi, Hiroyuki Sakai 0004, Shota Takeuchi, Tobias Cincarek, Hiromichi Kawanami, Kiyohiro Shikano
INTERSPEECH1
2008 Question and answer database optimization using speech recognition results
abstract
The aim of this research is a human-oriented spoken dialog system which provides replies to a variety of users' utterances. The example-based response generation method searches a question and answer database (QADB) for the example question most similar to a user utterance. With this method, the system can answer a question difficult for a model to express. A QADB is constructed from question and answer pairs (QA pairs) by employing a large corpus. In order to enhance robustness to recognition errors of inarticulate utterances such as children utterances, we propose to use speech recognition results, instead of manual transcriptions, as example questions. We also introduce an optimization method that removes inappropriate QA pairs from a QADB to maximize response accuracy. We show that our method improves the response accuracy of utterances especially for children utterances in the open test.
Shota Takeuchi, Tobias Cincarek, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2008 Maximum a posteriori adaptation for many-to-one eigenvoice conversion
abstract
Many-to-one eigenvoice conversion (EVC) allows the conversion from an arbitrary speaker's voice into the pre-determined target speaker's voice. In this method, a canonical eigenvoice Gaussian mixture model is effectively adapted to any source speaker using only a few utterances as the adaptation data. In this paper, we propose a many-to-one EVC based on maximum a posteriori (MAP) adaptation for further improving the robustness of the adaptation process to the amount of adaptation data. Results of objective and subjective evaluations demonstrate that the proposed method is the most effective among the other conventional many-to-one VC methods when using any amount of adaptation data (e.g., from 300 ms to 16 utterances).
Daisuke Tani, Tomoki Toda, Yamato Ohtani, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2008 An improved permutation solver for blind signal separation based front-ends in robot audition
abstract
The model of the human/machine hands-free speech interface is defined as a point source (the user voice) and a diffuse background noise. This situation is very different from the usual cocktail party model, separation of a mixture of speeches, that is usually treated in frequency domain blind signal separation (FD-BSS). In particular, the fast permutation solvers proposed for the cocktail party model results in poor separation performance in this case. In order to resolve the permutation more efficiently, this paper proposes a new approach that exploits the statistical discrepancy between the target speech and the diffuse background noise.
Jani Even, Hiroshi Saruwatari, Kiyohiro Shikano
IROS2
2008 Real-time implementation of blind spatial subtraction array for hands-free robot spoken dialogue system
abstract
In this paper, we construct a hands-free robot spoken dialogue system based on the real-time blind spatial subtraction array (BSSA) and evaluate the system. BSSA is the blind source extraction method, and the source extraction in BSSA is carried out by subtracting the power spectrum of the estimated noise signal by the independent component analysis from the power spectrum of the target speech partly enhanced signal. Although BSSA can reduce noise signal efficiently, ICA consumes huge amount of computational costs. Thus it is difficult to run BSSA in real-time. In this paper, we newly propose a real-time architecture of BSSA and construct a hands-free robot spoken dialogue system based on the real-time BSSA. In the hands-free robot spoken dialogue system with the real-time BSSA, 6% improvement of the speech recognition result can be seen compared with the conventional speech enhancement methods.
Yu Takahashi, Hiroshi Saruwatari, Kiyohiro Shikano
IROS2
2007 Development and portability of ASR and Q&A modules for real-environment speech-oriented guidance systems
abstract
In this paper, we investigate development and portability of ASR and Q&A modules of speech-oriented guidance systems for two different real environments. An initial prototype system has been constructed for a local community center using two years of human-labeled data collected by the system. Collection of real user data is required because ASR task and Q&A domain of a guidance system are defined by the target environment and potential users. However, since human preparation of data is always costly, most often only a relatively small amount real data will be available for system adaptation in practice. Therefore, the portability of the initial prototype system is investigated for a different environment, a local subway station. The purpose is to identify reusable system parts. The ASR module is found to be highly portable across the two environments. However, the portability of the Q&A module was only medium. From an objective analysis it became clear that this is mainly due to the environment-dependent domain differences between the two systems. This implicates that it will always be important to take the behavior of actual users under real conditions into account to build a system with high user satisfaction.
Tobias Cincarek, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
ASRU3
2007 High-Presence Hearing-Aid System using DSP-Based Real-Time Blind Source Separation Module
abstract
Real-time two-stage blind source separation (BSS) method for convolutive mixtures of speech is now being studied by the authors, in which a single-input multiple-output (SIMO)-model-based independent component analysis (ICA) and a SIMO-model-based binary masking are combined. In addition, we have developed a pocket-size real-time DSP module implementing the two-stage BSS method. In this paper, we introduce a high-presence hearing-aid system which can reduce the interference sound and reproduce the target sound while keeping the directivity, and realize the system with the real-time BSS module. To evaluate it, we carried out the objective and subjective experiments using 9 users. From these results, it is revealed that the decomposition performance and the directivity maintenance of the proposed system are superior to those of conventional methods.
Yoshimitsu Mori, Tomoya Takatani, Hiroshi Saruwatari, Kiyohiro Shikano, Takashi Hiekata, Takashi Morita 0003
ICASSP (4)3
2007 Efficient Blind Source Separation Combining Closed-Form Second-Order ICA and Nonclosed-Form Higher-Order ICA
abstract
In this paper, first, we propose a computational-cost efficient blind source separation combining closed-form 2nd-order independent component analysis (ICA) and nonclosed-form higher-order ICA. The closed-form solution of the 2nd-order ICA has been recently presented by one of the authors. This finding motivates us to combine the closed-form 2nd-order ICA and higher-order ICA, where the preceding closed-form ICA produces a good initial value and the following higher-order ICA updates the separation filters from the advantageous status. Secondly, we utilize the proposed architecture to address an essential question that which type of statistics is more beneficial to ICA among non-stationarity and non-Gaussianity. This can be conducted owing to the attractive property that the closed-form ICA can provide a good estimate of the theoretical upper limitation of the separation performance among 2nd-order ICAs without suffering from poor-convergence problems. Experimental results reveal that the non-Gaussianity-based ICA can outperform the non-stationarity-based ICA.
Kentaro Tachibana, Hiroshi Saruwatari, Yoshimitsu Mori, Shigeki Miyabe, Kiyohiro Shikano, Akira Tanaka
ICASSP (1)2
2007 Permutation-Robust Structure for ICA-Based Blind Source Extraction
abstract
In this paper, we investigate a new blind source separation (BSS) structure from a permutation-robustness viewpoint, to mitigate the permutation problem which commonly arises in frequency-domain independent component analysis (ICA). Permutation robustness means that how much the BSS method is not affected under a certain probability of arising permutation, unlike the conventional permutation-solving approaches. We address to analyze our previously proposed BSS architecture, so called blind spatial subtraction array (BSSA). In BSSA, source extraction is achieved by subtracting the power spectrum of the estimated noise via ICA from the power spectrum of partly speech-enhanced signal via delay-and-sum (DS) procedure. Indeed BSSA partially involves permutation problem in the ICA-based noise estimator part. However, BSSA can efficiently reduce the negative affection of the permutation owing to the over-subtraction in the spectral subtraction and defocusing properties in DS. Experiments using artificial and real-recording-based simulations reveal that the proposed method outperforms the conventional ICA.
Yu Takahashi, Tomoya Takatani, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP (1)3
2007 Development of preschool children subsystem for ASR and q&a in a real-environment speech-oriented guidance task
abstract
INTERSPEECH2007: 8th Annual Conference of the International Speech Communication Association, August 27-31, 2007, Antwerp, Belgium.
Tobias Cincarek, Izumi Shindo, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2007 Rapid unsupervised speaker adaptation using single utterance based on MLLR and speaker selection
abstract
INTERSPEECH2007: 8th Annual Conference of the International Speech Communication Association, August 27-31, 2007, Antwerp, Belgium.
Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2007 Impact of various small sound source signals on voice conversion accuracy in speech communication aid for laryngectomees
abstract
We proposed a speaking aid system using statistical voice conversion for laryngectomees, whose vocal folds have been removed. This paper investigates the influence of various small sound sources on the voice conversion accuracy. Spectral envelopes and power of sound sources are controlled independently. In total 8 different kinds of sound source signals, e.g. pulse train, sierra wave and so on, are examined. Results of objective and subjective evaluations demonstrate that for voice conversion, sound sources with various spectral envelopes and power in a large degree are acceptable unless the power of them is comparable to that of silence parts.
Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2007 Speaker adaptive training for one-to-many eigenvoice conversion based on Gaussian mixture model
abstract
One-to-many eigenvoice conversion (EVC) allows the conversion of a specific source speaker into arbitrary target speakers. Eigenvoice Gaussian mixture model (EV-GMM) is trained in advance with multiple parallel data sets consisting of the source speaker and many pre-stored target speakers. The EV-GMM is adapted for arbitrary target speakers using only a few utterances by estimating a small number of free parameters. Therefore, the initial EV-GMM directly affects the conversion performance of the adapted EV-GMM. In order to prepare a better initial model, this paper proposes Speaker Adaptive Training (SAT) of a canonical EV-GMM in one-to-many EVC. Results of objective and subjective evaluations demonstrate that SAT causes significant improvements in the performance of EVC.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2007 Study on speaker verification with non-audible murmur segments
abstract
We investigated a speaker verification method that uses non-audible murmur (NAM) segments using newly collected data and obtained several findings that will be useful when speaker verification systems are made in practice. NAM is recorded using a special microphone placed on the surface of the body, so it includes almost no external noise and is hard for other people to hear. By utilizing these properties, we have already reported a text-dependent method using NAM segments that can use a keyword phrase safely. This paper extends the examination with newly collected data consisting of NAM uttered by 18 male and 9 female imposter speakers and by 18 male and 10 female customer speakers. Experiments with various numbers of training utterances and sessions show that it is effective to use data recorded in multiple sessions. We also investigated the minimum number of training utterances needed in our method.
Hideki Okamoto, Mariko Kojima, Tomoko Matsui, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH5
2006 Improving Rapid Unsupervised Speaker Adaptation Based On Hmm Sufficient Statistics
abstract
In real-time speech recognition applications, there is a need to implement a fast and reliable adaptation algorithm. We propose a method to reduce adaptation time of the unsupervised speaker adaptation based on HMM-sufficient statistics. We use only a single arbitrary utterance without transcriptions in selecting the N-best speakers' sufficient statistics created offline to provide data for adaptation to a target speaker. Further reduction of N-best implies a reduction in adaptation time. However, it degrades recognition performance due to insufficiency of data needed to robustly adapt the model. Linear interpolation of the global HMM-sufficient statistics offsets this negative effect and achieves a 50% reduction in adaptation time without compromising the recognition performance. We have reduced the adaptation time from 10 sec to 5 sec without degradation of the word accuracy. Furthermore, we compared our method with vocal tract length normalization (VTLN), maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR). Moreover, we tested in office, car, crowd and booth noise environments in 10 dB, 15 dB, 20 dB and 25 dB SNRs
Randy Gomez, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP (1)3
2006 Double-Talk Free Spoken Dialogue Interface Combining Sound Field Control With Semi-Blind Source Separation
abstract
In this paper we introduce a new double-talk free spoken dialogue interface combining sound field control and a source separation technique based on independent component analysis (ICA). First, sound field control provides silent zones on the microphone elements and prevents the response sound from being observed. In the second step, we propose a novel semi-blind source separation algorithm to suppress the error caused by fluctuation of the room transfer function. By using a direct input of response sound signal to ICA, a source separation problem can be converted to a supervised learning problem. Since the problem becomes easier, the proposed method showed higher performances than the method using blind source separation
Shigeki Miyabe, Tomoya Takatani, Yoshimitsu Mori, Hiroshi Saruwatari, Kiyohiro Shikano, Yosuke Tatekura
ICASSP (1)4
2006 Blind Source Separation Combining Simo-Ica and Simo-Model-Based Binary Masking
abstract
A new two-stage blind source separation (BSS) for convolutive mixtures of speech is proposed, in which a single-input multiple-output (SIMO)-model-based ICA and a new SIMO-model-based binary mask processing are combined. SIMO-model-based ICA can separate the mixed signals, not into monaural source signals but into SIMO-model-based signals from independent sources as they are at the microphones. Thus, the separated signals of SIMO-model-based ICA can maintain the spatial qualities of each sound source. Owing to the attractive property, novel SIMO-model-based binary mask processing can be applied to efficiently remove the residual interference components after SIMO-model-based ICA. The experimental results reveal that the separation performance can be considerably improved by using the proposed method compared with the conventional BSS methods
Yoshimitsu Mori, Tomoya Takatani, Hiroshi Saruwatari, Takashi Hiekata, Takashi Morita 0003
ICASSP (5)3
2006 Acoustic modeling for spoken dialogue systems based on unsupervised utterance-based selective training
abstract
INTERSPEECH2006: the 9th International Conference on Spoken Language Processing (ICSLP), September 17-21, 2006, Pittsburgh, Pennsylvania, USA.
Tobias Cincarek, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2006 Speaker verification with non-audible murmur segments
abstract
We propose a speaker verification method using non-audible murmur (NAM) segments, which are different from normal speech and hard for other people to catch them. To use NAM, we therefore take a text-dependent verification strategy in which each user utters her/his own keyword phrase and utilize not only speaker-specific but also keyword-specific acoustic information. We expect this strategy to yield a relatively high performance. NAM segments, which consist of multiple short-term feature vectors, are used as input vectors to capture keyword-specific acoustic information well. To handle segments with a large number of dimensions, we use the support vector machine (SVM). In experiments using NAM data of 19 male and 10 female speakers recorded in three different sessions, we achieved equal error rates of 0.04% (male) and 1.1% (female) when using 145-ms-long NAM segments. These rates are half or less those obtained with 25-ms-long input vectors.
Mariko Kojima, Tomoko Matsui, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2006 Speaking aid system for total laryngectomees using voice conversion of body transmitted artificial speech
abstract
The aim of this paper is to improve the naturalness of speech using a medical device such as an electrolarynx. There are several problems associated with using existing electrolarynxes, such as the fact the loud volume of the electrolarynx itself might disturb smooth interpersonal communication, and that the generated speech is unnatural. We propose a novel speaking-aid system for total laryngectomees using a new sound source as an alternative to the existing electrolarynx and a statistical voice-conversion technique. The new sound-source unit outputs extremely small signals that cannot be heard by people around the speaker. Artificial speech is recorded with a NAM microphone through soft tissues of the head. From the result of voice conversion, the body-transmitted artificial speech is consistently converted to a more natural voice. We also demonstrate that the speech recognition performance of the proposed system substantially increases in terms of objective evaluation.
Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2006 Maximum likelihood voice conversion based on GMM with STRAIGHT mixed excitation
abstract
The performance of voice conversion has been considerably improved through statistical modeling of spectral sequences. However, the converted speech still contains traces of artificial sounds. To alleviate this, it is necessary to statistically model a source sequence as well as a spectral sequence. In this paper, we introduce STRAIGHT mixed excitation to a framework of the voice conversion based on a Gaussian Mixture Model (GMM) on joint probability density of source and target features. We convert both spectral and source feature sequences based on Maximum Likelihood Estimation (MLE). Objective and subjective evaluation results demonstrate that the proposed source conversion produces strong improvements in both the converted speech quality and the conversion accuracy for speaker individuality.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2006 Transcription Cost Reduction for Constructing Acoustic Models Using Acoustic Likelihood Selection Criteria
Tomoyuki Kato, Tomiki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
LREC3
2006 Blind source separation based on a fast-convergence algorithm combining ICA and beamforming
abstract
We propose a new algorithm for blind source separation (BSS), in which independent component analysis (ICA) and beamforming are combined to resolve the slow-convergence problem through optimization in ICA. The proposed method consists of the following three parts: (a) frequency-domain ICA with direction-of-arrival (DOA) estimation, (b) null beamforming based on the estimated DOA, and (c) integration of (a) and (b) based on the algorithm diversity in both iteration and frequency domain. The unmixing matrix obtained by ICA is temporally substituted by the matrix based on null beamforming through iterative optimization, and the temporal alternation between ICA and beamforming can realize fast- and high-convergence optimization. The results of the signal separation experiments reveal that the signal separation performance of the proposed algorithm is superior to that of the conventional ICA-based BSS method, even under reverberant conditions.
Hiroshi Saruwatari, Toshiya Kawamura, Tsuyoki Nishikawa, Akinobu Lee, Kiyohiro Shikano
IEEE Trans. Speech Audio Process.1
2005 Blind source separation combining SIMO-model-based ICA and adaptive beamforming
abstract
A new two-stage blind source separation (BSS) for convolutive mixtures of speech is proposed, in which a single-input multiple-output-model-based ICA (SIMO-ICA) and an adaptive beamforming (ABF) are combined. SIMO-ICA can separate the mixed signals, not into monaural source signals but into SIMO model-based signals from independent sources as they are at the microphones. Thus, the separated signals of SIMO-ICA can maintain the spatial qualities of each sound source, and the directions-of-arrival (DOAs) of the sources can be estimated after the separation by SIMO-ICA. Owing to the attractive property, the supervised ABF can be applied to removing the residual interference components efficiently after the SIMO-ICA and DOA estimation procedures. Experimental results reveal that separation performance can be considerably improved by using the proposed method. In addition, the proposed method outperforms the combination of the conventional SIMO-output-type ICA and ABF, as well as both the simple ICA and the simple ABF.
Satoshi Ukai, Tomoya Takatani, Tsuyoki Nishikawa, Hiroshi Saruwatari
ICASSP (3)4
2005 Rapid unsupervised speaker adaptation based on multi-template HMM sufficient statistics in noisy environments
abstract
INTERSPEECH2005: the 9th European Conference on Speech Communication and technology, September 4-8, 2005, Lisbon, Portugal.
Randy Gomez, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2005 Investigating the role of the Lombard reflex in non-audible murmur (NAM) recognition
abstract
In this paper, we report non-audible murmur (NAM) recognition results in noisy environments and investigate the effect of the Lombard reflex on non-audible murmur recognition. Non-Audible murmur is speech uttered very quietly and captured through body tissue by a special acoustic sensor (e.g., NAM microphone). A system based on non-audible murmur recognition can be applied in cases when privacy is preferable in human-machine communication. Moreover, due to direct body-transmission, the environmental noises do not affect the performance markedly. Previously, we reported non-audible murmur automatic recognition in a clean environment with very promising results. We also carried out experiments using clean models and simulated noisy data, showing that the performance did not change significantly. Using, however, real noisy test data, the performance decreased markedly. To investigate this problem, we studied the Lombard reflex and conducted non-audible murmur recognition experiments using Lombard data. Results show, that Lombard reflex affects non-audible murmur recognition.
Panikos Heracleous, Tomomi Kaino, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2005 Applications of NAM microphones in speech recognition for privacy in human-machine communication
abstract
In this paper, we present the use of stethoscope and silicon NAM microphones in automatic speech recognition. NAM microphones are special acoustic sensors, which are attached behind the talker's ear and can capture not only normal (audible) speech, but also very quietly uttered speech (non-audible murmur). As a result, NAM microphones can be applied in automatic speech recognition systems when privacy is desired. Previously, we presented speech recognition experiments for non-audible murmur captured by a stethoscope microphone. In this paper, we also present recognition results using a more advanced NAM microphone, the so-called silicon NAM microphone. Using adaptation techniques and a small amount of training data, we achieved a 93.9% word accuracy for non-audible murmur recognition. We also report experimental results in noisy environments showing the effectiveness of using a NAM microphone in noisy environments. In addition to a dictation task, we also present a keyword spotting experiment based on non-audible murmur.
Panikos Heracleous, Tomomi Kaino, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2005 Speech extraction in a car interior using frequency-domain ICA with rapid filter adaptations
abstract
This paper describes two new algorithms for blind source separation (BSS) based on frequency-domain independent component analysis (FDICA). One is FDICA with pre-filtering by a speech sub-band passing filter to slow down the learning speed in low signal-to-noise ratio (SNR) sub-bands. The other is FDICA with sub-band selection learning to reduce the number of iterations for those sub-bands. The results of speech recognition experiments show that each method can improve word accuracy by as much as 7% and that the second method can increase the speed by approximately 60%.
Daisuke Saitoh, Atsunobu Kaminuma, Hiroshi Saruwatari, Tsuyoki Nishikawa, Akinobu Lee
INTERSPEECH3
2005 Noise-robust hands-free speech recognition based on spatial subtraction array and known noise superimposition
abstract
We propose a spatial subtraction array (SSA) and known noise superimposition to achieve a noise-robust hands-free speech recognition which can be used in human-robot interaction. In the proposed SSA, noise reduction is achieved by subtracting the estimated noise power spectrum from the target speech power spectrum to be enhanced in the mel-scale filter bank domain. This offers a realization of error-robust spatial spectral subtraction with few computational complexities. In addition, we introduce known noise superimposition technique in the mel-scale filter bank domain, and utilize the matched acoustic model for the known noise. This can compensate the acoustic model mismatch and mask the residual noise component in SSA. The experimental results obtained under a real environment reveal that word accuracy of the proposed method is greater than that of the conventional method even when the target user moves between -10 and +10 degrees around the microphone array.
Yasuaki Ohashi, Tsuyoki Nishikawa, Hiroshi Saruwatari, Akinobu Lee, Kiyohiro Shikano
IROS3
2005 Two-stage blind source separation based on ICA and binary masking for real-time robot audition system
abstract
We newly propose a real-time two-stage blind source separation (BSS) for binaural mixed signals observed at the ears of humanoid robot, in which a single-input multiple-output (SIMO)-model-based independent component analysis (ICA) and binary mask processing are combined. SIMO-model-based ICA can separate the mixed signals, not into monaural source signals but into SIMO-model-based signals from independent sources as they are at the microphones. Thus, the separated signals of SIMO-model-based ICA can maintain the spatial qualities of each sound source, and this yields that binary mask processing can be applied to efficiently remove the residual interference components after SIMO-model-based ICA. The experimental results obtained with a human-like head reveal that the separation performance can be considerably improved by using the proposed method in comparison to the conventional ICA-based and binary-mask-based BSS methods.
Hiroshi Saruwatari, Yoshimitsu Mori, Tomoya Takatani, Satoshi Ukai, Kiyohiro Shikano, Takashi Hiekata, Takashi Morita 0003
IROS1
2005 Blind sound scene decomposition for robot audition using SIMO-model-based ICA
abstract
In this paper, we address a blind decomposition problem of binaural mixed signals observed at the ears of humanoid robot, and we introduce a novel blind signal decomposition algorithm using single-input multiple-output-model-based ICA (SIMO-ICA). The SIMO-ICA consists of multiple ICAs and a fidelity controller, and each ICA runs in parallel under the fidelity control of the entire separation system. The SIMO-ICA can separate the mixed signals, not into monaural source signals but into SIMO-model-based signals from independent sources as they are at the microphones in the robot ear. Thus, the separated signals of SIMO-ICA can maintain the spatial qualities of each sound source, i.e., they represent the decomposed sound scenes. Obviously the attractive feature of SIMO-ICA is highly applicable to not only speech recognition but also, e.g., humanoid-robot-based auditory tele-existence technology. The experimental results reveal that the spatial quality of the separated sound in SIMO-ICA is remarkably superior to that of the conventional method, particularly for the fidelity of the sound reproduction.
Tomoya Takatani, Satoshi Ukai, Tsuyoki Nishikawa, Hiroshi Saruwatari, Kiyohiro Shikano
IROS4
2005 Estimation of Shape Parameter of GGD Function by Negentropy Matching
Rajkishore Prasad, Hiroshi Saruwatari, Kiyohiro Shikano
Neural Process. Lett.2
2004 Overdetermined blind separation for convolutive mixtures of speech based on multistage ICA using subarray processing
abstract
We propose a new algorithm for overdetermined blind source separation based on multistage independent component analysis (MSICA). To improve the separation performance, we have proposed MSICA in which frequency-domain ICA and time-domain ICA are cascaded. In the original MSICA, the specific mixing model, where the number of microphones is equal to that of sources, was assumed. However, additional microphones are required to achieve an improved separation performance under reverberant environments. This leads to alternative problems, e.g., a complication of the permutation problem. In order to solve them, we propose a new extended MSICA using subarray processing, where the number of microphones and that of sources are set to be the same in every subarray. The experimental results obtained under the real environment reveal that the separation performance of the proposed MSICA is improved as the number of microphones is increased.
Tsuyoki Nishikawa, Hiroshi Abe, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP (1)3
2004 Public speech-oriented guidance system with adult and child discrimination capability
abstract
The Takemaru-kun system is a real world speech-oriented guidance system located at the Ikoma-City North Community Center. The system has been operated daily from November, 2002, to provide visitors a speech interface for information retrieval. This system also aims at the field test of a speech interface and collecting actual utterance data. By analyzing and evaluating the collected utterances, the flexible processing requirements are discovered according to the user's age group. It becomes impossible to disregard the increase of child users when the system is installed in a public place. The paper proposes an automatic approach discriminating speakers between adult and child users, which is based on statistical learning. This proposal realizes a flexible spoken dialogue to both adult and child users. As for parameter vectors in machine learning, acoustic and linguistic properties extracted from speech recognition logarithm likelihood scores are adopted to discriminate a user's age group. Although GMM-based recognition uses only acoustic properties, this method can also consider linguistic properties. In experiments with SVM-based screening, we obtained a 92.4% discrimination rate to the actual users' utterances. The advantage of using linguistic properties is also shown. The paper also describes an overview of the Takemaru-kun system and the data collection status from the field test. Child speech recognition performance is evaluated using the collected utterances.
Ryuichi Nisimura, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP (1)3
2004 Blind separation of binaural sound mixtures using SIMO-model-based independent component analysis
abstract
High-fidelity blind audio signal separation is addressed, adopting the extended ICA algorithm, single-input multiple-output (SIMO)-model-based ICA. The SIMO-ICA consists of multiple ICA parts and a fidelity controller, and each ICA runs in parallel under fidelity control of the entire separation system. SIMO-ICA can separate the mixed signals, not into monaural source signals, but into SIMO-model-based signals from independent sources as they are at the microphones. Thus, the separated signals of the SIMO-ICA can maintain the spatial qualities of each sound source. We apply the SIMO-ICA to the problem of blind separation of mixed binaural sounds, including the effect of the head-related transfer function (HRTF). Experimental results reveal that the performance of the proposed SIMO-ICA is superior to that of the conventional ICA-based method, and the separated signals of SIMO-ICA maintain the spatial qualities of each sound source.
Tomoya Takatani, Tsuyoki Nishikawa, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP (4)3
2004 Multistage SIMO-model-based blind source separation combining frequency-domain ICA and time-domain ICA
abstract
In this paper, single-input multiple-output (SIMO)-model-based blind source separation (BSS) is addressed, where unknown mixed source signals are detected at the microphones, and these signals can be separated, not into monaural source signals but into SIMO-model-based signals from independent sources as they are at the microphones. This technique is highly applicable to high-fidelity signal processing such as binaural signal processing. First, we provide an experimental comparison between two kinds of the SIMO-model-based BSS methods, namely, traditional frequency-domain ICA with projection-back processing (FDICA-PB), and SIMO-ICA recently proposed by the authors. Secondly, we propose a new combination technique of the FDICA-PB and SIMO-ICA, which can achieve a higher separation performance in comparison to two methods. The experimental results reveal that the accuracy of the separated SIMO signals in the simple SIMO-ICA is inferior to that of FDICA-PB, but the proposed combination technique can outperform both simple FDICA-PB and SIMO-ICA.
Satoshi Ukai, Hiroshi Saruwatari, Tomoya Takatani, Ryo Mukai, Hiroshi Sawada
ICASSP (4)2
2004 Interface for barge-in free spoken dialogue system using adaptive sound field control
abstract
This paper describes a new interface for a barge-in free spoken dialogue system combining an adaptive sound field control and a microphone array. In order to actualize robustness against the change of transfer functions due to the various interferences, the barge-in free spoken dialogue system which uses sound field control and a microphone array has been proposed by one of the authors. However, this method cannot follow the large change of transfer functions. To solve the problem, we introduce a new adaptive sound field control that follows the change of transfer functions. The experimental results reveal that the proposed method can improve the reduction accuracy of response sound in comparison with the conventional acoustic echo canceller as well as the previously proposed method which simply uses fixed sound field control system.
Tatsunori Asai, Shigeki Miyabe, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2004 Robust speech recognition with spectral subtraction in low SNR
abstract
ICSLP2004: the 8th International Conference on Spoken Language Processing, October 4-8, 2004, Jeju Island, Korea.
Randy Gomez, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2004 Non-audible murmur (NAM) speech recognition using a stethoscopic NAM microphone
abstract
In this paper, we introduce the Stethoscopic Non-Audible Murmur (NAM) microphone, and we focus on its application in automatic speech recognition systems. The NAM microphone is attached behind the talker's ear, and can capture very quietly uttered murmur (NAM speech). It is applicable in automatic speech recognition systems, when privacy is important in human-machine communication. Moreover, since the NAM microphone receives the speech signal directly from the body, it shows robustness against the environmental noises. In addition to these, it might be also used in special systems (speech recognition, speech transform, etc.) for sound-impaired people. By applying adaptation techniques, we performed automatic speech recognition experiments for NAM speech. Using Maximum A Posteriori (MAP) adaptation, and a combination with Maximum Likelihood Linear Regression (MLLR) adaptation we achieved for a 20k vocabulary dictation system a 93.5% word accuracy, which is a very promising result.
Panikos Heracleous, Yoshitaka Nakajima, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2004 Noise robust real world spoken dialogue system using GMM based rejection of unintended inputs
abstract
ICSLP2004: the 8th International Conference on Spoken Language Processing, October 4-8, 2004, Jeju Island, Korea.
Akinobu Lee, Keisuke Nakamura, Ryuichi Nisimura, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2004 Perceptual Evaluation of Quality Deterioration Owing to Prosody Modification
Kazuki Adachi, Tomoki Toda, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
LREC4
2003 Subband based blind source separation for convolutive mixtures of speech
abstract
Subband processing is applied to blind source separation (BSS) for convolutive mixtures of speech. This is motivated by the drawback of frequency-domain BSS, i.e., when a long frame with a fixed frame-shift is used to cover reverberation, the number of samples in each frequency decreases and the separation performance is degraded. In our proposed subband BSS, (1) by using a moderate number of subbands, a sufficient number of samples can be held in each subband, mid (2) by using FIR filters in each subband, we can handle long reverberation. Subband BSS achieves better performance than frequency-domain BSS. Moreover, we propose efficient separation procedures that take into consideration the frequency characteristics of room reverberation and speech signals. We achieve this (3) by using longer unmixing filters in low frequency bands, and (4) by adopting overlap-blockshift in BSS's batch adaptation in low frequency bands. Consequently, frequency-dependent subband processing is successfully realized in the proposed subband BSS.
Shoko Araki, Shoji Makino, Robert Aichner, Tsuyoki Nishikawa, Hiroshi Saruwatari
ICASSP (5)5
2003 Interface for barge-in free spoken dialogue system based on sound field control and microphone array
abstract
In this paper, a barge-in free spoken dialogue system using sound field control and microphone array is proposed. In the conventional spoken dialogue system using an acoustic echo canceller, it is indispensable to estimate and update the room transfer function, especially when the transfer function is changed by various interferences. However the estimation process for the transfer function prevents the user from speaking freely and simultaneously with speech responses from the system. In order to resolve the problem, we have already proposed a barge-in free spoken dialogue system that controls a sound field using multiple loudspeakers. In this paper, a microphone array for acquisition of user's speech is newly introduced in the previously proposed system. By introducing the microphone array, we can reduce the number of loudspeakers to be required in the system, and make the interface for spoken dialogue system more robust against the change of room transfer functions.
Yoichi Hinamoto, Kouichi Mino, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP (5)3
2003 Blind source separation based on binaural ICA
abstract
We newly propose a novel blind separation framework for binaural acoustic signals based on the extended ICA algorithm, binaural ICA (BICA). The BICA consists of multiple ICA and fidelity controller, and each ICA runs in parallel under the control of the fidelity of the whole separation system. The BICA can separate the mixed signals into not monaural source signals but binaurally heard signals of independent sources. Thus, the separated signals of BICA can maintain spatial qualities of each sound source. In order to evaluate its effectiveness, separation experiments are carried out under a reverberant condition. The experimental results reveal that (1) the signal separation performance of the proposed BICA is the same as that of the conventional ICA-based method; and (2) the spatial quality of the separated sound in BICA is remarkably superior to that of the conventional method, especially for the fidelity of the sound reproduction.
Tomoya Takatani, Tsuyoki Nishikawa, Hiroshi Saruwatari
ICASSP (5)3
2003 Parallel structured independent component analysis for SIMO-model-based blind separation and deconvolution of convolutive speech mixture
abstract
We propose a two-stage blind separation and deconvolution (BSD) algorithm for a convolutive mixture of temporally correlated signals, in which a new single-input multiple-output (SIMO)-model-based ICA (SIMO-ICA) and blind multichannel inverse filtering are combined. SIMO-ICA consists of multiple ICAs and a fidelity controller, and each ICA runs in parallel under fidelity control of the entire separation system. SIMO-ICA can separate the mixed signals, not into monaural source signals but into SIMO-model-based signals from independent sources as they are at the microphones. After the separation by SIMO-ICA, a simple blind deconvolution technique based on multichannel inverse filtering for the SIMO model can be applied even when the mixing system is the nonminimum phase system and each source signal is temporally correlated. The experimental results obtained under the reverberant condition reveal that the sound quality of the separated signals in the proposed method is superior to that in the conventional ICA-based BSD.
Hiroshi Saruwatari, Hiroaki Yamajo, Tomoya Takatani, Tsuyoki Nishikawa, Kiyohiro Shikano
IJCNN1
2003 GMM-based voice conversion applied to emotional speech synthesis
abstract
Voice conversion method is applied to synthesizing emotional speech from standard reading (neutral) speech. Pairs of neutral speech and emotional speech are used for conversion rule training. The conversion adopts GMM (Gaussian Mixture Model) with DFW (Dynamic Frequency Warping). We also adopt STRAIGHT, the high-quality speech analysis-synthesis algorithm. As conversion target emotions, (Hot) anger, (cold) sadness and (hot) happiness are used. The converted speech is evaluated objectively first using mel cepstrum distortion as a criterion. The result confirms the GMM-based voice conversion can reduce distortion between target speech and neutral speech. A subjective test is also carried out to investigate perceptual effect. From the viewpoint of influence of prosody, two kinds of prosody are used to synthesis. One is natural prosody extracted from neutral speech and the other is from emotional speech. The result shows that prosody mainly contribute to emotion and spectrum conversion can reinforce it. 1.
Hiromichi Kawanami, Yohei Iwami, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2003 Simple designing methods of corpus-based visual speech synthesis
abstract
EUROSPEECH2003: 8th European Conference on Speech Communication and Technology, September 1-4, 2003, Geneva, Switzerland.
Tatsuya Shiraishi, Tomoki Toda, Hiromichi Kawanami, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH4
2003 Unsupervised speaker adaptation based on HMM sufficient statistics in various noisy environments
abstract
Noise and speaker adaptation techniques are essential to realize robust speech recognition in noisy environments. In this paper, first, a noise robust speech recognition algorithm is implemented by superimposing a small quantity of noise data on spectral subtracted input speech. According to the recognition experiments, 30dB SNR noise superimposition on input speech after spectral subtraction increases the robustness against different noises significantly. Next, we apply this noise robust speech recognition to the unsupervised speaker adaptation algorithm based on HMM sufficient statistics in different noise environments. The HMM sufficient statistics for each speaker are calculated from 25dB SNR office noise added speech database beforehand. We evaluate successfully our proposed unsupervised speaker adaptation algorithm in noisy environments with 20k dictation task using 11 kinds of different noises, including office, car, exhibition, and crowd noises.
Shingo Yamade, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2003 Blind separation and deconvolution for convolutive mixture of speech using SIMO-model-based ICA and multichannel inverse filtering
abstract
We propose a new two-stage blind separation and deconvolution (BSD) algorithm for a convolutive mixture of speech, in which a new Single-Input Multiple-Output (SIMO)-model-based ICA (SIMO-ICA) and blind multichannel inverse filtering are combined. SIMO-ICA can separate the mixed signals, not into monaural source signals but into SIMO-model-based signals from independent sources as they are at the microphones. After SIMO-ICA, a simple blind deconvolution technique for the SIMO model can be applied even when each source signal is temporally correlated. The simulation results reveal that the proposed method can successfully achieve the separation and deconvolution for a convolutive mixture of speech.
Hiroaki Yamajo, Hiroshi Saruwatari, Tomoya Takatani, Tsuyoki Nishikawa, Kiyohiro Shikano
INTERSPEECH2
2003 The fundamental limitation of frequency domain blind source separation for convolutive mixtures of speech
abstract
Despite several recent proposals to achieve blind source separation (BSS) for realistic acoustic signals, the separation performance is still not good enough. In particular, when the impulse responses are long, performance is highly limited. In this paper, we consider a two-input, two-output convolutive BSS problem. First, we show that it is not good to be constrained by the condition T>P, where T is the frame length of the DFT and P is the length of the room impulse responses. We show that there is an optimum frame size that is determined by the trade-off between maintaining the number of samples in each frequency bin to estimate statistics and covering the whole reverberation. We also clarify the reason for the poor performance of BSS in long reverberant environments, highlighting that the framework of BSS works as two sets of frequency-domain adaptive beamformers. Although BSS can reduce reverberant sounds to some extent like adaptive beamformers, they mainly remove the sounds from the jammer direction. This is the reason for the difficulty of BSS in reverberant environments.
Shoko Araki, Ryo Mukai, Shoji Makino, Tsuyoki Nishikawa, Hiroshi Saruwatari
IEEE Trans. Speech Audio Process.5
2002 Equivalence between frequency domain blind source separation and frequency domain adaptive beamforming
abstract
Frequency domain Blind Source Separation (BSS) is shown to be equivalent to two sets of frequency domain adaptive microphone arrays, i.e., Adaptive Beamformers (ABFs). The minimization of the off-diagonal components in the BSS update equation can be viewed as the minimization of the mean square error in the ABF. The unmixing matrix of the BSS and the filter coefficients of the ABF converge to the same solution in the mean square error sense if the two source signals are ideally independent. Therefore, the performance of the BSS is limited by that of the ABF. This understanding. gives an interpretation of BSS from physical point of view.
Shoko Araki, Yoichi Hinamoto, Shoji Makino, Tsuyoki Nishikawa, Ryo Mukai, Hiroshi Saruwatari
ICASSP6
2002 Bund source separation based on Multi-Stage ICA combining frequency-domain ICA and time-domain ICA
abstract
We propose a new algorithm for blind source separation (BSS), in which frequency-domain independent component analysis (FDICA) and time-domain ICA (TDICA) are combined to achieve a superior source-separation performance under reverberant conditions. Generally speaking, the conventional TDICA fails to separate source signals under heavily reverberant conditions because of the low convergence in the iterative learning of the inverse of the mixing system. On the other hand, the separation performance of the conventional FDICA under reverberant conditions also degrades significantly because the independence assumption of narrowband signals collapses when the number of subbands increases. In the proposed method, the separated signals of FDICA are regarded as the input signals for TDICA, and we can remove the residual cross-talk components of FDICA by using TDICA. The experimental results under the reverberant condition reveal that the signal-separation performance of the proposed method is superior to that of the conventional ICA-based BSS methods.
Tsuyoki Nishikawa, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP2
2002 Blind source separation based on fast-convergence algorithm using ICA and beamforming for real convolutive mixture
abstract
We propose a new algorithm for blind source separation (BSS), in which independent component analysis (ICA) and beamforming are combined to resolve the low-convergence problem through optimization in ICA. The proposed method consists of the following three parts: (1) frequency-domain ICA with direction-of-arrival (DOA) estimation, (2) null beamforming based on the estimated DOA, and (3) integration of (1) and (2) based on the algorithm diversity in both iteration and frequency domain. The inverse of the mixing matrix obtained by ICA is temporally substituted by the matrix based on null beamforming through iterative optimization, and the temporal alternation between ICA and beamforming can realize fast- and high-convergence optimization. The results of the signal separation experiments reveal that the signal separation performance of the proposed algorithm is superior to that of the conventional ICA-based BSS method, even under reverberant conditions.
Hiroshi Saruwatari, Toshiya Kawamura, Katsuyuki Sawai, Atsunobu Kaminuma, Masao Sakata
ICASSP1
2002 Design and collection of acoustic sound data for hands-free speech recognition and sound scene understanding
abstract
The sound data for open evaluation is necessary for studies such as sound source localization, sound retrieval, sound recognition and hands-free speech recognition in real acoustic environments. This paper reports on our project for acoustic data collection. There are many kinds of sound scenes in real environments. The sound scene is specified by sound sources and room acoustics. The number of combinations of the sound sources, source positions and rooms is huge in real acoustic environments. We assumed that the sound in the environments can be simulated by convolution of the isolated sound sources and impulse responses. As an isolated sound source, hundred kinds of environment sounds and speech sounds are collected. The impulse responses are collected in various acoustic environments. Additionally we collected sounds from a moving source. In this paper, progress of our sound scene database collection project and application to environment sound recognition and hands-free speech recognition are described.
Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Yutaka Kaneda, Takeshi Yamada, Takanobu Nishiura, Tetsunori Kobayashi, Shiro Ise, Hiroshi Saruwatari
ICME (2)9
2002 Selective multi-path acoustic model based on database likelihoods
abstract
An efficient multi-path acoustic model based on database likelihoods for spontaneous speech recognition is presented. Although a multipath phone HMM that has several models for different target in parallel is considered effective to express multi-style or speed-variant nature of spontaneous speech, assuming various model to match at every time for all phones may cause mismatch of unintended mo- del, and spoil the model constraints. We propose defining a multipath model that has several different state resolutions only for the distortive phones selectively. The phone set is selected through an analysis of the likelihoods and duration times of phone segments in a spoken dialogue corpus using automatic viterbi alignment. Experiments on three test-sets showed that our multi-path model based on the phone selection can achieve better accuracy than a simple single-path model, whereas a full multi-path model without phone selection causes much degradation of accuracy.
Akinobu Lee, Yuichiro Mera, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2002 Speech enhancement in car environment using blind source separation
abstract
We propose a new algorithm for blind source separation (BSS), in which independent component analysis (ICA) and beamforming are combined to resolve the low-convergence problem through optimization in ICA. The proposed method consists of the following four parts: (1) frequency-domain ICA with direction-of-arrival (DOA) estimation, (2) null beamforming based on the estimated DOA, (3) diversity of (1) and (2) in both iteration and frequency domain, and (4) subband elimination (SBE) based on the independence among the separated signals. The temporal alternation between ICA and beamforming can realize fast- and high-convergence optimization. Also SBE enforcedly eliminates the subband components in which the separation could not be performed well. The experiment in a real car environment reveals that the proposed method can improve the qualities of the separated speech and word recognition rates for both directional and diffusive noises.
Hiroshi Saruwatari, Katsuyuki Sawai, Akinobu Lee, Kiyohiro Shikano, Atsunobu Kaminuma, Masao Sakata
INTERSPEECH1
2002 Spectral subtraction in noisy environments applied to speaker adaptation based on HMM sufficient statistics
abstract
Noise and speaker adaptation techniques are essential to realize robust speech recognition in real noisy environments . In this paper, we applied spectral subtraction to an unsupervised speaker adaptation algorithm in noisy environments. The adaptation algorithm consists of the following five steps. (1) Spectral subtraction is carried out for noise added database. (2) Noise matched acoustic models are trained by using noise added speech database. (3) HMM sufficient statistics for each speaker are calculated from noise added speech database, and stored. (4) According to one arbitrary utterance, speakers close to a test speaker are selected by using speaker GMMs. (5) Speaker adapted acoustic models are constructed from HMM sufficient statistics of the selected speakers. We evaluated our unsupervised speaker adaptation algorithm in noisy environments in the 20k dictation task. The recognition experiments show that our speaker adapted acoustic model can achieve 82% word accuracy in 20dB SNR, which is about 6% higher than that of the noise matched models trained by Forward-Backward algorithm.
Shingo Yamade, Kanako Matsunami, Akira Baba, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH5
2002 ASKA: receptionist robot with speech dialogue system
abstract
We implemented a humanoid robot, ASKA, in our university reception desk for the computerized university guidance. ASKA can recognize a user's question utterance, and answer the user's question by its text-to-speech voice, hand gesture and head movement. This paper describes the speech related parts of ASKA. ASKA can deal with a wide task domain of 20k large vocabulary using a word trigram model and an elaborated speaker-independent acoustic model. ASKA can also make a response with keyword and key-phrase detection in the N-best speech recognition results. The word recognition rate for the reception task is 90.9%, and the rate for the out-of-domain task is 78.9%. The correct response rate for the reception task is 61.7%. Users can enjoy their question-answering with ASKA.
Ryuichi Nisimura, Takashi Uchida, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano, Yoshio Matsumoto
IROS4
2001 Fundamental limitation of frequency domain blind source separation for convolutive mixture of speech
abstract
Despite several recent proposals to achieve blind source separation (BSS) for realistic acoustic signals, separation performance is still not good enough. In particular, when the length of impulse response is long, performance is highly limited. We show it is useless to be constrained by the condition, P /spl Lt/ T, where T is the frame size of FFT and P is the length of room impulse response. From our experiments. a frame size of 256 or 512 (32 or 64 ms at a sampling frequency of 8 kHz) is best even for the long room reverberation of T/sub R/ = 150 and 300 ms. We also clarified the reason for poor performance of BSS in a long reverberant environment, finding that separation is achieved chiefly for the sound from the direction of jammers because BSS cannot calculate the inverse of the room transfer function both for the target and jammer signals.
Shoko Araki, Shoji Makino, Tsuyoki Nishikawa, Hiroshi Saruwatari
ICASSP4
2001 Direction of arrival estimation based on nonlinear microphone array
abstract
This paper describes a new method for estimating the direction of arrival (DOA) using a nonlinear microphone array based on complementary beamforming. Complementary beamforming is based on two types of beamformers designed to obtain complementary directivity patterns with each other. In this system, since the resultant directivity pattern is proportional to the product of these directivity patterns, the proposed method can be used to estimate DOAs even when the number of sound sources is equal to or exceeds that of microphones. First, DOA-estimation experiments are performed using actual devices in real acoustic environments. The results clarify that DOA estimation for two sound sources can be accomplished by the proposed method with only two microphones. Also, by comparing the resolutions of DOA estimation by the proposed method and by the conventional minimum variance method, we can show that the performance of the proposed method is superior to that of the conventional method.
Hidekazu Kamiyanagida, Hiroshi Saruwatari, Kazuya Takeda, Fumitada Itakura
ICASSP2
2001 Blind source separation combining frequency-domain ICA and beamforming
abstract
We describe a new method of blind source separation (BSS) on a microphone array combining subband independent component analysis (ICA) and beamforming. The proposed array system consists of the following three sections: (1) subband-ICA-based BSS section with direction-of-arrival (DOA) estimation; (2) null beamforming section based on the estimated DOA information; and (3) integration of (1) and (2) based on the algorithm diversity. Using this technique, we can resolve the low-convergence problem through optimization in ICA. The results of the signal separation experiments reveal that a noise reduction rate (NRR) of about 18 dB is obtained under the nonreverberant condition, and NRR of 8 dB and 6 dB are obtained in the case that the reverberation times are 150 msec and 300 msec. These performances are superior to those of both simple ICA-based BSS and simple beamforming method.
Hiroshi Saruwatari, Satoshi Kurita, Kazuya Takeda
ICASSP1
2001 Voice conversion algorithm based on Gaussian mixture model with dynamic frequency warping of STRAIGHT spectrum
abstract
In the voice conversion algorithm based on the Gaussian Mixture Model (GMM) applied to STRAIGHT, quality of converted speech is degraded because the converted spectrum is exceedingly smooth. We propose the GMM-based algorithm with dynamic frequency warping to avoid the over-smoothing. We also propose an addition of the weighted residual spectrum, which is the difference between the GMM-based converted spectrum and the frequency-warped spectrum, to avoid the deterioration of conversion-accuracy on speaker individuality. Results of the evaluation experiments clarify that the converted speech quality is better than that of the GMM-based algorithm, and the conversion-accuracy on speaker individuality is the same as that of the GMM-based algorithm in the proposed method with the properly-weighted residual spectrum.
Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP2
2001 Equivalence between frequency domain blind source separation and frequency domain adaptive null beamformers
abstract
Frequency domain Blind Source Separation (BSS) is shown to be equivalent to two sets of frequency domain adaptive microphone arrays, that is, Adaptive Null Beamformers (ANB). The unmixing matrix of the BSS and the filter coefficients of the ANB converge to the same solution in the mean square error sense if the two source signals are ideally independent. This understanding clearly explains the poor performance of the BSS in a real room with long reverberation. The fundamental difference exists in the adaptation period when they should adapt. That is, the ANB can adapt in the presence of a jammer but the absence of a target, whereas the BSS can adapt in the presence of a target and jammer, and also in the presence of only a target.
Shoko Araki, Shoji Makino, Ryo Mukai, Hiroshi Saruwatari
INTERSPEECH4
2001 Automatic n-gram language model creation from web resources
abstract
EUROSPEECH2001: the 7th European Conference on Speech Communication and Technology, September 3-7, 2001, Aalborg, Denmark.
Ryuichi Nisimura, Kumiko Komatsu, Yuka Kuroda, Kentaro Nagatomo, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH6
2001 Blind source separation for speech based on fast-convergence algorithm with ICA and beamforming
abstract
We propose a new algorithm for blind source separation (BSS), in which independent component analysis (ICA) and beamforming are combined to resolve the low-convergence problem through optimization in ICA. The proposed method consists of the following three parts: (1) frequency-domain ICA with direction-of-arrival (DOA) estimation, (2) null beamforming based on the estimated DOA, and (3) integration of (1) and (2) based on the algorithm diversity in both iteration and frequency domain. The inverse of the mixing matrix obtained by ICA is temporally substituted by the matrix based on null beamforming through iterative optimization, and the temporal alternation between ICA and beamforming can realize fast- and high-convergence optimization. The results of the signal separation experiments reveal that the signal separation performance of the proposed algorithm is superior to that of the conventional ICA-based BSS method, even under reverberant conditions.
Hiroshi Saruwatari, Toshiya Kawamura, Kiyohiro Shikano
INTERSPEECH1
2001 High quality voice conversion based on Gaussian mixture model with dynamic frequency warping
abstract
In the voice conversion algorithm based on the Gaussian Mixture Model (GMM), quality of the converted speech is degraded because the converted spectrum is exceedingly smoothed. In this paper, we newly propose the GMM-based algorithm with the Dynamic Frequency Warping (DFW) to avoid the over-smoothing. We also propose that the converted spectrum is calculated by mixing the GMM-based converted spectrum and the DFW-based converted spectrum, to avoid the deterioration of conversion-accuracy on speaker individuality. Results of the evaluation experiments clarify that the converted speech quality is better than that of the GMM-based algorithm, and the conversionaccuracy on speaker individuality is the same as that of the GMM-based algorithm in the proposed algorithm with the proper weight for mixing spectra.
Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH2
2001 Unsupervised noisy environment adaptation algorithm using MLLR and speaker selection
abstract
An unsupervised acoustic model adaptation algorithm using MLLR and speaker selection for noisy environments is proposed. The proposed algorithm requires only one arbitrary utterance and environmental noise data. The adaptation procedure is composed of the following four steps. (1) Speaker selection from a large number of database speakers is carried out using GMM speaker models based on one arbitrary utterance. (2) Initial speaker adapted HMM acoustic models are calculated from the HMM sufficient statistics of the selected speakers, where the sufficient HMM statistics are pre-calculated and stored. (3) A small subset of the clean speech database from the selected speakers and the environment noise data are superimposed. (4) MLLR adaptation is carried out using the noise-superimposed speech database from the selected speakers. The proposed algorithm is evaluated in a 20k vocabulary dictation task for newspaper in noisy environments. We attain 85.7% word correct rate in 25dB SNR, which is slightly better than the matched model by the E-M training using noise superimposed whole speech database. The proposed algorithm is also 7% better than the HMM composition algorithm.
Miichi Yamada, Akira Baba, Shinichi Yoshizawa, Yuichiro Mera, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH6
2000 Evaluation of blind signal separation method using directivity pattern under reverberant conditions
abstract
This paper describes a new blind signal separation method using the directivity patterns of a microphone array. In this method, to deal with the arriving lags among each microphone, the inverses of the mixing matrices are calculated in the frequency domain so that the separated signals are mutually independent. Since the calculations are carried out in each frequency independently, the following problems arise: (1) permutation of each sound source, (2) arbitrariness of each source gain. In this paper, we propose a new solution that directivity patterns are explicitly used to estimate each sound source direction. As the results of signal separation experiments, it is shown that the proposed method improves the SNR of degraded speech by about 16 dB under non-reverberant condition. Also, the proposed method improves the SNR by 8.7 dB when the reverberation time is 184 ms, and by 5.1 dB when the reverberation time is 322 ms.
Satoshi Kurita, Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura
ICASSP2
2000 Speech enhancement using nonlinear microphone array with noise adaptive complementary beamforming
abstract
This paper describes an improved complementary beamforming microphone array with a new noise adaptation. Complementary beamforming is based on two types of beamformers designed to obtain complementary directivity patterns. In this system, two directivity patterns of the beamformers are adapted to the noise directions so that the expectation values of each noise power spectrum are minimized. Using this technique, we can realize the directional nulls for each noise even when the number of sound sources exceeds that of microphones. To evaluate the effectiveness, speech enhancement experiments are performed based on computer simulations with a two-element array and three sound sources. Compared with the conventional spectral subtraction method cascaded with the adaptive beamformer, it is shown that the proposed array improves the signal-to-noise ratio of degraded speech by more than 6 dB and performs more than 18% better in word recognition rates when the interfering noise is two speakers.
Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura
ICASSP1
2000 Blind source separation based on subband ICA and beamforming
abstract
This paper describes a new blind source separation (BSS) method on microphone array using the subband independent component analysis (ICA) and beamforming. The proposed array system consists of the following three sections: (1) subband-ICA-based BSS section, (2) null beamforming section, and (3) integration of (1) and (2) based on the algorithm diversity. Using this technique, we can resolve the low-convergence problem on optimization in ICA. Signal separation and speech recognition experiments clarify that the noise reduction rate (NRR) of about 18 dB is obtained under the nonreverberant condition, and NRRs of 8 dB and 6 dB are obtained in the case that the reverberation times are 150 msec and 300 msec. These performances are superior to those of both simple ICA-based BSS and simple beamforming method. Also, the improvements of the proposed method in word recognition rates are superior to those of the conventional ICA-based BSS method under all reverberant conditions.
Hiroshi Saruwatari, Satoshi Kurita, Kazuya Takeda, Fumitada Itakura, Kiyohiro Shikano
INTERSPEECH1
2000 Straight-based voice conversion algorithm based on Gaussian mixture model
abstract
The voice conversion algorithm based on the Gaussian mixture model (GMM) has also been proposed by Stylianou et al. In this algorithm, the acoustic space of a speaker is represented continuously. In this paper, we apply this GMM-based voice conversion algorithm to STRAIGHT proposed by Kawahara et al., which is recognized as a high quality vocoder. In order to evaluate this voice conversion algorithm, we perform subjective and objective experiments on speech quality and speaker individuality, comparing with the method based on the codebook mapping. As results, the performance of the GMM-based voice conversion algorithm is better than that of the codebook mapping method. Effects by the amount of training data for the voice conversion algorithms are also investigated, as well as the number of the Gaussian mixtures. These evaluation results clarify that the GMM-based voice conversion algorithm is successfully applied to STRAIGHT.
Tomoki Toda, Jinlin Lu, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
1999 Compensating of room acoustic transfer functions affected by change of room temperature
abstract
This paper proposes an efficient compensation method using a first-order approximation of time axis scaling for the variations of the room acoustic transfer function. The time axis scaling model is based on the fact that the change of the sound velocity due to the change of room temperature is a dominant factor for the variations of room impulse response affected by environmental conditions. In this paper, the effectiveness of the compensation method is evaluated using room impulse responses measured in the real environment. As the results, it is clarified that the variations of room impulse response can be modeled by the first-order approximated time axis scaling when the successive re-estimation is performed every small change of temperature. Furthermore, it is shown that the compensation method applied to an inverse filtering based dereverberation approach improves the intelligibility and speech recognition rates dramatically.
Michiaki Omura, Motohiko Yada, Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura
ICASSP3
1999 Speech enhancement using nonlinear microphone array with complementary beamforming
abstract
This paper describes an improved spectral subtraction method by using the complementary beamforming microphone array to enhance noisy speech signals for speech recognition. The complementary beamforming is based on two types of beamformers designed to obtain complementary directivity patterns with respect to each other. It is shown that the nonlinear subtraction processing with complementary beamforming can result in a kind of the spectral subtraction without the need for speech pause detection. In addition, the design of the optimization algorithm for the directivity pattern is also described. To evaluate the effectiveness, speech enhancement experiments and speech recognition experiments are performed based on computer simulations. In comparison with the optimized conventional delay-and-sum array, it is shown that the proposed array improves the signal-to-noise ratio of degraded speech by about 2 dB and performs about 10% better in word recognition rates under heavy noisy conditions.
Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura
ICASSP1
1999 Speech enhancement using nonlinear microphone array under nonstationary noise conditions
Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura
EUROSPEECH1