VLDB 2026 Research / reviewers in the wild / expert
Daisuke Saito
dblp:17/7825
· DBLP profile ↗
75ranked-venue papers
13as first author
30since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 48 · 8 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Benchmarking Prosody Encoding in Discrete Speech TokensabstractRecently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for various tasks. However, these discrete tokens are typically learned in advance, separately from the training of language models or downstream tasks. As a result, choices related to discretization, such as the SSL model used or the number of clusters, must be made heuristically. In particular, speech language models are expected to understand and generate responses that reflect not only the semantic content but also prosodic features. Yet, there has been limited research on the ability of discrete tokens to capture prosodic information. To address this gap, this study conducts a comprehensive analysis focusing on prosodic encoding based on their sensitivity to the artificially modified prosody, aiming to provide practical guidelines for designing discrete tokens. Kentaro Onda, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu |
ASRU | 3 |
| 2025 | A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
Haopeng Geng, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2025 | Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2025 | Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2024 | Evaluating Preschoolers' Block Programming Using Complexity and Personality TraitsabstractProgramming learning at an early age effectively fosters logical thinking and self-centeredness, but an appropriate evaluation method for young learners has yet to be established. Herein we propose a new learning evaluation method that incorporates problem constructs and the complexity metrics used in software engineering quality assessments. Specifically, we investigate the relationships between changes in block pro-gramming complexity, personality traits, and learning effects. Evaluation rubrics and log data assess complexity, while person-ality traits are based on Big-5. Then the learning effects of 34 kindergarten children participating in workshops are analyzed in terms of complexity and personality traits. After learning, first-time programmers tend to show a large increase in complexity. The correlation with the rubric score is$\rho=0.43$, and the correlation with log data is$\mathbf{r}s=0.92$. Furthermore, analysis using the Big-5 gives$\rho=0.609$for the rate of increase in extraversion and complexity, indicating a strong relationship between learning effects and personality traits. In the future, this data will be used to build AI tools for automatic evaluations, feedback, and learning curriculum recommendations. Yui Ono, Daisuke Saito, Hironori Washizaki |
CSEE&T | 2 |
| 2024 | Do Learned Speech Symbols Follow Zipf's Law?abstractIn this study, we investigate whether speech symbols, learned through deep learning, follow Zipf’s law, akin to natural language symbols. Zipf’s law is an empirical law that delineates the frequency distribution of words, forming fundamentals for statistical analysis in natural language processing. Natural language symbols, which are invented by humans to symbolize speech content, are recognized to comply with this law. On the other hand, recent breakthroughs in spoken language processing have given rise to the development of learned speech symbols; these are data-driven symbolizations of speech content. Our objective is to ascertain whether these datadriven speech symbols follow Zipf’s law, as the same as natural language symbols. Through our investigation, we aim to forge new ways for the statistical analysis of spoken language processing. Shinnosuke Takamichi, Hiroki Maeda, Joonyong Park, Daisuke Saito, Hiroshi Saruwatari |
ICASSP | 4 |
| 2024 | A ChatGPT-based oral Q&A practice system for first-time student participants in international conferences
Mayuko Aiba, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2024 | A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
Kentaro Onda, Joonyong Park, Nobuaki Minematsu, Daisuke Saito |
INTERSPEECH | 4 |
| 2024 | Acceleration of Posteriorgram-based DTW by Distilling the Class-to-class Distances Encoded in the Classifier Used to Calculate Posteriors
Haitong Sun, Jaehyun Choi, Nobuaki Minematsu, Daisuke Saito |
INTERSPEECH | 4 |
| 2024 | Analysis and Visualization of Directional Diversity in Listening Fluency of World Englishes Speakers in the Framework of Mutual Shadowing
Yu Tomita, Yingxiang Gao, Nobuaki Minematsu, Noriko Nakanishi, Daisuke Saito |
INTERSPEECH | 5 |
| 2024 | Enhancing Programming Education through Game-Based Learning: Design and Implementation of a Puyo Puyo-Inspired Teaching ToolabstractAlthough programming is part of primary school curricula in many countries, barriers persist for elementary students learning programming such as an insufficient understanding of the underlying mathematics, complex concepts, and purpose of programming. These challenges often lead to disinterest. Herein we present an innovative game-design-based programming education tool. Students progressively enhance a classic game, Puyo Puyo, using fundamental programming concepts and selecting the appropriate code. This engaging approach improves students' computational thinking abilities as they transform code into a functional game. Here, we describe the tool's background and structure. Then we detail a workshop using the tool, including analyzing the changes in students' programming skills, computational thinking, and interest in programming. Finally, we summarize the findings and future research directions. Ruochen Tian, Daisuke Saito, Hironori Washizaki, Yoshiaki Fukazawa, Hiroshi Kobayashi, Ayumi Tsuji |
SIGCSE (2) | 2 |
| 2023 | Programming Education for Young People using the Falling-Puzzle Game, "Puyo Puyo"abstractThis study utilizes the game rules of a falling-puzzle game, developed as a consumer-oriented digital game, in programming education for young people. When digital games are used in programming education, they are often designed specifically for that purpose. In this study, we focused on the game rules of the consumer falling-object puzzle game, “Puyo Puyo.” Learning impact was investigated on 23 Japanese elementary, junior high, and senior high school students. The results show that the concepts of branching and arraying in programming were taught effectively; thus, consumer digital game rules can be used as a case study for programming education. Daisuke Saito, Ruochen Tian, Hironori Washizaki, Yoshiaki Fukazawa |
EDUCON | 1 |
| 2023 | Multiple Acoustic Features Speech Emotion Recognition Using Cross-Attention TransformerabstractSpeech emotion recognition (SER) is a challenging task whose performance heavily relies on suitable affect-salient representations. Recently, transformer has exhibited outstanding qualities in learning relevant representations associated with this task. However, a normal transformer is only able to process the uni-source input, and there is often only one kind of input feature in a transformer-based SER system, which may cause limited knowledge. In this paper, we attempt to use the cross-attention transformer (CAT) to handle bi-source input. We propose a novel SER system to better fuse three types of acoustic features – raw waveform data, spectrogram, and MFCC using CAT. Experiments conducted on the IEMOCAP benchmark dataset have shown that our proposed system can achieve a 73.80% weighted accuracy (WA) and 74.25% unweighted accuracy (UA), which outperforms existing state-of-the-art approaches. Yurun He, Nobuaki Minematsu, Daisuke Saito |
ICASSP | 3 |
| 2023 | Automatic Prediction of Language Learners' Listenability Using Speech and Text Features Extracted from Listening Drills
Yingxiang Gao, Jaehyun Choi, Nobuaki Minematsu, Noriko Nakanishi, Daisuke Saito |
INTERSPEECH | 5 |
| 2023 | Gender Characteristics and Computational Thinking in ScratchabstractThis study investigates the Computational Thinking skill differences among novice programmers in relation to gender. Block-based visual programming languages such as Scratch particularly benefit K-12 programmers because they learn how to code intuitively. Our study analyzed 124 (62 males, 62 females) Scratch projects on the Scratch website, categorized projects on the basis of each user's gender and project type, and compared their Computational Thinking scores. The results of this study suggest that project types preferred by males require more programming construct reflected in the Computational Thinking score than that of females. Because gender differences appear by project type, project type presumably influences the gender gap in scores. Rose Niousha, Daisuke Saito, Hironori Washizaki, Yoshiaki Fukazawa |
SIGCSE (2) | 2 |
| 2023 | Improving Semi-Supervised Differentiable Synthesizer Sound Matching for Practical ApplicationsabstractWhile synthesizers have become commonplace in music production, many users find it difficult to control the parameters of a synthesizer to create a sound as they intended. In order to assist the user, thesound matchingtask aims to estimate synthesis parameters that produce a sound that is as close as possible to the query sound. Recently, neural networks have been employed for this task. These neural networks are trained on paired data of synthesis parameters and the corresponding output sound, optimizing a loss of synthesis parameters. However, query by the user usually consists of real-world sounds, different from the synthesizer output sounds used as training data. In a previous work, the authors presented a sound matching method where the synthesizer is implemented using differentiable DSP. The estimator network could then be trained by directly optimizing the spectral similarity between the original sound and the output sound. Furthermore, the network could be trained on real-world sounds whose ground-truth synthesis parameters are unavailable. This method was shown to improve the match quality in both objective and subjective measures. In this work, we experiment with different synthesizer configurations and extend this approach to a more practical synthesizer with effect modules and envelope generators. We propose a novel training strategy where the network is fully trained using both parameter loss and spectral loss. We show that models trained using this strategy is able to utilize the chorus effect effectively while models that switch completely to spectral loss underutilizes the chorus effect. Naotake Masuda, Daisuke Saito |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Quantifying Discriminability between NMF BasesabstractDiscriminative nonnegative matrix factorization (DNMF) has been investigated as a promising basis-learning method for monaural source separation. To the best of our knowledge, however, no good and sound discussion has been made on quantitative definition of discriminability and it is difficult to evaluate how discriminative DNMF is actually. This paper introduces a quantitative measure to calculate how discriminative two NMF bases are. From the viewpoint of our measure, we compare three basis-learning methods of plain NMF, DNMF, and minimum-volume (min-vol) NMF. Experimental results of monaural speech separation reveal that min-vol NMF actually learns as discriminative bases as DNMF and achieves the best separation performance. This is probably because min-vol NMF can learn the most compact basis possible that can cover training data. Eisuke Konno, Daisuke Saito, Nobuaki Minematsu |
ICASSP | 2 |
| 2022 | Text-to-speech synthesis using spectral modeling based on non-negative autoencoder
Takeru Gorai, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2022 | Detection of Learners' Listening Breakdown with Oral Dictation and Its Use to Model Listening Skill Improvement Exclusively Through Shadowing
Takuya Kunihara, Chuanbo Zhu 0001, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi |
INTERSPEECH | 3 |
| 2022 | Automatic Prediction of Intelligibility of Words and Phonemes Produced Orally by Japanese Learners of EnglishabstractThe practical goal for language learning is smooth communication with others, and many teachers have a strong focus on measurement of not accentedness but intelligibility, often regarded as correctness of actual understanding. However, automatic prediction of intelligibility has not been well developed especially for smaller units such as words and phonemes. This is mainly because of difficulty of measuring while-listening behaviors of listeners, and thus it was difficult to build an L2 speech corpus of a sufficient size with intelligibility annotation to train a network-based predictor. In this paper, we annotate intelligibility using oral dictation with a small delay, i.e., shadowing, to collect a large enough corpus from two raters with different language backgrounds. Since perceived intelligibility depends on their language background, inter-rater difference should be taken into account. Therefore with this corpus, a multi-rater neural model is built to predict each rater's intelligibility of the individual words and phonemes in L2 speech. Two tasks are examined, i.e., regression of intelligibility scores and classification of a given segment to be intelligible or not. Results show that our model has higher F1 scores than intra-rater agreements, indicating that our model can simulate the two raters accurately well although they have different language background. Chuanbo Zhu 0001, Takuya Kunihara, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi |
SLT | 3 |
| 2022 | Voice Conversion Based on Deep Neural Networks for Time-Variant Linear TransformationsabstractThis paper describes a novel framework of voice conversion to improve the conversion performance against the amount of training data. In voice conversion, deep neural networks are used as conversion models that map source to target features. In this framework, it generally needs a larger amount of training data and bigger models to build more accurate conversion models. This condition, however, will reduce the usability of voice conversion. In this paper, in order to improve the conversion performance versus the amount of training data, a top-down knowledge is introduced into models as prior. We expect that we can take advantage of top-down knowledge we have instead of preparing a large amount of data. In the proposed method, the conversion process of features is restricted to time-variant linear transformation on cepstral space. It explicitly utilizes an attribute of voice conversion i.e. homo-domain mapping, which is not common in automatic speech recognition or text-to-speech synthesis. In other words, in VC, the input and output are on the same feature domain. In addition, it also makes it possible to explicitly consider the physical difference between speakers such as the difference of vocal tract length. The assumption of the homo-domain mapping is related to conversion methods based on spectral differentials, and then the relation is discussed in the paper. Experiments demonstrate the effectiveness of our proposal and the way that the constraint of linear transformation works is investigated. Gaku Kotani, Daisuke Saito, Nobuaki Minematsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Singer Diarization for Polyphonic Music With Unison SingingabstractThis paper introduces a new framework for singer diarization, which is a technique to reveal who sings when in songs with multiple singers. Although various techniques have been developed to analyze and extract features of singing voices in musical audio signals, most of them assume that a song is sung by a single singer, and singer diarization for multiple singers has not been well studied in the field of singing information processing. To deal with multiple speakers in speech analysis, speaker diarization has been explored to handle overlapped speech voices, but cannot handle singing voices well because of acoustic differences between singing and speech voices. This paper therefore proposes a new diarization framework specialized in singing voices. To achieve high accuracy in overlap detection, this paper proposes a novel acoustic feature named Cosacorr score, which is helpful in estimating whether a song is sung by more than one singer. After extracting singing voices from polyphonic music by using a singing voice separation technique, the framework adopts an existing ArcFace technique to extract discriminative singer representations from short segments of the separated singing voices. The framework is evaluated by using a new private dataset of unison singing voices, which is constructed using commercially available compact discs (CDs). The experimental results show that the proposed framework outperformed the baseline method for speaker diarization in terms of diarization error rate (DER). Hitoshi Suda, Daisuke Saito, Satoru Fukayama, Tomoyasu Nakano, Masataka Goto |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Multi-Granularity Annotation of Instantaneous Intelligibility of Learners' Utterances Based on Shadowing TechniquesabstractThe practical goal of pronunciation training is to acquire an intelligible enough pronunciation, not a native-like pronunciation. In our studies [1], [2], we proposed a method that can annotate instantaneous intelligibility of a given L2 English utterance by monitoring listeners' listening behaviors. The listeners were asked to shadow the L2 utterance and the degree of being inarticulate in shadowing was automatically quantified to be used as scores of the instantaneous intelligibility, which were shown to be highly correlated with subjective intelligibility scores. In the present paper, we make an objective assessment of the proposed method, where the automatic scores are compared with those calculated objectively based on manual transcripts of the shadowings. Experiments show that the former scores have such a high correlation as 0.935 with the latter scores, which is higher than correlation obtained using another kind of automatic scores calculated by transcribing the shadowings with ASR. Further, since intelligibility is sometimes discussed in pronunciation training in such smaller units as phonemes, our method is experimentally applied to intelligibility annotation with multiple granularity. Experiments show a high validity of our method to calculate instantaneous intelligibility in units of word, syllable, phoneme, and frame. Chuanbo Zhu 0001, Ryo Hakoda, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi, Tazuko Nishimura |
ASRU | 3 |
| 2021 | Automated Educational Program Mapping on Learning Standards in Computer ScienceabstractThere are many lectures for students or working adults. Educational institutions have created a table that associates educational courses with learning standards in order to understand the content of these courses. Learning standards such as SFIA (Skills Framework for an Information Age) indicate the skills that learners should learn. The mapped table can make contents clear for students or other educational institutions. Koki Miura, Daisuke Saito, Hironori Washizaki, Yoshiaki Fukazawa |
COMPSAC | 2 |
| 2021 | Preliminary Literature Review of Machine Learning System Development PracticesabstractTo guide practitioners and researchers to design and research Machine Learning (ML) system development processes, we conduct a preliminary literature review on ML system development practices. We identified seven papers and two other papers determined in an ad-hoc review. Our findings include emphasized phases in ML system developments, frequently described ML-specific practices, and tailored traditional practices. Yasuhiro Watanabe, Hironori Washizaki, Kazunori Sakamoto, Daisuke Saito, Kiyoshi Honda, Naohiko Tsuda, Yoshiaki Fukazawa, Nobukazu Yoshioka |
COMPSAC | 4 |
| 2021 | Quality Diversity for Synthesizer Sound MatchingabstractIt is difficult to adjust the parameters of a complex synthesizer to create the desired sound. As such, sound matching, the estimation of synthesis parameters that can replicate a certain sound, is a task that has often been researched, utilizing optimization methods such as genetic algorithm (GA). In this paper, we introduce a novelty-based objective for GA-based sound matching. Our contribution is two-fold. First, we show that the novelty objective is able to improve the quality of sound matching by maintaining phenotypic diversity in the population. Second, we introduce a quality diversity approach to the problem of sound matching, aiming to find a diverse set of matching sounds. We show that the novelty objective is effective in producing high-performing solutions that are diverse in terms of specified audio features. This approach allows for a new way of discovering sounds and exploring the capabilities of a synthesizer. Naotake Masuda, Daisuke Saito |
DAFx | 2 |
| 2021 | Work-in-Progress: Analysis of the use of Mentoring with Online Mob ProgrammingabstractExtreme programming (XP) approaches such as pair and mob programming (PP and MP, respectively) have been introduced to the educational sector. However, higher education curricula have started implementing online courses as a method for teaching programming, which poses challenges when conducting XP. Because social factors are vital to learning, we examine a project-based learning method using online MP that introduces a high level of communication. It also maintains the advantages of PP through the “driver” and “navigator” approach. However, there are multiple complications, including discomfort, slow code generation, and interpersonal issues, when using conventional MP. Accordingly, we introduce a mentoring approach to the MP setup and examine the differences compared to conventional MP sessions. We create and analyze co-occurrence networks of codes found via open coding. In this paper, we explain our method and show its advantages based on the results of our analysis. Shota Kaieda, Daisuke Saito, Hironori Washizaki, Yoshiaki Fukazawa |
EDUCON | 2 |
| 2021 | Lexical Density Analysis of Word Productions in Japanese English Using Acoustic Word Embeddings
Shintaro Ando, Nobuaki Minematsu, Daisuke Saito |
Interspeech | 3 |
| 2021 | Optimized Prediction of Fluency of L2 English Based on Interpretable Network Using Quantity of Phonation and Quality of PronunciationabstractThis paper presents results of a joint project between an engineering team of a university and an educational team of another to develop an online fluency assessment system for Japanese learners of English. A picture description corpus of English spoken by 90 learners and 10 native speakers was used, where fluency was rated by other 10 native raters for each speaker manually. The assessment system was built to predict the averaged manual scores. For system development, a special focus was put on two separate purposes. The assessment system was trained in such an analytical way that teachers can know and discuss which speech features contribute more to fluency prediction, and in such a technical way that teachers' knowledge can be involved for training the system, which can be further optimized using an interpretable network. Experiments showed that quality-of-pronunciation features are much more helpful than quantity-of-phonation features, and the optimized system reached an extremely high correlation of 0.956 with the averaged manual scores, which is higher than the maximum of inter-rater correlations (0.910). Ayano Yasukagawa, Daisuke Saito, Nobuaki Minematsu, Kazuya Saito |
SLT | 3 |
| 2021 | Comparing Participants' Brainwaves During Solo, Pair, and Mob ProgrammingabstractAbstract Participants’ feelings and impressions utilizing electroencephalography (EEG) and the effectiveness of code are compared for different types of programming sessions. EEG information is obtained as an alternate viewpoint during three programming sessions (solo, pair, and mob programming). MindWave Mobile 2 (brainwave detector) is equipped to collect the attention levels, meditation levels, and EEG brainwaves. These data are utilized to distinguish efficiencies, weaknesses, and points of interest by programming session. The results provide preliminary information to distinguish between the three sessions, but further studies are necessary to make firm conclusions. Additionally, alternative methods or systems are required to analyze the collected data. Makoto Shiraishi, Hironori Washizaki, Daisuke Saito, Yoshiaki Fukazawa |
XP | 3 |
| 2020 | Attention-Based Speaker Embeddings for One-Shot Voice Conversion
Tatsuma Ishihara, Daisuke Saito |
INTERSPEECH | 2 |
| 2020 | Shadowability Annotation with Fine Granularity on L2 Utterances and its Improvement with Native Listeners' Script-Shadowing
Zhenchao Lin, Ryo Takashima, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi |
INTERSPEECH | 3 |
| 2020 | Discriminative Method to Extract Coarse Prosodic Structure and its Application for Statistical Phrase/Accent Command Estimation
Yuma Shirahata, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2020 | Nonparallel Training of Exemplar-Based Voice Conversion System Using INCA-Based Alignment Technique
Hitoshi Suda, Gaku Kotani, Daisuke Saito |
INTERSPEECH | 3 |
| 2019 | Analysis of Native Listeners' Facial Microexpressions While Shadowing Non-Native Speech - Potential of Shadowers' Facial Expressions for Comprehensibility Prediction
Tasavat Trisitichoke, Shintaro Ando, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2019 | Rubric to Evaluate Programming Learning of Elementary School StudentsabstractAs more children are exposed to computer science and programming, numerous indicators have been proposed to evaluate programming learning. Because these indicators are not divided by step in a learning goal, learner's growth cannot be evaluated in detail. Our research focuses on resolving this issue in the field of programming. Herein we propose a rubric to measure the progress of programming learning for elementary school students. This rubric, which is comprised of indices, aims to evaluate learning achievement of logical skills such as logical thinking and problem solving using a unified learning goal. Furthermore, this rubric consists of 30 evaluation items in 8 categories. Then we investigate whether this rubric can be adapted to existing workshops of programming learning. The workshops occurred in 2016 and 2017. in addition, A total of 101 students between 6 and 12 years old participated in the workshops. The rubric successfully evaluates programming learning workshops as it provides a unified evaluation that covers common learning objectives in existing indicators. Our rubric is available at https://g7programming.jp/plr/. Daisuke Saito, Hironori Washizaki, Yoshiaki Fukazawa, Mariko Tamura, Yuki Sakuragi |
SIGCSE | 1 |
| 2019 | Many-to-Many and Completely Parallel-Data-Free Voice Conversion Based on Eigenspace DNNabstractMedia conversion of image, text, speech, etc., generally requires a large amount of parallel data for training a conversion model. Recently, methods for training the model using no or a small amount of parallel data draw researchers' attention. In many-to-many voice conversion, since it is often hard to collect parallel data from every pair of speakers, the conversion models requiring no parallel data are desired. Conventional many-to-many voice conversion models required a large amount of prestored parallel data to acquire prior knowledge of the entire speaker space. Then, a specific model from an arbitrary speaker to another can be realized by adapting a few model parameters. Although these conversion models certainly do not use parallel data in an adaptation step, they still use parallel data for prior training. In this study, we aim at realizing completely parallel-data-free and many-to-many voice conversion. The proposed method uses both Eigenvoice Gaussian mixture models (EVGMM) and Deep neural network (DNN). EVGMM is a many-to-many conversion model that constructs the entire speaker space (called eigenspace) by analyzing mean vectors of Gaussian mixture models and it is used in our method to decompose training speakers' features into their eigenspace components. By using the speaker features and the obtained components as pseudo parallel data, multiple DNNs are trained to realize conversion between them. With these DNNs, features of any target speaker can be represented by a weighted sum of the components. It should be noted that all the processes of our proposal do not require any parallel data. A key technique is to estimate covariance terms of EVGMM with no parallel data. Experiments indicate that individuality scores of the proposed method using no parallel data are comparable enough to those of a baseline system trained with parallel data. Tetsuya Hashimoto, Daisuke Saito, Nobuaki Minematsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | A Study of Objective Measurement of Comprehensibility through Native Speakers' Shadowing of Learners' Utterances
Yusuke Inoue 0004, Suguru Kabashima, Daisuke Saito, Nobuaki Minematsu, Kumi Kanamura, Yutaka Yamauchi |
INTERSPEECH | 3 |
| 2018 | A Comparative Study of Statistical Conversion of Face to Voice Based on Their Subjective Impressions
Yasuhito Ohsugi, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2018 | DNN-Based Scoring of Language Learners' Proficiency Using Learners' Shadowings and Native Listeners' Responsive ShadowingsabstractThis paper investigates DNN-based scoring techniques when they are applied to two tasks related to foreign language education. One is a conventional task, which attempts to predict a language learner's overall proficiency of oral communication. For this aim, learners' shadowing utterances are assessed automatically. The other is a very new and novel task, which attempts to predict intelligibility or comprehensibility of a learner's pronunciation. In this task, native listeners' responsive shadowings are assessed. For both the tasks, similar technical frameworks are tested, where DNN-based phoneme posteriors, posteriogram-based DTW scores, ASR-based accuracies, shadowing latencies, etc are used to train regression models, which aim to predict manually rated scores. Experiments show that, in both the tasks, the correlation between the DNN-based predicted scores and the averaged human scores is higher than or at least comparable to the averaged correlation between the scores of human raters. This fact clearly indicates that our proposed automatic rating module can be introduced to language education as another human rater. Suguru Kabashima, Yusuke Inoue 0004, Daisuke Saito, Nobuaki Minematsu |
SLT | 3 |
| 2018 | Noise Reduction Method for Intra-Body Communication by Using Compensation ElectrodeabstractA noise reduction method using a compensation electrode and a capacitance in intra-body communication (IBC) is described. The problem with IBC is the influence of environmental noise on communication performance. A radiated noise through the human body is a significant problem. We propose a noise reduction method using a compensation electrode connected to a floor cold electrode with a variable capacitance. The radiated noise through the human body is canceled by the noise to the cold electrode by changing the capacitance. An experiment and equivalent circuit simulation were performed for validation and it was confirmed that the noise reduction method is effective for the IBC system. Yutaro Toyoshima, Yoshiki Matsui, Ryota Kato, Kenta Nezu, Mitsuru Shinagawa, Daisuke Saito, Ken Seo, Kyoji Oohashi |
TENCON | 6 |
| 2017 | Parallel-Data-Free Many-to-Many Voice Conversion Based on DNN Integrated with Eigenspace Using a Non-Parallel Speech Corpus
Tetsuya Hashimoto, Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2017 | Use of Global and Acoustic Features Associated with Contextual Factors to Adapt Language Models for Spontaneous Speech Recognition
Shohei Toyama, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2017 | Acoustic-to-Articulatory Mapping Based on Mixture of Probabilistic Canonical Correlation Analysis
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2017 | Automatic Scoring of Shadowing Speech Based on DNN Posteriors and Their DTW
Junwei Yue, Fumiya Shiozawa, Shohei Toyama, Yutaka Yamauchi, Kayoko Ito, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 6 |
| 2016 | Divergence estimation based on deep neural networks and its use for language identificationabstractIn this paper, we propose a method to estimate statistical divergence between probability distributions by a DNN-based discriminative approach and its use for language identification tasks. Since statistical divergence is generally defined as a functional of two probability density functions, these density functions are usually represented in a parametric form. Then, if a mismatch exists between the assumed distribution and its true one, the obtained divergence becomes erroneous. In our proposed method, by using Bayes' theorem, the statistical divergence is estimated by using DNN as discriminative estimation model. In our method, the divergence between two distributions is able to be estimated without assuming a specific form for these distributions. When the amount of data available for estimation is small, however, it becomes intractable to calculate the integral of the divergence function over all the feature space and to train neural networks. To mitigate this problem, two solutions are introduced; a model adaptation method for DNN and a sampling approach for integration. We apply this approach to language identification tasks, where the obtained divergences are used to extract a speech structure. Experimental results show that our approach can improve the performance of language identification by 10.85% relative compared to the conventional approach based on i-vector. Yosuke Kashiwagi, Congying Zhang, Daisuke Saito, Nobuaki Minematsu |
ICASSP | 3 |
| 2016 | Automatic Assessment and Error Detection of Shadowing Speech: Case of English Spoken by Japanese Learners
Shuju Shi, Yosuke Kashiwagi, Shohei Toyama, Junwei Yue, Yutaka Yamauchi, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 6 |
| 2016 | The Voice Conversion Challenge 2016abstractThis paper describes the Voice Conversion Challenge 2016 devised by the authors to better understand different voice conversion (VC) techniques by comparing their performance on a common dataset. The task of the challenge was speaker conversion, i.e., to transform the voice identity of a source speaker into that of a target speaker while preserving the linguistic content. Using a common dataset consisting of 162 utterances for training and 54 utterances for evaluation from each of 5 source and 5 target speakers, 17 groups working in VC around the world developed their own VC systems for every combination of the source and target speakers, i.e., 25 systems in total, and generated voice samples converted by the developed systems. These samples were evaluated in terms of target speaker similarity and naturalness by 200 listeners in a controlled environment. This paper summarizes the design of the challenge, its result, and a future plan to share views about unsolved problems and challenges faced by the current VC techniques. Tomoki Toda, Linghui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu 0001, Junichi Yamagishi |
INTERSPEECH | 3 |
| 2016 | Prediction of the Articulatory Movements of Unseen Phonemes of a Speaker Using the Speech Structure of Another Speaker
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2016 | Voice Conversion Based on Matrix Variate Gaussian Mixture Model Using Multiple Frame Features
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2016 | Speaker Representations for Speaker Adaptation in Multiple Speakers' BLSTM-RNN-Based Speech Synthesis
Yi Zhao 0006, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2016 | Influence of the Programming Environment on Programming EducationabstractAlthough both visual and text environments have been used to teach programming, the most appropriate method for beginners is unknown. Herein we research the most suitable programming environment to introduce programming to beginners using Minecraft to provide different programming learning environments (Visual or Text) via ComputerCraftEdu as an extended function. The learning effects between these two environments are compared using a lecture course. The results show that a visual environment is more suitable to introduce programming to beginners. Daisuke Saito, Hironori Washizaki, Yoshiaki Fukazawa |
ITiCSE | 1 |
| 2016 | Improved prediction of the accent gap between speakers of English for individual-based clustering of World EnglishesabstractThe term of “World Englishes” describes the current state of English and one of their main characteristics is a large diversity of pronunciation, called accents. In our previous studies, we developed several techniques to realize effective clustering and visualization of the diversity. For this aim, the accent gap between two speakers has to be quantified independently of extra-linguistic factors such as age and gender. To realize this, a unique representation of speech, called speech structure, which is theoretically invariant against these factors, was applied to represent pronunciation. In the current study, by controlling the degree of invariance, we attempt to improve accent gap prediction. Two techniques are tested: DNN-based model-free estimation of divergence and multi-stream speech structures. In the former, instead of estimating separability between two speech events based on some model assumptions, DNN-based class posteriors are utilized for estimation. In the latter, by deriving one speech structure for each sub-space of acoustic features, constrained invariance is realized. Our proposals are tested in terms of the correlation between reference accent gaps and the predicted and quantified gaps. Experiments show that the correlation is improved from 0.718 to 0.730. Fumiya Shiozawa, Daisuke Saito, Nobuaki Minematsu |
SLT | 2 |
| 2016 | Anti-Spoofing for Text-Independent Speaker Verification: An Initial Database, Comparison of Countermeasures, and Human PerformanceabstractIn this paper, we present a systematic study of the vulnerability of automatic speaker verification to a diverse range of spoofing attacks. We start with a thorough analysis of the spoofing effects of five speech synthesis and eight voice conversion systems, and the vulnerability of three speaker verification systems under those attacks. We then introduce a number of countermeasures to prevent spoofing attacks from both known and unknown attackers. Known attackers are spoofing systems whose output was used to train the countermeasures, while an unknown attacker is a spoofing system whose output was not available to the countermeasures during training. Finally, we benchmark automatic systems against human performance on both speaker verification and spoofing detection tasks. Zhizheng Wu 0001, Phillip L. De Leon, Cenk Demiroglu, Ali Khodabakhsh 0001, Simon King 0001, Zhen-Hua Ling, Daisuke Saito, Bryan Stewart, Tomoki Toda, Mirjam Wester, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2015 | SAS: A speaker verification spoofing database containing diverse attacksabstractThis paper presents the first version of a speaker verification spoofing and anti-spoofing database, named SAS corpus. The corpus includes nine spoofing techniques, two of which are speech synthesis, and seven are voice conversion. We design two protocols, one for standard speaker verification evaluation, and the other for producing spoofing materials. Hence, they allow the speech synthesis community to produce spoofing materials incrementally without knowledge of speaker verification spoofing and anti-spoofing. To provide a set of preliminary results, we conducted speaker verification experiments using two state-of-the-art systems. Without any anti-spoofing techniques, the two systems are extremely vulnerable to the spoofing attacks implemented in our SAS corpus. Zhizheng Wu 0001, Ali Khodabakhsh 0001, Cenk Demiroglu, Junichi Yamagishi, Daisuke Saito, Tomoki Toda, Simon King 0001 |
ICASSP | 5 |
| 2015 | Statistical acoustic-to-articulatory mapping unified with speaker normalization based on voice conversionabstractThis paper proposes a model of speaker-normalized acoustic-toarticulatory mapping using statistical voice conversion. A mapping function from acoustic parameters to articulatory parameters is usually developed with a single speaker’s parallel data. Hence the constructed mapping model can work appropriately only for this specific speaker, and applying this model to other speakers degrades the performance of acoustic-to-articulatory mapping. In this paper, two models of speaker conversion and acoustic-to-articulatory mapping are implemented using Gaussian Mixture Models (GMM), and by integrating these two models, we propose two methods of speaker-normalized acoustic-to-articulatory mapping. One is concatenating these models sequentially, and the other integrates the two models into a unified model, where acoustic parameters of a speaker can be converted directly to articulatory parameters of another speaker. Experiments show that both methods can improve the mapping accuracy and that the latter method works better than the former method. Especially in the case of velar stop consonants, the mapping accuracy is higher by 0.6 mm. Index Terms: acoustic-to-articulatory mapping, Gaussian mixture model, voice conversion, speaker normalization Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2014 | Improved and robust prediction of pronunciation distance for individual-basis clustering of World Englishes pronunciationabstractEnglish is the only language available for global communication and is used by approximately 1.5 billions of speakers. It is also known to have a large diversity of pronunciation due to the influence of speakers' mother tongue, called accents. Our project aims at creating a global and individual-basis map of English pronunciations to be used in teaching and learning World Englishes (WE) as well as research studies of WE [1, 2]. Creating the map mathematically requires a distance matrix in terms of pronunciation differences among all the speakers considered, and technically requires a method of predicting the pronunciation distance between any pair of the speakers only by using their speech samples. In our previous study [3], we combined invariant pronunciation structure analysis [4, 5, 6, 7] and Support Vector Regression (SVR) to predict the inter-speaker pronunciation distances. In this paper, several techniques are introduced and examined whether they can increase accuracy and robustness of prediction. Experiments show that the correlation between IPA-based reference distances and the predicted distances is increased from 0.805 to 0.903, which is over the correlation of 0.829 that is obtained by using the phoneme-based ground truth distances. Shun Kasahara, S. Kitahara, Nobuaki Minematsu, Han-Ping Shen, Takehiko Makino, Daisuke Saito, K. Hiorse |
ICASSP | 6 |
| 2014 | Semi-supervised noise dictionary adaptation for exemplar-based noise robust speech recognitionabstractThe exemplar-based approaches, which model signals as a sparse linear combination of exemplars of signals, are proved to have state-of-the-art performance in noise robust ASR, especially on low SNRs. However, since both the speech exemplars and noise exemplars are built from training data and are fixed throughout the process of enhancing speech features, the conventional approach is especially weak for unknown types of noise. Therefore, in this paper, we propose a semi-supervised approach which automatically adapt noise exemplars to the target noise, while keeping the speech exemplars fixed. Continuous digits recognition experiments show that this approach is much more robust for unknown noise. The recognition errors are reduced by 36.2%. Yi Luan, Daisuke Saito, Yosuke Kashiwagi, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 2 |
| 2014 | Application of matrix variate Gaussian mixture model to statistical voice conversionabstractThis paper describes a novel approach to construct a mapping function between a given speaker pair using probability density functions (PDF) of matrix variate. In voice conversion studies, two important functions should be realized: 1) precise modeling of both the source and target feature spaces, and 2) construction of a proper transform function between these spaces. Voice conversion based on Gaussian mixture model (GMM) is the de facto standard because of their flexibility and easiness in handling. In GMM-based approaches, a joint vector space of the source and target is first constructed, and the joint PDF of the two vectors is modeled as GMM in the joint vector space. The joint vector approach mainly focuses on precise modeling of the ‘joint’ feature space, and does not always construct a proper transform between two feature spaces. In contrast, the proposed method constructs the joint PDF as GMM in a matrix variate space whose row and column respectively correspond to the two functions, and it has potential to precisely model both the characteristics of the feature spaces and the relation between the source and target spaces. Daisuke Saito, Hidenobu Doi, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2013 | Discriminative piecewise linear transformation based on deep learning for noise robust automatic speech recognitionabstractIn this paper, we propose the use of deep neural networks to expand conventional methods of statistical feature enhancement based on piecewise linear transformation. Stereo-based piecewise linear compensation for environments (SPLICE), which is a powerful statistical approach for feature enhancement, models the probabilistic distribution of input noisy features as a mixture of Gaussians. However, soft assignment of an input vector to divided regions is sometimes done inadequately and the vector comes to go through inadequate conversion. Especially when conversion has to be linear, the conversion performance will be easily degraded. Feature enhancement using neural networks is another powerful approach which can directly model a non-linear relationship between noisy and clean feature spaces. In this case, however, it tends to suffer from over-fitting problems. In this paper, we attempt to mitigate this problem by reducing the number of model parameters to estimate. Our neural network is trained whose output layer is associated with the states in the clean feature space, not in the noisy feature space. This strategy makes the size of the output layer independent of the kind of a given noisy environment. Firstly, we characterize the distribution of clean features as a Gaussian mixture model and then, by using deep neural networks, estimate discriminatively the state in the clean space that an input noisy feature corresponds to. Experimental evaluations using the Aurora 2 dataset demonstrate that our proposed method has the best performance compared to conventional methods. Yosuke Kashiwagi, Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
ASRU | 2 |
| 2013 | Probabilistic speech F0 contour model incorporating statistical vocabulary model of phrase-accent command sequenceabstractWe have previously proposed a generative model of speech F0 contours, based on the discrete-time version of the Fujisaki model (a model of the mechanisim for controlling F0s through laryngeal muscles). One advantage of this model is that it allows us to apply statistical methods to estimate the Fujisakimodel parameters from speech F0 contours. This paper proposes a new generative model of speech F0 contours incorporating a vocabulary model of intonation patterns. A parameter inference algorithm for the present model is derived. We quantitatively evaluated the performance of our parameter inference algorithm. Tatsuma Ishihara, Hirokazu Kameoka, Kota Yoshizato, Daisuke Saito, Shigeki Sagayama |
INTERSPEECH | 4 |
| 2012 | A tandem connectionist model using combination of multi-scale spectro-temporal features for acoustic event detectionabstractAcoustic event detection systems supporting heterogeneous sets of events face the problem of having to characterize them when they have different acoustic properties (transient, stationary, both, etc.), observing this fact even within the acoustic event itself. Moreover, managing large feature vectors with features characterizing different properties of the signal is always difficult. This paper introduces the usage of spectro-temporal fluctuation features in a tandem connectionist approach, modified to generate posterior features separately for each fluctuation scale and then combine the streams to be fed to a classic GMM-HMM model. The experiments explore scale and event wise performance, as well as different stream combination methods, and show that the proposed method outperforms the GMM-HMM baseline as well as recent proposals in the CHIL 2007 evaluation campaign's related acoustic event detection tasks. Miquel Espi, Masakiyo Fujimoto, Daisuke Saito, Nobutaka Ono, Shigeki Sagayama |
ICASSP | 3 |
| 2012 | Effects of Speaker Adaptive Training on Tensor-based Arbitrary Speaker ConversionabstractThis paper introduces speaker adaptive training techniques to tensor-based arbitrary speaker conversion. In voice conversion studies, realization of conversion from/to an arbitrary speaker’s voice is one of the important objectives. For this purpose, eigen-voice conversion (EVC), which is based on an eigenvoice Gaus-sian mixture model (EV-GMM), was proposed. Although the EVC can effectively construct the conversion model for arbi-trary target speakers using only a few utterances, increase of the utterances used to construct the conversion model does not always improve the conversion performance. This is because the EV-GMMmethod has an inherent problem in representation of GMM supervectors. We previously proposed tensor-based speaker space as a solution for this problem, and realized more flexible control of speaker characteristics. In this paper, to aim larger improvement of the performance of VC, speaker adaptive training and tensor-based speaker representation are integrated. The proposed method can construct the flexible and precise con-version model, and experimental results of one-to-many voice conversion demonstrate the effectiveness of the proposed ap-proach. Index Terms: voice conversion, Gaussian mixture model, eigenvoice, Tucker decomposition, speaker adaptive training Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2012 | Hidden Markov Convolutive Mixture Model for Pitch Contour Analysis of SpeechabstractThis paper proposes a stochastic model of speech F0 contours, based on the stochastic formulation of the Fujisaki model. Our motivation for the stochastic formulation is twofold. Firstly, it allows us to derive a well-behaved algorithm for estimating the Fujisaki model parameters from a raw F0 contour. Secondly, it will open the door to incorporating the well-founded F0 contour model into various statistical speech processing problems. We quantitatively evaluated the performance of our method in terms of an Fujisaki-model parameter estimation accuracy using real speech data. Experimental results revealed that our method was superior to a state-of-the-art Fujisaki model parameter extractor. Index Terms: speech F0 contours, statistical model, Fujisaki model, hidden Markov model, EM algorithm Kota Yoshizato, Hirokazu Kameoka, Daisuke Saito, Shigeki Sagayama |
INTERSPEECH | 3 |
| 2012 | Statistical Voice Conversion Based on Noisy Channel ModelabstractThis paper describes a novel framework of voice conversion effectively using both a joint density model and a speaker model. In voice conversion studies, approaches based on the Gaussian mixture model (GMM) with probabilistic densities of joint vectors of a source and a target speakers are widely used to estimate a transform function between both the speakers. However, to achieve sufficient quality, these approaches require a parallel corpus which contains plenty of utterances with the same linguistic content spoken by both the speakers. In addition, the joint density GMM methods often suffer from overtraining effects when the amount of training data is small. To compensate for these problems, we propose a voice conversion framework, which integrates the speaker GMM of the target with the joint density model using a noisy channel model. The proposed method trains the joint density model with a few parallel utterances, and the speaker model with nonparallel data of the target, independently. It can ease the burden on the source speaker. Experiments demonstrate the effectiveness of the proposed method, especially when the amount of the parallel corpus is small. Daisuke Saito, Shinji Watanabe 0001, Atsushi Nakamura, Nobuaki Minematsu |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | High accurate model-integration-based voice conversion using dynamic features and model structure optimizationabstractThis paper combines a parameter generation algorithm and a model optimization approach with the model-integration-based voice con version (MIVC). We have proposed probabilistic integration of a joint density model and a speaker model to mitigate a requirement of the parallel corpus in voice conversion (VC) based on Gaussian Mixture Model (GMM). As well as the other VC methods, MIVC also suffers from the problems; the degradation of the perceptual quality caused by the discontinuity through the parameter trajectory, and the difficulty to optimize the model structure. To solve the problems, this paper proposes a parameter generation algorithm constrained by dynamic features for the first problem and an information criterion including mutual influences between the joint density model and the speaker model for the second problem. Experimental results show that the first approach improved the performance of VC and the second approach appropriately predicted the optimal number of mixtures of the speaker model for our MIVC. Daisuke Saito, Shinji Watanabe 0001, Atsushi Nakamura, Nobuaki Minematsu |
ICASSP | 1 |
| 2011 | Adaptation of Prosody in Speech Synthesis by Changing Command Values of the Generation Process Model of Fundamental Frequency
Keikichi Hirose, Keiko Ochi, Ryusuke Mihara, Hiroya Hashimoto, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 5 |
| 2011 | Gesture Design of Hand-to-Speech Converter Derived from Speech-to-Hand Converter Based on Probabilistic Integration ModelabstractWhen dysarthrics, individuals with speaking disabilities, try to communicate using speech, they often have no choice but to use speech synthesizers which require them to type word sym-bols or sound symbols. Input by this method often makes real-time communication troublesome and dysarthric users struggle to have smooth flowing conversations. In this study, we are developing a novel speech synthesizer where speech is gener-ated through hand motions rather than symbol input. In re-cent years, statistical voice conversion techniques have been proposed based on space mapping between given parallel ut-terances. By applying these methods, a hand space was mapped to a vowel space and a converter from hand motions to vowel transitions was developed. It reported that the proposed method is effective enough to generate the five Japanese vowels. In this paper, we discuss the expansion of this system to conso-nant generation. In order to create the gestures for consonants, a Speech-to-Hand conversion system is firstly developed using parallel data for vowels, in which consonants are not included. Then, we are able to automatically search for candidates for consonant gestures for a Hand-to-Speech system. Index Terms: Dysarthria, speech production, hand motions, media conversion, arrangement of gestures and vowels Aki Kunikoshi, Yu Qiao 0001, Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2011 | One-to-Many Voice Conversion Based on Tensor Representation of Speaker SpaceabstractThis paper describes a novel approach to flexible control of speaker characteristics using tensor representation of speaker space. In voice conversion studies, realization of conversion from/to an arbitrary speaker’s voice is one of the important objectives. For this purpose, eigenvoice conversion (EVC) based on an eigenvoice Gaussian mixture model (EV-GMM) was proposed. In the EVC, similarly to speaker recognition approaches, a speaker space is constructed based on GMM supervectors which are high-dimensional vectors derived by concatenating the mean vectors of each of the speaker GMMs. In the speaker space, each speaker is represented by a small number of weight parameters of eigen-supervectors. In this paper, we revisit construction of the speaker space by introducing the tensor analysis of training data set. In our approach, each speaker is represented as a matrix of which the row and the column respectively correspond to the Gaussian component and the dimension of the mean vector, and the speaker space is derived by the tensor analysis of the set of the matrices. Our approach can solve an inherent problem of supervector representation, and it improves the performance of voice conversion. Experimental results of oneto-many voice conversion demonstrate the effectiveness of the proposed approach. Index Terms: voice conversion, Gaussian mixture model, eigenvoice, tensor analysis, Tucker decomposition Daisuke Saito, Keisuke Yamamoto, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2010 | HMM-based sequence-to-frame mapping for voice conversionabstractVoice conversion can be reduced to a problem to find a transformation function between the corresponding speech sequences of two speakers. Perhaps the most voice conversions methods are GMM-based statistical mapping methods. However, the classical GMM-based mapping is frame-to-frame, and cannot take account of the contextual information existing over a speech sequence. It is well known that HMM yields an efficient method to model the density of a whole speech sequence and has found great successes in speech recognition and synthesis. Inspired by this fact, this paper studies how to use HMM for voice conversion. We derive an HMM-based sequence-to-frame mapping function with statistical analysis. Different from previous HMM-based voice conversion methods that used forced alignment for segmentation and transform frames aligned to a state with its associated linear transformation, our method has a soft mapping function as a weighted summation of linear transformations. The weights are calculated as the HMM posterior probabilities of frames. We also propose and compare two methods to learn the parameters of our mapping functions, namely least square error estimation and maximum likelihood estimation. We carried out experiments to examine the proposed HMM-based method for voice conversion. Yu Qiao 0001, Daisuke Saito, Nobuaki Minematsu |
ICASSP | 2 |
| 2010 | Probabilistic integration of joint density model and speaker model for voice conversionabstractThis paper describes a novel approach to voice conversion using both a joint density model and a speaker model. In voice con-version studies, approaches based on Gaussian Mixture Model (GMM) with probabilistic densities of joint vectors of a source and a target speakers are widely used to estimate a transfor-mation. However, for sufficient quality, they require a parallel corpus which contains plenty of utterances with the same lin-guistic content spoken by both the speakers. In addition, the joint density GMM methods often suffer from over-training ef-fects when the amount of training data is small. To compensate for these problems, we propose a novel approach to integrate the speaker GMM of the target with the joint density model using probabilistic formulation. The proposed method trains the joint density model with a few parallel utterances, and the speaker model with non-parallel data of the target, independently. It eases the burden on the source speaker. Experiments demon-strate the effectiveness of the proposed method, especially when the amount of the parallel corpus is small. Index Terms: voice conversion, joint density model, speaker model, probabilistic unification Daisuke Saito, Shinji Watanabe 0001, Atsushi Nakamura, Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2009 | Optimal event search using a structural cost function - improvement of structure to speech conversionabstractThis paper describes a new and improved method for the frame-work of structure to speech conversion we previously proposed. Most of the speech synthesizers take a phoneme sequence as input and generate speech by converting each of the phonemes into its corresponding sound. In other words, they simulate a human process of reading text out. However, infants usually acquire speech communication ability without text or phoneme sequences. Since their phonemic awareness is very immature, they can hardly decompose an utterance into a sequence of phones or phonemes. As developmental psychology claims, in-fants acquire the holistic sound patterns of words from the utter-ances of their parents, called word Gestalt, and they reproduce them with their vocal tubes. This behavior is called vocal im-itation. In our previous studies, the word Gestalt was defined physically and a method of extracting it from a word utterance was proposed. We already applied the word Gestalt to ASR, CALL, and also speech generation, which we call structure to speech conversion. Unlike reading machines, our framework simulates infants ’ vocal imitation. In this paper, a method for improving our speech generation framework based on a struc-tural cost function is proposed and evaluated. Index Terms: speech synthesis, the structural representation, vocal imitation, a structural cost function Daisuke Saito, Yu Qiao 0001, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2008 | Directional dependency of cepstrum on vocal tract lengthabstractIN this paper, we prove that the direction of cepstrum vectors strongly depends on vocal tract length and that this dependency is represented as rotation in the n dimensional cepstrum space. In speech recognition studies, vocal tract length normalization (VTLN) techniques are widely used to cancel age- and gender-differences. In VTLN, a frequency warping is often carried out and it can be implemented as a linear transformation in a cepstrum space; c = Ac. However, the geometric properties of this transformation matrix A have not been well discussed. In this study, its properties are made clear using n dimensional geometry and it is shown that the matrix rotates any cepstrum vector similarly and apparently. Experimental results using resynthesized speech demonstrate that cepstrum vectors extracted from a speaker of 180 [cm] in height and those from another speaker of 120 [cm] in height are reasonably orthogonal. This result makes clear one of the reasons why children's speech is very difficult for conventional speech recognizers to deal with adequately. Daisuke Saito, Ryo Matsuura, Satoshi Asakawa, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 1 |
| 2008 | Structure to speech conversion - speech generation based on infant-like vocal imitationabstractThis paper proposes a new framework of speech generation by imitating “infants ’ vocal imitation”. Most of the speech synthe-sizers take a phoneme sequence as input and generate speech by converting each of the phonemes into a sound sequentially. In other words, they simulate a human process of reading text out. However, infants usually acquire speech generation abil-ity without text or phoneme sequences. Since their phonemic awareness is very immature, they can hardly decompose a word utterance into a sequence of phones. In this situation, as devel-opmental psychology states, infants acquire the holistic sound pattern of words from the utterances of their parents, called word Gestalt, and they reproduce it with their vocal tubes. This behavior is called vocal imitation. In our previous studies, the word Gestalt was defined physically and a method of extract-ing it from an utterance was proposed and used successfully for ASR and CALL. In this paper, a method of converting the word Gestalt back to speech is proposed and evaluated. Unlike a read-ing machine, our proposal simulates infants ’ vocal imitation. Index Terms: speech synthesis, vocal imitation, word Gestalt, invariant structure, Bhattacharyya distance, searching problem Daisuke Saito, Satoshi Asakawa, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2008 | Decomposition of rotational distortion caused by VTL difference using eigenvalues of its transformation matrixabstractIn speech recognition studies, vocal tract length normalization (VTLN) techniques are widely used to cancel age- and gender-difference. In VTLN, the distortion is often modeled as a lin-ear transform in a cepstrum space; ĉ=Ac. In our previous study, the geometrical properties ofA were discussed and it was shown that the matrix can be approximated as rotation matrix. In this study, a new method of better approximating A is pro-posed. Using eigenvalues ofA, its quasi-rotational distortion is factorized into multiple rotation operations and multiple magni-fication operations. Using this method, the intrinsic ambiguity of the rotation angle used in our previous study is resolved. In-stead, multiple rotation angles are introduced to understand bet-ter what kind of geometrical distortionsA induces to cepstrum vectors. Experiments show the validity of the new method and a new speech feature is also derived by the new method. Index Terms: frequency warping, rotation matrix, vocal tract length, eigenvalue, rotational plane Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |