Hideki Kawahara

dblp:84/3249 · DBLP profile ↗
← Back
81ranked-venue papers
35as first author
7since 2021 · last 2023
0000-0001-9360-5700ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 77 · 35 first-author · 7 since 2021Artificial intelligence and machine learning · 54 · 24 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2023 Corrigendum to Modelling speaker-size discrimination with voiced and unvoiced speech sounds based on the effect of spectral lift, Speech Communication 136 (2022) 23-41
Toshie Matsui, Toshio Irino, Ryo Uemura, Kodai Yamamoto, Hideki Kawahara, Roy D. Patterson
Speech Commun.5
2022 An objective test tool for pitch extractors' response attributes
abstract
We propose an objective measurement method for pitch extractors' responses to frequency-modulated signals.It enables us to evaluate different pitch extractors with unified criteria.The method uses extended time-stretched pulses combined by binary orthogonal sequences.It provides simultaneous measurement results consisting of the linear and the non-linear timeinvariant responses and random and time-varying responses.We tested representative pitch extractors using fundamental frequencies spanning 80 Hz to 800 Hz with 1/48 octave steps and produced more than 2000 modulation frequency response plots.We found that making scientific visualization by animating these plots enables us to understand different pitch extractors' behavior at once.Such efficient and effortless inspection is impossible by inspecting all individual plots.The proposed measurement method with visualization leads to further improvement of the performance of one of the extractors mentioned above.In other words, our procedure turns the specific pitch extractor into the best reliable measuring equipment that is crucial for scientific research.We open-sourced MATLAB codes of the proposed objective measurement method and visualization procedure.
Hideki Kawahara, Kohei Yatabe, Ken-Ichi Sakakibara, Tatsuya Kitamura, Hideki Banno, Masanori Morise
INTERSPEECH1
2022 Perceptual Evaluation of Penetrating Voices through a Semantic Differential Method
Tatsuya Kitamura, Naoki Kunimoto, Hideki Kawahara, Shigeaki Amano
INTERSPEECH3
2022 Modelling speaker-size discrimination with voiced and unvoiced speech sounds based on the effect of spectral lift
Toshie Matsui, Toshio Irino, Ryo Uemura, Kodai Yamamoto, Hideki Kawahara, Roy D. Patterson
Speech Commun.5
2021 Cascaded All-Pass Filters with Randomized Center Frequencies and Phase Polarity for Acoustic and Speech Measurement and Data Augmentation
abstract
We introduce a new member of TSP (Time Stretched Pulse) for acoustic and speech measurement infrastructure, based on a simple all-pass filter and systematic randomization. This new infrastructure fundamentally upgrades our previous measurement procedure, which enables simultaneous measurement of multiple attributes, including non-linear ones without requiring extra filtering nor post-processing. Our new proposal establishes a theoretically solid, flexible, and extensible foundation in acoustic measurement. Moreover, it is general enough to provide versatile research tools for other fields, such as biological signal analysis. We illustrate using acoustic measurements and data augmentation as representative examples among various prospective applications. We open-sourced MATLAB implementation. It consists of an interactive and real-time acoustic tool, MATLAB functions, and supporting materials.
Hideki Kawahara, Kohei Yatabe
ICASSP1
2021 Mixture of Orthogonal Sequences Made from Extended Time-Stretched Pulses Enables Measurement of Involuntary Voice Fundamental Frequency Response to Pitch Perturbation
abstract
Auditory feedback plays an essential role in the regulation of the fundamental frequency of voiced sounds. The fundamental frequency also responds to auditory stimulation other than the speaker's voice. We propose to use this response of the fundamental frequency of sustained vowels to frequency-modulated test signals for investigating involuntary control of voice pitch. This involuntary response is difficult to identify and isolate by the conventional paradigm, which uses step-shaped pitch perturbation. We recently developed a versatile measurement method using a mixture of orthogonal sequences made from a set of extended time-stretched pulses (TSP). In this article, we extended our approach and designed a set of test signals using the mixture to modulate the fundamental frequency of artificial signals. For testing the response, the experimenter presents the modulated signal aurally while the subject is voicing sustained vowels. We developed a tool for conducting this test quickly and interactively. We make the tool available as an open-source and also provide executable GUI-based applications. Preliminary tests revealed that the proposed method consistently provides compensatory responses with about 100 ms latency, representing involuntary control. Finally, we discuss future applications of the proposed method for objective and non-invasive auditory response measurements.
Hideki Kawahara, Toshie Matsui, Kohei Yatabe, Ken-Ichi Sakakibara, Minoru Tsuzaki, Masanori Morise, Toshio Irino
Interspeech1
2021 Interactive and Real-Time Acoustic Measurement Tools for Speech Data Acquisition and Presentation: Application of an Extended Member of Time Stretched Pulses
Hideki Kawahara, Kohei Yatabe, Ken-Ichi Sakakibara, Mitsunori Mizumachi, Masanori Morise, Hideki Banno, Toshio Irino
Interspeech1
2019 Investigating the Physiological and Acoustic Contrasts Between Choral and Operatic Singing
Hiroko Terasawa, Kenta Wakasa, Hideki Kawahara, Ken-Ichi Sakakibara
INTERSPEECH3
2018 Frequency Domain Variants of Velvet Noise and Their Application to Speech Processing and Synthesis
Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda, Toshio Irino
INTERSPEECH1
2017 A Modulation Property of Time-Frequency Derivatives of Filtered Phase and its Application to Aperiodicity and fo Estimation
abstract
We introduce a simple and linear SNR (strictly speaking, periodic to random power ratio) estimator (0dB to 80dB without additional calibration/linearization) for providing reliable descriptions of aperiodicity in speech corpus. The main idea of this method is to estimate the background random noise level without directly extracting the background noise. The proposed method is applicable to a wide variety of time windowing functions with very low sidelobe levels. The estimate combines the frequency derivative and the time-frequency derivative of the mapping from filter center frequency to the output instantaneous frequency. This procedure can replace the periodicity detection and aperiodicity estimation subsystems of recently introduced open source vocoder, YANG vocoder. Source code of MATLAB implementation of this method will also be open sourced.
Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda
INTERSPEECH1
2017 A New Cosine Series Antialiasing Function and its Application to Aliasing-Free Glottal Source Models for Speech and Singing Synthesis
abstract
We formulated and implemented a procedure to generate aliasing-free excitation source signals. It uses a new antialiasing filter in the continuous time domain followed by an IIR digital filter for response equalization. We introduced a cosine-series-based general design procedure for the new antialiasing function. We applied this new procedure to implement the antialiased Fujisaki-Ljungqvist model. We also applied it to revise our previous implementation of the antialiased Fant-Liljencrants model. A combination of these signals and a lattice implementation of the time varying vocal tract model provides a reliable and flexible basis to test fo extractors and source aperiodicity analysis methods. MATLAB implementations of these antialiased excitation source models are available as part of our open source tools for speech science.
Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda, Toshio Irino
INTERSPEECH1
2017 The Effect of Spectral Tilt on Size Discrimination of Voiced Speech Sounds
abstract
A number of studies, with either voiced or unvoiced speech, have demonstrated that a speaker's geometric mean formant frequency (MFF) has a large effect on the perception of the speaker's size, as would be expected.One study with unvoiced speech showed that lifting the slope of the speech spectrum by 6 dB/octave also led to a reduction in the perceived size of the speaker.This paper reports an analogous experiment to determine whether lifting the slope of the speech spectrum by 6 dB/octave affects the perception of speaker size with voiced speech (words).The results showed that voiced speech with high-frequency enhancement was perceived to arise from smaller speakers.On average, the point of subjective equality in MFF discrimination was reduced by about 5%.However, there were large individual differences; some listeners were effectively insensitive to spectral enhancement of 6 dB/octave; others showed a consistent effect of the same enhancement.The results suggest that models of speaker size perception will need to include a listener specific parameter for the effect of spectral slope.
Toshie Matsui, Toshio Irino, Kodai Yamamoto, Hideki Kawahara, Roy D. Patterson
INTERSPEECH4
2016 SparkNG: Interactive MATLAB Tools for Introduction to Speech Production, Perception and Processing Fundamentals and Application of the Aliasing-Free L-F Model Component
Hideki Kawahara
INTERSPEECH1
2016 TUSK: A Framework for Overviewing the Performance of F0 Estimators
Masanori Morise, Hideki Kawahara
INTERSPEECH2
2015 How the slope of the speech spectrum affects the perception of speaker size
Kodai Yamamoto, Toshio Irino, Ryuichi Nisimura, Hideki Kawahara, Roy D. Patterson
INTERSPEECH4
2014 Vocal tract length estimation based on vowels using a database consisting of 385 speakers and a database with MRI-based vocal tract shape information
Hideki Kawahara, Tatsuya Kitamura, Hironori Takemoto, Ryuichi Nisimura, Toshio Irino
INTERSPEECH1
2014 Excitation source analysis for high-quality speech manipulation systems based on an interference-free representation of group delay with minimum phase response compensation
abstract
A group delay-based excitation source analysis and design method is introduced for extension of TANDEM-STRAIGHT, a speech analysis, modification and synthesis system. This introduction makes all components of the system be based on interference-free representations. They are power spectrum, instantaneous frequency and group delay representations. This unification has potential to solve the major weak point of VOCODER architecture for high-quality speech manipulation applications.
Hideki Kawahara, Masanori Morise, Tomoki Toda, Hideki Banno, Ryuichi Nisimura, Toshio Irino
INTERSPEECH1
2013 Higher order waveform symmetry measure and its application to periodicity detectors for speech and singing with fine temporal resolution
abstract
Another simple and high-speed F0 extractor with high temporal resolution based on our previous proposal has been developed by adding a higher-order symmetry measure. This extension made the proposed method significantly more robust than the previous one. The proposed method is a detector of the lowest prominent sinusoidal component. It can use several F0 refinement procedures when the signal is the sum of harmonic sinusoidal components. The refinement procedure presented here is based on a stable representation of instantaneous frequency of periodic signals. The whole procedure implemented by Matlab runs faster than realtime on usual PCs for 44,100 Hz sampled sounds. Application of the proposed algorithm revealed that rapid temporal modulations in both F0 trajectory and spectral envelope exist typically in expressive voices such as those those used in lively singing performance.
Hideki Kawahara, Masanori Morise, Ryuichi Nisimura, Toshio Irino
ICASSP1
2013 Beyond bandlimited sampling of speech spectral envelope imposed by the harmonic structure of voiced sounds
abstract
A new spectral envelope estimation procedure is proposed to recover details beyond band limitation imposed by the Shannon’s sampling theory when interpreting periodic excitation of voiced sounds as the sampling operation in the frequency domain. The proposed procedure is a hybrid of STRAIGHT, a F0-adaptive spectral envelope estimation and the auto regressive model parameter estimation. Wavelet analyses of these spectral models on the frequency domain enabled objective evaluation of this recovery procedure. The proposed procedure provides better speech quality especially when parameter manipulation is introduced. Index Terms: Speech analysis, envelope spectrum, sampling theory, speech modification, transfer function
Hideki Kawahara, Masanori Morise, Tomoki Toda, Ryuichi Nisimura, Toshio Irino
INTERSPEECH1
2013 Periodicity extraction for voiced sounds with multiple periodicity
Masanori Morise, Hideki Kawahara, Kenji Ozawa
INTERSPEECH2
2013 Controlling "shout" expression in a Japanese POP singing performance: analysis and suppression study
Yuri Nishigaki, Ken-Ichi Sakakibara, Masanori Morise, Ryuichi Nisimura, Toshio Irino, Hideki Kawahara
INTERSPEECH6
2012 Analysis and synthesis of strong vocal expressions: Extension and application of audio texture features to singing voice
abstract
Realistic reconstruction and manipulation of strong vocal expressions found in singing voices is a challenging and exciting topic. A speech analysis, modification and resynthesis framework based on interference-free power spectral and instantaneous frequency representations for periodic sounds is extended for handling such voices. Strong expressions are typically characterized by rapid variations in excitation timing and strength as well as complex structured excitation. Three types of excitation source extractors are revised and introduced to handle them. Preliminary tests successfully replicated strong vocal expressions. Also, additional attribute representations for modifying excitation and spectral information based on audio texture features are briefly discussed.
Hideki Kawahara, Masanori Morise
ICASSP1
2012 Deviation measure of waveform symmetry and its application to high-speed and temporally-fine F0 extraction for vocal sound texture manipulation
Hideki Kawahara, Masanori Morise, Ryuichi Nisimura, Toshio Irino
INTERSPEECH1
2012 Pitch-Scaled Analysis based Residual Reconstruction for Speech Analysis and Synthesis
abstract
The typical problem in LPC-like vocoder is buzzing sound which is mainly due to the simple pulse train or noise excitation model. One way to improve it is to reconstruct the residual obtained from inverse filtering. So a new parametric representation of speech based on pitch-scaled analysis is proposed in this paper. Pitch-scaled analysis is used to extract the periodic spectrum of residual with half pitch period length. Then these periodic spectrums are de-correlated by principal component analysis (PCA) to reduce their dimension. Aperiodic measure is defined as the harmonic-to-noise ratio in the frequency domain where voicing cut-off frequency (VCO) is used to control the smoothness of aperiodicity. Periodic spectrum and aperiodic measure together with F0 are indicated as excitation parameters in the proposed LPC vocoder. Experimental results show that this proposed vocoder can get a mean opinion score (MOS) of 4.1 for a female voice before dimensionality reduction and keep the high-quality property after parameter compression.
Zhengqi Wen, Hideki Kawahara, Jianhua Tao 0001
INTERSPEECH2
2012 Comparison of performance with voiced and whispered speech in word recognition and mean-formant-frequency discrimination
Toshio Irino, Yoshie Aoki, Hideki Kawahara, Roy D. Patterson
Speech Commun.3
2011 An interference-free representation of instantaneous frequency of periodic signals and its application to F0 extraction
abstract
An interference-free representation of the instantaneous frequency of constituent harmonic components of periodic signals is introduced. The power weighted average instantaneous frequency of a band-pass filter yields this property when the effective passband of the filter covers up to two harmonic components and the two windows used in averaging are separated by a half pitch period. The proposed representation eliminates the abrupt changes found in usual instantaneous frequency representations and is applicable to any periodic signals consisting of multiple harmonic components. An F0 extractor of voiced sounds based on this representation is introduced as an example of prospective applications.
Hideki Kawahara, Toshio Irino, Masanori Morise
ICASSP1
2011 Auditory Filterbank Improves Voice Morphing
Erika Okamoto, Toshio Irino, Ryuichi Nisimura, Hideki Kawahara
INTERSPEECH4
2010 High quality voice manipulation method based on the vocal tract area function obtained from sub-band LSP of straight spectrum
abstract
This paper describes a high-quality manipulation method of voice quality base on the vocal tract area function (VTAF) obtained from sub-band LSP of STRAIGHT spectrum. Our research group had developed the manipulation technique of voice quality based on VTAF that can generate natural formant transition. However, it is observed that the generated sound sometimes results in degradation when the input signal has a high sampling frequency. Therefore, we develop a new method that extracts VTAF properly from such input signal. This method firstly divides the input spectral envelope represented by STRAIGHT spectrum into lower and higher frequency bands, secondly extracts the Line spectrum pair (LSP) in each frequency band after spectral flattening that is appropriate for the frequency band, thirdly concatenates a pair of the sub-band LSP, and finally obtains VTAF from PARCOR coefficients converted from the concatenated LSP. A subjective experiment proved that the proposed method is high quality enough.
Ayanori Arakawa, Yoshinori Uchimura, Hideki Banno, Fumitada Itakura, Hideki Kawahara
ICASSP5
2010 High-quality and light-weight voice transformation enabling extrapolation without perceptual and objective breakdown
abstract
A voice transformation method that only relies on vowel information is proposed. The method is based on empirical cumulative distributions of perceptually relevant spectral distances, which are used to design mapping functions from distance to proximity. A set of operators are optimized in the design phase to implement on-the-fly compilation of executable transformations used in the transformation phase. Proximity of the current input parameters to the speaker's own vowel templates is used in this compilation. The proposed method deforms the source speaker's parameter space using a set of monotonic and continuous mapping functions. This smooth and topology-preserving mapping yields high-quality modification of existing speech resources.
Hideki Kawahara, Ryuichi Nisimura, Toshio Irino, Masanori Morise, Toru Takahashi 0004, Hideki Banno
ICASSP1
2010 Simplification and extension of non-periodic excitation source representations for high-quality speech manipulation systems
Hideki Kawahara, Masanori Morise, Toru Takahashi 0004, Hideki Banno, Ryuichi Nisimura, Toshio Irino
INTERSPEECH1
2009 Temporally variable multi-aspect auditory morphing enabling extrapolation without objective and perceptual breakdown
abstract
A generalized framework of auditory morphing based on the speech analysis, modification and resynthesis system STRAIGHT is proposed that enables each morphing rate of representational aspects to be a function of time, including the temporal axis itself. Two types of algorithms were derived: an incremental algorithm for real-time manipulation of morphing rates and a batch processing algorithm for off-line post-production applications. By defining morphing in terms of the derivative of mapping functions in the logarithmic domain, breakdown of morphing resynthesis found in the previous formulation in the case of extrapolations was eliminated. A method to alleviate perceptual defects in extrapolation is also introduced.
Hideki Kawahara, Ryuichi Nisimura, Toshio Irino, Masanori Morise, Toru Takahashi 0004, Hideki Banno
ICASSP1
2009 Observation of empirical cumulative distribution of vowel spectral distances and its application to vowel based voice conversion
abstract
A simple and fast voice conversion method based only on vowel information is proposed. The proposed method relies on empirical distribution of perceptual spectral distances between representative examples of each vowel segment extracted using TANDEM-STRAIGHT spectral envelope estimation procedure [1]. Mapping functions of vowel spectra are designed to preserve vowel space structure defined by the observed empirical distribution while transforming position and orientation of the structure in an abstract vowel spectral space. By introducing physiological constraints in vocal tract shapes and vocal tract length normalization, difficulties in careful frequency alignment between vowel template spectra of the source and the target speakers can be alleviated without significant degradations in converted speech. The proposed method is a framebased instantaneous method and is relevant for real-time processing. Applications of the proposed method in-cross language voice conversion are also discussed. Index Terms: voice conversion, STRAIGHT, vowel structure, line spectral pair, vocal tract length
Hideki Kawahara, Masanori Morise, Toru Takahashi 0004, Hideki Banno, Ryuichi Nisimura, Toshio Irino
INTERSPEECH1
2009 v.morish'09: A Morphing-Based Singing Design Interface for Vocal Melodies
Masanori Morise, Masato Onishi, Hideki Kawahara, Haruhiro Katayose
ICEC3
2008 Tandem-STRAIGHT: A temporally stable power spectral representation for periodic signals and applications to interference-free spectrum, F0, and aperiodicity estimation
abstract
A simple new method for estimating temporally stable power spectra is introduced to provide a unified basis for computing an interference-free spectrum, the fundamental frequency (F0), as well as aperiodicity estimation. F0 adaptive spectral smoothing and cepstral liftering based on consistent sampling theory are employed for interference-free spectral estimation. A perturbation spectrum, calculated from temporally stable power and interference-free spectra, provides the basis for both F0 and aperiodicity estimation. The proposed approach eliminates ad-hoc parameter tuning and the heavy demand on computational power, from which STRAIGHT has suffered in the past.
Hideki Kawahara, Masanori Morise, Toru Takahashi 0004, Ryuichi Nisimura, Toshio Irino, Hideki Banno
ICASSP1
2008 Spectral envelope recovery beyond the nyquist limit for high-quality manipulation of speech sounds
Hideki Kawahara, Masanori Morise, Hideki Banno, Toru Takahashi 0004, Ryuichi Nisimura, Toshio Irino
INTERSPEECH1
2008 Study on manipulation method of voice quality based on the vocal tract area function
Yoshinori Uchimura, Hideki Banno, Fumitada Itakura, Hideki Kawahara
INTERSPEECH4
2008 Speech-to-text input method for web system using JavaScript
abstract
We have developed a speech-to-text input method for web systems. The system is provided as a JavaScript library including an Ajax-like mechanism based on a Java applet, CGI programs, and dynamic HTML documents. It allows users to access voice-enabled web pages without requiring special browsers. Web developers can embed it on their web page by inserting only one line in the header field of an HTML document. This study also aims at observing natural spoken interactions in personal environments. We have succeeded in collecting 4,003 inputs during a period of seven months via our public Japanese ASR server. In order to cover out-of-vocabulary words to cope with some proper nouns, a web page to register new words into the language model are developed. As a result, we could obtain an improvement of 0.8% in the recognition accuracy. With regard to the acoustical conditions, an SNR of 25.3 dB was observed.
Ryuichi Nisimura, Jumpei Miyake, Hideki Kawahara, Toshio Irino
SLT3
2008 Vowel-based frequency alignment function design and recognition-based time alignment for automatic speech morphing
abstract
New design procedures of time-frequency alignment for automatic speech morphing are proposed. The frequency alignment function at a specific frame is represented as a weighted average of vowel alignment functions based on similarity to each vowel. Julian, an open source speech recognition system, was used to design a time alignment function. Objective and subjective tests were conducted to evaluate the proposed method, and test results indicated that the proposed method yields comparable naturalness to the manually morphed samples in terms of time alignment. The results also illustrated that the proposed frequency alignment provides significantly better naturalness than morphed samples without frequency alignment.
Masato Onishi, Toru Takahashi 0004, Toshio Irino, Hideki Kawahara
SLT4
2007 Discrimination and recognition of scaled word sounds
abstract
Smith et al. [2] and Ives et al. [3] demonstrated that humans could extract information about the size of a speaker's vocal tract from speech sounds (vowels and syllables, respectively). We have extended their discrimination and recognition experiments to naturally pronounced words. The Just Noticeable Difference (JND) for size discrimination was between 5.5% and 19% depending on the listener. The smallest JND is comparable to that of the syllable experiments; the average JND is comparable to that of the vowel experiments. The word recognition scores remain above 50% for speaker sizes beyond the normal range for humans. The fact that good performance extends over such a large range of acoustic scales supports Irino and Patterson’s hypothesis [1] that the auditory system segregates size and shape information at an early stage in the processing.
Toshio Irino, Yoshie Aoki, Yoshie Hayashi, Hideki Kawahara, Roy D. Patterson
INTERSPEECH4
2006 Analyzing dialogue data for real-world emotional speech classification
abstract
In order to obtain an understanding of the user’s emotion in human-machine dialogues, an analysis of dialogical utterances in the real world was performed. This work comprises three major steps. (1) The actual conditions of 16 basic emotions were evaluated using Japanese child voices, which were collected through the field test of the public spoken dialogue system. (2) Two factors were derived by a factor analysis. The factors were defined as fundamental psychological factors representing “delightful” and “hateable” emotions. (3) The relationships between the factors and the physical acoustic features were investigated to establish a capability to sense a user’s mental state for the dialogue system. In the experimental discriminations between the delightful and hateable emotions, a correct rate of 98.8% was achieved in classifying child’s utterances by the SVM (Support Vector Machine) with 11 acoustic features. Index Terms: emotional speech, classification, factor analysis, dialogue system, real world
Ryuichi Nisimura, Souji Omae, Hideki Kawahara, Toshio Irino
INTERSPEECH3
2006 Automatic assignment of anchoring points on vowel templates for defining correspondence between time-frequency representations of speech samples
Toru Takahashi 0004, Masashi Nishi, Toshio Irino, Hideki Kawahara
INTERSPEECH4
2006 Speech Segregation Using an Auditory Vocoder With Event-Synchronous Enhancements
abstract
We propose a new method to segregate concurrent speech sounds using an auditory version of a channel vocoder. The auditory representation of sound, referred to as an "auditory image," preserves fine temporal information, unlike conventional window-based processing systems. This makes it possible to segregate speech sources with an event synchronous procedure. Fundamental frequency information is used to estimate the sequence of glottal pulse times for a target speaker, and to repress the glottal events of other speakers. The procedure leads to robust extraction of the target speech and effective segregation even when the signal-to-noise ratio is as low as 0 dB. Moreover, the segregation performance remains high when the speech contains jitter, or when the estimate of the fundamental frequency F0 is inaccurate. This contrasts with conventional comb-filter methods where errors in F0 estimation produce a marked reduction in performance. We compared the new method to a comb-filter method using a cross-correlation measure and perceptual recognition experiments. The results suggest that the new method has the potential to supplant comb-filter and harmonic-selection methods for speech enhancement.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
IEEE Trans. Speech Audio Process.3
2005 Speech intelligibility derived from time-frequency and source smearing
Toshio Irino, Satoru Satou, Shunsuke Nomura, Hideki Banno, Hideki Kawahara
INTERSPEECH5
2005 Nearly defect-free F0 trajectory extraction for expressive speech modifications based on STRAIGHT
Hideki Kawahara, Alain de Cheveigné, Hideki Banno, Toru Takahashi 0004, Toshio Irino
INTERSPEECH1
2005 Voice and emotional expression transformation based on statistics of vowel parameters in an emotional speech database
Toru Takahashi 0004, Takeshi Fujii, Masashi Nishi, Hideki Banno, Toshio Irino, Hideki Kawahara
INTERSPEECH6
2004 Algorithm amalgam: morphing waveform based methods, sinusoidal models and STRAIGHT
abstract
A tool to investigate an important fundamental question in speech processing is proposed aiming to promote research on voice quality and para and non linguistic aspects of speech. The proposed method effectively emulates waveform-based methods, sinusoidal models and the high quality source filter model STRAIGHT The key idea that enables blending these seemingly disjoint algorithms is a group delay based representation of signal excitation. By using a STRAIGHT-based smoothed time-frequency representation that is shared by these three types of speech processing methods, a unified source representation is used to implement the proposed system. Informal listening tests using the proposed system indicated that phase manipulation introduces different timbre, but it does not need to reproduce the exact waveform to reproduce the same timbre. This may suggest that the possibility of further information reduction exists in synthesizing close to natural quality speech.
Hideki Kawahara, Hideki Banno, Toshio Irino, Parham Zolfaghari
ICASSP (1)1
2004 Intelligibility of degraded speech from smeared STRAIGHT spectrum
abstract
Intelligibility of degraded speech sounds has been investigated based on a new signal processing technique using a high-quality vocoder, STRAIGHT. This enables us to manipulate essential speech parameters for vocal tract filtering and glottal excitation. We report that the effect of spectral smearing on the intelligibility of Japanese fourmora words as an initial study. Results reveal that the intelligibility decreases as the degree of smearing increases. We also investigated the relationship between the phonetic and word intelligibilities and found that the word identification score was predicted as the power function of the phonetic score when the power value was about 5. This implies that the mora structure and prosodic information such as F0, timing, and duration also play an important role in speech perception.
Hideki Kawahara, Hideki Banno, Toshio Irino, Jiang Jin
INTERSPEECH1
2004 Procedure "senza vibrato": a key component for morphing singing
abstract
A procedure to remove vibrato from singing voice was proposed to enable auditory morphing between musical performances played under different conditions. Analyses of singing samples in the RWC music database using a speech analysis, modification and synthesis system STRAIGHT provided necessary information to implement “senza vibrato,” the procedure that removes vibrato. A preliminary subjective evaluation for artificially adding and removing vibrato indicated that the proposed procedure effectively control perceived vibrato while preserving naturalness of the original singing.
Hideki Kawahara, Yumi Hirachi, Masanori Morise, Hideki Banno
INTERSPEECH1
2004 A design of audio-visual talker tracking system based on CSP analysis and frame difference in real noisy environments
abstract
It is very important to capture the distant-talking speech with high-quality for voice-controlled systems or teleconferencing systems. A microphone array steering is an ideal candidate for this purpose. However, for the microphone array steering, it is necessary to track the target talker. Conventional talker tracking algorithms with audio signal only (ex. CSP (cross-power spectrum phase) analysis) have a difficulty estimating the target talker direction accurately in higher noisy environments. To overcome this problem, we propose a new target talker tracking algorithm that not only utilize the audio signal, but also utilize the visual signal. The proposed algorithm is based on integration of CSP analysis with audio signal and frame difference with visual signal. As a result of evaluation experiments in a real room, we confirmed that the proposed algorithm could track the target talker accurately than the conventional algorithm.
Nishiura Denda, Takanobu Nishiura, Hideki Kawahara, Toshio Irino
MMSP3
2003 Speech segregation using event synchronous auditory vocoder
abstract
We present a new auditory method to segregate concurrent speech sounds. The system is based on an auditory vocoder developed to resynthesize speech from an auditory Mellin representation using the vocoder STRAIGHT (Kawahara, H. et al., Speech Communication, vol.27, p.187-207, 1999). The quality of the transmitted sound is improved by introducing an event synchronous procedure to estimate glottal pulse times. The auditory representation preserves fine temporal information, unlike conventional window-based processing, which makes it possible to segregate the speech synchronously. The results show that the segregation is good even when the SNR is 0 dB; the extracted target speech was a little distorted but entirely intelligible (like telephone speech), whereas the distracter speech was reduced to a non-speech sound that was not perceptually disturbing. This auditory vocoder has potential for speech enhancement in applications such as hearing aids.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
ICASSP (5)3
2003 Auditory morphing based on an elastic perceptual distance metric in an interference-free time-frequency representation
abstract
An elastic spectral distance measure based on a F0 adaptive pitch synchronous spectral estimation and selective elimination of periodicity interferences, that was developed for a high-quality speech modification procedure STRAIGHT [1], is introduced to provide a basis for auditory morphing. The proposed measure is implemented on a low dimensional piecewise bilinear time-frequency mapping between the target and the original speech representations. A preliminary test results of morphing emotional speech samples indicated that proposed procedure provides perceptually monotonic and high-quality interpolation and extrapolation of CD quality speech samples.
Hideki Kawahara, Hisami Matsui
ICASSP (1)1
2003 Speech enhancement with microphone array and fourier / wavelet spectral subtraction in real noisy environments
Yuki Denda, Takanobu Nishiura, Hideki Kawahara
INTERSPEECH3
2003 Speech segregation based on fundamental event information using an auditory vocoder
abstract
We present a new auditory method to segregate concurrent speech sounds. The system is based on an auditory vocoder developed to resynthesize speech from an auditory Mellin representation using the vocoder STRAIGHT. The auditory representation preserves fine temporal information, unlike conventional window-based processing, and this makes it possible to segregate speech sources with an event synchronous procedure. We developed a method to convert fundamental frequency information to estimate glottal pulse times so as to facilitate robust extraction of the target speech. The results show that the segregation is good even when the SNR is 0 dB; the extracted target speech was a little distorted but entirely intelligible, whereas the distracter speech was reduced to a non-speech sound that was not perceptually disturbing. So, this auditory vocoder has potential for speech enhancement in applications such as hearing aids.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
INTERSPEECH3
2003 Influence of recording equipment on the identification of second language phoneme contrasts
Hiroaki Kato, Masumi Nukinay, Hideki Kawahara, Reiko Akahane-Yamada
INTERSPEECH3
2003 Investigation of emotionally morphed speech perception and its structure using a high quality speech manipulation system
Hisami Matsui, Hideki Kawahara
INTERSPEECH2
2003 Glottal closure instant synchronous sinusoidal model for high quality speech analysis/synthesis
Parham Zolfaghari, Tomohiro Nakatani, Toshio Irino, Hideki Kawahara, Fumitada Itakura
INTERSPEECH4
2002 Auditory VOCODER: Speech resynthesis from an auditory Mellin representation
abstract
We assume that speech rnorphing, noise suppression, and speech segregation would improve if they were more accurately based on human perception. Accordingly, an Auditory VOCODER was developed to resynthesize speech from an auditory Mellin representation used to explain human perception. The Auditory VOCODER has three modules: an Auditory Mellin Image model [9,10], a STRAIGHT VOCODER [2], and a mapping module consisting of warped-frequency cepstral analysis and nonlinear, multivariate regression analysis (MRA). We describe the modules and an evaluation of the system. Informal listening indicates that the sound quality is reasonable.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
ICASSP3
2002 On F0 trajectory optimization for very high-quality speech manipulation
abstract
cote interne IRCAM: Kawahara02a
Hideki Kawahara, Parham Zolfaghari, Alain de Cheveigné
INTERSPEECH1
2001 Comparative evaluation of F0 estimation algorithms
abstract
This paper reports the comparative evaluation of several speech F0 evaluation algorithms over a wide database of laryngographlabeled speech. Included are several classic algorithms that are available in software on the net, as well as two new algorithms that offer greatly reduced error rates. Particular attention is given to the methodology of evaluation.
Alain de Cheveigné, Hideki Kawahara
INTERSPEECH2
2001 Systematic F0 glitches around nasal-vowel transitions
abstract
High-resolution F0 analysis using a speech database with simultaneously recorded EGG (Electroglottogram) signals indicated that there are systematic F0 glitches around nasal-vowel transitions. The durations of the glitches are 10 to 20 ms and they introduce 5 to 10 Hz F0 shifts. A detailed series of analyses of these glitches indicated that the major contributing factor of these glitches is sudden changes of group delay values of the vocal tract transfer function in the vicinity of the fundamental frequency at nasal-vowel transitions. It is also suggested that the Doppler effects due to apparent changes of vocal tract length are marginal, even if they exist. Finally, issues in evaluating high resolution F0 extraction algorithms and applications to high quality speech manipulation methods are discussed.
Hideki Kawahara, Parham Zolfaghari
INTERSPEECH1
2000 Robust fundamental frequency estimation using instantaneous frequencies of harmonic components
abstract
This paper proposes a noise-tolerant method for fundamental frequency (F0) extraction. This method includes several new ideas, including the estimation of the instantaneous frequencies of the higher harmonic components, and the design of an adaptive weighting function based on a bandwidth equation that combines the F0 information in the harmonic components. To evaluate the proposed method, we constructed a relatively large database of simultaneous recordings of speech waveforms and EGG (Electro Glotto Graphy). The database consists of 30 sentences pronounced by 14 male and 14 female normal subjects, i.e., 840 sentences in total. The duration of the sound is about 35 minutes including about 20 minutes of voicing. The experiments were performed with additive noise for four pitch extraction methods, i.e., the proposed method, the original TEMPO, an improved cepstrum method, and a common F0 extraction program in ESPS. The results were as follows: 1) the proposed method is always better than any of the other methods when the SNR is greater than about 2 dB; 2) for high SNR values (> 15 dB), the correct rates of the proposed method and the original TEMPO are about 95% and much better than the improved cepstrum method (92%) and the ESPS function (89%); and 3) all of the methods degrade to less than 62% when the SNR is 0 dB. As a result, the proposed method improves the performance for low SNR values and also maintains high accuracy inherent from the original TEMPO for high SNR values.
Yoshinori Atake, Toshio Irino, Hideki Kawahara, Jinlin Lu, Satoshi Nakamura 0001, Kiyohiro Shikano
INTERSPEECH3
2000 Accurate vocal event detection method based on a fixed-point analysis of mapping from time to weighted average group delay
abstract
A new procedure for event detection and characterization is proposed based on group delay and fixed point analysis. This method enables the detection of precise timing and spread of speech events such as a vocal fold closure. A mapping from the center of a Gaussian time window to the mean time provides event locations as its fixed points. Refining these initial estimates using minimum phase group delay functions derived from the amplitude spectra provides accurate estimates of event locations and durations of excitations of each event. The proposed algorithm was tested using synthetic speech samples and natural speech database of simultaneously recorded sound waveforms and EGG signals. These tests revealed that the proposed method provides estimates of vocal fold closure instants with timing accuracy within 60 µ st o 210µs standard deviations. This algorithm is implemented to be suitable for real-time operation by making extensive use of FFTs without introducing any iterative procedures. It is potentially a very powerful tool for speech diagnosis and construction of very high quality speech manipulation systems.
Hideki Kawahara, Yoshinori Atake, Parham Zolfaghari
INTERSPEECH1
2000 Investigation of analysis and synthesis parameters of straight by subjective evaluation
abstract
The goal of this paper is to locate and understand the fine fundamental problems that exist in the representation of speech sounds by a very high quality speech analysis/synthesis engine namely STRAIGHT. The approach followed here is the evaluation of this system using subjective measures. We use the diagnostic rhyme test (DRT) to evaluate the intelligibility of speech analysed and synthesised by this system for various analysis frame-rates. Consequently we catagorise the fine problems and suggest possible improvements. The results from the DRT have indicated that STRAIGHT can produce speech with an average DRT score of 95 between 1-5 ms analysis frame-rate. In addition, a set of subjective quality measures using MOS and MNRU tests have been conducted. These tests have been carried out for three different versions of the STRAIGHT system: versions 17, 23 and 30. The DRT has been carried out using version 23 only. Based on the subjective evaluation results, a discussion of possible improvements to the STRAIGHT system is given.
Parham Zolfaghari, Yoshinori Atake, Kiyohiro Shikano, Hideki Kawahara
INTERSPEECH4
2000 A sinusoidal model based on frequency-to-instantaneous frequency mapping
Parham Zolfaghari, Hideki Kawahara
INTERSPEECH2
1999 Fixed point analysis of frequency to instantaneous frequency mapping for accurate estimation of F0 and periodicity
Hideki Kawahara, Haruhiro Katayose, Alain de Cheveigné, Roy D. Patterson
EUROSPEECH1
1999 Multiple period estimation and pitch perception model
Alain de Cheveigné, Hideki Kawahara
Speech Commun.2
1999 Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based F0 extraction: Possible role of a repetitive structure in sounds
Hideki Kawahara, Ikuyo Masuda-Katsuse, Alain de Cheveigné
Speech Commun.1
1999 Dynamic sound stream formation based on continuity of spectral change
Ikuyo Masuda-Katsuse, Hideki Kawahara
Speech Commun.2
1998 Efficient representation of short-time phase based on group delay
abstract
An efficient representation of short-time phase characteristics of speech sounds is proposed, based on findings which suggest the perceptual importance of phase characteristics. Subjective tests indicated that the synthesized speech sounds by the proposed method are indistinguishable from the original speech sounds with a moderate data compression. The proposed representation uses lower-order coefficients of the inverse Fourier transform of the group delay of speech. It also alleviates the voiced/unvoiced decision, which is an indispensable part in conventional speech coding algorithms. These features make our method potentially very useful in many applications like speech morphing.
Hideki Banno, Jinlin Lu, Satoshi Nakamura 0001, Kiyohiro Shikano, Hideki Kawahara
ICASSP5
1998 Brain Creators: Japanese Initiative to Create Computational Models of Brain Functions
Yasuji Sawada, Hideki Kawahara
ICONIP2
1998 Computer-based second language production training by using spectrographic representation and HMM-based speech recognition scores
abstract
How can we provide feedback to second language (L2) learners about the goodness of their productions in an automatic way? In this paper, we introduce our attempts to provide effective feedback when we train native speakers of Japanese to produce English /r/ and /l/. First, we adopted spectrographic representation overlayed with formant frequencies as feedback. Second, we investigated the correlation between human judgments of L2 production quality and acoustic scores produced by an HMM-based speech recognition system. We also adopted the HMM-based scores as feedback in the production training. Evaluation of the preand post-training productions by human judges showed that production abilities of the trainees improved in both training groups, suggesting that both spectrographic representation and HMM-based scores were useful and meaningful as feedback. These results are discussed in the context of optimizing L2 speech training.
Reiko Akahane-Yamada, Erik McDermott, Takahiro Adachi, Hideki Kawahara, John S. Pruitt
ICSLP4
1998 An instantaneous-frequency-based pitch extraction method for high-quality speech transformation: revised TEMPO in the STRAIGHT-suite
Hideki Kawahara, Alain de Cheveigné, Roy D. Patterson
ICSLP1
1998 An application of the Bayesian time series model and statistical system analysis for F0 control
Hiroko Kato, Hideki Kawahara
Speech Commun.2
1997 Speech representation and transformation using adaptive interpolation of weighted spectrum: vocoder revisited
abstract
A simple new procedure called STRAIGHT (speech transformation and representation using adaptive interpolation of weighted spectrum) has been developed. STRAIGHT uses pitch-adaptive spectral analysis combined with a surface reconstruction method in the time-frequency region, and an excitation source design based on phase manipulation. It preserves the bilinear surface in the time-frequency region and allows for over 600% manipulation of such speech parameters as pitch, vocal tract length, and speaking rate, without further degradation due to the parameter manipulation.
Hideki Kawahara
ICASSP1
1996 A neural matrix model for active tracking of frequency-modulated tones
Kiyoaki Aikawa, Hideki Kawahara, Minoru Tsuzaki
ICSLP2
1996 Effects of auditory feedback on F0 trajectory generation
abstract
In this paper, a method is proposed to evaluate contributions of auditory feedback to speech F0 trajectory generation. This method is based on data obtained in a series of new auditory feedback experiments (TAF: transformed auditory feedback) in which quantitative measurements were taken of interactions between speech perception and production under natural speech conditions. Experimental results revealed that the effects on power spectra of F0 trajectories vary among subjects and that the maximum magnitude exceeds 10 dB.
Hideki Kawahara, Hiroko Kato, J. C. Williams
ICSLP1
1994 Effects of natural auditory feedback on fundamental frequency control
Hideki Kawahara
ICSLP1
1993 A dynamic cepstrum incorporating time-frequency masking and its application to continuous speech recognition
Kiyoaki Aikawa, Harald Singer, Hideki Kawahara, Yoh'ichi Tohkura
ICASSP (2)3
1992 Signal reconstruction from modified wavelet transform-An application to auditory signal processing
abstract
A novel method of signal reconstruction from a modified auditory representation is presented. This consists of three parts: (1) an algorithm to reconstruct a signal from its modified wavelet transform with a general wavelet; (2) obtaining an auditory representation using an auditory wavelet transform whose analyzing wavelet is the impulse response of an auditory peripheral model; and (3) estimating the reconstruction algorithm both with and without data reduction. An example of its application to the time-scale modification of speech is presented. High-quality speech successfully generated by time-scale modification shows that the reconstruction method is suitable for various applications as well as making experimental auditory stimuli.>
Toshio Irino, Hideki Kawahara
ICASSP2
1990 A Method for Designing Neural Networks Using Nonlinear Multivariate Analysis: Application to Speaker-Independent Vowel Recognition
abstract
A nonlinear multiple logistic model and multiple regression analysis are described as a method for determining the weights for two-layer networks and are compared to error backpropagation. We also provide a method for constructing a three-layer network whose semilinear middle units are primarily provided to discriminate two categories. Experimental results on speaker-independent vowel recognition show that both multivariate methods provide stable weights with fewer iterations than backpropagation training started with random initial weights, but with slightly inferior performance. Backpropagation training with initial weights determined by a multiple logistic model after introduction of data distribution information gives a recognition rate of 98.2%, which is significantly better than average backpropagation with random initial weights.
Toshio Irino, Hideki Kawahara
Neural Comput.2
1988 Vowel-feature extraction from cochlear vibration using neural networks
Toshio Irino, Hideki Kawahara
Neural Networks2