Toshio Irino

dblp:72/3903 · DBLP profile ↗
← Back
71ranked-venue papers
24as first author
9since 2021 · last 2025
0000-0002-7691-4189ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 65 · 19 first-author · 9 since 2021Artificial intelligence and machine learning · 47 · 15 first-author · 6 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Predicting speech intelligibility in older adults for speech enhancement using the Gammachirp Envelope Similarity Index, GESI
Ayako Yamamoto, Fuki Miyazaki, Toshio Irino
Speech Commun.3
2024 Signal processing algorithm effective for sound quality of hearing loss simulators
Toshio Irino, Shintaro Doan, Minami Ishikawa
INTERSPEECH1
2023 Impact of Residual Noise and Artifacts in Speech Enhancement Errors on Intelligibility of Human and Machine
Shoko Araki, Ayako Yamamoto, Tsubasa Ochiai, Kenichi Arai, Atsunori Ogawa, Tomohiro Nakatani, Toshio Irino
INTERSPEECH7
2023 Corrigendum to Modelling speaker-size discrimination with voiced and unvoiced speech sounds based on the effect of spectral lift, Speech Communication 136 (2022) 23-41
Toshie Matsui, Toshio Irino, Ryo Uemura, Kodai Yamamoto, Hideki Kawahara, Roy D. Patterson
Speech Commun.2
2022 Speech intelligibility of simulated hearing loss sounds and its prediction using the Gammachirp Envelope Similarity Index (GESI)
abstract
In the present study, speech intelligibility (SI) experiments were performed using simulated hearing loss (HL) sounds in laboratory and remote environments to clarify the effects of peripheral dysfunction. Noisy speech sounds were processed to simulate the average HL of 70- and 80-year-olds using Wadai Hearing Impairment Simulator (WHIS). These sounds were presented to normal hearing (NH) listeners whose cognitive function could be assumed to be normal. The results showed that the divergence was larger in the remote experiments than in the laboratory ones. However, the remote results could be equalized to the laboratory ones, mostly through data screening using the results of tone pip tests prepared on the experimental web page. In addition, a newly proposed objective intelligibility measure (OIM) called the Gammachirp Envelope Similarity Index (GESI) explained the psychometric functions in the laboratory and remote experiments fairly well. GESI has the potential to explain the SI of HI listeners by properly setting HL parameters.
Toshio Irino, Honoka Tamaru, Ayako Yamamoto
INTERSPEECH1
2022 Modelling speaker-size discrimination with voiced and unvoiced speech sounds based on the effect of spectral lift
Toshie Matsui, Toshio Irino, Ryo Uemura, Kodai Yamamoto, Hideki Kawahara, Roy D. Patterson
Speech Commun.2
2021 Mixture of Orthogonal Sequences Made from Extended Time-Stretched Pulses Enables Measurement of Involuntary Voice Fundamental Frequency Response to Pitch Perturbation
abstract
Auditory feedback plays an essential role in the regulation of the fundamental frequency of voiced sounds. The fundamental frequency also responds to auditory stimulation other than the speaker's voice. We propose to use this response of the fundamental frequency of sustained vowels to frequency-modulated test signals for investigating involuntary control of voice pitch. This involuntary response is difficult to identify and isolate by the conventional paradigm, which uses step-shaped pitch perturbation. We recently developed a versatile measurement method using a mixture of orthogonal sequences made from a set of extended time-stretched pulses (TSP). In this article, we extended our approach and designed a set of test signals using the mixture to modulate the fundamental frequency of artificial signals. For testing the response, the experimenter presents the modulated signal aurally while the subject is voicing sustained vowels. We developed a tool for conducting this test quickly and interactively. We make the tool available as an open-source and also provide executable GUI-based applications. Preliminary tests revealed that the proposed method consistently provides compensatory responses with about 100 ms latency, representing involuntary control. Finally, we discuss future applications of the proposed method for objective and non-invasive auditory response measurements.
Hideki Kawahara, Toshie Matsui, Kohei Yatabe, Ken-Ichi Sakakibara, Minoru Tsuzaki, Masanori Morise, Toshio Irino
Interspeech7
2021 Interactive and Real-Time Acoustic Measurement Tools for Speech Data Acquisition and Presentation: Application of an Extended Member of Time Stretched Pulses
Hideki Kawahara, Kohei Yatabe, Ken-Ichi Sakakibara, Mitsunori Mizumachi, Masanori Morise, Hideki Banno, Toshio Irino
Interspeech7
2021 Comparison of Remote Experiments Using Crowdsourcing and Laboratory Experiments on Speech Intelligibility
abstract
Many subjective experiments have been performed to develop objective speech intelligibility measures, but the novel coronavirus outbreak has made it very difficult to conduct experiments in a laboratory. One solution is to perform remote testing using crowdsourcing; however, because we cannot control the listening conditions, it is unclear whether the results are entirely reliable. In this study, we compared speech intelligibility scores obtained in remote and laboratory experiments. The results showed that the mean and standard deviation (SD) of the remote experiments' speech reception threshold (SRT) were higher than those of the laboratory experiments. However, the variance in the SRTs across the speech-enhancement conditions revealed similarities, implying that remote testing results may be as useful as laboratory experiments to develop an objective measure. We also show that the practice session scores correlate with the SRT values. This is a priori information before performing the main tests and would be useful for data screening to reduce the variability of the SRT distribution.
Ayako Yamamoto, Toshio Irino, Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani
Interspeech2
2020 Predicting Intelligibility of Enhanced Speech Using Posteriors Derived from DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Toshio Irino
INTERSPEECH6
2020 Speech Clarity Improvement by Vocal Self-Training Using a Hearing Impairment Simulator and its Correlation with an Auditory Modulation Index
Toshio Irino, Soichi Higashiyama, Hanako Yoshigi
INTERSPEECH1
2020 GEDI: Gammachirp envelope distortion index for predicting intelligibility of enhanced speech
Katsuhiko Yamamoto, Toshio Irino, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani
Speech Commun.2
2019 Predicting Speech Intelligibility of Enhanced Speech Using Phone Accuracy of DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Katsuhiko Yamamoto, Toshio Irino
INTERSPEECH7
2018 Frequency Domain Variants of Velvet Noise and Their Application to Speech Processing and Synthesis
Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda, Toshio Irino
INTERSPEECH6
2018 Multi-resolution Gammachirp Envelope Distortion Index for Intelligibility Prediction of Noisy Speech
Katsuhiko Yamamoto, Toshio Irino, Narumi Ohashi, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani
INTERSPEECH2
2017 An Auditory Model of Speaker Size Perception for Voiced Speech Sounds
Toshio Irino, Eri Takimoto, Toshie Matsui, Roy D. Patterson
INTERSPEECH1
2017 A New Cosine Series Antialiasing Function and its Application to Aliasing-Free Glottal Source Models for Speech and Singing Synthesis
abstract
We formulated and implemented a procedure to generate aliasing-free excitation source signals. It uses a new antialiasing filter in the continuous time domain followed by an IIR digital filter for response equalization. We introduced a cosine-series-based general design procedure for the new antialiasing function. We applied this new procedure to implement the antialiased Fujisaki-Ljungqvist model. We also applied it to revise our previous implementation of the antialiased Fant-Liljencrants model. A combination of these signals and a lattice implementation of the time varying vocal tract model provides a reliable and flexible basis to test fo extractors and source aperiodicity analysis methods. MATLAB implementations of these antialiased excitation source models are available as part of our open source tools for speech science.
Hideki Kawahara, Ken-Ichi Sakakibara, Masanori Morise, Hideki Banno, Tomoki Toda, Toshio Irino
INTERSPEECH6
2017 The Effect of Spectral Tilt on Size Discrimination of Voiced Speech Sounds
abstract
A number of studies, with either voiced or unvoiced speech, have demonstrated that a speaker's geometric mean formant frequency (MFF) has a large effect on the perception of the speaker's size, as would be expected.One study with unvoiced speech showed that lifting the slope of the speech spectrum by 6 dB/octave also led to a reduction in the perceived size of the speaker.This paper reports an analogous experiment to determine whether lifting the slope of the speech spectrum by 6 dB/octave affects the perception of speaker size with voiced speech (words).The results showed that voiced speech with high-frequency enhancement was perceived to arise from smaller speakers.On average, the point of subjective equality in MFF discrimination was reduced by about 5%.However, there were large individual differences; some listeners were effectively insensitive to spectral enhancement of 6 dB/octave; others showed a consistent effect of the same enhancement.The results suggest that models of speaker size perception will need to include a listener specific parameter for the effect of spectral slope.
Toshie Matsui, Toshio Irino, Kodai Yamamoto, Hideki Kawahara, Roy D. Patterson
INTERSPEECH2
2017 Predicting Speech Intelligibility Using a Gammachirp Envelope Distortion Index Based on the Signal-to-Distortion Ratio
Katsuhiko Yamamoto, Toshio Irino, Toshie Matsui, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani
INTERSPEECH2
2016 Speech Intelligibility Prediction Based on the Envelope Power Spectrum Model with the Dynamic Compressive Gammachirp Auditory Filterbank
Katsuhiko Yamamoto, Toshio Irino, Toshie Matsui, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani
INTERSPEECH2
2015 How the slope of the speech spectrum affects the perception of speaker size
Kodai Yamamoto, Toshio Irino, Ryuichi Nisimura, Hideki Kawahara, Roy D. Patterson
INTERSPEECH2
2014 Vocal tract length estimation based on vowels using a database consisting of 385 speakers and a database with MRI-based vocal tract shape information
Hideki Kawahara, Tatsuya Kitamura, Hironori Takemoto, Ryuichi Nisimura, Toshio Irino
INTERSPEECH5
2014 Excitation source analysis for high-quality speech manipulation systems based on an interference-free representation of group delay with minimum phase response compensation
abstract
A group delay-based excitation source analysis and design method is introduced for extension of TANDEM-STRAIGHT, a speech analysis, modification and synthesis system. This introduction makes all components of the system be based on interference-free representations. They are power spectrum, instantaneous frequency and group delay representations. This unification has potential to solve the major weak point of VOCODER architecture for high-quality speech manipulation applications.
Hideki Kawahara, Masanori Morise, Tomoki Toda, Hideki Banno, Ryuichi Nisimura, Toshio Irino
INTERSPEECH6
2013 Higher order waveform symmetry measure and its application to periodicity detectors for speech and singing with fine temporal resolution
abstract
Another simple and high-speed F0 extractor with high temporal resolution based on our previous proposal has been developed by adding a higher-order symmetry measure. This extension made the proposed method significantly more robust than the previous one. The proposed method is a detector of the lowest prominent sinusoidal component. It can use several F0 refinement procedures when the signal is the sum of harmonic sinusoidal components. The refinement procedure presented here is based on a stable representation of instantaneous frequency of periodic signals. The whole procedure implemented by Matlab runs faster than realtime on usual PCs for 44,100 Hz sampled sounds. Application of the proposed algorithm revealed that rapid temporal modulations in both F0 trajectory and spectral envelope exist typically in expressive voices such as those those used in lively singing performance.
Hideki Kawahara, Masanori Morise, Ryuichi Nisimura, Toshio Irino
ICASSP4
2013 Beyond bandlimited sampling of speech spectral envelope imposed by the harmonic structure of voiced sounds
abstract
A new spectral envelope estimation procedure is proposed to recover details beyond band limitation imposed by the Shannon’s sampling theory when interpreting periodic excitation of voiced sounds as the sampling operation in the frequency domain. The proposed procedure is a hybrid of STRAIGHT, a F0-adaptive spectral envelope estimation and the auto regressive model parameter estimation. Wavelet analyses of these spectral models on the frequency domain enabled objective evaluation of this recovery procedure. The proposed procedure provides better speech quality especially when parameter manipulation is introduced. Index Terms: Speech analysis, envelope spectrum, sampling theory, speech modification, transfer function
Hideki Kawahara, Masanori Morise, Tomoki Toda, Ryuichi Nisimura, Toshio Irino
INTERSPEECH5
2013 Controlling "shout" expression in a Japanese POP singing performance: analysis and suppression study
Yuri Nishigaki, Ken-Ichi Sakakibara, Masanori Morise, Ryuichi Nisimura, Toshio Irino, Hideki Kawahara
INTERSPEECH5
2012 Deviation measure of waveform symmetry and its application to high-speed and temporally-fine F0 extraction for vocal sound texture manipulation
Hideki Kawahara, Masanori Morise, Ryuichi Nisimura, Toshio Irino
INTERSPEECH4
2012 Comparison of performance with voiced and whispered speech in word recognition and mean-formant-frequency discrimination
Toshio Irino, Yoshie Aoki, Hideki Kawahara, Roy D. Patterson
Speech Commun.1
2011 An interference-free representation of instantaneous frequency of periodic signals and its application to F0 extraction
abstract
An interference-free representation of the instantaneous frequency of constituent harmonic components of periodic signals is introduced. The power weighted average instantaneous frequency of a band-pass filter yields this property when the effective passband of the filter covers up to two harmonic components and the two windows used in averaging are separated by a half pitch period. The proposed representation eliminates the abrupt changes found in usual instantaneous frequency representations and is applicable to any periodic signals consisting of multiple harmonic components. An F0 extractor of voiced sounds based on this representation is introduced as an example of prospective applications.
Hideki Kawahara, Toshio Irino, Masanori Morise
ICASSP2
2011 Auditory Filterbank Improves Voice Morphing
Erika Okamoto, Toshio Irino, Ryuichi Nisimura, Hideki Kawahara
INTERSPEECH2
2010 High-quality and light-weight voice transformation enabling extrapolation without perceptual and objective breakdown
abstract
A voice transformation method that only relies on vowel information is proposed. The method is based on empirical cumulative distributions of perceptually relevant spectral distances, which are used to design mapping functions from distance to proximity. A set of operators are optimized in the design phase to implement on-the-fly compilation of executable transformations used in the transformation phase. Proximity of the current input parameters to the speaker's own vowel templates is used in this compilation. The proposed method deforms the source speaker's parameter space using a set of monotonic and continuous mapping functions. This smooth and topology-preserving mapping yields high-quality modification of existing speech resources.
Hideki Kawahara, Ryuichi Nisimura, Toshio Irino, Masanori Morise, Toru Takahashi 0004, Hideki Banno
ICASSP3
2010 Simplification and extension of non-periodic excitation source representations for high-quality speech manipulation systems
Hideki Kawahara, Masanori Morise, Toru Takahashi 0004, Hideki Banno, Ryuichi Nisimura, Toshio Irino
INTERSPEECH6
2010 Auditory speech processing for scale-shift covariance and its evaluation in automatic speech recognition
abstract
The syllables of speech contain information about the vocal tract length (VTL) of the speaker as well as the phonetic message. Ideally, the pre-processor used for automatic speech recognition (ASR) should segregate the phonetic message from the VTL information. This paper describes a method to calculate VTL-invariant auditory feature vectors from speech, using a method in which the message and the VTL are segregated. Spectra produced by an auditory filterbank are summarized by a Gaussian mixture model (GMM) to produce a low-dimensional feature vector. These features are evaluated for robustness in comparison with conventional mel-frequency cepstral coefficients (MFCCs) using a hidden-Markov-model (HMM) recognizer. A dynamic, compressive gammachirp (dcGC) auditory filterbank is also introduced. The dcGC provides a level-dependent spectral analysis, with near instantaneous compression, and two-tone suppression.
Roy D. Patterson, Thomas C. Walters, Jessica Monaghan, Christian Feldbauer, Toshio Irino
ISCAS5
2009 Temporally variable multi-aspect auditory morphing enabling extrapolation without objective and perceptual breakdown
abstract
A generalized framework of auditory morphing based on the speech analysis, modification and resynthesis system STRAIGHT is proposed that enables each morphing rate of representational aspects to be a function of time, including the temporal axis itself. Two types of algorithms were derived: an incremental algorithm for real-time manipulation of morphing rates and a batch processing algorithm for off-line post-production applications. By defining morphing in terms of the derivative of mapping functions in the logarithmic domain, breakdown of morphing resynthesis found in the previous formulation in the case of extrapolations was eliminated. A method to alleviate perceptual defects in extrapolation is also introduced.
Hideki Kawahara, Ryuichi Nisimura, Toshio Irino, Masanori Morise, Toru Takahashi 0004, Hideki Banno
ICASSP3
2009 Observation of empirical cumulative distribution of vowel spectral distances and its application to vowel based voice conversion
abstract
A simple and fast voice conversion method based only on vowel information is proposed. The proposed method relies on empirical distribution of perceptual spectral distances between representative examples of each vowel segment extracted using TANDEM-STRAIGHT spectral envelope estimation procedure [1]. Mapping functions of vowel spectra are designed to preserve vowel space structure defined by the observed empirical distribution while transforming position and orientation of the structure in an abstract vowel spectral space. By introducing physiological constraints in vocal tract shapes and vocal tract length normalization, difficulties in careful frequency alignment between vowel template spectra of the source and the target speakers can be alleviated without significant degradations in converted speech. The proposed method is a framebased instantaneous method and is relevant for real-time processing. Applications of the proposed method in-cross language voice conversion are also discussed. Index Terms: voice conversion, STRAIGHT, vowel structure, line spectral pair, vocal tract length
Hideki Kawahara, Masanori Morise, Toru Takahashi 0004, Hideki Banno, Ryuichi Nisimura, Toshio Irino
INTERSPEECH6
2009 Influences of vowel duration on speaker-size estimation and discrimination
abstract
Several experimental studies have shown that the human auditory system has a mechanism for extracting speaker-size information, using sufficiently long sounds. This paper investigated influence of vowel duration on the processing for size extraction using short vowels. In a size estimation experiment, listeners subjectively estimated the size (height) of the speaker for isolated vowels. The results showed that listeners’ perception of speaker size was highly correlated with the factor of vocaltract length in all the tested durations (from 16 ms to 256 ms). In a size discrimination experiment, listeners were presented with two vowels scaled the vocal-tract length and were asked which vowel was perceived to be spoken by a smaller speaker. The results showed that the just-noticeable differences (JNDs) in speaker size were almost the same for the durations longer than 32 ms. However, the JNDs rose considerably for 16-ms duration. These observations of the experiments suggest that the auditory system can extract speaker-size information even for 16-ms vowels although the precision of size extraction would deteriorate when the duration becomes less than 32 ms. Index Terms: auditory processing for speaker-size extraction, duration, size discrimination, subjective size estimation
Chihiro Takeshima, Minoru Tsuzaki, Toshio Irino
INTERSPEECH3
2008 Tandem-STRAIGHT: A temporally stable power spectral representation for periodic signals and applications to interference-free spectrum, F0, and aperiodicity estimation
abstract
A simple new method for estimating temporally stable power spectra is introduced to provide a unified basis for computing an interference-free spectrum, the fundamental frequency (F0), as well as aperiodicity estimation. F0 adaptive spectral smoothing and cepstral liftering based on consistent sampling theory are employed for interference-free spectral estimation. A perturbation spectrum, calculated from temporally stable power and interference-free spectra, provides the basis for both F0 and aperiodicity estimation. The proposed approach eliminates ad-hoc parameter tuning and the heavy demand on computational power, from which STRAIGHT has suffered in the past.
Hideki Kawahara, Masanori Morise, Toru Takahashi 0004, Ryuichi Nisimura, Toshio Irino, Hideki Banno
ICASSP5
2008 Spectral envelope recovery beyond the nyquist limit for high-quality manipulation of speech sounds
Hideki Kawahara, Masanori Morise, Hideki Banno, Toru Takahashi 0004, Ryuichi Nisimura, Toshio Irino
INTERSPEECH6
2008 Speech-to-text input method for web system using JavaScript
abstract
We have developed a speech-to-text input method for web systems. The system is provided as a JavaScript library including an Ajax-like mechanism based on a Java applet, CGI programs, and dynamic HTML documents. It allows users to access voice-enabled web pages without requiring special browsers. Web developers can embed it on their web page by inserting only one line in the header field of an HTML document. This study also aims at observing natural spoken interactions in personal environments. We have succeeded in collecting 4,003 inputs during a period of seven months via our public Japanese ASR server. In order to cover out-of-vocabulary words to cope with some proper nouns, a web page to register new words into the language model are developed. As a result, we could obtain an improvement of 0.8% in the recognition accuracy. With regard to the acoustical conditions, an SNR of 25.3 dB was observed.
Ryuichi Nisimura, Jumpei Miyake, Hideki Kawahara, Toshio Irino
SLT4
2008 Vowel-based frequency alignment function design and recognition-based time alignment for automatic speech morphing
abstract
New design procedures of time-frequency alignment for automatic speech morphing are proposed. The frequency alignment function at a specific frame is represented as a weighted average of vowel alignment functions based on similarity to each vowel. Julian, an open source speech recognition system, was used to design a time alignment function. Objective and subjective tests were conducted to evaluate the proposed method, and test results indicated that the proposed method yields comparable naturalness to the manually morphed samples in terms of time alignment. The results also illustrated that the proposed frequency alignment provides significantly better naturalness than morphed samples without frequency alignment.
Masato Onishi, Toru Takahashi 0004, Toshio Irino, Hideki Kawahara
SLT3
2008 A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments
Tomohiro Nakatani, Shigeaki Amano, Toshio Irino, Kentaro Ishizuka, Tadahisa Kondo
Speech Commun.3
2007 Discrimination and recognition of scaled word sounds
abstract
Smith et al. [2] and Ives et al. [3] demonstrated that humans could extract information about the size of a speaker's vocal tract from speech sounds (vowels and syllables, respectively). We have extended their discrimination and recognition experiments to naturally pronounced words. The Just Noticeable Difference (JND) for size discrimination was between 5.5% and 19% depending on the listener. The smallest JND is comparable to that of the syllable experiments; the average JND is comparable to that of the vowel experiments. The word recognition scores remain above 50% for speaker sizes beyond the normal range for humans. The fact that good performance extends over such a large range of acoustic scales supports Irino and Patterson’s hypothesis [1] that the auditory system segregates size and shape information at an early stage in the processing.
Toshio Irino, Yoshie Aoki, Yoshie Hayashi, Hideki Kawahara, Roy D. Patterson
INTERSPEECH1
2006 Dynamic, Compressive Gammachirp Auditory Filterbank for Perceptual Signal Processing
abstract
A gammachirp auditory filter was developed 1) to extend the domain of the gammatone auditory filter, 2) to simulate the changes in filter shape that occur with changes in stimulus level, 3) to explain a large body of simultaneous masking data, 4) to explain the compressive characteristics of the auditory filter system, and 5) to facilitate the development of a nonlinear, analysis/synthesis framework. What remains is to specify the dynamics of how the stimulus level controls the filter parameters. In this paper, we use psychophysical data involving compression to derive the details of the level control circuit for the dynamic version of the cGC (dcGC) filter and filterbank. The dcGC filterbank enhances spectral contrasts and reduces the dynamic range. This property with the analysis/synthesis framework should be useful in various forms of perceptual signal processing
Toshio Irino, Roy D. Patterson
ICASSP (5)1
2006 Analyzing dialogue data for real-world emotional speech classification
abstract
In order to obtain an understanding of the user’s emotion in human-machine dialogues, an analysis of dialogical utterances in the real world was performed. This work comprises three major steps. (1) The actual conditions of 16 basic emotions were evaluated using Japanese child voices, which were collected through the field test of the public spoken dialogue system. (2) Two factors were derived by a factor analysis. The factors were defined as fundamental psychological factors representing “delightful” and “hateable” emotions. (3) The relationships between the factors and the physical acoustic features were investigated to establish a capability to sense a user’s mental state for the dialogue system. In the experimental discriminations between the delightful and hateable emotions, a correct rate of 98.8% was achieved in classifying child’s utterances by the SVM (Support Vector Machine) with 11 acoustic features. Index Terms: emotional speech, classification, factor analysis, dialogue system, real world
Ryuichi Nisimura, Souji Omae, Hideki Kawahara, Toshio Irino
INTERSPEECH4
2006 Automatic assignment of anchoring points on vowel templates for defining correspondence between time-frequency representations of speech samples
Toru Takahashi 0004, Masashi Nishi, Toshio Irino, Hideki Kawahara
INTERSPEECH3
2006 A Dynamic Compressive Gammachirp Auditory Filterbank
abstract
It is now common to use knowledge about human auditory processing in the development of audio signal processors. Until recently, however, such systems were limited by their linearity. The auditory filter system is known to be level-dependent as evidenced by psychophysical data on masking, compression, and two-tone suppression. However, there were no analysis/synthesis schemes with nonlinear filterbanks. This paper describe18300060s such a scheme based on the compressive gammachirp (cGC) auditory filter. It was developed to extend the gammatone filter concept to accommodate the changes in psychophysical filter shape that are observed to occur with changes in stimulus level in simultaneous, tone-in-noise masking. In models of simultaneous noise masking, the temporal dynamics of the filtering can be ignored. Analysis/synthesis systems, however, are intended for use with speech sounds where the glottal cycle can be long with respect to auditory time constants, and so they require specification of the temporal dynamics of auditory filter. In this paper, we describe a fast-acting level control circuit for the cGC filter and show how psychophysical data involving two-tone suppression and compression can be used to estimate the parameter values for this dynamic version of the cGC filter (referred to as the "dcGC" filter). One important advantage of analysis/synthesis systems with a dcGC filterbank is that they can inherit previously refined signal processing algorithms developed with conventional short-time Fourier transforms (STFTs) and linear filterbanks.
Toshio Irino, Roy D. Patterson
IEEE Trans. Speech Audio Process.1
2006 Speech Segregation Using an Auditory Vocoder With Event-Synchronous Enhancements
abstract
We propose a new method to segregate concurrent speech sounds using an auditory version of a channel vocoder. The auditory representation of sound, referred to as an "auditory image," preserves fine temporal information, unlike conventional window-based processing systems. This makes it possible to segregate speech sources with an event synchronous procedure. Fundamental frequency information is used to estimate the sequence of glottal pulse times for a target speaker, and to repress the glottal events of other speakers. The procedure leads to robust extraction of the target speech and effective segregation even when the signal-to-noise ratio is as low as 0 dB. Moreover, the segregation performance remains high when the speech contains jitter, or when the estimate of the fundamental frequency F0 is inaccurate. This contrasts with conventional comb-filter methods where errors in F0 estimation produce a marked reduction in performance. We compared the new method to a comb-filter method using a cross-correlation measure and perceptual recognition experiments. The results suggest that the new method has the potential to supplant comb-filter and harmonic-selection methods for speech enhancement.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
IEEE Trans. Speech Audio Process.1
2005 Speech intelligibility derived from time-frequency and source smearing
Toshio Irino, Satoru Satou, Shunsuke Nomura, Hideki Banno, Hideki Kawahara
INTERSPEECH1
2005 Nearly defect-free F0 trajectory extraction for expressive speech modifications based on STRAIGHT
Hideki Kawahara, Alain de Cheveigné, Hideki Banno, Toru Takahashi 0004, Toshio Irino
INTERSPEECH5
2005 Voice and emotional expression transformation based on statistics of vowel parameters in an emotional speech database
Toru Takahashi 0004, Takeshi Fujii, Masashi Nishi, Hideki Banno, Toshio Irino, Hideki Kawahara
INTERSPEECH5
2004 Algorithm amalgam: morphing waveform based methods, sinusoidal models and STRAIGHT
abstract
A tool to investigate an important fundamental question in speech processing is proposed aiming to promote research on voice quality and para and non linguistic aspects of speech. The proposed method effectively emulates waveform-based methods, sinusoidal models and the high quality source filter model STRAIGHT The key idea that enables blending these seemingly disjoint algorithms is a group delay based representation of signal excitation. By using a STRAIGHT-based smoothed time-frequency representation that is shared by these three types of speech processing methods, a unified source representation is used to implement the proposed system. Informal listening tests using the proposed system indicated that phase manipulation introduces different timbre, but it does not need to reproduce the exact waveform to reproduce the same timbre. This may suggest that the possibility of further information reduction exists in synthesizing close to natural quality speech.
Hideki Kawahara, Hideki Banno, Toshio Irino, Parham Zolfaghari
ICASSP (1)3
2004 Intelligibility of degraded speech from smeared STRAIGHT spectrum
abstract
Intelligibility of degraded speech sounds has been investigated based on a new signal processing technique using a high-quality vocoder, STRAIGHT. This enables us to manipulate essential speech parameters for vocal tract filtering and glottal excitation. We report that the effect of spectral smearing on the intelligibility of Japanese fourmora words as an initial study. Results reveal that the intelligibility decreases as the degree of smearing increases. We also investigated the relationship between the phonetic and word intelligibilities and found that the word identification score was predicted as the power function of the phonetic score when the power value was about 5. This implies that the mora structure and prosodic information such as F0, timing, and duration also play an important role in speech perception.
Hideki Kawahara, Hideki Banno, Toshio Irino, Jiang Jin
INTERSPEECH3
2004 A design of audio-visual talker tracking system based on CSP analysis and frame difference in real noisy environments
abstract
It is very important to capture the distant-talking speech with high-quality for voice-controlled systems or teleconferencing systems. A microphone array steering is an ideal candidate for this purpose. However, for the microphone array steering, it is necessary to track the target talker. Conventional talker tracking algorithms with audio signal only (ex. CSP (cross-power spectrum phase) analysis) have a difficulty estimating the target talker direction accurately in higher noisy environments. To overcome this problem, we propose a new target talker tracking algorithm that not only utilize the audio signal, but also utilize the visual signal. The proposed algorithm is based on integration of CSP analysis with audio signal and frame difference with visual signal. As a result of evaluation experiments in a real room, we confirmed that the proposed algorithm could track the target talker accurately than the conventional algorithm.
Nishiura Denda, Takanobu Nishiura, Hideki Kawahara, Toshio Irino
MMSP4
2003 Speech segregation using event synchronous auditory vocoder
abstract
We present a new auditory method to segregate concurrent speech sounds. The system is based on an auditory vocoder developed to resynthesize speech from an auditory Mellin representation using the vocoder STRAIGHT (Kawahara, H. et al., Speech Communication, vol.27, p.187-207, 1999). The quality of the transmitted sound is improved by introducing an event synchronous procedure to estimate glottal pulse times. The auditory representation preserves fine temporal information, unlike conventional window-based processing, which makes it possible to segregate the speech synchronously. The results show that the segregation is good even when the SNR is 0 dB; the extracted target speech was a little distorted but entirely intelligible (like telephone speech), whereas the distracter speech was reduced to a non-speech sound that was not perceptually disturbing. This auditory vocoder has potential for speech enhancement in applications such as hearing aids.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
ICASSP (5)1
2003 Speech segregation based on fundamental event information using an auditory vocoder
abstract
We present a new auditory method to segregate concurrent speech sounds. The system is based on an auditory vocoder developed to resynthesize speech from an auditory Mellin representation using the vocoder STRAIGHT. The auditory representation preserves fine temporal information, unlike conventional window-based processing, and this makes it possible to segregate speech sources with an event synchronous procedure. We developed a method to convert fundamental frequency information to estimate glottal pulse times so as to facilitate robust extraction of the target speech. The results show that the segregation is good even when the SNR is 0 dB; the extracted target speech was a little distorted but entirely intelligible, whereas the distracter speech was reduced to a non-speech sound that was not perceptually disturbing. So, this auditory vocoder has potential for speech enhancement in applications such as hearing aids.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
INTERSPEECH1
2003 Dominance spectrum based v/UV classification and f_0 estimation
Tomohiro Nakatani, Toshio Irino, Parham Zolfaghari
INTERSPEECH2
2003 Glottal closure instant synchronous sinusoidal model for high quality speech analysis/synthesis
Parham Zolfaghari, Tomohiro Nakatani, Toshio Irino, Hideki Kawahara, Fumitada Itakura
INTERSPEECH3
2002 Auditory VOCODER: Speech resynthesis from an auditory Mellin representation
abstract
We assume that speech rnorphing, noise suppression, and speech segregation would improve if they were more accurately based on human perception. Accordingly, an Auditory VOCODER was developed to resynthesize speech from an auditory Mellin representation used to explain human perception. The Auditory VOCODER has three modules: an Auditory Mellin Image model [9,10], a STRAIGHT VOCODER [2], and a mapping module consisting of warped-frequency cepstral analysis and nonlinear, multivariate regression analysis (MRA). We describe the modules and an evaluation of the system. Informal listening indicates that the sound quality is reasonable.
Toshio Irino, Roy D. Patterson, Hideki Kawahara
ICASSP1
2002 Evaluation of a speech recognition / generation method based on HMM and straight
abstract
We propose a method for integrating speech recognition and generation within a unified framework. The method consists of STRAIGHT, warped-frequency DCT, and an HMM engine. The warped-frequency DCT i s used to derive a kind of mel-cepstral coefficient from the smoothed spectrum of STRAIGHT, which is known as a high-quality vocoder. This analysis/synthesis method has potential to improve the performance beyond a conventional method using the MFCC derived from the STFT. We evaluated the method by using speakerdependent speech recognition as well as by the perceptual evaluation of sounds generated by HMM text-to-speech. The recognition rate using the coefficients from the warped-DCT of the STRAIGHT spectrum was almost the same as that obtained using conventional MFCCs. The sound quality was sufficiently good for a fundamental system. 1.
Toshio Irino, Yasuhiro Minami, Tomohiro Nakatani, Minoru Tsuzaki, H. Tagawa
INTERSPEECH1
2002 Robust fundamental frequency estimation against background noise and spectral distortion
abstract
This paper presents a new method for robust fundamental frequency (F0) estimation in the presence of background noise and spectral distortion. We define degree of dominance and a dominance spectrum based on instantaneous frequencies. The degree of dominance allows us to evaluate the magnitude of individual harmonic components of speech signals relative to background noise while eliminating the influence of spectral distortion. The fundamental frequency is robustly estimated from reliable harmonic components easily selected from the dominance spectra. Experiments are performed using white and multi-talker background noise with and without spectral distortion produced by a SRAEN filter. Results show that the present method is better than the commonlyused methods in terms of correct F0 rates.
Tomohiro Nakatani, Toshio Irino
INTERSPEECH2
2002 Segregating information about the size and shape of the vocal tract using a time-domain auditory model: The stabilised wavelet-Mellin transform
Toshio Irino, Roy D. Patterson
Speech Commun.1
2000 Robust fundamental frequency estimation using instantaneous frequencies of harmonic components
abstract
This paper proposes a noise-tolerant method for fundamental frequency (F0) extraction. This method includes several new ideas, including the estimation of the instantaneous frequencies of the higher harmonic components, and the design of an adaptive weighting function based on a bandwidth equation that combines the F0 information in the harmonic components. To evaluate the proposed method, we constructed a relatively large database of simultaneous recordings of speech waveforms and EGG (Electro Glotto Graphy). The database consists of 30 sentences pronounced by 14 male and 14 female normal subjects, i.e., 840 sentences in total. The duration of the sound is about 35 minutes including about 20 minutes of voicing. The experiments were performed with additive noise for four pitch extraction methods, i.e., the proposed method, the original TEMPO, an improved cepstrum method, and a common F0 extraction program in ESPS. The results were as follows: 1) the proposed method is always better than any of the other methods when the SNR is greater than about 2 dB; 2) for high SNR values (> 15 dB), the correct rates of the proposed method and the original TEMPO are about 95% and much better than the improved cepstrum method (92%) and the ESPS function (89%); and 3) all of the methods degrade to less than 62% when the SNR is 0 dB. As a result, the proposed method improves the performance for low SNR values and also maintains high accuracy inherent from the original TEMPO for high SNR values.
Yoshinori Atake, Toshio Irino, Hideki Kawahara, Jinlin Lu, Satoshi Nakamura 0001, Kiyohiro Shikano
INTERSPEECH2
1999 Noise suppression using a time-varying, analysis/synthesis gamma chirp filterbank
abstract
Spectral subtraction has been cited most often as a noise suppression method for speech signals in steady background noise, because it is basically a non-parametric method and simple enough to implement for various applications using FFT. It has also been well known, however, that spectral subtraction produces so called "musical noise" in synthetic sounds. Since such musical noise, even at low levels, can often bother humans in speech perception, spectral subtraction has not been very successful in signal processing applications for human listeners. To suppress noise without producing musical noise, an alternative method has been developed using a time-varying, analysis/synthesis gammachirp filterbank; this was initially proposed as an auditory filterbank. The present method achieves about the same SNR improvement as spectral subtraction when using the same information on the non-speech interval. Moreover, the synthetic sounds only contain steady white-like noise at reduced levels when the original noise is white. This method is, therefore, advantageous in various applications for human listeners.
Toshio Irino
ICASSP1
1999 Stabilised wavelet mellin transform: an auditory strategy for normalising sound-source size
Toshio Irino, Roy D. Patterson
EUROSPEECH1
1998 A time-varying, analysis/synthesis auditory filterbank using the gammachirp
abstract
A time-varying, analysis/synthesis auditory filterbank has been developed using a new implementation of the "gammachirp", which has been shown to be an excellent function for the asymmetric, level-dependent auditory filter. The gammachirp filter is shown to be implemented through a combination of a gammatone filter and an IIR asymmetric compensation filter; which largely reduces the computational cost for time-varying filtering. The gammachirp filterbank is designed using a linear gammatone filterbank and a bank of time-varying asymmetric compensation filters controlled by the sound pressure level estimated at the output of the filterbank. Since the inverse filter of the asymmetric compensation filter is always stable, it is possible to resynthesize signals from time-varying, level-dependent auditory representations. The resynthesis error is only determined by the linear analysis/synthesis gammatone filterbank. The proposed filterbank is applicable to various types of signal processing required to model human auditory filtering.
Toshio Irino, Masashi Unoki
ICASSP1
1998 The Gammachirp for Optimal Auditory Filtering
Toshio Irino, Roy D. Patterson
ICONIP1
1996 A 'gammachirp' function as an optimal auditory filter with the Mellin transform
abstract
A 'gammachirp' function has been derived as an optimal auditory filter function in terms of minimal uncertainty in a joint time and modified-scale representation if the scale transform defined by Cohen (1989) is used in the auditory system. The gammatone function, which is widely used as the impulse response of a linear auditory filter, is a first-order approximation of the 'gammachirp' function consisting of a chirp carrier with an envelope that is a gamma distribution function. The optimality of the 'gammachirp' function is argued for the general Mellin transform since Cohen's scale transform is a specific example of the Mellin transform. A sample speech signal is analyzed to demonstrate the properties of a joint time and scale distribution derived with a short-time Mellin transform in comparison with a short-time Fourier spectrum.
Toshio Irino
ICASSP1
1994 A theory of asymmetric intensity enhancement around acoustic transients
Toshio Irino, Roy D. Patterson
ICSLP1
1992 Signal reconstruction from modified wavelet transform-An application to auditory signal processing
abstract
A novel method of signal reconstruction from a modified auditory representation is presented. This consists of three parts: (1) an algorithm to reconstruct a signal from its modified wavelet transform with a general wavelet; (2) obtaining an auditory representation using an auditory wavelet transform whose analyzing wavelet is the impulse response of an auditory peripheral model; and (3) estimating the reconstruction algorithm both with and without data reduction. An example of its application to the time-scale modification of speech is presented. High-quality speech successfully generated by time-scale modification shows that the reconstruction method is suitable for various applications as well as making experimental auditory stimuli.>
Toshio Irino, Hideki Kawahara
ICASSP1
1990 A Method for Designing Neural Networks Using Nonlinear Multivariate Analysis: Application to Speaker-Independent Vowel Recognition
abstract
A nonlinear multiple logistic model and multiple regression analysis are described as a method for determining the weights for two-layer networks and are compared to error backpropagation. We also provide a method for constructing a three-layer network whose semilinear middle units are primarily provided to discriminate two categories. Experimental results on speaker-independent vowel recognition show that both multivariate methods provide stable weights with fewer iterations than backpropagation training started with random initial weights, but with slightly inferior performance. Backpropagation training with initial weights determined by a multiple logistic model after introduction of data distribution information gives a recognition rate of 98.2%, which is significantly better than average backpropagation with random initial weights.
Toshio Irino, Hideki Kawahara
Neural Comput.1
1988 Vowel-feature extraction from cochlear vibration using neural networks
Toshio Irino, Hideki Kawahara
Neural Networks1