Takayuki Arai

dblp:04/1804 · DBLP profile ↗
← Back
76ranked-venue papers
34as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 74 · 32 first-author · 12 since 2021Artificial intelligence and machine learning · 62 · 31 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Vocal-tract model with two directions: Static design for a dummy head and dynamic design for a speaking machine
Takayuki Arai
INTERSPEECH1
2025 Using spatial sound reproduction for studying speech perception of listeners with different language immersion experiences
abstract
This study evaluates a research method for studying speech perception of listeners with different language background under practical acoustic environments. The proposed research method utilises spatial sound reproduction, an emerging technology that enables reproducing arbitrary acoustic environments in controlled laboratory settings, for testing participants recruited at multiple locations that are geographically distant from each other. To validate the research method, the current study conducted a listening test in a real seminar room and chapel as well as under a spherical harmonics-based spatial sound reproduction that reproduced the acoustics of the two venues up to the third order and investigates differences in the results collected from the two test types. Three groups of participants who had had different immersion level to New Zealand English were recruited in Auckland, New Zealand and Tokyo, Japan. The experimental results show that spatial sound reproduction is able to capture the advantage of first language (L1) listeners in terms of understanding speech in noise and reverberation correctly but is not sensitive enough to describe the subtle difference among second language (L2) listeners with different level of language immersion experiences. The research method is also partially able to describe how well listeners can benefit from spatial release from masking regardless of their language immersion experiences under room acoustics with higher speech clarity (C50), and may represent the effect of room acoustics in the real room within a certain range of room acoustics characterised by speech clarity. • Evaluates reproducibility of L2 speech perception using spatial sound reproduction. • L1 advantage in understanding speech in noise and reverberation correctly captured. • Not sensitive enough to describe difference by L2 language immersion experiences. • Spatial release from masking partially replicated when speech clarity is high. • Effect of real room’s acoustics within certain range of speech clarity reproduced.
Yusuke Hioka, C. T. Justine Hui, Hinako Masuda, Yunqi C. Zhang, Eri Osawa, Takayuki Arai
Speech Commun.6
2025 The impact of first and second formant variations on vowel identification among elderly Japanese listeners
abstract
Correct vowel identification is important for effective speech communication. While vowel perception largely requires accurate processing and detection of spectral information, spectral processing abilities (suprathreshold deficits such as difference limens for frequency) have been shown to degrade with age, and therefore potentially affecting elderly listeners’ vowel categorisation. The current study examined a group of near-normal hearing elderly listeners matched in age, hearing thresholds, auditory filter bandwidths and cognitive performance on their difference limens for frequency (DLFs) and vowel identification. Japanese has a five-vowel system, where/o/ and/e/ approximately differ in second formant (F2) frequencies and/a/ and/u/ approximately differ in first formant (F1) frequencies. We created two continua differing in F1 and F2, respectively, and examined elderly listeners’ identification of the vowels. We found differences in DLFs within the elderly listeners’ group, where the group with lower difference limens (LDL) to have similar DLFs to a control group of young listeners. Grouping the listeners according to their DLFs via k-means clustering with the young listeners’ results, we found the two groups to differ in their/o/ -/e/ (F2) vowel perception, where we observed a significant shift in the boundary between the two groups, but not when F1 was manipulated. This suggests that for elderly listeners, vowel identification that depends on higher formant cues may be affected by frequency discrimination abilities. However, the results did not suggest a difference in accuracy, but instead, a difference in how categories are assigned. • Japanese elderly listeners’ vowel perception varying in F1 and F2 was examined. • Listeners were grouped according to their frequency discrimination abilities (DLF). • A shift in perceptual boundary observed between the groups when F2 is manipulated. • No difference when F1 is manipulated.
C. T. Justine Hui, Takayuki Arai
Speech Commun.2
2025 Role of language familiarity in understanding speech in noise under various acoustic environments
abstract
We communicate in complex acoustic environments in everyday life but our familiarity with the language can affect how well we can understand speech in these environments. The current study examines the role of language familiarity in understanding speech in varying acoustic environments via a speech intelligibility test conducted under anechoic and reverberant conditions with various speech-noise separation angles. Four groups were recruited with differing level of language familiarity: first language (L1) New Zealand English (NZE) listeners, second language (L2) Japanese native listeners with exposure to NZE, L2 Japanese native listeners with overseas English experiences without exposure to NZE, and Japanese native listeners who have learnt English as a foreign language (FL) without overseas English experiences. The L1 group performed better in overall speech intelligibility performance compared to the 3 Japanese native groups. Contrary to previous literature where non-native listeners were found to have a similar benefit from spatial separation to native listeners, this was not the case for the FL group, suggesting that this benefit is only available for listeners with a certain level of language familiarity. While there were differences between L2 and FL groups in the anechoic condition, these differences become marginal in the reverberant conditions for the two groups with little exposure to NZE. This suggests that familiarity to the specific language variety has an advantage in acoustically adverse environments.
C. T. Justine Hui, Hinako Masuda, Eri Osawa, Takayuki Arai, Catherine I. Watson, Yusuke Hioka
Speech Commun.4
2024 Production of phrases by mechanical models of the human vocal tract
Takayuki Arai, Ryohei Suzuki, Chandler Earp, Shinya Tsuji, Keiko Ochi
INTERSPEECH1
2023 Comparing /b/ and /d/ with a Single Physical Model of the Human Vocal Tract to Visualize Droplets Produced while Speaking
Takayuki Arai, Tsukasa Yoshinaga, Akiyoshi Iida
INTERSPEECH1
2023 A Relationship Between Vocal Fold Vibration and Droplet Production
abstract
While some aerosol droplets causing airborne transmissions are argued to be produced by vocal fold vibrations, the detailed production mechanisms were unclear due to the difficulty of direct observation.In this study, by using a transparent acrylic vocal fold model and high-speed imaging, we observed vocal fold vibrations that produce droplets of artificial mucus to clarify the relationship between the vocal fold vibration and the droplet production.The vocal fold model was set on a lung model which has a manual diaphragm.The artificial mucus in between the vocal folds was lighted up by a laser sheet.The results showed that droplets were produced when the sound amplitudes were decreased or the fundamental frequency became unstable.We observed that mucus drops attached to the middle of the vocal fold wall splashed and formed a droplet that flew from the vocal fold wall.This suggests that the shear forces of turbulent airflow passing on the mucus mainly produced the droplets.
Tsukasa Yoshinaga, Takayuki Arai, Akiyoshi Iida
INTERSPEECH2
2022 Syllable sequence of /a/+/ta/ can be heard as /atta/ in Japanese with visual or tactile cues
Takayuki Arai, Miho Yamada, Megumi Okusawa
INTERSPEECH1
2021 Downsizing of Vocal-Tract Models to Line up Variations and Reduce Manufacturing Costs
Takayuki Arai
Interspeech1
2021 Vocal-Tract Models to Visualize the Airstream of Human Breath and Droplets While Producing Speech
Takayuki Arai
Interspeech1
2021 Comparison Between Lumped-Mass Modeling and Flow Simulation of the Reed-Type Artificial Vocal Fold
Rafia Inaam, Tsukasa Yoshinaga, Takayuki Arai, Hiroshi Yokoyama, Akiyoshi Iida
Interspeech3
2021 Effect of prior exposure on the perception of Japanese vowel length contrast in reverberation for nonnative listeners
abstract
While reverberation often degrades speech intelligibility, previous studies have shown that prior exposure to reverberation can reduce its adverse effects on speech perception. The current study investigated the effect of prior exposure to reverberation on the perception of nonnative speech sounds. We compared the results from two experiments, one in “blocked presentation” where the target words with the same amount of reverberation were presented to participants consistently, and another in “random presentation” where the amount of reverberation added to the target words changed between each trial. A Japanese minimal pair,/ie/ ‘house’ -/iie/ ‘no’, where vowel length creates a phonemic difference, was used as the target. The results for native listeners showed that their responses did not differ significantly between the blocked and random presentations. On the other hand, the effect of presentation type was significant in terms of the responses from the nonnative listeners. The results showed that nonnative listeners did not respond differently between the anechoic and reverberant conditions in the blocked presentation. However, there was a significant difference between the anechoic and reverberant conditions in the random presentation. The results from the nonnative listeners suggest that they try to obtain information of reverberation from the exposure since they could not use top-down processing effectively as much as native listeners.
Eri Osawa, C. T. Justine Hui, Yusuke Hioka, Takayuki Arai
Speech Commun.4
2020 Two Different Mechanisms of Movable Mandible for Vocal-Tract Model with Flexible Tongue
Takayuki Arai
INTERSPEECH1
2018 Flexible Tongue Housed in a Static Model of the Vocal Tract With Jaws, Lips and Teeth
Takayuki Arai
INTERSPEECH1
2017 Integrated Mechanical Model of [r]-[l] and [b]-[m]-[w] Producing Consonant Cluster [br]
Takayuki Arai
INTERSPEECH1
2017 Vocal-Tract Model with Static Articulators: Lips, Teeth, Tongue, and More
Takayuki Arai
INTERSPEECH1
2016 Mechanical Production of [b], [m] and [w] Using Controlled Labial and Velopharyngeal Gestures
Takayuki Arai
INTERSPEECH1
2015 Hands-on tool producing front vowels for phonetic education: aiming for pronunciation training with tactile sensation
Takayuki Arai
INTERSPEECH1
2015 Two extensions of umeda and teranishi's physical models of the human vocal tract
Takayuki Arai
INTERSPEECH1
2015 Perception of an existing and non-existing L2 English phoneme behind noise by Japanese native speakers
Mako Ishida, Takayuki Arai
INTERSPEECH2
2014 Retroflex and bunched English /r/ with physical models of the human vocal tract
Takayuki Arai
INTERSPEECH1
2013 Physical models of the vocal tract with a flapping tongue for flap and liquid sounds
abstract
Certain sounds are difficult for children to produce, even if the sounds are in their native language. For example, Japanese /r/ can be difficult for Japanese children to learn. Second language learners can also have difficulty acquiring certain sounds. For example, Japanese speakers learning English often have difficulty with English /r/ and /l/. To address this problem, we have developed two new physical models of the vocal tract: one for flap sounds (Model A) and another for liquid sounds (Model B). Each of them has a flapping tongue, and for Model B, the length of the tongue is variable. When the tongue is short, we can produce alveolar/retroflex approximants, and when the tongue is long we can produce lateral approximants. We recorded several sets of sounds produced by these models, analyzed the speech data, and used them for perceptual experiments. From the acoustic analysis and the perceptual experiments, we confirmed that the sounds produced by Model A were heard as Japanese /r/, and the sounds produced by Model B were heard as English /r/ and /l/. Furthermore, the models are helpful for practicing pronunciation because learners can see the tongue, alter tongue position manually, and hear the output sounds. Index Terms: speech production, physical models of the human vocal tract, tongue, flap/liquid sounds
Takayuki Arai
INTERSPEECH1
2013 On why Japanese /r/ sounds are difficult for children to acquire
abstract
Many studies have pointed out that the /r/ sounds in Japanese tend to be difficult for native children of Japanese to acquire. To verify this, we first investigated Japanese /r/ sounds uttered by two-year-old twins as a case study. The acoustic analysis of the recordings, which included several words with various /r/ sounds, revealed that certain /r/ sounds are difficult to produce and are often produced with speech errors. We also analyzed a set of utterances of Japanese /r/ spoken in a variety of phones pronounced by an adult male speaker. Then, for comparison, we synthesized Japanese /r/ sounds using four parameters. We conducted two perceptual experiments: one for the natural speech by the male speaker of Japanese, and another for the synthesized speech sounds based on the four parameters. The results showed that variation in pronunciation in adults was widely distributed. We discussed the reasons that it takes time for children to acquire /r/ sounds, and we concluded that it is possibly due to the combination of two factors: 1) some /r/ sounds themselves are difficult to produce, and 2) there is a wide distribution of pronunciation variation in adult speakers. Index Terms: Japanese /r/ sounds, flap sounds, children's speech, allophones, speech production
Takayuki Arai
INTERSPEECH1
2013 Weighting of acoustic cues shifts to frication duration in identification of fricatives/affricates when auditory properties are degraded due to aging
abstract
In previous studies, we conducted several experiments, including identification tests for young and elderly listeners using /shi/–/chi/ (CV) and /ishi/–/ichi/ (VCV) continua. For the CV stimuli, confusion of /shi/ as /chi/ increased when the frication had a long rise time, and /chi/ was confused with /shi/ when the frication had a short rise time. This was true for the group with the following auditory property degradation: 1) elevation of absolute threshold, 2) presence of loudness recruitment, and 3) deficit of auditory temporal resolution. When auditory property degradation was observed, the weighting of acoustic cues shifted to frication duration rather than the gradient of the amplitude of frication. The latter was calculated by dividing frication amplitude by rise time. In the VCV stimuli, confusion of /ichi/ as /ishi/ occurred for a long silent interval between the first V and C with auditory property degradation, and the weighting of acoustic cues shifted from the silent interval to frication duration. In the present study, we unified these findings into a single framework and found that degradation of auditory properties causes listeners to prefer duration of frication as a cue for identifying fricatives and affricates. Index Terms: elderly listeners, hearing impairments, aging, weighting shift of acoustic cues, trading relations, speech perception, fricatives/affricates
Keiichi Yasu, Takayuki Arai, Kei Kobayashi, Mitsuko Shindo
INTERSPEECH2
2012 Digital Pattern Playback for education in digital signal processing and speech science
abstract
We developed a digital version of Pattern Playback to convert a spectrographic representation of speech back into a speech signal. Pattern Playback was originally developed by Cooper and his colleagues from Haskins Laboratories in the late 1940s. We used our Digital Pattern Playback (DPP) for instruction in digital signal processing and speech science. The original DPP used two different algorithms: amplitude modulation and fast Fourier transform. The new DPP uses additive synthesis of sinusoidal harmonics, which is easier for undergraduate college students to understand. We also designed a scientific exhibition with DPP at a science museum for children and adults. DPP is educational for a wide variety of people, from children to technical students.
Takayuki Arai
ICASSP1
2012 Vowels Produced by Sliding Three-tube Model with Different Lengths
Takayuki Arai
INTERSPEECH1
2012 Intelligibility of speech spoken in noise/reverberation for older adults in reverberant environments
Nao Hodoshima, Takayuki Arai, Kiyohiro Kurisu
INTERSPEECH2
2012 Errata to "Using Steady-State Suppression to Improve Speech Intelligibility in Reverberant Environments for Elderly Listeners"
abstract
In the above titled paper (ibid., vol. 18, no. 7, pp. 1775-1780, Sep. 2010), several IPA fonts were displayed incorrectly. The correct fonts are presented here.
Takayuki Arai, Nao Hodoshima, Keiichi Yasu
IEEE Trans. Speech Audio Process.1
2011 Physical Models Producing Vowels with Pitch Variation
Takayuki Arai
INTERSPEECH1
2010 Mechanical vocal-tract models for speech dynamics
abstract
Arai has developed several physical models of the human vocal tract for education and has reported that they are intuitive and helpful for students of acoustics and speech science. We first reviewed dynamic models, including the sliding three-tube (S3T) model and the flexible-tongue model. We then developed a head-shaped model with a sliding tongue, which has the advantages of both the S3T and flexible-tongue models. We also developed a computer-controlled version of the Umeda & Teranishi model, as the original model was hard to manipulate precisely by hand. These models are useful when teaching the dynamic aspects of speech. Index Terms: vocal-tract model, speech dynamics, speech production, education in acoustics, speech science
Takayuki Arai
INTERSPEECH1
2010 Enhanced speech yielding higher intelligibility for all listeners and environments
abstract
The current paper discusses two approaches to enhanced speech in reverberation/noise: machine signal processing and human speech production. We reviewed the speech enhancement techniques, including steady-state suppression and compared the modulation spectra of speech signals before and after processing. We also introduced the Lombard-like effect of speech in reverberation, and compared the characteristics of speech signals, including the modulation spectra between speech signals uttered in quiet and reverberation. We found that the enhanced speech signals have distinct characteristics that yield higher speech intelligibility. Index Terms: speech enhancement, steady-state suppression, modulation spectrum, intelligibility of speech, reverberation, Lombard effect 1.
Takayuki Arai, Nao Hodoshima
INTERSPEECH1
2010 Perception of voiceless fricatives by Japanese listeners of advanced and intermediate level English proficiency
abstract
Numerous research has investigated how first language influences the perception of foreign sounds. The present study focuses on the perception of voiceless English fricatives by Japanese listeners with advanced and intermediate level English proficiency, and compares their results with that of English native listeners. Listeners identified consonants embedded in /a _ _ a / in quiet, multi-speaker babble and white noise (SNR=0 dB). Results revealed that intermediate level learners scored the lowest among all listener groups, and /th/-/s / confusions were unique to Japanese listeners. Confusions of /th/-/f / were observed among all listener groups, which suggest that those phoneme confusions may be universal.
Hinako Masuda, Takayuki Arai
INTERSPEECH2
2010 Using Steady-State Suppression to Improve Speech Intelligibility in Reverberant Environments for Elderly Listeners
abstract
Reverberation is a large problem for speech communication and it is known that strong reverberation affects speech intelligibility. This is especially true for people with hearing impairments and/or elderly people. Several approaches have been proposed and discussed to improve speech intelligibility degraded by reverberation. Steady-state suppression is one such approach, in which speech signals are processed before being radiated through the loudspeakers of a public address system to reduce overlap masking, which is one of the major causes of degradation in speech intelligibility. We investigated whether the steady-state suppression technique improves the intelligibility of speech in reverberant environments for elderly listeners. In both simulated and actual reverberant environments, elderly listeners performed worse than younger listeners. The performance in an actual hall was better than with simulated reverberation, and this was consistent with the results for younger listeners. Although the normal hearing group performed better than the presbycusis group, the steady-state suppression technique improved the intelligibility of speech for elderly listeners as was observed for younger listeners in both simulated and actual reverberant environments.
Takayuki Arai, Nao Hodoshima, Keiichi Yasu
IEEE Trans. Speech Audio Process.1
2009 Dialectal characteristics of osaka and tokyo Japanese: analyses of phonologically identical words
abstract
This study investigates the characteristics of the two major dialects of Japanese: Osaka and Tokyo dialects. We recorded the utterances of the speakers of both dialects, and analysed the differences that appear in the accentuation of the words at the phonetic-acoustic level. The Japanese words that are phonologically identical in both dialects were used as the analysis target. The results showed that the pitch patterns contained the dialect-dependent features of Osaka Japanese. Furthermore, these patterns could not be fully mimicked by speakers of Tokyo Japanese. These results show that there is a phonetics-phonology gap in the dialectal differences, and that we may exploit this gap for forensic purposes. Index Terms: pitch patterns, dialectal characteristics, phonetics-phonology gap, forensic phonetics
Kanae Amino, Takayuki Arai
INTERSPEECH2
2009 Sliding vocal-tract model and its application for vowel production
abstract
In a previous study, Arai implemented a sliding vocal-tract model based on Fant’s three-tube model and demonstrated its usefulness for education in acoustics and speech science. The sliding vocal-tract model consists of a long outer cylinder and a short inner cylinder, which simulates tongue constriction in the vocal tract. This model can produce different vowels by sliding the inner cylinder and changing the degree of constriction. In this study, we investigated the model’s coverage of vowels on the vowel space and explored its application for vowel production in the speech and hearing sciences. Index Terms: vocal-tract model, three-tube model, vowel production, education in acoustics, speech science
Takayuki Arai
INTERSPEECH1
2009 Simple physical models of the vocal tract for education in speech science
abstract
In the speech-related field, physical models of the vocal tract are effective tools for education in acoustics. Arai’s cylinder-type models are based on Chiba and Kajiyama’s measurement of vocal-tract shapes. The models quickly and effectively demonstrate vowel production. In this study, we developed physical models with simplified shapes as educational tools to illustrate how vocal-tract shape accounts for differences among vowels. As a result, the five Japanese vowels were produced by tube-connected models, where several uniform tubes with different cross-sectional areas and lengths are connected as Fant’s and Arai’s three-tube models. Index Terms: speech science, vocal-tract model, education in acoustics, vowel production, acoustic tube
Takayuki Arai
INTERSPEECH1
2008 Perceptual speaker identification using monosyllabic stimuli - effects of the nucleus vowels and speaker characteristics contained in nasals
abstract
The goal of our research is to find out the acoustical correlates of human perception of speaker identity. In this study we investigated the effects of the stimulus contents on perceptual speaker identification. Forty-eight monosyllables were used as the stimuli for identifying four male speakers. The results showed that the syllables containing a coronal nasal yielded higher identification accuracies than the syllables without it, and the syllables with a back vowel gained significantly better scores than those with a front vowel. We also found speaker-dependent characteristics in the velar movements in articulation of nasal consonants. Index Terms: perceptual speaker identification, speaker’s individuality, nasals, vowels, energy onset
Kanae Amino, Takayuki Arai
INTERSPEECH2
2008 Physical models of the human vocal tract with gel-type material
abstract
Beginning in 2001, we have been developing models of the vocal tract to promote a more intuitive understanding of the theories for speech science for technical and non-technical students. In this paper, we compared and contrasted four newer models of the talking heads: our original model with a gel-type tongue, a similar model including teeth and palate, a third model with teeth and palate having a default low tongue height, and a fourth ultra-malleable model made completely of gel material. Results are discussed regarding their strengths and weaknesses for educational purposes and articulatory training in speech pathology and language learning. Index Terms: speech production, vocal-tract model, flexible tongue, gel-type material
Takayuki Arai
INTERSPEECH1
2008 Science workshop with sliding vocal-tract model
abstract
In recent years, we have developed physical models of the vocal tract to promote education in the speech sciences for all ages of students. Beginning with our initial models based on Chiba and Kajiyama’s measurements, we have gone on to present many other models to the public. Through science fairs which attract all ages, including elementary through college level students, we have sought to increase awareness of the importance of the speech sciences in the public mind. In this paper we described two science workshops we organized at the National Science
Takayuki Arai
INTERSPEECH1
2008 Improving consonant identification in noise and reverberation by steady-state suppression as a preprocessing approach
abstract
Noise (N), reverberation (R), and a combination of N and R (NR) differently degrade speech intelligibility. The current study aims to improve speech intelligibility in public spaces by processing speech signals through public address systems (a preprocessing approach). As a preprocessing approach, we proposed steady-state suppression, and it has improved consonant identification in R. The current study tests the effect of steady-state suppression in N, R, and NR at three signal to noise ratios and at reverberation time of 0.9 s. Results showed that steady-state suppression significantly improved consonant identification by 21 young people in NR and in R. Furthermore, steady-state suppression improved consonant identification more in NR than in R. The results indicate that steady-state suppression may be applicable to public spaces having N and R. The results also indicate that an integration of N and R improves the performance of a preprocessing approach in a certain range of N and R.
Nao Hodoshima, Wataru Yoshida, Takayuki Arai
INTERSPEECH3
2008 Perception and production of consonant clusters in Japanese-English bilingual and Japanese monolingual speakers
abstract
Previous research has revealed that Japanese native speakers are more likely to perceive an ‘illusory vowel ’ within consonant clusters compared to French native speakers (Dupoux et al. 1998). The aim of this research is to investigate the differences of perception and production of consonant clusters in Japanese-English bilinguals and Japanese monolinguals. The stimuli of the two experiments consist of 36 pseudo-words that contain the two sequences, VCCV and VCVCV. Results of the perception and the production experiments on the two groups of participants revealed that bilinguals were more likely to achieve high scores in both experiments. Index Terms: consonant clusters, bilingual, speech perception, speech production
Hinako Masuda, Takayuki Arai
INTERSPEECH2
2006 Steady-state suppression in reverberation: a comparison of native and nonnative speech perception
abstract
This study investigated whether the steady-state suppression method proposed by Arai et al. (2001, 2002) improved consonant identification for nonnative listeners in reverberation. It also compared the effect of steady-state suppression on consonant identification by native and nonnative listeners in reverberation. We used steady-state suppression as a preprocessing technique which processes speech signals before they are radiated from loudspeakers in order to reduce the amount of overlap-masking. Participants were 24 native English (native listeners) and 24 Japanese speakers (nonnative listeners), both with normal hearing. A diotic Modified Rhyme Test was conducted with and without steady-state suppression for reverberation times of 0.4, 0.7 and 1.1 s and a nonreverberant condition. The results showed that native listeners performed better than nonnative listeners, and that the mean percentage of correct answers in initial consonants was higher than in final consonants. The results also showed that processed and unprocessed speech was comparable for word initial and final consonants. These findings indicate that parameters of steady-state suppression would need adjustment to accommodate speech materials and reverberant conditions. They also suggest that the difficulties that nonnative listeners have might not be due to the actual acoustic-phonetic information from the signal. Index Terms: speech enhancement, nonnative listeners, reverberation, steady-state suppression
Nao Hodoshima, Dawn M. Behne, Takayuki Arai
INTERSPEECH3
2005 The correspondences between the perception of the speaker individualities contained in speech sounds and their acoustic properties
abstract
This study investigates the correspondences between the differences among the phones in human speaker identification and their acoustic properties. In the speaker identification test, the Japanese CV syllables excerpted from the carrier sentences were used as the stimuli. As pointed out in the previous studies, the stimuli containing the nasal sounds were significantly effective for the identification of the speakers, compared to other stimuli containing only the oral sounds. In the acoustic analyses, we analysed the spectral properties of the stimuli in order to explain these differences in the perception test, and we found that the cepstral distances among the speakers were significantly larger in the nasal sounds than in the oral sounds. Also, there were correspondences between the rankings of the consonants in the identification test and in the cepstral distances.
Kanae Amino, Tsutomu Sugawara, Takayuki Arai
INTERSPEECH3
2005 Comparing tongue positions of vowels in oral and nasal contexts
abstract
We studied tongue positions of vowels in oral and nasal contexts. In the previous study [Arai, J. Acoust. Soc. Am., 115, p.2541 (2004)], formant frequencies were measured and bidirectional formant shifts in F1 frequency were observed: increasing F1 for high vowels and decreasing F1 for low vowels. Then, we tried to answer to the next question, that is, whether or not speakers and/or listeners compensate for the formant shifts. The perceptual experiment by Arai (2004) showed that compensation occurs when an isolated vowel has nasalization and is accompanied by formant transitions. This result agreed with the findings of Krakow et al. [J. Acoust. Soc. Am. 83, 1146-1158 (1988)]. The goal of this study is to examine the compensation effect for the formant shifts in production. In the EMMA experiment, the measurement of the positions of the articulators showed almost no compensation except for the lowest vowel /#/.
Takayuki Arai
INTERSPEECH1
2005 Steady-state pre-processing for improving speech intelligibility in reverberant environments: evaluation in a hall with an electrical reverberator
abstract
To improve speech intelligibility in reverberant environments, Arai et al. proposed the methods of “steady-state suppression” (Arai, 2002) and “steady-state zero-padding” (Arai, 2005) as a pre-processing method. We conducted two perceptual experiments to evaluate and compare these methods in a hall with an electrical reverberator. The advantage of using this electrical reverberator is that the same subjects are able to participate in the experiments with many different reverberant conditions (reverberation time was varied from 2.6-3.3s, in this study). As results, steady-state suppression showed the floor effect under these reverberant conditions, whereas steady-state zero-padding yielded significant improvements in terms of the speech intelligibility in the same reverberant conditions.
Nahoko Hayashi, Takayuki Arai, Nao Hodoshima, Yusuke Miyauchi, Kiyohiro Kurisu
INTERSPEECH2
2005 A preprocessing technique for improving speech intelligibility in reverberant environments: the effect of steady-state suppression on elderly people
abstract
In a large auditorium, perceiving speech may become difficult. One reason that reverberation degrades speech intelligibility is the effect of overlap-masking (Bolt and MacDonald, 1949; Nabelek and Robinette, 1978). Reverberation is a more critical issue for elderly people to perceive speech than it is for young people (Fitzgibbons and Gordon-Salant, 1999). Arai et al. suppressed steady-state portions of speech which have more energy but are less crucial for speech perception, and confirmed promising results for improving speech intelligibility (Arai et al., 2001, 2002). Hodoshima et al. conducted perceptual tests to confirm the effectiveness of steady-state suppression with several reverberation conditions, and obtained significant improvements with reverberation times of 0.7-1.3 s. In this study, we conducted an experiment for evaluating steady-state suppression with fifty elderly people and found that there were significant improvements. Also, steady-state suppression yielded better improvements in speech intelligibility for elderly people than it did for young people.
Yusuke Miyauchi, Nao Hodoshima, Keiichi Yasu, Nahoko Hayashi, Takayuki Arai, Mitsuko Shindo
INTERSPEECH5
2005 Modulation enhancement of speech by a pre-processing algorithm for improving intelligibility in reverberant environments
Akiko Amano-Kusumoto, Takayuki Arai, Keisuke Kinoshita, Nao Hodoshima, Nancy Vaughan
Speech Commun.2
2004 Perceptual discrimination of prosodic types and their preliminary acoustic analysis
abstract
A perceptual discrimination test was conducted to investigate whether humans can discriminate prosodic types solely based on suprasegmental acoustic cues. Excerpts from Chinese, English, Spanish, and Japanese, differing in lexical accent types and rhythm types, were used. From these excerpts, “source” signals of the source-filter model, differing in F0, intensity, and HNR, were created and used in a perceptual experiment. In general, the results indicated that humans can discriminate these prosodic types and that the discrimination is easier if more acoustic information is available. Further, the results showed that languages with similar rhythm types are difficult to discriminate (i.e., Chinese-English, EnglishSpanish, and Spanish-Japanese). As to accent types, tonal/nontonal contrast was easy to detect. We also conducted a preliminary acoustic analysis of the experimental stimuli and found that quick F0 fluctuations in Chinese contribute to the perceptual discrimination of tonal/non-tonal accents.
Masahiko Komatsu, Tsutomu Sugawara, Takayuki Arai
INTERSPEECH3
2003 Estimating number of speakers by the modulation characteristics of speech
abstract
A method for estimating number of speakers of mixed speech signals was proposed. The algorithm was based on the modulation characteristics of speech, specifically that a single speech utterance typically has a distinct modulation pattern with a peak around 4-5 Hz. Having observed that the modulation peak decreases as number of speakers increases, our estimation algorithm used the region of the modulation frequency between 2 and 8 Hz. We obtained a novel parameter we called "equivalent number of speakers" to estimate the number of simultaneous speakers when speech signals contain multiple speakers.
Takayuki Arai
ICASSP (2)1
2003 Improving speech intelligibility by steady-state suppression as pre-processing in small to medium sized halls
abstract
One of the reasons that reverberation degrades speech intelligibility is the effect of overlap-masking, in which segments of an acoustic signal are affected by reverberation components of previous segments [Bolt et al., 1949]. To reduce the overlap-masking, Arai et al. suppressed steady-state portions having more energy, but which are less crucial for speech perception, and confirmed promising results for improving speech intelligibility [Arai et al., 2002]. Our goal is to provide a pre-processing filter for each auditorium. To explore the relationship between the effect of a pre-processing filter and reverberation conditions, we conducted a perceptual test with steady-state suppression under various reverberation conditions. The results showed that processed stimuli performed better than unprocessed ones and clear improvements were observed for reverberation conditions of 0.8 - 1.0s. We certified that steady-state suppression was an effective pre-processing method for improving speech intelligibility under reverberant conditions and proved the effect of overlap-masking.
Nao Hodoshima, Takayuki Arai, Tsuyoshi Inoue, Keisuke Kinoshita, Akiko Amano-Kusumoto
INTERSPEECH2
2002 Duration and F0 as perceptual cues to Japanese vowel quantity
abstract
Vowel duration and local fundamental frequency changes are investigated as acoustical cues to vowel quantity identification by Japanese listeners. To examine the role of these factors, a perception experiment was carried out. The results indicate that, even though vowel duration serves as a dominant perceptual cue, when vowel quantity cannot be adequately cued by vowel duration alone, the F0 information within the vowel can be used to identify vowel quantity in Japanese.
Keisuke Kinoshita, Dawn M. Behne, Takayuki Arai
INTERSPEECH3
2002 Multi-dimensional analysis of sonority: perception, acoustics, and phonology
Masahiko Komatsu, Shinichi Tokuma, Won Tokuma, Takayuki Arai
INTERSPEECH4
2001 Prototype of a vocal-tract model for vowel production designed for education in speech science
Takayuki Arai, Nobuyuki Usuki, Yuji Murahara
INTERSPEECH1
2001 The relation between speech intelligibility and the complex modulation spectrum
abstract
The amplitude and phase components of the modulation spectrum were dissociated in order to ascertain the importance of cross-spectral, envelope-modulation phase information for understanding spoken language. The dissociation was effected via local time reversals of the speech waveform (i.e., flipping the signal on its horizontal axis) at intervals ranging between 0 and 180 ms. Intelligibility declines progressively as the length of the time-reversed segment increases, down to an asymptotic trough in performance at 100 ms (4% of the words correct). Intelligibility does not correlate highly with the amplitude component of the modulation spectrum, but does coincide closely with the contour of the complex modulation spectrum, a representation that integrates the cross-spectral modulation phase and the conventional (amplitude-based) modulation spectrum into a unified representation. The results imply that intelligibility is based on both the phase and amplitude components of the modulation spectrum.
Steven Greenberg, Takayuki Arai
INTERSPEECH2
2001 Human language identification with reduced segmental information: comparison between monolinguals and bilinguals
abstract
We conducted human language identification experiments using signals with reduced segmental information with Japanese and bilingual subjects. American English and Japanese excerpts from the OGI_TS Corpus were processed by spectral-envelope removal (SER), vowel extraction from SER (VES) and temporal-envelope modulation (TEM). With the SER signal, where the spectral-envelope is eliminated, humans could still identify the languages fairly successfully. With the VES signal, which retains only vowel sections of the SER signal, the identification score was low. With the TEM signal, composed of white-noise-driven intensity envelopes from several frequency bands, the identification score rose as the number of bands increased. Results varied depending on the stimulus language. Japanese and bilingual subjects demonstrated different scores from each other. These results indicate that humans can identify languages using a signal with drastically reduced segmental information. The results also suggest variation due to the phonetic attributes of languages and subjects’ knowledge.
Masahiko Komatsu, Kazuya Mori, Takayuki Arai, Yuji Murahara
INTERSPEECH3
2001 Modelling the perceptual identification of Japanese consonants from LPC cepstral distances
abstract
This study attempts to account for the perceptual phenomenon observed in Komatsu et al. [1] in terms of the spectral properties of the LPC re-synthesised stimuli. To implement this, LPC cepstral distances between re-synthesised samples and their original samples are measured. The results of the acoustic analysis and their comparison with the perceptual data indicate that there is a striking similarity in patterns between the spectral property of the Japanese consonants and their perceptual scores. This suggests that the role played by spectral information in the perception of Japanese consonants is significant across all consonant types, and also implies that even in its crudest form, it contributes significantly to their perception.
Masahiko Komatsu, Shinichi Tokuma, Won Tokuma, Takayuki Arai
INTERSPEECH4
2001 Using the modulation complex wavelet transform for feature extraction in automatic speech recognition
Yasunori Momomura, Kenji Okada, Takayuki Arai, Noboru Kanedera, Yuji Murahara
INTERSPEECH3
2000 The DSP experiments for under graduate students
abstract
This paper describes digital signal processing (DSP) microprocessor experiments designed for university juniors majoring in electric and electronic (E&E) engineering. At the time of enrollment, most students have only studied the curriculum for analog signal processing which is taken in the prior semester. The proposed DSP microprocessor experiments are included along with analog signal processing. This early-on introduction to DSP technology allows students to realize that, in comparison with analog circuits, DSP microprocessors can process the same signals in real-time with broader flexibility. Such an understanding is considered important to instill strong incentive for students to become interested in the field of DSP.
Y. Fuchiwaki, Nobuyuki Usuki, Takayuki Arai, Yuji Murahara
ICASSP3
2000 Modulation enhancement of speech as a preprocessing for reverberant chambers with the hearing-impaired
abstract
In this paper we report on a method for reducing the degradation of speech intelligibility in public halls caused severe reverberation. Hall reverberation makes speech more difficult to understand, particularly for the hearing-impaired. Our method involves processing the speech audio signal between a microphone and a loudspeaker that radiates the speech into the room. As there is a strong correlation between the modulation spectrum and the intelligibility of speech, we filtered the speech in the modulation frequency domain. Using several modulation filters, we conducted perceptual experiments with hearing-impaired subjects and asked their preference in a church. The experiments indicate that enhancing the modulation frequencies between 2 and 8 Hz improves intelligibility in reverberant environments. The four hearing-impaired subjects rated the processed speech easier to hear than the unprocessed speech.
Akiko Amano-Kusumoto, Takayuki Arai, Tomoko Kitamura, Yuji Murahara
ICASSP2
2000 The effect of polarity inversion of speech on human perception and data hiding as an application
abstract
In this paper we investigate how polarity inversion of speech signals effects human perception, and we apply this technique for data hiding. In most languages, glottal airflow during phonation is uni-directional, causing constant polarity of the speech waveform. On the other hand, the human auditory system cannot discriminate between speech signals with positive and negative polarity. Based on these facts, we developed an algorithm to hide data in speech signals. We assigned one bit to each syllable of speech, and inverted the polarity of the signal at every syllable according to the assigned bit. We performed a test using 20 sentences from the TIMIT corpus to determine both whether a human could distinguish between the original and polarity-inverted signal and whether we could automatically restore the embedded binary data. We found that we were able to successfully hide data and restore it automatically.
Shino Sakaguchi, Takayuki Arai, Yuji Murahara
ICASSP2
2000 Designing modulation filters for improving speech intelligibility in reverberant environments
abstract
In this paper, we propose a new technique to design modula-tion filters to reduce degradation of speech intelligibility in re-verberant environments. Using the inverse modulation trans-fer function, we design data-derived modulation filters for each speech frequency band. These filters preprocess speech signals between a microphone and a loudspeaker that radiates speech into a performance hall. Using our modulation fil-ters, we conducted perceptual experiments with one hearing-impaired subject and two subjects with normal hearing. Test results indicate that our proposed method improves the intel-ligibility of reverberant speech. 1.
Tomoko Kitamura, Keisuke Kinoshita, Takayuki Arai, Akiko Amano-Kusumoto, Yuji Murahara
INTERSPEECH3
2000 The effect of reduced spectral information on Japanese consonant perception: comparison between L1 and L2 listeners
abstract
We investigated how spectral information contributes to the perception of Japanese consonants, using re-synthesised samples that were created by (1) gradually reducing the order of LPC analysis in the residual excited LPC vocoder; and (2) gradually flattening the spectral peak in the frequency domain. The results of native Japanese speakers showed that the information in LPC residuals contributes significantly, if not sufficiently, to Japanese consonant perception, and that the minimum amount of spectral information is sufficient to achieve 90 % identification score. It was also found that, although the perceptual error patterns were different, there were striking similarities between Japanese and non-Japanese listeners in their averaged perception scores. The phonological feature analysis of the perceptual results indicated that the residuals provide broad phonotactic information such as major class features. 1.
Masahiko Komatsu, Won Tokuma, Shinichi Tokuma, Takayuki Arai
INTERSPEECH4
2000 Using the modulation wavelet transform for feature extraction in automatic speech recognition
Kenji Okada, Takayuki Arai, Noburu Kanederu, Yasunori Momomura, Yuji Murahara
INTERSPEECH2
1999 Effects of hoarseness on hypernasality ratings
abstract
Voice interfaces are not popular since they are neither useful nor user-friendly for non-specialist users. In this paper, EUROPA, a new framework for developing spoken dialogue systems, is introduced. In developing EUROPA, the authors focused on three points : (1) acceptance of spoken language, (2) portability in terms of domain and task, and (3) practical performance of the applied system. The framework is applied to prototyping a car navigation system called MINOS. MINOS is built on a portable PC, can process over 700 words of recognition vocabulary, and is able to respond to a user’s question within a few seconds.
Setsuko Imatomi, Takayuki Arai, Yuko Mimura, Masako Kato
EUROSPEECH2
1999 Human language identification with reduced spectral information
abstract
We conducted human language identification (LID) experiments using signals with reduced segmental information in pursuit of cues that humans use in their remarkable LID ability, which may be applicable to the development of robust automatic LID. American English and Japanese excerpts from the OGI-TS were processed by (1) spectral-envelope removal (SER) and (2) temporal-envelope modulation. With the SER signal, where the spectral-envelope is eliminated, humans could still identify the languages fairly successfully (85.2%). With the TEM signal, composed of white-noise driven, combined intensity envelopes from several frequency bands, the identification rate rose from 62.5 % to 93.8 % corresponding to the increasing number of bands from 1 to 4. These results, though with a limited number of languages, indicate that humans can identify languages using signal with its segmental information much reduced — in acoustic terms much reduced in spectral information. 1.
Kazuya Mori, Noriaki Toba, Tomoyuki Harada, Takayuki Arai, Masahiko Komatsu, Makiko Aoyagi, Yuji Murahara
EUROSPEECH4
1999 Temporal constraints on speech intelligibility as deduced from exceedingly sparse spectral representations
abstract
A novel means of quantifying the contribution of specific spectral bands for intelligibility is described.The spectrum of spoken English sentences is partitioned into one-third octave bands ("slits") and the contribution of each of four slits ascertained independently and in combination with other slits distributed across the spectrum.The intelligibility baseline (four concurrent slits) yields ca.85% intelligibility.The current study demonstrates that intelligibility progressively declines as the two central slits (2+3) are desynchronized between 25 and 250 ms.Beyond 250 ms intelligibility often declines even further but then begins to increase for greater degrees of asynchrony, suggesting the presence of a perceptual processing buffer of ca.200-300 ms in duration.The utility of the spectral slit technique is also demonstrated for estimating the contribution towards intelligibility of different regions of the modulation spectrum.The mid-frequency (10-25 Hz) modulations are shown to be of particular significance for encoding speech information above 1.5 kHz.These two experiments demonstrate the power and utility of using circumscribed portions of the spectrum for quantitative evaluation of the contribution made by specific spectrotemporal properties of the speech signal.
Rosaria Silipo, Steven Greenberg, Takayuki Arai
EUROSPEECH3
1999 On the relative importance of various components of the modulation spectrum for automatic speech recognition
Noboru Kanedera, Takayuki Arai, Hynek Hermansky, Misha Pavel
Speech Commun.2
1998 Speech intelligibility in the presence of cross-channel spectral asynchrony
abstract
The spectrum of spoken sentences was partitioned into quarter-octave channels and the onset of each channel shifted in time relative to the others so as to desynchronize spectral information across the frequency axis. Human listeners are remarkably tolerant of cross-channel spectral asynchrony induced in this fashion. Speech intelligibility remains relatively unimpaired until the average asynchrony spans three or more phonetic segments. Such perceptual robustness is correlated with the magnitude of the low-frequency (3-6 Hz) modulation spectrum and thus highlights the importance of syllabic segmentation and analysis for robust processing of spoken language. High-frequency channels (>1.5 kHz) play a particularly important role when the spectral asynchrony is sufficiently large as to significantly reduce the power in the low-frequency modulation spectrum (analogous to acoustic reverberation) and may thereby account for the deterioration of speech intelligibility among the hearing impaired under conditions of acoustic interference (such as background noise and reverberation) characteristic of the real world.
Takayuki Arai, Steven Greenberg
ICASSP1
1998 On properties of modulation spectrum for robust automatic speech recognition
abstract
We report on the effect of band-pass filtering of the time trajectories of spectral envelopes on speech recognition. Several types of filter (linear-phase FIR, DCT, and DFT) are studied. Results indicate the relative importance of different components of the modulation spectrum of speech for ASR. General conclusions are: (1) most of the useful linguistic information is in modulation frequency components from the range between 1 and 16 Hz, with the dominant component at around 4 Hz, (2) it is important to preserve the phase information in the modulation frequency domain, (3) the features which include components at around 4 Hz in the modulation spectrum outperform the conventional delta features, (4) the features which represent the several modulation frequency bands with appropriate center frequency and bandwidth increase recognition performance.
Noboru Kanedera, Hynek Hermansky, Takayuki Arai
ICASSP3
1998 Speech intelligibility derived from exceedingly sparse spectral information
abstract
Traditional models of speech assume that a detailed auditory analysis of the short-term acoustic spectrum is essential for understanding spoken language. The validity of this assumption was tested by partitioning the spectrum of spoken sentences into 1/3-octave channels ("slits") and measuring the intelligibility associated with each channel presented alone and in concert with the others. Four spectral channels, distributed over the speech-audio range (0.3-6 kHz) are sufficient for human listeners to decode sentential material with nearly 90 % accuracy although more than 70% of the spectrum is missing. Word recognition often remains relatively high (60-83%) when just two or three channels are presented concurrently, despite the fact that the intelligibility of these same slits, presented in isolation, is less than 9 % (Figure 2). Such data suggest that the intelligibility of spoken language is derived from a compound "image " of the modulation spectrum distributed across the frequency spectrum (Figures 1 and 3). Because intelligibility seriously degrades when slits are desynchronized by more than 25 ms (Figure 4) this compound image is probably derived from both the amplitude and phase components of the modulation spectrum, and implies that listeners ' sensitivity to the modulation phase is generally "masked " by the redundancy contained in full-spectrum speech (Figure 5). 1.
Steven Greenberg, Takayuki Arai, Rosaria Silipo
ICSLP2
1997 The temporal properties of spoken Japanese are similar to those of English
abstract
The languages of the world are generally classified into two types on the basis of their segmental timing. "Syllable-timed" languages, such as Japanese, are considered isochronous, exhibiting a highly regular pattern of syllabic duration. In contrast are the "stress-timed" languages, such as English, whose syllable timing varies greatly, both within and across sentential domains. The present study demonstrates that, even in a language as theoretically isochronous as Japanese, the duration of syllabic segments is as variable as their English counterparts. Moreover, the variability of moraic duration is as high as that observed for syllabic units. Two measures of segmental timing, syllable duration and the lowfrequency modulation spectrum, indicate that the coarse temporal characteristics of English and Japanese are remarkably similar. Such common properties may reflect inherent temporal characteristics of physiological mechanisms underlying the production and perception of speech that a...
Takayuki Arai, Steven Greenberg
EUROSPEECH1
1997 On the importance of various modulation frequencies for speech recognition
Noboru Kanedera, Takayuki Arai, Hynek Hermansky, Misha Pavel
EUROSPEECH2
1996 Intelligibility of speech with filtered time trajectories of spectral envelopes
Takayuki Arai, Misha Pavel, Hynek Hermansky, Carlos Avendaño
ICSLP1
1995 Analysis for palatalized articulation of [s] sounds using synthetic speech
abstract
Palatalized articulation (PA) is frequently observed in speech uttered by postoperative cleft palate patients. We analyzed the PA of [s] sounds and tested human perception of certain synthetic sounds to verify the characteristics of the PA of [s] sounds in Japanese. After analyzing the PA of [s] with linear predictive (LP) analysis, the mono-syllable /sa/, /su/ and /se/ were synthesized by an all-pole model. To synthesize the fricatives, we shifted the frequency of a complex-conjugate pole pair of a filter from 1000 to 3400 Hz. A perceptual experiment involving three speech therapists was carried out to analyze perception of the three syllables. From the results we concluded that fricatives having a peak around 1800 Hz tend to be identified as the PA of [s]. 1. INTRODUCTION After cleft palate surgery many cleft palate patients obtain normal velopharyngeal function; however some misarticulations may remain [1], such as palatalized articulation (PA) [2], nasopharyngeal articulation [3] ...
Takayuki Arai, Keiko Okazaki, Setsuko Imatomi, Yuichi Yoshida
EUROSPEECH1
1994 Analysis of phoneme-based features for language identification
abstract
This paper presents an analysis of the phonemic language identification system introduced previously (see Eurospeech, vol.2, p.1307, 1993), now extended to recognize German in addition to English and Japanese. In this system language identification is based on features derived from a superset of phonemes of all three languages. As we increase the number of languages, the need to reduce the feature space becomes apparent. Practical analysis of single-feature statistics in conjunction with linguistic knowledge leads to 90% reduction of the feature space with only a 5% loss in performance. Thus, the system discriminates between Japanese and English with 84.1% accuracy based on only 15 features compared to 84.6% based on the complete set of 318 phonemic features (or 83.6% using 333 broad-category features). Results indicate that a language identification system may be designed based on linguistic knowledge and then implemented with a neural network of appropriate complexity.>
Kay M. Berkling, Takayuki Arai, Etienne Barnard
ICASSP (1)2
1993 A comparison of approaches to automatic language identification using telephone speech
abstract
A variety of approaches to language identification, based on (a) acoustic features, (b) broad-category segmentation, and (c) fine phonetic classification, are introduced. These approaches are evaluated in terms of their ability to distinguish between English and Japanese utterances spoken over a telephone channel. It is found that the best performance (86.3 % accurate classification of utterances with a mean length of 13.4 sec) is obtained when fine phonetic features are employed. In addition, the results show the importance of discriminatory training rather than likelihood estimation. 1. INTRODUCTION As developments in telecommunications and long-distance travel cause national borders to become increasingly transparent, the ability to identify which language is being spoken is growing in importance. The utility of tasks such as directory assistance or automatic translation is, for instance, improved substantially by the availability of a means of identifying which language is being s...
Yeshwant K. Muthusamy, Kay M. Berkling, Takayuki Arai, Ronald A. Cole, Etienne Barnard
EUROSPEECH3