C. T. Justine Hui

dblp:246/4403 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-1411-8328ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Perception of Long and Short Vowel Contrast in Te Reo Māori in Clean and Everyday Listening Environments
abstract
Te reo Māori (the Māori language) is the language of the indigenous people in New Zealand and has a long-short vowel contrast. This study first investigates the cues used to perceive vowel length for Māori fluent users and learners. Secondly, it explores the effect of an everyday listening environment in the form of a traditional meeting house (wharenui) on cue-weighting between duration and stress for identifying long and short vowels when the stimuli are rendered with the wharenui room acoustics. An identification test was carried out with three pairs of words that differed in stress location and were manipulated in vowel duration. We found both groups to mainly use duration as a cue, even though the advanced listeners commented that they were listening for 'intonation'. For the wharenui acoustics, there were no differences observed between the groups and only a small categorical shift for the /a:/ vowel for the learners.
C. T. Justine Hui, Jenice Kuzhikombil, Isabella Shields, Hiraia Haami-Wells, Catherine I. Watson, Peter Keegan
INTERSPEECH1
2025 Effect of Noise Floor in Room Impulse Response on Speech Perception Under Spherical Harmonics-based Spatial Sound Reproduction
abstract
The current study investigates the effect of noise floor in measured room impulse responses (RIR) on the reproducibility of speech perception under spherical harmonics-based spatial sound reproduction. Subjective listening test measuring the intelligibility of speech in noise was conducted under the spatial sound reproduction implemented using practically measured RIR with varying level of noise floor. The same test was also conducted in the real rooms where the RIR were measured. The comparison of the experimental results from the spatial sound reproduction and the real room suggests using measured RIR with low noise floor contributes to reproducing speech perception in real rooms accurately when the room is highly reverberant. It also has an effect to improve the reproducibility when the sound sources are located at 5 m but not at 2 m. Truncating RIR to further remove the noise floor mostly did not help improve the reproducibility regardless of the acoustics of the room.
Yunqi C. Zhang, Dhruv Jagmohan, Hong Kit Li, C. T. Justine Hui, Yusuke Hioka
INTERSPEECH4
2025 Using spatial sound reproduction for studying speech perception of listeners with different language immersion experiences
abstract
This study evaluates a research method for studying speech perception of listeners with different language background under practical acoustic environments. The proposed research method utilises spatial sound reproduction, an emerging technology that enables reproducing arbitrary acoustic environments in controlled laboratory settings, for testing participants recruited at multiple locations that are geographically distant from each other. To validate the research method, the current study conducted a listening test in a real seminar room and chapel as well as under a spherical harmonics-based spatial sound reproduction that reproduced the acoustics of the two venues up to the third order and investigates differences in the results collected from the two test types. Three groups of participants who had had different immersion level to New Zealand English were recruited in Auckland, New Zealand and Tokyo, Japan. The experimental results show that spatial sound reproduction is able to capture the advantage of first language (L1) listeners in terms of understanding speech in noise and reverberation correctly but is not sensitive enough to describe the subtle difference among second language (L2) listeners with different level of language immersion experiences. The research method is also partially able to describe how well listeners can benefit from spatial release from masking regardless of their language immersion experiences under room acoustics with higher speech clarity (C50), and may represent the effect of room acoustics in the real room within a certain range of room acoustics characterised by speech clarity. • Evaluates reproducibility of L2 speech perception using spatial sound reproduction. • L1 advantage in understanding speech in noise and reverberation correctly captured. • Not sensitive enough to describe difference by L2 language immersion experiences. • Spatial release from masking partially replicated when speech clarity is high. • Effect of real room’s acoustics within certain range of speech clarity reproduced.
Yusuke Hioka, C. T. Justine Hui, Hinako Masuda, Yunqi C. Zhang, Eri Osawa, Takayuki Arai
Speech Commun.2
2025 The impact of first and second formant variations on vowel identification among elderly Japanese listeners
abstract
Correct vowel identification is important for effective speech communication. While vowel perception largely requires accurate processing and detection of spectral information, spectral processing abilities (suprathreshold deficits such as difference limens for frequency) have been shown to degrade with age, and therefore potentially affecting elderly listeners’ vowel categorisation. The current study examined a group of near-normal hearing elderly listeners matched in age, hearing thresholds, auditory filter bandwidths and cognitive performance on their difference limens for frequency (DLFs) and vowel identification. Japanese has a five-vowel system, where/o/ and/e/ approximately differ in second formant (F2) frequencies and/a/ and/u/ approximately differ in first formant (F1) frequencies. We created two continua differing in F1 and F2, respectively, and examined elderly listeners’ identification of the vowels. We found differences in DLFs within the elderly listeners’ group, where the group with lower difference limens (LDL) to have similar DLFs to a control group of young listeners. Grouping the listeners according to their DLFs via k-means clustering with the young listeners’ results, we found the two groups to differ in their/o/ -/e/ (F2) vowel perception, where we observed a significant shift in the boundary between the two groups, but not when F1 was manipulated. This suggests that for elderly listeners, vowel identification that depends on higher formant cues may be affected by frequency discrimination abilities. However, the results did not suggest a difference in accuracy, but instead, a difference in how categories are assigned. • Japanese elderly listeners’ vowel perception varying in F1 and F2 was examined. • Listeners were grouped according to their frequency discrimination abilities (DLF). • A shift in perceptual boundary observed between the groups when F2 is manipulated. • No difference when F1 is manipulated.
C. T. Justine Hui, Takayuki Arai
Speech Commun.1
2025 Role of language familiarity in understanding speech in noise under various acoustic environments
abstract
We communicate in complex acoustic environments in everyday life but our familiarity with the language can affect how well we can understand speech in these environments. The current study examines the role of language familiarity in understanding speech in varying acoustic environments via a speech intelligibility test conducted under anechoic and reverberant conditions with various speech-noise separation angles. Four groups were recruited with differing level of language familiarity: first language (L1) New Zealand English (NZE) listeners, second language (L2) Japanese native listeners with exposure to NZE, L2 Japanese native listeners with overseas English experiences without exposure to NZE, and Japanese native listeners who have learnt English as a foreign language (FL) without overseas English experiences. The L1 group performed better in overall speech intelligibility performance compared to the 3 Japanese native groups. Contrary to previous literature where non-native listeners were found to have a similar benefit from spatial separation to native listeners, this was not the case for the FL group, suggesting that this benefit is only available for listeners with a certain level of language familiarity. While there were differences between L2 and FL groups in the anechoic condition, these differences become marginal in the reverberant conditions for the two groups with little exposure to NZE. This suggests that familiarity to the specific language variety has an advantage in acoustically adverse environments.
C. T. Justine Hui, Hinako Masuda, Eri Osawa, Takayuki Arai, Catherine I. Watson, Yusuke Hioka
Speech Commun.1
2024 Performance of single-channel speech enhancement algorithms on Mandarin listeners with different immersion conditions in New Zealand English
abstract
Speech enhancement (SE) is a widely used technology to improve the quality and intelligibility of noisy speech. So far, SE algorithms were designed and evaluated on native listeners only, but not on non-native listeners who are known to be more disadvantaged when listening in noisy environments. This paper investigates the performance of five widely used single-channel SE algorithms on early-immersed New Zealand English (NZE) listeners and native Mandarin listeners with different immersion conditions in NZE under negative input signal-to-noise ratio (SNR) by conducting a subjective listening test in NZE sentences. The performance of the SE algorithms in terms of speech intelligibility in the three participant groups was investigated. The result showed that the early-immersed group always achieved the highest intelligibility. The late-immersed group outperformed the non-immersed group for higher input SNR conditions, possibly due to the increasing familiarity with the NZE accent, whereas this advantage disappeared at the lowest tested input SNR conditions. The SE algorithms tested in this study failed to improve and rather degraded the speech intelligibility, indicating that these SE algorithms may not be able to reduce the perception gap between early-, late- and non-immersed listeners, nor able to improve the speech intelligibility under negative input SNR in general. These findings have implications for the future development of SE algorithms tailored to Mandarin listeners, and for understanding the impact of language immersion on speech perception in noise.
Yunqi C. Zhang, Yusuke Hioka, C. T. Justine Hui, Catherine I. Watson
Speech Commun.3
2022 Differences between listeners with early and late immersion age in spatial release from masking in various acoustic environments
C. T. Justine Hui, Yusuke Hioka, Hinako Masuda, Catherine I. Watson
Speech Commun.1
2021 Comparing Speech Enhancement Techniques for Voice Adaptation-Based Speech Synthesis
abstract
This study investigates the use of speech enhancement techniques in creating text-to-speech voices with degraded or noisy speech. A number of synthetic voices were created using speech that was first degraded by different noise types at various signal-to-noise ratios (SNRs), then enhanced through four speech enhancement algorithms: Subspace, Wiener filter, SEGAN and a DNN-based method. Subjective listening tests show that the quality of the synthetic voices produced by subspace and the DNN-based method enhanced speech outperforms the quality of the voices created using Wiener filter or SEGAN enhanced speech at low SNRs, and speech enhanced by the subspace method results in higher quality synthetic speech at higher SNRs.
Nicholas Eng, C. T. Justine Hui, Yusuke Hioka, Catherine I. Watson
Interspeech2
2021 Effect of prior exposure on the perception of Japanese vowel length contrast in reverberation for nonnative listeners
abstract
While reverberation often degrades speech intelligibility, previous studies have shown that prior exposure to reverberation can reduce its adverse effects on speech perception. The current study investigated the effect of prior exposure to reverberation on the perception of nonnative speech sounds. We compared the results from two experiments, one in “blocked presentation” where the target words with the same amount of reverberation were presented to participants consistently, and another in “random presentation” where the amount of reverberation added to the target words changed between each trial. A Japanese minimal pair,/ie/ ‘house’ -/iie/ ‘no’, where vowel length creates a phonemic difference, was used as the target. The results for native listeners showed that their responses did not differ significantly between the blocked and random presentations. On the other hand, the effect of presentation type was significant in terms of the responses from the nonnative listeners. The results showed that nonnative listeners did not respond differently between the anechoic and reverberant conditions in the blocked presentation. However, there was a significant difference between the anechoic and reverberant conditions in the random presentation. The results from the nonnative listeners suggest that they try to obtain information of reverberation from the exposure since they could not use top-down processing effectively as much as native listeners.
Eri Osawa, C. T. Justine Hui, Yusuke Hioka, Takayuki Arai
Speech Commun.2
2019 Effects of sentence structure and word complexity on intelligibility in machine-to-human communications
C. T. Justine Hui, Sahil Jain, Catherine I. Watson
Comput. Speech Lang.1