Yusuke Hioka

dblp:79/4324 · DBLP profile ↗
← Back
32ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0003-3380-9677ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Effect of Noise Floor in Room Impulse Response on Speech Perception Under Spherical Harmonics-based Spatial Sound Reproduction
abstract
The current study investigates the effect of noise floor in measured room impulse responses (RIR) on the reproducibility of speech perception under spherical harmonics-based spatial sound reproduction. Subjective listening test measuring the intelligibility of speech in noise was conducted under the spatial sound reproduction implemented using practically measured RIR with varying level of noise floor. The same test was also conducted in the real rooms where the RIR were measured. The comparison of the experimental results from the spatial sound reproduction and the real room suggests using measured RIR with low noise floor contributes to reproducing speech perception in real rooms accurately when the room is highly reverberant. It also has an effect to improve the reproducibility when the sound sources are located at 5 m but not at 2 m. Truncating RIR to further remove the noise floor mostly did not help improve the reproducibility regardless of the acoustics of the room.
Yunqi C. Zhang, Dhruv Jagmohan, Hong Kit Li, C. T. Justine Hui, Yusuke Hioka
INTERSPEECH5
2025 Operating condition invariant representation learning for machine prognostics
abstract
Condition monitoring (CM) data can readily become complex as machines undergo continuous variations in operating conditions. This complexity poses a significant challenge to learning discriminative health-state representations. A standard solution to it is to incorporate operational parameters into learning framework, but doing so can be costly and often infeasible if such data are not available in practice. To this end, we propose a novel framework for learning health-state representations that are inherently invariant to changes in operating conditions-without relying on operational parameters. The core principle of our Operating Condition-Invariant Representation (OCIR) model is rooted in the intuition that learning to disentangle a factor of variation in data naturally leads to learning to encode representations that are invariant to the disentangled factor. We adopt an unsupervised generative model to disentangle operating condition factors at the observation level, thereby inducing invariance at the sequence level. Simultaneously, we leverage the generative model as a source of self-supervision and train a predictive model alongside it by enforcing cycle consistency in the transformation of knowledge between the two models. Experimental results demonstrate that the health-state representations learned through OCIR are highly competitive with those learned using operational parameters, while significantly outperforming methods that do not utilize such information. Additionally, we introduce a novel method for construction of virtually stationary trajectories directly from the raw CM data subject to varying operating conditions.
Yusuke Hioka, Michael Witbrock
Knowl. Based Syst.2
2025 Using spatial sound reproduction for studying speech perception of listeners with different language immersion experiences
abstract
This study evaluates a research method for studying speech perception of listeners with different language background under practical acoustic environments. The proposed research method utilises spatial sound reproduction, an emerging technology that enables reproducing arbitrary acoustic environments in controlled laboratory settings, for testing participants recruited at multiple locations that are geographically distant from each other. To validate the research method, the current study conducted a listening test in a real seminar room and chapel as well as under a spherical harmonics-based spatial sound reproduction that reproduced the acoustics of the two venues up to the third order and investigates differences in the results collected from the two test types. Three groups of participants who had had different immersion level to New Zealand English were recruited in Auckland, New Zealand and Tokyo, Japan. The experimental results show that spatial sound reproduction is able to capture the advantage of first language (L1) listeners in terms of understanding speech in noise and reverberation correctly but is not sensitive enough to describe the subtle difference among second language (L2) listeners with different level of language immersion experiences. The research method is also partially able to describe how well listeners can benefit from spatial release from masking regardless of their language immersion experiences under room acoustics with higher speech clarity (C50), and may represent the effect of room acoustics in the real room within a certain range of room acoustics characterised by speech clarity. • Evaluates reproducibility of L2 speech perception using spatial sound reproduction. • L1 advantage in understanding speech in noise and reverberation correctly captured. • Not sensitive enough to describe difference by L2 language immersion experiences. • Spatial release from masking partially replicated when speech clarity is high. • Effect of real room’s acoustics within certain range of speech clarity reproduced.
Yusuke Hioka, C. T. Justine Hui, Hinako Masuda, Yunqi C. Zhang, Eri Osawa, Takayuki Arai
Speech Commun.1
2025 Role of language familiarity in understanding speech in noise under various acoustic environments
abstract
We communicate in complex acoustic environments in everyday life but our familiarity with the language can affect how well we can understand speech in these environments. The current study examines the role of language familiarity in understanding speech in varying acoustic environments via a speech intelligibility test conducted under anechoic and reverberant conditions with various speech-noise separation angles. Four groups were recruited with differing level of language familiarity: first language (L1) New Zealand English (NZE) listeners, second language (L2) Japanese native listeners with exposure to NZE, L2 Japanese native listeners with overseas English experiences without exposure to NZE, and Japanese native listeners who have learnt English as a foreign language (FL) without overseas English experiences. The L1 group performed better in overall speech intelligibility performance compared to the 3 Japanese native groups. Contrary to previous literature where non-native listeners were found to have a similar benefit from spatial separation to native listeners, this was not the case for the FL group, suggesting that this benefit is only available for listeners with a certain level of language familiarity. While there were differences between L2 and FL groups in the anechoic condition, these differences become marginal in the reverberant conditions for the two groups with little exposure to NZE. This suggests that familiarity to the specific language variety has an advantage in acoustically adverse environments.
C. T. Justine Hui, Hinako Masuda, Eri Osawa, Takayuki Arai, Catherine I. Watson, Yusuke Hioka
Speech Commun.6
2024 Performance of single-channel speech enhancement algorithms on Mandarin listeners with different immersion conditions in New Zealand English
abstract
Speech enhancement (SE) is a widely used technology to improve the quality and intelligibility of noisy speech. So far, SE algorithms were designed and evaluated on native listeners only, but not on non-native listeners who are known to be more disadvantaged when listening in noisy environments. This paper investigates the performance of five widely used single-channel SE algorithms on early-immersed New Zealand English (NZE) listeners and native Mandarin listeners with different immersion conditions in NZE under negative input signal-to-noise ratio (SNR) by conducting a subjective listening test in NZE sentences. The performance of the SE algorithms in terms of speech intelligibility in the three participant groups was investigated. The result showed that the early-immersed group always achieved the highest intelligibility. The late-immersed group outperformed the non-immersed group for higher input SNR conditions, possibly due to the increasing familiarity with the NZE accent, whereas this advantage disappeared at the lowest tested input SNR conditions. The SE algorithms tested in this study failed to improve and rather degraded the speech intelligibility, indicating that these SE algorithms may not be able to reduce the perception gap between early-, late- and non-immersed listeners, nor able to improve the speech intelligibility under negative input SNR in general. These findings have implications for the future development of SE algorithms tailored to Mandarin listeners, and for understanding the impact of language immersion on speech perception in noise.
Yunqi C. Zhang, Yusuke Hioka, C. T. Justine Hui, Catherine I. Watson
Speech Commun.2
2023 Rotor Noise-Aware Noise Covariance Matrix Estimation for Unmanned Aerial Vehicle Audition
abstract
A noise covariance matrix (NCM) estimation method for unmanned aerial vehicle (UAV) audition is proposed with rotor noise reduction as its primary focus. The proposed NCM estimation method could be incorporated into audio processing algorithms using UAV-mounted microphone array systems. The NCM is formed through accurate estimation of the microphone array input signal's amplitude and phase by using a multi-sensory rotor noise power spectral density (PSD) estimator and a filter formed by exploiting the acoustical relationship between the microphone array and the rotor noise sources, respectively. The estimated NCM aims to be readily incorporable into several source enhancement algorithms to reduce the effects of rotor noise and improve the resultant audio quality. Experiment evaluation using real in-flight UAV rotor noise recordings shows that the estimated NCM significantly improves rotor noise reduction (of up to$\sim 28$dB) and the quality of the target sound.
Benjamin Yen 0001, Yameizhen Li, Yusuke Hioka
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Design of a low-cost passive acoustic monitoring system for animal localisation from calls
abstract
The field of bioacoustics is concerned with monitoring wild animals based on their vocalisations. Passive acoustic recorders are now commonly used to collect data of the soundscapes of our wild places. While the data they collect is extremely useful, the majority of the recorders use a single omnidirectional microphone, and thus cannot independently perform localisation of a calling animal. Localisation can be useful to differentiate between multiple calling animals, to improve statistical estimates of abundance, and to locate calling posts, which may be close to nests. In this paper, we consider the design of a low-cost, practical, passive directional acoustic recorder that will facilitate animal localisation, and present and evaluate a prototype system for this purpose.
Benjamin Yen 0001, Jemima Prins, Gian Schmid, Yusuke Hioka, Susan Ellis, Stephen R. Marsland
IROS4
2022 Differences between listeners with early and late immersion age in spatial release from masking in various acoustic environments
C. T. Justine Hui, Yusuke Hioka, Hinako Masuda, Catherine I. Watson
Speech Commun.2
2021 Comparing Speech Enhancement Techniques for Voice Adaptation-Based Speech Synthesis
abstract
This study investigates the use of speech enhancement techniques in creating text-to-speech voices with degraded or noisy speech. A number of synthetic voices were created using speech that was first degraded by different noise types at various signal-to-noise ratios (SNRs), then enhanced through four speech enhancement algorithms: Subspace, Wiener filter, SEGAN and a DNN-based method. Subjective listening tests show that the quality of the synthetic voices produced by subspace and the DNN-based method enhanced speech outperforms the quality of the voices created using Wiener filter or SEGAN enhanced speech at low SNRs, and speech enhanced by the subspace method results in higher quality synthetic speech at higher SNRs.
Nicholas Eng, C. T. Justine Hui, Yusuke Hioka, Catherine I. Watson
Interspeech3
2021 Effect of prior exposure on the perception of Japanese vowel length contrast in reverberation for nonnative listeners
abstract
While reverberation often degrades speech intelligibility, previous studies have shown that prior exposure to reverberation can reduce its adverse effects on speech perception. The current study investigated the effect of prior exposure to reverberation on the perception of nonnative speech sounds. We compared the results from two experiments, one in “blocked presentation” where the target words with the same amount of reverberation were presented to participants consistently, and another in “random presentation” where the amount of reverberation added to the target words changed between each trial. A Japanese minimal pair,/ie/ ‘house’ -/iie/ ‘no’, where vowel length creates a phonemic difference, was used as the target. The results for native listeners showed that their responses did not differ significantly between the blocked and random presentations. On the other hand, the effect of presentation type was significant in terms of the responses from the nonnative listeners. The results showed that nonnative listeners did not respond differently between the anechoic and reverberant conditions in the blocked presentation. However, there was a significant difference between the anechoic and reverberant conditions in the random presentation. The results from the nonnative listeners suggest that they try to obtain information of reverberation from the exposure since they could not use top-down processing effectively as much as native listeners.
Eri Osawa, C. T. Justine Hui, Yusuke Hioka, Takayuki Arai
Speech Commun.3
2019 Improving speech intelligibility using microphones on behind the ear hearing aids
abstract
Hearing aids have a great potential to facilitate better speech communication not only for hearing impaired users but also for people with normal hearing since the device will allow users to render sound with better speech intelligibility using signal processing techniques. This paper studies how a sound source separation technique would enable improving the intelligibility of a targeted speech when the technique is applied to the behind-the-ear hearing aids which has multiple microphones on each device. The sound source separation technique utilised in this study is based on the beamforming with post-filter framework and separates sound arriving from different direction. Experimental results using microphones attached to a head and torso simulator suggest the speech intelligibility can be improved by emphasising the target speech while suppressing sound from other angles,
Yusuke Hioka, Kei Kobayashi, Kenta Niwa
MMSP1
2018 End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain
abstract
This paper presents an end-to-end deep neural network (DNN)-based source enhancement on the basis of a time-frequency (T-F) mask processing in the modified discrete cosine transform (MDCT)-domain. To retrieve the target signal perfectly in the discrete Fourier transform (DFT)-domain, both amplitude and phase of the spectrum need to be manipulated. However, since it is difficult to deal with complex values by neural network straightforward way, a real-valued T-F mask is commonly estimated and only amplitude spectrum is manipulated. In this study, we use the MDCT instead of the DFT and estimate real-valued T-F masks in the MDCT-domain. The perfect retrieval can be achieved by manipulating only the real-valued MDCT-spectra. To reduce time-domain aliasing arises from manipulating the MDCT spectrum, we build an end-to-end DNN-based source enhancement using T-F mask and train the DNN to minimize an objective function defined in the time-domain. In experiments using several kinds of objective sound quality scores, we observed that the scores were significantly improved.
Yuma Koizumi, Noboru Harada, Youichi Haneda, Yusuke Hioka, Kazunori Kobayashi
ICASSP4
2018 DNN-Based Source Enhancement to Increase Objective Sound Quality Assessment Score
abstract
We propose a training method for deep neural network (DNN) based source enhancement to increase objective sound quality assessment (OSQA) scores such as the perceptual evaluation of speech quality. In many conventional studies, DNNs have been used as a mapping function to estimate time-frequency masks and trained to minimize an analytically tractable objective function such as the mean squared error (MSE). Since OSQA scores have been used widely for sound-quality evaluation, constructing DNNs to increase OSQA scores would be better than using the minimum MSE to create high-quality output signals. However, since most OSQA scores are not analytically tractable, i.e., they are black boxes, the gradient of the objective function cannot be calculated by simply applying backpropagation. To calculate the gradient of the OSQA-based objective function, we formulated a DNN optimization scheme on the basis of black-box optimization, which is used for training a computer that plays a game. For a black-box-optimization scheme, we adopt the policy gradient method for calculating the gradient on the basis of a sampling algorithm. To simulate output signals using the sampling algorithm, DNNs are used to estimate the probability density function of the output signals that maximize OSQA scores. The OSQA scores are calculated from the simulated output signals, and the DNNs are trained to increase the probability of generating the simulated output signals that achieve high OSQA scores. Through several experiments, we found that OSQA scores significantly increased by applying the proposed method, even though the MSE was not minimized.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Youichi Haneda
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Efficient Audio Rendering Using Angular Region-Wise Source Enhancement for 360° Video
abstract
In virtual reality, 360° video services provided through head-mounted displays or smartphones are widely available. Among these, some state-of-the-art devices are able to render varying auditory location of an object perceived by the user when the visual location of the object in the video moves along with the change of the user's looking direction. Nevertheless, an acoustic immersion technology that generates binaural sound to maintain a good match between the auditory and visual localization of an object in 360° video has not been studied sufficiently. This study focuses on an approach that synthesizes semibinaural sound being composed of virtual sources located in each angular region and the representative head related transfer functions of each angular region. To minimize the calculation cost on audio rendering and to reduce latency in downloading data from servers, the number of angular regions should be reduced while maintaining a good match between the auditory and visual localization of an object. In this paper, we investigate the minimum number of angular regions at which it is possible to maintain a good match by conducting subjective tests using a 360° video viewing system composed of virtual images and sound sources. From the subjective tests, it was confirmed that the acoustic field should be divided into more than six equispaced angular regions so as to achieve natural auditory localization that matches an object's location in 360° video.
Kenta Niwa, Yusuke Hioka, Hisashi Uematsu
IEEE Trans. Multim.2
2017 DNN-based source enhancement self-optimized by reinforcement learning using sound quality measurements
abstract
We investigated whether a deep neural network (DNN)-based source enhancement function can be self-optimized by reinforcement learning (RL). The use of a DNN is a powerful approach to describing the relationship between two sets of variables and can be useful for source enhancement function design. By training the DNN using a huge amount of training data, sound quality of output signals are improved. However, collecting a huge amount of training data is often difficult in practice. To use limited training data efficiently, we focus on the “self-optimization” of DNN-based source enhancement function in which RL is commonly utilized in the development of game playing computers. As a reward for RL, quantitative metrics that reflect a human's perceptual score (perceptual score), e.g., perceptual evaluation methods for audio source separation (PEASS), are utilized. To investigate whether the sound quality is improved by RL-based source enhancement, subjective tests were conducted. It was confirmed that the output sound quality of the RL-based source enhancement function improved as the number of iterations was increased and finally outperformed the conventional method.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Youichi Haneda
ICASSP3
2017 Supervised source enhancement composed of nonnegative auto-encoders and complementarity subtraction
abstract
A method for constructing deep neural networks (DNNs) for accurate supervised source enhancement is proposed. Attempts were made in previous studies to estimate the power spectral densities (PSDs) of sound sources, which are used to estimate Wiener filters for source enhancement, from the output of multiple beamformings using DNNs. Although performance improved, it was not possible to guarantee accurate PSD estimation since the trained DNNs were treated as black boxes. The proposed DNN construction method uses non-negative auto-encoders and complementarity subtraction. This study also reveals that auto-encoders whose weights are non-negative correspond to non-negative matrix factorization (NMF), which decomposes source PSDs into non-negative spectral bases and their activations. It further introduces a complementarity subtraction method for estimating PSDs accurately. Through several experiments, it was confirmed that the signal-to-interference plus noise ratio improved by approximately 12 dB for datasets captured in various noisy/reverberant rooms.
Kenta Niwa, Yuma Koizumi, Tomoko Kawase, Kazunori Kobayashi, Yusuke Hioka
ICASSP5
2017 Informative Acoustic Feature Selection to Maximize Mutual Information for Collecting Target Sources
abstract
An informative acoustic-feature-selection method for collecting target sources in noisy environments is proposed. Wiener filtering is a powerful framework for sound-source enhancement. For Wiener-filter estimation, statistical-mapping functions, such as deep neural network based or Gaussian mixture model based mappings, have been used. In this framework, it is essential to find informative acoustic features that provide effective cues for Wiener-filter estimation. In this study, we measured the informativeness of acoustic features using mutual information between acoustic features and supervised Wiener-filter parameters, e.g., prior signal-to-noise ratios, and developed a method for automatically selecting informative acoustic features from a large number of feature candidates. To automatically select optimum features, we derived a differentiable objective function in proportion to mutual information based on the kernel method. Since the higher order correlations between acoustic features and Wiener-filter parameters are calculated using the kernel method, the statistical dependence of these variables is accurately calculated; thus, only meaningful acoustic features are selected. Through several experiments conducted on a mock sports field, we confirmed that the signal-to-distortion ratio score improved when various types of target sources were surrounded by loud cheering noise.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Hitoshi Ohmuro
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Estimating direct-to-reverberant ratio mapped from power spectral density using deep neural network
abstract
A new attempt for estimating the direct-to-reverberant ratio (DRR) by mapping the power spectral density (PSD) of the direct sound and reverberation using the deep neural network is reported. The method finds the correct DRR from the PSD estimated with an algorithm using a microphone array. The experimental results using a recording of a reverberant speech signal, which included various environmental noise, reveal that the proposed method is effective in improving the accuracy of DRR estimation and robust against various noise.
Yusuke Hioka, Kenta Niwa
ICASSP1
2016 Integrated approach of feature extraction and sound source enhancement based on maximization of mutual information
abstract
We investigated informative acoustic feature extraction based on dimension reduction for collecting target sources on a noisy sports field. Although a Wiener filter is often used for sound source enhancement, it is difficult to accurately design the Wiener filter by simply using spatial cues because the noise on a sports field (e.g., cheering from spectators) arrives from the same direction as that of the targeted source. A statistical approach is used to estimate the Wiener filter by using pre-trained acoustic feature models. However, an informative acoustic feature, which provides a powerful clue for clear extraction of the target source, is unknown. For this study, we developed a method for optimizing a projection matrix for dimension reduction by maximizing the mutual information between acoustic features and the Wiener filter. Through experiments using two-directional microphones on a mock sports field, we confirmed that the proposed method outperformed previous methods in terms of both the noise reduction and quality of the recovered sound sources.
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Hitoshi Ohmuro
ICASSP3
2016 Pinpoint extraction of distant sound source based on DNN mapping from multiple beamforming outputs to prior SNR
abstract
We propose a method for estimating the prior signal-to-noise ratio (SNR), which is used for calculating the Wiener filter for distant sound source extraction, from output signals of beamforming using statistical mapping based on the deep neural network (DNN). Since informative features to estimate the prior SNR are included in multiple beamforming outputs, the SNR can be accurately estimated by this mapping using the DNN. The proposed method was applied to a large microphone array, the design of which was optimized to form effective directivity patterns to extract distant sound sources. Experimental results proved that the target source was clearly extracted with the proposed method.
Kenta Niwa, Yuma Koizumi, Tomoko Kawase, Kazunori Kobayashi, Yusuke Hioka
ICASSP5
2016 Optimal Microphone Array Observation for Clear Recording of Distant Sound Sources
abstract
We propose the principle for deriving an optimum design for a microphone array that uses mutual information to segregate distant sound sources. Many conventional studies on array signal processing have focused on methods for estimating sound sources from array observations. To record distant sound sources clearly, designing an optimum array structure to segregate a target from other noise is also necessary. In this study, we reveal that the optimum array observation was achieved by receiving signals that are physically decorrelated between microphones, which homogenizes the eigenvalues of the spatial correlation matrix. We theoretically explain this underlying principle using mutual information between sound sources and microphone observations whose relation to the existing minimum mean square error criterion for source separation is also discussed. The implementation of such a microphone array is possible by placing microphones in front of parabolic reflectors since the phase/amplitude around the focal point of the reflectors drastically varies with small perturbation of the microphone position. Crosscorrelation between observed signals can be reduced by optimally placing microphones. An array structure based on our proposed principle was tested by implementing minimum variance distortion-less response beamforming and postfiltering in the observations of a prototype microphone array. We experimentally confirmed that 1) the eigenvalues of the spatial correlation matrix were asymptotically homogenized and 2) the target source could be extracted clearly even when the sound sources were positioned 16.5 m from the array.
Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Underdetermined Sound Source Separation Using Power Spectrum Density Estimated by Combination of Directivity Gain
abstract
A method for separating underdetermined sound sources based on a novel power spectral density (PSD) estimation is proposed. The method enables up toM(M-1)+1 sources to be separated when we use a microphone array ofMsensors and a Wiener post-filter calculated by the estimated PSDs. The PSD of a beamformer's output is modelled by a mixture of source PSDs multiplied by the beamformer's directivity gain in the particular angle where each source is located. Based on this model, the PSD of each sound source is estimated from the PSD of multiple fixed beamformers' outputs using the difference in the combination of directivity gains. Simulation results proved that the proposed method effectively separated up toM(M-1)+1 sound sources if the fixed beamformers were appropriately selected. Experiments were also conducted in a reverberant chamber to ensure the proposed method was also effective in practical use.
Yusuke Hioka, Ken'ichi Furuya, Kazunori Kobayashi, Kenta Niwa, Youichi Haneda
IEEE Trans. Speech Audio Process.1
2013 Diffused Sensing for Sharp Directive Beamforming
abstract
We generalized our previously proposed diffused sensing for a microphone array design to achieve sharp directive beamforming to enable various filter design methods to be applied. In the conventional microphone array, various filter design methods have been studied to narrow the directivity beam width. However, it is difficult to minimize the power of interference sources in the beamforming output (output interference power) over a broad frequency range since the cross-correlation between transfer functions from sound sources to microphones increases in some frequencies. With the diffused sensing, the cross-correlation is minimized by physically varying the transfer functions. We investigated how a microphone array should be designed in order to minimize the cross-correlation between transfer functions and found that placing the array in a diffuse acoustic field produces optimum results. Because the transfer functions are known a priori, this finding makes it possible to narrow the directivity beam width over a broad frequency range. This observation can be practically achieved by placing microphones inside a reflective enclosure, part of which is open to let sound waves enter. We conducted experiments using 24 microphones and confirmed that the output interference power was reduced over a broad frequency range and the beam width was narrowed by using the diffused sensing.
Kenta Niwa, Yusuke Hioka, Ken'ichi Furuya, Youichi Haneda
IEEE Trans. Speech Audio Process.2
2012 Efficient crosstalk canceler design with impulse response shortening filters
abstract
An impulse response shortening approach is used to perform acoustic crosstalk cancellation. Crosstalk canceler filters are traditionally designed using least squares, with an approach that equalizes all room reverberation. However, depending upon end application, some reverberation may be permissible in the delivered signals. This idea is used to create more efficient crosstalk cancellation filters. The filter design is formulated as a minimax problem solvable with linear programming methods. Penalty functions on crosstalk levels and detrimental reverberation are introduced, which allow control of the reverberant tails and crosstalk levels. Shorter crosstalk cancellation filters are designed, by leaving in early echoes and/or allowing a slower decay of the late reverberant tail.
Terence Betlehem, Paul D. Teal, Yusuke Hioka
ICASSP3
2012 Telescopic microphone array using reflector for segregating target source from noises in same direction
abstract
A spatial sensitivity control method for segregating the sound sources in the same direction by using an acoustic reflector is proposed. Our goal is to clearly pick up the target source at an arbitrary position using a microphone array. Though many methods have been studied for spatial sensitivity control, it is difficult to robustly suppress the power of noise sources in the same direction of the target source in a room. To overcome this problem, we attach a reflector to a microphone array to capture the reflected sounds whose characteristics vary depending on the distance from the array to the source. Assuming that the acoustical properties of the reflector are known e.g., measuring the transfer functions, those reflected sounds can be used as effective clues for segregating the sound sources in the same direction. With the proposed method, a filter for minimizing the output noise power is derived by taking into consideration of the acoustic properties of the reflector. Experiments were conducted in an actual room by using 96 microphones and a large reflector. We confirmed that the spatial sensitivities for segregating the target source at an arbitrary position from noise sources can be achieved by using the proposed method.
Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda
ICASSP2
2011 Distributed blind source separation with an application to audio signals
abstract
A scalable blind source separation paradigm aimed at sensor networks is described. The approach facilitates an unlimited number of sensors and sources and does not require a fusion centre. It is based on a so-called ownership principle, where each network node aims to extract (own) a source signal that is not already extracted (owned) by another network node. Nodes that own a source signal broadcast that signal to user nodes outside the network. Nodes that do not currently own a source signal do not transmit information and can be active intermittently. A natural application of the method is a distributed microphone network in a multi-talker environment, with as user nodes hearing aids or telephone interface devices. Such a network can stretch across buildings or neighbourhoods. Simulations using independent component analysis (ICA) indicate the validity of the principles of the method.
Yusuke Hioka, W. Bastiaan Kleijn
ICASSP1
2011 Estimating Direct-to-Reverberant Energy Ratio Using D/R Spatial Correlation Matrix Model
abstract
We present a method for estimating the direct-to-reverberant energy ratio (DRR) that uses a direct and reverberant sound spatial correlation matrix model (Hereafter referred to as the spatial correlation model). This model expresses the spatial correlation matrix of an array input signal as two spatial correlation matrices, one for direct sound and one for reverberation. The direct sound propagates from the direction of the sound source but the reverberation arrives from every direction uniformly. The DRR is calculated from the power spectra of the direct sound and reverberation that are estimated from the spatial correlation matrix of the measured signal using the spatial correlation model. The results of experiment and simulation confirm that the proposed method gives mostly correct DRR estimates unless the sound source is far from the microphone array, in which circumstance the direct sound picked up by the microphone array is very small. The method was also evaluated using various scales in simulated and actual acoustical environments, and its limitations revealed. We estimated the sound source distance using a small microphone array, which is an example of application of the proposed DRR estimation method.
Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda
IEEE ACM Trans. Audio Speech Lang. Process.1
2010 Estimating direct-to-reverberant energy ratio based on spatial correlation model segregating direct sound and reverberation
abstract
A new approach for estimating the direct-to-reverberant energy ratio (DRR) using a microphone array is proposed. The method is based on a model of a spatial correlation matrix that segregates direct sound and reverberation. It estimates DRR from the power spectra of both components, which are derived from the correlation matrix of the observed signal. In experiments performed in simulated and actual reverberant environments, the proposed method mostly succeeded in estimating DRR accurately. We also present speech enhancement using binary masking as an example of an application of the estimated DRR. By utilization of the DRR as a factor to discriminate the distances of speakers, separation of speech signals whose sources were located in the same direction but at different distances was achieved.
Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda
ICASSP1
2010 Estimation of sound source orientation using eigenspace of spatial correlation matrix
abstract
We propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of spatial correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method.
Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda
ICASSP2
2009 Speech enhancement in a 2-dimensional area based on power spectrum estimation of multiple areas with investigation of existence of active sources
Yusuke Hioka, Ken'ichi Furuya, Youichi Haneda, Akitoshi Kataoka
INTERSPEECH1
2004 Separate estimation of azimuth and elevation DOA using microphones located at apices of regular tetrahedron
abstract
We propose a method for DOA (direction of arrival) estimation of speech signals in both azimuth and elevation directions. Our previous DOA estimation method achieves high precision and uniform spatial resolution by integrating the frequency array data (Hioka, Y. et al., J. Sig. Process., vol.7, no.1, p.105-9, 2003) generated from microphone pairs in an equilateral-triangular microphone array (Hioka and Hamada, N., IEICE Technical Report, EA2003-44, p.9-16, 2003). We now extend the method using four microphones located at the apices of a regular tetrahedron to enable the method to estimate the elevation angle from the array plane as well. Furthermore, we introduce an idea for separate estimation of azimuth and elevation to reduce the computational loads.
Yusuke Hioka, Nozomu Hamada
ICASSP (2)1
2003 DOA estimation of speech signal using equilateral-triangular microphone array
abstract
In this contribution, we propose a DOA (Direction Of Arrival) estimation method of speech signal whose angular resolution is almost uniform with respect to DOA. Our previous DOA estimation method[1] achieves high precision with only two microphones, however its resolution degrades as the propagating direction apart from the array broadside. In the proposed method, the equilateral-triangular microphone array is adopted, and the subspace analysis is applied. The efficiency of the proposed method is shown both from the simulation and experimental results.
Yusuke Hioka, Nozomu Hamada
INTERSPEECH1