Shoichi Koyama

dblp:01/10499 · DBLP profile ↗
← Back
41ranked-venue papers
14as first author
14since 2021 · last 2025
0000-0003-2283-0884ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Past, Present, and Future of Spatial Audio and Room Acoustics
abstract
The study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have been developed based both on theoretical advancements and practical innovations. We highlight historical achievements, initiative activities, recent advancements, and future outlooks in the research area of spatial audio recording and reproduction, and room acoustic simulation, modeling, analysis, and control.
Shoichi Koyama, Enzo De Sena, Prasanga N. Samarasinghe, Mark R. P. Thomas, Fabio Antonacci
ICASSP1
2024 Localizing Acoustic Energy in Sound Field Synthesis by Directionally Weighted Exterior Radiation Suppression
abstract
A method for synthesizing the desired sound field while suppressing the exterior radiation power with directional weighting is proposed. The exterior radiation from the loudspeakers in sound field synthesis systems can be problematic in practical situations. Although several methods to suppress the exterior radiation have been proposed, suppression in all outward directions is generally difficult, especially when the number of loudspeakers is not sufficiently large. We propose the directionally weighted exterior radiation representation to prioritize the suppression directions by incorporating it into the optimization problem of sound field synthesis. By using the proposed representation, the exterior radiation in the prioritized directions can be significantly reduced while maintaining high interior synthesis accuracy, owing to the relaxed constraint on the exterior radiation. Its performance is evaluated with the application of the proposed representation to amplitude matching in numerical experiments.
Yoshihide Tomita, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2024 Sound Field Estimation Based on Physics-Constrained Kernel Interpolation Adapted to Environment
abstract
A sound field estimation method based on kernel interpolation with an adaptive kernel function is proposed. The kernel-interpolation-based sound field estimation methods enable physics-constrained interpolation from pressure measurements of distributed microphones with a linear estimator, which constrains interpolation functions to satisfy the Helmholtz equation. However, a fixed kernel function would not be capable of adapting to the acoustic environment in which the measurement is performed, limiting their applicability. To make the kernel function adaptive, we represent it with a sum of directed and residual trainable kernel functions. The directed kernel is defined by a weight function composed of a superposition of exponential functions to capture highly directional components. The weight function for the residual kernel is represented by neural networks to capture unpredictable spatial patterns of the residual components. Experimental results using simulated and real data indicate that the proposed method outperforms the current kernel-interpolation-based methods and a method based on physics-informed neural networks.
Juliano G. C. Ribeiro, Shoichi Koyama, Ryosuke Horiuchi, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Spatial Active Noise Control Method Based on Sound Field Interpolation from Reference Microphone Signals
abstract
A spatial active noise control (ANC) method based on the interpolation of a sound field from reference microphone signals is proposed. In most current spatial ANC methods, a sufficient number of error microphones are required to reduce noise over the target region because the sound field is estimated from error microphone signals. However, in practical applications, it is preferable that the number of error microphones is as small as possible to keep a space in the target region for ANC users. We propose to interpolate the sound field using reference microphones, which are normally placed outside the target region, instead of the error microphones. We derive a fixed filter for spatial noise reduction on the basis of the kernel ridge regression for sound field interpolation. Furthermore, to compen-sate for estimation errors, we combine the proposed fixed filter with multichannel ANC based on a transition of the control filter using the error microphone signals. Numerical experimental results indicate that regional noise can be sufficiently reduced by the proposed methods even when the number of error microphones is particularly small.
Kazuyuki Arikawa, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2023 Kernel Interpolation of Acoustic Transfer Functions with Adaptive Kernel for Directed and Residual Reverberations
abstract
An interpolation method for region-to-region acoustic transfer functions (ATFs) based on kernel ridge regression with an adaptive kernel is proposed. Most current ATF interpolation methods do not incorporate the acoustic properties for which measurements are performed. Our proposed method is based on a separate adaptation of directional weighting functions to directed and residual reverberations, which are used for adapting kernel functions. Thus, the proposed method can not only impose constraints on fundamental acoustic properties, but can also adapt to the acoustic environment. Numerical experimental results indicated that our proposed method outperforms the current methods in terms of interpolation accuracy, especially at high frequencies.
Juliano G. C. Ribeiro, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2023 Amplitude Matching for Multizone Sound Field Control
abstract
A multizone sound field control method, called amplitude matching, is proposed. The objective of amplitude matching is to synthesize a desired amplitude (or magnitude) distribution over a target region with multiple loudspeakers, whereas the phase distribution is arbitrary. Most of the current multizone sound field control methods are intended to synthesize a specific sound field including phase or to control acoustic potential energy inside the target region. In amplitude matching, a specific desired amplitude distribution can be set, ignoring sound propagation directions. Although the optimization problem of amplitude matching does not have a closed-form solution, our proposed algorithm based on the alternating direction method of multipliers (ADMM) allows us to accurately and efficiently synthesize the desired amplitude distribution. We also introduce the differential-norm penalty for a time-domain filter design with a small filter length. The experimental results indicated that the proposed method outperforms current multizone sound field control methods in terms of accuracy of the synthesized amplitude distribution.
Takumi Abe, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Spatial Active Noise Control Based on Individual Kernel Interpolation of Primary and Secondary Sound Fields
abstract
A spatial active noise control (ANC) method based on the individual kernel interpolation of primary and secondary sound fields is proposed. Spatial ANC is aimed at cancelling unwanted primary noise within a continuous region by using multiple secondary sources and microphones. A method based on the kernel interpolation of a sound field makes it possible to attenuate noise over the target region with flexible array geometry. Furthermore, by using the kernel function with directional weighting, prior information on primary noise source directions can be taken into consideration. However, whereas the sound field to be interpolated is a superposition of primary and secondary sound fields, the directional weight for the primary noise source was applied to the total sound field in previous work; therefore, the performance improvement was limited. We propose a method of individually interpolating the primary and secondary sound fields and formulate a normalized least-mean-square algorithm based on this interpolation method. Experimental results indicate that the proposed method outperforms the method based on total kernel interpolation.
Kazuyuki Arikawa, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2022 Variable Span Trade-Off Filter for Sound Zone Control with Kernel Interpolation Weighting
abstract
A sound zone control method is proposed, based on the frequency domain variable span trade-off filter (VAST). Existing VAST methods optimizes the sound field at a set of discrete points, while the proposed method uses kernel interpolation to instead optimize the sound field over a continuous region. When the loudspeaker positions are known, the performance can be improved further by applying a directional weighting to the interpolation procedure. The proposed method is evaluated by simulating broadband sound in a reverberant environment, focusing on the case when microphone placement is restricted. The proposed method with directional weighting outperforms the pointwise VAST over the full bandwidth of the signal, and the proposed method without directional weighting outperforms the pointwise VAST at low frequencies.
Jesper Brunnström, Shoichi Koyama, Marc Moonen
ICASSP2
2022 Region-to-Region Kernel Interpolation of Acoustic Transfer Function with Directional Weighting
abstract
A method of interpolating the acoustic transfer function (ATF) between regions that takes into account both the physical properties of the ATF and the directionality of region configurations is proposed. Most spatial ATF interpolation methods are limited to estimation in the region of receivers. A kernel method for region-to-region ATF interpolation makes it possible to estimate the ATFs for both source and receiver regions from a discrete set of ATF measurements. We newly formulate the reproducing kernel Hilbert space and associated kernel function incorporating directional weight to enhance the interpolation accuracy. We also investigate hyperparameter optimization methods for this kernel function. Numerical experiments indicate that the proposed method outperforms the method without the use of directional weighting.
Juliano G. C. Ribeiro, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2022 Region-to-Region Kernel Interpolation of Acoustic Transfer Functions Constrained by Physical Properties
abstract
A method to interpolate the acoustic transfer function (ATF) between regions using kernel ridge regression (KRR) is proposed. Conventionally, the ATF interpolation problem is strongly restricted and situational, depending on knowledge of environmental conditions while not accounting for source position variation. We derive our interpolation function as the solution of an optimization problem defined on a function space where every element holds the acoustic properties of the ATF. By making the space a reproducing kernel Hilbert space (RKHS), we can guarantee that our problem has a known and unique optimizer. The generality of the formulation of this method enables region-to-region estimations, with variable source and receiver within the assigned bounds. The definition of a RKHS also allows for the use of kernel principal component analysis, thereby efficiently providing greater noise robustness to our interpolation function. Our proposed method is compared with a previously established region-to-region interpolation method in numerical simulations where the advantages of the KRR approach are confirmed, showing lower error and greater stability for higher frequencies.
Juliano G. C. Ribeiro, Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Kernel-Interpolation-Based Filtered-X Least Mean Square for Spatial Active Noise Control In Time Domain
abstract
Time-domain spatial active noise control (ANC) algorithms based on kernel interpolation of a sound field are proposed. Spatial ANC aims to suppress a primary noise field in a region by synthesizing an anti-noise field using secondary loudspeakers. In contrast to the multipoint pressure control by conventional ANC methods, it is necessary to estimate a sound field from microphone measurements at discrete positions. A promising approach is the kernel-interpolation-based spatial ANC method, which allows for flexible array configurations. However, most of the current methods are formulated only in the frequency domain, which cannot handle broadband noise with straightforward implementation. We formulate two time-domain spatial ANC algorithms based on kernel interpolation using a spatial interpolation filter defined by the kernel function. Numerical simulation results indicate that our filtered-x least-mean-square-based algorithms are effective for reducing nonstationary broadband noise in a region, as compared with multipoint pressure control.
Jesper Brunnström, Shoichi Koyama
ICASSP2
2021 Amplitude Matching: Majorization-Minimization Algorithm for Sound Field Control Only with Amplitude Constraint
abstract
A sound field control method for synthesizing a desired amplitude distribution inside a target region, amplitude matching, is proposed. In the conventional pressure matching, a desired sound field is set as a pressure distribution including amplitude and phase. In personal audio applications, it is sometimes not necessary to synthesize a specific phase distribution, but a certain acoustic power level should be controlled inside the target region. Since the optimization problem to achieve amplitude matching becomes nonlinear, there is no closed-form solution and numerical optimization algorithms are generally applied. We derive an efficient algorithm for the amplitude matching based on the majorization–minimization algorithm. Numerical experiments indicated that high control accuracy over the target region can be achieved with low computational cost by using the proposed algorithm.
Shoichi Koyama, Takashi Amakasu, Natsuki Ueno, Hiroshi Saruwatari
ICASSP1
2021 Spatial Active Noise Control Based on Kernel Interpolation of Sound Field
abstract
An active noise control (ANC) method to reduce noise over a region in space based on kernel interpolation of sound field is proposed. Current methods of spatial ANC are largely based on spherical or circular harmonic expansion of the sound field, where the geometry of the error microphone array is restricted to a simple one such as a sphere or circle. We instead apply the kernel interpolation method, which allows for the estimation of a sound field in a continuous region with flexible array configurations. The interpolation scheme is used to derive adaptive filtering algorithms for minimizing the acoustic potential energy inside a target region. A practical time-domain algorithm is also developed together with its computationally efficient block-based equivalent. We conduct experiments to investigate the achievable level of noise reduction in a two-dimensional free space, as well as adaptive broadband noise control in a three-dimensional reverberant space. The experimental results indicated that the proposed method outperforms the multipoint-pressure-control-based method in terms of regional noise reduction.
Shoichi Koyama, Jesper Brunnström, Hayato Ito, Natsuki Ueno, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Multichannel Blind Source Separation Based on Evanescent-Region-Aware Non-Negative Tensor Factorization in Spherical Harmonic Domain
abstract
There is growing interest in new audio formats in the context of virtual reality (VR), and higher-order ambisonics (HOA) is preferred for VR systems to transmit recorded scenes owing to its transmission efficiency and its flexibility to work with different loudspeaker setups. However, the conversion between another well-known format, i.e., object format, and the HOA format is not fully addressed in the literature. To address this issue, blind source separation in a spherical harmonic (SH) domain can be considered as the best way to extract objects in terms of efficiency, i.e., decoding HOA signals for separation can be omitted. A few authors attempted to extract objects from encoded HOA signals directly by using multichannel non-negative matrix factorization (MNMF), but these approaches either assume only far-field sources or do not take array characteristics into account, which make these methods difficult to use for VR in practical situations where singers or speakers often perform close to microphones. Furthermore, MNMF generally requires a huge computational cost, although dimensional reduction to the SH domain is performed. In this work, we also model near-field sources by estimating the model parameters of non-negative tensor factorization (NTF) in the SH domain assuming that microphone signals can be obtained with a rigid spherical array. We propose a masking scheme to exclude noisy evanescent regions in the SH domain from the NTF cost function. Evaluations show that our method outperforms existing methods devised for the HOA format and that our masking approach is effective in improving the separation quality.
Yuki Mitsufuji, Norihiro Takamune, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Mutual-Information-Based Sensor Placement for Spatial Sound Field Recording
abstract
A sensor (microphone) placement method based on mutual information for spatial sound field recording is proposed. The sound field recording methods using distributed sensors enable the estimation of the sound field inside a target region of arbitrary shape; however, it is a difficult task to find the best placement of sensors. We focus on the mutual-information-based sensor placement method in which spatial phenomena are modeled as a Gaussian process (GP). We propose the use of the sound-field-interpolation kernel for the covariance of measurements in a GP model to obtain the sensor placement suitable for sound field recording. We also extend the method to treat broadband signals and derive an efficient algorithm based on block matrix inversion. Numerical simulation results indicated that the proposed method achieves accurate sound field estimation compared with a method using the generally used Gaussian kernel.
Kentaro Ariga, Tomoya Nishida, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP3
2020 Spatial Active Noise Control Based on Kernel Interpolation with Directional Weighting
abstract
A spatial active noise control (ANC) method taking prior information on the approximate direction of primary noise sources into consideration is proposed. ANC aims to cancel incoming primary noise using secondary loudspeakers. Conventional multipoint ANC does not guarantee the reduction of noise between multiple discrete control points; therefore, several attempts have been made to reduce the noise over an entire target region, i.e., by spatial ANC. We have recently proposed a spatial ANC method based on kernel ridge regression for sound field interpolation using distributed microphones and loudspeakers, where the cost function is formulated on the basis of the regional power obtained by kernel interpolation. In this study, we incorporate prior knowledge on the noise source direction into spatial ANC based on the kernel interpolation with directional weighting. Numerical simulation results indicate that the proposed method can achieve larger regional noise reduction than the methods without the information on noise source direction.
Hayato Ito, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP2
2020 Binaural Rendering From Distributed Microphone Signals Considering Loudspeaker Distance in Measurements
abstract
A method of binaural rendering from distributed microphone recordings that takes loudspeaker distance for measuring head-related transfer function (HRTF) into consideration is proposed. In general, to reproduce the binaural signals from the signals captured by multiple microphones in the recording area, the captured sound field is represented by plane-wave decomposition. Thus, HRTF is approximated as a transfer function from a plane-wave source in binaural rendering. To incorporate the distance in HRTF measurements, we propose a method based on the spherical-wave decomposition of a sound field, in which the HRTF is assumed to be measured from a point source. Result of experiments using HRTFs calculated by the boundary element method indicated that the accuracy of binaural signal reproduction by the proposed method based on the spherical-wave decomposition was higher than that by the plane-wave-decomposition-based method. We also evaluate the performance of signal conversion from distributed microphone measurements into binaural signals.
Naoto Iijima, Shoichi Koyama, Hiroshi Saruwatari
MMSP2
2020 Reciprocity gap functional in spherical harmonic domain for gridless sound field decomposition
abstract
A sound field decomposition method based on the reciprocity gap functional (RGF) in the spherical harmonic domain is proposed. To estimate and reconstruct a continuous sound field including sources by using multiple microphones, an intuitive and powerful strategy is to decompose the sound field into Green’s functions. Sparse-representation algorithms have been applied to this decomposition problem; however, it requires the discretization of the target region into grid points to construct a dictionary matrix. Discretization-based methods lead to decomposition errors of off-grid sources and high computational cost of sparse representation. We apply the RGF to sparse sound field decomposition, which makes it possible to decompose the sound field as a closed-form solution without discretization. In addition, the formulation in the spherical harmonic domain enables the flexible arrangement of microphones under the assumption of the spherical target region. Numerical simulation results indicated that high decomposition and reconstruction accuracies can be achieved by the proposed method, especially at low frequencies, with a low computational cost.
Yuhta Takida, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
Signal Process.2
2020 Optimizing Source and Sensor Placement for Sound Field Control: An Overview
abstract
In order to control an acoustic field inside a target region, it is important to choose suitable positions of secondary sources (loudspeakers) and sensors (control points/microphones). This article provides an overview of state-of-the-art source and sensor placement methods in sound field control. Although the placement of both sources and sensors greatly affects control accuracy and filter stability, their joint optimization has not been thoroughly investigated in the acoustics literature. In this context, we reformulate five general source and/or sensor placement methods that can be applied for sound field control. We compare the performance of these methods through extensive numerical simulations in both narrowband and broadband scenarios.
Shoichi Koyama, Gilles Chardon, Laurent Daudet
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Multichannel Non-Negative Matrix Factorization Using Banded Spatial Covariance Matrices in Wavenumber Domain
abstract
Blind source separation exploiting multichannel information has long been a popular topic, and recently proposed methods based on the local Gaussian model have shown promising results despite its high computational cost for the case of many microphone signals. The low updating speed for such a model is mainly due to the inversion of a spatial covariance matrix, for which the complexity increases with the number of microphones, M, and is generally of order O(M3). Several projection-based approaches that attempt to concentrate energy on the diagonal part of the spatial covariance matrix have been introduced to circumvent the matrix inversion, which can reduce the complexity to O(M). In this article, we focus on the fast Fourier transform as a projection method because the energy concentration on the diagonal can be efficiently achieved compared with other projection-based methods. For the case where the diagonalization is imperfect, for example, owing to discontinuities at the edge of a linear array, we also developed a more robust algorithm approximating the tri-diagonal part of the spatial covariance matrix, which requires a complexity of O(M2) for the inversion by applying the Thomas algorithm. To remove the ad-hoc integration of post clustering after the decomposition, we also examine a self-clustering algorithm. Our evaluation shows better results than other previously proposed methods in terms of the separation quality under reverberant conditions as well as higher efficiency than multichannel non-negative matrix factorization.
Yuki Mitsufuji, Stefan Uhlich, Norihiro Takamune, Daichi Kitamura, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Feedforward Spatial Active Noise Control Based on Kernel Interpolation of Sound Field
abstract
A method for feedforward active noise control (ANC) over a spatial region is proposed. Conventional multipoint ANC aims to reduce the noise at multiple discrete positions; therefore, the noise reduction in the region between these points cannot be guaranteed. Recent studies revealed the possibility of spatial ANC, i.e., noise control in a continuous target region. These methods are essentially based on spherical/circular harmonic decomposition of the sound field by using spherical/circular arrays and have mainly been investigated for feedback control under the assumption of periodicity of the noise. We apply a sound field interpolation method based on kernel ridge regression to feedforward spatial ANC to control spatial nonstationary noise using distributed arrays. Numerical simulation results indicated that a large regional noise reduction is achieved by the proposed method compared with feedforward multipoint ANC.
Hayato Ito, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP2
2019 Robust Gridless Sound Field Decomposition Based on Structured Reciprocity Gap Functional in Spherical Harmonic Domain
abstract
A sound field reconstruction method for a region including sources is proposed. Under the assumption of spatial sparsity of the sources, this reconstruction problem has been solved by using sparse decomposition algorithms with the discretization of the target region. Since this discretization leads to the off-grid problem, we previously proposed a gridless sound field decomposition method based on the reciprocity gap functional in the spherical harmonic domain. Even though this method allows efficient estimation with a closed-form solution while avoiding the off-grid problem, the estimation using a single time-frequency bin can be greatly affected by measurement errors. We formulate an optimization problem using the identical structure of source locations in multiple time-frequency bins and derive an algorithm based on an annihilating filter. Numerical simulation results indicated that robustness against noise can be improved by the proposed method.
Yuhta Takida, Shoichi Koyama, Natsuki Ueno, Hiroshi Saruwatari
ICASSP2
2019 Three-Dimensional Sound Field Reproduction Based on Weighted Mode-Matching Method
abstract
A sound field reproduction method based on the spherical wavefunction expansion of sound fields is proposed, which can be flexibly applied to various array geometries and directivities. First, we formulate sound field synthesis as a minimization problem of some norm on the difference between the desired and synthesized sound fields, and then the optimal driving signals are derived by using the spherical wavefunction expansion of the sound fields. This formulation is closely related to the mode-matching method; a major advantage of the proposed method is the optimal weight on the mode determined according to the norm to be minimized instead of the empirical truncation in the mode-matching method. We also provide some examples of norms and their corresponding weights in analytical forms. Both interior and exterior sound field reproduction are considered in the proposed method, and some applications, such as multizone reproduction and interior reproduction with exterior cancellation, are also discussed. Numerical simulation results indicated that higher reproduction accuracy can be achieved by the proposed method than by the current pressure-matching and mode-matching methods.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Joint Source and Sensor Placement for Sound Field Control Based on Empirical Interpolation Method
abstract
This study proposes a principled method to jointly determine the placement of acoustic sources (loudspeakers) and sensors (control points/microphones) in sound field control. The goal of this setup is to efficiently produce a sound field using multiple loudspeakers, approximately matching a target sound field over a region of interest. Therefore, the loudspeaker and control-point placement problem can be seen as the problem of finding interpolating functions (associated with individual loudspeaker sound fields) and sampling points (corresponding to control points or microphones) to approximate the target sound field in the given domain. We here solve this problem using the empirical interpolation method, originally developed for the numerical analysis of partial differential equations. The proposed method enables a joint determination of loudspeaker and control-point placement, from a large set of candidate locations, independently of the desired sound field. Numerical simulation results indicate that accurate and stable sound field control can be achieved by the proposed method, with significantly better results than with random and regular placements.
Shoichi Koyama, Gilles Chardon, Laurent Daudet
ICASSP1
2018 Sound Field Reproduction with Exterior Cancellation Using Analytical Weighting of Harmonic Coefficients
abstract
A method for sound field reproduction with the suppression of exterior radiation is proposed, which makes it possible to synthesize a desired sound field in a reverberant environment without prior knowledge of the transfer functions of the multiple loudspeakers. The objective function used to achieve this is formulated as the weighted sum of the interior reproduction error and exterior radiation power. The optimal driving signals are derived by harmonic expansion of both the interior and exterior sound fields. In contrast to the empirical coefficient truncation in the state of the art, in the proposed method, an optimal weighting of the harmonic coefficients is derived analytically. Numerical simulation results indicated that high interior reproduction accuracy and exterior power suppression can be achieved by the proposed method compared with the mode-matching method using harmonic-order truncation owing to the optimal weighting.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2018 Sound Field Recording Using Distributed Microphones Based on Harmonic Analysis of Infinite Order
abstract
A sound field recording method based on spherical or circular harmonic analysis for arbitrary array geometry and directivity of microphones is proposed. In current methods based on harmonic analysis, a sound field is decomposed into harmonic functions with a center given in advance, which is called a global origin, and their coefficients are obtained up to a certain truncation order using microphone measurements. However, the accuracy of the reconstructed sound field depends on the predefined position of the global origin and the truncation order, which makes it difficult to apply this technique to an asymmetric array since the criterion to determine the position of the global origin and the truncation order is not obvious. We formulate an estimate of the harmonic coefficients on the basis of infinite-order analysis. This formulation enables us to estimate the harmonic coefficients at an arbitrary desired position independently of the position of the global origin without truncation errors. Numerical simulation results indicated that the proposed method makes it possible to avoid performance degradation caused by inappropriate setting of the global origin.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
IEEE Signal Process. Lett.2
2017 Spatio-temporal sparse sound field decomposition considering acoustic source signal characteristics
abstract
We propose a sound field decomposition method that takes into consideration spatio-temporal sparsity. It has been proved that sparse representation of a sound field is effective in reducing errors originating from spatial aliasing artifacts compared with conventional plane wave decomposition. In most current methods of sparse sound field decomposition, the spatial sparsity of the sound source distribution is only assumed. However, it is known that the temporal structure of the source signal to be decomposed can also be sparse in the time-frequency domain. We formulate an objective function for sparse sound field decomposition by using the ℓp,q-norm to simultaneously induce sparsity in the space and time domains. An optimization algorithm on the auxiliary function method is derived to solve it. Numerical simulations of acoustic holography indicate that the reconstruction accuracy can be improved by controlling the parameter of temporal sparsity. We also demonstrate that a statistical measure of the source signals can be used as an indicator to determine a nearly optimal parameter.
Naoki Murata, Shoichi Koyama, Norihiro Takamune, Hiroshi Saruwatari
ICASSP2
2017 Listening-area-informed sound field reproduction based on circular harmonic expansion
abstract
A sound field reproduction method that exploits prior information on listening areas is proposed. Most current methods are aimed at reproducing the sound field over the entire space or around listener locations. We formulate the objective function for this problem as the expectation minimization of the spatial squared error of the sound pressure inside the listening areas. The optimal driving signals are obtained by circular harmonic expansion. Comparing the proposed method with the mode-matching method, the advantage of the proposed method appears in the optimal weighting matrix of the circular harmonics, which depends on the location and range of the listening area. Numerical simulations indicated that high reproduction accuracy compared with other current methods can be maintained in the listening areas at high frequencies by using the proposed method.
Natsuki Ueno, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2016 Sound field decomposition in reverberant environment using sparse and low-rank signal models
abstract
A sound field decomposition method for a reverberant environment is proposed. Sound field decomposition is the foundation of various acoustic signal processing applications and enables the estimation of the entire sound field from pressure measurements. Although spatial Fourier analysis of the sound field has been widely used, sparse decomposition of the sound field has recently been proved to be effective in several applications. However, in current methods, no constraints are imposed on ambiance components, whereas source components are assumed to be sparsely distributed in the space. This results in inaccurate decomposition in a reverberant environment. The proposed method is based on sparse and low-rank signal models, which are used for simultaneous decomposition of the observed signals into source and ambiance components. Numerical simulation results indicated that the decomposition accuracy is superior to that of current methods.
Shoichi Koyama, Hiroshi Saruwatari
ICASSP1
2016 Multichannel blind source separation based on non-negative tensor factorization in wavenumber domain
abstract
Multichannel non-negative matrix factorization based on a spatial covariance model is one of the most promising techniques for blind source separation. However, this approach is not tractable for a large number of microphones, M, because the computational cost is of order O(M3) per time-frequency bin. To circumvent this drawback, we propose non-negative tensor factorization in the wavenumber domain, which reduces the cost to the order O(M). It transforms microphone signals into the spatial frequency domain, a technique that is commonly used for soundfield reconstruction. The proposed method is compared to several blind source separation (BSS) methods in terms of separation quality and computational cost.
Yuki Mitsufuji, Shoichi Koyama, Hiroshi Saruwatari
ICASSP2
2016 Sparse sound field decomposition with multichannel extension of complex NMF
abstract
A sparse sound field decomposition method using prior information on source signals in the time-frequency domain is proposed. Sparse sound field decomposition has been proved to be effective for various acoustic signal processing applications. Current methods for sparse decomposition are based only on the spatial sparsity of the source distribution. However, it can be assumed that possible source signals to be decomposed are approximately known in advance. To exploit this prior information, we incorporated the complex nonnegative factorization model into sparse sound field decomposition. Since the magnitude spectrum of the possible source signals can be trained in advance, accuracy of the sparse decomposition can be improved even when the source signals are highly correlated and the sources are in a highly noisy environment. In addition, the proposed decomposition algorithm is derived using the auxiliary function method. Numerical experiments indicated that the sparse decomposition performance was significantly improved using the proposed method.
Naoki Murata, Shoichi Koyama, Hirokazu Kameoka, Norihiro Takamune, Hiroshi Saruwatari
ICASSP2
2015 Structured sparse signal models and decomposition algorithm for super-resolution in sound field recording and reproduction
abstract
A method for achieving super-resolution of sound field recording and reproduction is proposed. To obtain driving signals of loudspeakers for reproduction from received signals of microphones, sparse signal decomposition makes it possible to reduce spatial aliasing artifacts when the number of microphones is less than that of loudspeakers. For more accurate and robust signal decomposition, we propose three types of group sparse signal model based on the physical properties of a sound field. In addition, a decomposition algorithm is derived to address these signal models as an extension of M-FOCUSS. In the simulation experiments, the accuracy of the sparse decomposition was significantly improved compared with that of M-FOCUSS. Furthermore, the accuracy of sound field reproduction using our proposed method was higher than that using current methods, especially at frequencies above the spatial Nyquist frequency.
Shoichi Koyama, Naoki Murata, Hiroshi Saruwatari
ICASSP1
2015 Statistical modeling of binaural signal and its application to binaural source separation
abstract
This paper addresses a new statistical model of binaural signals and its application to efficient binaural source separation. Binaural source separation is always required to retain a spatial cue of the separated sound, such as a head-related transfer function (HRTF). However, the direct use of an HRTF is not realistic because this information is normally not known in advance. To cope with this problem, first, we focus on the difference between signal probability density functions at both ears, which can be blindly estimated by using our previous work on higher-order statistics. Next, we derive a sound-localization-preserved generalized minimum mean-square error short-time spectral amplitude estimator. Objective and subjective experiments show the efficacy of the proposed method in terms of spatial quality.
Yuki Murota, Daichi Kitamura, Shoichi Koyama, Hiroshi Saruwatari, Satoshi Nakamura 0001
ICASSP3
2014 Close-talking spherical microphone array using sound pressure interpolation based on spherical harmonic expansion
abstract
We propose a novel close-talking spherical microphone array that uses the residual signal between the observed sound pressure and the interpolated sound pressure at the center of the spherical array. The interpolated sound is obtained from the sound pressures observed on the surface of a sphere on the basis of the spherical harmonic expansion, assuming that the sound originates from the outside of the array. If the sound source is close to the spherical array, the array cannot express the spherical wave correctly because the number of microphones is limited. As a result, the residual signal increases. This method is a modified form of the conventional method, which interpolates the sound pressure by using the Kirchhoff integral equation. In contrast with the conventional method, we interpolate the sound at the center of the sphere by using only the average value of the sound pressures on the spherical array surface. The computer simulations were conducted using a 12-element spherical microphone array with radius of 5 cm. These results showed that the performances of both methods were almost equivalent, although the proposed method used half the number of microphones as the conventional method.
Youichi Haneda, Ken'ichi Furuya, Shoichi Koyama, Kenta Niwa
ICASSP3
2014 Sparse sound field representation in recording and reproduction for reducing spatial aliasing artifacts
abstract
We propose a sound-pressure-to-driving-signal (SP-DS) conversion method for sound field reproduction based on sparse sound field representation. The most important problem in sound field reproduction is how to calculate driving signals of loudspeakers to reproduce desired sound fields. In common recording and reproduction systems, sound pressures at multiple positions obtained in a recording area are only known as the desired sound field; therefore, SP-DS conversion algorithms are necessary. Current SP-DS conversion methods do not take into account sound sources to be reproduced, which results in severe spatial aliasing artifacts. Our proposed method decomposes the received sound pressure distribution based on the generative model of the sound field. Numerical simulation results indicate that the proposed method can achieve higher reproduction accuracy compared to the current methods, especially in higher frequencies above the spatial Nyquist frequency.
Shoichi Koyama, Suehiro Shimauchi, Hitoshi Ohmuro
ICASSP1
2014 Wave Field Reconstruction Filtering in Cylindrical Harmonic Domain for With-Height Recording and Reproduction
abstract
For sound field reproduction that includes height (with-height reproduction), it is more efficient to record and reproduce the sound field with lower resolution in elevation than in azimuth due to the spatial abilities of human auditory perception. We propose a sound field reproduction method using horizontally arranged cylindrical arrays of microphones and loudspeakers, which is based on the wave field reconstruction (WFR) filter. With the use of cylindrical array configurations, it is possible to reproduce sound waves arriving from upper and lower directions with a smaller number of array elements at angular positions. The WFR filter is analytically derived in the cylindrical harmonic domain and allows direct transformation from the received signals of the microphones into the driving signals of the loudspeakers. A model in which microphones are mounted on a rigid cylindrical baffle is introduced to stabilize the WFR filter. Numerical simulation results indicated that the reproduction accuracy in the neighboring region along the central axis of the cylindrical array was better preserved when using the proposed method than when the method with planar arrays was used.
Shoichi Koyama, Ken'ichi Furuya, Yusuke Hiwasaki, Youichi Haneda, Yôiti Suzuki
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Sound field reproduction using multiple linear arrays based on wave field reconstruction filtering in helicalwave spectrum domain
abstract
For with-height reproduction of a sound field, an efficient technique is to record and reproduce the sound field at lower resolution at an elevation angle than that at a horizontal angle based on auditory perception. To achieve this, developing a method using multiple horizontal linear arrays of microphones and loudspeakers is necessary. We propose a sound field reproduction method for a cylindrical array configuration which is based on a wave field reconstruction (WFR) filter analytically derived in the helical wave spectrum domain. The filter is stabilized by introducing a model in which microphones are mounted on a rigid cylindrical baffle. Numerical simulation results indicated that the reproduction accuracy of the proposed method was well preserved in the near-field region along the central axis of the cylinder even when the number of elements at angular position is small.
Shoichi Koyama, Ken'ichi Furuya, Yusuke Hiwasaki, Youichi Haneda, Yôiti Suzuki
ICASSP1
2013 Improvement using circular harmonics beamforming on reverberation problem of wave field reconstruction filtering
abstract
In real-time sound field transmission systems, the driving signals of a loudspeaker array should be obtained using only the received signals of a microphone array. For efficient transformation of signals of planar or linear arrays, we previously proposed a method of applying a transform filter in the spatio-temporal frequency domain; this filter is defined as the wave field reconstruction (WFR) filter. In linear array configurations, the major artifact of this method is an increase in reverberation time and a decrease in the direct-to-reverberant energy ratio (DRR) because reflections from above and below the microphone array can not be distinguished. We propose a method combining circular harmonics beamforming with the WFR filter as a preprocess in order to match the reproduced DRR to the original one at the time the direct sound wave is properly reproduced. Simulation results indicated that the DRR reproduced using the proposed method was much closer to that reproduced by the method without beamforming.
Shoichi Koyama, Timothy Lee, Ken'ichi Furuya, Yusuke Hiwasaki, Youichi Haneda
ICASSP1
2013 Analytical Approach to Wave Field Reconstruction Filtering in Spatio-Temporal Frequency Domain
abstract
For transmission of a physical sound field in a large area, it is necessary to transform received signals of a microphone array into driving signals of a loudspeaker array to reproduce the sound field. We propose a method for transforming these signals by using planar or linear arrays of microphones and loudspeakers. A continuous transform equation is analytically derived based on the physical equation of wave propagation in the spatio-temporal frequency domain. By introducing spatial sampling, the uniquely determined transform filter, called a wave field reconstruction filter (WFR filter), is derived. Numerical simulations show that the WFR filter can achieve the same performance as that obtained using the conventional least squares (LS) method. However, since the proposed WFR filter is represented as a spatial convolution, it has many advantages in filter design, filter size, computational cost, and filter stability over the transform filter designed by the LS method.
Shoichi Koyama, Ken'ichi Furuya, Yusuke Hiwasaki, Youichi Haneda
IEEE Trans. Speech Audio Process.1
2012 Design of transform filter for reproducing arbitrarily shifted sound field using phase-shift of spatio-temporal frequency
abstract
For real-time sound field transmission systems from a far-end to a near-end, the driving signals of a loudspeaker array at the near-end need to be calculated by using only received signals obtained by a microphone array at the far-end. Additionally, having the capability to control the location of the sound field to be reproduced in order to adjust it to the visual images is advantageous. The goal of this study was to develop a method to transform received signals of a microphone array into driving signals of a loudspeaker array in order to reproduce arbitrary shifted sound fields. We analytically derive a transform filter in the spatio-temporal frequency domain. The location of the sound field to be reproduced is controllable only by phase-shift of the transform filter. The proposed method was found to be computationally efficient compared to the conventional method based on a least squares algorithm, and numerical simulation results indicated that reproduction accuracies were almost the same in both methods.
Shoichi Koyama, Ken'ichi Furuya, Yusuke Hiwasaki, Youichi Haneda
ICASSP1
2012 Reproducing Virtual Sound Sources in Front of a Loudspeaker Array Using Inverse Wave Propagator
abstract
It has been possible to reproduce point sound sources between listeners and a loudspeaker array by using the focused-source method. However, this method requires physical parameters of the sound sources to be reproduced, such as source positions, directions, and original signals. This fact makes it difficult to apply the method to real-time reproduction systems because decomposing received signals into such parameters is not a trivial task. This paper proposes a method for recreating virtual sound sources in front of a planar or linear loudspeaker array. The method is based on wave field synthesis but extended to include the inverse wave propagator often used in acoustical holography. Virtual sound sources can be placed between listeners and a loudspeaker array even when the received signals of a microphone array equally aligned with the loudspeaker array are only known. Numerical simulation results are presented to compare the proposed and focused-source methods. A system was constructed consisting of linear microphone and loudspeaker arrays and measurement experiments were conducted in an anechoic room. When comparing the sound field reproduced using the proposed method with that using the focused-source method, it was found that the proposed method could reproduce the sound field at almost the same accuracy.
Shoichi Koyama, Ken'ichi Furuya, Yusuke Hiwasaki, Youichi Haneda
IEEE Trans. Speech Audio Process.1