EDBT 2026 Demo / reviewers in the wild / expert
Kenta Niwa
dblp:64/1008
· DBLP profile ↗
47ranked-venue papers
18as first author
14since 2021 · last 2026
0000-0002-6911-0238ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 first-author · 4 since 2021Artificial intelligence and machine learning · 21 · 6 first-author · 13 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedPM: Federated Learning Using Second-order Optimization with Preconditioned Mixing of Local ParametersabstractWe propose Federated Preconditioned Mixing (FedPM), a novel Federated Learning (FL) method that leverages second-order optimization. Prior methods - such as LocalNewton, LTDA, and FedSophia - have incorporated second-order optimization in FL by performing iterative local updates on clients and applying simple mixing of local parameters on the server. However, these methods often suffer from drift in local preconditioners, which significantly disrupts the convergence of parameter training, particularly in heterogeneous data settings. To overcome this issue, we refine the update rules by decomposing the ideal second-order update - computed using globally preconditioned global gradients - into parameter mixing on the server and local parameter updates on clients. As a result, our FedPM introduces preconditioned mixing of local parameters on the server, effectively mitigating drift in local preconditioners. We provide a theoretical convergence analysis demonstrating a superlinear rate for strongly convex objectives in scenarios involving a single local update. To demonstrate the practical benefits of FedPM, we conducted extensive experiments. The results showed significant improvements with FedPM in the test accuracy compared to conventional methods incorporating simple mixing, fully leveraging the potential of second-order optimization. Hiro Ishii, Kenta Niwa, Hiroshi Sawada, Akinori Fujino, Noboru Harada, Rio Yokota |
AAAI | 2 |
| 2025 | Plausible Token Amplification for Improving Accuracy of Differentially Private In-Context Learning Based on Implicit Bayesian InferenceabstractWe propose Plausible Token Amplification (PTA) to improve the accuracy of Differentially Private In-Context Learning (DP-ICL) using DP synthetic demonstrations. While Tang et al. empirically improved the accuracy of DP-ICL by limiting vocabulary space during DP synthetic demonstration generation, its theoretical basis remains unexplored. By interpreting ICL as implicit Bayesian inference on a concept underlying demonstrations, we not only provide theoretical evidence supporting Tang et al.’s empirical method but also introduce PTA, a refined method for modifying next-token probability distribution. Through the modification, PTA highlights tokens that distinctly represent the ground-truth concept underlying the original demonstrations. As a result, generated DP synthetic demonstrations guide the Large Language Model to successfully infer the ground-truth concept, which improves the accuracy of DP-ICL. Experimental evaluations on both synthetic and real-world text-classification datasets validated the effectiveness of PTA. Yusuke Yamasaki, Kenta Niwa, Daiki Chijiwa, Takumi Fukami, Takayuki Miura |
ICML | 2 |
| 2025 | Revisiting 1-peer exponential graph for enhancing decentralized learning efficiencyabstractFor communication-efficient decentralized learning, it is essential to employ dynamic graphs designed to improve the expected spectral gap by reducing deviations from global averaging. The $1$-peer exponential graph demonstrates its finite-time convergence property--achieved by maximizing the expected spectral gap--but only when the number of nodes $n$ is a power of two. However, its efficiency across any $n$ and the commutativity of mixing matrices remain unexplored. We delve into the principles underlying the $1$-peer exponential graph to explain its efficiency across any $n$ and leverage them to develop new dynamic graphs. We propose two new dynamic graphs: the $k$-peer exponential graph and the null-cascade graph. Notably, the null-cascade graph achieves finite-time convergence for any $n$ while ensuring commutativity. Our experiments confirm the effectiveness of these new graphs, particularly the null-cascade graph, in most test settings. Kenta Niwa, Yuki Takezawa, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
NeurIPS | 1 |
| 2024 | Optimal Transport with Cyclic SymmetryabstractWe propose novel fast algorithms for optimal transport (OT) utilizing a cyclic symmetry structure of input data. Such OT with cyclic symmetry appears universally in various real-world examples: image processing, urban planning, and graph processing. Our main idea is to reduce OT to a small optimization problem that has significantly fewer variables by utilizing cyclic symmetry and various optimization techniques. On the basis of this reduction, our algorithms solve the small optimization problem instead of the original OT. As a result, our algorithms obtain the optimal solution and the objective function value of the original OT faster than solving the original OT directly. In this paper, our focus is on two crucial OT formulations: the linear programming OT (LOT) and the strongly convex-regularized OT, which includes the well-known entropy-regularized OT (EROT). Experiments show the effectiveness of our algorithms for LOT and EROT in synthetic/real-world data that has a strict/approximate cyclic symmetry structure. Through theoretical and experimental results, this paper successfully introduces the concept of symmetry into the OT research field for the first time. Shoichiro Takeda, Yasunori Akagi, Naoki Marumo, Kenta Niwa |
AAAI | 4 |
| 2024 | On Accelerating Diffusion-Based Sampling Processes via Improved Integration ApproximationabstractA popular approach to sample a diffusion-based generative model is to solve an ordinary differential equation (ODE). In existing samplers, the coefficients of the ODE solvers are pre-determined by the ODE formulation, the reverse discrete timesteps, and the employed ODE methods. In this paper, we consider accelerating several popular ODE-based sampling processes (including EDM, DDIM, and DPM-Solver) by optimizing certain coefficients via improved integration approximation (IIA). We propose to minimize, for each time step, a mean squared error (MSE) function with respect to the selected coefficients. The MSE is constructed by applying the original ODE solver for a set of fine-grained timesteps, which in principle provides a more accurate integration approximation in predicting the next diffusion state. The proposed IIA technique does not require any change of a pre-trained model, and only introduces a very small computational overhead for solving a number of quadratic optimization problems. Extensive experiments show that considerably better FID scores can be achieved by using IIA-EDM, IIA-DDIM, and IIA-DPM-Solver than the original counterparts when the neural function evaluation (NFE) is small (i.e., less than 25). Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
ICLR | 2 |
| 2024 | Simple Minimax Optimal Byzantine Robust Algorithm for Nonconvex Objectives with Uniform Gradient HeterogeneityabstractIn this study, we consider nonconvex federated learning problems with the existence of Byzantine workers. We propose a new simple Byzantine robust algorithm called Momentum Screening. The algorithm is adaptive to the Byzantine fraction, i.e., all its hyperparameters do not depend on the number of Byzantine workers. We show that our method achieves the best optimization error of $O(\delta^2\zeta_\mathrm{max}^2)$ for nonconvex smooth local objectives satisfying $\zeta_\mathrm{max}$-uniform gradient heterogeneity condition under $\delta$-Byzantine fraction, which can be better than the best known error rate of $O(\delta\zeta_\mathrm{mean}^2)$ for local objectives satisfying $\zeta_\mathrm{mean}$-mean heterogeneity condition when $\delta \leq (\zeta_\mathrm{max}/\zeta_\mathrm{mean})^2$. Furthermore, we derive an algorithm independent lower bound for local objectives satisfying $\zeta_\mathrm{max}$-uniform gradient heterogeneity condition and show the minimax optimality of our proposed method on this class. In numerical experiments, we validate the superiority of our method over the existing robust aggregation algorithms and verify our theoretical results. Tomoya Murata, Kenta Niwa, Takumi Fukami, Iifan Tyou |
ICLR | 2 |
| 2024 | Parameter-free Clipped Gradient Descent Meets PolyakabstractGradient descent and its variants are de facto standard algorithms for training machine learning models. As gradient descent is sensitive to its hyperparameters, we need to tune the hyperparameters carefully using a grid search. However, the method is time-consuming, particularly when multiple hyperparameters exist. Therefore, recent studies have analyzed parameter-free methods that adjust the hyperparameters on the fly. However, the existing work is limited to investigations of parameter-free methods for the stepsize, and parameter-free methods for other hyperparameters have not been explored. For instance, although the gradient clipping threshold is a crucial hyperparameter in addition to the stepsize for preventing gradient explosion issues, none of the existing studies have investigated parameter-free methods for clipped gradient descent. Therefore, in this study, we investigate the parameter-free methods for clipped gradient descent. Specifically, we propose Inexact Polyak Stepsize, which converges to the optimal solution without any hyperparameters tuning, and its convergence rate is asymptotically independent of $L$ under $L$-smooth and $(L_0, L_1)$-smooth assumptions of the loss function, similar to that of clipped gradient descent with well-tuned hyperparameters. We numerically validated our convergence results using a synthetic function and demonstrated the effectiveness of our proposed methods using LSTM, Nano-GPT, and T5. Yuki Takezawa, Han Bao 0002, Ryoma Sato, Kenta Niwa, Makoto Yamada |
NeurIPS | 4 |
| 2024 | DP-Norm: Differential Privacy Primal-Dual Algorithm for Decentralized Federated LearningabstractA novel algorithm is proposed for highly privacy-preserving decentralized federated learning (FL). Several studies have reported security risks in decentralized FL by reconstructing data even from model update differences. A common approach to overcome this issue is to use the diffusion process following differential privacy (DP), i.e., message passing between nodes is hidden by noise. However, this often makes the learning process unstable, leading to degraded results compared to without using DP diffusion process. In this paper, we propose a primal-dual DP algorithm with denoising normalization (DP-Norm) for less sensitivity to noise/interference, such as DP diffusion and heterogeneous data allocation. For DP-Norm, privacy analysis to determine minimal noise level and convergence analysis are conducted. Through image classification benchmark tests, we confirmed that DP-Norm performed close to the single-node reference score, even when statistically heterogeneous data was allocated on six nodes. Takumi Fukami, Tomoya Murata, Kenta Niwa, Iifan Tyou |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | SSFG: Stochastically Scaling Features and Gradients for Regularizing Graph Convolutional NetworksabstractGraph convolutional networks (GCNs) have been successfully applied in various graph-based tasks. In a typical graph convolutional layer, node features are updated by aggregating neighborhood information. Repeatedly applying graph convolutions can cause the oversmoothing issue, i.e., node features at deep layers converge to similar values. Previous studies have suggested that oversmoothing is one of the major issues that restrict the performance of GCNs. In this article, we propose a stochastic regularization method to tackle the oversmoothing problem. In the proposed method, we stochastically scale features and gradients (SSFG) by a factor sampled from a probability distribution in the training procedure. By explicitly applying a scaling factor to break feature convergence, the oversmoothing issue is alleviated. We show that applying stochastic scaling at the gradient level is complementary to that applied at the feature level to improve the overall performance. Our method does not increase the number of trainable parameters. When used together with ReLU, our SSFG can be seen as a stochastic ReLU activation function. We experimentally validate our SSFG regularization method on three commonly used types of graph networks. Extensive experimental results on seven benchmark datasets for four graph-based tasks demonstrate that our SSFG regularization is effective in improving the overall performance of the baseline graph networks. The code is available at https://github.com/vailatuts/SSFG-regularization. Haimin Zhang 0001, Min Xu 0001, Guoqiang Zhang 0003, Kenta Niwa |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Lookahead Diffusion Probabilistic Models for Refining Mean EstimationabstractWe propose lookahead diffusion probabilistic models (LA-DPMs) to exploit the correlation in the outputs of the deep neural networks (DNNs) over subsequent timesteps in diffusion probabilistic models (DPMs) to refine the mean estimation of the conditional Gaussian distributions in the backward process. A typical DPM first obtains an estimate of the original data sample$x$by feeding the most recent state$z_{i}$and index$i$into the DNN model and then computes the mean vector of the conditional Gaussian distribution for$z_{i-1}$. We propose to calculate a more accurate estimate for$x$by performing extrapolation on the two estimates of$x$that are obtained by feeding ($z_{i+1},i+1$) and ($z_{i},i$) into the DNN model. The extrapolation can be easily integrated into the backward process of existing DPMs by introducing an additional connection over two consecutive timesteps, and fine-tuning is not required. Extensive experiments showed that plugging in the additional connection into DDPM, DDIM, DEIS, S-PNDM, and high-order DPM-Solvers leads to a significant performance gain in terms of Fréchet inception distance (FID) score. Our implementation is available at https://github.com/guoqiang-zhang-x/LA-DPM. Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
CVPR | 2 |
| 2023 | Beyond Exponential Graph: Communication-Efficient Topologies for Decentralized Learning via Finite-time ConvergenceabstractDecentralized learning has recently been attracting increasing attention for its applications in parallel computation and privacy preservation. Many recent studies stated that the underlying network topology with a faster consensus rate (a.k.a. spectral gap) leads to a better convergence rate and accuracy for decentralized learning. However, a topology with a fast consensus rate, e.g., the exponential graph, generally has a large maximum degree, which incurs significant communication costs. Thus, seeking topologies with both a fast consensus rate and small maximum degree is important. In this study, we propose a novel topology combining both a fast consensus rate and small maximum degree called the Base-$\left(k+1\right)$ Graph. Unlike the existing topologies, the Base-$\left(k+1\right)$ Graph enables all nodes to reach the exact consensus after a finite number of iterations for any number of nodes and maximum degree $k$. Thanks to this favorable property, the Base-$\left(k+1\right)$ Graph endows Decentralized SGD (DSGD) with both a faster convergence rate and more communication efficiency than the exponential graph. We conducted experiments with various topologies, demonstrating that the Base-$\left(k+1\right)$ Graph enables various decentralized learning methods to achieve higher accuracy with better communication efficiency than the existing topologies. Our code is available at https://github.com/yukiTakezawa/BaseGraph. Yuki Takezawa, Ryoma Sato, Han Bao 0002, Kenta Niwa, Makoto Yamada |
NeurIPS | 4 |
| 2022 | Bilateral Video Magnification FilterabstractEulerian video magnification (EVM) has progressed to magnify subtle motions with a target frequency even under the presence of large motions of objects. However, existing EVM methods often fail to produce desirable results in real videos due to (1) misextracting subtle motions with a non-target frequency and (2) collapsing results when large de/acceleration motions occur (e.g., objects suddenly start, stop, or change direction). To enhance EVM performance on real videos, this paper proposes a bilateral video magnification filter (BVMF) that offers simple yet robust temporal filtering. BVMF has two kernels; (I) one kernel performs temporal bandpass filtering via a Laplacian of Gaussian whose passband peaks at the target frequency with unity gain and (II) the other kernel excludes large motions outside the magnitude of interest by Gaussian filtering on the intensity of the input signal via the Fourier shift theorem. Thus, BVMF extracts only subtle motions with the target frequency while excluding large motions outside the magnitude of interest, regardless of motion dynamics. In addition, BVMF runs the two kernels in the temporal and intensity domains simultaneously like the bilateral filter does in the spatial and intensity domains. This simplifies implementation and, as a secondary effect, keeps the memory usage low. Experiments conducted on synthetic and real videos show that BVMF outperforms state-of-the-art methods. Shoichiro Takeda, Kenta Niwa, Mariko Isogawa, Shinya Shimizu, Kazuki Okami, Yushi Aono |
CVPR | 2 |
| 2021 | Asynchronous Decentralized Optimization With Implicit Stochastic Variance ReductionabstractA novel asynchronous decentralized optimization method that follows Stochastic Variance Reduction (SVR) is proposed. Average consensus algorithms, such as Decentralized Stochastic Gradient Descent (DSGD), facilitate distributed training of machine learning models. However, the gradient will drift within the local nodes due to statistical heterogeneity of the subsets of data residing on the nodes and long communication intervals. To overcome the drift problem, (i) Gradient Tracking-SVR (GT-SVR) integrates SVR into DSGD and (ii) Edge-Consensus Learning (ECL) solves a model constrained minimization problem using a primal-dual formalism. In this paper, we reformulate the update procedure of ECL such that it implicitly includes the gradient modification of SVR by optimally selecting a constraint-strength control parameter. Through convergence analysis and experiments, we confirmed that the proposed ECL with Implicit SVR (ECL-ISVR) is stable and approximately reaches the reference performance obtained with computation on a single-node using full data set. Kenta Niwa, Guoqiang Zhang 0003, W. Bastiaan Kleijn, Noboru Harada, Hiroshi Sawada, Akinori Fujino |
ICML | 1 |
| 2021 | Ambisonic Signal Processing DNNs Guaranteeing Rotation, Scale and Time Translation EquivarianceabstractWe propose a novel framework to design Ambisonic signal processing deep neural networks (DNNs) that guarantee physical symmetries. In general, spatial acoustic signal processing DNNs for, e.g., sound event detection, ought to perform with the equivalent accuracy regardless of the directions of arrival of sound sources. This property is well known as rotation symmetry in natural science. However, in most conventional multichannel signal processing DNNs, rotation symmetry has not been explicitly incorporated into the model structure, and pseudo rotation symmetry has been acquired by training models with a large amount of signal datasets arriving from various directions. Therefore, the conventional methods will not perform sufficiently when the training dataset is relatively small scale or statistically biased, e.g., the distribution of the arriving directions of the sound events is inhomogeneous. Furthermore, in order to efficiently handle acoustic signals in DNNs, it is necessary to consider several additional symmetries, such as amplitude scaling and time translation of the signals. In this paper, we integratedly formulate these symmetry assumptions, which are called equivariance, in the form of constraints for our targeted DNN design. We propose a new DNN design method called Clebsch-Gordan Nets with Scale and Time translation Symmetry (CGNets-STS), which guarantees to simultaneously satisfy three types of equivariance (3D rotation, amplitude scaling, and time translation). As an instance of this method, we design a DNN model for sound event localization and detection tasks from Ambisonic signals. Experimental results show that this model is highly robust against spatial rotations for input data. Ryotaro Sato, Kenta Niwa, Kazunori Kobayashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Projected Weight Regularization to Improve Neural Network GeneralizationabstractGeneralization of a deep neural network (DNN) is one major concern when employing the deep learning approach for solving practical problems. In this paper we propose a new technique, named projected weight regularization (PWR), to improve the generalization capacity of a DNN model. Consider a weight matrix W from a particular neural layer in the model. Our objective is to make the eigenvalues of the matrix product WWThave comparable or roughly the same magnitudes while allowing the DNN model to fit the training data sufficiently accurate. Intuitively speaking, by doing so, it would prevent the W matrix from matching the training data too well. Specifically, at each iteration, we first project the W matrix to a number of vectors along randomly generated directions. After that, we build an objective function of the projected vectors to regularize their behaviours towards comparable eigenvalue magnitudes of WWT. Experimental results on training VGG16 for CIFAR10 show that PWR combined with centered weight normalization (CWN) yields promising validation performance compared to orthonormal regularisation combined with CWN. Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2020 | Edge-consensus Learning: Deep Learning on P2P Networks with Nonhomogeneous DataabstractAn effective Deep Neural Network (DNN) optimization algorithm that can use decentralized data sets over a peer-to-peer (P2P) network is proposed. In applications such as medical data analysis, the aggregation of data in one location may not be possible due to privacy issues. Hence, we formulate an algorithm to reach a global DNN model that does not require transmission of data among nodes. An existing solution for this issue is gossip stochastic gradient descend (SGD), which updates by averaging node models over a P2P network. However, in practical situations where the data are statistically heterogeneous across the nodes and/or where communication is asynchronous, gossip SGD often gets trapped in local minimum since the model gradients are noticeably different. To overcome this issue, we solve a linearly constrained DNN cost minimization problem, which results in variable update rules that restrict differences among all node models. Our approach can be based on the Primal-Dual Method of Multipliers (PDMM) or the Alternating Direction Method of Multiplier (ADMM), but the cost function is linearized to be suitable for deep learning. It facilitates asynchronous communication. The results of our numerical experiments using CIFAR-10 indicate that the proposed algorithms converge to a global recognition model even though statistically heterogeneous data sets are placed on the nodes. Kenta Niwa, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
KDD | 1 |
| 2020 | Microphone Array Wiener Post Filtering Using Monotone Operator SplittingabstractFor array-based acoustic source enhancement, variants of multi-channel Wiener filters are commonly used. The approach includes a Wiener post-filter that requires the simultaneous estimation of the power spectral density (PSD) of the target source and of noise sources for each time-frame. Conventional methods generally do not exploit prior knowledge, such as sparsity of the source, in solving this simultaneous estimation problem. We show that, for common scenarios, the simultaneous PSD estimation with consideration of prior knowledge can be formulated as a convex optimization problem with linear constraints. We use monotone operator splitting (MOS) to solve the constrained optimization problem. Our experiments confirm that the proposed method improves the accuracy of the noise PSD estimation, and that the resulting enhanced target signal is of higher quality. Kenta Niwa, Hironobu Chiba, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Fast Edge-consensus Computing Based on Bregman Monotone Operator SplittingabstractEdge-consensus computing is a framework to optimize a global cost function when distributed nodes observe distinct data sets. The distributed primal-dual method of multipliers (PDMM) and distributed alternating direction method of multipliers (ADMM) find network-global optima for edge-consensus algorithms by exchanging variables rather than data sets among the nodes. Since the distributed PDMM follows traditional Peaceman-Rachford splitting, it has a faster convergence rate than the distributed ADMM. To further speed up the convergence rate, we propose a new edge-consensus computing algorithm based on Bregman Peaceman-Rachford splitting. In traditional Peaceman-Rachford splitting, the variable update is defined based on a Euclidean metric and the convergence rate and is a form of first-order gradient descent. By generalizing the metric to a Bregman divergence and designing the divergence adaptively, our fast edge-consensus computing algorithm corresponds to the Newton or an accelerated gradient descent method. The results of our experiments confirm that the proposed algorithm can significantly improve the convergence rate of edge-consensus computing over state-of-the-art algorithms. Kenta Niwa, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2019 | Non-negative Matrix Factorization Using Bregman Monotone Operator SplittingabstractA non-negative matrix factorization (NMF) algorithm based on Bregman monotone operator splitting (B-MOS) is proposed. Several commonly used NMF algorithms, such as the multiplicative update method, are often used in source separation for speech and image signals. To improve the convergence rate in the tail, applying the alternating direction method of multipliers (ADMM) is reported to be effective. However, a fixed step-size parameter has to be carefully chosen for fast and stable convergence. Our main idea to overcome this issue is to adaptively modify the variable space metric so that it matches the cost convexity. Besides this, selecting an appropriate MOS (e.g., Peaceman-Rachford splitting) instead of the Douglas-Rachford splitting used in the ADMM may effectively improve the convergence rate further. To realize these ideas w.r.t. adaptive metric modification and appropriate operator splitting selection, we apply B-MOS to the NMF problem and obtain a new NMF solver in this paper. Results of numerical experiments demonstrate that the proposed NMF solver with B-MOS improved the convergence rate in the tail. Kenta Niwa, Noboru Harada |
ICASSP | 1 |
| 2019 | Function Designable Beamformer Based on Probabilistic Assumptions on Filter and Its Auxiliary VariablesabstractWe propose a novel beamformer design method that exploits probabilistic assumptions on auxiliary variables derived from filters and observed signals. Many conventional beamformer design methods can be understood in the context of optimization problems for some probabilistic cost functions. However, the class of cost functions used with these methods is quite limited to reflect multiple pieces of information and our demands on the filter, such as the sparsity assumption of the source signals and the low-latency constraint of the filter. We propose a method to design cost functions that incorporate multiple probabilistic assumptions. The assumptions are expressed as the sum of many convex terms, and every term has different auxiliary variables that are linearly constrained. Such cost functions can be optimized by iteratively optimizing with regard to each term alternately. This method enables us to more arbitrarily tune the beamformer. We conducted numerical simulations showing that our method effectively improves the performance from multiple perspectives. Ryotaro Sato, Kenta Niwa, Noboru Harada |
ICASSP | 2 |
| 2019 | Improving speech intelligibility using microphones on behind the ear hearing aidsabstractHearing aids have a great potential to facilitate better speech communication not only for hearing impaired users but also for people with normal hearing since the device will allow users to render sound with better speech intelligibility using signal processing techniques. This paper studies how a sound source separation technique would enable improving the intelligibility of a targeted speech when the technique is applied to the behind-the-ear hearing aids which has multiple microphones on each device. The sound source separation technique utilised in this study is based on the beamforming with post-filter framework and separates sound arriving from different direction. Experimental results using microphones attached to a head and torso simulator suggest the speech intelligibility can be improved by emphasising the target speech while suppressing sound from other angles, Yusuke Hioka, Kei Kobayashi, Kenta Niwa |
MMSP | 3 |
| 2018 | DNN-Based Source Enhancement to Increase Objective Sound Quality Assessment ScoreabstractWe propose a training method for deep neural network (DNN) based source enhancement to increase objective sound quality assessment (OSQA) scores such as the perceptual evaluation of speech quality. In many conventional studies, DNNs have been used as a mapping function to estimate time-frequency masks and trained to minimize an analytically tractable objective function such as the mean squared error (MSE). Since OSQA scores have been used widely for sound-quality evaluation, constructing DNNs to increase OSQA scores would be better than using the minimum MSE to create high-quality output signals. However, since most OSQA scores are not analytically tractable, i.e., they are black boxes, the gradient of the objective function cannot be calculated by simply applying backpropagation. To calculate the gradient of the OSQA-based objective function, we formulated a DNN optimization scheme on the basis of black-box optimization, which is used for training a computer that plays a game. For a black-box-optimization scheme, we adopt the policy gradient method for calculating the gradient on the basis of a sampling algorithm. To simulate output signals using the sampling algorithm, DNNs are used to estimate the probability density function of the output signals that maximize OSQA scores. The OSQA scores are calculated from the simulated output signals, and the DNNs are trained to increase the probability of generating the simulated output signals that achieve high OSQA scores. Through several experiments, we found that OSQA scores significantly increased by applying the proposed method, even though the MSE was not minimized. Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Youichi Haneda |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Efficient Audio Rendering Using Angular Region-Wise Source Enhancement for 360° VideoabstractIn virtual reality, 360° video services provided through head-mounted displays or smartphones are widely available. Among these, some state-of-the-art devices are able to render varying auditory location of an object perceived by the user when the visual location of the object in the video moves along with the change of the user's looking direction. Nevertheless, an acoustic immersion technology that generates binaural sound to maintain a good match between the auditory and visual localization of an object in 360° video has not been studied sufficiently. This study focuses on an approach that synthesizes semibinaural sound being composed of virtual sources located in each angular region and the representative head related transfer functions of each angular region. To minimize the calculation cost on audio rendering and to reduce latency in downloading data from servers, the number of angular regions should be reduced while maintaining a good match between the auditory and visual localization of an object. In this paper, we investigate the minimum number of angular regions at which it is possible to maintain a good match by conducting subjective tests using a 360° video viewing system composed of virtual images and sound sources. From the subjective tests, it was confirmed that the acoustic field should be divided into more than six equispaced angular regions so as to achieve natural auditory localization that matches an object's location in 360° video. Kenta Niwa, Yusuke Hioka, Hisashi Uematsu |
IEEE Trans. Multim. | 1 |
| 2017 | DNN-based source enhancement self-optimized by reinforcement learning using sound quality measurementsabstractWe investigated whether a deep neural network (DNN)-based source enhancement function can be self-optimized by reinforcement learning (RL). The use of a DNN is a powerful approach to describing the relationship between two sets of variables and can be useful for source enhancement function design. By training the DNN using a huge amount of training data, sound quality of output signals are improved. However, collecting a huge amount of training data is often difficult in practice. To use limited training data efficiently, we focus on the “self-optimization” of DNN-based source enhancement function in which RL is commonly utilized in the development of game playing computers. As a reward for RL, quantitative metrics that reflect a human's perceptual score (perceptual score), e.g., perceptual evaluation methods for audio source separation (PEASS), are utilized. To investigate whether the sound quality is improved by RL-based source enhancement, subjective tests were conducted. It was confirmed that the output sound quality of the RL-based source enhancement function improved as the number of iterations was increased and finally outperformed the conventional method. Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Youichi Haneda |
ICASSP | 2 |
| 2017 | Supervised source enhancement composed of nonnegative auto-encoders and complementarity subtractionabstractA method for constructing deep neural networks (DNNs) for accurate supervised source enhancement is proposed. Attempts were made in previous studies to estimate the power spectral densities (PSDs) of sound sources, which are used to estimate Wiener filters for source enhancement, from the output of multiple beamformings using DNNs. Although performance improved, it was not possible to guarantee accurate PSD estimation since the trained DNNs were treated as black boxes. The proposed DNN construction method uses non-negative auto-encoders and complementarity subtraction. This study also reveals that auto-encoders whose weights are non-negative correspond to non-negative matrix factorization (NMF), which decomposes source PSDs into non-negative spectral bases and their activations. It further introduces a complementarity subtraction method for estimating PSDs accurately. Through several experiments, it was confirmed that the signal-to-interference plus noise ratio improved by approximately 12 dB for datasets captured in various noisy/reverberant rooms. Kenta Niwa, Yuma Koizumi, Tomoko Kawase, Kazunori Kobayashi, Yusuke Hioka |
ICASSP | 1 |
| 2017 | Music staging AIabstractThrough smartphones, user enables to download/listen music anytime and anywhere. As a concept of a future audio player, we propose a framework of "music staging artificial intelligence (AI)". In that framework, audio object signals, e.g. vocal, guitar, bass, drums and keyboards, are assumed to be extracted from stereo music signals. To visualize music as if live performance is virtually conducted, playing motion sequence is estimated by using separated signals. After adjusting the spatial arrangement of audio objects so as to each user prefers it, audio/visual rendering is conducted. We constructed two types of demonstration systems for music staging AI. In the smartphone-based implementation, each user enables to change the spatial arrangement through sliderbar dragging. Since information of user preferable spatial arrangement can be sent from each smartphone to server, it would enable to predict/recommend the user preferable spatial arrangement. In another implementation, head mount display (HMD) was utilized to dive into virtual music live performance. Each user enables to walk/teleport anywhere and audio is then changing corresponding to the user view. Kenta Niwa, Kento Ohtani, Kazuya Takeda |
ICASSP | 1 |
| 2017 | Software defined media: Virtualization of audio-visual servicesabstractInternet-native audio-visual services are witnessing rapid development. Among these services, object-based audiovisual services are gaining importance. In 2014, we established the Software Defined Media (sDM) consortium to target new research areas and markets involving object-based digital media and Internet-by-design audio-visual environments. In this paper, we introduce the SDM architecture that virtualizes networked audio-visual services along with the development of smart buildings and smart cities using Internet of Things (IoT) devices and smart building facilities. Moreover, we design the SDM architecture as a layered architecture to promote the development of innovative applications on the basis of rapid advancements in software-defined networking (SDN). Then, we implement a prototype system based on the architecture, present the system at an exhibition, and provide it as an SDM API to application developers at hackathons. Various types of applications are developed using the API at these events. An evaluation of SDM API access shows that the prototype SDM platform effectively provides 3D audio reproducibility and interactiveness for SDM applications. Manabu Tsukada, Keiko Ogawa, Masahiro Ikeda, Takuro Sone, Kenta Niwa, Shoichiro Saito, Takashi Kasuya, Hideki Sunahara, Hiroshi Esaki |
ICC | 5 |
| 2017 | Informative Acoustic Feature Selection to Maximize Mutual Information for Collecting Target SourcesabstractAn informative acoustic-feature-selection method for collecting target sources in noisy environments is proposed. Wiener filtering is a powerful framework for sound-source enhancement. For Wiener-filter estimation, statistical-mapping functions, such as deep neural network based or Gaussian mixture model based mappings, have been used. In this framework, it is essential to find informative acoustic features that provide effective cues for Wiener-filter estimation. In this study, we measured the informativeness of acoustic features using mutual information between acoustic features and supervised Wiener-filter parameters, e.g., prior signal-to-noise ratios, and developed a method for automatically selecting informative acoustic features from a large number of feature candidates. To automatically select optimum features, we derived a differentiable objective function in proportion to mutual information based on the kernel method. Since the higher order correlations between acoustic features and Wiener-filter parameters are calculated using the kernel method, the statistical dependence of these variables is accurately calculated; thus, only meaningful acoustic features are selected. Through several experiments conducted on a mock sports field, we confirmed that the signal-to-distortion ratio score improved when various types of target sources were surrounded by loud cheering noise. Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Hitoshi Ohmuro |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Estimating direct-to-reverberant ratio mapped from power spectral density using deep neural networkabstractA new attempt for estimating the direct-to-reverberant ratio (DRR) by mapping the power spectral density (PSD) of the direct sound and reverberation using the deep neural network is reported. The method finds the correct DRR from the PSD estimated with an algorithm using a microphone array. The experimental results using a recording of a reverberant speech signal, which included various environmental noise, reveal that the proposed method is effective in improving the accuracy of DRR estimation and robust against various noise. Yusuke Hioka, Kenta Niwa |
ICASSP | 2 |
| 2016 | Real-time integration of statistical model-based speech enhancement with unsupervised noise PSD estimation using microphone arrayabstractWe propose a technique of multi-channel speech enhancement based on integration of beamforming and statistical model-based speech enhancement to clearly extract the target speech, even in very noisy environments. Conventional microphone array-based techniques estimate speech and noise power spectral densities (PSDs) from the spatial cues of the sound sources; however, their estimation errors dramatically increase when there are many noise sources. We integrated clean speech models trained in advance and the noise PSDs estimated in beamspace to compose observation models and designed a precise Wiener filter. Experiments under adverse noise conditions showed that the proposed technique significantly improved the signal-to-noise ratios (SNRs) compared with the conventional microphone array processing technique. Tomoko Kawase, Kenta Niwa, Masakiyo Fujimoto, Noriyoshi Kamado, Kazunori Kobayashi, Shoko Araki, Tomohiro Nakatani |
ICASSP | 2 |
| 2016 | Integrated approach of feature extraction and sound source enhancement based on maximization of mutual informationabstractWe investigated informative acoustic feature extraction based on dimension reduction for collecting target sources on a noisy sports field. Although a Wiener filter is often used for sound source enhancement, it is difficult to accurately design the Wiener filter by simply using spatial cues because the noise on a sports field (e.g., cheering from spectators) arrives from the same direction as that of the targeted source. A statistical approach is used to estimate the Wiener filter by using pre-trained acoustic feature models. However, an informative acoustic feature, which provides a powerful clue for clear extraction of the target source, is unknown. For this study, we developed a method for optimizing a projection matrix for dimension reduction by maximizing the mutual information between acoustic features and the Wiener filter. Through experiments using two-directional microphones on a mock sports field, we confirmed that the proposed method outperformed previous methods in terms of both the noise reduction and quality of the recovered sound sources. Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, Hitoshi Ohmuro |
ICASSP | 2 |
| 2016 | Pinpoint extraction of distant sound source based on DNN mapping from multiple beamforming outputs to prior SNRabstractWe propose a method for estimating the prior signal-to-noise ratio (SNR), which is used for calculating the Wiener filter for distant sound source extraction, from output signals of beamforming using statistical mapping based on the deep neural network (DNN). Since informative features to estimate the prior SNR are included in multiple beamforming outputs, the SNR can be accurately estimated by this mapping using the DNN. The proposed method was applied to a large microphone array, the design of which was optimized to form effective directivity patterns to extract distant sound sources. Experimental results proved that the target source was clearly extracted with the proposed method. Kenta Niwa, Yuma Koizumi, Tomoko Kawase, Kazunori Kobayashi, Yusuke Hioka |
ICASSP | 1 |
| 2016 | Binaural sound generation corresponding to omnidirectional video view using angular region-wise source enhancementabstractWeb applications for watching omnidirectional video through head-mounted displays (HMDs) or smartphones have been widely distributed. The goal of this study was to generate binaural sounds corresponding to the user viewpoint. Assuming that a microphone array is used for sound recording, the enhanced signal for each angular region can be extracted. By convolving head-related transfer functions (HRTFs) and enhanced signals and re-synthesizing them, binaural sounds corresponding to the user viewpoint can be virtually generated. In this paper, we propose a method for achieving angular region-wise source enhancement by generating a multichannel Wiener filter based on the power spectral density (PSD)-estimation-in-beamspace method. To measure user localization when watching omnidirectional video through an HMD, we used a system that enables the generation of binaural sounds corresponding to the user viewpoint in real time. Through subjective tests, we confirmed that sound localization corresponding to the user viewpoint can be obtained when applying about a 40-degree angular region-wise source enhancement. Kenta Niwa, Yuma Koizumi, Kazunori Kobayashi, Hisashi Uematsu |
ICASSP | 1 |
| 2016 | Optimal Microphone Array Observation for Clear Recording of Distant Sound SourcesabstractWe propose the principle for deriving an optimum design for a microphone array that uses mutual information to segregate distant sound sources. Many conventional studies on array signal processing have focused on methods for estimating sound sources from array observations. To record distant sound sources clearly, designing an optimum array structure to segregate a target from other noise is also necessary. In this study, we reveal that the optimum array observation was achieved by receiving signals that are physically decorrelated between microphones, which homogenizes the eigenvalues of the spatial correlation matrix. We theoretically explain this underlying principle using mutual information between sound sources and microphone observations whose relation to the existing minimum mean square error criterion for source separation is also discussed. The implementation of such a microphone array is possible by placing microphones in front of parabolic reflectors since the phase/amplitude around the focal point of the reflectors drastically varies with small perturbation of the microphone position. Crosscorrelation between observed signals can be reduced by optimally placing microphones. An array structure based on our proposed principle was tested by implementing minimum variance distortion-less response beamforming and postfiltering in the observations of a prototype microphone array. We experimentally confirmed that 1) the eigenvalues of the spatial correlation matrix were asymptotically homogenized and 2) the target source could be extracted clearly even when the sound sources were positioned 16.5 m from the array. Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Microphone array for increasing mutual information between sound sources and observation signalsabstractWe investigated the basic principle of how spatial signals should be captured with a microphone array to estimate each source signal and its practical implementation. Most conventional studies on array signal processing have been focused on the design of beamforming and Wiener filters. To achieve further effective noise reduction, designing an optimum array structure to segregate a target from other noises is necessary. We found the optimum structure of the spatial correlation matrix to estimate each source signal. This is achieved by receiving signals whose eigenvalues of the spatial correlation matrix are homogenized. To homogenize the eigenvalues of the spatial correlation matrix while maintaining a short impulse response length, we propose an array structure composed of parabolic reflectors and 96 microphones. Through experiments using the proposed array structure, we confirmed that the eigenvalues of the spatial correlation matrix was asymptotically homogenized and that sharp directivity could be formed. Kenta Niwa, Tatsuya Kako, Kazunori Kobayashi |
ICASSP | 1 |
| 2015 | Dive into Remote Events: Omnidirectional Video Streaming with Acoustic ImmersionabstractWe propose a system that can provide the physical presence of remote events through a head mount display (HMD) and a headphone. It can stream omnidirectional video within a limited network bandwidth at a high bitrate without sending regions that users are not viewing. It can also reproduce binaural sounds by convoluting head related transfer functions and angular region-wise separated signals. Technical demos of the system using an Oculus Rift HMD with a headphone will be performed to enable users to experience the visual and acoustic immersion it provides. Daisuke Ochi, Kenta Niwa, Akio Kameda, Yutaka Kunita, Akira Kojima |
ACM Multimedia | 2 |
| 2014 | Close-talking spherical microphone array using sound pressure interpolation based on spherical harmonic expansionabstractWe propose a novel close-talking spherical microphone array that uses the residual signal between the observed sound pressure and the interpolated sound pressure at the center of the spherical array. The interpolated sound is obtained from the sound pressures observed on the surface of a sphere on the basis of the spherical harmonic expansion, assuming that the sound originates from the outside of the array. If the sound source is close to the spherical array, the array cannot express the spherical wave correctly because the number of microphones is limited. As a result, the residual signal increases. This method is a modified form of the conventional method, which interpolates the sound pressure by using the Kirchhoff integral equation. In contrast with the conventional method, we interpolate the sound at the center of the sphere by using only the average value of the sound pressures on the spherical array surface. The computer simulations were conducted using a 12-element spherical microphone array with radius of 5 cm. These results showed that the performances of both methods were almost equivalent, although the proposed method used half the number of microphones as the conventional method. Youichi Haneda, Ken'ichi Furuya, Shoichi Koyama, Kenta Niwa |
ICASSP | 4 |
| 2013 | Underdetermined Sound Source Separation Using Power Spectrum Density Estimated by Combination of Directivity GainabstractA method for separating underdetermined sound sources based on a novel power spectral density (PSD) estimation is proposed. The method enables up toM(M-1)+1 sources to be separated when we use a microphone array ofMsensors and a Wiener post-filter calculated by the estimated PSDs. The PSD of a beamformer's output is modelled by a mixture of source PSDs multiplied by the beamformer's directivity gain in the particular angle where each source is located. Based on this model, the PSD of each sound source is estimated from the PSD of multiple fixed beamformers' outputs using the difference in the combination of directivity gains. Simulation results proved that the proposed method effectively separated up toM(M-1)+1 sound sources if the fixed beamformers were appropriately selected. Experiments were also conducted in a reverberant chamber to ensure the proposed method was also effective in practical use. Yusuke Hioka, Ken'ichi Furuya, Kazunori Kobayashi, Kenta Niwa, Youichi Haneda |
IEEE Trans. Speech Audio Process. | 4 |
| 2013 | Diffused Sensing for Sharp Directive BeamformingabstractWe generalized our previously proposed diffused sensing for a microphone array design to achieve sharp directive beamforming to enable various filter design methods to be applied. In the conventional microphone array, various filter design methods have been studied to narrow the directivity beam width. However, it is difficult to minimize the power of interference sources in the beamforming output (output interference power) over a broad frequency range since the cross-correlation between transfer functions from sound sources to microphones increases in some frequencies. With the diffused sensing, the cross-correlation is minimized by physically varying the transfer functions. We investigated how a microphone array should be designed in order to minimize the cross-correlation between transfer functions and found that placing the array in a diffuse acoustic field produces optimum results. Because the transfer functions are known a priori, this finding makes it possible to narrow the directivity beam width over a broad frequency range. This observation can be practically achieved by placing microphones inside a reflective enclosure, part of which is open to let sound waves enter. We conducted experiments using 24 microphones and confirmed that the output interference power was reduced over a broad frequency range and the beam width was narrowed by using the diffused sensing. Kenta Niwa, Yusuke Hioka, Ken'ichi Furuya, Youichi Haneda |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Estimating sound source depth using a small-size arrayabstractA method for estimating the sound source depth, i.e., the distance between a source and receiver, using a small-size array is proposed. The proposed method uses the spatial distribution pattern of quasi-independent signal components obtained by the frequency-domain independent component analysis (FDICA) as the cue for depth estimation. The quasi-independent components are calculated by applying FDICA to array signals with very high redundancy, for example, 60 microphone signals for a pair of sources; therefore, signal components associated with reflection signals are obtained even though they are correlated with the direct signal. Experimental evaluation using a small-size microphone array with a large number of elements confirms that the average (RMS) estimation error of the proposed method is 0.33 m, which is sufficiently accurate for our applications. Satoshi Esaki, Kenta Niwa, Takanori Nishino, Kazuya Takeda |
ICASSP | 2 |
| 2012 | Telescopic microphone array using reflector for segregating target source from noises in same directionabstractA spatial sensitivity control method for segregating the sound sources in the same direction by using an acoustic reflector is proposed. Our goal is to clearly pick up the target source at an arbitrary position using a microphone array. Though many methods have been studied for spatial sensitivity control, it is difficult to robustly suppress the power of noise sources in the same direction of the target source in a room. To overcome this problem, we attach a reflector to a microphone array to capture the reflected sounds whose characteristics vary depending on the distance from the array to the source. Assuming that the acoustical properties of the reflector are known e.g., measuring the transfer functions, those reflected sounds can be used as effective clues for segregating the sound sources in the same direction. With the proposed method, a filter for minimizing the output noise power is derived by taking into consideration of the acoustic properties of the reflector. Experiments were conducted in an actual room by using 96 microphones and a large reflector. We confirmed that the spatial sensitivities for segregating the target source at an arbitrary position from noise sources can be achieved by using the proposed method. Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda |
ICASSP | 1 |
| 2012 | Diffused sensing for sharp directivity microphone arrayabstractWe propose a method for achieving sharp directivity by sensing signals in a diffuse acoustic field. Directivity control based on a beamforming method has been studied to make it possible to extract the waveform and location of an identified target source even if there are many noise sources. Sharp directivity can be achieved by minimizing the output noise power of a beamforming filter. However, it is difficult to minimize the output noise power over a broad frequency ranges. Our approach for minimizing the output noise power is to control the spatial properties of the transfer functions and the spatial correlation matrix, by using a reflector that surrounds a microphone array. We investigated the relationships between the output noise power and the structure of the spatial correlation matrix and found that it was possible to minimize the output noise power by sensing diffuse acoustic signals and by designing filters taking the diffuseness of the acoustic field into consideration. In experiments, we observed diffusely reflected signals by placing a truncated-octahedral reflector near a spherical microphone array. We designed filters by using measured transfer functions and confirmed that the proposed method was effective for reducing the output noise power and forming a sharp directivity beamforming filter. Kenta Niwa, Sumitaka Sakauchi, Ken'ichi Furuya, Manabu Okamoto, Youichi Haneda |
ICASSP | 1 |
| 2011 | Estimating Direct-to-Reverberant Energy Ratio Using D/R Spatial Correlation Matrix ModelabstractWe present a method for estimating the direct-to-reverberant energy ratio (DRR) that uses a direct and reverberant sound spatial correlation matrix model (Hereafter referred to as the spatial correlation model). This model expresses the spatial correlation matrix of an array input signal as two spatial correlation matrices, one for direct sound and one for reverberation. The direct sound propagates from the direction of the sound source but the reverberation arrives from every direction uniformly. The DRR is calculated from the power spectra of the direct sound and reverberation that are estimated from the spatial correlation matrix of the measured signal using the spatial correlation model. The results of experiment and simulation confirm that the proposed method gives mostly correct DRR estimates unless the sound source is far from the microphone array, in which circumstance the direct sound picked up by the microphone array is very small. The method was also evaluated using various scales in simulated and actual acoustical environments, and its limitations revealed. We estimated the sound source distance using a small microphone array, which is an example of application of the proposed DRR estimation method. Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | Estimating direct-to-reverberant energy ratio based on spatial correlation model segregating direct sound and reverberationabstractA new approach for estimating the direct-to-reverberant energy ratio (DRR) using a microphone array is proposed. The method is based on a model of a spatial correlation matrix that segregates direct sound and reverberation. It estimates DRR from the power spectra of both components, which are derived from the correlation matrix of the observed signal. In experiments performed in simulated and actual reverberant environments, the proposed method mostly succeeded in estimating DRR accurately. We also present speech enhancement using binary masking as an example of an application of the estimated DRR. By utilization of the DRR as a factor to discriminate the distances of speakers, separation of speech signals whose sources were located in the same direction but at different distances was achieved. Yusuke Hioka, Kenta Niwa, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda |
ICASSP | 2 |
| 2010 | Estimation of sound source orientation using eigenspace of spatial correlation matrixabstractWe propose a method for estimating the sound source orientation by using the reflection sounds. The sound source orientation is important spatial information for promoting communication using teleconference systems. We assume that the observed signals captured using several microphones in a reverberant room are used for estimating the sound source orientation. Since the power of each reflection sound depends on the sound source orientation, the transfer functions between a sound source and multiple microphones are varied corresponding to the sound source orientation. We found that the eigenspace of spatial correlation constructed from the observed signals has a characteristic shape corresponding to the sound source orientation. We also proposed an efficient method for estimating the sound source orientation by matching the eigenspace of observed signals with pre-learned eigenspace models for every sound source orientation. In numerical experiments, we obtained about 80% accuracy. We confirmed the effectiveness of the proposed method. Kenta Niwa, Yusuke Hioka, Sumitaka Sakauchi, Ken'ichi Furuya, Youichi Haneda |
ICASSP | 1 |
| 2008 | Encoding large array signals into a 3D sound field representation for selective listening point audio based on blind source separationabstractABSTRACT A sound field reproduction method which uses blind source separation and head-related transfer function is proposed. In the proposed system, multichannel acoustic signals captured at the distant microphones are encoded to a set of location/signal pairs of virtual sound sources based on frequency-domain ICA. After estimating the locations and the signals of the virtual sources, by convolving the controlled acoustic transfer functions with each signal, the spatial sound at the selected point is constructed. In the evaluation, the sound field made by 6 sound sources is captured using 48 distant microphones and is encoded into set of virtual sound sources. Subjective evaluation shows that there is no significant difference between natural and reconstructed sound when more than 6 virtual sources are used. Therefore the effectiveness of the encoding algorithm as well as the virtual source representation is confirmed. Kenta Niwa, Takanori Nishino, Kazuya Takeda |
ICASSP | 1 |
| 2008 | 3DAV integrated system featuring arbitrary listening-point and viewpoint generationabstractIn this paper, we propose two novel methods for arbitrary listening-point generation for 3D audio-video (3DAV) integration in a large-scale multipoint cameras and microphones system with abilities to process, and display information of any recorded 3D scene in realtime. With this system, users are able to control their own viewpoint/listening-point position, freely. Arbitrary listening-point can be generated by either (i) ray-space representation of sound wave field (i.e. source sound independent) for multi frequency layers, or (ii) acoustic transfer function estimation (i.e. source sound dependent) and blind separation of sources of sounds. Arbitrary viewpoint generation is based on ray-space method, which is enhanced by using multipass dynamic programming for geometry compensation. Integration is done by either (i) ray-space representation of sound wave and image together, or (ii) integrating each camera video signal and acoustic transfer function of the same location as integrated 3DAV data. The prototype system of integrated audio-visual viewer achieves both good image and sound qualities with 15 frames/second. Mehrdad Panahpour Tehrani, Kenta Niwa, Norishige Fukushima, Yasushi Hirano, Toshiaki Fujii, Masayuki Tanimoto, Kazuya Takeda, Kenji Mase, Akio Ishikawa, Shigeyuki Sakazawa, Atsushi Koike |
MMSP | 2 |