Yasuhiro Oikawa

dblp:116/7393 · DBLP profile ↗
← Back
39ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-9078-1714ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Extension of Sound Field Image Denoising to High-Frequency Sound Fields by Considering Wavenumber Spectral Loss
abstract
Recent development in optical technology enables the capture of sound fields as image sequences. However, these images contain heavy noises derived from measurement devices. Previously, a deep-learning-based denoising method, called deep sound-field denoiser (DSFD), has been proposed to remove noises from sound-field images obtained by optical sound measurement. Although DSFD can remove noises superior to conventional filtering methods, its performance decreases where the frequencies of sound are high. In this paper, we propose a method to improve performance even in high-frequency sound by introducing the wavenumber spectrum obtained by a two-dimensional Fourier transform of the sound-field images. We performed denoising experiments for both simulated and measured data. The results showed that the proposed method was superior to DSFD.
Kazuma Kishida, Risako Tanigawa, Kenji Ishikawa, Yasuhiro Oikawa
ICIP4
2025 SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
abstract
Development of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acoustooptic sensing enables understanding of the interaction between sound and objects such as reflection and diffraction. Moreover, it is expected to be used an advanced measurement technology for sonars in self-driving vehicles and assistive robots. However, the low sound-pressure sensitivity of the acousto-optic sensing results in high intensity of noise on images. Therefore, denoising is an essential task to visualize and analyze the sound fields. In addition to denoising, segmentation of sound and object silhouette is also required to analyze interactions between them. In this paper, we propose sound-field-images-with-object-silhouette denoising and segmentation (SoundSil-DS) that jointly perform denoising and segmentation for sound fields and object silhouettes on a visualized image. We developed a new model based on the current state-of-the-art denoising network. We also created a dataset to train and evaluate the proposed method through acoustic simulation. The proposed method was evaluated using both simulated and measured data. We confirmed that our method can applied to experimentally measured data. These results suggest that the proposed method may improve the post-processing for sound fields, such as physical model-based three-dimensional reconstruction since it can remove unwanted noise and separate sound fields and other object silhouettes. Our code is available at https://github.com/httcslab/soundsil-ds.
Risako Tanigawa, Kenji Ishikawa, Noboru Harada, Yasuhiro Oikawa
WACV4
2024 PHAIN: Audio Inpainting via Phase-Aware Optimization With Instantaneous Frequency
abstract
Audio inpainting restores locally corrupted parts of digital audio signals. Sparsity-based methods achieve this by promoting sparsity in the time-frequency (T-F) domain, assuming short-time audio segments consist of a few sinusoids. However, such sparsity promotion reduces the magnitudes of the resulting waveforms; moreover, it often ignores the temporal connections of sinusoidal components. To address these problems, we propose a novel phase-aware audio inpainting method. Our method minimizes the time variations of a particular T-F representation calculated using the time derivative of the phase. This promotes sinusoidal components that coherently fit in the corrupted parts without directly suppressing the magnitudes. Both objective and subjective experiments confirmed the superiority of the proposed method compared with state-of-the-art methods.
Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Improving Phase-Vocoder-Based Time Stretching by Time-Directional Spectrogram Squeezing
abstract
Time stretching of music signals has a crucial problem, i.e., smearing of percussive sounds. Some time stretching algorithms have addressed this problem by detecting percussive components and manipulating them differently from the other components. However, conventional methods cause artifacts. In this paper, to prevent percussion smearing, we propose a preprocessing for time stretching. The proposed algorithm aims to preserve time scale of percussive components while stretching the rest of components in the ordinary way. To do so, time-frequency bins dominated by percussive components are squeezed in time direction so as to preserve the shape of spectrogram of percussive components. Our experiment showed that our method could improve sound quality for long stretching.
Natsuki Akaishi, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2023 UPGLADE: Unplugged Plug-and-Play Audio Declipper Based on Consensus Equilibrium of DNN and Sparse Optimization
abstract
In this paper, we propose a novel audio declipping method that fuses sparse-optimization-based and deep neural network (DNN)– based methods. The two methods have contrasting characteristics, depending on clipping level. Sparse-optimization-based audio de-clipping can preserve reliable samples, being suitable for precise restoration of small clipping. Besides, DNN-based methods are potent for recovering large clipping thanks to their data-driven approaches. Therefore, if these two methods are properly combined, audio declipping effective for a wide range of clipping levels can be realized. In the proposed method, we use a framework called consensus equilibrium to fuse the above two methods. Our experiments confirmed that the proposed method was superior to both conventional sparse-optimization-based and DNN-based methods.
Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2023 Online Phase Reconstruction via DNN-Based Phase Differences Estimation
abstract
This paper presents a two-stage online phase reconstruction framework using causal deep neural networks (DNNs). Phase reconstruction is a task of recovering phase of the short-time Fourier transform (STFT) coefficients only from the corresponding magnitude. However, phase is sensitive to waveform shifts and not easy to estimate from the magnitude even with a DNN. To overcome this problem, we propose to use DNNs for estimating differences of phase between adjacent time-frequency bins. We show that convolutional neural networks are suitable for phase difference estimation, according to the theoretical relation between partial derivatives of STFT phase and magnitude. The estimated phase differences are used for reconstructing phase by solving a weighted least squares problem in a frame-by-frame manner. In contrast to existing DNN-based phase reconstruction methods, the proposed framework is causal and does not require any iterative procedure. The experiments showed that the proposed method outperforms existing online methods and a DNN-based method for phase reconstruction.
Yoshiki Masuyama, Kohei Yatabe, Kento Nagatomo, Yasuhiro Oikawa
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Harmonic and Percussive Sound Separation Based on Mixed Partial Derivative of Phase Spectrogram
abstract
Harmonic and percussive sound separation (HPSS) is a widely applied pre-processing tool that extracts distinct (harmonic and percussive) components of a signal. In the previous methods, HPSS has been performed based on the structural properties of magnitude (or power) spectrograms. However, such approach does not take advantage of phase that contains useful information of the waveform. In this paper, we propose a novel HPSS method named MipDroP that relies only on phase and does not use information of magnitude spectrograms. The proposed MipDroP algorithm effectively examines phase through its mixed partial derivative and constructs a pair of masks for the separation. Our experiments showed that MipDroP can extract percussive components better than the other methods.
Natsuki Akaishi, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2022 Acoustic Application of Phase Reconstruction Algorithms in Optics
abstract
Phase reconstruction from amplitude spectrograms has attracted attention in recent acoustics because of its potential applications in speech synthesis and enhancement. The most well-known algorithm in acoustics is based on alternating projection and called Griffin– Lim algorithm (GLA). At the same time, GLA is known as the Gerchberg–Saxton algorithm in optics, and a lot of its variants have been proposed independently of those in acoustics. In this paper, we propose to apply phase reconstruction algorithms developed in the optics community to acoustic applications and evaluate them using acoustical metrics. Specifically, we propose to apply the averaged alternating reflections (AAR), relaxed AAR (RAAR), and hybrid input-output (HIO) algorithms to acoustic signals. Our experimental results suggested that RAAR has enough potential for acoustic applications because it clearly outperformed GLA.
Tomoki Kobayashi, Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa
ICASSP4
2022 Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around Head
abstract
Sound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although ambisonic microphones are popular in the literature of SELD, they might limits the range of applications due to their predetermined geometry. Some applications (including those for pedestrians that perform SELD while walking) require a wearable microphone array whose geometry can be designed to suit the task. In this paper, for development of such a wearable SELD, we propose a dataset named Wearable SELD dataset. It consists of data recorded by 24 microphones placed on a head and torso simulators (HATS) with some accessories mimicking wearable devices (glasses, earphones, and headphones). We also provide experimental results of SELD using the proposed dataset and SELDNet to investigate the effect of microphone configuration.
Kento Nagatomo, Masahiro Yasuda, Kohei Yatabe, Shoichiro Saito, Yasuhiro Oikawa
ICASSP5
2022 APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse Optimization
abstract
In this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been developed upon constrained optimization, they are adjustable and well-studied in theory. However, they always uniformly promote sparsity and ignore the individual properties of a signal. Deep neural network (DNN)– based methods can learn the properties of target signals and use them for audio declipping. Still, they cannot perform well if the training data have mismatches and/or constraints in the time domain are not imposed. In the proposed method, we use a DNN in an optimization algorithm. It is inspired by an idea called plug-and-play (PnP) and enables us to promote sparsity based on the learned information of data, considering constraints in the time domain. Our experiments confirmed that the proposed method is stable and robust to mismatches between training and test data.
Tomoro Tanaka, Kohei Yatabe, Masahiro Yasuda, Yasuhiro Oikawa
ICASSP4
2021 Sparse Time-Frequency Representation Via Atomic Norm Minimization
abstract
Nonstationary signals are commonly analyzed and processed in the time-frequency (T-F) domain that is obtained by the discrete Gabor transform (DGT). The T-F representation obtained by DGT is spread due to windowing, which may degrade the performance of T-F domain analysis and processing. To obtain a well-localized T-F representation, sparsity-aware methods using ℓ1-norm have been studied. However, they need to discretize a continuous parameter onto a grid, which causes a model mismatch. In this paper, we propose a method of estimating a sparse T-F representation using atomic norm. The atomic norm enables sparse optimization without discretization of continuous parameters. Numerical experiments show that the T-F representation obtained by the proposed method is sparser than the conventional methods.
Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2020 Maximally Energy-Concentrated Differential Window for Phase-Aware Signal Processing Using Instantaneous Frequency
abstract
The short-time Fourier transform (STFT) is widely employed in non-stationary signal analysis, whose property depends on window functions. Instantaneous frequency in STFT, the time-derivative of phase, is recently applied to many applications including spectrogram reassignment. The computation of instantaneous frequency requires STFT with the window and STFT with the (time-)differential window, i.e., the computation of instantaneous frequency depends on both the window function and its time derivative. To obtain the instantaneous frequency accurately, the sidelobe of frequency response of differential window should be reduced because the side-lobe causes mixing of multiple components. In this paper, we propose window functions suitable for computing the instantaneous frequency which are designed based on minimizing the sidelobe energy of the frequency response of the differential window.
Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2020 Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks
abstract
Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase reconstruction methods. However, the training of a DNN for phase reconstruction is not an easy task because phase is sensitive to the shift of a waveform. To overcome this problem, we propose a DNN-based two-stage phase reconstruction method. In the proposed method, DNNs estimate phase derivatives instead of phase itself, which allows us to avoid the sensitivity problem. Then, phase is recursively estimated based on the estimated derivatives, which is named recurrent phase unwrapping (RPU). The experimental results confirm that the proposed method outperformed the direct phase estimation by a DNN.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP4
2020 Real-Time Speech Enhancement Using Equilibriated RNN
abstract
We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capability of effectively modelling time-sequential data like speech. In particular, the long short-term memory (LSTM) is often used to alleviate the vanishing/exploding gradient problem which makes the training of an RNN difficult. However, the number of parameters of LSTM is increased as the price of mitigating the difficulty of training, which requires more computational resources. For real-time speech enhancement, it is preferable to use a smaller network without losing the performance. In this paper, we propose to use the equilibriated recurrent neural network (ERNN) for avoiding the vanishing/exploding gradient problem without increasing the number of parameters. The proposed structure is causal, which requires only the information from the past, in order to apply it in real-time. Compared to the uni- and bi-directional LSTM networks, the proposed method achieved the similar performance with much fewer parameters.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP4
2020 Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement
abstract
We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-time Fourier transform (STFT), and estimates a T-F mask using DNN. On the other hand, some methods have considered end-to-end networks which directly estimate the enhanced signals without T-F transform. While end-to-end methods have shown promising results, they are black boxes and hard to understand. Therefore, some end-to-end methods used a DNN to learn the linear T-F transform which is much easier to understand. However, the learned transform may not have a property important for ordinary signal processing. In this paper, as the important property of the T-F transform, perfect reconstruction is considered. An invertible nonlinear T-F transform is constructed by DNNs and learned from data so that the obtained transform is perfectly reconstructing filterbank.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP4
2020 Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
abstract
Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living environments, labeling all sounding objects is impossible in practice. This calls for self-supervised learning which does not require manual labeling. Most of conventional self-supervised learning uses monaural audio signals and images and cannot distinguish sound source objects having similar appearances due to poor spatial information in audio signals. To solve this problem, this paper presents a self-supervised training method using 360° images and multichannel audio signals. By incorporating with the spatial information in multichannel audio signals, our method trains deep neural networks (DNNs) to distinguish multiple sound source objects. Our system for localizing sound source objects in the image is composed of audio and visual DNNs. The visual DNN is trained to localize sound source candidates within an input image. The audio DNN verifies whether each candidate actually produces sound or not. These DNNs are jointly trained in a self-supervised manner based on a probabilistic spatial audio model. Experimental results with simulated data showed that the DNNs trained by our method localized multiple speakers. We also demonstrate that the visual DNN detected objects including talking visitors and specific exhibits from real data recorded in a science museum.
Yoshiki Masuyama, Yoshiaki Bando, Kohei Yatabe, Yoko Sasaki, Masaki Onishi, Yasuhiro Oikawa
IROS6
2020 Joint Amplitude and Phase Refinement for Monaural Source Separation
abstract
Monaural source separation is often conducted by manipulating the amplitude spectrogram of a mixture (e.g., via time-frequency masking and spectral subtraction). The obtained amplitudes are converted back to the time domain by using the phase of the mixture or by applying phase reconstruction. Although phase reconstruction performs well for the true amplitudes, its performance is degraded when the amplitudes contain error. To deal with this problem, we propose an optimization-based method to refine both amplitudes and phases based on the given amplitudes. It aims to find time-domain signals whose amplitude spectrograms are close to the given ones in terms of the generalized alpha-beta divergences. To solve the optimization problem, the alternating direction method of multipliers (ADMM) is utilized. We confirmed the effectiveness of the proposed method through speech-nonspeech separation in various conditions.
Yoshiki Masuyama, Kohei Yatabe, Kento Nagatomo, Yasuhiro Oikawa
IEEE Signal Process. Lett.4
2019 Deep Griffin-Lim Iteration
abstract
This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude spectrogram, the corresponding phase is required. One of the popular phase reconstruction methods is the Griffin-Lim algorithm (GLA), which is based on the redundancy of the short-time Fourier transform. However, GLA often involves many iterations and produces low-quality signals owing to the lack of prior knowledge of the target signal. In order to address these issues, in this study, we propose an architecture which stacks a sub-block including two GLA-inspired fixed layers and a DNN. The number of stacked sub-blocks is adjustable, and we can trade the performance and computational load based on requirements of applications. The effectiveness of the proposed method is investigated by reconstructing phases from amplitude spectrograms of speeches.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP4
2019 Low-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio Processing
abstract
Low-rankness of amplitude spectrograms has been effectively utilized in audio signal processing methods including non-negative matrix factorization. However, such methods have a fundamental limitation owing to their amplitude-only treatment where the phase of the observed signal is utilized for resynthesizing the estimated signal. In order to address this limitation, we directly treat a complex-valued spectrogram and show a complex-valued spectrogram of a sum of sinusoids can be approximately low-rank by modifying its phase. For evaluating the applicability of the proposed low-rank representation, we further propose a convex prior emphasizing harmonic signals, and it is applied to audio denoising.
Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2019 Phase-aware Harmonic/percussive Source Separation via Convex Optimization
abstract
Decomposition of an audio mixture into harmonic and percussive components, namely harmonic/percussive source separation (HPSS), is a useful pre-processing tool for many audio applications. Popular approaches to HPSS exploit the distinctive source-specific structures of power spectrograms. However, such approaches consider only power spectrograms, and the phase remains intact for resynthesizing the separated signals. In this paper, we propose a phase-aware HPSS method based on the structure of the phase of harmonic components. It is formulated as a convex optimization problem in the time domain, which enables the simultaneous treatment of both amplitude and phase. The numerical experiment validates the effectiveness of the proposed method.
Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2019 Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement
abstract
We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when a simple cost function as mean-squared error (MSE) is utilized comparing to some advanced cost such as objective sound quality assessments. However, such a simple cost function inherits strong assumptions on the statistics of the target and/or noise which is often not satisfied, and the mismatch of assumption results in degraded performance. In this paper, we propose to design the frequency scale of PRFB from training data so that the assumption on MSE is satisfied. For designing the frequency scale, the warped filterbank frame (WFBF) is considered as PRFB. The frequency characteristic of learned WFBF was in between STFT and the wavelet transform, and its effectiveness was confirmed by comparison with a standard STFT-based DNN whose input feature is compressed into the mel scale.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP4
2019 Guided-spatio-temporal Filtering for Extracting Sound from Optically Measured Images Containing Occluding Objects
abstract
Recent development of optical interferometry enables us to measure sound without placing any device inside the sound field. In particular, parallel phase-shifting interferometry (PPSI) has realized advanced measurement of refractive index of air. Its novel application investigated very recently is simultaneous visualization of flow and sound, which had been difficult until PPSI enabled high-speed and accurate measurement several years ago. However, for understanding aerodynamic sound, separation of air flow and sound is necessary since they are mixed up in the observed video. In this paper, guided-spatio-temporal filtering is proposed to separate sound from the optically measured images. Guided filtering is combined with a physical-model-based spatio-temporal filterbank for extracting sound-related information without the undesired effect caused by the image boundary or occluding objects. Such image boundary and occluding objects are typical difficulty arose in signal processing of an optically measured sound filed.
Risako Tanigawa, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2019 Griffin-Lim Like Phase Recovery via Alternating Direction Method of Multipliers
abstract
Recovering a signal from its amplitude spectrogram, or phase recovery, exhibits many applications in acoustic signal processing. When only an amplitude spectrogram is available and no explicit information is given for the phases, the Griffin-Lim algorithm (GLA) is one of the most utilized methods for phase recovery. However, GLA often requires many iterations and results in low perceptual quality in some cases. In this letter, we propose two novel algorithms based on GLA and the alternating direction method of multipliers (ADMM) for better recovery with fewer iteration. Some interpretation of the existing methods and their relation to the proposed method are also provided. Evaluations are performed with both objective measure and subjective test.
Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa
IEEE Signal Process. Lett.3
2018 Audio Hotspot Attack: An Attack on Voice Assistance Systems Using Directional Sound Beams
abstract
We propose a novel attack named "Audio Hotspot Attack'', which performs an inaudible malicious voice command attack, targeting voice assistance systems, e.g., smart speakers or in-car navigation systems. This attack leverages directional sound beams generated from parametric loudspeakers, which emit AM-modulated ultrasounds that will be self-demodulated in the air. It can succeed in the attack on a long distance (2--4 meters in a small room and 10+ meter in a long hallway). To evaluate the feasibility of the attack, we performed extensive in-lab experiments and a user study involving 20 participants. The results demonstrate that the attack is feasible in a real-world setting.
Ryo Iijima, Shota Minami, Yunao Zhou, Tatsuya Takehisa, Takeshi Takahashi 0001, Yasuhiro Oikawa, Tatsuya Mori 0003
CCS6
2018 Parametric Approximation of Piano Sound Based on Kautz Model with Sparse Linear Prediction
abstract
The piano is one of the most popular and attractive musical instruments that leads to a lot of research on it. To synthesize the piano sound in a computer, many modeling methods have been proposed from full physical models to approximated models. The focus of this paper is on the latter, approximating piano sound by an IIR filter. For stably estimating parameters, the Kautz model is chosen as the filter structure. Then, the selection of poles and excitation signal rises as the questions which are typical to the Kautz model that must be solved. In this paper, sparsity based construction of the Kautz model is proposed for approximating piano sound.
Kenji Kobayashi, Daiki Takeuchi, Mio Iwamoto, Kohei Yatabe, Yasuhiro Oikawa
ICASSP5
2018 Envelope Estimation by Tangentially Constrained Spline
abstract
Estimating envelope of a signal has various applications including empirical mode decomposition (EMD) in which the cubic C2-spline based envelope estimation is generally used. While such functional approach can easily control smoothness of an estimated envelope, the so-called undershoot problem often occurs that violates the basic requirement of envelope. In this paper, a tangentially constrained spline with tangential points optimization is proposed for avoiding the undershoot problem while maintaining smoothness. It is defined as a quartic C2-spline function constrained with first derivatives at tangential points that effectively avoids undershoot. The tangential points optimization method is proposed in combination with this spline to attain optimal smoothness of the estimated envelope.
Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2018 Modal Decomposition of Musical Instrument Sound Via Alternating Direction Method of Multipliers
abstract
For a musical instrument sound containing partials, or modes, the behavior of modes around the attack time is particularly important. However, accurately decomposing it around the attack time is not an easy task, especially when the onset is sharp. This is because spectra of the modes are peaky while the sharp onsets need a broad one. In this paper, an optimization-based method of modal decomposition is proposed to achieve accurate decomposition around the attack time. The proposed method is formulated as a constrained optimization problem to enforce the perfect reconstruction property which is important for accurate decomposition. For optimization, the alternating direction method of multipliers (ADMM) is utilized, where the update of variables is calculated in closed form. The proposed method realizes accurate modal decomposition in the simulation and real piano sounds.
Yoshiki Masuyama, Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP4
2018 Individual Difference of Ultrasonic Transducers for Parametric Array Loudspeaker
abstract
A parametric array loudspeaker (PAL) consists of a lot of ultrasonic transducers in most cases and is driven by an ultrasonic which is modulated by audible sound. Because each ultrasonic transducer has each difference resonant frequency, there is the individual difference in ultrasonic transducers of a PAL in a manufacturing process. In this paper, two PALs are made of each set of transducers with large and small variance of resonant frequencies. Quality factor of PAL with the large variance of resonant frequencies is smaller than that of PAL with small variance, and the demodulated audible sound pressure level (SPL) is large and almost flat to 3 kHz in PAL with the large variance of resonant frequencies.
Shota Minami, Jun Kuroda, Yasuhiro Oikawa
ICASSP3
2018 Realizing Directional Sound Source in FDTD Method by Estimating Initial Value
abstract
Wave-based acoustic simulation methods are studied actively for predicting acoustical phenomena. Finite-difference time-domain (FDTD) method is one of the most popular methods owing to its straightforwardness of calculating an impulse response. In an FDTD simulation, an omnidirectional sound source is usually adopted, which is not realistic because the real sound sources often have specific directivities. However, there is very little research on imposing a directional sound source into FDTD methods. In this paper, a method of realizing a directional sound source in FDTD methods is proposed. It is formulated as an estimation problem of the initial value so that the estimated result corresponds to the desired directivity. The effectiveness of the proposed method is illustrated through some numerical experiments.
Daiki Takeuchi, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2018 Phase Corrected Total Variation for Audio Signals
abstract
In optimization-based signal processing, the so-called prior term models the desired signal, and therefore its design is the key factor to achieve a good performance. For audio signals, the time-directional total variation applied to a spectrogram in combination with phase correction has been proposed recently to model sinusoidal components of the signal. Although it is a promising prior, its applicability might be restricted to some extent because of the mismatch of the assumption to the signal. In this paper, based upon the previously proposed one, an improved prior for audio signals named instantaneous phase corrected total variation (iPCTV) is proposed. It can handle wider range of audio signals owing to the instantaneous phase correction term calculated from the observed signal.
Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2017 Infinite-dimensional SVD for analyzing microphone array
abstract
Nowadays, various types of microphone array are used in many applications. However, it is not easy to compare arrays of different types because each array has been treated by a specific theory depending on the type of an array. Although several criteria have been proposed for microphone arrays for evaluating and/or designing an array, most of them are application-oriented criteria and the best configuration for some criterion may not be a better one in the other criterion. Therefore, an analysis and comparing method for microphone arrays which does not depend on an array configuration and application are necessary. In this paper, infinite-dimensional SVD is proposed for analyzing and comparing properties of arrays. The singular values and functions obtained by proposed method show sampling property of an array and can be unified criterion.
Yuji Koyano, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2017 Coherence-adjusted monopole dictionary and convex clustering for 3D localization of mixed near-field and far-field sources
abstract
In this paper, 3D sound source localization method for simultaneously estimating both direction-of-arrival (DOA) and distance from the microphone array is proposed. For estimating distance, the off-grid problem must be overcome because the range of distance to be considered is quite broad and even not bounded. The proposed method estimates positions based on an extension of the convex clustering method combined with sparse coefficients estimation. A method for constructing a suitable monopole dictionary based on coherence is also proposed so that the convex clustering based method appropriately estimate distance of sound sources. Numerical experiments of distance estimation and 3D localization show possibility of the proposed method.
Tomoya Tachikawa, Kohei Yatabe, Yasuhiro Oikawa
ICASSP3
2016 Physical-model based efficient data representation for many-channel microphone array
abstract
Recent development of microphone arrays which consist of more than several tens or hundreds microphones enables acquisition of rich spatial information of sound. Although such information possibly improve performance of any array signal processing technique, the amount of data will increase as the number of microphones increases; for instance, a 1024 ch MEMS microphone array, as in Fig. 1, generates data more than 10 GB per minute. In this paper, a method constructing an orthogonal basis for efficient representation of sound data obtained by the microphone array is proposed. The proposed method can obtain a basis for arrays with any configuration including rectangle, spherical, and random microphone array. It can also be utilized for designing a microphone array because it offers a quantitative measure for comparing several array configurations.
Yuji Koyano, Kohei Yatabe, Yusuke Ikeda, Yasuhiro Oikawa
ICASSP4
2015 Visualization of sound field by means of Schlieren method with spatio-temporal filtering
abstract
Visualization of sound field using Schlieren technique provides many advantages. It enables us to investigate the change of the sound field in real-time from every point of the observing region. However, since the density gradient of air caused by the disturbance of acoustic field is very small, it is difficult to observe the audible sound field from the raw Schlieren video. In this paper, to enhance visibility of the audible sound fields from the Schlieren videos, we propose to use spatio-temporal filters for extracting sound information and for noise removal. We have utilized different filtering techniques such as the FIR bandpass filter, the Gaussian filter, the Wiener filter and the 3D Gabor filter, to do this. The results indicate that the data observed after using these signal processing methods are clearer than the raw Schlieren videos.
Nachanant Chitanont, Keita Yaginuma, Kohei Yatabe, Yasuhiro Oikawa
ICASSP4
2015 Optically visualized sound field reconstruction based on sparse selection of point sound sources
abstract
Visualization is an effective way to understand the behavior of a sound field. There are several methods for such observation including optical measurement technique which enables a non-destructive acoustical observation by detecting density variation of the medium. For audible sound propagating through the air, however, smallness of the variation requires high sensitivity of the measuring system that causes problematic noise contamination. In this paper, a method for reconstructing two-dimensional audible sound fields from noisy optical observation is proposed.
Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2014 PDE-based interpolation method for optically visualized sound field
abstract
An effective way to understand the behavior of a sound field is to visualize it. An optical measurement method is a suitable option for this as it enables contactless non-destructive measurement. After measuring a sound field, interpolation of the data is necessary for a smooth visualization. However, conventional interpolation methods cannot provide a physically meaningful result especially when the condition of the measurement causes moiré effect. In this paper, a special interpolation method for an optically visualized sound field based on the Kirchhoff-Helmholtz integral equation is proposed.
Kohei Yatabe, Yasuhiro Oikawa
ICASSP2
2012 Extraction of sound field information from flowing dust captured with high-speed camera
abstract
In this paper, we propose a measuring method of the sound field from high-speed movie of dust. The movements of dust in the sound field are affected by the sound vibration. We observe the dust using high-speed cameras. The movie is recorded by one high-speed camera in order to get the information of 2-D sound field and two high-speed cameras for 3-D. The influence of air current is reduced from the movement of dust so that the sound field information is extracted. The experimental results indicate that this method is effective to observe the sound field especially composed of low frequency components.
Mariko Akutsu, Yasuhiro Oikawa
ICASSP2
2012 Development of a Broadcast Sound Receiver for Elderly Persons
Tomoyasu Komori, Atsushi Imai, Nobumasa Seiyama, Reiko Takou, Tohru Takagi, Yasuhiro Oikawa
ICCHP (1)6
2005 Sound field measurements based on reconstruction from laser projections
abstract
In this paper, we describe some new sound field measurement methods by using a laser Doppler vibrometer (LDV). By irradiating the reflection wall with a laser, we can observe the light velocity change that is caused by the refractive index change from the change in air density. It means that it is possible to observe the change of the sound pressure. We measured a sound field projection on a 2D plane using a scanning laser Doppler vibrometer (SVM) which can visualize a sound field. And we made a 3D sound field reconstruction from some 2D laser projections based on computed tomography (CT) techniques. We made the reconstructed image for the sound field near the loudspeaker or in the room.
Yasuhiro Oikawa, Makoto Goto, Yusuke Ikeda, Toshikazu Takizawa, Yoshio Yamasaki
ICASSP (4)1