EDBT 2026 Demo / reviewers in the wild / expert
Yasuhiro Oikawa
dblp:116/7393
· DBLP profile ↗
39ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-9078-1714ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Extension of Sound Field Image Denoising to High-Frequency Sound Fields by Considering Wavenumber Spectral LossabstractRecent development in optical technology enables the capture of sound fields as image sequences. However, these images contain heavy noises derived from measurement devices. Previously, a deep-learning-based denoising method, called deep sound-field denoiser (DSFD), has been proposed to remove noises from sound-field images obtained by optical sound measurement. Although DSFD can remove noises superior to conventional filtering methods, its performance decreases where the frequencies of sound are high. In this paper, we propose a method to improve performance even in high-frequency sound by introducing the wavenumber spectrum obtained by a two-dimensional Fourier transform of the sound-field images. We performed denoising experiments for both simulated and measured data. The results showed that the proposed method was superior to DSFD. Kazuma Kishida, Risako Tanigawa, Kenji Ishikawa, Yasuhiro Oikawa |
ICIP | 4 |
| 2025 | SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with SilhouettesabstractDevelopment of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acoustooptic sensing enables understanding of the interaction between sound and objects such as reflection and diffraction. Moreover, it is expected to be used an advanced measurement technology for sonars in self-driving vehicles and assistive robots. However, the low sound-pressure sensitivity of the acousto-optic sensing results in high intensity of noise on images. Therefore, denoising is an essential task to visualize and analyze the sound fields. In addition to denoising, segmentation of sound and object silhouette is also required to analyze interactions between them. In this paper, we propose sound-field-images-with-object-silhouette denoising and segmentation (SoundSil-DS) that jointly perform denoising and segmentation for sound fields and object silhouettes on a visualized image. We developed a new model based on the current state-of-the-art denoising network. We also created a dataset to train and evaluate the proposed method through acoustic simulation. The proposed method was evaluated using both simulated and measured data. We confirmed that our method can applied to experimentally measured data. These results suggest that the proposed method may improve the post-processing for sound fields, such as physical model-based three-dimensional reconstruction since it can remove unwanted noise and separate sound fields and other object silhouettes. Our code is available at https://github.com/httcslab/soundsil-ds. Risako Tanigawa, Kenji Ishikawa, Noboru Harada, Yasuhiro Oikawa |
WACV | 4 |
| 2024 | PHAIN: Audio Inpainting via Phase-Aware Optimization With Instantaneous FrequencyabstractAudio inpainting restores locally corrupted parts of digital audio signals. Sparsity-based methods achieve this by promoting sparsity in the time-frequency (T-F) domain, assuming short-time audio segments consist of a few sinusoids. However, such sparsity promotion reduces the magnitudes of the resulting waveforms; moreover, it often ignores the temporal connections of sinusoidal components. To address these problems, we propose a novel phase-aware audio inpainting method. Our method minimizes the time variations of a particular T-F representation calculated using the time derivative of the phase. This promotes sinusoidal components that coherently fit in the corrupted parts without directly suppressing the magnitudes. Both objective and subjective experiments confirmed the superiority of the proposed method compared with state-of-the-art methods. Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Improving Phase-Vocoder-Based Time Stretching by Time-Directional Spectrogram SqueezingabstractTime stretching of music signals has a crucial problem, i.e., smearing of percussive sounds. Some time stretching algorithms have addressed this problem by detecting percussive components and manipulating them differently from the other components. However, conventional methods cause artifacts. In this paper, to prevent percussion smearing, we propose a preprocessing for time stretching. The proposed algorithm aims to preserve time scale of percussive components while stretching the rest of components in the ordinary way. To do so, time-frequency bins dominated by percussive components are squeezed in time direction so as to preserve the shape of spectrogram of percussive components. Our experiment showed that our method could improve sound quality for long stretching. Natsuki Akaishi, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2023 | UPGLADE: Unplugged Plug-and-Play Audio Declipper Based on Consensus Equilibrium of DNN and Sparse OptimizationabstractIn this paper, we propose a novel audio declipping method that fuses sparse-optimization-based and deep neural network (DNN)– based methods. The two methods have contrasting characteristics, depending on clipping level. Sparse-optimization-based audio de-clipping can preserve reliable samples, being suitable for precise restoration of small clipping. Besides, DNN-based methods are potent for recovering large clipping thanks to their data-driven approaches. Therefore, if these two methods are properly combined, audio declipping effective for a wide range of clipping levels can be realized. In the proposed method, we use a framework called consensus equilibrium to fuse the above two methods. Our experiments confirmed that the proposed method was superior to both conventional sparse-optimization-based and DNN-based methods. Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2023 | Online Phase Reconstruction via DNN-Based Phase Differences EstimationabstractThis paper presents a two-stage online phase reconstruction framework using causal deep neural networks (DNNs). Phase reconstruction is a task of recovering phase of the short-time Fourier transform (STFT) coefficients only from the corresponding magnitude. However, phase is sensitive to waveform shifts and not easy to estimate from the magnitude even with a DNN. To overcome this problem, we propose to use DNNs for estimating differences of phase between adjacent time-frequency bins. We show that convolutional neural networks are suitable for phase difference estimation, according to the theoretical relation between partial derivatives of STFT phase and magnitude. The estimated phase differences are used for reconstructing phase by solving a weighted least squares problem in a frame-by-frame manner. In contrast to existing DNN-based phase reconstruction methods, the proposed framework is causal and does not require any iterative procedure. The experiments showed that the proposed method outperforms existing online methods and a DNN-based method for phase reconstruction. Yoshiki Masuyama, Kohei Yatabe, Kento Nagatomo, Yasuhiro Oikawa |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Harmonic and Percussive Sound Separation Based on Mixed Partial Derivative of Phase SpectrogramabstractHarmonic and percussive sound separation (HPSS) is a widely applied pre-processing tool that extracts distinct (harmonic and percussive) components of a signal. In the previous methods, HPSS has been performed based on the structural properties of magnitude (or power) spectrograms. However, such approach does not take advantage of phase that contains useful information of the waveform. In this paper, we propose a novel HPSS method named MipDroP that relies only on phase and does not use information of magnitude spectrograms. The proposed MipDroP algorithm effectively examines phase through its mixed partial derivative and constructs a pair of masks for the separation. Our experiments showed that MipDroP can extract percussive components better than the other methods. Natsuki Akaishi, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2022 | Acoustic Application of Phase Reconstruction Algorithms in OpticsabstractPhase reconstruction from amplitude spectrograms has attracted attention in recent acoustics because of its potential applications in speech synthesis and enhancement. The most well-known algorithm in acoustics is based on alternating projection and called Griffin– Lim algorithm (GLA). At the same time, GLA is known as the Gerchberg–Saxton algorithm in optics, and a lot of its variants have been proposed independently of those in acoustics. In this paper, we propose to apply phase reconstruction algorithms developed in the optics community to acoustic applications and evaluate them using acoustical metrics. Specifically, we propose to apply the averaged alternating reflections (AAR), relaxed AAR (RAAR), and hybrid input-output (HIO) algorithms to acoustic signals. Our experimental results suggested that RAAR has enough potential for acoustic applications because it clearly outperformed GLA. Tomoki Kobayashi, Tomoro Tanaka, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 4 |
| 2022 | Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around HeadabstractSound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although ambisonic microphones are popular in the literature of SELD, they might limits the range of applications due to their predetermined geometry. Some applications (including those for pedestrians that perform SELD while walking) require a wearable microphone array whose geometry can be designed to suit the task. In this paper, for development of such a wearable SELD, we propose a dataset named Wearable SELD dataset. It consists of data recorded by 24 microphones placed on a head and torso simulators (HATS) with some accessories mimicking wearable devices (glasses, earphones, and headphones). We also provide experimental results of SELD using the proposed dataset and SELDNet to investigate the effect of microphone configuration. Kento Nagatomo, Masahiro Yasuda, Kohei Yatabe, Shoichiro Saito, Yasuhiro Oikawa |
ICASSP | 5 |
| 2022 | APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse OptimizationabstractIn this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been developed upon constrained optimization, they are adjustable and well-studied in theory. However, they always uniformly promote sparsity and ignore the individual properties of a signal. Deep neural network (DNN)– based methods can learn the properties of target signals and use them for audio declipping. Still, they cannot perform well if the training data have mismatches and/or constraints in the time domain are not imposed. In the proposed method, we use a DNN in an optimization algorithm. It is inspired by an idea called plug-and-play (PnP) and enables us to promote sparsity based on the learned information of data, considering constraints in the time domain. Our experiments confirmed that the proposed method is stable and robust to mismatches between training and test data. Tomoro Tanaka, Kohei Yatabe, Masahiro Yasuda, Yasuhiro Oikawa |
ICASSP | 4 |
| 2021 | Sparse Time-Frequency Representation Via Atomic Norm MinimizationabstractNonstationary signals are commonly analyzed and processed in the time-frequency (T-F) domain that is obtained by the discrete Gabor transform (DGT). The T-F representation obtained by DGT is spread due to windowing, which may degrade the performance of T-F domain analysis and processing. To obtain a well-localized T-F representation, sparsity-aware methods using ℓ1-norm have been studied. However, they need to discretize a continuous parameter onto a grid, which causes a model mismatch. In this paper, we propose a method of estimating a sparse T-F representation using atomic norm. The atomic norm enables sparse optimization without discretization of continuous parameters. Numerical experiments show that the T-F representation obtained by the proposed method is sparser than the conventional methods. Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2020 | Maximally Energy-Concentrated Differential Window for Phase-Aware Signal Processing Using Instantaneous FrequencyabstractThe short-time Fourier transform (STFT) is widely employed in non-stationary signal analysis, whose property depends on window functions. Instantaneous frequency in STFT, the time-derivative of phase, is recently applied to many applications including spectrogram reassignment. The computation of instantaneous frequency requires STFT with the window and STFT with the (time-)differential window, i.e., the computation of instantaneous frequency depends on both the window function and its time derivative. To obtain the instantaneous frequency accurately, the sidelobe of frequency response of differential window should be reduced because the side-lobe causes mixing of multiple components. In this paper, we propose window functions suitable for computing the instantaneous frequency which are designed based on minimizing the sidelobe energy of the frequency response of the differential window. Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2020 | Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural NetworksabstractPhase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase reconstruction methods. However, the training of a DNN for phase reconstruction is not an easy task because phase is sensitive to the shift of a waveform. To overcome this problem, we propose a DNN-based two-stage phase reconstruction method. In the proposed method, DNNs estimate phase derivatives instead of phase itself, which allows us to avoid the sensitivity problem. Then, phase is recursively estimated based on the estimated derivatives, which is named recurrent phase unwrapping (RPU). The experimental results confirm that the proposed method outperformed the direct phase estimation by a DNN. Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada |
ICASSP | 4 |
| 2020 | Real-Time Speech Enhancement Using Equilibriated RNNabstractWe propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capability of effectively modelling time-sequential data like speech. In particular, the long short-term memory (LSTM) is often used to alleviate the vanishing/exploding gradient problem which makes the training of an RNN difficult. However, the number of parameters of LSTM is increased as the price of mitigating the difficulty of training, which requires more computational resources. For real-time speech enhancement, it is preferable to use a smaller network without losing the performance. In this paper, we propose to use the equilibriated recurrent neural network (ERNN) for avoiding the vanishing/exploding gradient problem without increasing the number of parameters. The proposed structure is causal, which requires only the information from the past, in order to apply it in real-time. Compared to the uni- and bi-directional LSTM networks, the proposed method achieved the similar performance with much fewer parameters. Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada |
ICASSP | 4 |
| 2020 | Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech EnhancementabstractWe propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-time Fourier transform (STFT), and estimates a T-F mask using DNN. On the other hand, some methods have considered end-to-end networks which directly estimate the enhanced signals without T-F transform. While end-to-end methods have shown promising results, they are black boxes and hard to understand. Therefore, some end-to-end methods used a DNN to learn the linear T-F transform which is much easier to understand. However, the learned transform may not have a property important for ordinary signal processing. In this paper, as the important property of the T-F transform, perfect reconstruction is considered. An invertible nonlinear T-F transform is constructed by DNNs and learned from data so that the obtained transform is perfectly reconstructing filterbank. Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada |
ICASSP | 4 |
| 2020 | Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial ModelingabstractDetecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living environments, labeling all sounding objects is impossible in practice. This calls for self-supervised learning which does not require manual labeling. Most of conventional self-supervised learning uses monaural audio signals and images and cannot distinguish sound source objects having similar appearances due to poor spatial information in audio signals. To solve this problem, this paper presents a self-supervised training method using 360° images and multichannel audio signals. By incorporating with the spatial information in multichannel audio signals, our method trains deep neural networks (DNNs) to distinguish multiple sound source objects. Our system for localizing sound source objects in the image is composed of audio and visual DNNs. The visual DNN is trained to localize sound source candidates within an input image. The audio DNN verifies whether each candidate actually produces sound or not. These DNNs are jointly trained in a self-supervised manner based on a probabilistic spatial audio model. Experimental results with simulated data showed that the DNNs trained by our method localized multiple speakers. We also demonstrate that the visual DNN detected objects including talking visitors and specific exhibits from real data recorded in a science museum. Yoshiki Masuyama, Yoshiaki Bando, Kohei Yatabe, Yoko Sasaki, Masaki Onishi, Yasuhiro Oikawa |
IROS | 6 |
| 2020 | Joint Amplitude and Phase Refinement for Monaural Source SeparationabstractMonaural source separation is often conducted by manipulating the amplitude spectrogram of a mixture (e.g., via time-frequency masking and spectral subtraction). The obtained amplitudes are converted back to the time domain by using the phase of the mixture or by applying phase reconstruction. Although phase reconstruction performs well for the true amplitudes, its performance is degraded when the amplitudes contain error. To deal with this problem, we propose an optimization-based method to refine both amplitudes and phases based on the given amplitudes. It aims to find time-domain signals whose amplitude spectrograms are close to the given ones in terms of the generalized alpha-beta divergences. To solve the optimization problem, the alternating direction method of multipliers (ADMM) is utilized. We confirmed the effectiveness of the proposed method through speech-nonspeech separation in various conditions. Yoshiki Masuyama, Kohei Yatabe, Kento Nagatomo, Yasuhiro Oikawa |
IEEE Signal Process. Lett. | 4 |
| 2019 | Deep Griffin-Lim IterationabstractThis paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude spectrogram, the corresponding phase is required. One of the popular phase reconstruction methods is the Griffin-Lim algorithm (GLA), which is based on the redundancy of the short-time Fourier transform. However, GLA often involves many iterations and produces low-quality signals owing to the lack of prior knowledge of the target signal. In order to address these issues, in this study, we propose an architecture which stacks a sub-block including two GLA-inspired fixed layers and a DNN. The number of stacked sub-blocks is adjustable, and we can trade the performance and computational load based on requirements of applications. The effectiveness of the proposed method is investigated by reconstructing phases from amplitude spectrograms of speeches. Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada |
ICASSP | 4 |
| 2019 | Low-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio ProcessingabstractLow-rankness of amplitude spectrograms has been effectively utilized in audio signal processing methods including non-negative matrix factorization. However, such methods have a fundamental limitation owing to their amplitude-only treatment where the phase of the observed signal is utilized for resynthesizing the estimated signal. In order to address this limitation, we directly treat a complex-valued spectrogram and show a complex-valued spectrogram of a sum of sinusoids can be approximately low-rank by modifying its phase. For evaluating the applicability of the proposed low-rank representation, we further propose a convex prior emphasizing harmonic signals, and it is applied to audio denoising. Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2019 | Phase-aware Harmonic/percussive Source Separation via Convex OptimizationabstractDecomposition of an audio mixture into harmonic and percussive components, namely harmonic/percussive source separation (HPSS), is a useful pre-processing tool for many audio applications. Popular approaches to HPSS exploit the distinctive source-specific structures of power spectrograms. However, such approaches consider only power spectrograms, and the phase remains intact for resynthesizing the separated signals. In this paper, we propose a phase-aware HPSS method based on the structure of the phase of harmonic components. It is formulated as a convex optimization problem in the time domain, which enables the simultaneous treatment of both amplitude and phase. The numerical experiment validates the effectiveness of the proposed method. Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2019 | Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source EnhancementabstractWe propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when a simple cost function as mean-squared error (MSE) is utilized comparing to some advanced cost such as objective sound quality assessments. However, such a simple cost function inherits strong assumptions on the statistics of the target and/or noise which is often not satisfied, and the mismatch of assumption results in degraded performance. In this paper, we propose to design the frequency scale of PRFB from training data so that the assumption on MSE is satisfied. For designing the frequency scale, the warped filterbank frame (WFBF) is considered as PRFB. The frequency characteristic of learned WFBF was in between STFT and the wavelet transform, and its effectiveness was confirmed by comparison with a standard STFT-based DNN whose input feature is compressed into the mel scale. Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada |
ICASSP | 4 |
| 2019 | Guided-spatio-temporal Filtering for Extracting Sound from Optically Measured Images Containing Occluding ObjectsabstractRecent development of optical interferometry enables us to measure sound without placing any device inside the sound field. In particular, parallel phase-shifting interferometry (PPSI) has realized advanced measurement of refractive index of air. Its novel application investigated very recently is simultaneous visualization of flow and sound, which had been difficult until PPSI enabled high-speed and accurate measurement several years ago. However, for understanding aerodynamic sound, separation of air flow and sound is necessary since they are mixed up in the observed video. In this paper, guided-spatio-temporal filtering is proposed to separate sound from the optically measured images. Guided filtering is combined with a physical-model-based spatio-temporal filterbank for extracting sound-related information without the undesired effect caused by the image boundary or occluding objects. Such image boundary and occluding objects are typical difficulty arose in signal processing of an optically measured sound filed. Risako Tanigawa, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2019 | Griffin-Lim Like Phase Recovery via Alternating Direction Method of MultipliersabstractRecovering a signal from its amplitude spectrogram, or phase recovery, exhibits many applications in acoustic signal processing. When only an amplitude spectrogram is available and no explicit information is given for the phases, the Griffin-Lim algorithm (GLA) is one of the most utilized methods for phase recovery. However, GLA often requires many iterations and results in low perceptual quality in some cases. In this letter, we propose two novel algorithms based on GLA and the alternating direction method of multipliers (ADMM) for better recovery with fewer iteration. Some interpretation of the existing methods and their relation to the proposed method are also provided. Evaluations are performed with both objective measure and subjective test. Yoshiki Masuyama, Kohei Yatabe, Yasuhiro Oikawa |
IEEE Signal Process. Lett. | 3 |
| 2018 | Audio Hotspot Attack: An Attack on Voice Assistance Systems Using Directional Sound BeamsabstractWe propose a novel attack named "Audio Hotspot Attack'', which performs an inaudible malicious voice command attack, targeting voice assistance systems, e.g., smart speakers or in-car navigation systems. This attack leverages directional sound beams generated from parametric loudspeakers, which emit AM-modulated ultrasounds that will be self-demodulated in the air. It can succeed in the attack on a long distance (2--4 meters in a small room and 10+ meter in a long hallway). To evaluate the feasibility of the attack, we performed extensive in-lab experiments and a user study involving 20 participants. The results demonstrate that the attack is feasible in a real-world setting. Ryo Iijima, Shota Minami, Yunao Zhou, Tatsuya Takehisa, Takeshi Takahashi 0001, Yasuhiro Oikawa, Tatsuya Mori 0003 |
CCS | 6 |
| 2018 | Parametric Approximation of Piano Sound Based on Kautz Model with Sparse Linear PredictionabstractThe piano is one of the most popular and attractive musical instruments that leads to a lot of research on it. To synthesize the piano sound in a computer, many modeling methods have been proposed from full physical models to approximated models. The focus of this paper is on the latter, approximating piano sound by an IIR filter. For stably estimating parameters, the Kautz model is chosen as the filter structure. Then, the selection of poles and excitation signal rises as the questions which are typical to the Kautz model that must be solved. In this paper, sparsity based construction of the Kautz model is proposed for approximating piano sound. Kenji Kobayashi, Daiki Takeuchi, Mio Iwamoto, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 5 |
| 2018 | Envelope Estimation by Tangentially Constrained SplineabstractEstimating envelope of a signal has various applications including empirical mode decomposition (EMD) in which the cubic C2-spline based envelope estimation is generally used. While such functional approach can easily control smoothness of an estimated envelope, the so-called undershoot problem often occurs that violates the basic requirement of envelope. In this paper, a tangentially constrained spline with tangential points optimization is proposed for avoiding the undershoot problem while maintaining smoothness. It is defined as a quartic C2-spline function constrained with first derivatives at tangential points that effectively avoids undershoot. The tangential points optimization method is proposed in combination with this spline to attain optimal smoothness of the estimated envelope. Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2018 | Modal Decomposition of Musical Instrument Sound Via Alternating Direction Method of MultipliersabstractFor a musical instrument sound containing partials, or modes, the behavior of modes around the attack time is particularly important. However, accurately decomposing it around the attack time is not an easy task, especially when the onset is sharp. This is because spectra of the modes are peaky while the sharp onsets need a broad one. In this paper, an optimization-based method of modal decomposition is proposed to achieve accurate decomposition around the attack time. The proposed method is formulated as a constrained optimization problem to enforce the perfect reconstruction property which is important for accurate decomposition. For optimization, the alternating direction method of multipliers (ADMM) is utilized, where the update of variables is calculated in closed form. The proposed method realizes accurate modal decomposition in the simulation and real piano sounds. Yoshiki Masuyama, Tsubasa Kusano, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 4 |
| 2018 | Individual Difference of Ultrasonic Transducers for Parametric Array LoudspeakerabstractA parametric array loudspeaker (PAL) consists of a lot of ultrasonic transducers in most cases and is driven by an ultrasonic which is modulated by audible sound. Because each ultrasonic transducer has each difference resonant frequency, there is the individual difference in ultrasonic transducers of a PAL in a manufacturing process. In this paper, two PALs are made of each set of transducers with large and small variance of resonant frequencies. Quality factor of PAL with the large variance of resonant frequencies is smaller than that of PAL with small variance, and the demodulated audible sound pressure level (SPL) is large and almost flat to 3 kHz in PAL with the large variance of resonant frequencies. Shota Minami, Jun Kuroda, Yasuhiro Oikawa |
ICASSP | 3 |
| 2018 | Realizing Directional Sound Source in FDTD Method by Estimating Initial ValueabstractWave-based acoustic simulation methods are studied actively for predicting acoustical phenomena. Finite-difference time-domain (FDTD) method is one of the most popular methods owing to its straightforwardness of calculating an impulse response. In an FDTD simulation, an omnidirectional sound source is usually adopted, which is not realistic because the real sound sources often have specific directivities. However, there is very little research on imposing a directional sound source into FDTD methods. In this paper, a method of realizing a directional sound source in FDTD methods is proposed. It is formulated as an estimation problem of the initial value so that the estimated result corresponds to the desired directivity. The effectiveness of the proposed method is illustrated through some numerical experiments. Daiki Takeuchi, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2018 | Phase Corrected Total Variation for Audio SignalsabstractIn optimization-based signal processing, the so-called prior term models the desired signal, and therefore its design is the key factor to achieve a good performance. For audio signals, the time-directional total variation applied to a spectrogram in combination with phase correction has been proposed recently to model sinusoidal components of the signal. Although it is a promising prior, its applicability might be restricted to some extent because of the mismatch of the assumption to the signal. In this paper, based upon the previously proposed one, an improved prior for audio signals named instantaneous phase corrected total variation (iPCTV) is proposed. It can handle wider range of audio signals owing to the instantaneous phase correction term calculated from the observed signal. Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 2 |
| 2017 | Infinite-dimensional SVD for analyzing microphone arrayabstractNowadays, various types of microphone array are used in many applications. However, it is not easy to compare arrays of different types because each array has been treated by a specific theory depending on the type of an array. Although several criteria have been proposed for microphone arrays for evaluating and/or designing an array, most of them are application-oriented criteria and the best configuration for some criterion may not be a better one in the other criterion. Therefore, an analysis and comparing method for microphone arrays which does not depend on an array configuration and application are necessary. In this paper, infinite-dimensional SVD is proposed for analyzing and comparing properties of arrays. The singular values and functions obtained by proposed method show sampling property of an array and can be unified criterion. Yuji Koyano, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2017 | Coherence-adjusted monopole dictionary and convex clustering for 3D localization of mixed near-field and far-field sourcesabstractIn this paper, 3D sound source localization method for simultaneously estimating both direction-of-arrival (DOA) and distance from the microphone array is proposed. For estimating distance, the off-grid problem must be overcome because the range of distance to be considered is quite broad and even not bounded. The proposed method estimates positions based on an extension of the convex clustering method combined with sparse coefficients estimation. A method for constructing a suitable monopole dictionary based on coherence is also proposed so that the convex clustering based method appropriately estimate distance of sound sources. Numerical experiments of distance estimation and 3D localization show possibility of the proposed method. Tomoya Tachikawa, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 3 |
| 2016 | Physical-model based efficient data representation for many-channel microphone arrayabstractRecent development of microphone arrays which consist of more than several tens or hundreds microphones enables acquisition of rich spatial information of sound. Although such information possibly improve performance of any array signal processing technique, the amount of data will increase as the number of microphones increases; for instance, a 1024 ch MEMS microphone array, as in Fig. 1, generates data more than 10 GB per minute. In this paper, a method constructing an orthogonal basis for efficient representation of sound data obtained by the microphone array is proposed. The proposed method can obtain a basis for arrays with any configuration including rectangle, spherical, and random microphone array. It can also be utilized for designing a microphone array because it offers a quantitative measure for comparing several array configurations. Yuji Koyano, Kohei Yatabe, Yusuke Ikeda, Yasuhiro Oikawa |
ICASSP | 4 |
| 2015 | Visualization of sound field by means of Schlieren method with spatio-temporal filteringabstractVisualization of sound field using Schlieren technique provides many advantages. It enables us to investigate the change of the sound field in real-time from every point of the observing region. However, since the density gradient of air caused by the disturbance of acoustic field is very small, it is difficult to observe the audible sound field from the raw Schlieren video. In this paper, to enhance visibility of the audible sound fields from the Schlieren videos, we propose to use spatio-temporal filters for extracting sound information and for noise removal. We have utilized different filtering techniques such as the FIR bandpass filter, the Gaussian filter, the Wiener filter and the 3D Gabor filter, to do this. The results indicate that the data observed after using these signal processing methods are clearer than the raw Schlieren videos. Nachanant Chitanont, Keita Yaginuma, Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 4 |
| 2015 | Optically visualized sound field reconstruction based on sparse selection of point sound sourcesabstractVisualization is an effective way to understand the behavior of a sound field. There are several methods for such observation including optical measurement technique which enables a non-destructive acoustical observation by detecting density variation of the medium. For audible sound propagating through the air, however, smallness of the variation requires high sensitivity of the measuring system that causes problematic noise contamination. In this paper, a method for reconstructing two-dimensional audible sound fields from noisy optical observation is proposed. Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 2 |
| 2014 | PDE-based interpolation method for optically visualized sound fieldabstractAn effective way to understand the behavior of a sound field is to visualize it. An optical measurement method is a suitable option for this as it enables contactless non-destructive measurement. After measuring a sound field, interpolation of the data is necessary for a smooth visualization. However, conventional interpolation methods cannot provide a physically meaningful result especially when the condition of the measurement causes moiré effect. In this paper, a special interpolation method for an optically visualized sound field based on the Kirchhoff-Helmholtz integral equation is proposed. Kohei Yatabe, Yasuhiro Oikawa |
ICASSP | 2 |
| 2012 | Extraction of sound field information from flowing dust captured with high-speed cameraabstractIn this paper, we propose a measuring method of the sound field from high-speed movie of dust. The movements of dust in the sound field are affected by the sound vibration. We observe the dust using high-speed cameras. The movie is recorded by one high-speed camera in order to get the information of 2-D sound field and two high-speed cameras for 3-D. The influence of air current is reduced from the movement of dust so that the sound field information is extracted. The experimental results indicate that this method is effective to observe the sound field especially composed of low frequency components. Mariko Akutsu, Yasuhiro Oikawa |
ICASSP | 2 |
| 2012 | Development of a Broadcast Sound Receiver for Elderly Persons
Tomoyasu Komori, Atsushi Imai, Nobumasa Seiyama, Reiko Takou, Tohru Takagi, Yasuhiro Oikawa |
ICCHP (1) | 6 |
| 2005 | Sound field measurements based on reconstruction from laser projectionsabstractIn this paper, we describe some new sound field measurement methods by using a laser Doppler vibrometer (LDV). By irradiating the reflection wall with a laser, we can observe the light velocity change that is caused by the refractive index change from the change in air density. It means that it is possible to observe the change of the sound pressure. We measured a sound field projection on a 2D plane using a scanning laser Doppler vibrometer (SVM) which can visualize a sound field. And we made a 3D sound field reconstruction from some 2D laser projections based on computed tomography (CT) techniques. We made the reconstructed image for the sound field near the loudspeaker or in the room. Yasuhiro Oikawa, Makoto Goto, Yusuke Ikeda, Toshikazu Takizawa, Yoshio Yamasaki |
ICASSP (4) | 1 |