EDBT 2026 Demo / reviewers in the wild / expert
Masashi Unoki
dblp:73/1580
· DBLP profile ↗
71ranked-venue papers
9as first author
21since 2021 · last 2025
0000-0002-6605-2052ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 41 · 6 first-author · 15 since 2021Security and privacy · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-tuning TitaNet-Large Model for Speaker Anonymization Attacker SystemsabstractSpeaker anonymization techniques are crucial for safeguarding user privacy in voice-based applications. However, these methods are susceptible to adversarial attacks that can compromise their effectiveness. This paper proposes attacker systems that leverage the power of fine-tuned TitaNet-Large and ECAPA-TDNN models to identify the original speaker from anonymized speech generated by various anonymization methods. Both pre-trained models are renowned for their state-of-the-art ability to extract robust speaker embeddings. Finetuning these models with anonymized speech enables them to identify underlying patterns in anonymized speech. We evaluated the proposed attacker systems against multiple anonymization techniques that performed effectively in a series of voice privacy challenges. Our experimental results underscore the effectiveness of the fine-tuned TitaNet-Large model in breaking through these anonymization methods, as indicated by the reduced equal error rate (EER). This highlights the importance of robust and adaptive anonymization strategies to counter such emerging semiinformed threats. Candy Olivia Mawalim, Aulia Adila, Masashi Unoki |
ICASSP | 3 |
| 2025 | Robust Multilingual Audio Deepfake Detection Through Hybrid ModelingabstractThe increasing sophistication of AI-generated human voice poses a significant threat, demanding robust detection systems that can generalize effectively across diverse linguistic environments and synthesis techniques.In response to the SAFE Challenge, this paper introduces a novel approach to multilingual audio deepfake detection.Our primary contribution lies in the comprehensive study of deepfake detection using a multilingual speech corpus encompassing 17 languages and a broad spectrum of synthesis methods and acoustic conditions, designed to enable more realistic and challenging evaluations.To optimally utilize this diverse data, we propose a hybrid detection model that synergistically combines the strengths of end-to-end RawNet and AASIST architectures with language-agnostic representations learned from a multilingual selfsupervised learning model.Additionally, we explore the efficacy of RawBoost data augmentation in enhancing robustness against realworld noise.Our experimental evaluation demonstrates promising generalization in generated audio detection, achieving approximately 73% balanced accuracy across multilingual data and unseen synthesis algorithms. Candy Olivia Mawalim, Aulia Adila, Shogo Okada, Masashi Unoki |
IH&MMSec | 5 |
| 2024 | Are Recent Deep Learning-Based Speech Enhancement Methods Ready to Confront Real-World Noisy Environments?abstractRecent advancements in speech enhancement techniques have ignited interest in improving speech quality and intelligibility.However, the effectiveness of recently proposed methods is unclear.In this paper, a comprehensive analysis of modern deep learning-based speech enhancement approaches is presented.Through evaluations using the Deep Suppression Noise and Clarity Enhancement Challenge datasets, we assess the performances of three methods: Denoiser, DeepFilterNet3, and FullSubNet+.Our findings reveal nuanced performance differences among these methods, with varying efficacy across datasets.While objective metrics offer valuable insights, they struggle to represent complex scenarios with multiple noise sources.Leveraging ASR-based methods for these scenarios shows promise but may induce critical hallucination effects.Our study emphasizes the need for ongoing research to refine techniques for diverse real-world environments. Candy Olivia Mawalim, Shogo Okada, Masashi Unoki |
INTERSPEECH | 3 |
| 2024 | Robust voice activity detection using an auditory-inspired masked modulation encoder based convolutional attention network
Longbiao Wang, Meng Ge, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
Speech Commun. | 4 |
| 2023 | An Improved Optimal Transport Kernel Embedding Method with Gating Mechanism for Singing Voice Separation and Speaker IdentificationabstractSinging voice separation (SVS) and speaker identification (SI) are two classic problems in speech signal processing. Deep neural networks (DNNs) solve these two problems by extracting effective representations of the target signal from the input mixture. Since essential features of a signal can be well reflected on its latent geometric structure of the feature distribution, a natural way to address SVS/SI is to extract the geometry-aware and distribution-related features of the target signal. To do this, this work introduces the concept of optimal transport (OT) to SVS/SI and proposes an improved optimal transport kernel embedding (iOTKE) to extract the target-distribution-related features. The iOTKE learns an OT from the input signal to the target signal on the basis of a reference set learned from all training data. Thus it can maintain the feature diversity and preserve the latent geometric structure of the distribution for the target signal. To further improve the feature selection ability, we extend the proposed iOTKE to a gated version, i.e., gated iOTKE (G-iOTKE), by incorporating a lightweight gating mechanism. The gating mechanism controls effective information flow and enables the proposed method to select important features for a specific input signal. We evaluated the proposed G-iOTKE on SVS/SI. Experimental results showed that the proposed method provided better results than other models. Weitao Yuan, Yuren Bian, Shengbei Wang, Masashi Unoki, Wenwu Wang 0001 |
ICASSP | 4 |
| 2023 | Consonant-emphasis Method Incorporating Robust Consonant-section Detection to Improve Intelligibility of Bone-conducted speech
Yasufumi Uezu, Teruki Toya, Masashi Unoki |
INTERSPEECH | 4 |
| 2023 | Music Theory-Inspired Acoustic Representation for Speech Emotion RecognitionabstractThis research presents a music theory-inspired acoustic representation (hereafter, MTAR) to address improved speech emotion recognition. The recognition of emotion in speech and music is developed in parallel, yet a relatively limited understanding of MTAR for interpreting speech emotions is involved. In the present study, we use music theory to study representative acoustics associated with emotion in speech from vocal emotion expressions and auditory emotion perception domains. In experiments assessing the role and effectiveness of the proposed representation in classifying discrete emotion categories and predicting continuous emotion dimensions, it shows promising performance compared with extensively used features for emotion recognition based on the spectrogram, Mel-spectrogram, Mel-frequency cepstral coefficients, VGGish, and the large baseline feature sets of the INTERSPEECH challenges. This proposal opens up a novel research avenue in developing a computational acoustic representation of speech emotion via music theory. Xingfeng Li 0001, Desheng Hu, Qingchen Zhang 0001, Zhengxia Wang, Masashi Unoki, Masato Akagi |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2023 | A Discriminative Feature Representation Method Based on Cascaded Attention Network With Adversarial Strategy for Speech Emotion RecognitionabstractCurrently, speech emotion recognition models still could not show satisfactory performance due to the complexity of emotions. In most of the previous studies, there is a common problem that some of the particular emotions are severely misclassified. In this article, we propose a novel framework integrating cascaded attention network and adversarial joint loss strategy for speech emotion recognition, aiming at discriminating the confusions by emphasizing more on the emotions which are difficult to be correctly classified. First, we extract log-Mels, deltas and delta-deltas of log-Mels as 3D features to effectively reduce the interference of external factors. Next, we introduce a cascaded attention network to extract effective emotional features, where spatiotemporal attention selectively locates the targeted emotional regions from the input features. In these targeted regions, the self attention with head fusion captures the long-distance dependence of temporal features. Finally, an adversarial joint loss strategy is proposed to distinguish the emotional embeddings with high similarity by the generated hard triplets in an adversarial fashion. To evaluate our proposed method, experiments are performed with the IEMOCAP, CASIA, and EMODB corpora. The experimental results demonstrate that our proposed method significantly outperforms the state-of-the-art approaches on all datasets. Yang Liu 0262, Haoqin Sun, Wenbo Guan, Yuqi Xia, Masashi Unoki, Zhen Zhao 0006 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Unsupervised Deep Unfolded Representation Learning for Singing Voice SeparationabstractLearning effective vocal representations from a waveform mixture is a crucial but challenging task for deep neural network (DNN)-based singing voice separation (SVS). Successful representation learning (RL) depends heavily on well-designed neural architectures and effective general priors. However, DNNs for RL in SVS are mostly built on generic architectures without general priors being systematically considered. To address these issues, we introduce deep unfolding to RL and propose two RL-based models for SVS, deep unfolded representation learning (DURL) and optimal transport DURL (OT-DURL). In both models, we formulate RL as a sequence of optimization problems for signal reconstruction, where three general priors, synthesis, non-negative, and our novel analysis, are incorporated. In DURL and OT-DURL, we take different approaches in penalizing the analysis prior. DURL uses the Euclidean distance as its penalty, while OT-DURL uses a more sophisticated penalty known as the OT distance. We address the optimization problems in DURL and OT-DURL with the first-order operator splitting algorithm and unfold the obtained iterative algorithms to novel encoders, by mapping the synthesis/analysis/non-negative priors to different interpretable sublayers of the encoders. We evaluated these DURL and OT-DURL encoders in the unsupervised informed SVS and supervised Open-Unmix frameworks. Experimental results indicate that (1) the OT-DURL encoder is better than the DURL encoder and (2) both encoders can considerably improve the vocal-signal-separation performance compared with those of the baseline model. Weitao Yuan, Shengbei Wang, Masashi Unoki, Wenwu Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | An Improved Stimulus Reconstruction Method for EEG-Based Short-Time Auditory Attention Detection
Gaoyan Zhang, Masashi Unoki, Jianwu Dang 0001, Longbiao Wang |
ICONIP (5) | 4 |
| 2022 | Vector-quantized Variational Autoencoder for Phase-aware Speech EnhancementabstractSpeech-enhancement methods based on the complex ideal ratio mask (cIRM) have achieved promising results.These methods often deploy a deep neural network to jointly estimate the real and imaginary components of the cIRM defined in the complex domain.However, the unbounded property of the cIRM poses difficulties when it comes to effectively training a neural network.To alleviate this problem, this paper proposes a phase-aware speech-enhancement method through estimating the magnitude and phase of a complex adaptive Wiener filter.With this method, a noise-robust vector-quantized variational autoencoder is used for estimating the magnitude of the Wiener filter by using the Itakura-Saito divergence on the time-frequency domain, while the phase of the Wiener filter is estimated using a convolutional recurrent network using the scale-invariant signal-to-noise-ratio constraint in the time domain.The proposed method was evaluated on the open Voice Bank+DEMAND dataset to provide a direct comparison with other speech-enhancement methods and achieved a Perceptual Evaluation of Speech Quality score of 2.85 and ShortTime Objective Intelligibility score of 0.94, which is better than the stateof-art method based on cIRM estimation during the 2020 Deep Noise Challenge. Tuan Vu Ho, Masato Akagi, Masashi Unoki |
INTERSPEECH | 4 |
| 2022 | Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio DetectionabstractFake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%. Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki |
INTERSPEECH | 10 |
| 2022 | Global Signal-to-noise Ratio Estimation Based on Multi-subband Processing Using Convolutional Neural Network
Meng Ge, Longbiao Wang, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2022 | Automatic Mean Opinion Score Estimation with Temporal Modulation Features on Gammatone Filterbank for Speech Assessment
Kai Li 0018, Masashi Unoki |
INTERSPEECH | 3 |
| 2022 | Method for improving the word intelligibility of presented speech using bone-conduction headphones
Teruki Toya, Wenyu Zhu, Maori Kobayashi, Kenichi Nakamura, Masashi Unoki |
INTERSPEECH | 5 |
| 2022 | Speaker anonymization by modifying fundamental frequency and x-vector singular valueabstractSpeaker anonymization is a method of protecting voice privacy by concealing individual speaker characteristics while preserving linguistic information. The VoicePrivacy Challenge 2020 was initiated to generalize the task of speaker anonymization. In the challenge, two frameworks for speaker anonymization were introduced; in this study, we propose a method of improving the primary framework by modifying the state-of-the-art speaker individuality feature (namely, x-vector) in a neural waveform speech synthesis model. Our proposed method is constructed based on x-vector singular value modification with a clustering model. We also propose a technique of modifying the fundamental frequency and speech duration to enhance the anonymization performance. To evaluate our method, we carried out objective and subjective tests. The overall objective test results show that our proposed method improves the anonymization performance in terms of the speaker verifiability, whereas the subjective evaluation results show improvement in terms of the speaker dissimilarity. The intelligibility and naturalness of the anonymized speech with speech prosody modification were slightly reduced (less than 5% of word error rate) compared to the results obtained by the baseline system. Candy Olivia Mawalim, Kasorn Galajit, Jessada Karnjana, Shunsuke Kidani, Masashi Unoki |
Comput. Speech Lang. | 5 |
| 2021 | Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural NetworkabstractVoice activity detection (VAD) based on deep learning has achieved remarkable success. However, when the traditional features (e.g., raw waveforms and MFCCs) are directly fed to the deep neural network model, the performance decreases because of noise interference. Here, we propose a robust VAD approach using a masked auditory encoder based convolutional neural network (M-AECNN). First, we analyze the effectiveness of using auditory features as deep learning encoder. These features can roughly simulate the transmission of sound to human inner-ear hair cells; thus, they are more robust than the raw waveform and frequency domain features designed as encoders. Second, similar to the human ear’s masking effect for different speech frequencies, the proposed auditory encoder can further improve the robustness of VAD by increasing the gain for cleaner speech frequencies. Extensive experimental results demonstrate that this approach achieves about 10.5% absolute improvement in the area under the curve on the AURORA-2J dataset compared with a VAD method based on a CNN and MFCCs. Longbiao Wang, Masashi Unoki, Sheng Li 0010, Rui Wang 0102, Meng Ge, Jianwu Dang 0001 |
ICASSP | 3 |
| 2021 | Synchronous Multi-Bit Audio Watermarking Based on Phase ShiftingabstractAudio watermarking has been developed to protect the copyright of audio signals. We considered the use of the distribution of the phase spectrum and propose an effective multi-bit audio watermarking method based on phase shifting. The proposed method is implemented on the basis of a frame-wise framework. The phase spectrum of each frame is first calculated, then, in accordance with the embedded bit number of each frame, the phase bins are segmented into several subsets. An artificial phase pattern is constructed in one of these subsets to carry a specific multi-bit watermark. In the water-mark extraction process, each frame is synchronized by identifying the phase pattern, then watermarks are extracted. We investigate the performance of the proposed method, and the results indicated that the proposed method balances the trade-off of inaudibility and robustness. From comparative evaluations, the proposed method also performed better than several other audio watermarking methods. Shengbei Wang, Weitao Yuan, Masashi Unoki |
ICASSP | 5 |
| 2021 | Crossfire Conditional Generative Adversarial Networks for Singing Voice ExtractionabstractGenerative adversarial networks (GANs) and Conditional GANs (cGANs) have recently been applied for singing voice extraction (SVE), since they can accurately model the vocal distributions and effectively utilize a large amount of unlabelled datasets.However, current GANs/cGANs based SVE frameworks have no explicit mechanism to eliminate the mutual interferences between different sources.In this work, we introduce a novel 'crossfire' criterion into GANs to complement its standard adversarial training, which forms a dual-objective GANs, namely Crossfire GANs (Cr-GANs).In addition, we design a Generalized Projection Method (GPM) for cGANs based frameworks to extract more effective conditional information for SVE.Using the proposed GPM, we extend our Cr-GANs to conditional version, i.e., Crossfire Conditional GANs (Cr-cGANs).The proposed methods were evaluated on the DSD100 and CCMixter datasets.The numerical results have shown that the 'crossfire' criterion and GPM are beneficial to each other and considerably improve the separation performance of existing GANs/cGANs based SVE methods. Weitao Yuan, Shengbei Wang, Xiangrui Li, Masashi Unoki, Wenwu Wang 0001 |
Interspeech | 4 |
| 2021 | Multi-resolution modulation-filtered cochleagram feature for LSTM-based dimensional emotion recognition from speech
Zhichao Peng, Jianwu Dang 0001, Masashi Unoki, Masato Akagi |
Neural Networks | 3 |
| 2021 | Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice SeparationabstractMonaural singing voice separation (MSVS) is a challenging task and has been extensively studied. Deep neural networks (DNNs) are current state-of-the-art methods for MSVS. However, they are often designed manually, which is time-consuming and error-prone. They are also pre-defined, thus cannot adapt their structures to the training data. To address these issues, we first designed a multi-resolution convolutional neural network (CNN) for MSVS called multi-resolution pooling CNN (MRP-CNN), which uses various-sized pooling operators to extract multi-resolution features. We then introduced Neural Architecture Search (NAS) to extend the MRP-CNN to the evolving MRP-CNN (E-MRP-CNN) to automatically search for effective MRP-CNN structures using genetic algorithms optimized in terms of a single objective taking into account only separation performance and multiple objectives taking into account both separation performance and model complexity. The E-MRP-CNN using the multi-objective algorithm gives a set of Pareto-optimal solutions, each providing a trade-off between separation performance and model complexity. Evaluations on the MIR-1 K, DSD100, and MUSDB18 datasets were used to demonstrate the advantages of the E-MRP-CNN over several recent baselines. Weitao Yuan, Bofei Dong, Shengbei Wang, Masashi Unoki, Wenwu Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | X-Vector Singular Value Modification and Statistical-Based Decomposition with Ensemble Regression Modeling for Speaker Anonymization System
Candy Olivia Mawalim, Kasorn Galajit, Jessada Karnjana, Masashi Unoki |
INTERSPEECH | 4 |
| 2020 | Cortical Oscillatory Hierarchy for Natural Sentence Processing
Jianwu Dang 0001, Gaoyan Zhang, Masashi Unoki |
INTERSPEECH | 4 |
| 2020 | Multi-Subspace Echo Hiding Based on Time-Frequency Similarities of Audio SignalsabstractAudio watermarking plays an important role in copyright protection. Echo hiding, one of the most effective techniques for audio watermarking, has been studied for decades. However, the conventional echo hiding has been criticized for its weak security, as watermarks can be easily extracted by means of cepstrum analysis even without any prior knowledge. This article explores the time-frequency (T-F) characteristics of the repetition structures in an audio signal to improve the security of conventional echo hiding. In our approach, the original audio signal is first converted into a high-dimensional T-F representation, and then by clustering the time frames that have similar T-F characteristics into the same subspace, the original audio is decomposed into a union of subspaces with each corresponding to one time-domain subsignal. Paired and opposite echo kernels are applied to energy-balanced subsignals for watermark embedding, which significantly improves the security. In the watermark extraction process, the subspaces are recovered on the basis of the T-F similarities and cepstrum analysis is utilized to extract watermarks. The proposed embedding and extraction schemes thus offer a new approach for echo hiding. The results of experiments demonstrate the effectiveness of our approach with respect to inaudibility, security, and robustness. Shengbei Wang, Weitao Yuan, Masashi Unoki |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Inaudible Speech Watermarking Based on Self-compensated Echo-hiding and Sparse Subspace ClusteringabstractThe method reported here realizes an inaudible echo-hiding based speech watermarking by using sparse subspace clustering (SSC). Speech signal is first analyzed with SSC to obtain its sparse and low-rank components. Watermarks are embedded as the echoes of the sparse component for robust extraction. Self-compensated echoes consisting of two independent echo kernels are designed to have similar delay offsets but opposite amplitudes. A one-bit watermark is embedded by separately performing the echo kernels on the sparse and low-rank components. As a result, the sound distortion caused by one echo signal can be quickly compensated by the other echo signal, which enables better inaudibility. Since the embedded echoes have the same sparsity as the sparse component, watermarks can be extracted with a basic cepstrum analysis even if the echo kernels are not directly performed on the original speech. The evaluation results verify the feasibility and effectiveness of this method. Shengbei Wang, Weitao Yuan, Masashi Unoki |
ICASSP | 4 |
| 2019 | Proximal Deep Recurrent Neural Network for Monaural Singing Voice SeparationabstractThe recent deep learning methods can offer state-of-the-art performance for Monaural Singing Voice Separation (MSVS). In these deep methods, the recurrent neural network (RNN) is widely employed. This work proposes a novel type of Deep RNN (DRNN), namely Proximal DRNN (P-DRNN) for MSVS, which improves the conventional Stacked RNN (S-RNN) by introducing a novel interlayer structure. The interlayer structure is derived from an optimization problem for Monaural Source Separation (MSS). Accordingly, this enables a new hierarchical processing in the proposed P-DRNN with the explicit state transfers between different layers and the skip connections from the inputs, which are efficient for source separation. Finally, the proposed approach is evaluated on the MIR-IK dataset to verify its effectiveness. The numerical results show that the P-DRNN performs better than the conventional S-RNN and several recent MSVS methods. Weitao Yuan, Shengbei Wang, Xiangrui Li, Masashi Unoki, Wenwu Wang 0001 |
ICASSP | 4 |
| 2019 | Data Augmentation for Monaural Singing Voice Separation Based on Variational Autoencoder-Generative Adversarial NetworkabstractRandom mixing and circularly shifting for augmenting the training set are used to improve the separation effect of deep neural network (DNN)-based monaural singing voice separation (MSVS). However, these manual methods are based on unrealistic assumptions that two sources in the mixture are independent of each other, which limits the separation effect. This paper proposes a data augmentation method based on variational autoencoder (VAE) and generative adversarial network (GAN), which is called as VAE-GAN. The VAE models the observed spectra of sources (vocal and music) separately and reconstructs new spectra from the latent space. The GAN's discriminator is introduced to measure the correlation between the latent variables of the vocal and music generated by the VAE probability encoder. This adversarial mechanism in VAE's latent space could learn the synthetic likelihood and ultimately decode high quality spectra samples, which further improves the separation effect of general MSVS networks. Boxin He, Shengbei Wang, Weitao Yuan, Masashi Unoki |
ICME | 5 |
| 2019 | Detection of speech tampering using sparse representations and spectral manipulations based information hiding
Shengbei Wang, Weitao Yuan, Masashi Unoki |
Speech Commun. | 4 |
| 2019 | Enhanced feature network for monaural singing voice separation
Weitao Yuan, Boxin He, Shengbei Wang, Masashi Unoki |
Speech Commun. | 5 |
| 2019 | A Skip Attention Mechanism for Monaural Singing Voice SeparationabstractThis work proposes a simple but effective attention mechanism, namely Skip Attention (SA), for monaural singing voice separation (MSVS). First, the SA, embedded in the convolutional encoder-decoder network (CEDN), realizes an attention-driven and dependency modeling for the repetitive structures of the music source. Second, the SA, replacing the popular skip connection in the CEDN, effectively controls the flow of the low-level (vocal and musical) features to the output and improves the feature sensitivity and accuracy for MSVS. Finally, we implement the proposed SA on the Stacked Hourglass Network (SHN), namely Skip Attention SHN (SA-SHN). Quantitative and qualitative evaluation results have shown that the proposed SA-SHN achieves significant performance improvement on the MIR-1K dataset (compared to the state-of-the-art SHN) and competitive MSVS performance on the DSD100 dataset (compared to the state-of-the-art DenseNet), even without using any data augmentation methods. Weitao Yuan, Shengbei Wang, Xiangrui Li, Masashi Unoki, Wenwu Wang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Method of Estimating Direction of Arrival of Sound Source for Monaural Hearing Based on Temporal Modulation PerceptionabstractAlthough humans are capable of using monaural and modulation cues for sound localization, it is not yet clear how they can use that information to estimate the direction of arrival (DOA) of a sound source in 3D space. Our previous study revealed that the head-related modulation transfer function (HR-MTF) contains significant trends and features, which can be used for DOA estimation. This paper proposes a method of estimating the DOA in a 3D space by using the monaural modulation spectrum (MMS), based on the concept of modulation transfer function (MTF) and auditory perception of temporal modulation. We carried out over 51, 840 simulations with several signal types and multiple subjects to simultaneously estimate the azimuth and the elevation of an incoming sound source. The root mean square error (RMSE) was derived to evaluate the accuracy of monaural DOA estimates. Our results indicated that the proposed method could adequately estimate the DOA in 3D space with an overall mean RMSE of 21.9 degrees. Nguyen Khanh Bui, Daisuke Morikawa, Masashi Unoki |
ICASSP | 3 |
| 2018 | Speech Watermarking Based on Robust Principal Component Analysis and Formant ManipulationsabstractThis paper proposes a watermarking method for speech signals based on Robust Principal Component Analysis (RPCA) and formant manipulations. As the spectrogram of speech has a relatively sparse structure, the core information of speech is extracted into a sparse matrix using RPCA so that formants can be estimated with Linear Prediction (LP) more accurately even under noise/interferences, which significantly improves the robustness of proposed method. We investigate how the formants can be controlled and manipulated to make the watermarking method effective. Watermarks are embedded into speech by controlling the shape and power of formants using the stable and robust parameter, i.e., line spectral frequencies (LSFs). Evaluations regarding inaudibility and robustness are carried out and the results suggest that the proposed method can not only satisfy inaudibility but also provide good robustness against general processing and different speech codecs which is better than the other methods. Shengbei Wang, Weitao Yuan, Masashi Unoki |
ICASSP | 4 |
| 2018 | Auditory-Inspired End-to-End Speech Emotion Recognition Using 3D Convolutional Recurrent Neural Networks Based on Spectral-Temporal RepresentationabstractThe human auditory system has far superior emotion recognition abilities compared with recent speech emotion recognition systems, so research has focused on designing emotion recognition systems by mimicking the human auditory system. Psychoacoustic and physiological studies indicate that the human auditory system decomposes speech signals into acoustic and modulation frequency components, and further extracts temporal modulation cues. Speech emotional states are perceived from temporal modulation cues using the spectral and temporal receptive field of the neuron. This paper proposes an emotion recognition system in an end-to-end manner using three-dimensional convolutional recurrent neural networks (3D-CRNNs) based on temporal modulation cues. Temporal modulation cues contain four-dimensional spectral-temporal (ST) integration representations directly as the input of 3D-CRNNs. The convolutional layer is used to extract high-level multiscale ST representations, and the recurrent layer is used to extract long-term dependency for emotion recognition. The proposed method was verified on the IEMOCAP database. The results show that our proposed method can exceed the recognition accuracy compared to that of the state-of-the-art systems. Zhichao Peng, Masashi Unoki, Jianwu Dang 0001, Masato Akagi |
ICME | 3 |
| 2017 | Robust Method for Estimating F0 of Complex Tone Based on Pitch Perception of Amplitude Modulated Signal
Kenichiro Miwa, Masashi Unoki |
INTERSPEECH | 2 |
| 2016 | Investigations into vowel and consonant structures in articulatory and auditory spaces using Laplacian eigenmapsabstractMany studies have investigated the relationship between the articulatory and auditory features for isolated speech sound and vowels. For fully understanding the mechanisms of speech production and perception, it is necessary to investigate the consonants in the same way. For this reason, in this study, we investigate the manifolds of vowels and consonants out of Japanese reading speech using Laplacian eigenmaps. We constructed uniform articulatory and auditory spaces based on the vowels and consonants to investigate their manifolds. It is found that the distribution of consonants in articulatory space could be classified into labial and lingual groups which reflected their articulatory properties, while in auditory space their distribution was clustered according to voiced and unvoiced, plosive and fricative properties. In vowel-consonant acoustic space, the consonants distributed as a hoe-like shape, with voiced consonants located on the blade of the hoe and fused with vowels. We defined average correlation coefficients to measure the similarity of manifold between three speakers. The results indicated that the vowel/consonant structures had high consistency among the three speakers. Jianwu Dang 0001, Shengbei Wang, Masashi Unoki |
ICASSP | 3 |
| 2016 | Modulation Spectral Features for Predicting Vocal Emotion Recognition by Simulated Cochlear Implants
Ryota Miyauchi, Yukiko Araki, Masashi Unoki |
INTERSPEECH | 4 |
| 2016 | Speech enhancement of instantaneous amplitude and phase for applications in noisy reverberant environments
Yang Liu 0052, Naushin Nower, Shota Morita, Masashi Unoki |
Speech Commun. | 4 |
| 2015 | Robust and reliable audio watermarking based on phase codingabstractThis paper proposes a novel robust audio watermarking method based on phase coding. The quantization index modulation technique is employed for embedding watermarks into the phase spectrum of audio signals. To increase robustness of the proposed method, the region of phase spectrum that is resistant against attacks is selected for embedding. We experimentally analyzed the phase spectrum to find out which region is not distorted under attacks. On the other hand, the quantization step size is suitably selected so that the modification of phase does not cause severe distortion in sound quality. The experimental results show that the watermarks could be kept inaudible in audio signals and robust against attacks. The proposed method has the ability to embed watermarks into audio signals up to 400 bits per second with a bit error rate of less than 1%. Nhut Minh Ngo, Masashi Unoki |
ICASSP | 2 |
| 2015 | Complex tensor factorization in modulation frequency domain for single-channel speech enhancement
Shogo Masaya, Masashi Unoki |
INTERSPEECH | 2 |
| 2015 | Restoration scheme of instantaneous amplitude and phase using Kalman filter with efficient linear prediction for speech enhancement
Naushin Nower, Yang Liu 0052, Masashi Unoki |
Speech Commun. | 3 |
| 2014 | Restoration of instantaneous amplitude and phase using Kalman filter for speech enhancementabstractThis paper proposes a restoration scheme for the instantaneous amplitudes and phases in sub-bands by using a Kalman filter with linear prediction (LP). A few important studies have already proved that phase spectrum on the short term Fourier transform plays an important role in speech enhancement. Thus, the proposed scheme concentrates on restoring of both the instantaneous amplitudes and phases simultaneously. In this scheme, the Kalman filter is used for both instantaneous amplitudes and phases in the sub-band representation to remove the effect of noise. Thus, it can sufficiently reduce the noise effects in both. Simulations were carried out in various noisy environments to evaluate the effectiveness of the proposed scheme. The signal to error ratio (SER), perceptual evaluation of speech quality (PESQ), and SNR loss were used as objective measures. Results showed that the proposed scheme can effectively improve these objective measures more than conventional methods. Naushin Nower, Yang Liu 0052, Masashi Unoki |
ICASSP | 3 |
| 2014 | Formant enhancement based speech watermarking for tampering detectionabstractUnauthorized tampering in speech signals has brought serious problems when verifying the originality and integrity of speech signals. Digital watermarking can effectively check if the original signals have been tampered by embedding digital data into them. This paper proposes a tampering detection scheme for speech signals based on formant enhancement-based watermarking. Watermarks are embedded as slight enhancement of formant by symmetrically controlling a pair of linear spectral frequencies (LSFs) of corresponding formant. We evaluated the proposed scheme with objective evaluations concerning three criteria that are required for tampering detection scheme: (i) inaudibility to human auditory system, (ii) robustness against meaningful processing, and (iii) fragility against tampering. The evaluation results showed that the proposed scheme could provide satisfactory performance in all the criteria and had the ability to detect tampering in speech signals. Index Terms: tampering detection, speech watermarking, formant enhancement, inaudibility, robustness, fragility Shengbei Wang, Masashi Unoki, Nam Soo Kim |
INTERSPEECH | 2 |
| 2014 | An Audio Watermarking Scheme Based on Singular-Spectrum Analysis
Jessada Karnjana, Masashi Unoki, Pakinee Aimmanee, Chai Wutiwiwatchai |
IWDW | 2 |
| 2014 | Watermarking for Digital Audio Based on Adaptive Phase Modulation
Nhut Minh Ngo, Masashi Unoki |
IWDW | 2 |
| 2013 | Concurrent processing of voice activity detection and noise reduction using empirical mode decomposition and modulation spectrum analysis
Yasuaki Kanai, Shota Morita, Masashi Unoki |
INTERSPEECH | 3 |
| 2011 | Adaptive Regularization Framework for Robust Voice Activity Detection
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2011 | Voice Activity Detection in MTF-Based Power Envelope Restoration
Masashi Unoki, Xugang Lu, Rico Petrick, Shota Morita, Masato Akagi, Rüdiger Hoffmann |
INTERSPEECH | 1 |
| 2011 | Sub-band temporal modulation envelopes and their normalization for automatic speech recognition in reverberant environments
Xugang Lu, Masashi Unoki, Satoshi Nakamura 0001 |
Comput. Speech Lang. | 2 |
| 2011 | Temporal modulation normalization for robust speech feature extraction and recognition
Xugang Lu, Shigeki Matsuda, Masashi Unoki, Satoshi Nakamura 0001 |
Multim. Tools Appl. | 3 |
| 2010 | Voice activity detection in a reguarized reproducing kernel hilbert space
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2010 | Methods for robust speech recognition in reverberant environments: a comparisonabstractIn this article the authors continue previous studies regarding the investigation of methods that aim to improve the decreased recognition rate (RR) in reverberant environments of automatic speech recognition (ASR) systems. Previously threerobust front-end methods are tested, the harmonicity based feature analysis (HFA), the temporal power envelope feature analysis(TPEFA) and their combination (HFA+TPEFA). This paper additionally introduces two well-known methods into the comparison. These are the dereverberation method using the inverse modulation transfer function (IMTF) and the delay-and-sum beamformer (DSB). Recognition experiments are accomplished for command word recognition, the reverberant environmentsare comprehensive chosen as functions of the reverberation time T_60 and the speaker to microphone distance (SMD) as the most important parameters to describe reverberant distortions.The results of this first comparison of such methodsprove experimentally some drawn assumptions, e. g. the IMTF method achieves robustness only in the far field, the DSB improves the RR slightly but is outperformed by the HFA due to its indirectivity at low frequencies. Rico Petrick, Thomas Fehér, Masashi Unoki, Rüdiger Hoffmann |
INTERSPEECH | 3 |
| 2010 | Temporal contrast normalization and edge-preserved smoothing of temporal modulation structures of speech for robust speech recognition
Xugang Lu, Shigeki Matsuda, Masashi Unoki, Satoshi Nakamura 0001 |
Speech Commun. | 3 |
| 2009 | Temporal contrast normalization and edge-preserved smoothing on temporal modulation structure for robust speech recognitionabstractIn this paper, we propose a two-step processing algorithm which adaptively normalizes the temporal modulation of speech to extract robust speech feature for automatic speech recognition systems. The first step processing is to normalize the temporal modulation contrast (TMC) of the cepstral time series for both clean and noisy speech. The second step processing is to smooth the normalized temporal modulation structure to reduce the artifacts due to noise while preserving the speech modulation events (edges). We tested our algorithm on speech recognition experiments in additive noise condition (AURORA-2J data corpus), reverberant noise condition (convolution of clean speech utterances from AURORA-2J with a smart room impulse response), and noisy condition with both reverberant and additive noise (air conditioner noise in a smart room). For comparison, the ETSI advanced front-end (AFE) algorithm was used. Our results showed that the algorithm provided: (1) for additive noise condition, 57.26% relative word error reduction (RWER) rate for clean conditional training (59.37% for AFE), and 33.52% RWER rate for multi-conditional training (35.77% for AFE), (2) for reverberant condition, 51.28% RWER rate (10.17% for AFE) and (3) for noisy condition with both reverberant and additive noise, 71.74% RWER rate (48.86% for AFE). Xugang Lu, Shigeki Matsuda, Masashi Unoki, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2009 | Subband temporal modulation spectrum normalization for automatic speech recognition in reverberant environments
Xugang Lu, Masashi Unoki, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2008 | Comparative evaluations of robust and accurate F0 estimates in reverberant environmentsabstractThis paper reports comparative evaluations of the method we previously proposed of estimating fundamental frequency (F0) based on complex cepstrum analysis with nine typical methods over huge speech-sound datasets in both artificial and realistic reverberant environments (in room acoustics). They involve several classic algorithms (Cepstrum, AMDF, LPC, and modified autocorrelation) and a few modern algorithms (TEMPO, YIN, and PHIA). The comparative results revealed that the percentage correct rates of the estimated FOs using them were drastically reduced as the reverberation time increased while Foestimated with the proposed method was completely robust and accurate. They also demonstrated that homomorphic analysis and the concept of a source-filter model were relatively effective for estimating Fo. The results also demonstrated that it was much better than the previously reported methods in terms of robustness and providing accurate Foestimates in both artificial and realistic reverberant environments. Masashi Unoki, Toshihiro Hosorogiya, Yuichi Ishimoto |
ICASSP | 1 |
| 2008 | Robust front end processing for speech recognition in reverberant environments: utilization of speech characteristicsabstractThis paper proposes two methods for robust automatic speech recognition (ASR) in reverberant environments. Unlike other methods which mostly apply inverse filtering by blindly estimated room impulse responses to achieve dereverberation, theproposed methods are based on the utilization of the characteristics of speech. The first method - Harmonicity based Feature Analysis – takes advantage of the harmonic componentsof speech, which are assumed to be undistorted. The second method - Temporal Power Envelope Feature Analysis – utilizes the temporal modulation structure of speech, representing the phoneme level temporal events which contain most intelligibility information. Both methods increase the recognition performance remarkably in a different way. Combining both of them connects their individual advantages. In order to examine theperformance of utilizing harmonicity and modulation temporal structure for reverberant ASR, the methods are tested in clean and reverberant training. As results show, even in strong reverberantconditions both methods obtain practical applicableperformance for reverberant training. In addition, besides testing their performance in dependency on the reverberation time, their performance considering the speaker-to-microphone distanceis tested, which is another new contributions in this paper. Rico Petrick, Xugang Lu, Masashi Unoki, Masato Akagi, Rüdiger Hoffmann |
INTERSPEECH | 3 |
| 2008 | A comprehensive study on the effects of room reverberation on fundamental frequency estimation
Rico Petrick, Masashi Unoki, Anish Mittal, Carlos Segura, Rüdiger Hoffmann |
INTERSPEECH | 2 |
| 2007 | Vocal conversion from speaking voice to singing voice using STRAIGHT
Takeshi Saitou, Masataka Goto, Masashi Unoki, Masato Akagi |
INTERSPEECH | 3 |
| 2007 | Method of LP-based blind restoration for improving intelligibility of bone-conducted speech
Thang Tat Vu, Germine Seide, Masashi Unoki, Masato Akagi |
INTERSPEECH | 3 |
| 2006 | A robust feature extraction based on the MTF concept for speech recognition in reverberant environment
Xugang Lu, Masashi Unoki, Masato Akagi |
INTERSPEECH | 2 |
| 2005 | A model for selective segregation of a target instrument sound from the mixed sound of various instruments
Masashi Unoki, Masaaki Kubo, Atsushi Haniu, Masato Akagi |
INTERSPEECH | 1 |
| 2005 | Development of an F0 control model based on F0 dynamic characteristics for singing-voice synthesis
Takeshi Saitou, Masashi Unoki, Masato Akagi |
Speech Commun. | 2 |
| 2004 | Analysis of acoustic features affecting "singing-ness" and its application to singing-voice synthesis from speaking-voiceabstractTo construct a natural singing-voice synthesis system, it is important to adequately control acoustic features such as fundamental frequency (F0), spectrum shapes, and phoneme duration in the synthesis method. This paper reveals acoustic features affecting singing-voice perception by comparative analyzing singing- and speaking-voices, and then proposes a transforming method from speaking-voice into singing-voice using STRAIGHT [1]. This method is composed of an F0 control model for generating F0 contours of singing-voices, a spectral sequence control model for modifying spectral shapes in speaking-voice, and a duration control model based on rhythm. Results showed that the proposed system could synthesize a natural singing-voice, whose sound quality is almost the same as that of real one. 1. Takeshi Saitou, Naoya Tsuji, Masashi Unoki, Masato Akagi |
INTERSPEECH | 3 |
| 2003 | A method based on the MTF concept for dereverberating the power envelope from the reverberant signalabstractThis paper proposes a method for dereverberating the power envelope from the reverberant signal. This method is based on the modulation transfer function (MTF) and does not require that the impulse response of an environment be measured. It improves upon the basic model proposed by Hirobayashi et al. (1998) regarding the following problems: (i) how to precisely extract the power envelope from the observed signal; (ii) how to determine the parameters of the impulse response of the room; and (iii) a lack of consideration as to whether the MTF concept can be applied to a more realistic signal. We have shown that the proposed method can accurately dereverberate the power envelope from the reverberant signal. Masashi Unoki, Masashi Furukawa, Keigo Sakata, Masato Akagi |
ICASSP (1) | 1 |
| 2003 | A speech dereverberation method based on the MTF concept
Masashi Unoki, Keigo Sakata, Masato Akagi |
INTERSPEECH | 1 |
| 2001 | A fundamental frequency estimation method for noisy speech based on instantaneous amplitude and frequencyabstractThis paper proposes a robust and accurate F0 estimation method for noisy speech. This method uses two different principles: (1) an F0 estimation based on periodicity and harmonicity of instantaneous amplitude for a robust estimation in noisy environments, and (2) an F0 estimation based on stability of instantaneous frequency as an accurate estimation method. The proposed method also uses a comb filter with controllable passbands to combine the two estimation methods. Simulation results showed that: (1) the proposed method can estimate F0s for clean speech as accurate as the method using only instantaneous frequency, (2) the proposed method can robustly estimate F0s for speech with aperiodic noise in comparison with the other methods such as the cepstrum method, and (3) the proposed method had the capability of estimating F0s for speech with periodic noise. 1. Yuichi Ishimoto, Masashi Unoki, Masato Akagi |
INTERSPEECH | 2 |
| 1999 | Segregation of vowel in background noise using the model of segregating two acoustic sources based on auditory scene analysis
Masashi Unoki, Masato Akagi |
EUROSPEECH | 1 |
| 1999 | A method of signal extraction from noisy signal based on auditory scene analysis
Masashi Unoki, Masato Akagi |
Speech Commun. | 1 |
| 1998 | A time-varying, analysis/synthesis auditory filterbank using the gammachirpabstractA time-varying, analysis/synthesis auditory filterbank has been developed using a new implementation of the "gammachirp", which has been shown to be an excellent function for the asymmetric, level-dependent auditory filter. The gammachirp filter is shown to be implemented through a combination of a gammatone filter and an IIR asymmetric compensation filter; which largely reduces the computational cost for time-varying filtering. The gammachirp filterbank is designed using a linear gammatone filterbank and a bank of time-varying asymmetric compensation filters controlled by the sound pressure level estimated at the output of the filterbank. Since the inverse filter of the asymmetric compensation filter is always stable, it is possible to resynthesize signals from time-varying, level-dependent auditory representations. The resynthesis error is only determined by the linear analysis/synthesis gammatone filterbank. The proposed filterbank is applicable to various types of signal processing required to model human auditory filtering. Toshio Irino, Masashi Unoki |
ICASSP | 2 |
| 1998 | Signal extraction from noisy signal based on auditory scene analysis
Masashi Unoki, Masato Akagi |
ICSLP | 1 |
| 1997 | A method of signal extraction from noisy signal
Masashi Unoki, Masato Akagi |
EUROSPEECH | 1 |