EDBT 2026 Demo / reviewers in the wild / expert
Masato Akagi
dblp:18/2260
· DBLP profile ↗
85ranked-venue papers
9as first author
12since 2021 · last 2023
0000-0003-2450-6754ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 52 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Music Theory-Inspired Acoustic Representation for Speech Emotion RecognitionabstractThis research presents a music theory-inspired acoustic representation (hereafter, MTAR) to address improved speech emotion recognition. The recognition of emotion in speech and music is developed in parallel, yet a relatively limited understanding of MTAR for interpreting speech emotions is involved. In the present study, we use music theory to study representative acoustics associated with emotion in speech from vocal emotion expressions and auditory emotion perception domains. In experiments assessing the role and effectiveness of the proposed representation in classifying discrete emotion categories and predicting continuous emotion dimensions, it shows promising performance compared with extensively used features for emotion recognition based on the spectrogram, Mel-spectrogram, Mel-frequency cepstral coefficients, VGGish, and the large baseline feature sets of the INTERSPEECH challenges. This proposal opens up a novel research avenue in developing a computational acoustic representation of speech emotion via music theory. Xingfeng Li 0001, Desheng Hu, Qingchen Zhang 0001, Zhengxia Wang, Masashi Unoki, Masato Akagi |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2022 | Speak Like a Professional: Increasing Speech Intelligibility by Mimicking Professional Announcer Voice with Voice ConversionabstractIn most of practical scenarios, the announcement system must deliver speech messages in a noisy environment, in which the background noise cannot be cancelled out. The local noise reduces speech intelligibility and increases listening effort of the listener, hence hamper the effectiveness of announcement system. There has been reported that voices of professional announcers are clearer and more comprehensive than that of nonexpert speakers in noisy environment. This finding suggests that the speech intelligibility might be related to the speaking style of professional announcer, which can be adapted using\nvoice conversion method. Motivated by this idea, this paper proposes a speech intelligibility enhancement in noisy environment by applying voice conversion method on non-professional voice. We discovered that the professional announcers and nonprofessional speakers are clusterized into different clusters on the speaker embedding plane. This implies that the speech intelligibility can be controlled as an independent feature of speaker individuality. To examine the advantage of converted voice in noisy environment, we experimented using test words masked in pink noise at different SNR levels. The results of objective\nand subjective evaluations confirm that the speech intelligibility of converted voice is higher than that of original voice in low SNR conditions. Tuan Vu Ho, Maori Kobayashi, Masato Akagi |
INTERSPEECH | 3 |
| 2022 | Vector-quantized Variational Autoencoder for Phase-aware Speech EnhancementabstractSpeech-enhancement methods based on the complex ideal ratio mask (cIRM) have achieved promising results.These methods often deploy a deep neural network to jointly estimate the real and imaginary components of the cIRM defined in the complex domain.However, the unbounded property of the cIRM poses difficulties when it comes to effectively training a neural network.To alleviate this problem, this paper proposes a phase-aware speech-enhancement method through estimating the magnitude and phase of a complex adaptive Wiener filter.With this method, a noise-robust vector-quantized variational autoencoder is used for estimating the magnitude of the Wiener filter by using the Itakura-Saito divergence on the time-frequency domain, while the phase of the Wiener filter is estimated using a convolutional recurrent network using the scale-invariant signal-to-noise-ratio constraint in the time domain.The proposed method was evaluated on the open Voice Bank+DEMAND dataset to provide a direct comparison with other speech-enhancement methods and achieved a Perceptual Evaluation of Speech Quality score of 2.85 and ShortTime Objective Intelligibility score of 0.94, which is better than the stateof-art method based on cIRM estimation during the 2020 Deep Noise Challenge. Tuan Vu Ho, Masato Akagi, Masashi Unoki |
INTERSPEECH | 3 |
| 2022 | Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio DetectionabstractFake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%. Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki |
INTERSPEECH | 4 |
| 2022 | DDAM '22: 1st International Workshop on Deepfake Detection for Audio MultimediaabstractOver the last few years, the technology of speech synthesis and voice conversion has made significant improvement with the development of deep learning. The models can generate realistic and human-like speech. It is difficult for most people to distinguish the generated audio from the real. However, this technology also poses a great threat to the global political economy and social stability if some attackers and criminals misuse it with the intent to cause harm. In this workshop, we aim to bring together researchers from the fields of audio deepfake detection, audio deep synthesis, audio fake game and adversarial attacks to further discuss recent research and future directions for detecting deepfake and manipulated audios in multimedia. Jianhua Tao 0001, Jiangyan Yi, Cunhang Fan, Ruibo Fu, Shan Liang 0007, Pengyuan Zhang, Haizhou Li 0001, Helen M. Meng, Dong Yu 0001, Masato Akagi |
ACM Multimedia | 10 |
| 2022 | Survey on bimodal speech emotion recognition from acoustic and linguistic information fusionabstractSpeech emotion recognition (SER) is traditionally performed using merely acoustic information. Acoustic features, commonly are extracted per frame, are mapped into emotion labels using classifiers such as support vector machines for machine learning or multi-layer perceptron for deep learning. Previous research has shown that acoustic-only SER suffers from many issues, mostly on low performances. On the other hand, not only acoustic information can be extracted from speech but also linguistic information. The linguistic features can be extracted from the transcribed text by an automatic speech recognition system. The fusion of acoustic and linguistic information could improve the SER performance. This paper presents a survey of the works on bimodal emotion recognition fusing acoustic and linguistic information. Five components of bimodal SER are reviewed: emotion models, datasets, features, classifiers, and fusion methods. Some major findings, including state-of-the-art results and their methods from the commonly used datasets, are also presented to give insights for the current research and to surpass these results. Finally, this survey proposes the remaining issues in the bimodal SER research for future research directions. Bagus Tris Atmaja, Akira Sasou, Masato Akagi |
Speech Commun. | 3 |
| 2022 | Acoustic features correlated to perceived urgency in evacuation announcements
Maori Kobayashi, Yasuhiro Hamada, Masato Akagi |
Speech Commun. | 3 |
| 2021 | Acoustic and articulatory analysis and synthesis of shouted vowels
Yawen Xue, Michael Marxen, Masato Akagi, Peter Birkholz |
Comput. Speech Lang. | 3 |
| 2021 | Multi-resolution modulation-filtered cochleagram feature for LSTM-based dimensional emotion recognition from speech
Zhichao Peng, Jianwu Dang 0001, Masashi Unoki, Masato Akagi |
Neural Networks | 4 |
| 2021 | Two-stage dimensional emotion recognition by fusing predictions of acoustic and text networks using SVM
Bagus Tris Atmaja, Masato Akagi |
Speech Commun. | 2 |
| 2021 | Increasing speech intelligibility and naturalness in noise based on concepts of modulation spectrum and modulation transfer function
Thuan Van Ngo, Rieko Kubo, Masato Akagi |
Speech Commun. | 3 |
| 2021 | $F_0$-Noise-Robust Glottal Source and Vocal Tract Analysis Based on ARX-LF ModelabstractThis paper proposes a robust automatic speech analysis method based on a source-filter model constructed of an Auto-Regressive eXogenous (ARX) model and the Liljencrants-Fant (LF) model. The proposed method estimates glottal source waveform and vocal tract shape parameters using an analysis-by-synthesis approach. Structurally, the first step is to initialize the glottal source parameters using the inverse filter method, and the second step is to simultaneously estimate the glottal source waveform and the vocal tract shape parameters using an analysis-by-synthesis approach with an iterative algorithm. The proposed method was verified on synthetic voices with different glottal noise (signal to noise ratio) from 0 dB to 50 dB and different fundamental frequency ( F_0 ) from 80 Hz to 320 Hz levels. The results show that the proposed method achieved a much higher estimation accuracy than that of the state-of-the-art inverse filtering methods on both different glottal noise and different F_0 levels. Jianhua Tao 0001, Donna Erickson, Bin Liu 0041, Masato Akagi |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Multitask Learning and Multistage Fusion for Dimensional Audiovisual Emotion RecognitionabstractDue to its ability to accurately predict emotional state using multimodal features, audiovisual emotion recognition has recently gained more interest from researchers. This paper proposes two methods to predict emotional attributes from audio and visual data using a multitask learning and a fusion strategy. First, multitask learning is employed by adjusting three parameters for each attribute to improve the recognition rate. Second, a multistage fusion is proposed to combine results from various modalities' final prediction. Our approach used multitask learning, employed at unimodal and early fusion methods, shows improvement over single-task learning with an average CCC score of 0.431 compared to 0.297. A multistage method, employed at the late fusion approach, significantly improved the agreement score between true and predicted values on the development set of data (from [0.537, 0.565, 0.083] to [0.68, 0.656, 0.443]) for arousal, valence, and liking. Bagus Tris Atmaja, Masato Akagi |
ICASSP | 2 |
| 2020 | Segment-Level Effects of Gender, Nationality and Emotion Information on Text-Independent Speaker VerificationabstractSpeaker embeddings extracted from neural network (NN) achieve excellent performance on general speaker verification (SV) missions. Most current SV systems use only speaker labels. Therefore, the interaction between different types of domain information decrease the prediction accuracy of SV. To overcome this weakness and improve SV performance, four effective SV systems were proposed by using gender, nationality, and emotion information to add more constraints in the NN training stage. More specifically, multitask learning-based systems which including multitask gender (MTG), multitask nationality (MTN) and multitask gender and nationality (MTGN) were used to enhance gender and nationality information learning. Domain adversarial training-based system which including emotion domain adversarial training (EDAT) was used to suppress different emotions information learning. Experimental results indicate that encouraging gender and nationality information and suppressing emotion information learning improve the performance of SV. In the end, our proposed systems achieved 16.4 and 22.9% relative improvements in the equal error rate for MTL- and DAT-based systems, respectively. Kai Li 0018, Masato Akagi, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2020 | Comparison of Glottal Source Parameter Values in Emotional VowelsabstractSince glottal source plays an important role for expressing emotions in speech, it is crucial to compare a set of glottal source parameter values to find differences in these expressions of emotions for emotional speech recognition and synthesis. This paper focuses on comparing a set of glottal source parameter values among varieties of emotional vowels /a/ (joy, neutral, anger, and sadness) using an improved ARX-LF model algorithm. The set of glottal source parameters included in the comparison were T_p, T_e, T_a, E_e, and F_0(1/T_0) in the LF model; parameter values were divided into 5 levels according to that of neutral vowel. Results showed that each emotion has its own levels for each set of the glottal source parameter value. These findings could be used for emotional speech recognition and synthesis. Jianhua Tao 0001, Bin Liu 0041, Donna Erickson, Masato Akagi |
INTERSPEECH | 5 |
| 2020 | On The Differences Between Song and Speech Emotion Recognition: Effect of Feature Sets, Feature Types, and ClassifiersabstractIn this paper, we argue that singing voice (song) is more emotional than speech. We evaluate different features sets, feature types, and classifiers on both song and speech emotion recognition. Three feature sets: GeMAPS, pyAudioAnalysis, and LibROSA; two feature types, low-level descriptors and high-level statistical functions; and four classifiers: multilayer perceptron, LSTM, GRU, and convolution neural networks; are examined on both songand speech data with the same parameter values. The results show no remarkable difference between song and speech data on using the same method. Comparisons of two results reveal that song is more emotional than speech. In addition, high-level statistical functions of acoustic features gained higher performance than low-level descriptors in this classification task. This result strengthens the previous finding on the regression task which reported the advantage use of high-level features. Bagus Tris Atmaja, Masato Akagi |
TENCON | 2 |
| 2020 | Predicting Valence and Arousal by Aggregating Acoustic Features for Acoustic-Linguistic Information FusionabstractThis paper presents an evaluation of acoustic feature aggregation and acoustic-linguistic features combination for valence and arousal prediction within a speech. First, acoustic features were aggregated from chunk-based processing for story-based processing. We evaluated mean and maximum aggregation methods for those acoustic features and compared the results with the baseline, which used majority voting aggregation. Second, the extracted acoustic features are combined with linguistic features for predicting valence and arousal categories: low, medium, or high. The unimodal result using acoustic features aggregation showed an improvement over the baseline majority voting on development partition for the same acoustic feature set. The bimodal results (by combining acoustic and linguistic information at the feature level) improved both development and test scores over the official baseline. This combination of acoustic-linguistic information targeted speech-based applications where acoustic and linguistic features can be extracted from the sole speech modality. Bagus Tris Atmaja, Yasuhiro Hamada, Masato Akagi |
TENCON | 3 |
| 2020 | Combining F0 and non-negative constraint robust principal component analysis for singing voice separation
Masato Akagi |
Signal Process. | 2 |
| 2020 | Effect of articulatory and acoustic features on the intelligibility of speech in noise: An articulatory synthesis study
Thuan Van Ngo, Masato Akagi, Peter Birkholz |
Speech Commun. | 2 |
| 2019 | The Contribution of Acoustic Features Analysis to Model Emotion Perceptual Process for Language DiversityabstractThe multi-layered perceptual process of emotion in human speech plays an essential role in the field of affective computing for underlying a speaker’s state. However, a comprehensive process analysis of emotion perception is still challenging due to the lack of powerful acoustic features allowing accurate inference of emotion across speaker and language diversities. Most previous research works study acoustic features mostly using Fourier transform, short time Fourier transform or linear predictive coding. Even though these features may be useful for stationary signal within short frames, they may not capture the localized event adequately as speech transmits emotion information dynamically over time. This case introduces a set of acoustic features via wavelet transform analysis of the speech signal, and specifically, models the perceptual process of emotion for language diversity. For this aim, the proposed features are analyzed in a three-layer emotion perception model across multiple languages. Experiments show that the proposed acoustic features significantly enhance the perceptual process of emotion and render a better result in multilingual emotion recognition when compared it to the widely used prosodic and spectral features, as well as their combination in literature. Xingfeng Li 0001, Masato Akagi |
INTERSPEECH | 2 |
| 2019 | Blind monaural singing voice separation using rank-1 constraint robust principal component analysis and vocal activity detection
Masato Akagi |
Neurocomputing | 2 |
| 2019 | Improving multilingual speech emotion recognition by combining acoustic features in a three-layer model
Xingfeng Li 0001, Masato Akagi |
Speech Commun. | 2 |
| 2018 | Auditory-Inspired End-to-End Speech Emotion Recognition Using 3D Convolutional Recurrent Neural Networks Based on Spectral-Temporal RepresentationabstractThe human auditory system has far superior emotion recognition abilities compared with recent speech emotion recognition systems, so research has focused on designing emotion recognition systems by mimicking the human auditory system. Psychoacoustic and physiological studies indicate that the human auditory system decomposes speech signals into acoustic and modulation frequency components, and further extracts temporal modulation cues. Speech emotional states are perceived from temporal modulation cues using the spectral and temporal receptive field of the neuron. This paper proposes an emotion recognition system in an end-to-end manner using three-dimensional convolutional recurrent neural networks (3D-CRNNs) based on temporal modulation cues. Temporal modulation cues contain four-dimensional spectral-temporal (ST) integration representations directly as the input of 3D-CRNNs. The convolutional layer is used to extract high-level multiscale ST representations, and the recurrent layer is used to extract long-term dependency for emotion recognition. The proposed method was verified on the IEMOCAP database. The results show that our proposed method can exceed the recognition accuracy compared to that of the state-of-the-art systems. Zhichao Peng, Masashi Unoki, Jianwu Dang 0001, Masato Akagi |
ICME | 5 |
| 2018 | A Three-Layer Emotion Perception Model for Valence and Arousal-Based Detection from Multilingual SpeechabstractAutomated emotion detection from speech has recently shifted from monolingual to multilingual tasks for human-like interaction in real-life where a system can handle more than a single input language. However, most work on monolingual emotion detection is difficult to generalize in multiple languages, because the optimal feature sets of the work differ from one language to another. Our study proposes a framework to design, implement and validate an emotion detection system using multiple corpora. A continuous dimensional space of valence and arousal is first used to describe the emotions. A three-layer model incorporated with fuzzy inference systems is then used to estimate two dimensions. Speech features derived from prosodic, spectral and glottal waveform are examined and selected to capture emotional cues. The results of this new system outperformed the existing state-of-the-art system by yielding a smaller mean absolute error and higher correlation between estimates and human evaluators. Moreover, results for speaker independent validation are comparable to human evaluators. Xingfeng Li 0001, Masato Akagi |
INTERSPEECH | 2 |
| 2018 | Voice conversion for emotional speech: Rule-based synthesis with degree of emotion controllable in dimensional space
Yawen Xue, Yasuhiro Hamada, Masato Akagi |
Speech Commun. | 3 |
| 2017 | Weighted Robust Principal Component Analysis with Gammatone Auditory Filterbank for Singing Voice Separation
Masato Akagi |
ICONIP (6) | 2 |
| 2016 | Multilingual Speech Emotion Recognition System Based on a Three-Layer Model
Xingfeng Li 0001, Masato Akagi |
INTERSPEECH | 2 |
| 2013 | Comparative investigation of objective speech intelligibility prediction measures for noise-reduced signals in Mandarin and JapaneseabstractIn this paper, eight state-of-the-art objective speech intelligibility prediction measures are comparatively investigated for noisy signals before and after noise-reduction processing between Mandarin and Japanese. Clean speech signals (Chinese words and Japanese words) were first corrupted by three types of noise at two signal-to-noise ratios and then processed by normal-hearing listeners for recognition, whose intelligibility was subsequently predicted by objective measures. Further investigations were conducted for objective measures in predicting speech intelligibility of noise-reduced signals between subjective evaluation scores and objective prediction results, and of noisy signals before and after noise-reduction processing, in terms of correlation analysis and prediction errors. Results showed that the majority of objective measures behave differently for Mandarin and Japanese in predicting the subjective ratings, and the STOI measure consistently provided the best ability in predicting the effect on speech intelligibility of the noise-reduction processing for both Mandarin and Japanese. Fei Chen 0011, Masato Akagi, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2012 | Evaluation of objective intelligibility prediction measures for noise-reduced signals in mandarinabstractIn this paper, the performance of eight state-of-the-art objective measures is evaluated in terms of predicting speech intelligibility in Mandarin of the processed signals by noise-reduction algorithms. The speech signals were first corrupted by three types of noises at two signal-to-noise ratios and subsequently processed by four classes of noise reduction algorithms, followed by objective intelligibility prediction. The subjective intelligibility ratings were obtained through a set of listening tests. Further investigation was conducted for objective measures in predicting speech intelligibility of noisy signals before and after noise-reduction processing in terms of correlation analysis and prediction errors. The analysis results reported here do provide valuable hints for analyzing and optimizing noise-reduction algorithms for Mandarin. Risheng Xia, Masato Akagi, Yonghong Yan 0002 |
ICASSP | 3 |
| 2011 | Voice Activity Detection in MTF-Based Power Envelope Restoration
Masashi Unoki, Xugang Lu, Rico Petrick, Shota Morita, Masato Akagi, Rüdiger Hoffmann |
INTERSPEECH | 5 |
| 2011 | Two-stage binaural speech enhancement with Wiener filter for high-quality speech communication
Shuichi Sakamoto, Satoshi Hongo, Masato Akagi, Yôiti Suzuki |
Speech Commun. | 4 |
| 2010 | A DOA estimation algorithm based on equalization-cancellation theoryabstractDirection of arrival (DOA) estimation plays an important role in multi-channel (binaural) speech enhancement systems and auditory humanoid robots. A number of localization methods have been presented, however, most of them require a large array of microphones, or cannot adapt to some special conditions, e.g., humanoid robot with the effect of head-related transfer function (HRTF). In this paper, we propose a two-microphone DOA estimation algorithm, namely EC-Beam, which applies equalization-cancellation (EC) model to DOA estimation through beamformer-based technique. Specifically, the EC model is integrated into beamforming to remove the signal components from a given direction and yield the energy of the remaining signals from other directions. Through searching several DOA candidates in the space, the estimation of DOA is finally determined as the direction at which the energy of the remaining signal reaches the minimum. Interpolation method is further exploited in EC-Beam to estimate non-beamformed directions. Experimental results showed that the EC-Beam with only two microphones is able to estimate accurately the DOA of target signal in various noise conditions, and well adapted to binaural hearing systems. Duc Thanh Chau, Masato Akagi |
INTERSPEECH | 3 |
| 2009 | Psychoacoustically-motivated adaptive beta-order generalized spectral subtraction for cochlear implant patientsabstractMany cochlear implant (CI) users are able to understand speech in quiet listening conditions, however, CI users' speech recognition deteriorates rapidly as the level of background noise increases. To make CI more applicable in reallife environments, noise reduction is needed in CI processor. Recently, we presented a psychoacoustically-motivated adaptive β-order generalized spectral subtraction (GSS) which deals with the weakness of the traditional SS algorithms [9, 10]. To apply this adaptive β-order GSS into CI processor, in this paper, we investigate the effects of noise estimation approaches and residual noise components for the proposed adaptive β-order GSS. Word-in-sentence recognition in steady white noise and speech babble noise was measured in four CI users. Experimental results showed that 1) noise estimation significantly affected performance of the proposed algorithm, 2) the algorithm with the least residual noise components was preferred by CI subjects, and 3) the proposed psychoacoustically-motivated adaptive β-order GSS outperformed the traditional SS algorithms. Qian-Jie Fu, Hui Jiang 0001, Masato Akagi |
ICASSP | 4 |
| 2009 | Efficient modeling of temporal structure of speech for applications in voice transformationabstractAims of voice transformation are to change styles of given utterances. Most voice transformation methods process speech signals in a time-frequency domain. In the time domain, whenprocessing spectral information, conventional methods do not consider relations between neighboring frames. If unexpected modifications happen, there are discontinuities between frames,which lead to the degradation of the transformed speech quality. This paper proposes a new modeling of temporal structure of speech to ensure the smoothness of the transformed speech for improving the quality of transformed speech in the voice transformation. In our work, we propose an improvement of the temporal decomposition (TD) technique, which decomposes a speech signal into event targets and event functions, to modelthe temporal structure of speech. The TD is used to control the spectral dynamics and to ensure the smoothness of transformed speech. We investigate the TD in two applications, concatenative speech synthesis and spectral voice conversion. Experimental results confirm the effectiveness of TD in terms of improving the quality of the transformed speech. Binh Phu Nguyen, Masato Akagi |
INTERSPEECH | 2 |
| 2008 | Psychoacoustically-motivated adaptive β-order generalized spectral subtraction based on data-driven optimizationabstractTo mitigate the performance limitations caused by the constant spectral order β in the traditional spectral subtraction methods, we previously presented an adaptive β-order generalizedspectral subtraction (GSS) in which the spectral order β is updated in a heuristic way [10]. In this paper, we propose a psychoacoustically-motivated adaptive β-order GSS, by considering that different frequency bands contribute different amounts to speech intelligibility (i.e., the bandimportancefunction). Specifically, in this proposed adaptiveβ-order GSS, the tendency of spectral order β to change with the input local signal-to-noise ratio (SNR) is quantitatively approximated by a sigmoid function, which is derived through a data-driven optimization procedure by minimizing the intelligibility-weighted distance between the desired speech spectrum and its estimate. The inherent parameters of the sigmoid function are further optimized with the data-driven optimizationprocedure. Experimental results indicate that theproposed psychoacoustically-motivated adaptive β-order GSS yields great improvements over the traditional spectral subtraction methods with the intelligibility-weighted measures. Hui Jiang 0001, Masato Akagi |
INTERSPEECH | 3 |
| 2008 | High-quality analysis/synthesis method based on temporal decomposition for speech modificationabstractThe challenge of speech modification is to flexibly modify the speech without degrading speech quality. The conventional methods are limited by their inability to flexibly control speech signals in time and frequency domains. This causes degradationof the quality of modified speech. This paper proposes a highquality analysis/synthesis method for speech modification. To control the temporal evolution, we use a speech analysis techniquecalled temporal decomposition (TD), which decomposes a speech signal into event targets and event functions. The same event functions evaluated for the spectral parameters are alsoused to model the temporal evolution of the excitation parameters. The event functions describe the temporal evolution of the spectral and excitation parameters, and the event targets represent the “ideal” spectral parameters. To flexibly control speech signals in both time and frequency domains, we propose new methods to model the event functions and the event targets. The experimental results show that our proposed analysis/synthesis method produces high-quality synthesized speech, and allows the flexibility to modify speech signals. Binh Phu Nguyen, Takeshi Shibata, Masato Akagi |
INTERSPEECH | 3 |
| 2008 | Robust front end processing for speech recognition in reverberant environments: utilization of speech characteristicsabstractThis paper proposes two methods for robust automatic speech recognition (ASR) in reverberant environments. Unlike other methods which mostly apply inverse filtering by blindly estimated room impulse responses to achieve dereverberation, theproposed methods are based on the utilization of the characteristics of speech. The first method - Harmonicity based Feature Analysis – takes advantage of the harmonic componentsof speech, which are assumed to be undistorted. The second method - Temporal Power Envelope Feature Analysis – utilizes the temporal modulation structure of speech, representing the phoneme level temporal events which contain most intelligibility information. Both methods increase the recognition performance remarkably in a different way. Combining both of them connects their individual advantages. In order to examine theperformance of utilizing harmonicity and modulation temporal structure for reverberant ASR, the methods are tested in clean and reverberant training. As results show, even in strong reverberantconditions both methods obtain practical applicableperformance for reverberant training. In addition, besides testing their performance in dependency on the reverberation time, their performance considering the speaker-to-microphone distanceis tested, which is another new contributions in this paper. Rico Petrick, Xugang Lu, Masashi Unoki, Masato Akagi, Rüdiger Hoffmann |
INTERSPEECH | 4 |
| 2008 | Adaptive beta-order generalized spectral subtraction for speech enhancement
Shuichi Sakamoto, Satoshi Hongo, Masato Akagi, Yôiti Suzuki |
Signal Process. | 4 |
| 2008 | A three-layered model for expressive speech perception
Chun-Fang Huang, Masato Akagi |
Speech Commun. | 2 |
| 2007 | A rule-based speech morphing for verifying a expressive speech perception model
Chun-Fang Huang, Masato Akagi |
INTERSPEECH | 2 |
| 2007 | Noise reduction based on adaptive β-order generalized spectral subtraction for speech enhancement
Shuichi Sakamoto, Satoshi Hongo, Masato Akagi, Yôiti Suzuki |
INTERSPEECH | 4 |
| 2007 | A flexible spectral modification method based on temporal decomposition and Gaussian mixture modelabstractManipulating spectral structure often leads to degradation of speech quality, which is mainly due to insufficient smoothness of the modified spectra between frames, and ineffective spectral modification. This paper presents a new spectral modification method to improve the quality of modified speech. If frames are processed independently, discontinuous features may be generated. Therefore, a speech analysis technique called temporal decomposition (TD), which decomposes speech into event targets and event functions, is used to model the spectral evolution effectively. Instead of modifying the speech spectra frame by frame, we only need to modify event targets and event functions. This feature leads to easy modification of the speech spectra, and the smoothness of modified speech is ensured by the shape of event functions. To improve spectral modification, we explore Gaussian mixture model parameters (spectral-GMM parameters) to model the spectral envelope of each event target, and develop a new algorithm for modifying spectral-GMM parameters in accordance with formant scaling factors. We first evaluate the effectiveness of our proposed method in spectra modeling, and then apply it to two areas which require different amounts of spectral modification, emotional speech synthesis and voice gender conversion. Experimental results show that the effectiveness of our proposed method is verified for spectra modeling and spectral modification. Binh Phu Nguyen, Masato Akagi |
INTERSPEECH | 2 |
| 2007 | Vocal conversion from speaking voice to singing voice using STRAIGHT
Takeshi Saitou, Masataka Goto, Masashi Unoki, Masato Akagi |
INTERSPEECH | 4 |
| 2007 | Method of LP-based blind restoration for improving intelligibility of bone-conducted speech
Thang Tat Vu, Germine Seide, Masashi Unoki, Masato Akagi |
INTERSPEECH | 4 |
| 2007 | Limited error based event localizing temporal decomposition and its application to variable-rate speech coding
Phu Chien Nguyen, Masato Akagi, Binh Phu Nguyen |
Speech Commun. | 2 |
| 2006 | Improved hybrid microphone array post-filter by integrating a robust speech absence probability estimator for speech enhancement
Masato Akagi, Yôiti Suzuki |
INTERSPEECH | 2 |
| 2006 | A robust feature extraction based on the MTF concept for speech recognition in reverberant environment
Xugang Lu, Masashi Unoki, Masato Akagi |
INTERSPEECH | 3 |
| 2006 | Communication Between Speech Production and Perception Within the Brain-Observation and Simulation
Jianwu Dang 0001, Masato Akagi, Kiyoshi Honda |
J. Comput. Sci. Technol. | 2 |
| 2006 | A noise reduction system based on hybrid noise estimation technique and post-filtering in arbitrary noise environments
Masato Akagi |
Speech Commun. | 2 |
| 2005 | Toward a Rule-Based Synthesis of Emotional Speech on Linguistic Descriptions of Perception
Chun-Fang Huang, Masato Akagi |
ACII | 2 |
| 2005 | A noise reduction system in arbitrary noise environments and its applications to speech enhancement and speech recognitionabstractThe paper proposes a novel noise reduction system in arbitrary noise environments, consisting of localized and nonlocalized noises, where few existing systems work well. In the proposed system, localized noises are estimated and reduced by the hybrid noise estimation technique we previously proposed and spectral subtraction. Non-localized noises are reduced by a post-filter whose performance is further improved by a novel estimator for the a priori speech absence probability calculated under the assumption of a diffuse noise field. Experimental results show that the proposed system results in significant improvements in terms of speech quality measures and speech recognition performance in various noise conditions. Xugang Lu, Masato Akagi |
ICASSP (3) | 3 |
| 2005 | A multi-layer fuzzy logical model for emotional speech perceptionabstractPerceiving emotion from speech plays an important role in human communication. In the filed of computer science, previous works tried to build the relationship between categories of emotion in speech and acoustic features by statistics. But they do not provide ways and means of both perception and expression. According to an observation of a phenomenon how we perceive emotion from speech, a threelayer model is proposed in this paper. A new layer primitive features which is a linguistic form of adjectives for describing speech voice is added. The purpose of this study is building and proving the model by a two-phase approach. In this paper, we introduce to the model and describe the first phase of building of the model using a top-down method to build two relationships. Firstly, the building of the first relationship between the emotional speech and the primitive feature is described. Three experiments to select suitable primitive features and the application of fuzzy inference system to build the relationship are then described. Finally, the building of the second relationship between the primitive feature and the acoustic features are demonstrated. The resultant relationships show significance of the proposed three-layer model. Chun-Fang Huang, Masato Akagi |
INTERSPEECH | 2 |
| 2005 | A hybrid microphone array post-filter in a diffuse noise fieldabstractIn this paper, a hybrid post-filter for microphone arrays with the assumption of a diffuse noise field is proposed to suppress correlated as well as uncorrelated noise.In the proposed post-filter, a modified Zelinski post-filter, which is estimated using the signals on the microphone pairs on which noises are uncorrelated by considering the correlation characteristics of noise impinging on different microphone pairs, is applied to the high frequencies to suppress spatially uncorrelated noise; a single-channel Wiener post-filter is applied to the low frequencies for cancellation of spatially correlated noise.In theory, the proposed post-filter is a Wiener post-filter.In practice, experiments using multi-channel recordings were conducted, and experimental results demonstrate the usefulness and superiority of the proposed post-filter compared to other post-filters using speech quality measures and speech recognition rate. Masato Akagi |
INTERSPEECH | 2 |
| 2005 | A model for selective segregation of a target instrument sound from the mixed sound of various instruments
Masashi Unoki, Masaaki Kubo, Atsushi Haniu, Masato Akagi |
INTERSPEECH | 4 |
| 2005 | Development of an F0 control model based on F0 dynamic characteristics for singing-voice synthesis
Takeshi Saitou, Masashi Unoki, Masato Akagi |
Speech Commun. | 3 |
| 2004 | Noise reduction using hybrid noise estimation technique and post-filteringabstractIn this paper, a novel noise reduction method using hybrid noise estimation technique and post-filtering is proposed to suppress both localized and non-localized noise components which can not be dealt with by the traditional methods [2][3][4]. To do this, a hybrid noise estimation approach is proposed by combining our previously constructed multichannel noise estimation approach and a single-channel estimation approach to improve estimation accuracy for localized noise components. The non-localized noise components are suppressed by a single-channel post-filter based on an optimally modified log spectral amplitude (OM-LSA) estimator. To verify the superiorities of the proposed hybrid noise estimation approach and noise reduction system, they are compared to the multi-channel and single-channel scheme based systems under various noise conditions. Masato Akagi |
INTERSPEECH | 2 |
| 2004 | Analysis of acoustic features affecting "singing-ness" and its application to singing-voice synthesis from speaking-voiceabstractTo construct a natural singing-voice synthesis system, it is important to adequately control acoustic features such as fundamental frequency (F0), spectrum shapes, and phoneme duration in the synthesis method. This paper reveals acoustic features affecting singing-voice perception by comparative analyzing singing- and speaking-voices, and then proposes a transforming method from speaking-voice into singing-voice using STRAIGHT [1]. This method is composed of an F0 control model for generating F0 contours of singing-voices, a spectral sequence control model for modifying spectral shapes in speaking-voice, and a duration control model based on rhythm. Results showed that the proposed system could synthesize a natural singing-voice, whose sound quality is almost the same as that of real one. 1. Takeshi Saitou, Naoya Tsuji, Masashi Unoki, Masato Akagi |
INTERSPEECH | 4 |
| 2003 | Temporal decomposition: a promising approach to VQ-based speaker identificationabstractA new set of features is proposed that has been found to improve the performance of automatic speaker identification systems. The new set of features is referred to as "event targets". The new features have been derived from line spectral frequency (LSF) parameters using the so-called "temporal decomposition" (TD) technique. The number of feature vectors required for both the training and testing phases has been reduced by one-fifth compared to that of the traditional Mel-frequency cepstrum coefficients (MFCC) features, while the identification results obtained are comparable or even better. Also, we introduce one more application of TD (speaker recognition) in addition to speech coding, speech segmentation, and speech recognition. It shows that the event targets in TD can convey information about the identity of a speaker. Phu Chien Nguyen, Masato Akagi |
ICASSP (1) | 2 |
| 2003 | A method based on the MTF concept for dereverberating the power envelope from the reverberant signalabstractThis paper proposes a method for dereverberating the power envelope from the reverberant signal. This method is based on the modulation transfer function (MTF) and does not require that the impulse response of an environment be measured. It improves upon the basic model proposed by Hirobayashi et al. (1998) regarding the following problems: (i) how to precisely extract the power envelope from the observed signal; (ii) how to determine the parameters of the impulse response of the room; and (iii) a lack of consideration as to whether the MTF concept can be applied to a more realistic signal. We have shown that the proposed method can accurately dereverberate the power envelope from the reverberant signal. Masashi Unoki, Masashi Furukawa, Keigo Sakata, Masato Akagi |
ICASSP (1) | 4 |
| 2003 | Temporal decomposition: a promising approach to VQ-based speaker identificationabstractIn this paper, a new set of features is proposed that has been found to improve the performance of automatic speaker identification systems. The new set of features is referred to as "even targets". The new features have been derived from line spectral frequency (LSF) parameters using the so-called "temporal decomposition" (TD) technique. The number of feature vectors required for both training and testing phases has been reduced by one-fifth compared to that of the traditional mel-frequency cepstrum coefficients (MFCC) features, while the identification results obtained are comparable or even better. Also, this work introduces one more application of TD (speaker recognition) in addition to speech coding, speech segmentation, and speech recognition. It shows that the event targets in TD can convey information about the identity of a speaker. Phu Chien Nguyen, Masato Akagi |
ICME | 2 |
| 2003 | Efficient quantization of speech excitation parameters using temporal decomposition
Phu Chien Nguyen, Masato Akagi |
INTERSPEECH | 2 |
| 2003 | A speech dereverberation method based on the MTF concept
Masashi Unoki, Keigo Sakata, Masato Akagi |
INTERSPEECH | 3 |
| 2002 | Noise reduction using a small-scale microphone array in multi noise source environmentabstractTo construct a front-end for ASR systems using a small-scale microphone array in real environments, robustness for unstable sudden-noises, multi-noises and near-field sound sources are required. This paper proposes a front-end method for enhancing target signals that subtracts estimated noise from noisy signals by using paired microphones in each sub-band. The proposed method assumes one integrated noise source exists in each narrow sub-band, estimates its noise spectrum correctly using the cancellation method, and subtracts it from the noise-added sound spectrum using the SS. Thus, the proposed method can be used in multi noise source and near-field conditions, although it uses a small-scale microphone array consisting of only three microphones. Masato Akagi, Takashi Kago |
ICASSP | 1 |
| 2002 | Improvement of the restricted temporal decomposition method for line spectral frequency parametersabstractLine spectral frequency (LSF) parameters have not been used for the temporal decomposition (TD) method due to the stability problems in the linear preditive coding (LPC) model. To overcome this deficiency, Kim et al. introduced a method called restricted temporal decomposition (RTD)[4]. This method enforces a constraint on the event vectors to preserve their LSF ordering property, which is needed to ensure the stability of the corresponding LPC synthesis filter. Also, additional constraints are enforced on the event functions to reduce the computational cost of TD. The RTD method, however, has not completely guaranteed the LSF ordering property for the event vectors. This paper proposes an improved algorithm named modified RTD (MRTD) to solve this problem. Moreover, basing on the geometric interpretation of RTD method we find it necessary to have a new property imposed on the event functions namely the well-shapedness property, which is satisfied by adding one more constraint on them. It is also shown by experiments that spectral information of speech can be coded efficiently using MRTD based vector quantization. Phu Chien Nguyen, Masato Akagi |
ICASSP | 2 |
| 2002 | Coding speech at very low rates using straight and temporal decomposition
Phu Chien Nguyen, Takao Ochi, Masato Akagi |
INTERSPEECH | 3 |
| 2001 | A fundamental frequency estimation method for noisy speech based on instantaneous amplitude and frequencyabstractThis paper proposes a robust and accurate F0 estimation method for noisy speech. This method uses two different principles: (1) an F0 estimation based on periodicity and harmonicity of instantaneous amplitude for a robust estimation in noisy environments, and (2) an F0 estimation based on stability of instantaneous frequency as an accurate estimation method. The proposed method also uses a comb filter with controllable passbands to combine the two estimation methods. Simulation results showed that: (1) the proposed method can estimate F0s for clean speech as accurate as the method using only instantaneous frequency, (2) the proposed method can robustly estimate F0s for speech with aperiodic noise in comparison with the other methods such as the cepstrum method, and (3) the proposed method had the capability of estimating F0s for speech with periodic noise. 1. Yuichi Ishimoto, Masashi Unoki, Masato Akagi |
INTERSPEECH | 3 |
| 2001 | Spectral stability based event localizing temporal decomposition
A. C. R. Nandasena, Phu Chien Nguyen, Masato Akagi |
Comput. Speech Lang. | 3 |
| 2000 | Perception of synthesized singing voices with fine fluctuations in their fundamental frequency contours
Masato Akagi, Hironori Kitakaze |
INTERSPEECH | 1 |
| 2000 | Design of robust subtractive beamformer for noisy speech recognition
Mitsunori Mizumachi, Masato Akagi, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 1999 | An objective distortion estimator for hearing aids and its application to noise reduction
Mitsunori Mizumachi, Masato Akagi |
EUROSPEECH | 2 |
| 1999 | Segregation of vowel in background noise using the model of segregating two acoustic sources based on auditory scene analysis
Masashi Unoki, Masato Akagi |
EUROSPEECH | 2 |
| 1999 | A method of signal extraction from noisy signal based on auditory scene analysis
Masashi Unoki, Masato Akagi |
Speech Commun. | 2 |
| 1998 | Noise reduction by paired-microphones using spectral subtractionabstractThis paper proposes a method of noise reduction by paired microphones as a front-end processor for speech recognition systems. This method estimates noise using a subtractive microphone array and subtracts them from the noisy speech signal using spectral subtraction (SS). Since this method can estimate noise analytically and frame by frame, it is easy to estimate noise not depending on these acoustic properties. Therefore, this method can also reduce non-stationary noise, for example sudden noise when a door has just closed, which cannot be reduced by other SS methods. The results of computer simulations and experiments in a real environment show that this method can reduce LPC log spectral envelope distortions. Mitsunori Mizumachi, Masato Akagi |
ICASSP | 2 |
| 1998 | Spectral stability based event localizing temporal decompositionabstractA new approach to temporal decomposition (TD) of speech, called "spectral stability based event localizing temporal decomposition", abbreviated S/sup 2/ BEL-TD, is presented. The original method of TD proposed by Atal (1983) is known to have the drawbacks of high computational cost, and the instability of the number and locations of events. In S/sup 2/ BEL-TD, the event localization is performed based on a maximum spectral stability criterion. This overcomes the instability problem of events of the Atal's method. Also, S/sup 2/ BEL-TD avoids the use of the computationally costly singular value decomposition routine used in the Atal's method, thus resulting in a computationally simpler algorithm of TD. Simulation results show that an average spectral distortion of about 1.5 dB can be achieved with LSF as the spectral parameter. Also, we have shown that the temporal pattern of the speech excitation parameters can also be well described using the S/sup 2/ BEL-TD technique. A. C. R. Nandasena, Masato Akagi |
ICASSP | 2 |
| 1998 | Fundamental frequency fluctuation in continuous vowel utterance and its perception
Masato Akagi, Mamoru Iwaki, Tomoya Minakawa |
ICSLP | 1 |
| 1998 | Spectral sequence compensation based on continuity of spectral sequence
Masato Akagi, Mamoru Iwaki, Noriyoshi Sakaguchi |
ICSLP | 1 |
| 1998 | Signal extraction from noisy signal based on auditory scene analysis
Masashi Unoki, Masato Akagi |
ICSLP | 2 |
| 1997 | Noise reduction by paired microphones
Masato Akagi, Mitsunori Mizumachi |
EUROSPEECH | 1 |
| 1997 | A method of signal extraction from noisy signal
Masashi Unoki, Masato Akagi |
EUROSPEECH | 2 |
| 1996 | Modeling of contextual effects and its application to word spottingabstractWe propose a model of spectral contextual eects to simulate the superior recognition ability of humans and apply it to a front-end processor for word spotting.This model assumes that perceived spectra are inuenced by adjacent spectral peaks and that the magnitude of the inuence can be estimated by the minimum classication error criterion.Three experiments were carried out to evaluate the performance of the model.The results show that the model can compensate for neutralized spectra and bring them to their typical patterns.This improves word spotting accuracy.1. Yuji Yonezawa, Masato Akagi |
ICSLP | 2 |
| 1995 | Speaker individualities in fundamental frequency contours and its control
Masato Akagi, Taw Ienaga |
EUROSPEECH | 1 |
| 1994 | Perception of central vowel with pre- and post-anchors
Masato Akagi, Astrid van Wieringen, Louis C. W. Pols |
ICSLP | 1 |
| 1994 | Speaker individualities in speech spectral envelopesabstractThe aim of the three psychoacoustic experiments described here was to clarify whether there are speaker individualities in the spectral envelopes, in which frequency bands such individualities exist, and how frequency bands having speaker individualities can be manipulated. The LMA analysis-synthesis system was used to prepare stimuli varied specific frequency bands, and the frequency bands having speaker individualities were estimated expermentally. The results indicate that (1) speaker individualities exist in spectral envelopes, (2) these individualities are mainly at frequencies higher than 22 ERB rate (2212Hz) and vowel characteristics exist from 12 ERB rate (603Hz) to 22 ERB rate, and (3) the voice quality can be controlled by replacing the higher frequency band of one talker with that of other talkers. The replace point is the adjacent spectral local minimum below the spectral local maximum around 23 ERB rate in the spectral envelopes. Tatsuya Kitamura, Masato Akagi |
ICSLP | 2 |
| 1990 | Contextual effect models and psycho acoustic evidence for the models
Masato Akagi |
ICSLP | 1 |
| 1988 | On the application of spectrum target prediction model to speech recognitionabstractA preprocessing method is proposed for automatic speech recognition that uses a spectrum target prediction model to cope with coarticulation, one of the most serious problems in automatic speech recognition. The method is evaluated by three measures: spectral stability with respect to measuring predicted spectrum variation, and intracategory variation. Experimental results indicate that predicted spectra throughout the model are stabilized in each phoneme portion by eliminating variations of original spectra without prediction. The results also indicate that by using the preprocessing method, intracategory variation decreases and intercategory variation increases. Consequently, the spectrum target prediction model implemented as a speech-recognition preprocessor improves automatic speech recognition performance.> Masato Akagi, Yoh'ichi Tohkura |
ICASSP | 1 |