VLDB 2026 Research / reviewers in the wild / expert
Tatsuya Komatsu
dblp:136/5113
· DBLP profile ↗
37ranked-venue papers
12as first author
25since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 12 first-author · 24 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Music Tagging with Classifier Group ChainsabstractWe propose music tagging with classifier chains that model the interplay of music tags. Most conventional methods estimate multiple tags independently by treating them as multiple independent binary classification problems. This treatment overlooks the conditional dependencies among music tags, leading to suboptimal tagging performance. Unlike most music taggers, the proposed method sequentially estimates each tag based on the idea of the classifier chains. Beyond the naive classifier chains, the proposed method groups the multiple tags by category, such as genre, and performs chains by unit of groups, which we call classifier group chains. Our method allows the modeling of the dependence between tag groups. We evaluate the effectiveness of the proposed method for music tagging performance through music tagging experiments using the MTG-Jamendo dataset. Furthermore, we investigate the effective order of chains for music tagging. Takuya Hasumi, Tatsuya Komatsu, Yusuke Fujita |
ICASSP | 2 |
| 2025 | Pre-training with Synthetic Patterns for AudioabstractIn this paper, we propose to pre-train audio encoders using synthetic patterns instead of real audio data. Our proposed framework consists of two key elements. The first one is Masked Autoencoder (MAE), a self-supervised learning framework that learns from reconstructing data from randomly masked counterparts. MAEs tend to focus on low-level information such as visual patterns and regularities within data. Therefore, it is unimportant what is portrayed in the input, whether it be images, audio mel-spectrograms, or even synthetic patterns. This leads to the second key element, which is synthetic data. Synthetic data, unlike real audio, is free from privacy and licensing infringement issues. By combining MAEs and synthetic patterns, our framework enables the model to learn generalized feature representations without real data, while addressing the issues related to real audio. To evaluate the efficacy of our framework, we conduct extensive experiments across a total of 13 audio tasks and 17 synthetic datasets. The experiments provide insights into which types of synthetic patterns are effective for audio. Our results demonstrate that our framework achieves performance comparable to models pre-trained on AudioSet-2M and partially outperforms image-based pre-training methods. Yuchi Ishikawa, Tatsuya Komatsu, Yoshimitsu Aoki |
ICASSP | 2 |
| 2025 | Aligned Contrastive Learning for Text-to-Music RetrievalabstractThis paper proposes aligned contrastive learning for text-to-music retrieval. The proposed method introduces a new similarity measure, 'aligned similarity', which captures the frame-level and token-level correspondence within text and audio sequences. Unlike traditional approaches that aggregate sequence into clip-level and sentence-level embeddings, our method aligns the text token exhibiting the highest cosine similarity with each temporal frame of the audio sequence and averages these maximum similarity values across the entire sequence. This approach enables the capture of fine-grained relationships between audio and text that are often overlooked when sequences are aggregated into a single embedding. Retrieval experiments show significant performance improvements, with a notable gain being a 17.8% increase in Recall@5. Moreover, the alignment elucidates how specific audio frames correlate with textual tokens, enhancing the model's transparency and interpretability. Tatsuya Komatsu, Hokuto Munakata, Takuya Hasumi, Yusuke Fujita |
ICASSP | 1 |
| 2025 | Language-based Audio Moment RetrievalabstractIn this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict relevant moments in untrimmed long audio based on a text query. Given the lack of prior work in AMR, we first build a dedicated dataset, Clotho-Moment, consisting of large-scale simulated audio recordings with moment annotations. We then propose a Detection Transformer-based model, named Audio Moment DETR (AM-DETR), as a fundamental framework for AMR tasks. This model captures temporal dependencies within audio features, inspired by similar video moment retrieval tasks, thus surpassing conventional clip-level audio retrieval methods. Additionally, we provide manually annotated datasets to properly measure the effectiveness and robustness of our methods on real data. Experimental results show that AM-DETR, trained with Clotho-Moment, outperforms a baseline model that applies a clip-level audio retrieval method with a sliding window on all metrics, particularly improving [email protected] by 9.00 points. Our datasets and code are publicly available in https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval. Hokuto Munakata, Taichi Nishimura, Shota Nakada, Tatsuya Komatsu |
ICASSP | 4 |
| 2025 | DETECLAP: Enhancing Audio-Visual Representation Learning with Object InformationabstractCurrent audio-visual representation learning can capture rough object categories (e.g., "animals" and "instruments"), but it lacks the ability to recognize fine-grained details, such as specific categories like "dogs" and "flutes" within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification. Shota Nakada, Taichi Nishimura, Hokuto Munakata, Masayoshi Kondo, Tatsuya Komatsu |
ICASSP | 5 |
| 2025 | Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Yuchi Ishikawa, Shota Nakada, Hokuto Munakata, Kazuhiro Saito, Tatsuya Komatsu, Yoshimitsu Aoki |
INTERSPEECH | 5 |
| 2025 | Leveraging Unlabeled Audio for Audio-Text Contrastive Learning via Audio-Composed Text Features
Tatsuya Komatsu, Hokuto Munakata, Yuchi Ishikawa |
INTERSPEECH | 1 |
| 2024 | Keep Decoding Parallel With Effective Knowledge Distillation From Language Models To End-To-End Speech RecognisersabstractThis study presents a novel approach for knowledge distillation (KD) from a BERT teacher model to an automatic speech recognition (ASR) model using intermediate layers. To distil the teacher’s knowledge, we use an attention decoder that learns from BERT’s token probabilities. Our method shows that language model (LM) information can be more effectively distilled into an ASR model using both the intermediate layers and the final layer. By using the intermediate layers as distillation target, we can more effectively distil LM knowledge into the lower network layers. Using our method, we achieve better recognition accuracy than with shallow fusion of an external LM, allowing us to maintain fast parallel decoding. Experiments on the LibriSpeech dataset demonstrate the effectiveness of our approach in enhancing greedy decoding with connectionist temporal classification (CTC). Michael Hentschel, Yuta Nishikawa, Tatsuya Komatsu, Yusuke Fujita |
ICASSP | 3 |
| 2024 | Audio Difference Learning for Audio CaptioningabstractThis study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a feature representation space that preserves the relationship between audio, enabling the generation of captions that detail intricate audio information. This method employs a reference audio along with the input audio, both of which are transformed into feature representations via a shared encoder. Captions are then generated from these differential features to describe their differences. Furthermore, a unique technique is proposed that involves mixing the input audio with additional audio, and using the additional audio as a reference. This results in the difference between the mixed audio and the reference audio reverting back to the original input audio. This allows the original input’s caption to be used as the caption for their difference, eliminating the need for additional annotations for the differences. In the experiments using the Clotho and ESC50 datasets, the proposed method demonstrated an improvement in the SPIDEr score by 7% compared to conventional methods. Tatsuya Komatsu, Yusuke Fujita, Kazuya Takeda, Tomoki Toda |
ICASSP | 1 |
| 2024 | PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language DescriptionsabstractWe propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framework, we introduce the concept of speaker prompt, which describes voice characteristics (e.g., gender-neutral, young, old, and muffled) designed to be approximately independent of speaking style. Since there is no large-scale dataset containing speaker prompts, we first construct a dataset based on the LibriTTS-R corpus with manually annotated speaker prompts. We then employ a diffusion-based acoustic model with mixture density networks to model diverse speaker factors in the training data. Unlike previous studies that rely on style prompts describing only a limited aspect of speaker individuality, such as pitch, speaking speed, and energy, our method utilizes an additional speaker prompt to effectively learn the mapping from natural language descriptions to the acoustic features of diverse speakers. Our subjective evaluation results show that the proposed method can better control speaker characteristics than the methods without the speaker prompt. Audio samples are available at https://reppy4620.github.io/demo.promptttspp/. Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, Kentaro Tachibana |
ICASSP | 6 |
| 2024 | Audio Fingerprinting with Holographic Reduced Representations
Yusuke Fujita, Tatsuya Komatsu |
INTERSPEECH | 2 |
| 2024 | Universal Score-based Speech Enhancement with High Content Preservation
Robin Scheibler, Yusuke Fujita, Yuma Shirahata, Tatsuya Komatsu |
INTERSPEECH | 4 |
| 2023 | Neural Diarization with Non-Autoregressive Intermediate AttractorsabstractEnd-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single neural network. While the EEND model can produce all frame-level speaker labels simultaneously, it disregards output label dependency. In this work, we propose a novel EEND model that introduces the label dependency between frames. The proposed method generates non-autoregressive intermediate attractors to produce speaker labels at the lower layers and conditions the subsequent layers with these labels. While the proposed model works in a non-autoregressive manner, the speaker labels are refined by referring to the whole sequence of intermediate labels. The experiments with the two-speaker CALLHOME dataset show that the intermediate labels with the proposed non-autoregressive intermediate attractors boost the diarization performance. The proposed method with the deeper net-work benefits more from the intermediate labels, resulting in better performance and training throughput than EEND-EDA. Yusuke Fujita, Tatsuya Komatsu, Robin Scheibler, Yusuke Kida, Tetsuji Ogawa |
ICASSP | 2 |
| 2023 | Analysis of Biological Data of Cows and Development of Detection Systems for Calving PhaseabstractIn this paper, we analyze body surface temperature data of cows, and detect for the calving phase of cows. Firstly, appropriate preprocessing is applied to waveforms for analysis. Specifically, the waveform to be analyzed is extracted from the cow's body surface temperature data, and next outlier in the body surface temperature data is removed. Secondly approximate waveforms based on Fourier series expansion are generated. Additionally, a nominal waveform which does not correspond to calving phase is reconstructed, and feature parameters are extracted from the nominal waveform. Finally, by using the Mahalanobis distance, the proposed detect system for calving phase is developed. Furthermore, the extraction range of the waveform to be analyzed is discussed, and we show the effectiveness of the proposed detection system. Hiroto Noma, Tatsuya Komatsu, Kansei Matsumoto, Hidetoshi Oya, Ryotaro Miura, Koji Yoshioka, Yoshikatsu Hoshi |
IECON | 2 |
| 2023 | Target Vocabulary Recognition Based on Multi-Task Learning with Decomposed Teacher Sequences
Aoi Ito, Tatsuya Komatsu, Yusuke Fujita, Yusuke Kida |
INTERSPEECH | 2 |
| 2022 | Non-Autoregressive ASR with Self-Conditioned Folded EncodersabstractThis paper proposes CTC-based non-autoregressive ASR with self-conditioned folded encoders. The proposed method realizes non-autoregressive ASR with fewer parameters by folding the conventional stack of encoders into only two blocks; base encoders and folded encoders. The base en-coders convert the input audio features into a neural representation suitable for recognition. This is followed by the folded encoders applied repeatedly for further refinement. Applying the CTC loss to the outputs of all encoders enforces the consistency of the input-output relationship. Thus, folded encoders learn to perform the same operations as an encoder with deeper distinct layers. In experiments, we investigate how to set the number of layers and the number of iterations for the base and folded encoders. The results show that the proposed method achieves a performance comparable to that of the conventional method using only 38% as many parameters. Furthermore, it outperforms the conventional method when increasing the number of iterations. Tatsuya Komatsu |
ICASSP | 1 |
| 2022 | Self-Supervised Learning Method Using Multiple Sampling Strategies for General-Purpose Audio RepresentationabstractWe propose a self-supervised learning method using multiple sampling strategies to obtain general-purpose audio representation. Multiple sampling strategies are used in the proposed method to construct contrastive losses from different perspectives and learn representations based on them. In this study, in addition to the widely used clip-level sampling strategy, we introduce two new strategies, a frame-level strategy and a task-specific strategy. The proposed multiple strategies improve the performance of frame-level classification and other tasks like pitch detection, which are not the focus of the conventional single clip-level sampling strategy. We pre-trained the method on a subset of Audioset and applied it to a downstream task with frozen weights. The proposed method improved clip classification, sound event detection, and pitch detection performance by 25 %, 20 %, and 3.6 %. Ibuki Kuroyanagi, Tatsuya Komatsu |
ICASSP | 2 |
| 2022 | Better Intermediates Improve CTC InferenceabstractThis paper proposes a method for improved CTC inference with searched intermediates and multi-pass conditioning.The paper first formulates self-conditioned CTC as a probabilistic model with an intermediate prediction as a latent representation and provides a tractable conditioning framework.We then propose two new conditioning methods based on the new formulation:(1) Searched intermediate conditioning that refines intermediate predictions with beam-search, (2) Multi-pass conditioning that uses predictions of previous inference for conditioning the next inference.These new approaches enable better conditioning than the original self-conditioned CTC during inference and improve the final performance.Experiments with the LibriSpeech dataset show relative 3%/12% performance improvement at the maximum in test clean/other sets compared to the original selfconditioned CTC. Tatsuya Komatsu, Yusuke Fujita, Jaesong Lee, Lukas Lee, Shinji Watanabe 0001, Yusuke Kida |
INTERSPEECH | 1 |
| 2022 | InterAug: Augmenting Noisy Intermediate Predictions for CTC-based ASRabstractThis paper proposes InterAug: a novel training method for CTC-based ASR using augmented intermediate representations for conditioning.The proposed method exploits the conditioning framework of self-conditioned CTC to train robust models by conditioning with "noisy" intermediate predictions.During the training, intermediate predictions are changed to incorrect intermediate predictions, and fed into the next layer for conditioning.The subsequent layers are trained to correct the incorrect intermediate predictions with the intermediate losses.By repeating the augmentation and the correction, iterative refinements, which generally require a special decoder, can be realized only with the audio encoder.To produce noisy intermediate predictions, we also introduce new augmentation: intermediate feature space augmentation and intermediate token space augmentation that are designed to simulate typical errors.The combination of the proposed InterAug framework with new augmentation allows explicit training of the robust audio encoders.In experiments using augmentations simulating deletion, insertion, and substitution error, we confirmed that the trained model acquires robustness to each error, boosting the speech recognition performance of the strong self-conditioned CTC baseline. Yu Nakagome, Tatsuya Komatsu, Yusuke Fujita, Shuta Ichimura, Yusuke Kida |
INTERSPEECH | 2 |
| 2022 | Alternate Intermediate Conditioning with Syllable-Level and Character-Level Targets for Japanese ASR
Yusuke Fujita, Tatsuya Komatsu, Yusuke Kida |
SLT | 2 |
| 2022 | Interdecoder: using Attention Decoders as Intermediate Regularization for CTC-Based Speech RecognitionabstractWe propose InterDecoder: a new non-autoregressive automatic speech recognition (NAR-ASR) training method that injects the advantage of token-wise autoregressive decoders while keeping the efficient non-autoregressive inference. The NAR-ASR models are often less accurate than autoregressive models such as Transformer decoder, which predict tokens conditioned on previously predicted tokens. The Inter-Decoder regularizes training by feeding intermediate encoder outputs into the decoder to compute the token-level prediction errors given previous ground-truth tokens, whereas the widely used Hybrid CTC/Attention model uses the decoder loss only at the final layer. In combination with Self-conditioned CTC, which uses the Intermediate CTC predictions to condition the encoder, performance is further improved. Experiments on the Librispeech and Tedlium2 dataset show that the proposed method shows a relative 6% WER improvement at the maximum compared to the conventional NAR-ASR methods. Tatsuya Komatsu, Yusuke Fujita |
SLT | 1 |
| 2021 | A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text GenerationabstractNon-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to autoregressive baselines. Showing great potential for real-time applications, an increasing number of NAR models have been explored in different fields to mitigate the performance gap against AR models. In this work, we conduct a comparative study of various NAR modeling methods for end-to-end automatic speech recognition (ASR). Experiments are performed in the state-of-the-art setting using ESPnet. The results on various tasks provide interesting findings for developing an understanding of NAR ASR, such as the accuracy-speed trade-off and robustness against long-form utterances. We also show that the techniques can be combined for further improvement and applied to NAR end-to-end speech translation. All the implementations are publicly available to encourage further research in NAR speech processing. Yosuke Higuchi, Nanxin Chen, Yuya Fujita, Hirofumi Inaguma, Tatsuya Komatsu, Jaesong Lee, Jumon Nozaki, Tianzi Wang, Shinji Watanabe 0001 |
ASRU | 5 |
| 2021 | Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTSabstractWe propose a method for obtaining disentangled speaker and language representations via mutual information minimization and domain adaptation for cross-lingual text-to-speech (TTS) synthesis. The proposed method extracts speaker and language embeddings from acoustic features by a speaker encoder and a language encoder. Then the proposed method applies domain adaptation on the two embeddings to obtain language-invariant speaker embedding and speaker-invariant language embedding. To get more disentangled representations, the proposed method further uses mutual information minimization between the two embeddings to remove entangled information within each embedding. Disentangled representations of speaker and language are critical for cross-lingual TTS synthesis since entangled representations make it difficult to maintain speaker identity information when changing the language representation and consequently causes performance degradation. We evaluate the proposed method using English and Japanese multi-speaker datasets with a total of 207 speakers. Experimental result demonstrates that the proposed method significantly improves the naturalness and speaker similarity of both intra-lingual and cross-lingual TTS synthesis. Furthermore, we show that the proposed method has a good capability of maintaining the speaker identity between languages. Detai Xin, Tatsuya Komatsu, Shinnosuke Takamichi, Hiroshi Saruwatari |
ICASSP | 2 |
| 2021 | Acoustic Event Detection with Classifier ChainsabstractThis paper proposes acoustic event detection (AED) with classifier chains, a new classifier based on the probabilistic chain rule. The proposed AED with classifier chains consists of a gated recurrent unit and performs iterative binary detection of each event one by one. In each iteration, the event's activity is estimated and used to condition the next output based on the probabilistic chain rule to form classifier chains. Therefore, the proposed method can handle the interdependence among events upon classification, while the conventional AED methods with multiple binary classifiers with a linear layer and sigmoid function have placed an assumption of conditional independence. In the experiments with a real-recording dataset, the proposed method demonstrates its superior AED performance to a relative 14.80% improvement compared to a convolutional recurrent neural network baseline system with the multiple binary classifiers. Tatsuya Komatsu, Shinji Watanabe 0001, Koichi Miyazaki, Tomoki Hayashi |
Interspeech | 1 |
| 2021 | Relaxing the Conditional Independence Assumption of CTC-Based ASR by Conditioning on Intermediate PredictionsabstractThis paper proposes a method to relax the conditional independence assumption of connectionist temporal classification (CTC)-based automatic speech recognition (ASR) models.We train a CTC-based ASR model with auxiliary CTC losses in intermediate layers in addition to the original CTC loss in the last layer.During both training and inference, each generated prediction in the intermediate layers is summed to the input of the next layer to condition the prediction of the last layer on those intermediate predictions.Our method is easy to implement and retains the merits of CTC-based ASR: a simple model architecture and fast decoding speed.We conduct experiments on three different ASR corpora.Our proposed method improves a standard CTC model significantly (e.g., more than 20 % relative word error rate reduction on the WSJ corpus) with a little computational overhead.Moreover, for the TEDLIUM2 corpus and the AISHELL-1 corpus, it achieves a comparable performance to a strong autoregressive model with beam search, but the decoding speed is at least 30 times faster. Jumon Nozaki, Tatsuya Komatsu |
Interspeech | 2 |
| 2020 | Scene-Dependent Acoustic Event Detection with Scene Conditioning and Fake-Scene-Conditioned LossabstractIn this paper, we propose scene-dependent acoustic event detection (AED) with scene conditioning and fake-scene-conditioned loss. The proposed method employs a multitask network, that has not only AED part but also acoustic scene classification (ASC). The scenes predicted by ASC are employed as an additional feature for scene conditioning of AED to learn the relationship between scenes and events. For efficient training, the proposed method incorporates a new AED loss function, which is the fake-scene-conditioned loss, in addition to the conventional AED loss. Upon training, the AED part is conditioned with fake scenes as well as predicted and true scenes. The fake-scene-conditioned loss is calculated between the fake-scene-conditioned AED results and labels of events that do not exist in the fake scenes are removed. Whereas training with combinations of true scenes/events, i.e., the conventional AED loss, only reveals that an event is present in a scene, with fake-scene-conditioned loss, the proposed method can learn that an event is absent in a scene. Experimental results show that the proposed method improves the AED performance compared with the baseline; an increase in the f1 score of 23% and a decrease in the false alarm rate of 56% for scenes where no event exists. Tatsuya Komatsu, Keisuke Imoto, Masahito Togami |
ICASSP | 1 |
| 2020 | Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural NetworksabstractThis paper proposes a deep neural network (DNN)–based multichannel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often conducted in the time-frequency (T-F) domain because spatial filtering can be efficiently implemented in the T-F domain. In such a case, ordinary objective functions are computed on the estimated T-F mask or spectrogram. However, the estimated spectrogram is often inconsistent, and its amplitude and phase may change when the spectrogram is converted back to the time-domain. That is, the objective function does not evaluate the enhanced time-domain signal properly. To address this problem, we propose to use an objective function defined on the reconstructed time-domain signal. Specifically, speech enhancement is conducted by multi-channel Wiener filtering in the T-F domain, and its result is converted back to the time-domain. We propose two objective functions computed on the reconstructed signal where the first one is defined in the time-domain, and the other one is defined in the T-F domain. Our experiment demonstrates the effectiveness of the proposed system comparing to T-F masking and mask-based beamforming. Yoshiki Masuyama, Masahito Togami, Tatsuya Komatsu |
ICASSP | 3 |
| 2020 | Weakly-Supervised Sound Event Detection with Self-AttentionabstractIn this paper, we propose a novel sound event detection (SED) method that incorporates a self-attention mechanism of the Transformer for a weakly-supervised learning scenario. The proposed method utilizes the Transformer encoder, which consists of multiple self-attention modules, allowing to take both local and global context information of the input feature sequence into account. Furthermore, inspired by the great success of BERT in the natural language processing field, the proposed method introduces a special tag token into the input sequence for weak label prediction, which enables the aggregation of the whole sequence information. To demonstrate the performance of the proposed method, we conduct the experimental evaluation using the DCASE2019 Task4 dataset. The experimental results demonstrate that the proposed method outperforms the DCASE2019 Task4 baseline method, which is based on the convolutional recurrent neural network, and the self-attention mechanism effectively works for SED. Koichi Miyazaki, Tatsuya Komatsu, Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Kazuya Takeda |
ICASSP | 2 |
| 2020 | Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss FunctionabstractIn this paper, we propose a multi-channel speech source separation method with a deep neural network (DNN) which is trained under the condition that no clean signal is available. As an alternative to a clean signal, the proposed method adopts an estimated speech signal by an unsupervised speech source separation method which leverages a statistical model. As a statistical model of microphone input signal, we adopts a time-varying spatial covariance matrix (SCM) model which includes reverberation and background noise submodels so as to achieve robustness against reverberation and background noise. The DNN infers intermediate variables which are needed for constructing the time-varying SCM. Separation is performed in a probabilistic manner so as to avoid overfitting to separation error. Since there are multiple intermediate variables, a loss function which evaluates a single intermediate variable is not applicable. Instead, the proposed method adopts a loss function which evaluates the output probabilistic signal directly based on Kullback-Leibler Divergence (KLD). The gradient of the loss function can be back-propagated into the DNN through all the intermediate variables. Experimental results under reverberant conditions show that the proposed method achieves better results than the conventional methods. Masahito Togami, Yoshiki Masuyama, Tatsuya Komatsu, Yu Nakagome |
ICASSP | 3 |
| 2019 | Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vectorabstractThis paper proposes a scene-dependent anomalous acoustic-event detection based on conditional WaveNet and i-vector. The WaveNet builds normal acoustic event models by exhaustive learning of time-domain signals in the public space to provide scene-independent anomaly detection. I-vectors are used as additional features to describe acoustic scenes, where the input signals are observed, to complement the WaveNet. The proposed method can detect anomalous acoustic-events in environments whose acoustic scenes vary depending on time, location, and surrounding environment. Evaluations with data recorded from the real environment demonstrate that the proposed method achieved as much as 15 pt higher F-measure than LSTM and AE. The difference in F-measure by the WaveNet with and without i-vector turned out to be 1.5 pt. Tatsuya Komatsu, Tomoki Hayashi, Reishi Kondo, Tomoki Toda, Kazuya Takeda |
ICASSP | 1 |
| 2019 | Bayesian Non-parametric Multi-source Modelling Based Determined Blind Source SeparationabstractThis paper proposes a determined blind source separation method using Bayesian non-parametric modelling of sources. Conventionally source signals are separated from a given set of mixture signals by modelling them using non-negative matrix factorization (NMF). However in NMF, a latent variable signifying model complexity must be appropriately specified to avoid over-fitting or under-fitting. As real-world sources can be of varying and unknown complexities, we propose a Bayesian non-parametric framework which is invariant to such latent variables. We show that our proposed method adapts to different source complexities, while conventional methods require parameter tuning for optimal separation. Chaitanya Narisetty, Tatsuya Komatsu, Reishi Kondo |
ICASSP | 2 |
| 2019 | Multichannel Loss Function for Supervised Speech Source Separation by Mask-Based BeamformingabstractIn this paper, we propose two mask-based beamforming methods using a deep neural network (DNN) trained by multichannel loss functions. Beamforming technique using time-frequency (TF)-masks estimated by a DNN have been applied to many applications where TF-masks are used for estimating spatial covariance matrices. To train a DNN for mask-based beamforming, loss functions designed for monaural speech enhancement/separation have been employed. Although such a training criterion is simple, it does not directly correspond to the performance of mask-based beamforming. To overcome this problem, we use multichannel loss functions which evaluate the estimated spatial covariance matrices based on the multichannel Itakura--Saito divergence. DNNs trained by the multichannel loss functions can be applied to construct several beamformers. Experimental results confirmed their effectiveness and robustness to microphone configurations. Yoshiki Masuyama, Masahito Togami, Tatsuya Komatsu |
INTERSPEECH | 3 |
| 2019 | Variational Bayesian Multi-Channel Speech Dereverberation Under Noisy Environments with Probabilistic Convolutive Transfer Function
Masahito Togami, Tatsuya Komatsu |
INTERSPEECH | 2 |
| 2017 | Detection of anomaly acoustic scenes based on a temporal dissimilarity modelabstractThis paper proposes detection of anomaly acoustic scenes based on a temporal dissimilarity model. The periodicity in the temporal variation of acoustic scenes is first pointed out and then used to build a new stochastic model. In the new model, the temporal variation is expressed by dissimilarity between current and previous acoustic scenes. Anomaly acoustic scenes are detected based on the 24-hour periodic dissimilarity model. Evaluation results using 40-day (1000-hour) data show that the proposed method can detect unknown anomaly acoustic scenes with 82.3% F-measure in 0 dB signal-to-noise-ratio conditions. Tatsuya Komatsu, Reishi Kondo |
ICASSP | 1 |
| 2016 | Acoustic event detection based on non-negative matrix factorization with mixtures of local dictionaries and activation aggregationabstractThis paper proposes a new non-negative matrix factorization (NMF) based acoustic event detection (AED) method with mixtures of local dictionaries (MLD) and activation aggregation. One of the key problems of conventional NMF-based methods is instability of activations due to redundancy of a region spanned by the bases of dictionaries. Sounds inside the redundant region are often decomposed into undesired combinations of bases and activations that cause failure of detection. The proposed method employs MLD for allocating sub-groups of basis dictionaries to acoustic elements to minimize redundancy in the region and obtain controlled activations. In order to make activations more stable, the proposed method also introduces activation aggregation which combines basis-wise activations into acoustic-element-wise activations. Much more stable activations by the proposed method lead to significant improvement in F-measure by up to 60% compared to an ordinary convolutive-NMF-based method. The proposed method also outperforms a latest alternative which is not based on NMF. Tatsuya Komatsu, Yuzo Senda, Reishi Kondo |
ICASSP | 1 |
| 2013 | Modeling head-related transfer functions via spatial-temporal Gaussian processabstractWe propose a novel application of a family of non-parametric statistical models to estimate head-related transfer functions (HRTFs) using spatial-temporal Gaussian processes (GPs). In this approach, we model the head-related impulse response (HRIR) utilizing non-parametric regression via a GP. The challenge posed by this problem involves accurate modeling of the spatial correlation structure jointly with the temporal correlation structure at each spatial location for the HRIR. We solve this problem by constructing a joint spatial-temporal kernel characterizing the GP regression model. To perform inference, we estimate the hyper-parameters of the GP regression kernel via maximum signal-to-deviation-ratio estimation on the basis of a real experimental setup in which we collected observations of the HRIR using two head-and-torso simulators (HATSs): KEMAR and B&K. We also perform cross validation of the model by training on the KEMAR system and assessing the generalization of our model and its out-of-sample predictive power for HRIRs at any locations that we predict by the model assessed on the B&K system. The corresponding HRTFs are obtained as the Fourier transform of the HRIRs. In the experiments, we show that our method is robust against variation in the azimuth interval needed to perform high-accuracy interpolation and has the expressive power to handle the individual characteristics of each HATS. Tatsuya Komatsu, Takanori Nishino, Gareth W. Peters, Tomoko Matsui, Kazuya Takeda |
ICASSP | 1 |
| 2013 | Computationally efficient single channel dereverberation based on complementary wiener filterabstractA single-channel dereverberation method with low computational complexity is proposed. We introduce the complementary Wiener filter which can suppress a late reverberation during silence intervals via theoretical analysis and numerical calculation. An implementation represents reductions both of memory consumption and operative calculations compared to a conventional method; each reduction is almost by half. Dereverberation performance is evaluated by an experimental simulation using speech signals and measured room impulse responses. The performance under several hundreds of msec of the reverberation time is similar to the conventional method: 6 [dB] reverberation reduction and 3 [dB] improvement of target-to-interference ratio. Kazunobu Kondo, Yu Takahashi, Tatsuya Komatsu, Takanori Nishino, Kazuya Takeda |
ICASSP | 3 |