EDBT 2026 Demo / reviewers in the wild / expert
Keitaro Tanaka
dblp:294/5996
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0005-4338-5224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Formula-Supervised Sound Event Detection: Pre-Training Without Real DataabstractIn this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effectiveness for sound event detection (SED). The SED task, which involves estimating the types and timings of sound events, is particularly challenged by the difficulty of acquiring a sufficient quantity of accurately labeled training data. Moreover, it is well known that manually annotated labels often contain noises and are significantly influenced by the subjective judgment of annotators. To address these challenges, we propose a novel pretraining method that utilizes a synthetic dataset, Formula-SED, where acoustic data are generated solely based on mathematical formulas. The proposed method enables large-scale pre-training by using the synthesis parameters applied at each time step as ground truth labels, thereby eliminating label noise and bias. We demonstrate that large-scale pre-training with Formula-SED significantly enhances model accuracy and accelerates training, as evidenced by our results in the DESED dataset used for DCASE2023 Challenge Task 4. The project page is at https://yutoshibata07.github.io/Formula-SED/. Yuto Shibata, Keitaro Tanaka, Yoshiaki Bando, Keisuke Imoto, Hirokatsu Kataoka, Yoshimitsu Aoki |
ICASSP | 2 |
| 2025 | Cross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech Recognition
Shunsuke Mitsumori, Sara Kashiwagi, Keitaro Tanaka, Shigeo Morishima |
INTERSPEECH | 3 |
| 2025 | Training Onset-and-Offset-Aware Sound Event Detection on a Heterogeneous Dataset via Probabilistic Sequential Modeling
Tomoya Yoshinaga, Yoshiaki Bando, Keitaro Tanaka, Keisuke Imoto, Masaki Onishi, Shigeo Morishima |
INTERSPEECH | 3 |
| 2025 | SyncViolinist: Music-Oriented Violin Motion Generation Based on Bowing and FingeringabstractAutomatically generating realistic musical performance motion can greatly enhance digital media production, often involving collaboration between professionals and musicians. However, capturing the intricate body, hand, and finger movements required for accurate musical performances is challenging. Existing methods often fall short due to the complex mapping between audio and motion, typically requiring additional inputs like scores or MIDI data. In this work, we present SyncViolinist, a multi-stage end-to-end framework that generates synchronized violin performance motion solely from audio input. Our method overcomes the challenge of capturing both global and finegrained performance features through two key modules: a bowing/fingering module and a motion generation module. The bowing/fingering module extracts detailed playing information from the audio, which the motion generation module uses to create precise, coordinated body motions reflecting the temporal granularity and nature of the violin performance. We demonstrate the effectiveness of SyncViolinist with significantly improved qualitative and quantitative results from unseen violin performance audio, outperforming state-of-the-art methods. Extensive subjective evaluations involving professional violinists further validate our approach. The code and dataset are available at https://github.com/Kakanat/SyncViolinist. Hiroki Nishizawa, Keitaro Tanaka, Asuka Hirata, Shugo Yamaguchi, Masatoshi Hamanaka, Shigeo Morishima |
WACV | 2 |
| 2025 | Onset-and-Offset-Aware Sound Event Detection via Differentiable Frame-to-Event MappingabstractThis paper presents a sound event detection (SED) method that handles sound event boundaries in a statistically principled manner. A typical approach to SED is to train a deep neural network (DNN) in a supervised manner such that the model predicts frame-wise event activities. Since the predicted activities often contain fine insertion and deletion errors due to their temporal fluctuations, post-processing has been applied to obtain more accurate onset and offset boundaries. Existing post-processing methods are, however, non-differentiable and prohibit end-to-end (E2E) training. In this paper, we propose an E2E detection method based on a probabilistic formulation of sound event sequences called a hidden semi-Markov model (HSMM). The HSMM is utilized to transform frame-wise features predicted by a DNN into posterior probabilities of sound events represented by their class labels and temporal boundaries. We jointly train the DNN and HSMM in a supervised E2E manner by maximizing the event-wise posterior probabilities of the HSMM. This objective is a differentiable function thanks to the forward-backward algorithm of the HSMM. Experimental results with real recordings show that our method outperforms baseline systems with standard post-processing methods. Tomoya Yoshinaga, Keitaro Tanaka, Yoshiaki Bando, Keisuke Imoto, Shigeo Morishima |
IEEE Signal Process. Lett. | 2 |
| 2023 | Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric LearningabstractThis paper presents a novel metric learning approach to address the performance gap between normal and silent speech in visual speech recognition (VSR).The difference in lip movements between the two poses a challenge for existing VSR models, which exhibit degraded accuracy when applied to silent speech.To solve this issue and tackle the scarcity of training data for silent speech, we propose to leverage the shared literal content between normal and silent speech and present a metric learning approach based on visemes.Specifically, we aim to map the input of two speech types close to each other in a latent space if they have similar viseme representations.By minimizing the Kullback-Leibler divergence of the predicted viseme probability distributions between and within the two speech types, our model effectively learns and predicts viseme identities.Our evaluation demonstrates that our method improves the accuracy of silent VSR, even when limited training data is available. Sara Kashiwagi, Keitaro Tanaka, Shigeo Morishima |
INTERSPEECH | 2 |
| 2021 | Pitch-Timbre Disentanglement Of Musical Instrument Sounds Based On Vae-Based Metric LearningabstractThis paper describes a representation learning method for disentangling an arbitrary musical instrument sound into latent pitch and timbre representations. Although such pitch-timbre disentanglement has been achieved with a variational autoencoder (VAE), especially for a predefined set of musical instruments, the latent pitch and timbre representations are outspread, making them hard to interpret. To mitigate this problem, we introduce a metric learning technique into a VAE with latent pitch and timbre spaces so that similar (different) pitches or timbres are mapped close to (far from) each other. Specifically, our VAE is trained with additional contrastive losses so that the latent distances between two arbitrary sounds of the same pitch or timbre are minimized, and those of different pitches or timbres are maximized. This training is performed under weak supervision that uses only whether the pitches and timbres of two sounds are the same or not, instead of their actual values. This improves the generalization capability for unseen musical instruments. Experimental results show that the proposed method can find better-structured disentangled representations with pitch and timbre clusters even for unseen musical instruments. Keitaro Tanaka, Ryo Nishikimi, Yoshiaki Bando, Kazuyoshi Yoshii, Shigeo Morishima |
ICASSP | 1 |
| 2021 | Manifold-Aware Deep Clustering: Maximizing Angles Between Embedding Vectors Based on Regular SimplexabstractThis paper presents a new deep clustering (DC) method called manifold-aware DC (M-DC) that can enhance hyperspace utilization more effectively than the original DC.The original DC has a limitation in that a pair of two speakers has to be embedded having an orthogonal relationship due to its use of the one-hot vector-based loss function, while our method derives a unique loss function aimed at maximizing the target angle in the hyperspace based on the nature of a regular simplex.Our proposed loss imposes a higher penalty than the original DC when the speaker is assigned incorrectly.The change from DC to M-DC can be easily achieved by rewriting just one term in the loss function of DC, without any other modifications to the network architecture or model parameters.As such, our method has high practicability because it does not affect the original inference part.The experimental results show that the proposed method improves the performances of the original DC and its expansion method. Keitaro Tanaka, Ryosuke Sawata, Shusuke Takahashi |
Interspeech | 1 |