EDBT 2026 Demo / reviewers in the wild / expert
Hokuto Munakata
dblp:313/1381
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-3624-5838ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Aligned Contrastive Learning for Text-to-Music RetrievalabstractThis paper proposes aligned contrastive learning for text-to-music retrieval. The proposed method introduces a new similarity measure, 'aligned similarity', which captures the frame-level and token-level correspondence within text and audio sequences. Unlike traditional approaches that aggregate sequence into clip-level and sentence-level embeddings, our method aligns the text token exhibiting the highest cosine similarity with each temporal frame of the audio sequence and averages these maximum similarity values across the entire sequence. This approach enables the capture of fine-grained relationships between audio and text that are often overlooked when sequences are aggregated into a single embedding. Retrieval experiments show significant performance improvements, with a notable gain being a 17.8% increase in Recall@5. Moreover, the alignment elucidates how specific audio frames correlate with textual tokens, enhancing the model's transparency and interpretability. Tatsuya Komatsu, Hokuto Munakata, Takuya Hasumi, Yusuke Fujita |
ICASSP | 2 |
| 2025 | Language-based Audio Moment RetrievalabstractIn this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict relevant moments in untrimmed long audio based on a text query. Given the lack of prior work in AMR, we first build a dedicated dataset, Clotho-Moment, consisting of large-scale simulated audio recordings with moment annotations. We then propose a Detection Transformer-based model, named Audio Moment DETR (AM-DETR), as a fundamental framework for AMR tasks. This model captures temporal dependencies within audio features, inspired by similar video moment retrieval tasks, thus surpassing conventional clip-level audio retrieval methods. Additionally, we provide manually annotated datasets to properly measure the effectiveness and robustness of our methods on real data. Experimental results show that AM-DETR, trained with Clotho-Moment, outperforms a baseline model that applies a clip-level audio retrieval method with a sliding window on all metrics, particularly improving [email protected] by 9.00 points. Our datasets and code are publicly available in https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval. Hokuto Munakata, Taichi Nishimura, Shota Nakada, Tatsuya Komatsu |
ICASSP | 1 |
| 2025 | DETECLAP: Enhancing Audio-Visual Representation Learning with Object InformationabstractCurrent audio-visual representation learning can capture rough object categories (e.g., "animals" and "instruments"), but it lacks the ability to recognize fine-grained details, such as specific categories like "dogs" and "flutes" within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification. Shota Nakada, Taichi Nishimura, Hokuto Munakata, Masayoshi Kondo, Tatsuya Komatsu |
ICASSP | 3 |
| 2025 | Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Yuchi Ishikawa, Shota Nakada, Hokuto Munakata, Kazuhiro Saito, Tatsuya Komatsu, Yoshimitsu Aoki |
INTERSPEECH | 3 |
| 2025 | Leveraging Unlabeled Audio for Audio-Text Contrastive Learning via Audio-Composed Text Features
Tatsuya Komatsu, Hokuto Munakata, Yuchi Ishikawa |
INTERSPEECH | 2 |
| 2024 | Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis FrameworkabstractWe propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization.Our proposed model converts song data with choral singing commonly contained in popular music and unsuitable for generating a simulated dataset to the solo singing data.This cleansing is based on NANSY++, which is a framework trained to reconstruct an input non-overlapped audio signal.We exploit the pretrained NANSY++ to convert choral singing into clean, nonoverlapped audio.This cleansing process mitigates the mislabeling of choral singing to solo singing and helps the effective training of EEND models even when the majority of available song data contains choral singing sections.We experimentally evaluated the EEND model trained with a dataset using our proposed method using annotated popular duet songs.As a result, our proposed method improved 14.8 points in diarization error rate. Hokuto Munakata, Ryo Terashima, Yusuke Fujita |
INTERSPEECH | 1 |
| 2023 | Recursive Sound Source Separation with Deep Learning-based Beamforming for Unknown Number of Sources
Hokuto Munakata, Ryu Takeda, Kazunori Komatani |
INTERSPEECH | 1 |
| 2023 | Joint Separation and Localization of Moving Sound Sources Based on Neural Full-Rank Spatial Covariance AnalysisabstractThis paper presents an unsupervised multichannel method that can separate moving sound sources based on an amortized variational inference (AVI) of joint separation and localization. A recently proposed blind source separation (BSS) method called neural full-rank spatial covariance analysis (FCA) trains a neural separation model based on a nonlinear generative model of multichannel mixtures and can precisely separate unseen mixture signals. This method, however, assumes that the sound sources hardly move, and thus its performance is easily degraded by the source movements. In this paper, we solve this problem by introducing time-varying spatial covariance matrices and directions of arrival of sources into the nonlinear generative model of the neural FCA. This generative model is used for training a neural network to jointly separate and localize moving sources by using only multichannel mixture signals and array geometries. The training objective is derived as a lower bound on the log-marginal posterior probability in the framework of AVI. Experimental results obtained with mixture signals of moving sources show that our method outperformed an existing joint separation and localization method and standard BSS methods. Hokuto Munakata, Yoshiaki Bando, Ryu Takeda, Kazunori Komatani, Masaki Onishi |
IEEE Signal Process. Lett. | 1 |
| 2022 | Training Data Generation with DOA-based Selecting and Remixing for Unsupervised Training of Deep Separation Models
Hokuto Munakata, Ryu Takeda, Kazunori Komatani |
INTERSPEECH | 1 |