EDBT 2026 Demo / reviewers in the wild / expert
Masahiro Yasuda
dblp:117/6601
· DBLP profile ↗
21ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event AnalysisabstractThis article addresses a new task: distributed multimedia sensor event analysis (DiMSEA). DiMSEA aims to analyze a series of human and machine activities (called “events” in this article) in complex and extensive real-world environments. Since an observation from a single sensor is often missing or fragmented in such an environment, observations from multiple locations and modalities should be integrated to analyze events comprehensively. However, a learning method has yet to be established to extract joint representations that effectively combine such distributed observations. Therefore, we propose guided masked self-distillation modeling (Guided-MELD) for inter-sensor relationship modeling. The basic idea of Guided-MELD is to learn to supplement the information from the masked sensor with information from other sensors needed to detect the event. Guided-MELD is expected to effectively distill fragmented target event information from sensors without over-relying on any specific sensors. To validate the effectiveness of the proposed method in DiMSEA, we recorded two new datasets: MM-Store and MM-Office. These datasets consist of human activities in a convenience store and an office, recorded using distributed cameras and microphones. Experimental results show that the proposed Guided-MELD improves event tagging and detection performance and outperforms conventional inter-sensor relationship modeling methods. Furthermore, the proposed method performed robustly even when sensors were reduced. Masahiro Yasuda, Noboru Harada, Yasunori Ohishi, Shoichiro Saito, Akira Nakayama, Nobutaka Ono |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Collision-less and Balanced Sampling for Language-Queried Audio Source SeparationabstractLanguage-queried audio source separation (LASS) is an emerging research field that has recently received increasing attention. This task aims to isolate individual sources from a mixture of signals using natural language descriptions, enabling applications in various areas such as automatic audio editing. While conventional methods focus on the system architecture, the important aspect of data processing has been overlooked. The data for training LASS are typically created by mixing various audio signals in the dataset to form a mixture. One signal is then used as the target, whereas the others are regarded as interference. However, sound events in the target signal could overlap with those in the interference signals, which may cause confusion that instructs the model to both retain and suppress the same sound events within a single training example. In addition, training LASS with large-scale datasets may suffer from the data imbalance problem, where some sound events appear too frequently while others are rare. In this paper, we address these problems by using data sampling techniques. Specifically, the interference signals are sampled so that their audio tags do not conflict with those of the target signal, where the tags are generated using an audio tagging model. To balance the data, we consider several balanced sampling approaches using tag or caption embedding. By leveraging their distribution information, we use either weighted or group sampling to boost the occurrence of underrepresented samples while reducing the presence of overrepresented ones. Experimental results show the superiority of the proposed method over state-of-the-art LASS systems in DCASE 2024 Challenge Task 9. Pre-trained model is available at: https://github.com/tucothien/LASS-CLBS. Binh Thien Nguyen, Daiki Takeuchi, Masahiro Yasuda, Daisuke Niizumi, Noboru Harada |
ICASSP | 3 |
| 2025 | Sound Source Distance Estimation Utilizing Physics-informed Prior for Sound Event Localization and DetectionabstractSound Event Localization and Detection (SELD) is the combined task of detecting sound events and estimating their spatial locations. We propose a Sound source Distance Estimation (SDE) method for SELD that utilizes a physics-informed prior. The conventional data-driven approach of SDE for SELD can handle complex situations where sound sources move or overlap, thanks to multitask learning. Deep Neural Network (DNN)-based SELD systems are generally pre-trained with synthesized data to compensate for the lack of real data. However, in the context of recent SELD tasks, the performance of DNN models, which are pre-trained with synthetic data, significantly degrades on real data. One possible cause of this is that the DNN models overfit sound characteristics that do not exist in the real data, i.e., are not physically reasonable. Therefore, we propose a hybrid SDE method that utilizes a physics-informed prior in a data-driven SELD system. Focusing on the specificity of the expected sound PoWer Level (PWL) of the sound sources depending on the class, we set the typical PWL for each class as a prior. To improve real-world applicability, we also adopt the sound attenuation model as a physics-informed prior for explicitly utilizing physical laws in SDE. Experimental results suggested the effectiveness of utilizing a physics-informed prior in SDE for SELD to improve its applicability to the real world. Nao Sato, Masahiro Yasuda, Shoichiro Saito, Noboru Harada |
ICASSP | 2 |
| 2025 | Spatial Annotation-free Training for Sound Event Localization and DetectionabstractSound Event Localization and Detection (SELD) is the task of estimating the class, duration, and direction of arrival (DOA) of sound events. State-of-the-art SELD systems use a data-driven approach based on Deep Neural Networks (DNNs) to deal with complex situations involving overlapping and moving sound sources. Such systems need to be trained on real-recording data to succeed in real-world situations. However, annotation of real-world data, especially spatial annotation, incurs huge costs, and the amount of labeled data is currently insufficiently small. Therefore, we are introducing spatial annotation-free training for SELD, which trains a SELD system using only sound class and duration labels. As a first attempt at this task, we propose Beam-based Multiple Instance Learning (Beam-MIL). Beam-MIL first instantiates acoustic signals for each DOA by beamforming. Then, the sound events contained in each instance are trained indirectly by MIL without using each instance’s ground truth information, i.e., spatial annotation. Experimental results show that Beam-MIL can effectively train a valid DOA estimator without spatial annotations. Moreover, when the small amount of the annotated data is available, enlarging data size by adding annotation-free data significantly improved the performance of the system. Masahiro Yasuda, Shoichiro Saito, Nao Sato, Noboru Harada |
ICASSP | 1 |
| 2025 | Towards Pre-training an Effective Respiratory Audio Foundation Model
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen, Yasunori Ohishi, Noboru Harada |
INTERSPEECH | 3 |
| 2025 | CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
Daiki Takeuchi, Binh Thien Nguyen, Masahiro Yasuda, Yasunori Ohishi, Daisuke Niizumi, Noboru Harada |
INTERSPEECH | 3 |
| 2025 | Real-time TSE demonstration via SoundBeam with KD
Keigo Wakayama, Tomoko Kawase, Takafumi Moriya, Marc Delcroix, Hiroshi Sato 0002, Tsubasa Ochiai, Masahiro Yasuda, Shoko Araki |
INTERSPEECH | 7 |
| 2024 | Online Target Sound Extraction with Knowledge Distillation from Partially Non-Causal TeacherabstractTarget Sound Extraction (TSE) is a technique for extracting sound events belonging to a target sound class in a mixture using a Deep Neural Network (DNN). Offline TSE that uses non-causal models has achieved high extraction performance. However, many applications require online processing. Simply converting the non-causal TSE model architecture to a causal one leads to significant performance degradation. To mitigate this problem, we propose using Knowledge Distillation (KD) from a non-causal teacher to a causal student for TSE. In particular, we investigate different options for the non-causal teacher. We identify that a causal network with a non-causal layer normalization provides a strong teacher from which it is easier to transfer knowledge to the student. We conduct experiments with simulated sound mixtures and show that training a causal TSE with the proposed KD scheme can improve the signal-to-distortion ratio (SDR) by 0.9 dB compared to a baseline causal system. Keigo Wakayama, Tsubasa Ochiai, Marc Delcroix, Masahiro Yasuda, Shoichiro Saito, Shoko Araki, Akira Nakayama |
ICASSP | 4 |
| 2024 | 6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on Self-Motioning HumanabstractWe aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with microphone arrays located in static positions. However, self-motion with three rotational and three translational degrees of freedom (6DoF) shall be considered for wearable microphone arrays. A system trained only with a dataset using microphone arrays in a fixed position would be unable to adapt to the fast relative motion of sound events associated with self-motion, resulting in the degradation of SELD performance. To address this, we designed 6DoF SELD Dataset1for wearable systems, the first SELD dataset considering the self-motion of microphones. Furthermore, we proposed a multi-modal SELD system that jointly utilizes audio and motion tracking sensor signals. These sensor signals are expected to help the system find useful acoustic cues for SELD on the basis of the current self-motion state. Experimental results on our dataset show that the proposed method effectively improves SELD performance with a mechanism to extract acoustic features conditioned by sensor signals. Masahiro Yasuda, Shoichiro Saito, Akira Nakayama, Noboru Harada |
ICASSP | 1 |
| 2024 | M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Masahiro Yasuda, Shunsuke Tsubaki, Keisuke Imoto |
INTERSPEECH | 5 |
| 2022 | Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around HeadabstractSound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although ambisonic microphones are popular in the literature of SELD, they might limits the range of applications due to their predetermined geometry. Some applications (including those for pedestrians that perform SELD while walking) require a wearable microphone array whose geometry can be designed to suit the task. In this paper, for development of such a wearable SELD, we propose a dataset named Wearable SELD dataset. It consists of data recorded by 24 microphones placed on a head and torso simulators (HATS) with some accessories mimicking wearable devices (glasses, earphones, and headphones). We also provide experimental results of SELD using the proposed dataset and SELDNet to investigate the effect of microphone configuration. Kento Nagatomo, Masahiro Yasuda, Kohei Yatabe, Shoichiro Saito, Yasuhiro Oikawa |
ICASSP | 2 |
| 2022 | APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse OptimizationabstractIn this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been developed upon constrained optimization, they are adjustable and well-studied in theory. However, they always uniformly promote sparsity and ignore the individual properties of a signal. Deep neural network (DNN)– based methods can learn the properties of target signals and use them for audio declipping. Still, they cannot perform well if the training data have mismatches and/or constraints in the time domain are not imposed. In the proposed method, we use a DNN in an optimization algorithm. It is inspired by an idea called plug-and-play (PnP) and enables us to promote sparsity based on the learned information of data, considering constraints in the time domain. Our experiments confirmed that the proposed method is stable and robust to mismatches between training and test data. Tomoro Tanaka, Kohei Yatabe, Masahiro Yasuda, Yasuhiro Oikawa |
ICASSP | 3 |
| 2022 | Echo-Aware Adaptation of Sound Event Localization and Detection in Unknown EnvironmentsabstractOur goal is to develop a sound event localization and detection (SELD) system that works robustly in unknown environments. A SELD system trained on known environment data is degraded in an unknown environment due to environmental effects such as reverberation and noise not contained in the training data. Previous studies on related tasks have shown that domain adaptation methods are effective when data on the environment in which the system will be used is available even without labels. However adaptation to unknown environments remains a difficult task. In this study, we propose echo-aware feature refinement (EAR) for SELD, which suppresses environmental effects at the feature level by using additional spatial cues of the unknown environment obtained through measuring acoustic echoes. FOA-MEIR1, an impulse response dataset containing over 100 environments, was recorded to validate the proposed method. Experiments on FOA-MEIR show that the EAR effectively improves SELD performance in unknown environments. Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito |
ICASSP | 1 |
| 2022 | Multi-View And Multi-Modal Event Detection Utilizing Transformer-Based Multi-Sensor FusionabstractWe tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task, distributed sensors are utilized complementarily to capture events that are difficult to capture with a single sensor, such as a series of actions of people moving in an intricate room, or communication between people located far apart in a room. For sensors to cooperate effectively in such a situation, the system should be able to exchange information among sensors and combines information that is useful for identifying events in a complementary manner. For such a mechanism, we propose a Transformer-based multi-sensor fusion (MultiTrans) which combines multi-sensor data on the basis of the relationships between features of different viewpoints and modalities. In the experiments using a dataset1newly collected for this task, our proposed method using MultiTrans improved the event detection performance and outperformed comparatives. Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, Noboru Harada |
ICASSP | 1 |
| 2020 | Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene LabelsabstractSound event detection (SED) and acoustic scene classification (ASC) are major tasks in environmental sound analysis. Considering that sound events and scenes are closely related to each other, some works have addressed joint analyses of sound events and acoustic scenes based on multitask learning (MTL), in which the knowledge of sound events and scenes can help in estimating them mutually. The conventional MTL-based methods utilize one-hot scene labels to train the relationship between sound events and scenes; thus, the conventional methods cannot model the extent to which sound events and scenes are related. However, in the real environment, common sound events may occur in some acoustic scenes; on the other hand, some sound events occur only in a limited acoustic scene. In this paper, we thus propose a new method for SED based on MTL of SED and ASC using the soft labels of acoustic scenes, which enable us to model the extent to which sound events and scenes are related. Experiments conducted using TUT Sound Events 2016/2017 and TUT Acoustic Scenes 2016 datasets show that the proposed method improves the SED performance by 3.80% in F-score compared with conventional MTL-based SED. Keisuke Imoto, Noriyuki Tonami, Yuma Koizumi, Masahiro Yasuda, Ryosuke Yamanishi, Yoichi Yamashita |
ICASSP | 4 |
| 2020 | SPIDERnet: Attention Network For One-Shot Anomaly Detection In SoundsabstractWe propose a similarity function for one-shot anomaly detection in sounds (ADS) called SPecific anomaly IDentifiER network (SPIDERnet). In ADS systems, since overlooking an anomaly may result in serious incidents, we need to update such systems using an (often only one) overlooked anomalous sample. A previous study proposed the use of memory-based one-shot learning. A problem with this previous method is that it can detect only short anomalous sounds such as collision sounds because its similarity function is based on a naive mean-squared-error between the input and memorized spectrogram. To detect various anomalous sounds, SPIDERnet consists of (i) a neural network-based feature extractor for measuring similarity in embedded space and (ii) attention mechanisms for absorbing time-frequency stretching. Experimental results on two public datasets indicate that SPIDERnet outperforms conventional methods and robustly detects various anomalous sounds. Yuma Koizumi, Masahiro Yasuda, Shin Murata, Shoichiro Saito, Hisashi Uematsu, Noboru Harada |
ICASSP | 2 |
| 2020 | Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source SeparationabstractWe propose a direction-of-arrival (DOA) estimation method for Sound Event Localization and Detection (SELD). Direct estimation of DOA using a deep neural network (DNN), i.e. completely-datadriven approach, achieves high accuracy. However, there is a gap in the accuracy between DOA estimation for single and overlapping sources because they cannot incorporate physical knowledge. Meanwhile, although the accuracy of physics-based approaches is inferior to DNN-based approaches, it is robust for overlapping-source. In this study, we consider a combination of physics-based and DNN-based approaches; the sound intensity vectors (IVs) for physics-based DOA estimation is refined based on DNN-based denoising and source separation. This method enables the accurate DOA estimation for both single and overlapping sources using a spherical microphone array. Experimental results show that the proposed method achieves state-of-the-art DOA estimation accuracy on an open dataset of the SELD. Masahiro Yasuda, Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, Keisuke Imoto |
ICASSP | 1 |
| 2020 | A Transformer-Based Audio Captioning Model with Keyword EstimationabstractOne of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene.Since one acoustic event/scene can be described with several words, it results in a combinatorial explosion of possible captions and difficulty in training.To solve this problem, we propose a Transformer-based audio-captioning model with keyword estimation called TRACKE.It simultaneously solves the word-selection indeterminacy problem with the main task of AAC while executing the sub-task of acoustic event detection/acoustic scene classification (i.e., keyword estimation).TRACKE estimates keywords, which comprise a word set corresponding to audio events/scenes in the input audio, and generates the caption while referring to the estimated keywords to reduce word-selection indeterminacy.Experimental results on a public AAC dataset indicate that TRACKE achieved state-ofthe-art performance and successfully estimated both the caption and its keywords. Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda, Shoichiro Saito |
INTERSPEECH | 4 |
| 2020 | Crossmodal Sound Retrieval Based on Specific Target Co-Occurrence Denoted with Weak Labels
Masahiro Yasuda, Yasunori Ohishi, Yuma Koizumi, Noboru Harada |
INTERSPEECH | 1 |
| 2017 | Analysis of the changes in listening trends of a music streaming serviceabstractPeople frequently change their music listening behaviors to fit their mood and particular situation. Such changes can be interpreted at various scales, e.g., the genres, artists, or specific tracks. Users of music streaming services expect to discover new music. However, discovering new music is not always a pleasant experience. Herein, we show that small changes in listening trends are good for users of such services. In contrast, large changes are bad. We modeled user listening trends using a hidden Markov model, which was applied hierarchically to analyze the user trends at multiple scales. We evaluated the changes in user listening trends to find good changes. Additionally, we analyzed the relationships between user listening trends and user lifestyle. Masanori Takano, Hiroki Mizukami, Fujio Toriumi, Makoto Takeuchi, Kazuya Wada, Masahiro Yasuda, Ichiro Fukiida |
IEEE BigData | 6 |
| 2016 | Relationships between EEGs and eye movements in response to facial expressionsabstractTo determine the relationship between brain activity and eye movements when activated by images of facial expressions, electroencephalograms (EEGs) and eye movements based on electrooculograms (EOGs) were measured and analyzed. Typical facial expressions from a photo database were grouped into two clusters by subjective evaluation and designated as either "Pleasant" or "Unpleasant" facial images. Regarding chronological analysis, the correlation coefficients of frequency powers between EEGs at a central area and eye movements monotonically increased throughout the time course when "Unpleasant" images were presented. Both the definite relationships and these dependencies on images of facial expressions were confirmed. Minoru Nakayama, Masahiro Yasuda |
ETRA | 2 |