Shoichiro Saito

dblp:29/6412 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-8712-0464ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorComputer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event Analysis
abstract
This article addresses a new task: distributed multimedia sensor event analysis (DiMSEA). DiMSEA aims to analyze a series of human and machine activities (called “events” in this article) in complex and extensive real-world environments. Since an observation from a single sensor is often missing or fragmented in such an environment, observations from multiple locations and modalities should be integrated to analyze events comprehensively. However, a learning method has yet to be established to extract joint representations that effectively combine such distributed observations. Therefore, we propose guided masked self-distillation modeling (Guided-MELD) for inter-sensor relationship modeling. The basic idea of Guided-MELD is to learn to supplement the information from the masked sensor with information from other sensors needed to detect the event. Guided-MELD is expected to effectively distill fragmented target event information from sensors without over-relying on any specific sensors. To validate the effectiveness of the proposed method in DiMSEA, we recorded two new datasets: MM-Store and MM-Office. These datasets consist of human activities in a convenience store and an office, recorded using distributed cameras and microphones. Experimental results show that the proposed Guided-MELD improves event tagging and detection performance and outperforms conventional inter-sensor relationship modeling methods. Furthermore, the proposed method performed robustly even when sensors were reduced.
Masahiro Yasuda, Noboru Harada, Yasunori Ohishi, Shoichiro Saito, Akira Nakayama, Nobutaka Ono
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Sound Source Distance Estimation Utilizing Physics-informed Prior for Sound Event Localization and Detection
abstract
Sound Event Localization and Detection (SELD) is the combined task of detecting sound events and estimating their spatial locations. We propose a Sound source Distance Estimation (SDE) method for SELD that utilizes a physics-informed prior. The conventional data-driven approach of SDE for SELD can handle complex situations where sound sources move or overlap, thanks to multitask learning. Deep Neural Network (DNN)-based SELD systems are generally pre-trained with synthesized data to compensate for the lack of real data. However, in the context of recent SELD tasks, the performance of DNN models, which are pre-trained with synthetic data, significantly degrades on real data. One possible cause of this is that the DNN models overfit sound characteristics that do not exist in the real data, i.e., are not physically reasonable. Therefore, we propose a hybrid SDE method that utilizes a physics-informed prior in a data-driven SELD system. Focusing on the specificity of the expected sound PoWer Level (PWL) of the sound sources depending on the class, we set the typical PWL for each class as a prior. To improve real-world applicability, we also adopt the sound attenuation model as a physics-informed prior for explicitly utilizing physical laws in SDE. Experimental results suggested the effectiveness of utilizing a physics-informed prior in SDE for SELD to improve its applicability to the real world.
Nao Sato, Masahiro Yasuda, Shoichiro Saito, Noboru Harada
ICASSP3
2025 Spatial Annotation-free Training for Sound Event Localization and Detection
abstract
Sound Event Localization and Detection (SELD) is the task of estimating the class, duration, and direction of arrival (DOA) of sound events. State-of-the-art SELD systems use a data-driven approach based on Deep Neural Networks (DNNs) to deal with complex situations involving overlapping and moving sound sources. Such systems need to be trained on real-recording data to succeed in real-world situations. However, annotation of real-world data, especially spatial annotation, incurs huge costs, and the amount of labeled data is currently insufficiently small. Therefore, we are introducing spatial annotation-free training for SELD, which trains a SELD system using only sound class and duration labels. As a first attempt at this task, we propose Beam-based Multiple Instance Learning (Beam-MIL). Beam-MIL first instantiates acoustic signals for each DOA by beamforming. Then, the sound events contained in each instance are trained indirectly by MIL without using each instance’s ground truth information, i.e., spatial annotation. Experimental results show that Beam-MIL can effectively train a valid DOA estimator without spatial annotations. Moreover, when the small amount of the annotated data is available, enlarging data size by adding annotation-free data significantly improved the performance of the system.
Masahiro Yasuda, Shoichiro Saito, Nao Sato, Noboru Harada
ICASSP2
2024 Online Target Sound Extraction with Knowledge Distillation from Partially Non-Causal Teacher
abstract
Target Sound Extraction (TSE) is a technique for extracting sound events belonging to a target sound class in a mixture using a Deep Neural Network (DNN). Offline TSE that uses non-causal models has achieved high extraction performance. However, many applications require online processing. Simply converting the non-causal TSE model architecture to a causal one leads to significant performance degradation. To mitigate this problem, we propose using Knowledge Distillation (KD) from a non-causal teacher to a causal student for TSE. In particular, we investigate different options for the non-causal teacher. We identify that a causal network with a non-causal layer normalization provides a strong teacher from which it is easier to transfer knowledge to the student. We conduct experiments with simulated sound mixtures and show that training a causal TSE with the proposed KD scheme can improve the signal-to-distortion ratio (SDR) by 0.9 dB compared to a baseline causal system.
Keigo Wakayama, Tsubasa Ochiai, Marc Delcroix, Masahiro Yasuda, Shoichiro Saito, Shoko Araki, Akira Nakayama
ICASSP5
2024 6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on Self-Motioning Human
abstract
We aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with microphone arrays located in static positions. However, self-motion with three rotational and three translational degrees of freedom (6DoF) shall be considered for wearable microphone arrays. A system trained only with a dataset using microphone arrays in a fixed position would be unable to adapt to the fast relative motion of sound events associated with self-motion, resulting in the degradation of SELD performance. To address this, we designed 6DoF SELD Dataset1for wearable systems, the first SELD dataset considering the self-motion of microphones. Furthermore, we proposed a multi-modal SELD system that jointly utilizes audio and motion tracking sensor signals. These sensor signals are expected to help the system find useful acoustic cues for SELD on the basis of the current self-motion state. Experimental results on our dataset show that the proposed method effectively improves SELD performance with a mechanism to extract acoustic features conditioned by sensor signals.
Masahiro Yasuda, Shoichiro Saito, Akira Nakayama, Noboru Harada
ICASSP2
2022 Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around Head
abstract
Sound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with the sound signals observed by a microphone array. Although ambisonic microphones are popular in the literature of SELD, they might limits the range of applications due to their predetermined geometry. Some applications (including those for pedestrians that perform SELD while walking) require a wearable microphone array whose geometry can be designed to suit the task. In this paper, for development of such a wearable SELD, we propose a dataset named Wearable SELD dataset. It consists of data recorded by 24 microphones placed on a head and torso simulators (HATS) with some accessories mimicking wearable devices (glasses, earphones, and headphones). We also provide experimental results of SELD using the proposed dataset and SELDNet to investigate the effect of microphone configuration.
Kento Nagatomo, Masahiro Yasuda, Kohei Yatabe, Shoichiro Saito, Yasuhiro Oikawa
ICASSP4
2022 CNN-Transformer with Self-Attention Network for Sound Event Detection
abstract
In sound event detection (SED), the representation ability of deep neural network (DNN) models must be increased to significantly improve the accuracy or increase the number of classifiable classes. When building large-scale DNN models, a highly parameter-efficient DNN architecture should preferably be adopted. In image recognition, there has been a proposal to replace a convolutional neural network (CNN) extracting high-level features with a highly parameter-efficient DNN architecture, i.e., a self-attention network (SAN). The high-level features are essential information that contributes to prediction. In SED, we find that a model that exceeds the prediction accuracy of CNN-Transformer is difficult to build simply by replacing CNN with SAN, in the process of our experiments. To construct a model with high prediction accuracy while capturing the properties of acoustic signals well, we propose an architecture called a CNN-SAN-Transformer, which retains CNN in the blocks close to the input and uses SAN in all remaining blocks. Experimental results suggest that the proposed method has the same or higher prediction accuracy with a smaller number of parameters than the CNN-Transformer and higher prediction accuracy with a similar number of parameters to the CNN-Transformer and that the proposed method may be a parameter-efficient architecture.
Keigo Wakayama, Shoichiro Saito
ICASSP2
2022 Echo-Aware Adaptation of Sound Event Localization and Detection in Unknown Environments
abstract
Our goal is to develop a sound event localization and detection (SELD) system that works robustly in unknown environments. A SELD system trained on known environment data is degraded in an unknown environment due to environmental effects such as reverberation and noise not contained in the training data. Previous studies on related tasks have shown that domain adaptation methods are effective when data on the environment in which the system will be used is available even without labels. However adaptation to unknown environments remains a difficult task. In this study, we propose echo-aware feature refinement (EAR) for SELD, which suppresses environmental effects at the feature level by using additional spatial cues of the unknown environment obtained through measuring acoustic echoes. FOA-MEIR1, an impulse response dataset containing over 100 environments, was recorded to validate the proposed method. Experiments on FOA-MEIR show that the EAR effectively improves SELD performance in unknown environments.
Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito
ICASSP3
2022 Multi-View And Multi-Modal Event Detection Utilizing Transformer-Based Multi-Sensor Fusion
abstract
We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task, distributed sensors are utilized complementarily to capture events that are difficult to capture with a single sensor, such as a series of actions of people moving in an intricate room, or communication between people located far apart in a room. For sensors to cooperate effectively in such a situation, the system should be able to exchange information among sensors and combines information that is useful for identifying events in a complementary manner. For such a mechanism, we propose a Transformer-based multi-sensor fusion (MultiTrans) which combines multi-sensor data on the basis of the relationships between features of different viewpoints and modalities. In the experiments using a dataset1newly collected for this task, our proposed method using MultiTrans improved the event detection performance and outperformed comparatives.
Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, Noboru Harada
ICASSP3
2020 SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds
abstract
We propose a similarity function for one-shot anomaly detection in sounds (ADS) called SPecific anomaly IDentifiER network (SPIDERnet). In ADS systems, since overlooking an anomaly may result in serious incidents, we need to update such systems using an (often only one) overlooked anomalous sample. A previous study proposed the use of memory-based one-shot learning. A problem with this previous method is that it can detect only short anomalous sounds such as collision sounds because its similarity function is based on a naive mean-squared-error between the input and memorized spectrogram. To detect various anomalous sounds, SPIDERnet consists of (i) a neural network-based feature extractor for measuring similarity in embedded space and (ii) attention mechanisms for absorbing time-frequency stretching. Experimental results on two public datasets indicate that SPIDERnet outperforms conventional methods and robustly detects various anomalous sounds.
Yuma Koizumi, Masahiro Yasuda, Shin Murata, Shoichiro Saito, Hisashi Uematsu, Noboru Harada
ICASSP4
2020 Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source Separation
abstract
We propose a direction-of-arrival (DOA) estimation method for Sound Event Localization and Detection (SELD). Direct estimation of DOA using a deep neural network (DNN), i.e. completely-datadriven approach, achieves high accuracy. However, there is a gap in the accuracy between DOA estimation for single and overlapping sources because they cannot incorporate physical knowledge. Meanwhile, although the accuracy of physics-based approaches is inferior to DNN-based approaches, it is robust for overlapping-source. In this study, we consider a combination of physics-based and DNN-based approaches; the sound intensity vectors (IVs) for physics-based DOA estimation is refined based on DNN-based denoising and source separation. This method enables the accurate DOA estimation for both single and overlapping sources using a spherical microphone array. Experimental results show that the proposed method achieves state-of-the-art DOA estimation accuracy on an open dataset of the SELD.
Masahiro Yasuda, Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, Keisuke Imoto
ICASSP3
2020 A Transformer-Based Audio Captioning Model with Keyword Estimation
abstract
One of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene.Since one acoustic event/scene can be described with several words, it results in a combinatorial explosion of possible captions and difficulty in training.To solve this problem, we propose a Transformer-based audio-captioning model with keyword estimation called TRACKE.It simultaneously solves the word-selection indeterminacy problem with the main task of AAC while executing the sub-task of acoustic event detection/acoustic scene classification (i.e., keyword estimation).TRACKE estimates keywords, which comprise a word set corresponding to audio events/scenes in the input audio, and generates the caption while referring to the estimated keywords to reduce word-selection indeterminacy.Experimental results on a public AAC dataset indicate that TRACKE achieved state-ofthe-art performance and successfully estimated both the caption and its keywords.
Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda, Shoichiro Saito
INTERSPEECH5
2019 SNIPER: Few-shot Learning for Anomaly Detection to Minimize False-negative Rate with Ensured True-positive Rate
abstract
In anomaly detection systems, overlooking anomalies may result in serious incidents. Thus, when a system overlooks an anomaly, we need to update the system to never overlook the observed type of anomalies twice. There are roughly two possible approaches to solve this problem; re-training the whole system using all training data, or cascading a new specific detector for the overlooked anomaly. The first approach is the most effective solution; however, a huge computational cost and an amount of anomalous training data are required to re-train the system when it consists of a deep-learning-based anomaly detector. We focused on the latter approach and propose a training method for a cascaded specific anomaly detector using few-shot (just 1 to 3) samples. To suppress the false-negative rate of the overlooked anomaly, the proposed method works to decrease the false-positive rate under the constraint of true-positive rate equaling 1. Experimental results show that the proposed method outperformed conventional cross-entropy-based few-shot learning methods.
Yuma Koizumi, Shin Murata, Noboru Harada, Shoichiro Saito, Hisashi Uematsu
ICASSP4
2019 Unsupervised Detection of Anomalous Sound Based on Deep Learning and the Neyman-Pearson Lemma
abstract
This paper proposes a novel optimization principle and its implementation for unsupervised anomaly detection in sound (ADS) using an autoencoder (AE). The goal of the unsupervised-ADS is to detect unknown anomalous sounds without training data of anomalous sounds. The use of an AE as a normal model is a state-of-the-art technique for the unsupervised-ADS. To decrease the false positive rate (FPR), the AE is trained to minimize the reconstruction error of normal sounds, and the anomaly score is calculated as the reconstruction error of the observed sound. Unfortunately, since this training procedure does not take into account the anomaly score for anomalous sounds, the true positive rate (TPR) does not necessarily increase. In this study, we define an objective function based on the Neyman-Pearson lemma by considering the ADS as a statistical hypothesis test. The proposed objective function trains the AE to maximize the TPR under an arbitrary low FPR condition. To calculate the TPR in the objective function, we consider that the set of anomalous sounds is the complementary set of normal sounds and simulate anomalous sounds by using a rejection sampling algorithm. Through experiments using synthetic data, we found that the proposed method improved the performance measures of the ADS under low FPR conditions. In addition, we confirmed that the proposed method could detect anomalous sounds in real environments.
Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, Yuta Kawachi, Noboru Harada
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Software defined media: Virtualization of audio-visual services
abstract
Internet-native audio-visual services are witnessing rapid development. Among these services, object-based audiovisual services are gaining importance. In 2014, we established the Software Defined Media (sDM) consortium to target new research areas and markets involving object-based digital media and Internet-by-design audio-visual environments. In this paper, we introduce the SDM architecture that virtualizes networked audio-visual services along with the development of smart buildings and smart cities using Internet of Things (IoT) devices and smart building facilities. Moreover, we design the SDM architecture as a layered architecture to promote the development of innovative applications on the basis of rapid advancements in software-defined networking (SDN). Then, we implement a prototype system based on the architecture, present the system at an exhibition, and provide it as an SDM API to application developers at hackathons. Various types of applications are developed using the API at these events. An evaluation of SDM API access shows that the prototype SDM platform effectively provides 3D audio reproducibility and interactiveness for SDM applications.
Manabu Tsukada, Keiko Ogawa, Masahiro Ikeda, Takuro Sone, Kenta Niwa, Shoichiro Saito, Takashi Kasuya, Hideki Sunahara, Hiroshi Esaki
ICC6
2008 Specmurt Analysis of Polyphonic Music Signals
abstract
This paper introduces a new music signal processing method to extract multiple fundamental frequencies, which we call specmurt analysis. In contrast with cepstrum which is the inverse Fourier transform of log-scaled power spectrum with linear frequency, specmurt is defined as the inverse Fourier transform of linear power spectrum with log-scaled frequency. Assuming that all tones in a polyphonic sound have a common harmonic pattern, the sound spectrum can be regarded as a sum of linearly stretched common harmonic structures along frequency. In the log-frequency domain, it is formulated as the convolution of a common harmonic structure and the distribution density of the fundamental frequencies of multiple tones. The fundamental frequency distribution can be found by deconvolving the observed spectrum with the assumed common harmonic structure, where the common harmonic structure is given heuristically or quasi-optimized with an iterative algorithm. The efficiency of specmurt analysis is experimentally demonstrated through generation of a piano-roll-like display from a polyphonic music signal and automatic sound-to-MIDI conversion. Multipitch estimation accuracy is evaluated over several polyphonic music signals and compared with manually annotated MIDI data.
Shoichiro Saito, Hirokazu Kameoka, Keigo Takahashi, Takuya Nishimoto, Shigeki Sagayama
IEEE Trans. Speech Audio Process.1