Satoshi Tamura

dblp:32/4043 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Few-Shot Anomalous Sound Detection Based on Anomaly Map Estimation Using Pseudo Abnormal Data
abstract
This paper proposes a novel anomalous sound detection based on anomaly map estimation. Different from conventional autoencoder-based schemes, given a log-mel spectrogram image, our proposed method predicts an anomaly score map that indicates an anomalous part in the image. We generate pseudo abnormal data from normal data by CutPaste, then exploit them for model training. In real-world environments, the small number of real anomalies may be often available. Therefore, we also employ a few-shot learning architecture, so that not only normal data but also anomalies can be used for model training. Experiments were conducted to clarify the effectiveness of our methods, using a common machine-sound dataset. It is shown that the proposed methods outperform baseline methods. We also compared our method with few-shot learning to the state-of-the-art method. It is finally found that the proposed scheme has significant effectiveness in anomalous sound detection.
Ryosuke Tanaka, Satoshi Tamura
ICASSP2
2024 Speech Recognition for Indigenous Language Using Self-Supervised Learning and Natural Language Processing
Satoshi Tamura, Tomohiro Hattori, Yusuke Kato, Naoki Noguchi
ICPRAM1
2023 Speech Recognition for Minority Languages Using HuBERT and Model Adaptation
Tomohiro Hattori, Satoshi Tamura
ICPRAM2
2022 Efficient Multi-angle Audio-visual Speech Recognition using Parallel WaveGAN based Scene Classifier
Shinnosuke Isobe, Satoshi Tamura, Yuuto Gotoh, Masaki Nose
ICPRAM2
2022 Visual-only Voice Activity Detection using Human Motion in Conference Video
Keisuke Yamazaki, Satoshi Tamura, Yuuto Gotoh, Masaki Nose
ICPRAM2
2021 Speech Recognition using Deep Canonical Correlation Analysis in Noisy Environments
Shinnosuke Isobe, Satoshi Tamura, Satoru Hayamizu
ICPRAM2
2018 Audio-visual Voice Conversion Using Deep Canonical Correlation Analysis for Deep Bottleneck Features
Satoshi Tamura, Kento Horio, Hajime Endo, Satoru Hayamizu, Tomoki Toda
INTERSPEECH1
2016 Investigation of clinical process visualization using EMR data in clinics
Kodai Nakajima, Satoshi Tamura, Satoru Hayamizu, Takashi Ichinomiya, Yasutomi Kinosada
AMIA2
2015 Integration of deep bottleneck features for audio-visual speech recognition
Hiroshi Ninomiya, Norihide Kitaoka, Satoshi Tamura, Yurie Iribe, Kazuya Takeda
INTERSPEECH3
2014 Improvement of utterance clustering by using employees' sound and area data
abstract
In this paper, we propose to use staying area data toward the estimation of serving time for customers. To classify utterances enables us to estimate conversation types between speakers. However, its performance becomes lower in real environments. We propose a method using area data with sound data to solve this problem. We also propose a method to estimate the conversation types using the decision trees. They were tested with the data recorded in a Japanese restaurant. In the experiment to classify utterances, the proposed method performed better than the method using only sound data. In the experiment to estimate the conversation types, we succeeded to recover 70% of the mis-classified conversations using both of sound and area data.
Tetsuya Kawase, Masanori Takehara, Satoshi Tamura, Satoru Hayamizu, Ryuhei Tenmoku, Takeshi Kurata
ICASSP3
2014 Audio-visual voice conversion using noise-robust features
abstract
Voice Conversion (VC) is a technique to convert speech data of source speaker into ones of target speaker. VC has been investigated and statistical VC is used for various purposes. Conventional VC uses acoustic features, however, the audio-only VC has suffered from the degradation in noisy or real environments. This paper proposes an AudioVisual VC (AVVC) method using not only audio features but also visual information, i.e. lip images. Eigenlip feature is employed in our scheme as visual feature. We also propose a feature selection approach for audio-visual features. Experiments were conducted to evaluate our AVVC scheme comparing with audio-only VC, using noisy data. The results show that AVVC can improve the performance even in noisy environments, by properly selecting audio and visual parameters. It is also found that visual VC is also successful. Furthermore, it is observed that visual dynamic features are more effective than visual static information.
Kohei Sawada, Masanori Takehara, Satoshi Tamura, Satoru Hayamizu
ICASSP3
2013 Time-series analysis of health checkup data using Hidden-Markov model
Alwis Nazir, Ryouhei Kawamoto, Keiko Yamamoto, Satoshi Tamura, Takashi Ichinomiya, Satoru Hayamizu, Yasutomi Kinosada
AMIA4
2013 Probabilistic expression of Polynomial Semantic Indexing and its application for classification
Kentaro Minoura, Satoshi Tamura, Satoru Hayamizu
Pattern Recognit. Lett.2
2012 Sparse representation of audio features for sputum detection from lung sounds
Tatsuya Yamashita, Satoshi Tamura, Kenji Hayashi, Yutaka Nishimoto, Satoru Hayamizu
ICPR2
2010 Template-based spectral estimation using microphone array for speech recognition
Satoshi Tamura, Eriko Hishikawa, Wataru Taguchi, Satoru Hayamizu
INTERSPEECH1
2010 A robust audio-visual speech recognition using audio-visual voice activity detection
Satoshi Tamura, Masato Ishikawa, Takashi Hashiba, Shin'ichi Takeuchi, Satoru Hayamizu
INTERSPEECH1
2008 CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environments
abstract
In this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework
Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001
INTERSPEECH11
2008 Evaluation Framework for Distant-talking Speech Recognition under Reverberant Environments: newest Part of the CENSREC Series -
Takanobu Nishiura, Masato Nakayama, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001
LREC11
2007 Development of VAD evaluation framework CENSREC-1-C and investigation of relationship between VAD and speech recognition performance
abstract
Voice activity detection (VAD) plays an important role in speech processing including speech recognition, speech enhancement, and speech coding in noisy environments. We developed an evaluation framework for VAD in such environments, called corpus and environment for noisy speech recognition 1 concatenated (CENSREC-1-C). This framework consists of noisy continuous digit utterances and evaluation tools for VAD results. By adoptiong two evaluation measures, one for frame-level detection performance and the other for utterance-level detection performance, we provide the evaluation results of a power-based VAD method as a baseline. When using VAD in speech recognizer, the detected speech segments are extended to avoid the loss of speech frames and the pause segments are then absorbed by a pause model. We investigate the balance of an explicit segmentation by VAD and an implicit segmentation by a pause model using an experimental simulation of segment extension and show that a small extension improves speech recognition.
Norihide Kitaoka, Kazumasa Yamamoto, Tomohiro Kusamizu, Seiichi Nakagawa, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Takanobu Nishiura, Masato Nakayama, Yuki Denda, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001
ASRU13
2007 GEMSIS - a novel application of speech recognition to emergency and disaster medicine
Satoshi Tamura, Kunihiko Takamatsu, Shinji Ogura, Satoru Hayamizu
INTERSPEECH1
2006 Automatic metadata generation and video editing based on speech and image recognition for medical education contents
Satoshi Tamura, Koji Hashimoto, Jiong Zhu, Satoru Hayamizu, Hirotsugu Asai, Hideki Tanahashi, Makoto Kanagawa
INTERSPEECH1
2005 A Stream-Weight Optimization Method for Multi-Stream HMMS Based on Likelihood Value Normalization
abstract
In the field of audio-visual speech recognition, multi-stream HMM are widely used, thus how to automatically and properly determine stream weight factors using a small data set becomes an important research issue. This paper proposes a new stream-weight optimization method based on an output likelihood normalization criterion. In this method, the stream weights are adjusted to equalize the mean values of log likelihood for all HMM based on likelihood-ratio maximization which achieved significant improvement by using a large optimization data set. The new method is evaluated using Japanese connected digit speech recorded in real-world environments. Using 10 seconds speech data for stream-weight optimization, a 10% absolute accuracy improvement is achieved compared to the result before optimization. By additionally applying the MLLR (maximum likelihood linear regression) adaptation, a 23% improvement is obtained over the audio-only scheme.
Satoshi Tamura, Koji Iwano, Sadaoki Furui
ICASSP (1)1
2004 A stream-weight optimization method for audio-visual speech recognition using multi-stream HMMs
abstract
For multi-stream HMM that are widely used in audio-visual speech recognition, it is important to automatically and properly adjust stream weights. This paper proposes a stream-weight optimization technique based on a likelihood-ratio maximization criterion. In our audiovisual speech recognition system, video signals are captured and converted into visual features using HMM-based techniques. Extracted acoustic and visual features are concatenated into an audio-visual vector. A multi-stream HMM is obtained from audio and visual HMM. Experiments are conducted using Japanese connected digit speech recorded in real-world environments. Applying the MLLR (maximum likelihood linear regression) adaptation and our optimization method, we achieve a 29% absolute accuracy improvement and a 76% relative error rate reduction compared with the audio-only scheme.
Satoshi Tamura, Koji Iwano, Sadaoki Furui
ICASSP (1)1
2002 Robust bi-modal speech recognition based on state synchronous modeling and stream weight optimization
abstract
There have been higher demands recently for Automatic Speech Recognition (ASR) systems able to operate robustly in acoustically noisy environments. This paper proposes a method to effectively integrate audio and visual information in audio-visual (bi-modal) ASR systems. Such integration inevitably necessitates modeling of the synchronization of the audio and visual information. To address the time lag and correlation problems in individual features between speech and lip movements, we introduce a type of integrated HMM modeling of audio-visual information based on a family of HMM composition. The proposed model can represent state synchronicity not only within a phoneme but also between phonemes. Furthermore, we also propose a rapid stream weight optimization based on GPD algorithm for noisy bi-modal speech recognition. Evaluation experiments show that the proposed method improves the recognition accuracy for noisy speech. In SNR=0dB our proposed method attained 16% higher performance compared to a product HMMs without the synchronicity re-estimation.
Satoshi Nakamura 0001, Ken'ichi Kumatani, Satoshi Tamura
ICASSP3
2002 Multi-Modal Temporal Asynchronicity Modeling by Product HMMs for Robust
abstract
The demand for audio-visual speech recognition (AVSR) has increased in order to make speech recognition systems robust to acoustic noise. There are two kinds of research issue in audio-visual speech recognition, such as integration modeling considering asynchronicity between modalities and adaptive information weighting according information reliability. This paper proposes a method to effectively integrate audio and visual information. Such integration, inevitably, necessitates modeling the synchronization and asynchronization of audio and visual information. To address the time lag and correlation problems in individual features between speech and lip movements, we introduce a type of integrated HMM modeling of audio-visual information based on a family of a product HMM. The proposed model can represent state synchronicity not only within a phoneme, but also between phonemes. Furthermore, we also propose a rapid stream weight optimization based on the GPD algorithm for noisy, bimodal speech recognition. Evaluation experiments show that the proposed method improves the recognition accuracy for noisy speech. When SNR=0 dB our proposed method attained 16% higher performance compared to a product HMM without synchronicity re-estimation.
Satoshi Nakamura 0001, Ken'ichi Kumatani, Satoshi Tamura
ICMI3
2001 Ubiquitous speech processing
abstract
In the ubiquitous (pervasive) computing era, it is expected that everybody will access information services anytime anywhere, and these services are expected to augment various human intelligent activities. Speech recognition technology can play an important role in this era by providing: (a) conversational systems for accessing information services and (b) systems for transcribing, understanding and summarizing ubiquitous speech documents such as meetings, lectures, presentations and voicemails. In the former systems, robust conversation using wireless handheld/hands-free devices in the real mobile computing environment will be crucial as will multimodal speech recognition technology. To create the latter systems, the ability to understand and summarize speech documents is one of the key requirements. The paper presents technological perspectives and introduces several research activities being conducted from these standpoints in our research group.
Sadaoki Furui, Koji Iwano, Chiori Hori, Takahiro Shinozaki, Yohei Saito, Satoshi Tamura
ICASSP6