VLDB 2026 Research / reviewers in the wild / expert
Tetsuji Ogawa
dblp:10/2533
· DBLP profile ↗
73ranked-venue papers
10as first author
30since 2021 · last 2025
0000-0002-7316-2073ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 47 · 6 first-author · 19 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model for Guiding End-to-End Speech RecognitionabstractWe propose to utilize an instruction-tuned large language model (LLM) for guiding the text generation process in automatic speech recognition (ASR). Modern LLMs are adept at performing various text generation tasks through zero-shot learning, prompted with instructions designed for specific objectives. This paper explores the potential of LLMs to derive linguistic information that can facilitate text generation in end-to-end ASR models. Specifically, we instruct an LLM to correct grammatical errors in an ASR hypothesis and use the LLM-derived representations to refine the output further. The proposed model is built on the joint CTC and attention architecture, with the LLM serving as a front-end feature extractor for the decoder. The ASR hypothesis, subject to correction, is obtained from the encoder via CTC decoding and fed into the LLM along with a specific instruction. The decoder subsequently takes as input the LLM output to perform token predictions, combining acoustic information from the encoder and the powerful linguistic information provided by the LLM. Experimental results show that the proposed LLM-guided model achieves a relative gain of approximately 13% in word error rates across major benchmarks. Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 2 |
| 2025 | Video-Based Vibration Analysis for Predictive Maintenance: A Motion Magnification and Random Forest Approach
Walid Gomaa 0001, Abdelrahman Wael Ammar, Ismael Abbo, Mohamed Galal Nassef, Tetsuji Ogawa, Mohab Hossam |
ICINCO (1) | 5 |
| 2025 | End-to-End Speech Translation Guided by Robust Translation Capability of Large Language Model
Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2025 | Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
Asahi Sakuma, Hiroaki Sato, Ryuga Sugano, Tadashi Kumano, Yoshihiko Kawai, Tetsuji Ogawa |
INTERSPEECH | 6 |
| 2025 | Analysis of the Correlation Between Theory of Mind and Dialogue Ability to Identify Essential ToM for Dialogue Systems
Haruhisa Iseno, Atsumoto Ohashi, Tetsuji Ogawa, Shinnosuke Takamichi, Ryuichiro Higashinaka |
PACLIC | 3 |
| 2024 | Parody Detection Using Source-Target Attention with Teacher-Forced LyricsabstractWe propose an approach to detect parodies in singing voices, analyzing attention weights derived from an encoder-decoder-based automatic speech recognition (ASR) model. Here, parodies involve modifying and singing existing lyrics written for songs. Sharing such modified singing voices on the internet carries the potential risk of copyright infringement, posing the need of an automatic parody detection system. Given that songs typically comprise fixed lyrics, the pair of speech and its corresponding transcription can be used to analyze singing voices. In this work, we feed singing voices into an encoder-decoder-based ASR system and perform the decoding process using the corresponding lyrics in a teacher-forcing manner. Here, when the ASR model encounters a singing voice that includes a parody segment, there is a potential for the attention weights between the singing voice and the correct lyrics become collapsed. By identifying such misalignments in the attention weights, we attempt to detect parodies in singing voices. Experimental comparisons using real karaoke singing voice data demonstrate that the developed system achieves highly accurate parody detection performance by effectively identifying misalignments. Tomoki Ariga, Yosuke Higuchi, Kazutoshi Hayasaka, Naoki Okamoto, Tetsuji Ogawa |
ICASSP | 5 |
| 2024 | WindVibraTransformer: A Foundational Model for Precise and Robust Wind Turbine Condition Monitoring via Vibration SignalsabstractWe aimed to develop WindVibraTransformer, a foundational model based on Transformers applied to vibration spectrograms, to derive precise and robust feature representations for wind turbine condition monitoring. In anomaly detection for wind turbine equipment, it is common to employ inlier modeling, where models are built solely on data from normal operation, with deviations classified as anomalies. Acquiring data for wind turbines, especially from the specific turbine being monitored, is challenging, often resulting in limited availability. However, this limitation can increase susceptibility to false alarms due to environmental variations. Thus, in this study, we explored the use of a vision Transformer trained with self-supervised learning to predict masked spectra conditioned on unmasked ones from vibration spectrograms, serving as the precise and robust feature extractor. Since this model aims to capture relationships within the context of the surroundings rather than solely the overview of the spectrogram, the resulting feature representations are expected to be complex yet less directly influenced by environmental variations. The complexity of these feature representations is addressed by constructing the normal state model using highly expressive normalizing flows. Experimental comparisons conducted on anomaly detection using real wind turbine vibration signals demonstrated that our approach enables highly reliable detection compared to existing unsupervised auto-encoder-based and supervised wind turbine classifier-based feature extraction methods. Takuya Wakayama, Taiki Inoue, Jun Ogata, Makoto Iida, Tetsuji Ogawa |
ICMLA | 5 |
| 2024 | Leveraging Data from Vast Unexplored Seas: Positive Unlabeled Learning for Refining Prediction Area in Good Fishing Ground Prediction
Haruki Konii, Teppei Nakano, Yasumasa Miyazawa, Tetsuji Ogawa |
ICPR (10) | 4 |
| 2024 | Hierarchical Multi-Task Learning with CTC and Recursive Operation
Nahomi Kusunoki, Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2023 | A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, And ExtractionabstractWe propose a multi-task universal speech enhancement (MUSE) model that can perform five speech enhancement (SE) tasks: dereverberation, denoising, speech separation (SS), target speaker extraction (TSE), and speaker counting. This is achieved by integrating two modules into an SE model: 1) an internal separation module that does both speaker counting and separation; and 2) a TSE module that extracts the target speech from the internal separation outputs using target speaker cues. The model is trained to perform TSE if the target speaker cue is given and SS otherwise. By training the model to remove noise and reverberation, we allow the model to tackle the five tasks mentioned above with a single model, which has not been accomplished yet. Evaluation results demonstrate that the proposed MUSE model can successfully handle multiple tasks with a single model. Kohei Saijo, Wangyou Zhang, Zhongqiu Wang 0001, Shinji Watanabe 0001, Tetsunori Kobayashi, Tetsuji Ogawa |
ASRU | 6 |
| 2023 | Neural Diarization with Non-Autoregressive Intermediate AttractorsabstractEnd-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single neural network. While the EEND model can produce all frame-level speaker labels simultaneously, it disregards output label dependency. In this work, we propose a novel EEND model that introduces the label dependency between frames. The proposed method generates non-autoregressive intermediate attractors to produce speaker labels at the lower layers and conditions the subsequent layers with these labels. While the proposed model works in a non-autoregressive manner, the speaker labels are refined by referring to the whole sequence of intermediate labels. The experiments with the two-speaker CALLHOME dataset show that the intermediate labels with the proposed non-autoregressive intermediate attractors boost the diarization performance. The proposed method with the deeper net-work benefits more from the intermediate labels, resulting in better performance and training throughput than EEND-EDA. Yusuke Fujita, Tatsuya Komatsu, Robin Scheibler, Yusuke Kida, Tetsuji Ogawa |
ICASSP | 5 |
| 2023 | Intermpl: Momentum Pseudo-Labeling With Intermediate CTC LossabstractThis paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision. Momentum PL (MPL) trains a connectionist temporal classification (CTC)-based model on unlabeled data by continuously generating pseudo-labels on the fly and improving their quality. In contrast to autoregressive formulations, such as the attention-based encoder-decoder and transducer, CTC is well suited for MPL, or PL-based semi-supervised ASR in general, owing to its simple/fast inference algorithm and robustness against generating collapsed labels. However, CTC generally yields inferior performance than the autoregressive models due to the conditional independence assumption, thereby limiting the performance of MPL. We propose to enhance MPL by introducing intermediate loss, inspired by the recent advances in CTC-based modeling. Specifically, we focus on self-conditional and hierarchical conditional CTC, that apply auxiliary CTC losses to intermediate layers such that the conditional independence assumption is explicitly relaxed. We also explore how pseudo-labels should be generated and used as supervision for intermediate losses. Experimental results in different semi-supervised settings demonstrate that the proposed approach outperforms MPL and improves an ASR model by up to a 12.1% absolute performance gain. In addition, our detailed analysis validates the importance of the intermediate loss. Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2023 | BECTRA: Transducer-Based End-To-End ASR with Bert-Enhanced EncoderabstractWe present BERT-CTC-Transducer (BECTRA), a novel end-to-end automatic speech recognition (E2E-ASR) model formulated by the transducer with a BERT-enhanced encoder. Integrating a large-scale pre-trained language model (LM) into E2E-ASR has been actively studied, aiming to utilize versatile linguistic knowledge for generating accurate text. One crucial factor that makes this integration challenging lies in the vocabulary mismatch; the vocabulary constructed for a pre-trained LM is generally too large for E2E-ASR training and is likely to have a mismatch against a target ASR domain. To overcome such an issue, we propose BECTRA, an extended version of our previous BERT-CTC, that realizes BERT-based E2E-ASR using a vocabulary of interest. BECTRA is a transducer-based model, which adopts BERT-CTC for its encoder and trains an ASR-specific decoder using a vocabulary suitable for a target task. With the combination of the transducer and BERT-CTC, we also propose a novel inference algorithm for taking advantage of both autoregressive and non-autoregressive decoding. Experimental results on several ASR tasks, varying in amounts of data, speaking styles, and languages, demonstrate that BECTRA outperforms BERT-CTC by effectively dealing with the vocabulary mismatch while exploiting BERT knowledge. Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2023 | Self-Remixing: Unsupervised Speech Separation VIA Separation and RemixingabstractWe present Self-Remixing, a novel self-supervised speech separation method, which refines a pre-trained separation model in an unsupervised manner. Self-Remixing consists of a shuffler module and a solver module, and they grow together through separation and remixing processes. Specifically, the shuffler first separates observed mixtures and makes pseudo-mixtures by shuffling and remixing the separated signals. The solver then separates the pseudo-mixtures and remixes the separated signals back to the observed mixtures. The solver is trained using the observed mixtures as supervision, while the shuffler’s weights are updated by taking the moving average with the solver’s, generating the pseudo-mixtures with fewer distortions. Our experiments demonstrate that Self-Remixing gives better performance over existing remixing-based self-supervised methods with the same or less training costs under unsupervised setup. Self-Remixing also outperforms baselines in semi-supervised domain adaptation, showing effectiveness in multiple setups. Kohei Saijo, Tetsuji Ogawa |
ICASSP | 2 |
| 2023 | Conversation-Oriented ASR with Multi-Look-Ahead CBS ArchitectureabstractDuring conversations, humans are capable of inferring the intention of the speaker at any point of the speech to prepare the following action promptly. Such ability is also the key for conversational systems to achieve rhythmic and natural conversation. To perform this, the automatic speech recognition (ASR) used for transcribing the speech in real-time must achieve high accuracy without delay. In streaming ASR, high accuracy is assured by attending to look-ahead frames, which leads to delay increments. To tackle this trade-off issue, we propose a multiple latency streaming ASR to achieve high accuracy with zero look-ahead. The proposed system contains two encoders that operate in parallel, where a primary encoder generates accurate outputs utilizing look-ahead frames, and the auxiliary encoder recognizes the look-ahead portion of the primary encoder without look-ahead. The proposed system is constructed based on contextual block streaming (CBS) architecture, which leverages block processing and has a high affinity for the multiple latency architecture. Various methods are also studied for architecting the system, including shifting the network to perform as different encoders; as well as generating both encoders’ outputs in one encoding pass. Huaibo Zhao, Shinya Fujie, Tetsuji Ogawa, Jin Sakuma, Yusuke Kida, Tetsunori Kobayashi |
ICASSP | 3 |
| 2023 | Masry: A Text-to-Speech System for the Egyptian Arabic
Ahmed Hammad Azab, Ahmed Bayoumy Zaki, Tetsuji Ogawa, Walid Gomaa 0001 |
ICINCO (2) | 3 |
| 2023 | Learning Discriminative Feature Representations via Metric Learning for Early Operation of Wind Turbine Anomaly Detection SystemsabstractTo achieve robust wind turbine anomaly detection, we attempted to incorporate metric learning into the process of learning discriminative feature representations. In the context of anomaly detection based on inlier modeling, where anomalies are identified as inputs deviating from the normal state distribution, the key to building a high-performance and robust system, even with limited training data, lies in eliminating the influence of environmental differences from the inputs and obtaining a compact normal state distribution. To address this, we designed a deep neural network-based feature extractor to mitigate the impact of environmental changes. Additionally, we integrated metric learning into its learning process to obtain a more compact distribution by embedding inputs with similar properties close to one another. We validated the effectiveness of the developed network as a feature extractor for constructing a normal model with limited data through experimental comparisons using vibration data collected from sensors. Specifically, we utilized MobileNet as a discriminative feature extractor to identify 37 wind turbines. By introducing metric learning based on t-SNE into its training process, we achieved a notable 50 % improvement in the AUC. Remarkably, this significant improvement was accomplished using only a small amount of data, approximately 13 minutes, obtained from the monitored turbine. Taiki Inoue, Jun Ogata, Makoto Iida, Tetsuji Ogawa |
ICMLA | 4 |
| 2023 | Thermal Gait Dataset for Deep Learning-Oriented Gait RecognitionabstractThis study attempted to construct a thermal dataset of human gait in diverse environments suitable for building and evaluating sophisticated deep learning models (e.g., vision transformers) for gait recognition. Gait is a behavioral biometric to identify a person and requires no cooperation from the person, making it suitable for security and surveillance applications. For security purposes, it is desirable to be able to recognize a person in darkness or other inadequate lighting conditions, in which thermal imagery is advantageous over visible light imagery. Despite the importance of such nighttime person identification, available thermal gait datasets captured in the dark are scarce. This study, therefore, collected a relatively large set of thermal gait data in both indoor and outdoor environments with several walking styles, e.g., walking normally, walking while carrying a bag, and walking fast. This dataset was utilized in multiple gait recognition tasks, such as gender classification and person verification, using legacy convolutional neural networks (CNNs) and modern vision transformers (ViTs). Experiments using this dataset revealed the effective training method for person ver-ification, the effectiveness of ViT on gait recognition, and the robustness of the models against the difference in walking styles; it suggests that the developed dataset enables various studies on gait recognition using state-of-the-art deep learning models. Fatma Youssef, Ahmed El-Mahdy 0002, Tetsuji Ogawa, Walid Gomaa 0001 |
IJCNN | 3 |
| 2023 | Remixing-based Unsupervised Source Separation from Scratch
Kohei Saijo, Tetsuji Ogawa |
INTERSPEECH | 2 |
| 2022 | PostMe: Unsupervised Dynamic Microtask Posting For Efficient and Reliable CrowdsourcingabstractEven after over a decade of many crowdsourcing researches, we have no standard framework for low-cost quality assurance in crowdsourced data annotation. This paper proposes an unsupervised learning method for dynamic microtask posting which allows each microtask to adjust their own number of collected responses based on the data difficulty. Since crowdsourced data labels are likely to contain errors, researchers often employ majority voting that aggregates responses from multiple workers to calculate a final l abel. T his t echnique, h owever, i nvolves a trade-off between label accuracy and cost. This paper presents a dynamic microtask posting model that reduces the total number of collected responses while maintaining the labeling accuracy; we also aim to obtain the model with an “unsupervised” approach, which does not require training through experience of microtask posting for data labeled with ground-truths. Our simulation in annotating livestock surveillance images demonstrated that our approach achieved i) comparable learning performance to that of the supervised approach that required model training with labeled data, and ii) a significant c ost r eduction without degrading accuracy in comparison to simple majority voting. Ryo Yanagisawa, Teppei Nakano, Tetsunori Kobayashi, Tetsuji Ogawa |
IEEE Big Data | 5 |
| 2022 | Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword UnitsabstractIn end-to-end automatic speech recognition (ASR), a model is expected to implicitly learn representations suitable for recognizing a word-level sequence. However, the huge abstraction gap between input acoustic signals and output linguistic tokens makes it challenging for a model to learn the representations. In this work, to promote the word-level representation learning in end-to-end ASR, we propose a hierarchical conditional model that is based on connectionist temporal classification (CTC). Our model is trained by auxiliary CTC losses applied to intermediate layers, where the vocabulary size of each target subword sequence is gradually increased as the layer becomes close to the word-level output. Here, we make each level of sequence prediction explicitly conditioned on the previous sequences predicted at lower levels. With the proposed approach, we expect the proposed model to learn the word-level representations effectively by exploiting a hierarchy of linguistic structures. Experimental results on LibriSpeech-{100h, 960h} and TEDLIUM2 demonstrate that the proposed model improves over a standard CTCbased model and other competitive models from prior work. We further analyze the results to confirm the effectiveness of the intended representation learning with our model. Yosuke Higuchi, Keita Karube, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 3 |
| 2022 | Remix-Cycle-Consistent Learning on Adversarially Learned Separator for Accurate and Stable Unsupervised Speech SeparationabstractA new learning algorithm for speech separation networks is designed to explicitly reduce residual noise and artifacts in the separated signal in an unsupervised manner. Generative adversarial networks are known to be effective in constructing separation networks when the ground truth for the observed signal is inaccessible. Still, weak objectives aimed at distribution-to-distribution mapping make the learning unstable and limit their performance. This study introduces the remix-cycle-consistency loss as a more appropriate objective function and uses it to fine-tune adversarially learned source separation models. The remix-cycle-consistency loss is de-fined as the difference between the mixed speech observed at microphones and the pseudo-mixed speech obtained by alternating the process of separating the mixed sound and remixing its outputs with another combination. The minimization of this loss leads to an explicit reduction in the distortions in the output of the separation network. Experimental comparisons with multichannel speech separation demonstrated that the proposed method achieved high separation accuracy and learning stability comparable to supervised learning. Kohei Saijo, Tetsuji Ogawa |
ICASSP | 2 |
| 2022 | Can Humans Correct Errors From System? Investigating Error Tendencies in Speaker Identification Using Crowdsourcing
Yuta Ide, Teppei Nakano, Tetsuji Ogawa |
INTERSPEECH | 4 |
| 2022 | Confusion Detection for Adaptive Conversational Strategies of An Oral Proficiency Assessment Interview Agent
Mao Saeki, Kotoka Miyagi, Shinya Fujie, Shungo Suzuki, Tetsuji Ogawa, Tetsunori Kobayashi, Yoichi Matsuyama |
INTERSPEECH | 5 |
| 2022 | Unsupervised Training of Sequential Neural Beamformer Using Coarsely-separated and Non-separated Signals
Kohei Saijo, Tetsuji Ogawa |
INTERSPEECH | 2 |
| 2022 | Text-Only Domain Adaptation Based on Intermediate CTC
Hiroaki Sato, Tomoyasu Komori, Takeshi Mishima, Yoshihiko Kawai, Takahiro Mochizuki, Shoei Sato, Tetsuji Ogawa |
INTERSPEECH | 7 |
| 2021 | Improved Mask-CTC for Non-Autoregressive End-to-End ASRabstractFor real-world deployment of automatic speech recognition (ASR), the system is desired to be capable of fast inference while relieving the requirement of computational resources. The recently proposed end-to-end ASR system based on mask-predict with connectionist temporal classification (CTC), Mask-CTC, fulfills this demand by generating tokens in a non-autoregressive fashion. While Mask-CTC achieves remarkably fast inference speed, its recognition performance falls behind that of conventional autoregressive (AR) systems. To boost the performance of Mask-CTC, we first propose to enhance the encoder network architecture by employing a recently proposed architecture called Conformer. Next, we propose new training and decoding methods by introducing auxiliary objective to predict the length of a partial target sequence, which allows the model to delete or insert tokens during inference. Experimental results on different ASR tasks show that the proposed approaches improve Mask-CTC significantly, outperforming a standard CTC model (15.5% → 9.1% WER on WSJ). Moreover, Mask-CTC now achieves competitive results to AR models with no degradation of inference speed (< 0.1 RTF using CPU). We also show a potential application of Mask-CTC to end-to-end speech translation. Yosuke Higuchi, Hirofumi Inaguma, Shinji Watanabe 0001, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 4 |
| 2021 | Efficient and Stable Adversarial Learning Using Unpaired Data for Unsupervised Multichannel Speech Separation
Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi |
Interspeech | 3 |
| 2021 | VocalTurk: Exploring Feasibility of Crowdsourced Speaker Identification
Yuta Ide, Teppei Nakano, Tetsuji Ogawa |
Interspeech | 4 |
| 2021 | Analysis of Multimodal Features for Speaking Proficiency Scoring in an Interview DialogueabstractThis paper analyzes the effectiveness of different modalities in automated speaking proficiency scoring in an online dialogue task of non-native speakers. Conversational competence of a language learner can be assessed through the use of multimodal behaviors such as speech content, prosody, and visual cues. Although lexical and acoustic features have been widely studied, there has been no study on the usage of visual features, such as facial expressions and eye gaze. To build an automated speaking proficiency scoring system using multi-modal features, we first constructed an online video interview dataset of 210 Japanese English-learners with annotations of their speaking proficiency. We then examined two approaches for incorporating visual features and compared the effectiveness of each modality. Results show the end-to-end approach with deep neural networks achieves a higher correlation with human scoring than one with handcrafted features. Modalities are effective in the order of lexical, acoustic, and visual features. Mao Saeki, Yoichi Matsuyama, Satoshi Kobashikawa, Tetsuji Ogawa, Tetsunori Kobayashi |
SLT | 4 |
| 2020 | Exploiting Narrative Context and A Priori Knowledge of Categories in Textual Emotion ClassificationabstractRecognition of the mental state of a human character in text is a major challenge in natural language processing. In this study, we investigate the efficacy of the narrative context in recognizing the emotional states of human characters in text and discuss an approach to make use of a priori knowledge regarding the employed emotion category system. Specifically, we experimentally show that the accuracy of emotion classification is substantially increased by encoding the preceding context of the target sentence using a BERT-based text encoder. We also compare ways to incorporate a priori knowledge of emotion categories by altering the loss function used in training, in which our proposal of multi-task learning that jointly learns to classify positive/negative polarity of emotions is included. The experimental results suggest that, when using Plutchik’s Wheel of Emotions, it is better to jointly classify the basic emotion categories with positive/negative polarity rather than directly exploiting its characteristic structure in which eight basic categories are arranged in a wheel. Hikari Tanabe, Tetsuji Ogawa, Tetsunori Kobayashi, Yoshihiko Hayashi |
COLING | 2 |
| 2020 | Deep Speech Extraction with Time-Varying Spatial Filtering Guided By Desired Direction AttractorabstractIn this investigation, a deep neural network (DNN) based speech extraction method is proposed to enhance a speech signal propagating from the desired direction. The proposed method integrates knowledge based on a sound propagation model and the time-varying characteristics of a speech source, into a DNN-based separation framework. This approach outputs a separated speech source using time-varying spatial filtering, which achieves superior speech extraction performance compared with time-invariant spatial filtering. Given that the gradient of all modules can be calculated, back-propagation can be performed to maximize the speech quality of the output signal in an end-to-end manner. Guided information is also modeled based on the sound propagation model, which facilitates disentangled representations of the target speech source and noise signals. The experimental results demonstrate that the proposed method can extract the target speech source more accurately than conventional DNN-based speech source separation and conventional speech extraction using time-invariant spatial filtering. Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 3 |
| 2020 | Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short UtterancesabstractThis paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and testing time. However, many studies have shown that incorporating phoneme information is quite effective to improve the performance of the speaker recognition system. One reasonable explanation for this counter-intuitive result is that the pooling mechanism of segment-based speaker embedding can focus on the specific phonemes which contain rich speaker information, and phoneme information may help this. From this insight, we hypothesize that the pooling mechanism and phoneme-aware training are harmful to extract the speaker embeddings from extremely short utterances. To verify this hypothesis, an adversarial framework is introduced to remove phoneme-variability from the frame-wise speaker embeddings. The experimental results on the Librispeech corpus confirm that our frame-wise, phoneme-adversarial approach outperforms the conventional segment-wise, phoneme-aware approach for short utterances of less than about 1.4 seconds. Naohiro Tawara, Atsunori Ogawa, Tomoharu Iwata, Marc Delcroix, Tetsuji Ogawa |
ICASSP | 5 |
| 2020 | Feature Representation Learning for Calving Detection of Cows Using Video FramesabstractData-driven feature extraction is examined to realize accurate and robust calving detection. Automatic calving sign detection systems can support farmers' decision making. In this paper, neural networks are designed to extract information relevant to calving signs, which can be observed from video frames, such as the frequency in pre-calving postures, statistics in movement, and statistics in rotation. Experimental comparisons using surveillance videos demonstrate that the proposed feature extraction methods contribute to reducing false positives and explaining the basis of the prediction compared to the end-to-end calving detection system. Ryosuke Hyodo, Teppei Nakano, Tetsuji Ogawa |
ICPR | 3 |
| 2020 | Toward Building a Data-Driven System For Detecting Mounting Actions of Black Beef CattleabstractThis paper tackles on building a pattern recognition system that detects whether a pair of Japanese black beefs captured in a given image region is in a “mounting” action, which is known to be a sign critically important to be detected for cattle farmers before artificial insemination. The “mounting” action refers to a cattle's action where a cow bends over another cow usually when either cow is in estrus. Although a pattern recognition-based approach for detecting such an action would be appreciated as being low-cost and robust, it had not been discussed much due to the complexity of the system architecture, unavailability of datasets, etc. This study presents i) our image dataset construction technique that exploits both object detection algorithm and crowdsourcing for collecting cattle pair images with labels of either “mounting” or not; and ii) a system for detecting the mounting action from any given image of a cattle pair, developed based on the dataset. Starting with an algorithm for extracting regions of cattle pairs from a video frame based on intersection of single cattle regions, we then designed our crowdsourcing microtask in which crowd workers were given simple guidelines to annotate mounting-action-relevant labels to the extracted regions, to finally obtain a dataset. We also introduce our tandem-layered pattern recognition system trained with the dataset. The system is comprised of two serially-connected machine learning components, and is capable of more robustly detecting mounting actions even with a small amount of training data than a normal end-to-end neural network. Experimental comparisons demonstrated that our detection system was capable of detecting estrus with a precision rate of 80% and a recall rate of 76%. Yuriko Kawano, Teppei Nakano, Ikumi Kondo, Ryota Yamazaki, Hiromi Kusaka, Minoru Sakaguchi, Tetsuji Ogawa |
ICPR | 8 |
| 2020 | Crowdsourced Verification for Operating Calving Surveillance Systems at an Early StageabstractThis study attempts to use crowdsourcing to facilitate the operation of pattern-recognition-based video surveillance systems at an early stage. Target events (i.e. events to be detected during surveillance) are not frequently observed in recorded video, so achieving reliable surveillance on the basis of machine learning requires a sufficient amount of target data. Acquiring sufficient data is time-consuming. However, operating unreliable surveillance systems can induce many false alarms. Crowdsourcing is introduced to address this problem by verifying the unreliable results in data-driven surveillance. Experimental simulation conducted using monitoring video of Japanese black beef cattle demonstrates that crowdsourced verification successfully reduced false alarms in calving detection systems. Yusuke Okimoto, Soshi Kawata, Teppei Nakano, Tetsuji Ogawa |
ICPR | 5 |
| 2020 | Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask PredictabstractWe present Mask CTC, a novel non-autoregressive end-to-end automatic speech recognition (ASR) framework, which generates a sequence by refining outputs of the connectionist temporal classification (CTC).Neural sequence-to-sequence models are usually autoregressive: each output token is generated by conditioning on previously generated tokens, at the cost of requiring as many iterations as the output length.On the other hand, non-autoregressive models can simultaneously generate tokens within a constant number of iterations, which results in significant inference time reduction and better suits end-toend ASR model for real-world scenarios.In this work, Mask CTC model is trained using a Transformer encoder-decoder with joint training of mask prediction and CTC.During inference, the target sequence is initialized with the greedy CTC outputs and low-confidence tokens are masked based on the CTC probabilities.Based on the conditional dependence between output tokens, these masked low-confidence tokens are then predicted conditioning on the high-confidence tokens.Experimental results on different speech recognition tasks show that Mask CTC outperforms the standard CTC model (e.g., 17.9% → 12.1% WER on WSJ) and approaches the autoregressive model, requiring much less inference time using CPUs (0.07 RTF in Python implementation).All of our codes are publicly available at https://github.com/espnet/espnet. Yosuke Higuchi, Shinji Watanabe 0001, Nanxin Chen, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2020 | Mentoring-Reverse Mentoring for Unsupervised Multi-Channel Speech Source Separation
Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2019 | Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware TrainingabstractAn adversarial denoising autoencoder (ADAE) with noise-aware training is proposed and successfully applied to post-filtering for linear noise reduction. The ADAE is effective for attenuating interference sounds, however, it is difficult to learn to handle its various unexpected harmful effects (e.g., various types of noise) using a single network. Legacy speech enhancement was introduced as a pre-processor to make it possible to efficiently train the ADAEs by reducing the unexpected variabilities in the inputs to the ADAEs. Time-frequency masking performed well to suppress the variabilities, however, it induced unpleasant distortion, which is difficult for the ADAE to complement. In this paper, a minimum variance distortionless response (MVDR) beam-former, which can avoid troublesome non-linear distortions, is exploited as a preprocessor, and the MVDR outputs are used as the inputs to the ADAE-based post-filter. In addition, noise-dominant signals derived from the MVDR beamformer can improve the accuracy of the ADAE-based post-filter because the residual noise depends on the original noise signals. Experimental comparisons conducted using multichannel speech enhancement demonstrate that ADAE-based post-filtering yields significant improvements over the MVDR-and ADAE-based speech enhancement systems, and noise-aware training of ADAE works well. Naohiro Tawara, Hikari Tanabe, Tetsunori Kobayashi, Masaru Fujieda, Kazuhiro Katagiri, Takashi Yazu, Tetsuji Ogawa |
ICASSP | 7 |
| 2019 | Speaker Adversarial Training of DPGMM-Based Feature Extractor for Zero-Resource Languages
Yosuke Higuchi, Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa |
INTERSPEECH | 4 |
| 2019 | Multi-Channel Speech Enhancement Using Time-Domain Convolutional Denoising Autoencoder
Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa |
INTERSPEECH | 3 |
| 2018 | Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific RepresentationsabstractTraining recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to a target domain, have been proposed for low-resource language modeling. While there are the commonalities and discrepancies between the source and target domains in terms of the statistics of words and their contexts, these methods for domain adaptation make the commonalities and discrepancies jumbled. We propose novel domain adaptation techniques for RNNLM by introducing domain-shared and domain-specific word embedding and contextual features. This explicit modeling of the commonalities and discrepancies would improve the language modeling performance. Experimental comparisons using multiparty conversation data as the target domain augmented by lecture data from the source domain demonstrate that the proposed domain adaptation method exhibits improvements in the perplexity and word error rate over the long short-term memory based language model (LSTMLM) trained using the source and target domain data. Tsuyoshi Morioka, Naohiro Tawara, Tetsuji Ogawa, Atsunori Ogawa, Tomoharu Iwata, Tetsunori Kobayashi |
ICASSP | 3 |
| 2018 | Speaker Invariant Feature Extraction for Zero-Resource Languages with Adversarial LearningabstractWe introduce a novel type of representation learning to obtain a speaker invariant feature for zero-resource languages. Speaker adaptation is an important technique to build a robust acoustic model. For a zero-resource language, however, conventional model-dependent speaker adaptation methods such as constrained maximum likelihood linear regression are insufficient because the acoustic model of the target language is not accessible. Therefore, we introduce a model-independent feature extraction based on a neural network. Specifically, we introduce a multi-task learning to a bottleneck feature-based approach to make bottleneck feature invariant to a change of speakers. The proposed network simultaneously tackles two tasks: phoneme and speaker classifications. This network trains a feature extractor in an adversarial manner to allow it to map input data into a discriminative representation to predict phonemes, whereas it is difficult to predict speakers. We conduct phone discriminant experiments in Zero Resource Speech Challenge 2017. Experimental results showed that our multi-task network yielded more discriminative features eliminating the variety in speakers. Taira Tsuchiya, Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 3 |
| 2018 | Sequential Fish Catch Forecasting Using Bayesian State Space ModelsabstractA new state space model suitable for fixed shore net fishing is proposed and successfully applied to daily fish catch forecasting. Accurate prediction of daily fish catches makes it possible to support fishery workers with decision-making for efficient operations. For that purpose, the predictive model should be intuitive to the fishery workers and provide an estimate with a confidence. In the present paper, a fish catch forecasting method is developed using a state space model that emulates the process of fixed shore net fishing. In this method, the parameter estimation and prediction are sequentially performed using the Hamiltonian Monte Carlo method. The experimental comparisons using actual fish catch data and public meteorological information demonstrated that the proposed forecasting system yielded significant reductions in predictive errors over the systems based on decision-trees and legacy state-space models. Yuya Kokaki, Naohiro Tawara, Tetsunori Kobayashi, Kazuo Hashimoto, Tetsuji Ogawa |
ICPR | 5 |
| 2017 | Associative Memory Model-Based Linear Filtering and Its Application to Tandem Connectionist Blind Source SeparationabstractWe propose a blind source separation method that yields high-quality speech with low distortion. Time-frequency (TF) masking can effectively reduce interference, but it produces nonlinear distortion. By contrast, linear filtering using a separation matrix such as independent vector analysis (IVA) can avoid nonlinear distortion, but the separation performance is reduced under reverberant conditions. The tandem connectionist approach combines several separation methods and it has been used frequently to compensate for the disadvantages of these methods. In this study, we propose associative memory model (AMM)-based linear filtering and a tandem connectionist framework, which applies TF masking followed by linear filtering. By using AMM trained with speech spectra to optimize the separation matrix, the proposed linear filtering method considers the properties of speech that are not considered explicitly in IVA, such as the harmonic components of spectra. TF masking is applied in the proposed tandem connectionist framework to reduce unwanted components that hinder the optimization of the separation matrix, and it is approximated by using a linear separation matrix to reduce nonlinear distortion. The results obtained in simultaneous speech separation experiments demonstrate that although the proposed linear filtering method can increase the signal-to-distortion ratio (SDR) and signal-to-interference ratio (SIR) compared with IVA, the proposed tandem connectionist framework can obtain greater increases in SDR and SIR, and it reduces the phoneme error rate more than the proposed linear filtering method. Motoi Omachi, Tetsuji Ogawa, Tetsunori Kobayashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Real-Time Large-Scale Map Matching Using Mobile Phone DataabstractWith the wide spread use of mobile phones, cellular mobile big data is becoming an important resource that provides a wealth of information with almost no cost. However, the data generally suffers from relatively high spatial granularity, limiting the scope of its application. In this article, we consider, for the first time, the utility of actual mobile big data for map matching allowing for “microscopic” level traffic analysis. The state-of-the-art in map matching generally targets GPS data, which provides far denser sampling and higher location resolution than the mobile data. Our approach extends the typical Hidden-Markov model used in map matching to accommodate for highly sparse location trajectories, exploit the large mobile data volume to learn the model parameters, and exploit the sparsity of the data to provide for real-time Viterbi processing. We study an actual, anonymised mobile trajectories data set of the city of Dakar, Senegal, spanning a year, and generate a corresponding road-level traffic density, at an hourly granularity, for each mobile trajectory. We observed a relatively high correlation between the generated traffic intensities and corresponding values obtained by the gravity and equilibrium models typically used in mobility analysis, indicating the utility of the approach as an alternative means for traffic analysis. Essam Algizawy, Tetsuji Ogawa, Ahmed El-Mahdy 0002 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2016 | A new efficient measure for accuracy prediction and its application to multistream-based unsupervised adaptationabstractA new efficient measure for predicting estimation accuracy is proposed and successfully applied to multistream-based unsupervised adaptation of ASR systems to address data uncertainty when the ground-truth is unknown. The proposed measure is an extension of the M-measure, which predicts confidence in the output of a probability estimator by measuring the divergences of probability estimates spaced at specific time intervals. In this study, the M-measure was extended by considering the latent phoneme information, resulting in an improved reliability. Experimental comparisons carried out in a multistream-based ASR paradigm demonstrated that the extended M-measure yields a significant improvement over the original M-measure, especially under narrow-band noise conditions. Tetsuji Ogawa, Sri Harish Reddy Mallidi, Emmanuel Dupoux, Jordan Cohen, Naomi Feldman, Hynek Hermansky |
ICPR | 1 |
| 2015 | Uncertainty estimation of DNN classifiersabstractNew efficient measures for estimating uncertainty of deep neural network (DNN) classifiers are proposed and successfully applied to multistream-based unsupervised adaptation of ASR systems to address uncertainty derived from noise. The proposed measure is the error from associative memory models trained on outputs of a DNN. In the present study, an attempt is made to use autoencoders for remembering the property of data. Another measure proposed is an extension of the M-measure, which computes the divergences of probability estimates spaced at specific time intervals. The extended measure results in an improved reliability by considering the latent information of phoneme duration. Experimental comparisons carried out in a multistream-based ASR paradigm demonstrates that the proposed measures yielded improvements over the multistyle trained system and system selected based on existing measures. Fusion of the proposed measures achieved almost the same performance as the oracle system selection. Sri Harish Reddy Mallidi, Tetsuji Ogawa, Hynek Hermansky |
ASRU | 2 |
| 2015 | Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial WorkshopabstractA group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which the estimator was trained. The paper describes the problem and summarizes approaches that were taken by the group1. Hynek Hermansky, Lukás Burget, Jordan Cohen, Emmanuel Dupoux, Naomi Feldman, John Godfrey, Sanjeev Khudanpur, Matthew Maciejewski, Sri Harish Reddy Mallidi, Anjali Menon, Tetsuji Ogawa, Vijayaditya Peddinti, Richard C. Rose, Richard M. Stern, Matthew Wiesner, Karel Veselý |
ICASSP | 11 |
| 2015 | A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditionsabstractThe present paper dealt with speaker clustering for speech corrupted by noise. In general, the performance of speaker clustering significantly depends on how well the similarities between speech utterances can be measured. The recently proposed i-vector-based cosine similarity has yielded the state-of-the-art performance in speaker clustering systems. However, this similarity often fails to capture the speaker similarity under noisy conditions. Therefore, we attempted to examine the efficiency of spectral clustering on i-vector-based similarity for speech corrupted by noise because spectral clustering can yield robustness against noise by non-linear projection. Experimental comparisons demonstrated that spectral clustering yielded significant improvement from conventional methods, such as agglomerative clustering and k-means clustering, under non-stationary noise conditions. Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 2 |
| 2015 | Autoencoder based multi-stream combination for noise robust speech recognition
Sri Harish Reddy Mallidi, Tetsuji Ogawa, Karel Veselý, Phani S. Nidadavolu, Hynek Hermansky |
INTERSPEECH | 2 |
| 2015 | Bilinear map of filter-bank outputs for DNN-based speech recognition
Tetsuji Ogawa, Kenshiro Ueda, Kouichi Katsurada, Tetsunori Kobayashi, Tsuneo Nitta |
INTERSPEECH | 1 |
| 2014 | Effect of frequency weighting on MLP-based speaker canonicalization
Yuichi Kubota, Motoi Omachi, Tetsuji Ogawa, Tetsunori Kobayashi, Tsuneo Nitta |
INTERSPEECH | 3 |
| 2013 | Stream selection and integration in multistream ASR using GMM-based performance monitoringabstractA moderately deep and rather wide artificial neural net is applied in phoneme recognition of noisy speech. The net is formed by first estimating posterior probabilities of phonemes in 21 band-limited streams covering the whole speech spectrum. These 21 band-limited streams are subdivided into three seven band-limited stream subsets, by differently sub-sampling the original 21 band-limited streams. In the second processing stage, all non-empty combinations of seven band-limited streams from each subset are formed as inputs to 127 artificial neural nets that are again trained to yield phoneme posteriors. In this way, 127 × 3 = 381 processing streams are formed. A novel technique for finding the best combination of the resulting 381 parallel processing streams, which uses the likelihood of a single-state Gaussian mixture model of the final classifier output is applied to selecting the most efficient streams. The technique is efficient in phoneme recognition of speech that is corrupted by realistic additive noise. Tetsuji Ogawa, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 1 |
| 2012 | Fully Bayesian inference of multi-mixture Gaussian model and its evaluation using speaker clusteringabstractThis study aims to verify effective optimization methods for estimating parametric, fully Bayesian models in speech processing. For that purpose, we investigate the impact of the difference in optimization methods for the multi-scale Gaussian mixture model, which is suitable for speaker clustering, on the clustering accuracy. The Markov chain Monte Carlo (MCMC)-based method was compared with the variational Bayesian method in the speaker clustering experiment; with a small amount of data, the MCMC-based method was more effective; with large scale data (more than one million samples), the difference between these methods in terms of the clustering accuracy decreased and the MCMC-based method was computationally efficient. Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe 0001, Tetsunori Kobayashi |
ICASSP | 2 |
| 2012 | An improved entropy-based multiple kernel learning
Hideitsu Hino, Tetsuji Ogawa |
ICPR | 2 |
| 2012 | Fully Bayesian speaker clustering based on hierarchically structured utterance-oriented Dirichlet process mixture model
Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2011 | Speaker recognition using multiple kernel learning based on conditional entropy minimizationabstractWe applied a multiple kernel learning (MKL) method based on information-theoretic optimization to speaker recognition. Most of the kernel methods applied to speaker recognition systems require a suitable kernel function and its parameters to be determined for a given data set. In contrast, MKL eliminates the need for strict determination of the kernel function and parameters by using a convex combination of element kernels. In the present paper, we describe an MKL algorithm based on conditional entropy minimization (MCEM). We experimentally verified the effectiveness of MCEM for speaker classification; this method reduced the speaker error rate as compared to conventional methods. Tetsuji Ogawa, Hideitsu Hino, Nima Reyhani, Noboru Murata, Tetsunori Kobayashi |
ICASSP | 1 |
| 2011 | Speaker Verification Robust to Talking Style Variation Using Multiple Kernel Learning Based on Conditional Entropy Minimization
Tetsuji Ogawa, Hideitsu Hino, Noboru Murata, Tetsunori Kobayashi |
INTERSPEECH | 1 |
| 2011 | Spatial Filter Calibration Based on Minimization of Modified LSD
Nobuaki Tanaka, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2011 | Speaker Clustering Based on Utterance-Oriented Dirichlet Process Mixture ModelabstractThis paper provides the analytical solution and algorithm of UO-DPMM based on a non-parametric Bayesian manner, and thus realizes fully Bayesian speaker clustering. We carried out preliminary speaker clustering experiments by using a TIMIT database to compare the proposed method with the conventional Bayesian Information Criterion (BIC) based method, which is an approximate Bayesian approach. The results showed that the proposed method outperformed the conventional one in terms of both computational cost and robustness to changes in tuning parameters. Naohiro Tawara, Shinji Watanabe 0001, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2009 | Robot auditory system using head-mounted square microphone arrayabstractA new noise reduction method suitable for autonomous mobile robots was proposed and applied to preprocessing of a hands-free spoken dialogue system. When a robot talks with a conversational partner in real environments, not only speech utterances by the partner but also various types of noise, such as directional noise, diffuse noise, and noise from the robot, are observed at microphones. We attempted to remove these types of noise simultaneously with small and light-weighted devices and low-computational-cost algorithms. We assumed that the conversational partner of the robot was in front of the robot. In this case, the aim of the proposed method is extracting speech signals coming from the frontal direction of the robot. The proposed noise reduction system was evaluated in the presence of various types of noise: the number of word errors was reduced by 69% as compared to the conventional methods. The proposed robot auditory system can also cope with the case in which a conversational partner (i.e., a sound source) moves from the front of the robot: the sound source was localized by face detection and tracking using facial images obtained from a camera mounted on an eye of the robot. As a result, various types of noise could be reduced in real time, irrespective of the sound source positions, by combining speech information with image information. Kosuke Hosoya, Tetsuji Ogawa, Tetsunori Kobayashi |
IROS | 2 |
| 2008 | A Kalman filter based restoration method for in-vehicle camera images in foggy conditionsabstractThis paper proposes a Kalman filter based restoration method for images obtained by in-vehicle camera in foggy conditions. The proposed method introduces two novel approaches into the Kalman filter based restoration. The first one is an automatic determination of a fog deterioration model. A vanishing point in the foggy image is estimated by using cross ratio of lane marking, and automatic determination of all parameters of the fog deterioration model is realized. Furthermore, the obtained model is introduced into the Kalman filter. Specifically, our method regards each frame as a state variable and its observation model is defined by the fog deterioration model. Then, since the correlation between successive frame can be effectively utilized by the Kalman filter, the accurate restoration of foggy images is achieved. Experimental results show that the proposed method achieves higher performance than the traditional method based on the fog deterioration model. Tomoki Hiramatsu, Tetsuji Ogawa, Miki Haseyama |
ICASSP | 2 |
| 2008 | Kernel PCA-based resolution enhancement approach of still images using different levels of pyramid structureabstractThis paper presents a kernel PCA-based adaptive resolution enhancement method of still images. The proposed method introduces two novel approaches into the kernel PCA-based reconstruction of high-frequency components missed from a high-resolution (HR) image. First, since local images between two different resolution levels of a pyramid structure are similar to each other, nonlinear eigenspaces of local images in the target low-resolution (LR) image are utilized as those of local images in the HR image. Further, in the kernel PCA-based reconstruction process of the high-frequency components, our method monitors errors caused in the known low-frequency components and realizes the selection of the optimal eigenspace. Then, since the missing high-frequency components can be adaptively estimated, the accurate HR image can be obtained. Tetsuji Ogawa, Miki Haseyama |
ICASSP | 1 |
| 2008 | Speech enhancement using square microphone array for mobile devicesabstractIn this paper, we propose a new type of speech enhancement method that is suitable for mobile devices used in noisy environments. For the sake of achieving high-performance speech recognition and auditory perception in the mobile devices, disturbance noises have to be removed under the requirements of a space-saving microphone arrangement and a low computational cost. The proposed method can reduce both the directional and the diffuse noises under the requirements for the mobile devices by applying the square microphone array and the low-cost processing that consists of multiple null beamforming, their minimum power channel selection and Wiener filtering. The effectiveness of the proposed method is clarified for speech recognition accuracies and speech qualities under the condition in which both the directional and the diffuse noises exist simultaneously: it reduced 40% of recognition errors and improved PESQ-based MOS value by 0.75 point. Shintaro Takada, Tetsuji Ogawa, Kenzo Akagiri, Tetsunori Kobayashi |
ICASSP | 2 |
| 2008 | CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environmentsabstractIn this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
INTERSPEECH | 12 |
| 2007 | Adequacy Analysis of Simulation-Based Assessment of Speech Recognition SystemabstractThe adequacies of the simulation-based assessment of speech recognition systems under noisy conditions are investigated and discussed. To evaluate the speech recognition systems in various environments, it is desirable to collect the test data uttered in the corresponding environments but it is not realistic since enormous works are required. To conduct evaluations of the speech recognition systems properly, it is promising to simulate evaluation experiments in the target environments as described below: comparatively small test data are collected, and test data of the target environment are generated by computing convolution of the impulse response of the target environment with the collected data. However, it is well known that changes of the acoustic characteristics are caused by the Lombard effect, and so it is not necessarily obvious whether the simulation can precisely approximate the experiment in actual environment. This paper clarifies the condition to perform effective simulations of the noisy speech recognition, focusing on the influence of impulse response accuracies and Lombard effects on the speech recognition performance. Tetsuji Ogawa, Satoshi Kanba, Tetsunori Kobayashi |
ICASSP (4) | 1 |
| 2006 | Manifold HLDA and its application to robust speech recognition
Toshiaki Kubo, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2005 | Optimizing the structure of partly-hidden Markov models using weighted likelihood-ratio maximization criterion
Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 1 |
| 2004 | Recognition of three simultaneous utterance of speech by four-line directivity microphone mounted on head of robotabstractA sound source separation method using four-line directivity microphones mounted on a head of a robot is proposed and applied to speech recognition under existence of two disturbances of speech. Sound source separation methods using microphones mounted on robot heads generally used strict head-related transfer functions(HRTF). We propose a robust sound source separation that does not require an estimate of a strict HRTF. Our method takes advantage of a sound pressure difference with the robot head acting as a sound barrier. The enhancement of the difference in the target speech is performed by signal processing of three layers:two-line SAFIA, twoline Spectral Subtraction and their integration. The experimental results of three simultaneous utterance recognition with vocabulary of 20K show that the proposed method is effective in achieving 71% error reduction. Naoya Mochiki, Tetsunori Kobayashi, Toshiyuki Sekiya, Tetsuji Ogawa |
INTERSPEECH | 4 |
| 2003 | Hybrid modeling of PHMM and HMM for speech recognitionabstractA hybrid acoustic model of partly hidden Markov model (PHMM) and HMM is proposed. PHMM was proposed in our previous work to deal with the complicated temporal changes of acoustic features (Ogawa, T. and Kobayashi, T, Proc. ICSLP2002, p.2673-6, 2002). It can realized observation dependent behaviors in both observations and state transitions. It achieved good performance but some errors with different trends from HMM still remained. We have designed a new acoustic model on the basis of PHMM, in which the observation and state transition probabilities are defined by the geometric means of PHMM-based ones and HMM-based ones. In this framework, if a word hypothesis is given a low score by either PHMM or HMM, it almost loses the possibility of being a probable candidate. Since many errors are due to high-scores of incorrect categories rather than low-score of the correct category, this property contributes to reducing errors. Moreover, the proposed model is more stable than PHMM because the higher order statistics of PHMM, which is generally accurate but sometimes less reliable, are smoothed by the lower order statistics of HMM, which is not so accurate, but robust. Experimental results show the effectiveness of the proposed model: it reduces the word errors by 25% compared with HMM. Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP (1) | 1 |
| 2003 | Speech recognition of double talk using SAFIA-based audio segregation
Toshiyuki Sekiya, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2002 | Generalization of state-observation-dependency in partly hidden Markov models
Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 1 |