EDBT 2026 Demo / reviewers in the wild / expert
Koichi Shinoda
dblp:74/4623
· DBLP profile ↗
95ranked-venue papers
12as first author
16since 2021 · last 2025
0000-0003-1095-3203ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 48 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion Pretraining for Gait Recognition in the WildabstractRecently, diffusion models have garnered much attention for their remarkable generative capabilities. Yet, their application for representation learning remains largely unexplored. In this paper, we explore the potential of diffusion models to pretrain the backbone of a deep learning model for a specific application—gait recognition in the wild. To do so, we condition a latent diffusion model on the output of a gait recognition model backbone. Our pretraining experiments on the Gait3D and GREW datasets reveal an interesting phenomenon: diffusion pretraining causes the gait recognition backbone to separate gait sequences belonging to different subjects further apart than those belonging to the same subjects. Subsequently, our transfer learning experiments on Gait3D and GREW show that the pretrained backbone can serve as an effective initialization for the downstream gait recognition task, improving gait recognition accuracies by as much as 7.9% on Gait3D and 4.2% on GREW. Wei Ming Neo, Koichi Shinoda, Tat-Jen Cham |
ICIP | 2 |
| 2025 | SepVAC: Multitask Learning of Speaker Separation, Speaker Localization, Microphone Array Localization, and Room Acoustic Parameter Estimation in Various Acoustic ConditionsabstractThis paper proposes a multitask learning method for speech separation, that Separates speech and estimates the recording conditions in Various Acoustic Conditions (SepVAC) jointly. Unlike the previous methods that aim to achieve robustness against the uncertainty caused by noise and reverberation, this method explicitly estimates speaker & microphone location and room acoustic parameters to disambiguate them from speech features. We introduce curriculum learning to learn the model parameters stably. In our evaluation using SMS-WSJ-Plus dataset, it outperforms the state-of-the-art SpatialNet baseline by 0.67 points in word error rate (WER). Roland Hartanto, Sakriani Sakti, Koichi Shinoda |
INTERSPEECH | 3 |
| 2025 | Diffusion-Based Generative Regularization for Supervised Discriminative LearningabstractEnsuring the quality and quantity of labeled training data has long been a challenge in training deep neural networks for discriminative tasks. One solution to this problem is to use a generative model to augment training data and learn a discriminative model with it. For image classification, with the recent development of diffusion models, it has become possible to generate a variety synthetic images, and there are high expectations for their use as training data. However, to obtain high-quality labeled synthetic images, the hyperparameters and prompts often need to be manually tuned, and the accuracy of the trained image classification model is highly dependent on them. To address this issue, this paper proposes diffusion-based generative regularization, a supervised discriminative learning framework that utilizes a diffusion-based image generation model as a regularizer to robustly learn discriminative representations without the need to synthesize images. Our experiments using vision transformers and stable diffusion models on ImageNet-1k demonstrate that the proposed framework improves classification accuracy on both in-distribution and distribution-shifted data. Takuya Asakura, Nakamasa Inoue, Koichi Shinoda |
WACV | 3 |
| 2025 | ContextualCoder: Adaptive In-Context Prompting for Programmatic Visual Question AnsweringabstractVisual Question Answering (VQA) presents a challenging task at the intersection of computer vision and natural language processing, aiming to bridge the semantic gap between visual perception and linguistic comprehension. Traditional VQA approaches do not distinguish between data processing and reasoning, limiting their interpretability and generalizability in complex and diverse scenarios. Conversely, Programmatic Visual Question Answering (PVQA) models leverage large language models (LLMs) to generate executable codes, providing answers with detailed and interpretable reasoning processes. However, existing PVQA models typically rely on simplistic input-output prompting, which struggles to elicit domain-specific knowledge from LLMs and often produces unclear or extraneous outputs. Furthermore, PVQA models typically rely on a basic in-context example (ICE) selection methodology that is heavily influenced by individual word similarity rather than the overall sentence context. This leads to suboptimal ICE selection and a reliance on dataset-specific ICE candidates. In this paper, we propose ContextualCoder, a novel prompting framework tailored for PVQA models. ContextualCoder leverages frozen LLMs for code generation and pre-trained visual models for code execution, eliminating the need for extensive training and enhancing model flexibility. By incorporating an innovative prompting methodology and a novel ICE selection strategy, ContextualCoder facilitates the use of diverse in-context information for code generation, thereby improving the performance of PVQA models. Our approach surpasses state-of-the-art models, as evidenced by comprehensive experiments across diverse VQA datasets, including multilingual scenarios. Ruoyue Shen, Nakamasa Inoue, Dayan Guan, Rizhao Cai, Alex Chichung Kot, Koichi Shinoda |
IEEE Trans. Multim. | 6 |
| 2024 | Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question AnsweringabstractVisual question answering (VQA) is the task of providing accurate answers to natural language questions based on visual input. Programmatic VQA (PVQA) models have been gaining attention recently. These use large language models (LLMs) to formulate executable programs that address questions requiring complex visual reasoning. However, there are challenges in enabling LLMs to comprehend the usage of image processing modules and generate relevant code. To overcome these challenges, this paper introduces PyramidCoder, a novel prompting framework for PVQA models. PyramidCoder consists of three hierarchical levels, each serving a distinct purpose: query rephrasing, code generation, and answer aggregation. Notably, PyramidCoder utilizes a single frozen LLM and pre-defined prompts at each level, eliminating the need for additional training and ensuring flexibility across various LLM architectures. Compared to the state-of-the-art PVQA model, our approach improves accuracy by at least 0.5% on the GQA dataset, 1.4% on the VQAv2 dataset, and 2.9% on the NLVR2 dataset. Ruoyue Shen, Nakamasa Inoue, Koichi Shinoda |
ICIP | 3 |
| 2024 | MSDET: Multitask Speaker Separation and Direction-of-Arrival Estimation TrainingabstractThe information on the spatial location of speakers can be effectively used for multi-channel speaker separation. For example, Location-Based Training (LBT) uses the order of azimuth angles and distances of speakers to solve the permutation ambiguity problem. This location information can be used to improve the separation performance further. This paper proposes a multitask learning approach, Multitask Speaker Separation and Direction-of-Arrival Estimation Training (MSDET), jointly optimizing speaker separation and Direction-of-Arrival (DoA) estimation. In our evaluation using SMS-WSJ dataset, it outperforms LBT by 0.13 points in SI-SDR and 0.35 points in ESTOI. Roland Hartanto, Sakriani Sakti, Koichi Shinoda |
INTERSPEECH | 3 |
| 2024 | Co-speech Gesture Generation with Variational Auto Encoder
Shinichi Ka, Koichi Shinoda |
MMM (3) | 2 |
| 2024 | CAMOT: Camera Angle-aware Multi-Object TrackingabstractThis paper proposes CAMOT, a simple camera angle estimator for multi-object tracking to tackle two problems: 1) occlusion and 2) inaccurate distance estimation in the depth direction. Under the assumption that multiple objects are located on a flat plane in each video frame, CAMOT estimates the camera angle using object detection. In addition, it gives the depth of each object, enabling pseudo-3D MOT. We evaluated its performance by adding it to various 2D MOT methods on the MOT17 and MOT20 datasets and confirmed its effectiveness. Applying CAMOT to ByteTrack, we obtained 63.8% HOTA, 80.6% MOTA, and 78.5% IDF1 in MOT17, which are state-of-the-art results. Its computational cost is significantly lower than the existing deep-learning-based depth estimators for tracking. Felix Limanta, Kuniaki Uto, Koichi Shinoda |
WACV | 3 |
| 2023 | Synthesizing Speech from ECoG with a Combination of Transformer-Based Encoder and Neural VocoderabstractThis paper reports on a novel invasive brain–computer interface (BCI) paradigm that has successfully reconstructed spoken sentences from invasive electrocorticogram (ECoG) signals using deep-neural-network-based encoders and a pre-trained neural vocoder. We recorded ECoG signals while 13 participants were speaking short sentences. Our BCI could map the ECoG recording to the log-mel spectrograms of the spoken sentences using a bidirectional long short-term memory (BLSTM) or a Transformer. The estimated log-mel spectrograms were used in Parallel WaveGAN to synthesize speech waveforms. An evaluation of the model performance revealed that the Transformer model significantly outperformed (Wilcoxon signed-rank test, p < 0.001) the BLSTM in terms of mean square error loss and Pearson correlation. Kai Shigemi, Shuji Komeiji, Takumi Mitsuhashi, Yasushi Iimura, Hiroharu Suzuki, Hidenori Sugano, Koichi Shinoda, Kohei Yatabe, Toshihisa Tanaka 0001 |
ICASSP | 7 |
| 2023 | EvIs-Kitchen: Egocentric Human Activities Recognition with Video and Inertial Sensor Data
Yuzhe Hao, Kuniaki Uto, Asako Kanezaki, Ikuro Sato, Rei Kawakami, Koichi Shinoda |
MMM (1) | 6 |
| 2023 | Text-Guided Object Detector for Multi-modal Video Question AnsweringabstractVideo Question Answering (Video QA) is a task to answer a text-format question based on the understanding of linguistic semantics, visual information, and also linguistic-visual alignment in the video. In Video QA, an object detector pre-trained with large-scale datasets, such as Faster R-CNN, has been widely used to extract visual representations from video frames. However, it is not always able to precisely detect the objects needed to answer the question be-cause of the domain gaps between the datasets for training the object detector and those for Video QA. In this paper, we propose a text-guided object detector (TGOD), which takes text question-answer pairs and video frames as inputs, detects the objects relevant to the given text, and thus provides intuitive visualization and interpretable results. Our experiments using the STAGE framework on the TVQA+ dataset show the effectiveness of our proposed detector. It achieves a 2.02 points improvement in accuracy of QA, 12.13 points improvement in object detection (mAP50), 1.1 points improvement in temporal location, and 2.52 points improvement in ASA over the STAGE original detector. Ruoyue Shen, Nakamasa Inoue, Koichi Shinoda |
WACV | 3 |
| 2022 | Implicit Neural Representations for Variable Length Human Motion Generation
Pablo Cervantes, Yusuke Sekikawa, Ikuro Sato, Koichi Shinoda |
ECCV (17) | 4 |
| 2022 | Transformer-Based Estimation of Spoken Sentences Using ElectrocorticographyabstractInvasive brain–machine interfaces (BMIs) are a promising neurotechnological venture for achieving direct speech communication from a human brain, but it faces many challenges. In this paper, we measured the invasive electrocorticogram (ECoG) signals from seven participating epilepsy patients as they spoke a sentence consisting of multiple phrases. A Transformer encoder was incorporated into a "sequence-to-sequence" model to decode spoken sentences from the ECoG. The decoding test revealed that the use of the Transformer model achieved a minimum phrase error rate (PER) of 16.4%, and the median (±standard deviation) across seven participants was 31.3% (±10.0%). Moreover, the proposed model with the Transformer achieved significantly better decoding accuracy than a conventional long short-term memory model. Shuji Komeiji, Kai Shigemi, Takumi Mitsuhashi, Yasushi Iimura, Hiroharu Suzuki, Hidenori Sugano, Koichi Shinoda, Toshihisa Tanaka 0001 |
ICASSP | 7 |
| 2022 | RI-DC: Rotation-Invariant Detection and Classification for Wheat Head DetectionabstractWe propose a novel automatic detection method of wheat heads from images taken from above wheat fields. The automatic detection of heads is useful for predicting yields. For this purpose, deep learning based methods has proved to be effective recently, but they are not robust against the variations of head directions. To tackle this problem, we utilize a two-step approach which first carries out object detection for augmented test images rotated in many directions, and then classifies the detected objects by using a classifier trained with many rotated images. It was evaluated by using GWHD dataset and proved to be effective. Takeru Ito, Kuniaki Uto, Koichi Shinoda |
IGARSS | 3 |
| 2022 | MSR-DARTS: Minimum Stable Rank of Differentiable Architecture SearchabstractIn neural architecture search (NAS), differentiable architecture search (DARTS) has recently attracted much attention due to its high efficiency. However, this method finds a model with the weights converging faster than the others, and such a model with fastest convergence often leads to overfitting. Accordingly, the resulting model cannot always be well-generalized. To overcome this problem, we propose a method called minimum stable rank DARTS (MSR-DARTS), for finding a model with the best generalization error by replacing architecture optimization with the selection process using the minimum stable rank criterion. Specifically, a convolution operator is represented by a matrix, and MSR-DARTS selects the one with the smallest stable rank. We evaluated MSR-DARTS on CIFAR-10 and ImageNet datasets. It achieves an error rate of 2.54% with 4.0M parameters within 0.3 GPU-days on CIFAR-10, and a top-1 error rate of 23.9% on ImageNet. Kengo Machida, Kuniaki Uto, Koichi Shinoda, Taiji Suzuki |
IJCNN | 3 |
| 2021 | Multimodal Emotion Recognition with High-Level Speech and Text FeaturesabstractAutomatic emotion recognition is one of the central concerns of the Human-Computer Interaction field as it can bridge the gap between humans and machines. Current works train deep learning models on low-level data representations to solve the emotion recognition task. Since emotion datasets often have a limited amount of data, these approaches may suffer from overfitting, and they may learn based on superficial cues. To address these issues, we propose a novel cross-representation speech model, inspired by disentangle-ment representation learning, to perform emotion recognition on wav2vec 2.0 speech features. We also train a CNN-based model to recognize emotions from text features extracted with Transformer-based models. We further combine the speech-based and text-based results with a score fusion approach. Our method is evaluated on the IEMOCAP dataset in a 4-class classification problem, and it surpasses current works on speech-only, text-only, and multimodal emotion recognition. Mariana Rodrigues Makiuchi, Kuniaki Uto, Koichi Shinoda |
ASRU | 3 |
| 2020 | Estimation of Leaf Angle Distribution Based on Statistical Properties of Leaf Shading DistributionabstractLeaf angle distribution is an important phenotype parameter that is related to photosynthesis. Thanks to the recent advent of drones and high-resolution imaging devices, leaf-scale aerial images with high spectral and spatial resolution are available. This work is the first attempt to utilize a single leaf-scale image to differentiate plants with different leaf angle distribution. First, assuming that a rice leaf surface resembles a section of a hemiellipsoid surface, a collection of rice leaf surfaces is approximated by a hemiellipsoid surface. Time-series of shading distributions on the hemiellipsoids with different structural parameters under different direct sunlight directions are generated. By investigating the statistical properties, i.e., skewness, kurtosis and the most probable intensity, of the frequencies of the simulated shading intensity that well-differentiate hemiellipsoids with different structural parameters, we identified an appropriate time slot, i.e., 11: 00-12:30, for image acquisitions. Then, time-series leaf-scale images and depth maps of rice plants with/without silicate fertilizer under sunlight were collected. Based on the depth maps, it was confirmed that silicate fertilizer dosed leaves are more upright than leaves from non treated plants. It was demonstrated that 89% and 100% of kurtosis and the most probable intensity of the leaf-scale images during the appropriate time slot showed consistent relations with the simulations, which indicates that the proposed method is useful to distinguish different leaf angle distributions based on the frequency of shading intensity of rice leaf images. Kuniaki Uto, Mauro Dalla Mura, Yuka Sasaki, Koichi Shinoda |
IGARSS | 4 |
| 2020 | NEC-TT Speaker Verification System for SRE'19 CTS Challenge
Kong-Aik Lee, Koji Okabe, Hitoshi Yamamoto, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Keisuke Ishikawa, Koichi Shinoda |
INTERSPEECH | 9 |
| 2020 | NEC-TT System for Mixed-Bandwidth and Multi-Domain Speaker Recognition
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda |
Comput. Speech Lang. | 8 |
| 2019 | Sequence-level Knowledge Distillation for Model Compression of Attention-based Sequence-to-sequence Speech RecognitionabstractWe investigate the feasibility of sequence-level knowledge distillation of Sequence-to-Sequence (Seq2Seq) models for Large Vocabulary Continuous Speech Recognition (LVCSR). We first use a pre-trained larger teacher model to generate multiple hypotheses per utterance with beam search. With the same input, we then train the student model using these hypotheses generated from the teacher as pseudo labels in place of the original ground truth labels. We evaluate our proposed method using Wall Street Journal (WSJ) corpus. It achieved up to 9.8× parameter reduction with accuracy loss of up to 7.0% word-error rate (WER) increase. Raden Muaz, Nakamasa Inoue, Koichi Shinoda |
ICASSP | 3 |
| 2019 | Estimation of Diffuse Component of Global Radiation Based on Leaf-Scale Crop ImagesabstractAmong direct and diffuse components that compose photosynthetically active radiation (PAR), diffuse component of sunlight is important to evaluate fraction absorbed PAR (FAPAR) of plants because diffuse flux penetrates more deeply than direct flux and the peak of photosynthetic photon flux density (PPFD) use efficiency occurs at low to medium PPFD. Shading distribution in leaf-scale aerial images of plants by low-altitude measurement via UAVs contains sunlight information as well as plant structure and leaf pigments. In this work, we investigate the relationship between statistical properties of leaf-scale images, share of diffuse flux (SDF) in global radiation and solar zenith angles (SZA). Higher-order statistics (HOSs) were calculated from ground-based close-range images of wheat leaves under various sunlight conditions. SDF were measured based on two field spectrometers. We confirmed that (1) images under clear sky can be distinguished from those under cloudy sky based on HOSs, and (2) it is possible to estimate SZA based on HOSs of leaf-scale images under clear sky. Kuniaki Uto, Mauro Dalla Mura, Jocelyn Chanussot, Koichi Shinoda |
IGARSS | 4 |
| 2019 | The NEC-TT 2018 Speaker Verification System
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda |
INTERSPEECH | 8 |
| 2019 | A Modified Algorithm for Multiple Input Spectrogram InversionabstractWe propose a new algorithm to estimate the phase of speech signal in the mixture of audio sources under the assumption that the magnitude spectrum of each source is given. The pre- vious method, multiple input spectrogram inversion algorithm (MISI), often performs poorly when the magnitude spectro- grams estimated are not accurate. This may be because it im- poses a strict constraint that the summation of source wave- forms should be exactly the same as the mixture waveform. Our proposing algorithm employs a new objective function in which this constraint is relaxed. In this objective function, the difference between the summation of source waveforms and the mixture waveform is the target to be minimized. The perfor- mance of our method, modified MISI is evaluated on two dif- ferent experimental settings. In both settings it improves the audio source separation performance compared to MISI. Dongxiao Wang, Hirokazu Kameoka, Koichi Shinoda |
INTERSPEECH | 3 |
| 2019 | Recurrent out-of-vocabulary word detection based on distribution of features
Taichi Asami, Ryo Masumura, Yushi Aono, Koichi Shinoda |
Comput. Speech Lang. | 4 |
| 2018 | A Fine-to-Coarse Convolutional Neural Network for 3D Human Action Recognition
Thao Le Minh, Nakamasa Inoue, Koichi Shinoda |
BMVC | 3 |
| 2018 | Multi-Task Autoencoder for Noise-Robust Speech RecognitionabstractFor speech recognition in noisy environments, we propose a multi-task autoencoder which estimates not only clean speech features but also noise features from noisy speech. We introduce the deSpeeching autoencoder, which excludes speech signals from noisy speech, and combine it with the conventional denoising autoencoder to form a unified multi-task au-toencoder (MTAE). We evaluate it using the Aurora 2 dataset and CHIME 3 dataset. It reduced WER by 15.7% from the conventional denoising autoencoder in the Aurora 2 test set A. Haoyi Zhang, Conggui Liu, Nakamasa Inoue, Koichi Shinoda |
ICASSP | 4 |
| 2018 | Deep Learning Based Multi-modal Addressee Recognition in Visual Scenes with UtterancesabstractWith the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including those outdoors, in cafeterias, and hospitals. Because previous studies typically focused only on pre-specified tasks with limited conversational situations such as controlling smart homes, we created a mock dataset called Addressee Recognition in Visual Scenes with Utterances (ARVSU) that contains a vast body of image variations in visual scenes with an annotated utterance and a corresponding addressee for each scenario. We also propose a multi-modal deep-learning-based model that takes different human cues, specifically eye gazes and transcripts of an utterance corpus, into account to predict the conversational addressee from a specific speaker's view in various real-life conversational scenarios. To the best of our knowledge, we are the first to introduce an end-to-end deep learning model that combines vision and transcripts of utterance for addressee recognition. As a result, our study suggests that future addressee recognition can reach the ability to understand human intention in many social situations previously unexplored, and our modality dataset is a first step in promoting research in this field. Thao Le Minh, Nobuyuki Shimizu, Takashi Miyazaki, Koichi Shinoda |
IJCAI | 4 |
| 2018 | Attentive Statistics Pooling for Deep Speaker EmbeddingabstractThis paper proposes attentive statistics pooling for deep speaker embedding in text-independent speaker verification. In conventional speaker embedding, frame-level features are averaged over all the frames of a single utterance to form an utterance-level feature. Our method utilizes an attention mechanism to give different weights to different frames and generates not only weighted means but also weighted standard deviations. In this way, it can capture long-term variations in speaker characteristics more effectively. An evaluation on the NIST SRE 2012 and the VoxCeleb data sets shows that it reduces equal error rates (EERs) from the conventional method by 7.5% and 8.1%, respectively. Koji Okabe, Takafumi Koshinaka, Koichi Shinoda |
INTERSPEECH | 3 |
| 2018 | Detecting Alzheimer's Disease Using Gated Convolutional Neural Network from Audio DataabstractWe propose an automatic detection method of Alzheimer's diseases using a gated convolutional neural network (GCNN) from speech data. This GCNN can be trained with a relatively small amount of data and can capture the temporal information in audio paralinguistic features. Since it does not utilize any linguistic features, it can be easily applied to any languages. We evaluated our method using Pitt Corpus. The proposed method achieved the accuracy of 73.6%, which is better than the conventional sequential minimal optimization (SMO) by 7.6 points. Tifani Warnita, Nakamasa Inoue, Koichi Shinoda |
INTERSPEECH | 3 |
| 2018 | I-vector Transformation Using Conditional Generative Adversarial Networks for Short Utterance Speaker VerificationabstractI-vector based text-independent speaker verification (SV) systems often have poor performance with short utterances, as the biased phonetic distribution in a short utterance makes the extracted i-vector unreliable.This paper proposes an i-vector compensation method using a generative adversarial network (GAN), where its generator network is trained to generate a compensated i-vector from a short-utterance i-vector and its discriminator network is trained to determine whether an i-vector is generated by the generator or the one extracted from a long utterance.Additionally, we assign two other learning tasks to the GAN to stabilize its training and to make the generated ivector more speaker-specific.Speaker verification experiments on the NIST SRE 2008 "10sec-10sec" condition show that after applying our method, the equal error rate reduced by 11.3% from the conventional i-vector and PLDA system. Jiacen Zhang, Nakamasa Inoue, Koichi Shinoda |
INTERSPEECH | 3 |
| 2018 | Few-Shot Adaptation for Multimedia Semantic IndexingabstractWe propose a few-shot adaptation framework, which bridges zero-shot learning and supervised many-shot learning, for semantic indexing of image and video data. Few-shot adaptation provides robust parameter estimation with few training examples, by optimizing the parameters of zero-shot learning and supervised many-shot learning simultaneously. In this method, first we build a zero-shot detector, and then update it by using the few examples. Our experiments show the effectiveness of the proposed framework on three datasets: TRECVID Semantic Indexing 2010, 2014, and ImageNET. On the ImageNET dataset, we show that our method outperforms recent few-shot learning methods. On the TRECVID 2014 dataset, we achieve 15.19~% and 35.98~% in Mean Average Precision under the zero-shot condition and the supervised condition, respectively. To the best of our knowledge, these are the best results on this dataset. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 2 |
| 2017 | Boredom Recognition Based on Users' Spontaneous Behaviors in Multiparty Human-Robot Interactions
Yasuhiro Shibasaki, Kotaro Funakoshi, Koichi Shinoda |
MMM (1) | 3 |
| 2017 | Cross-view human action recognition from depth maps using spectral graph sequences
Tommi Kerola, Nakamasa Inoue, Koichi Shinoda |
Comput. Vis. Image Underst. | 3 |
| 2016 | Recurrent Out-of-Vocabulary Word Detection Using Distribution of FeaturesabstractThe repeated use of out-of-vocabulary (OOV) words in a spo-\nken document seriously degrades a speech recognizer’s perfor-\nmance. This paper provides a novel method for accurately de-\ntecting such recurrent OOV words. Standard OOV word de-\ntection methods classify each word segment into in-vocabulary\n(IV) or OOV. This word-by-word classification tends to be af-\nfected by sudden vocal irregularities in spontaneous speech,\ntriggering false alarms. To avoid this sensitivity to the irreg-\nularities, our proposal focuses on consistency of the repeated\noccurrence of OOV words. The proposed method preliminar-\nily detects recurrent segments, segments that contain the same\nword, in a spoken document by open vocabulary spoken term\ndiscovery using a phoneme recognizer. If the recurrent seg-\nments are OOV words, features for OOV detection in those\nsegments should exhibit consistency. We capture this consis-\ntency by using the mean and variance (distribution) of features\n(DOF) derived from the recurrent segments, and use the DOF\nfor IV/OOV classification. Experiments illustrate that the pro-\nposed method’s use of the DOF significantly improves its per-\nformance in recurrent OOV word detection.\nIndex Terms: speech recognition, OOV word detection, recur-\nrent OOV words, distribution of features Taichi Asami, Ryo Masumura, Yushi Aono, Koichi Shinoda |
INTERSPEECH | 4 |
| 2016 | Adaptation of Word Vectors using Tree Structure for Visual SemanticsabstractWe propose a framework of word-vector adaptation, which makes vectors of visually similar concepts close to each other. Here, word vectors are real-valued vector representation of words, e.g., word2vec representation. Our basic idea is to assume that each concept has some hypernyms that are important to determine its visual features. For example, for a concept Swallow with hypernyms Bird, Animal and Entity, we believe Bird is the most important since birds have common visual features with their feathers etc. Adapted word vectors are obtained for each word by taking a weighted sum of a given original word vector and its hypernym word vectors. Our weight optimization makes vectors of visually similar concepts close to each other, by giving a large weight for such important hypernyms. We apply the adapted word vectors to zero-shot learning on the TRECVID 2014 semantic indexing dataset. We achieved 0.083 of Mean Average Precision, which is the best performance without using TRECVID training data to the best of our knowledge. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 2 |
| 2016 | Robust discriminative training against data insufficiency in PLDA-based speaker verification
Johan Rohdin, Sangeeta Biswas, Koichi Shinoda |
Comput. Speech Lang. | 3 |
| 2016 | Fast Coding of Feature Vectors Using Neighbor-to-Neighbor SearchabstractSearching for matches to high-dimensional vectors using hard/soft vector quantization is the most computationally expensive part of various computer vision algorithms including the bag of visual word (BoW). This paper proposes a fast computation method, Neighbor-to-Neighbor (NTN) search [1] , which skips some calculations based on the similarity of input vectors. For example, in image classification using dense SIFT descriptors, the NTN search seeks similar descriptors from a point on a grid to an adjacent point. Applications of the NTN search to vector quantization, a Gaussian mixture model, sparse coding, and a kernel codebook for extracting image or video representation are presented in this paper. We evaluated the proposed method on image and video benchmarks: the PASCAL VOC 2007 Classification Challenge and the TRECVID 2010 Semantic Indexing Task. NTN-VQ reduced the coding cost by 77.4 percent, and NTN-GMM reduced it by 89.3 percent, without any significant degradation in classification performance. Nakamasa Inoue, Koichi Shinoda |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Vocabulary Expansion Using Word Vectors for Video Semantic IndexingabstractWe propose vocabulary expansion for video semantic indexing. From many semantic concept detectors obtained by using training data, we make detectors for concepts not included in training data. First, we introduce Mikolov's word vectors to represent a word by a low-dimensional vector. Second, we represent a new concept by a weighted sum of concepts in training data in the word vector space. Finally, we use the same weighting coefficients for combining detectors to make a new detector. In our experiments, we evaluate our methods on the TRECVID Video Semantic Indexing (SIN) Task. We train our models with Google News text documents and ImageNET images to generate new semantic detectors for SIN task. We show that our method performs as well as SVMs trained with 100 TRECVID ex- ample videos. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 2 |
| 2015 | Autonomous selection of i-vectors for PLDA modelling in speaker verification
Sangeeta Biswas, Johan Rohdin, Koichi Shinoda |
Speech Commun. | 3 |
| 2014 | Spectral Graph Skeletons for 3D Action Recognition
Tommi Kerola, Nakamasa Inoue, Koichi Shinoda |
ACCV (4) | 3 |
| 2014 | Constrained discriminative PLDA training for speaker verificationabstractMany studies have proven the effectiveness of discriminative training for speaker verification based on probabilistic linear discriminative analysis (PLDA) with i-vectors as features. Most of them directly optimize the log-likelihood ratio score function of the PLDA model instead of explicitly train the PLDA model. But this optimization process removes some of the constraints that normally are imposed on the PLDA log likelihood ratio score function. This may deteriorate the verification performance when the amount of training data is limited. In this paper, we first show two constraints which the score function should follow, and then we propose a new constrained discriminative training algorithm which keeps these constraints. Our experiments show that our method obtained significant improvements in the verification performance in the male trials of the telephone speaker verification tasks of NIST SRE08 and SRE10. Johan Rohdin, Sangeeta Biswas, Koichi Shinoda |
ICASSP | 3 |
| 2014 | Simple gesture-based error correction interface for smartphone speech recognitionabstractConventional error correction interfaces for speech recog-
nition require a user to first mark an error region and choose
the correct word from a candidate list. Taking the user’s effort
and the limited user interface available in a smartphone use into
account, this operation should be simpler. In this paper, we pro-
pose an interface where users mark the error region once, and
then the word will be replaced by another candidate. Assuming
that the words preceding/succeeding the error region are vali-
dated by the user, we search the Web n-grams for long word
sequences matched to such a context. The acoustic features of
the error region are also utilized to rerank the candidate words.
The experimental result proved the effectiveness of our method.
30.2% of the error words were corrected by a single operation. Koji Iwano, Koichi Shinoda |
INTERSPEECH | 3 |
| 2014 | n-gram Models for Video Semantic IndexingabstractWe propose n-gram modeling of shot sequences for video semantic indexing, in which semantic concepts are extracted from a video shot. Most previous studies for this task have assumed that video shots in a video clip are independent from each other. We model the time-dependency between them assuming that n-consecutive video shots are dependent. Our models improve the robustness against occlusion and camera-angle changes by effectively using information from the previous video shots. In our experiments on the TRECVID 2012 Semantic Indexing Benchmark, we applied the proposed models to a system using Gaussian mixture models and support vector machines. Mean average precision was improved from 30.62% to 32.14%, which is the best performance on the TRECVID 2012 Semantic Indexing to the best of our knowledge. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 2 |
| 2014 | Event Detection by Velocity Pyramid
Zhuolin Liang, Nakamasa Inoue, Koichi Shinoda |
MMM (1) | 3 |
| 2014 | An efficient error correction interface for speech recognition on mobile touchscreen devicesabstractCorrecting speech recognition errors on a mobile touchscreen device is an unavoidable but time-consuming task that requires a lot of user effort. To reduce this user effort, we previously proposed an error correction method using long context match with Web N-gram, which we combined with a simple gesture-based user interface. This method automatically replaces an error word with its corresponding correct word. However, it was evaluated only substitution errors in sentences, each of which involves only one error. In this paper, we extend this method to be used for more general cases when a sentence has more than one error. It recovers not only substitution errors but also deletion errors and insertion errors. For recovering deletion errors, it predicts a deleted word based on the phonemes and the part-of-speech tags of its surrounding words. Our experimental results show that the proposed method recovered the errors more accurately with less user effort than the conventional Word Confusion Network based error correction interface. Koji Iwano, Koichi Shinoda |
SLT | 3 |
| 2014 | Speaker adaptation of deep neural networks using a hierarchy of output layersabstractDeep neural networks (DNN) used for acoustic modeling in speech recognition often have a very large number of output units corresponding to context dependent (CD) triphone HMM states. The amount of data available for speaker adaptation is often limited so a large majority of these CD states may not be observed during adaptation. In this case, the posterior probabilities of unseen CD states are only pushed towards zero during DNN speaker adaptation and the ability to predict these states can be degraded relative to the speaker independent network. We address this problem by appending an additional output layer which maps the original set of DNN output classes to a smaller set of phonetic classes (e.g. monophones) thereby reducing the occurrences of unseen states in the adaptation data. Adaptation proceeds by backpropagation of errors from the new output layer, which is disregarded at recognition time when posterior probabilities over the original set of CD states are used. We demonstrate the benefits of this approach over adapting the network with the original set of CD states using experiments on a Japanese voice search task and obtain 5.03% relative reduction in character error rate with approximately 60 seconds of adaptation data. Ryan Price, Ken-ichi Iso, Koichi Shinoda |
SLT | 3 |
| 2013 | Neighbor-to-Neighbor Search for Fast Coding of Feature VectorsabstractAssigning a visual code to a low-level image descriptor, which we call code assignment, is the most computationally expensive part of image classification algorithms based on the bag of visual word (BoW) framework. This paper proposes a fast computation method, Neighbor-to-Neighbor (NTN) search, for this code assignment. Based on the fact that image features from an adjacent region are usually similar to each other, this algorithm effectively reduces the cost of calculating the distance between a codeword and a feature vector. This method can be applied not only to a hard codebook constructed by vector quantization (NTN-VQ), but also to a soft codebook, a Gaussian mixture model (NTN-GMM). We evaluated this method on the PASCAL VOC 2007 classification challenge task. NTN-VQ reduced the assignment cost by 77.4% in super-vector coding, and NTN-GMM reduced it by 89.3% in Fisher-vector coding, without any significant degradation in classification performance. Nakamasa Inoue, Koichi Shinoda |
ICCV | 2 |
| 2013 | Combining deep speaker specific representations with GMM-SVM for speaker verificationabstractThis study combines a Gaussian mixture model support vector machine (GMM-SVM) system with a nonlinear feature transformation, discriminatively trained to extract speaker specific features from MFCCs. Separation of the speaker information component and non-speaker related information in the speech signal is accomplished using a regularized siamese deep network (RSDN). RSDN learns a hidden representation that well characterizes speaker information by training a subset of the hidden units using pairs of speech segments. MFCC features are input to a trained RSDN and a subset of hidden layer outputs are used as new input features in a GMM-SVM system. We demonstrate the potential of this approach for text-independent speaker verification by applying it to a subset of the NIST SRE 2006 1conv4w-1conv4w task. The hybrid RSDN GMM-SVM system achieves about 5% relative improvement over the baseline GMM-SVM system. Index Terms: speaker verification, neural networks, feature extraction, GMM-SVM Ryan Price, Sangeeta Biswas, Koichi Shinoda |
INTERSPEECH | 3 |
| 2013 | q-Gaussian mixture models for image and video semantic indexing
Nakamasa Inoue, Koichi Shinoda |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Feature normalization based on non-extensive statistics for speech recognition
Hilman Ferdinandus Pardede, Koji Iwano, Koichi Shinoda |
Speech Commun. | 3 |
| 2013 | Detection of overlapped speech using lapel microphones in meeting
Ryo Yokoyama, Yu Nasu, Koji Iwano, Koichi Shinoda |
Speech Commun. | 4 |
| 2012 | q-Gaussian Mixture Models Based on Non-extensive Statistics for Image and Video Semantic Indexing
Nakamasa Inoue, Koichi Shinoda |
ACCV (2) | 2 |
| 2012 | Multimedia event detection using GMM supervectors and SVMSabstractIn multimedia event detection, complex target events are extracted from a large set of consumer-generated videos taken in unconstrained environments. We devised a multimedia event detection method based on GMM supervectors and support vector machines (SVMs) using multiple features. A GMM supervector consists of the parameters of a Gaussian mixture model (GMM) for the distribution of local features extracted from a video clip. A GMM is regarded as an extension of the Bag-of-Words (BoW) to a probabilistic framework, and thus, it can be expected to be robust against the data insufficiency problem. This method outperformed previous methods including BoW in experiments using the dataset of the multimedia event detection task in TRECVID2010 and 2011. Yusuke Kamishima, Nakamasa Inoue, Koichi Shinoda, Shunsuke Sato |
ICIP | 3 |
| 2012 | Q-Gaussian based spectral subtraction for robust speech recognitionabstractSpectral subtraction (SS) is derived using maximum likelihood estimation assuming both noise and speech follow Gaussian distributions and are independent from each other.Under this assumption, noisy speech, speech contaminated by noise, also follows a Gaussian distribution.However, it is well known that noisy speech observed in real situations often follows a heavytailed distribution, not a Gaussian distribution.In this paper, we introduce a q-Gaussian distribution in non-extensive statistics to represent the distribution of noisy speech and derive a new spectral subtraction method based on it.In our analysis, the q-Gaussian distribution fits the noisy speech distribution better than the Gaussian distribution does.Our speech recognition experiments showed that the proposed method, q-spectral subtraction (q-SS), outperformed the conventional SS method using the Aurora-2 database. Hilman Ferdinandus Pardede, Koichi Shinoda, Koji Iwano |
INTERSPEECH | 2 |
| 2012 | Overlapped Speech Detection in Meeting Using Cross-Channel Spectral Subtraction and Spectrum SimilarityabstractWe propose an overlapped speech detection method for speech recognition and speaker diarization of meetings, where each speaker wears a lapel microphone.Two novel features are utilized as inputs for a GMM-based detector.One is speech power after cross-channel spectral subtraction which reduces the power from the other speakers.The other is an amplitude spectral cosine correlation coefficient which effectively extracts the correlation of spectral components in a rather quiet condition.We evaluated our method using a meeting speech corpus of four speakers.The accuracy of our proposed method, 74.1%, was significantly better than that of the conventional method, 67.0%, which uses raw speech power and power spectral Pearson's correlation coefficient. Ryo Yokoyama, Yu Nasu, Koichi Shinoda, Koji Iwano |
INTERSPEECH | 3 |
| 2012 | A Fast and Accurate Video Semantic-Indexing System Using Fast MAP Adaptation and GMM SupervectorsabstractWe propose a fast maximum a posteriori (MAP) adaptation method for video semantic indexing that uses Gaussian mixture model (GMM) supervectors. In this method, a tree-structured GMM is utilzed to decrease the computational cost, where only the output probabilities of mixture components close to an input sample are precisely calculated. Experimental evaluation on the TRECVID 2010 dataset demonstrates the effectiveness of the proposed method. The calculation time of the MAP adaptation step is reduced by 76.2% compared with that of a conventional method. The total calculation time is reduced by 56.6% while keeping the same level of the accuracy. Nakamasa Inoue, Koichi Shinoda |
IEEE Trans. Multim. | 2 |
| 2011 | Designing text corpus using phone-error distribution for acoustic modelingabstractIt is expensive to prepare a sufficient amount of training data for acoustic modeling for developing large vocabulary continuous speech recognition systems. This is a serious problem especially for resource-deficient languages. We propose an active learning method that effectively reduces the amount of training data without any degradation in recognition performance. It is used to design a text corpus for read speech collection. It first estimates phone-error distribution using a small amount of fully transcribed speech data. Second, it constructs a sentence set whose phone-occurrence distribution is close to the phone-error distribution and collects its speech data. It then extends this process to diphones and triphones and collects more speech data. We evaluated our method with simulation experiments using the Corpus of Spontaneous Japanese. It required only 76 h of speech data to achieve word accuracy of 74.7%, while the conventional training method required 152 h of data to achieve the same rate. Hiroko Murakami, Koichi Shinoda, Sadaoki Furui |
ASRU | 2 |
| 2011 | Structural MAP adaptation in GMM-supervector based speaker recognitionabstractIn recent years, adaptation techniques have been given special focus in speaker recognition tasks, mainly targeting speaker and session variation disentangling under the Maximum a Posteriori (MAP) criterion. For these techniques, unseen mixtures are usually adapted in a global manner, if ever. In this paper, we explore Structural MAP (SMAP), Maximum a Posteriori adaptation using hierarchical structures of the acoustic space that allow data scarceness issues to be tackled with different precision levels. We explore this approach in a speaker verification system using a Support Vector Machine (SVM) classifier and Gaussian mean supervectors (GMM-SVM). We show that this is an effective approach that considerably outperforms its relevance MAP counterpart in the 2006 NIST Speaker Recognition Evaluation. We also show that using a speaker-adapted Universal Background Model can improve the stability of the clustering algorithm besides obtaining further improvements. Marc Ferras, Koichi Shinoda, Sadaoki Furui |
ICASSP | 2 |
| 2011 | Cross-Channel Spectral Subtraction for meeting speech recognitionabstractWe propose Cross-Channel Spectral Subtraction (CCSS), a source separation method for recognizing meeting speech where one microphone is prepared for each speaker. The method quickly adapts to changes in transfer functions and uses spectral subtraction to suppress the speech of other speakers. Compared with conventional source separation methods based on independent component analysis (ICA) or that use binary masks, it requires less computational costs and the resulting speech signals have less distortion. In a recognition task of computer-simulated, partially-overlapped speech, CCSS improved the word accuracy from 66.5% to 77.7%. It also significantly improved the recognition accuracy of speech data in actual meetings. Yu Nasu, Koichi Shinoda, Sadaoki Furui |
ICASSP | 2 |
| 2011 | Acoustic Forest for SMAP-Based Speaker VerificationabstractIn speaker verification, structural maximum-a-posteriori (SMAP) adaptation for Gaussian mixture model (GMM) has been proven effective, especially when the speech segment is very short.In SMAP adaptation, an acoustic tree of Gaussian components is constructed to represent the hierarchical acoustic space.Until now, however, there has been no clear way to automatically find the optimal tree structure for a given speaker.In this paper, we propose using an acoustic forest, which is a set of trees, for SMAP adaptation, instead of a single tree.In this approach, we combine the results of SMAP adaptation systems with different acoustic trees.A key issue is how to combine the trees.We explore three score fusion techniques, and evaluate our approach in the text-independent speaker verification task of the NIST 2006 SRE plan using 10-second speech segments.Our proposed method decreased EER by 3.2% from the relevant MAP adaptation and by 1.6% from the conventional SMAP with a single tree. Sangeeta Biswas, Marc Ferras, Koichi Shinoda, Sadaoki Furui |
INTERSPEECH | 3 |
| 2011 | Structural Joint Factor Analysis for Speaker Recognition
Marc Ferras, Koichi Shinoda, Sadaoki Furui |
INTERSPEECH | 2 |
| 2011 | Generalized-Log Spectral Mean Normalization for Speech RecognitionabstractMost compensation methods for robust speech recognition against noise assume independency between speech, additive and convolutive noise.However, the nonlinear nature distortion caused by noise may introduce correlation between noise and speech.To tackle this issue, we propose generalized-log spectral mean normalization (GLSMN) in which log spectral mean normalization (LSMN) is carried out in the q-logarithmic domain.Experiments on the Aurora-2 database show that GLSMN improved speech recognition accuracies by 20% compared to cepstral mean normalization (CMN) in mel-frequency domain. Hilman Ferdinandus Pardede, Koichi Shinoda |
INTERSPEECH | 2 |
| 2011 | A fast MAP adaptation technique for gmm-supervector-based video semantic indexing systemsabstractWe propose a fast maximum a posteriori (MAP) adaptation technique for a GMM-supervectors-based video semantic indexing system.The use of GMM supervectors is one of the state-of-the-art methods in which MAP adaptation is needed for estimating the distribution of local features extracted from video data. The proposed method cuts the calculation time of the MAP adaptation step. With the proposed method, a tree-structured GMM is constructed to quickly calculate posterior probabilities for each mixture component of a GMM. The basic idea of the tree-structured GMM is to cluster Gaussian components and approximate them with a single Gaussian. Leaf nodes of the tree correspond to the mixture components, and each non-leaf node has a single Gaussian that approximates its descendant Gaussian distributions. Experimental evaluation on the TRECVID 2010 dataset demonstrates the effectiveness of the proposed method. The calculation time of the MAP adaptation step is reduced by 76.2% compared to that of a conventional method and resulting accuracy (in terms of Mean average precision) was 10.2%. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 2 |
| 2011 | Semi-synchronous speech and pen input for mobile user interfaces
Koichi Shinoda, Yasushi Watanabe, Kenji Iwata, Ryuta Nakagawa, Sadaoki Furui |
Speech Commun. | 1 |
| 2010 | Speech modeling based on committee-based active learningabstractWe propose a committee-based active learning method for large vocabulary continuous speech recognition. In this approach, multiple recognizers are prepared beforehand, and the recognition results obtained from them are used for selecting utterances. Here, a progressive search method is used for aligning sentences, and voting entropy is used as a measure for selecting utterances. We apply our method not only to acoustic models but also to language models and their combination. Our method was evaluated by using 190-hour speech data in the Corpus of Spontaneous Japanese. It proved to be significantly better than random selection. It only required 63 h of data to achieve a word accuracy of 74%, while standard training (i.e., random selection) required 97 h of data. The recognition accuracy of our proposed method was also better than that of the conventional uncertainty sampling method using word posterior probabilities as the confidence measure for selecting sentences. Yuzo Hamanaka, Koichi Shinoda, Sadaoki Furui, Tadashi Emori, Takafumi Koshinaka |
ICASSP | 2 |
| 2010 | Robust Gait Recognition Against Speed VariationabstractVariations in walking speed have a strong impact on the recognition of gait. We propose a method of recognition of gait that is robust against walking-speed variations. It is established on a combination of Fisher discriminant analysis (FDA)-based cubic higher-order local auto-correlation (CHLAC) and the statistical framework provided by hidden Markov models (HMMs). The HMMs in this method identify the phase of each gait even when walking speed changes nonlinearly, and the CHLAC features capture the within-phase spatio-temporal characteristics of each individual. We compared the performance of our method with other conventional methods in our evaluation using three different databases, i.e., USH, USF-NIST, and Tokyo Tech DB. Ours was equal or better than the others when the speed did not change too much, and was significantly better when the speed varied across and within a gait sequence. Muhammad Rasyid Aqmar, Koichi Shinoda, Sadaoki Furui |
ICPR | 2 |
| 2010 | High-Level Feature Extraction Using SIFT GMMs and Audio ModelsabstractWe propose a statistical framework for high-level feature extraction that uses SIFT Gaussian mixture models (GMMs) and audio models. SIFT features were extracted from all the image frames and modeled by a GMM. In addition, we used mel-frequency cepstral coefficients and ergodic hidden Markov models to detect high-level features in audio streams. The best result obtained by using SIFT GMMs in terms of mean average precision on the TRECVID 2009 corpus was 0.150 and was improved to 0.164 by using audio information. Nakamasa Inoue, Tatsuhiko Saito, Koichi Shinoda, Sadaoki Furui |
ICPR | 3 |
| 2010 | Dynamic language model adaptation using keyword category classification
Hitoshi Yamamoto, Ken Hanazawa, Kiyokazu Miki, Koichi Shinoda |
INTERSPEECH | 4 |
| 2009 | Independent component analysis for noisy speech recognitionabstractIndependent component analysis (ICA) is not only popular for blind source separation but also for unsupervised learning when the observations can be decomposed into some independent components. These components represent the specific speaker, gender, accent, noise or environment, and act as the basis functions to span the vector space of the human voices in different conditions. Different from eigenvoices built by principal component analysis, the proposed independent voices are estimated by ICA algorithm, and are applied for efficient coding of an adapted acoustic model. Since the information redundancy is significantly reduced in independent voices, we effectively calculate a coordinate vector in independent voice space, and estimate the hidden Markov models (HMMs) for speech recognition. In the experiments, we build independent voices from HMMs under different noise conditions, and find that these voices attain larger redundancy reduction than eigenvoices. The noise adaptive HMMs generated by independent voices achieve better recognition performance than those by eigenvoices. Hsin-Lung Hsieh, Jen-Tzung Chien, Koichi Shinoda, Sadaoki Furui |
ICASSP | 3 |
| 2009 | Online speaker clustering using incremental learning of an ergodic hidden Markov modelabstractA novel online speaker clustering method suitable for real-time applications is proposed. Using an ergodic hidden Markov model, it employs incremental learning based on a variational Bayesian framework and provides probabilistic (non-deterministic) decisions for each input utterance, directly considering the specific history of preceding utterances. It makes possible more robust cluster estimation and precise classification of utterances than do conventional online methods. Experiments on meeting-speech data show that the proposed method produces 70-80% fewer errors than a conventional method does. Takafumi Koshinaka, Kentaro Nagatomo, Koichi Shinoda |
ICASSP | 3 |
| 2009 | Speaker adaptation based on two-step active learningabstractWe propose a two-step active learning method for supervised speaker adaptation. In the first step, the initial adaptation data is collected to obtain a phone error distribution. In the second step, those sentences whose phone distributions are close to the error distribution are selected, and their utterances are collected as the additional adaptation data. We evaluated the method us- ing a Japanese speech database and maximum likelihood linear regression (MLLR) as the speaker adaptation algorithm. We confirmed that our method had a significant improvement over a method using randomly chosen sentences for adaptation. Index Terms: speech recognition ,speaker adaptation, active learning nition accuracies in comparison with the conventional adapta- tion frameworks when the initial speaker-independent acoustic model has already been tuned to the given recognition task. In this paper, we propose a two-step active learning method for supervised speaker adaptation. In the first step, our method collects a small amount of utterances from a user to obtain his/her tendency in speech recognition errors. In the second step, it selects those sentences rich in phonetic units in the er- rors from a sentence pool, and it collects their utterances as ad- ditional data for adaptation. Since our method directly aims at decreasing recognition errors, it is expected to be highly dis- criminative. Our method has two critical issues: One is how to relate the recognition errors to the selection criterion of adap- tation sentences. The other is how to set the size of the initial adaptation data in the first step; it should be as small as possi- ble in order to decrease the user's effort, but it must be sufficient enough to estimatethe tendency of errors precisely. Wedescribe our approach to resolving these issues in this paper. This paper is organized as follows. Section 2 explains our method, and Section 3 briefly explains the MLLR speaker adap- tation method. Section 4 reports our evaluation experiments us- ing a Japanese speech database, and Section 5 concludes the paper. Koichi Shinoda, Hiroko Murakami, Sadaoki Furui |
INTERSPEECH | 1 |
| 2008 | Robust spoken term detection using combination of phone-based and word-based recognitionabstractWe propose a robust spoken term detection method against word recognition errors using a combination of phone-based and word-based recognition.Conventional methods based on similar frameworks are problematic because phone-based recognition produces a large number of insertion errors.In our method, different substitution penalties are assigned for phone pairs to reduce such errors.We evaluated our method using the corpus of spontaneous Japanese.When recall was fixed at 50%, precision improved to 4.4 points above detection using only word-based recognition.We also report here on the effectiveness of optimization of the combination weight for each keyword. Kenji Iwata, Koichi Shinoda, Sadaoki Furui |
INTERSPEECH | 2 |
| 2008 | Improvement of eigenvoice-based speaker adaptation by parameter space clusteringabstractThe segmental eigenvoice method has been proposed to provide rapid speaker adaptation with limited amounts of adaptation data.In this method, the speaker-vector space is clustered to several subspaces and PCA is applied to each of the resulting subspaces.In this paper, we propose two new techniques to improve the performance of this segmental eigenvoice approach.First, we propose a soft-clustering method in which each element in a speaker vector can be assigned to more than one cluster.Second, those elements far apart from any of the clusters are removed.Our experiments using the JNAS and S-JNAS databases show that the proposed method outperforms both the original eigenvoice and the segmental eigenvoice methods, e.g., 3.3% average improvement when only 10 utterances are used for adaptation. Shutaro Tanji, Koichi Shinoda, Sadaoki Furui, Antonio Ortega |
INTERSPEECH | 2 |
| 2008 | Time-lag adaptation for semi-synchronous speech and pen inputabstractIn a previous study, we developed an interface using semisynchronous speech and pen input.In this interface, a user speaks while writing, and the pen input complements the speech, enabling a higher recognition performance than with speech alone.When a user inputs speech and pen, there is a time lag between the two modes, and the lag differs among users.We propose a method for adapting to the different time lags of individual users.This method was evaluated in a Japanese continuous speech recognition task with three different pen-input interfaces including a QWERTY keyboard interface.The time-lag adaptation improved recognition accuracies by up to 0.5 point. Yasushi Watanabe, Koichi Shinoda, Sadaoki Furui |
INTERSPEECH | 2 |
| 2007 | Speech Recognition using FHMMS Robust Against Nonstationary NoiseabstractWe focus on the problem of speech recognition in the presence of nonstationary sudden noise, which is very likely to happen in home environments. As a model compensation method for this problem, we investigated the use of factorial hidden Markov model (FHMM) architecture developed from a clean-speech hidden Markov model (HMM) and a sudden-noise HMM. While in conventional studies this architecture is defined only for static features of the observation vector, we extended it to dynamic features. A database recorded by a personal robot called PaPeRo in home environments was used for the evaluation of the proposed method under noisy conditions. While we presented a recognition system using isolated-word FHMMs in our previous work, here we evaluated the effectiveness of the phoneme FHMMs. Agnieszka Betkowska Cavalcante, Koichi Shinoda, Sadaoki Furui |
ICASSP (4) | 2 |
| 2007 | Semi-Synchronous Speech and Pen InputabstractThis paper proposes a new interface method using semi-synchronous speech and pen input for mobile environments. In this interface, a user speaks while writing, where pen input complements speech to achieve higher recognition performance than speech alone. A multimodal recognition algorithm that can handle the asynchronicity of the two modes using a segment-based unification scheme is proposed. This method is evaluated under noisy conditions with four different pen-input interfaces: character, stroke, pen-touch, and point-to-character, each of which is assumed to be given for a phrase unit in speech. It is confirmed that the recognition accuracy is improved by the proposed method in comparison with that by speech alone in all the pen-input conditions. Yasushi Watanabe, Kenji Iwata, Ryuta Nakagawa, Koichi Shinoda, Sadaoki Furui |
ICASSP (4) | 4 |
| 2007 | Predictive minimum Bayes risk classification for robust speech recognitionabstractThis paper presents a new Bayes classification rule towards minimizing the predictive Bayes risk for robust speech recognition. Conventionally, the plug-in maximum a posteriori (MAP) classification is constructed by adopting nonparametric loss function and deterministic model parameters. Speech recognition performance is limited due to the environmental mismatch and the ill-posed model. Concerning these issues, we develop the predictive minimum Bayes risk (PMBR) classification where the predictive distributions are inherent in Bayes risk. More specifically, we exploit the Bayes loss function and the predictive word posterior probability for Bayes classification. Model mismatch and randomness are compensated to improve generalization capability in speech recognition. In the experiments on car speech recognition, we estimate the prior densities of hidden Markov model parameters from adaptation data. With the prior knowledge of new environment and model uncertainty, PMBR classification is realized and evaluated to be better than MAP, MBR and Bayesian predictive classification. Jen-Tzung Chien, Koichi Shinoda, Sadaoki Furui |
INTERSPEECH | 2 |
| 2007 | Automatic estimation of scaling factors among probabilistic models in speech recognitionabstractWe propose an efficient new method for estimating scaling factors among probabilistic models in speech recognition. Most speech recognition systems consist of an acoustic and a language model, and require scaling factors to balance probabilities among them. The scaling factors are conventionally optimized in recognition tests. In our proposed method, the scaling factors are regarded as parameters of a log-linear model, and they are estimated using a gradient-ascent method based on the maximum a posteriori probability criterion. Posterior probability is computed using word-lattices. We employ an iteration technique which repeats a word-lattice-generation/scalingfactor-estimation process, and the resulting scaling factor estimation is robust with respect to the changes in initial values. In experiments, estimated scaling factors were nearly identical to optimal values obtained in a greedy grid search, and they changed little with variations in initial values. Index Terms: speech recognition, scaling factor, log-linear model, word lattice Tadashi Emori, Yoshifumi Onishi, Koichi Shinoda |
INTERSPEECH | 3 |
| 2007 | Dynamic language model adaptation using presentation slides for lecture speech recognitionabstractWe propose a dynamic language model adaptation method that uses the temporal information from lecture slides for lecture speech recognition. The proposed method consists of two steps. First, the language model is adapted with the text information extracted from all the slides of a given lecture. Next, the text information of a given slide is extracted based on temporal information and used for local adaptation. Hence, the language model, used to recognize speech associated with the given slide changes dynamically from one slide to the next. We evaluated the proposed method with the speech data from four Japanese lecture courses. Our experiments show the effectiveness of our Hiroki Yamazaki, Koji Iwano, Koichi Shinoda, Sadaoki Furui, Haruo Yokota |
INTERSPEECH | 3 |
| 2006 | Towards Optimal Bayes Decision for Speech RecognitionabstractThis paper presents a new speech recognition framework towards fulfilling optimal Bayes decision theory, which is essential for general pattern recognition. The recognition procedure is developed through minimizing the Bayes risk, or equivalently the expected loss due to classification action. Typically, loss function measures the penalty/evidence of choosing a candidate hypothesis. This function was manually specified or empirically calculated. Here, we exploit a novel Bayes loss function via testing the hypotheses whether the classification action produces loss or not. A Bayes factor is derived to measure loss in a statistical and meaningful way. Attractively, Bayes loss function using predictive distributions is robust to the uncertainty of environments. Also, optimizing this Bayes criterion equals to minimizing classification errors of test data. The relation between the minimum classification error (MCE) classifier and the proposed optimal Bayes classifier (OBC) is bridged. Specifically, the logarithm of Bayes factor in OBC is analogous to the misclassification measure in MCE when using predictive distribution as the discriminant function. We accordingly build a robust and discriminative classification for large vocabulary continuous speech recognition. In the experiments on broadcast news transcription, the new OBC rule significantly outperforms traditional maximum a posteriori classification. Jen-Tzung Chien, Chih-Hsien Huang, Koichi Shinoda, Sadaoki Furui |
ICASSP (1) | 3 |
| 2005 | Robust highlight extraction using multi-stream hidden Markov models for baseball videoabstractThis paper proposes a robust statistical framework to extract highlights from a baseball broadcast video. We applied multi-stream hidden Markov models (HMMs) to control the weights among different features. To achieve robustness against new highlights, we used a common simple structure for all the HMMs. In addition, scene segmentation and unsupervised adaptation were applied to achieve more robustness against the differences of environmental conditions among games. The precision rate of high-light extracting experiments for eight kinds of highlights from 4.5 hours of digest data was 77.4% and was increased to 78.7% by applying scene segmentation. Furthermore, the unsupervised adaptation method improved precision by 2.7 points to 81.4%. These results confirm the effectiveness of our framework. Nguyen Huu Bach, Koichi Shinoda, Sadaoki Furui |
ICIP (3) | 2 |
| 2002 | Efficient reduction of Gaussian components using MDL criterion for HMM-based speech recognitionabstractA method is proposed to reduce the number of Gaussian components in continuous density hidden Markov models (HMMs). As its initial model, the method employs a well-trained, large-sized HMM in which the components of each state's Gaussian mixture probability density function are clustered into a binary tree. For each state, a subset of Gaussian components is chosen from the Gaussian tree on the basis of the minimum description length (MDL) criterion. By varying the penalty coefficient for large size models in the MDL criterion, it is possible to obtain the total number of Gaussian components desired for smaller models. In our experimental evaluations, the proposed method successfully reduced the number of Gaussian components by 75%, with only 1% degradation in recognition accuracy. Koichi Shinoda, Ken-ichi Iso |
ICASSP | 1 |
| 2001 | Rapid vocal tract length normalization using maximum likelihood estimationabstractRecently, vocal tract length normalization (VTLN) techniques have been developed for speaker normalization in speech recognition.This paper proposes a new VTLN method, in which the vocal tract length is normalized in the cepstrum space by means of linear mapping whose parameter is derived using maximumlikelihood estimation.The computational costs of this method are much lower than that of such conventional methods as ML-VTLN, in which the parameter for mapping is selected from among several parameters.Further, the new method offers greater precision in determining parameters for individual speakers.Experimental use of the method resulted in an error reduction rate of 7.1%.A combination of the proposed method with cepstrum mean normalization (CMN) method was also examined and found to reduce the error rate even more, by 14.6%. Tadashi Emori, Koichi Shinoda |
INTERSPEECH | 2 |
| 2001 | A structural Bayes approach to speaker adaptationabstractMaximum a posteriori (MAP) estimation has been successfully applied to speaker adaptation in speech recognition systems using hidden Markov models. When the amount of data is sufficiently large, MAP estimation yields recognition performance as good as that obtained using maximum-likelihood (ML) estimation. This paper describes a structural maximum a posteriori (SMAP) approach to improve the MAP estimates obtained when the amount of adaptation data is small. A hierarchical structure in the model parameter space is assumed and the probability density functions for model parameters at one level are used as priors for those of the parameters at adjacent levels. Results of supervised adaptation experiments using nonnative speakers' utterances showed that SMAP estimation reduced error rates by 61% when ten utterances were used for adaptation and that it yielded the same accuracy as MAP and ML estimation when the amount of data was sufficiently large. Furthermore, the recognition results obtained in unsupervised adaptation experiments showed that SMAP estimation was effective even when only one utterance from a new speaker was used for adaptation. An effective way to combine rapid supervised adaptation and on-line unsupervised adaptation was also investigated. Koichi Shinoda |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | A family of Hadamard matrices of dihedral group type
Koichi Shinoda, Mieko Yamada |
Discret. Appl. Math. | 1 |
| 1998 | Unsupervised adaptation using structural Bayes approachabstractIt is well-known that the performance of recognition systems is often largely degraded when there is a mismatch between the training and testing environment. It is desirable to compensate for the mismatch when the system is in operation without any supervised learning. Previously, a structural maximum a posteriori (SMAP) adaptation approach, in which a hierarchical structure in the parameter space is assumed, was proposed. In this paper, this SMAP method is applied to unsupervised adaptation. A novel normalization technique is also introduced as a front end for the adaptation process. The recognition results showed that the proposed method was effective even when only one utterance from a new speaker was used for adaptation. Furthermore, an effective way to combine the supervised adaptation and the unsupervised adaptation was investigated to reduce the need for a large amount of supervised learning data. Koichi Shinoda |
ICASSP | 1 |
| 1997 | Acoustic modeling based on the MDL principle for speech recognitionabstractACOUSTIC MODELING BASED ON THE MDLPRINCIPLE FOR SPEECH RECOGNITIONKoichi Shinoda and Takao WatanabeNEC Corp oration4-1-1 Miyazaki, Miyamae-ku, Kawasaki 216, JAPANfshino da,watanab [email protected] .nec.co.jpABSTRACTRecently context-dep endent phone units, such as tri-phones, have b een used to mo del subword units in sp eechrecognition based on Hidden Markov Mo dels (HMMs).While most such metho ds employ clustering of theHMM parameters(e.g., subword clustering, state cluster-ing, etc.), to control HMM size so as to avoid p o or recogni-tion accuracy due to an insuciency of training data, noneof them provide any e ective criterion for the optimal de-gree of clustering that should b e p erformed. This pap erprop oses a metho d in which state clustering is accom-plished byway of phonetic decision trees and in which theMDL criterion is used to optimize the degree of cluster-ing. Large-vo cabulary Japanese recognition exp erimentsshow that the mo dels obtained by this metho d achievedthe highest accuracy among the mo dels of various sizesobtained with conventional clustering approaches.1.INTRODUCTIONOver the past few years, extensive studies have b een car-ried out on sp eaker-indep end ent sp eech recognition us-ing continuous density Hidden Markov Mo dels (HMMs).It is well known that in most such systems, the use ofcontext-dep endent(CD) phone mo dels instead of context-indep endent(CI) phone mo dels(monophon es), improvesrecognition accuracy[1-7].Since the numb er of CD mo dels is usually much largerthan that of CI mo dels, using CD mo dels b etter capturesvariations in sp eech data. However, the amountof aail-able training data is likely to b e insucient to supp ortthe use of such a large numb er of CD mo dels. It is oftenimpractical to prepare such a large amount of data. Fur-thermore, the frequency with which a CD phone app earsin training data usually di ers substantiall y in the set ofCD phones; in most case, the frequencies for some CDphones are so small that those CD phones do not app earin training data even if a large amount of data is pro-vided. This data insuciency often causes serious degra-dation in sp eech recognition p erformance. Most recogni-tion systems using CD mo dels employ clustering of mo delparameters to try to alleviate part of the problem.Various clustering metho ds have b een develop ed for thispurp ose. First, there are several choices for the units towhich clustering is carried out; K.F. Leeet al.[1], for ex-ample, use subword clustering, Hwanget al.[2] use stateclustering, and Digalakis et al.[3] cluster the mixture com-p onents of the HMMs with Gaussian-mixture state ob-servation densities.Second, there are several metho dsto select the acoustically-si mi lar units to b e clustered.Some metho ds use only the acoustic characteristics of thedata and the merging of the units are carried out in ab ottom-up manner[4 , 2, 3 ]. The other metho ds, in addi-tion, utilizea prioriknowledge ab out acoustic similariti esbetween the units, which are mostly represented by deci-sion trees[1, 5, 6, 7]. In most of the latter metho ds, split-ting of the units of CI mo dels is carried out in a top-downmanner, instead of merging the units of CD mo dels.In these clustering metho ds, it is imp ortant to prop-erly measure the acoustic similariti es b etween the unitsutilizing training data, in order to select the units tob e clustered.One of the most successful approachis the approach based on the maximum-likel ihood(ML)criterion(e.g.,[7 ]).In the following,for simplicity,the splitting metho d(top-down clustering) is explained,though the similar explanation is also applicable to themerging metho d(b ottom-up clustering). In this approach,the increase of the likeliho o d by splitting is calculated foreach unit in the unit set, and the unit that has the largestincrease is selected and split.However, this ML approach has one drawback. In mostcase, the likelihood becomes larger as the numb er of unitsb ecomes larger.In the nal stage of the splitting, themo del set b ecomes almost identical to the set of CD mo d-els without clustering. Therefore, this approach requiresan external parameter to control the degree of clustering.Most metho ds limit splitting using a threshold on the in-crease in the likeliho o d or on the numb er of units. Thesethresholds needs to b e optimized through a series of recog-nition exp eriments using test data or by a cross-validationmetho d. These optimization pro cesses are computation-ally exp ensive, need more data, and have no strong theo-retical justi cation.In this pap er we prop ose a new approach in whicha minimum description length(MDL) criterion, insteadof the ML criterion, is used for clustering.The MDLapproach[9 ] is based on an information theoretic criterion,which has b een used for selecting the probabilisti c mo delwith an appropriate complexity for the given amountofdata. This MDL criterion is e ective not only for select-ing the units to b e split, but also for deciding whether tostop splitting. Therefore, no other external parameter isneeded to control the degree of clustering. We apply thiscriterion to state splitting using phonetic decision tree.2.MDL CRITERIONMDL[9] is an information criterion which has b een provento b e e ective in selecting the optimal mo del from amongvarious probabilis tic mo dels. The MDL criterion selectsthe mo del with the minimum description length for thegiven data as the optimal mo del from among a set of mo d-els. When a set of mo delsf1; :::;i;:::;Igis given, the de-scription length,li(xN), of the data,f=1;:::;xNg,together with an underlying mo deliis given by,l(i)=logP^(i)xN)+i2N+ logI(1)whereiis the dimensionali ty (the numb er of free param-eters) of mo deli, and^(i)is the maximum likeliho o d es-timates for the parameters(i)=(1;:::;i)ofmodeli. The rst term in (1) is the co de length for the dataxNwhen mo deliis used as a probabili stic mo del. This term Koichi Shinoda, Takao Watanabe |
EUROSPEECH | 1 |
| 1996 | Speaker adaptation with autonomous model complexity control by MDL principleabstractA speaker adaptation method for continuous density HMMs, which performs well for any amount of data for adaptation, is proposed. This method estimates shift parameters for the means of Gaussian mixture components in the HMM. Each shift parameter is shared by more than one Gaussian components. Many sets of shift parameters with various degree of sharing are prepared, and the set with the appropriate complexity for the gives amount of data is selected using minimum description length (MDL) principle. Unlike previous similar works, the proposed method needs no control parameters for selecting models. A series of 5000-word recognition experiments have demonstrated the effectiveness of this new method. Koichi Shinoda, Takao Watanabe |
ICASSP | 1 |
| 1996 | Unsupervised and incremental speaker adaptation under adverse environmental conditions
Keizaburo Takagi, Koichi Shinoda, Hiroaki Hattori, Takao Watanabe |
ICSLP | 2 |
| 1995 | High speed speech recognition using tree-structured probability density functionabstractThis paper proposes a new speech recognition method using a tree-structured probability density function (PDF) to realize high speed HMM based speech recognition. In order to reduce the likelihood calculation for a PDF set composed of the Gaussian PDFs for all mixture components, all states and all recognition units, it is coarsely done for the element PDF whose likelihood is not likely to be large. The PDF set is expressed as a tree-structured form. In the recognition process, the likelihood set is calculated by searching the tree; by calculating the likelihood from the cluster PDF at the node and traversing the nodes with the largest likelihood from the root. Experimental results showed that the computation load was drastically reduced with little reduction in the recognition accuracy, in both speaker-independent and speaker-adaptive cases. The algorithm was applied to a personal computer speech recognition software without using special hardware. Takao Watanabe, Koichi Shinoda, Keizaburo Takagi, Ken-ichi Iso |
ICASSP | 2 |
| 1995 | Speaker adaptation with autonomous control using tree structure
Koichi Shinoda, Takao Watanabe |
EUROSPEECH | 1 |
| 1994 | Unsupervised speaker adaptation for speech recognition using demi-syllable HMM
Koichi Shinoda, Takao Watanabe |
ICSLP | 1 |
| 1994 | Speech recognition using tree-structured probability density function
Takao Watanabe, Koichi Shinoda, Keizaburo Takagi, Eiko Yamada |
ICSLP | 2 |
| 1991 | Speaker adaptation for demi-syllable based continuous density HMMabstractA novel speaker adaptation method for a speech recognition system which uses a continuous density HMM (hidden Markov model) is proposed. It is a supervised adaptation method in which the HMM parameters are modified for new speakers. It is effective not only for recognition units for which there are training samples available, but also for recognition units for which there are no training samples, since the parameters for these units without training samples are estimated by an interpolation technique which are often used in unsupervised adaptation. The effectiveness of the proposed method was evaluated by large vocabulary word recognition experiments, which were carried out under a demi-syllable-based speaker-dependent speech recognition system. The proposed method is shown to be effective when applied to a speaker independent system, under which the recognition accuracy improved by an average of 2.9% for 50 words of training data.> Koichi Shinoda, Ken-ichi Iso, Takao Watanabe |
ICASSP | 1 |
| 1990 | Speaker adaptation for demi-syllable based speech recognition using continuous HMM
Koichi Shinoda, Ken-ichi Iso, Takao Watanabe |
ICSLP | 1 |