VLDB 2026 Research / reviewers in the wild / expert
Katunobu Itou
dblp:35/4885 · also Katsunobu Itou
· DBLP profile ↗
61ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 41 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Selective Privacy Protection in Speech Data: Suppressing Age Inference While Preserving Speaker Utility
Aoi Ito, Katunobu Itou, Ryuichi Nisimura |
COMPSAC | 2 |
| 2025 | Dialogue-Pseudo: A Speaker Pseudonymization Framework for Privacy Protection in Dialogue Speech DataabstractWe present “Dialogue-Pseudo”, a dialogue-aware pseudonymization framework that (i) maps each real speaker to a distinct pseudo-speaker via a dialogue-aware Partition-Assignment-Selection procedure to preserve inter-speaker separability, (ii) replaces personally identifiable information with syllable- and accent-matched surrogates to maintain phonorhythmic timing, and (iii) applies bounded prosody control that normalizes global pitch/energy while preserving relative contours and turn-taking cues. Experiments on the RWCP conversational corpus with a pseudo-speaker pool derived from Common Voice show improved privacy under an ECAPA-TDNN speaker verification model: Equal Error Rate (EER) increases from 37.38% (original) to$\mathbf{4 2. 2 2 \%}$(pseudonymized). Prosody-oriented statistics (F0 mean/variance, RMS energy, pause ratio, turn length) remain close across conditions, and UMAP visualizations indicate that pseudo-speakers are well separated from both the sources and from one another. These results support Dialogue-Pseudo as a practical approach to privacy-preserving dialogue analytics that extends pseudonymization beyond timbre conversion to dialogue-relevant linguistic and prosodic cues, thereby balancing unlinkability with the retention of task-critical information. Aoi Ito, Katunobu Itou |
ISM | 2 |
| 2024 | Speaker Pseudonymization for Japanese Speech Using Duration EmbeddingsabstractIn order to safely expand the use of speech data, research is being conducted on speaker pseudonymization as a method to protect the privacy of speech data. Speaker pseudonymization is a technology that processes speech data to preserve information that can be used for subsequent tasks such as speech recognition while concealing the speaker. In machine learning-based methods, the mainstream approach is to process x-vectors created from MFCCs to change information about speech quality. In this study, we focus on duration information of phoneme or mora as an element that can identify individuals other than speech quality, and aim to expand the scope of pseudonymization of speech data by processing duration as a feature that represents individuality. The proposed pseudonymization based on duration and its conversion reduced the accuracy rate of speaker recognition models by 22% and expanded the scope of pseudonymization of speech data. Aoi Ito, Katunobu Itou |
ISM | 2 |
| 2024 | Homophonic Music Composition Using a GAN and LSTM Pipeline for Melody and Harmony GenerationabstractOver the years, various efforts have been made to create music using procedural methods. However, few of these attempts have focused on generating structured music with melodies and chords that naturally complement each other. This research aims to develop a pipeline of neural networks to produce homophonic music, including a melody and accompanying chords, without external input. The goal is for the generated music to feature a natural-sounding melody with harmonies that provide structural coherence, akin to the approach of a human composer. To enhance the quality of the generated music, the research employs a Generative Adversarial Network (GAN) for melody generation and a hierarchical Long Short-Term Memory (LSTM) network for harmony generation. The GAN is used to produce melodies that exhibit natural motifs and variations, while the hierarchical LSTM is utilized to generate chord progressions that contextually support and enrich the melodies. This step-by-step approach ensures that the generated music is both structurally sound and expressive. By focusing on naturalness and consistency in music, this research aims to closely mimic human creativity. The symbolic models used are unconstrained by specific musical guidelines, allowing for the procedural generation of music that closely resembles human composition. Clément Saint-Marc, Katunobu Itou |
ISM | 2 |
| 2024 | Instrumentality Classification Evaluation System for Natural Sounds*abstractExploring the captivating realm of natural sounds in music creation typically presents unique challenges. Although the simplicity of producing melodies with glasses filled with varying amounts of water is enchanting, replicating such versatility with other natural sounds poses difficulties. The limitations lie in the inherent properties of most natural sounds, which lack the pitch modulation freedom akin to glass instruments, resulting in homogeneity or complexity that may not be conducive to direct musical application. This study addresses these challenges by proposing an evaluation method that leverages neural networks and deep learning techniques. Our method uses a dataset of natural sounds to train a deep neural network model, extracting feature vectors that encapsulate their sonic characteristics. Subsequently, these extracted features are applied to instrumental sound datasets, enabling the modeling of natural sound dimensions within the instrumental realm. Through a classification model, the obtained features facilitate the comparison of natural and instrumental ensembles, offering insights into the suitability of natural sounds as musical instruments. By applying this methodology, our study illustrates how natural sounds such as cowbells can be evaluated in comparison with conventional instruments such as the marimba. This approach enhances our understanding of natural sound potential in music composition and performance, offering new avenues for creative exploration and innovation in sound design. Yuhuan Wang, Katunobu Itou |
ISM | 2 |
| 2022 | Cross-Lingual Transfer Learning Approach to Phoneme Error Detection via Latent Phonetic Representation
Jovan M. Dalhouse, Katunobu Itou |
INTERSPEECH | 2 |
| 2022 | Homophonic Music Composition Using Pipelined LSTMs for Melody and Harmony GenerationabstractThroughout the years, many attempts have been made at creating music procedurally. However, very few of those attempts were concerned with the actual “meaning” of the music that was generated. The goal of this research is to implement a pipelined model of neural networks capable of generating homophonic music without input. The generated music should be “meaningful”, that is to say, it should sound like it has a purpose. This is similar to how a human composer would write music. This idea of “meaning” makes a composition tell a story or express feelings. This is the reason humans write music, as well as create other forms of arts. As the goal of artificial creativity is to approach human creativity as close as possible, Artificial Intelligence should try to imitate humans as close as possible. Therefore, it is important for a music generating AI to understand “meaning” in music. In order to introduce meaning in AI composition, the process is approached step-by-step. Dedicated neural models trained on melodies that express purpose through the use of motifs are first used to generate a melody. That melody serves as input to another neural model, trained on chord progressions that contextualize the melodies they accompany, to generate harmony. Besides the model being symbolic, there are no musical constraints or guidelines for generation. By using this pipelined but unconstrained approach, it is possible to procedurally generate music that sounds as if composed by a human being. Clément Saint-Marc, Katunobu Itou |
SMC | 2 |
| 2018 | Automatic Electronic Organ Reduction System Based on Melody Clustering Considering Melodic and Instrumental CharacteristicsabstractTo maintain the artistic side of music, it is important that many people play it. However, musical arrangement is difficult and time-consuming for amateurs, so even good music is not arranged for instruments that relatively few people play. In this study, we propose an arrangement system that can easily generate scores for various musical instruments, with the aim of helping to keep music alive. Here, we target electronic organ arrangements that can represent a full score. Since such scores include a large number of instrumental parts, we summarize the melodies that play the same role using four features (pitch, harmony, note length, and timbre of the instrument). Next, we select a melody cluster for each part, considering the connections between the melodies and the characteristics of electronic organs. Finally, we adjust the score to make it easier to play on an electronic organ. We then conduct experiments to confirm whether the system can create arrangements that maintain the original piece's overall impression. Here, we compare our new arrangements of four pieces with the existing electronic organ scores. The results showed that the proposed arrangement system can create an electronic organ score from the full score. Katunobu Itou, Daiki Tanaka |
ISM | 1 |
| 2014 | Intra-note segmentation via sticky HMM with DP emissionabstractThis paper presents an intra-note segmentation method for mono-phonic recordings based on acoustic feature variation; each musical note is separated into onset, steady and offset states. The task of intra-note segmentation from audio signals is detecting change points of acoustic feature. In proposed method, the Markov process is assumed on state transition, and time-varying acoustic feature is represented by three Dirichlet processes (DP) that are emitted by the each state. In order to express the generative process, the sticky hidden Markov model (HMM) with DP emission is employed. This modeling allows us to automatically estimate the state transition while avoiding the model selection problem by assuming countably infinite of possible acoustic feature in musical notes. Experimental result shows that the detection accuracy of onset-to-steady and steady-to-offset were improved 2.3 points and 20.7 points from previous method, respectively. Yuma Koizumi, Katunobu Itou |
ICASSP | 2 |
| 2010 | Speaker model updating by the conversational sounds in speaker verificationabstractRecent advancements in information technology have enabled handheld devices to process many types of information. Most security locks used in these devices are managed through the use of a password. However, this feature has attracted only a few users because conventional security is generally viewed as troublesome. Speaker verification is a possible alternative because it offers a feature that converts the user's password into a security lock. Because this is a unobtrusive user certification method, it is less cumbersome for the user. However, speaker verification is not robust to changes in the user's voices. Therefore, in this study, we propose to update the speaker model used for speaker verification in smart-phones. The sound samples comprise Japanese phoneme-balanced sentences recorded from 20 speakers over a 1-month period. The speaker models are updated by putting an old sound sample and a new sound sample together. As a result of the model update, the equal error rate (EER) decreases from 9.84% to 5.46%, representing a drop of approximately 44.52%. Keita Yamamuro, Katunobu Itou |
iiWAS | 2 |
| 2009 | The use of acoustically detected filled and silent pauses in spontaneous speech recognitionabstractIn recognizing spontaneous speech, the performance of typical speech recognizers tends to be degraded by filled and silent pauses, which are hesitation phenomena frequently occurred in such speech. In this paper, we present a method for improving the performance of a speech recognizer by detecting and handling both filled pauses (lengthened vowels) and silent (unfilled) pauses. Our method automatically detects these pauses by using a bottom-up acoustical analysis in parallel with a typical speech decoding process, and then incorporates the detected results into the decoding process. From the results of experiments conducted using the CIAIR spontaneous speech corpus, the effectiveness of the proposed method was confirmed. Jun Ogata, Masataka Goto, Katunobu Itou |
ICASSP | 3 |
| 2008 | Test Collections for Spoken Document Retrieval from Lecture Audio Data
Tomoyosi Akiba, Kiyoaki Aikawa, Yoshiaki Itoh 0001, Tatsuya Kawahara, Hiroaki Nanjo, Hiromitsu Nishizaki, Norihito Yasuda, Yoichi Yamashita, Katunobu Itou |
LREC | 9 |
| 2008 | In-car Speech Data Collection along with Various Multimodal Signals
Akira Ozaki, Sunao Hara, Takashi Kusakawa, Chiyomi Miyajima, Takanori Nishino, Norihide Kitaoka, Katunobu Itou, Kazuya Takeda |
LREC | 7 |
| 2007 | Statistical segmentation and recognition of fingertip trajectories for a gesture interfaceabstractThis paper presents a virtual push button interface created by drawing a shape or line in the air with a fingertip. As an example of such a gesture-based interface, we developed a four-button interface for entering multi-digit numbers by pushing gestures within an invisible 2x2 button matrix inside a square drawn by the user. Trajectories of fingertip movements entering randomly chosen multi-digit numbers are captured with a 3D position sensor mounted on the the forefinger's tip. We propose a statistical segmentation method for the trajectory of movements and a normalization method that is associated with the direction and size of gestures. The performance of the proposed method is evaluated in HMM-based gesture recognition. The recognition rate of 60.0% was improved to 91.3% after applying the normalization method. Kazuhiro Morimoto, Chiyomi Miyajima, Norihide Kitaoka, Katunobu Itou, Kazuya Takeda |
ICMI | 4 |
| 2007 | Driver Modeling Based on Driving Behavior and Its Evaluation in Driver IdentificationabstractAll drivers have habits behind the wheel. Different drivers vary in how they hit the gas and brake pedals, how they turn the steering wheel, and how much following distance they keep to follow a vehicle safely and comfortably. In this paper, we model such driving behaviors as car-following and pedal operation patterns. The relationship between following distance and velocity mapped into a two-dimensional space is modeled for each driver with an optimal velocity model approximated by a nonlinear function or with a statistical method of a Gaussian mixture model (GMM). Pedal operation patterns are also modeled with GMMs that represent the distributions of raw pedal operation signals or spectral features extracted through spectral analysis of the raw pedal operation signals. The driver models are evaluated in driver identification experiments using driving signals collected in a driving simulator and in a real vehicle. Experimental results show that the driver model based on the spectral features of pedal operation signals efficiently models driver individual differences and achieves an identification rate of 76.8% for a field test with 276 drivers, resulting in a relative error reduction of 55% over driver models that use raw pedal operation signals without spectral analysis Chiyomi Miyajima, Yoshihiro Nishiwaki, Koji Ozawa, Toshihiro Wakita, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
Proc. IEEE | 5 |
| 2006 | Development of Micro-Dodecahedral Loudspeaker for Measuring Head-Related Transfer Functions in The Proximal regionabstractThis paper describes new equipment for measuring head-related transfer functions (HRTFs) near a listener's head. 3D sounds in headphones are generated by the convolution of sound signals and an HRTF, which is defined as the acoustical transfer function between a point sound source and the entrance to the ear canal. A loudspeaker is usually used for HRTF measurements, and a distance of more than 1 m separates the loudspeaker and the subject. The region within 1 m of the head is called the 'proximal region,' where a small loudspeaker is needed for accurately measuring HRTF, that is, a conventional loudspeaker cannot be used. In our study, a micro-dodecahedral loudspeaker with twelve piezoelectric ceramic devices is used for HRTF measurements at a diameter is 38 mm. Our experiments examined the characteristics of this loudspeaker. From the results, our developed loudspeaker provides similar performance to a point source, and it is very effective for measuring the HRTFs in the proximal region. Seiichiro Hosoe, Takanori Nishino, Katunobu Itou, Kazuya Takeda |
ICASSP (5) | 3 |
| 2006 | Adaptive Regression Based Framework for In-Car Speech RecognitionabstractWe address issues for improving hands-free speech recognition performance in different car environments using a single distant microphone. In our previous work, we proposed a regression based enhancement method for in-car speech recognition. In this paper, we describe recent improvements and propose a data-driven adaptive regression based speech recognition system, in which both feature enhancement and model compensation are performed. Based on isolated word recognition experiments conducted in 15 real car environments, the proposed adaptive regression approach shows an advantage in average relative word error rate (WER) reductions of 52.5% and 14.8%, compared to original noisy speech and ETSI advanced front-end, respectively. Weifeng Li 0001, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
ICASSP (1) | 2 |
| 2006 | Cepstral Analysis of Driving Behavioral Signals for Driver IdentificationabstractSpectral analysis is applied to such driving behavioral signals as gas and brake pedal operation signals for extracting drivers' characteristics while accelerating or decelerating. Cepstral features of each driver obtained through spectral analysis of driving signals are modeled with a Gaussian mixture model (GMM). A GMM driver model based on cepstral features is evaluated in driver identification experiments using driving signals collected in a driving simulator and in a real vehicle on a city road. Experimental results show that the driver model based on cepstral features achieves a driver identification rate of 89.6% for driving simulator and 76.8% for real vehicle, resulting in 61 % and 55 % error reduction, respectively, over a conventional driver model that uses raw driving signals without spectral analysis Chiyomi Miyajima, Yoshihiro Nishiwaki, Koji Ozawa, Toshihiro Wakita, Katunobu Itou, Kazuya Takeda |
ICASSP (5) | 5 |
| 2006 | Statistical Analysis for Thesaurus Construction using an Encyclopedic Corpus
Yasunori Ohishi, Katunobu Itou, Kazuya Takeda, Atsushi Fujii |
LREC | 2 |
| 2006 | LODEM: A system for on-demand video lectures
Atsushi Fujii, Katunobu Itou, Tetsuya Ishikawa |
Speech Commun. | 2 |
| 2005 | Analysis of a large in-car speech corpus and its application to the multimodel ASRabstractIn-car ASR performance improvement, utilizing a large in-car speech corpus, consisting of the utterances of more than five hundred drivers under real driving conditions, is discussed. A subset design method for efficient cross validations in large-scale speech recognition experiments is proposed. The factor analysis of the results of the recognition experiments show the relationship between word accuracy and utterance characteristics, i.e., SNR, entropy and speaking rates. Based on the factor analysis results, a multimodel approach which uses the utterance duration and subband SNRs as the model selection measures for acoustic and language models, respectively, is proposed. By the proposed multimodel approach, a relative error reduction of 16% is obtained. Hiroshi Fujimura, Chiyomi Miyajima, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
ICASSP (1) | 3 |
| 2005 | Two-stage Noise Spectra Estimation and Regression based In-car Speech Recognition using Single Distant MicrophoneabstractWe present a two-stage noise spectra estimation approach. After the first-stage noise estimation using the improved minima controlled recursive averaging (IMCRA) method, the second-stage noise estimation is performed by employing a maximum a posteriori (MAP) noise amplitude estimator. We also develop a regression-based speech enhancement system by approximating the clean speech with the estimated noise and the original noisy speech. Evaluation experiments show that the proposed two-stage noise estimation method results in lower estimation error for all test noise types. Compared to the original noisy speech, the proposed regression-based approach obtains an average relative word error rate (WER) reduction of 65% in our isolated word recognition experiments conducted in 12 real car environments. Weifeng Li 0001, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
ICASSP (1) | 2 |
| 2005 | Subjective and objective quality assessment of regression-enhanced speech in real car environments
Weifeng Li 0001, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 2 |
| 2005 | Discrimination between singing and speaking voicesabstractDiscriminating between singing and speaking voices by using the local and global characteristics of voice signals is discussed. From the results of subjective experiments, we show that human beings can discriminate singing and speaking voices with more than 70 % and 95 % accuracy from 300 ms and one second long signals, respectively. From the subjective experiment results, assuming that different features are effective for shortterm and long-term signals, we designed two measures using a spectral envelope (MFCC) and the fundamental frequency (F0, perceived as pitch) contour. Experimental results show that the F0 measure performs better than the spectral envelope measure when the input voice signals are longer than one second. Particularly, it can discriminate singing and speaking voices with more than 80 % accuracy with two-second signals. On the other hand, when the input signals are shorter than one second, the spectral envelope measure performs better than the F0 measure. Finally, by simply combining the two measures, more than 90% accuracy is obtained for two-second signals. 1. Yasunori Ohishi, Masataka Goto, Katunobu Itou, Kazuya Takeda |
INTERSPEECH | 3 |
| 2005 | Data collection and evaluation of speech recognition for motorbike ridersabstractAbstract Speech recognition should be as an eyes-free and hands-free interface. To realise this technology, we need to clar-ify acoustics in a helmet and determine how much high-level riding noise affects captured speech data. This pa-per describes the acoustics in a helmet and transfer func-tions of the microphone position. We constructed a datacollection system and collected the speech data of motor-bike riders on city roads and express highways. Speechrecognition experiments were conducted and we obtaineda recognition rate high of 83.1%. 1. Introduction Riding a motorbike requires more care than driving a car.Even when a rider idles his/her motorbike, such as whenhe/she waits at a red light, button operations are incon-venient because the rider needs to remove his/her glovesto push the buttons. Therefore, for motorbike riders, aneyes-free and hands-free interface is required for operat-ing information appliances, such as a cellular phone anda route navigation system. Thus, speech recognition is avery important technology.This study investigated the feasibility of speechrecognition for motorbike riders. On a motorbike, rid-ers are exposed directly to high-level noises such as windnoise, engine noise, and road noise. It is known thatexposed noise level is varied by various factors such asspeed, riding position, and helmets[1, 2]. In order to re-alize speech recognition on a motorbike, we first need toinvestigate how much such factors degrade conventionalspeech recognition performance.Riders must put on a helmet when they ride motor-bikes. Helmets are designed to reduce noise level, how-ever, we need to clarify how much this reduction con-tributes to speech recognition using microphones insidethe side of helmets. Moreover, we need to investigateacoustics in a helmet, because a helmet has a very smallcavity.In this paper, we measured acoustics in a helmetto determine microphone positions for collecting riders’speech corpus. We then collected the speech data utteredby motorbike riders riding on a highway. We also providean analysis of the corpus and the results of our speechrecognition evaluation. Hiroshi Fujimura, Chiyomi Miyajima, Takanori Nishino, Katunobu Itou, Kazuya Takeda |
INTERSPEECH | 5 |
| 2004 | Biometric identification using driving behavioral signalsabstractWe investigate the uniqueness of driver behavior in vehicles and the possibility of using it for personal identification with the objectives of achieving safer driving, of assisting the driver in case of emergencies, and of being a part of a multi-mode biometric signature for driver identification. We use Gaussian mixture models (GMM) for modeling the individualities of the accelerator and brake pedal pressures, and focus on not only the static features, but also the dynamics of the pedal pressures. Experimental results show that the dynamic features significantly improve the performance of driver identification. Kei Igarashi, Chiyomi Miyajima, Katunobu Itou, Kazuya Takeda, Fumitada Itakura, Hüseyin Abut |
ICME | 3 |
| 2004 | Speech recognition using synchronization between speech and finger tapping
Hiromitsu Ban, Chiyomi Miyajima, Katunobu Itou, Fumitada Itakura, Kazuya Takeda |
INTERSPEECH | 3 |
| 2004 | Unsupervised topic adaptation for lecture speech retrievalabstractWe are developing a cross-media information retrieval system, in which users can view specific segments of lecture videos by submitting text queries. To produce a text index, the audio track is extracted from a lecture video and a transcription is generated by automatic speech recognition. In this paper, to improve the quality of our retrieval system, we extensively investigate the effects of adapting acoustic and language models on speech recognition. We perform an MLLR-based method to adapt an acoustic model. To obtain a corpus for language model adaptation, we use the textbook for a target lecture to search a Web collection for the pages associated with the lecture topic. We show the effectiveness of our method by means of experiments. 1. Atsushi Fujii, Tetsuya Ishikawa, Katunobu Itou, Tomoyosi Akiba |
INTERSPEECH | 3 |
| 2004 | Analysis of in-car speech recognition experiments using a large-scale multi-mode dialogue corpusabstractThe dependency of conversational utterances on themode of dialogue is analyzed. A speech corpus of 800 speak-ers collected under three different modes, i.e., talking to a human operator, an WOZ system and an ASR system, is used for analysis. Some characteristics such as sen-tence complexity loudness of the voice and speaking-rate are found to be significantly different among the dialogue modes. Linear regression analysis results also clarify the relative importance of those characteristics on speech recognition accuracy. 1. Hiroshi Fujimura, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 2 |
| 2004 | Speech spotter: on-demand speech recognition in human-human conversation on the telephone or in face-to-face situationsabstractThis paper describes a novel speech-interface function, called “speech spotter”,whichenablesausertoentervoicecommands into a speech recognizer in the midst of natural human-human conversation. In the past, it has been difficult to use automatic speech recognition in human-human conversation since it was not easy to judge, from only microphone input, whether a user was speaking to another person or a speech recognizer. We solve this problem by using two kinds of nonverbal speech information: a filled pause (a vowel-lengthening hesitation like “er...”) and voice pitch. Only when a user utters a voice command with a high pitch just after a filled pause is the voice command accepted by the speech recognizer. By using this speechspotter function, we have built two application systems: an ondemand information system for assisting human-human conversation and a music-playback system for enriching telephone conversation. The results from using these systems have shown thatthespeech-spotter functionisrobustandconvenientenough to be used in face-to-face or cellular-phone conversations. Masataka Goto, Koji Kitayama, Katunobu Itou, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2004 | Effects of language modeling on speech-driven question answeringabstractWe integrate automatic speech recognition (ASR) and question answering (QA) to realize a speech-driven QA system, and evaluate its performance. We adapt an Ngram language model to natural language questions, so that the input of our system can be recognized with a high accuracy. We target WH-questions which consist of the topic part and fixed phrase used to ask about something. We first produce a general N-gram model intended to recognize the topic and emphasize the counts of the N-grams that correspond to the fixed phrases. Given a transcription by the ASR engine, the QA engine extracts the answer candidates from target documents. We propose a passage retrieval method robust against recognition errors in the transcription. We use the QA test collection produced in NTCIR, which is a TREC-style evaluation workshop, and show the effectiveness of our method by means of experiments. Katunobu Itou, Atsushi Fujii, Tomoyosi Akiba |
INTERSPEECH | 1 |
| 2004 | Recent progress of open-source LVCSR engine julius and Japanese model repositoryabstractContinuous Speech Recognition Consortium (CSRC) was founded for further enhancement of Japanese Dictation Toolkit that had been developed by the support of a Japanese agency. Overview of its product software is reported in this paper. The open-source LVCSR (large vocabulary continuous speech recognition) engine Julius has been improved both in performance and functionality, and it is also ported to Microsoft Windows in compliance with SAPI (Speech API). The software is now used for not a few languages and plenty of applications. For plug-and-play speech recognition in various applications, we have also compiled a repository of acoustic and language models for Japanese. Especially, the acoustic model set realizes wider coverage of user generations and speech-input environments. Tatsuya Kawahara, Akinobu Lee, Kazuya Takeda, Katunobu Itou, Kiyohiro Shikano |
INTERSPEECH | 4 |
| 2004 | Collecting Spontaneously Spoken Queries for Information Retrieval
Tomoyosi Akiba, Atsushi Fujii, Katunobu Itou |
LREC | 3 |
| 2003 | Adapting language models for frequent fixed phrases by emphasizing n-gram subsetsabstractIn support of speech-driven question answering, we propose a method to construct N-gram language models for recognizing spoken questions with high accuracy. Question-answering sys-tems receive queries that often consist of two parts: one conveys the query topic and the other is a fixed phrase used in query sentences. A language model constructed by using a target col-lection of QA, for example, newspaper articles, can model the former part, but cannot model the latter part appropriately. We tackle this problem as task adaptation from language models ob-tained from background corpora (e.g., newspaper articles) to the fixed phrases, and propose a method that does not use the task-specific corpus, which is often difficult to obtain, but instead uses only manually listed fixed phrases. The method empha-sizes a subset of N-grams obtained from a background corpus that corresponds to fixed phrases specified by the list. Theoret-ically, this method can be regarded as maximizing a posteriori probability (MAP) estimation using the subset of the N-grams as a posteriori distribution. Some experiments show the effec-tiveness of our method. 1. Tomoyosi Akiba, Katunobu Itou, Atsushi Fujii |
INTERSPEECH | 2 |
| 2003 | Building a test collection for speech-driven web retrievalabstractThis paper describes a test collection (benchmark data) for retrieval systems driven by spoken queries. This collection was produced in the subtask of the NTCIR-3 Web retrieval task, which was performed in a TREC-style evaluation workshop. The search topics and document collection for the Web retrieval task were used to produce spoken queries and language models for speech recognition, respectively. We used this collection to evaluate the performance of our retrieval system. Experimental results showed that (a) the use of target documents for language modeling and (b) enhancement of the vocabulary size in speech recognition were effective in improving the system performance. Atsushi Fujii, Katunobu Itou |
INTERSPEECH | 2 |
| 2003 | A cross-media retrieval system for lecture videosabstractWe propose a cross-media lecture-on-demand system, in which users can selectively view specific segments of lecture videos by submitting text queries. Users can easily formulate queries by using the textbook associated with a target lecture, even if they cannot come up with effective keywords. Our system extracts the audio track from a target lecture video, generates a transcription by large vocabulary continuous speech recognition, and produces a text index. Experimental results showed that by adapting speech recognition to the topic of the lecture, the recognition accuracy increased and the retrieval accuracy was comparable with that obtained by human transcription. 1. Atsushi Fujii, Katunobu Itou, Tomoyosi Akiba, Tetsuya Ishikawa |
INTERSPEECH | 2 |
| 2003 | Speech shift: direct speech-input-mode switching through intentional control of voice pitchabstractThis paper describes a speech-input interface function, called speech shift, that enables a user to specify a speech-input mode by simply changing (shifting) voice pitch. While current speech-input interfaces have used only verbal information, we aimed at building a more user-friendly speech interface by making use of nonverbal information, the voice pitch. By intentionally controlling the pitch, a user can enter the same word with it having different meanings (functions) without explicitly changing the speech-input mode. Our speech-shift function implemented on a voice-enabled word processor, for example, can distinguish an utterance with a high pitch from one with a normal (low) pitch, and regard the former as voice-command-mode input(suchasfile-menuandedit-menucommands)andthelatter as regular dictation-mode text input. Our experimental results from twenty subjects showed that the speech-shift function is effective, easy to use, and a labor-saving input method. Masataka Goto, Yukihiro Omoto, Katunobu Itou, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2003 | Speech starter: noise-robust endpoint detection by using filled pausesabstractIn this paper we propose a speech interface function, called speech starter, that enables noise-robust endpoint (utterance) detection for speech recognition. When current speech recognizers are used in a noisy environment, a typical recognition error is caused by incorrect endpoints because their automatic detection is likely to be disturbed by non-stationary noises. The speech starter function enables a user to specify the beginning of each utterance by uttering a filler with a filled pause, which is used as a trigger to start speech-recognition processes. Since filled pauses can be detected robustly in a noisy environment, practical endpoint detection is achieved. Speech starter also offers the advantage of providing a hands-free speech interface and it is user-friendly because a speaker tends to utter filled pauses (e.g., “er...”) at the beginning of utterances when hesitating in human-human communication. Experimental results from a 10-dB-SNR noisy environment show that the recognition error rate with speech starter was lower than with conventional endpoint-detection methods. 1. Koji Kitayama, Masataka Goto, Katunobu Itou, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2002 | A Method for Open-Vocabulary Speech-Driven Text RetrievalabstractWhile recent retrieval techniques do not limit the number of index terms, out-ofvocabulary (OOV) words are crucial in speech recognition.Aiming at retrieving information with spoken queries, we fill the gap between speech recognition and text retrieval in terms of the vocabulary size.Given a spoken query, we generate a transcription and detect OOV words through speech recognition.We then correspond detected OOV words to terms indexed in a target collection to complete the transcription, and search the collection for documents relevant to the completed transcription.We show the effectiveness of our method by way of experiments. Atsushi Fujii, Katunobu Itou, Tetsuya Ishikawa |
EMNLP | 2 |
| 2002 | Selective back-off smoothing for incorporating grammatical constraints into the n-gram language modelabstractSpoken queries submitted to question answering systems usually consist of query contents (e.g. about newspaper articles) and frozen patterns (e.g. WH-words), which can be modeled with N-gram models and grammar-based models, respectively. We propose a method to integrate those different types of models into a single N-gram model. We represent the two types of language models in a single word network. However, common smoothing methods, which are effective for N-gram models, decrease grammatical constraints for frozen patterns. For this problem, we propose a selective back-off smoothing method, which controls a degree to which smoothing is applied depending the network fragment. Additionally, resulting models are compatible with the conventional back-off N-gram models, and thus existing N-gram decoders can easily be used. We show the effectiveness of our method by way of experiments. 1. Tomoyosi Akiba, Katunobu Itou, Atsushi Fujii, Tetsuya Ishikawa |
INTERSPEECH | 2 |
| 2002 | Speech completion: on-demand completion assistance using filled pauses for speech input interfacesabstractThis paper describes a novel speech interface function, called speech completion, that helps a user enter a word or phrase by completing (filling in the rest of) a phrase fragment uttered by the user. Although the concept of completion is widely used in text-based interfaces, there have been no reports of completion being effectively applied to speech. By using a filled pause, we enable a user to effortlessly invoke the speech-completion function which helps the user recall uncertain phrases and saves labor when the input phrase is long. When a user hesitates by lengthening a vowel (a filled pause is uttered) during a phrase, our system immediately displays completion candidates whose beginnings acoustically resemble the uttered fragment so that the user can select the correct one. In our experiments with a system that included a filled-pause detector and a speech recognizer capable of listing candidates, the effectiveness of speech completion was confirmed. Masataka Goto, Katunobu Itou, Satoru Hayamizu |
INTERSPEECH | 2 |
| 2002 | Producing a Large-scale Encyclopedic Corpus over the Web
Atsushi Fujii, Katunobu Itou, Tetsuya Ishikawa |
LREC | 2 |
| 2002 | Continuous Speech Recognition Consortium an Open Repository for CSR Tools and Models
Akinobu Lee, Tatsuya Kawahara, Kazuya Takeda, Masato Mimura, Atsushi Yamada, Akinori Ito, Katunobu Itou, Kiyohiro Shikano |
LREC | 7 |
| 2001 | A structured statistical language model conditioned by arbitrarily abstracted grammatical categories based on GLR parsingabstractThis paper presents a new statistical language model for speech recognition, based on Generalized LR parsing. The proposed model, the Abstracted Probabilistic GLR (APGLR) model, is an extension of the existing structured language model known as the Probabilistic GLR (PGLR) model. It can predict next words from arbitrarily abstracted categories. The APGLR model is also a generalization of the original PGLR model, because PGLR can be considered to be a special case of APGLRs that predict the next words from the least abstracted grammatical categories, namely the terminal symbols. The selection of the Tomoyosi Akiba, Katunobu Itou |
INTERSPEECH | 2 |
| 2001 | Real-time sound source localization and separation system and its application to automatic speech recognitionabstractA real-time sound localization/separation system for near-field sound sources was constructed and evaluated in a real office environment. As for the sound localization, the experimental results showed that the direction of the two sources was estimated with high accuracy while the range of the sources was estimated with moderate accuracy. As for the sound separation, a recognition rate of 70 % for an on-line recognizer on a network and of 90% for an off-line recognizer were achieved, respectively. Futoshi Asano, Masataka Goto, Katunobu Itou, Hideki Asoh |
INTERSPEECH | 3 |
| 2001 | Spoken Language Interface of the Jijo-2 Office Robot
Toshihiro Matsui, Hideki Asoh, Futoshi Asano, John Fry, Isao Hara, Yoichi Motomura, Katunobu Itou |
ISRR | 7 |
| 2000 | Semi-automatic language model acquisition without large corpora
Tomoyosi Akiba, Katunobu Itou |
INTERSPEECH | 2 |
| 2000 | Free software toolkit for Japanese large vocabulary continuous speech recognitionabstractA sharable software repository for Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) is introduced. It is designed as a baseline platform for research and developed by researchers of different academic institutes under a governmental support. The repository consists of a recognition engine (Julius), Japanese acoustic models and statistical language models as well as Japanese morphological analysis tools. These modules can be easily integrated and replaced under a plug-and-play framework, which makes it possible to fairly evaluate components and to develop specific application systems. Assessment of these modules and systems in a 20000-word dictation task is reported. The software repository is freely available to the public. Tatsuya Kawahara, Akinobu Lee, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Shigeki Sagayama, Katunobu Itou, Akinori Ito, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
INTERSPEECH | 7 |
| 2000 | IPA Japanese Dictation Free Software Project
Katunobu Itou, Kiyohiro Shikano, Tatsuya Kawahara, Kazuya Takeda, Atsushi Yamada, Akinori Ito, Takehito Utsuro, Tetsunori Kobayashi, Nobuaki Minematsu, Mikio Yamamoto, Shigeki Sagayama, Akinobu Lee |
LREC | 1 |
| 1999 | A real-time filled pause detection system for spontaneous speech recognitionabstractThis paper describes a method for automatically detecting filled (vocalized) pauses, which are one of the hesitation phenomena that current speech recognizers typically cannot handle. The detection of these pauses is important in spontaneous speech dialogue systems because they play valuable roles, such as helping a speaker keep a conversational turn, in oral communication. Although a few speech recognition systems have processed filled pauses within subword-based connected word recognition or word-spotting frameworks, they did not detect the pauses individually and consequently could not consider their roles. In this paper we propose a method that detects filled pauses and word lengthening on the basis of small fundamental frequency transition and small spectral envelope deformation under the assumption that speakers do not change articulator parameters during filled pauses. Experimental results for a Japanese spoken dialogue corpus show that our real-time filled-pause-detection system yielded a recall rate of 84.9 % and a precision rate of 91.5%. Masataka Goto, Katunobu Itou, Satoru Hayamizu |
EUROSPEECH | 2 |
| 1998 | The design of the newspaper-based Japanese large vocabulary continuous speech recognition corpusabstractIn this paper we present the first public Japanese speech corpus for large vocabulary continuous speech recognition (LVCSR) technology, which we have titled JNAS (Japanese Newspaper Article Sentences). We designed it to be comparable to the corpora used in the American and European LVCSR projects. The corpus contains speech recordings (60 hrs.) and their orthographic transcriptions for 306 speakers (153 males and 153 females) reading excerpts from the newspaper's articles and phonetically balanced (PB) sentences. This corpus contains utterances of about 45,000 sentences as a whole with each speaker reading about 150 sentences. JNAS is being distributed on 16 CD-ROMs. Katunobu Itou, Mikio Yamamoto, Kazuya Takeda, Toshiyuki Takezawa, Tatsuo Matsuoka, Tetsunori Kobayashi, Kiyohiro Shikano, Shuichi Itahashi |
ICSLP | 1 |
| 1998 | Sharable software repository for Japanese large vocabulary continuous speech recognitionabstractThe project of Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) platform is introduced. It is a collaboration of researchers of different academic institutes and intended to develop a sharable software repository of not only databases but also models and programs. The platform consists of a standard recognition engine, Japanese phone models and Japanese statistical language models. A set of Japanese phone HMMs are trained with ASJ (Acoustic Society of Japan) databases of 20K sentence utterances per each gender. Japanese word N-gram (2-gram and 3-gram) models are constructed with a corpus of Mainichi newspaper of four years. The recognition engine JULIUS is developed for assessment of both acoustic and language models. The modules are integrated as a Japanese LVCSR system and evaluated on 5000-word dictation task. The software repository is available to the public. Tatsuya Kawahara, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Katunobu Itou, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
ICSLP | 5 |
| 1996 | RWC multimodal database for interactions by integration of spoken language and visual information
Satoru Hayamizu, Osamu Hasegawa, Katunobu Itou, Katsuhiko Sakaue, Kazuyo Tanaka, Shigeki Nagaya, Masayuki Nakazawa, T. Endoh, Fumio Togawa, Kenji Sakamoto, Kazuhiko Yamamoto |
ICSLP | 3 |
| 1995 | Active Agent Oriented Multimodal Interface System
Osamu Hasegawa, Katunobu Itou, Takio Kurita, Satoru Hayamizu, Kazuyo Tanaka, Kazuhiko Yamamoto, Nobuyuki Otsu |
IJCAI | 2 |
| 1994 | Collecting and analyzing nonverbal elements for maintenance of dialog using a wizard of oz simulation
Katunobu Itou, Tomoyosi Akiba, Osamu Hasegawa, Satoru Hayamizu, Kazuyo Tanaka |
ICSLP | 1 |
| 1994 | Annotating illocutionary force types and phonological features into a spontaneous dialogue corpus: an experimental study
Kazuyo Tanaka, Kanae Kinebuchi, Naoko Houra, Kazuyuki Takagi, Shuichi Itahashi, Katunobu Itou, Satoru Hayamizu |
ICSLP | 6 |
| 1993 | Detection of unknown words in large vocabulary speech recognitionabstractThis paper describes the relation between vocabulary sizes and detection errors of unknown words in large vocabulary speech recognition through recognition and detection experiments.Although the relation between vocabulary sizes and recognition performances has been reported, the relation between vocabulary sizes and detection performances has not yet been studied.Especially, it has not for the cases of vocabulary sizes of over 1,000 words.Experiments were conducted using the speech material of speaker MAU's ATR word speech database.The entries of the dictionary used is 40,000 words from the Shinmeikai Japanese Language Dictionary.It is shown that when the vocabulary size increases from 1,000 words to 40,000 words, the relation between vocabulary sizes and detection errors has a similar tendency with the relation between vocabulary sizes and recognition errors.And increases of detection errors caused by increases of vocabulary sizes are shown to be small for the case of within vocabulary, compared with increases of detection errors for the case of out of vocabulary.These results should be taken into accounts in designing large vocabulary speech recognition systems including unknown word processing. Satoru Hayamizu, Katunobu Itou, Kazuyo Tanaka |
EUROSPEECH | 2 |
| 1992 | Continuous speech recognition by context-dependent phonetic HMM and an efficient algorithm for finding N-Best sentence hypothesesabstractA continuous speech recognition system 'niNja' (Natural language INterface in JApanese), is presented. Efficient search algorithms are proposed to get high accuracy and to reduce the required computations. First, an LR parsing algorithm with context-dependent phone models is proposed. Second, scores of the same phone models in different hypotheses at the phone-level are represented by the single score of the best hypotheses. The system is tested for the task with a 113 word vocabulary, with a word perplexity of 4.1. It produces a sentence accuracy of 97.3% for the 10 open speakers' 110 sentences and the error reduction is as much as 77% compared with using context independent phone models.> Katunobu Itou, Satoru Hayamizu, Hozumi Tanaka |
ICASSP | 1 |
| 1992 | A spoken language dialogue system for automatic collection of spontaneous speech
Satoru Hayamizu, Katunobu Itou, Masafumi Tamoto, Kazuyo Tanaka |
ICSLP | 2 |
| 1992 | Detection of unknown words and automatic estimation of their transcriptions in continuous speech recognition
Katunobu Itou, Satoru Hayamizu, Hozumi Tanaka |
ICSLP | 1 |
| 1990 | Japanese phonetic typewriter using HMM phone units and syllable trigrams
Takeshi Kawabata, Toshiyuki Hanazawa, Katunobu Itou, Kiyohiro Shikano |
ICSLP | 3 |