Carlos Toshinori Ishi

dblp:85/222 · DBLP profile ↗
← Back
84ranked-venue papers
31as first author
24since 2021 · last 2025
0000-0001-8130-1048ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 69 · 27 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 18 first-author · 10 since 2021Systems, architecture and hardware · 24 · 10 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 SignFlow: End-to-End Sign Language Generation for One-to-Many Modeling using Conditional Flow Matching
Khan Nabeela Khanum, Bowen Wu 0002, Sihan Tan, Carlos Toshinori Ishi, Kazuhiro Nakadai
ICMI4
2025 MultiGAU: Real Time Sign Language Generation Using Multimodal Gated Attention
Khan Nabeela Khanum, Bowen Wu 0002, Carlos Toshinori Ishi, Kazuhiro Nakadai
IEA/AIE (1)3
2025 What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
Kiyotada Mori, Seiya Kawano, Carlos Toshinori Ishi, Angel F. Garcia Contreras, Koichiro Yoshino
INTERSPEECH4
2025 RoboDJ: Live Commentary Robots System Driven by Physical- and Cyber-World Observations
Yasutomo Kawanishi, Yutaka Nakamura, Taiken Shintani, Carlos Toshinori Ishi, Seiya Kawano, Koichiro Yoshino, Takashi Minato, Michihiko Minoh
MMM (5)4
2025 HAM-GNN: A hierarchical attention-based multi-dimensional edge graph neural network for dialogue act classification
Changzeng Fu, Yikai Su, Kaifeng Su, Yinghao Liu, Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro
Expert Syst. Appl.8
2025 Facial action units guided graph representation learning for multimodal depression detection
Changzeng Fu, Fengkui Qian, Yikai Su, Kaifeng Su, Siyang Song, Mingyue Niu, Zhigang Liu 0014, Carlos Toshinori Ishi, Hiroshi Ishiguro
Neurocomputing10
2025 HiMul-LGG: A hierarchical decision fusion-based local-global graph neural network for multimodal emotion recognition in conversation
Changzeng Fu, Fengkui Qian, Kaifeng Su, Yikai Su, Zhigang Liu 0014, Carlos Toshinori Ishi
Neural Networks9
2025 EffIntentGCN: An Efficient Graph Convolutional Network for Skeleton-Based Pedestrian Crossing Intention Prediction
abstract
Understanding pedestrian behavior and predicting their intentions near roads are crucial for enhancing road safety and traffic efficiency. In intelligent transportation systems, especially where computational and energy resources are limited, real-time inference is vital. To address this, our study focuses on pedestrian skeletal information, which offers lower dimensionality and reduces computational demands. We introduce EffIntentGCN, a lightweight graph convolutional network designed to analyze pedestrian poses using a part-based graph and an adaptive graph, emphasizing features crucial for intention prediction and enhancing the learning of local and global interactions. An expansion layer added before each convolution block increases feature dimensionality for a broader representation, without excessive parameter increase. We further incorporate spatial-temporal joint attention, focusing on crucial joints in dynamic skeletal sequences. The model also integrates second-order information, representing joint connections, to provide complementary information. To balance model complexity with computational efficiency, essential for applications in resource-constrained environments like autonomous vehicles, we employ depthwise separable convolution and low-rank approximation techniques, aimed to reduce parameter count and maintain computational efficiency. The model is evaluated on the JAAD dataset, demonstrating its effectiveness in accurately predicting pedestrian crossing intentions while optimizing computational resources.
Carlos Toshinori Ishi, Hiroshi Ishiguro
IEEE Trans. Intell. Transp. Syst.3
2024 X-E-Speech: Joint Training Framework of Non-Autoregressive Cross-lingual Emotional Text-to-Speech and Voice Conversion
Houjian Guo, Carlos Toshinori Ishi, Hiroshi Ishiguro
INTERSPEECH3
2024 Retargeting Human Facial Expression to Human-like Robotic Face through Neural Network Surrogate-based Optimization
abstract
Facial mimicry is crucial for human-like robots in human-robot interaction. The challenge is that the high diversity of facial expressions proposes difficulties in programming a robotic face to mimic human facial expressions using traditional methods. In this paper, we present a data-driven method to retarget human facial expressions to robotic faces without human effort. Our data collection is fully automatic, where only a robotic face and Apple ARKit are involved to sample actuator commands and record the resulting facial blendshape values. We trained a neural network that predicts blendshape values from commands, which is then used as a surrogate model to optimize command values to resemble given facial expressions. Experiments show that the proposed method has achieved lower error in terms of facial blendshape values than baselines. Moreover, the response time can be reduced to 0.2 seconds via TCP/IP through WiFi, offering great potential for real-time application. Our method is a novel framework for retargeting facial expressions to robotic faces, which can be incorporated into various human-robot interaction systems.
Bowen Wu 0002, Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro
IROS3
2024 Age and Spatial Cue Effects on User Performance for an Adaptable Verbal Wayfinding System
abstract
This study aims on developing an interactive verbal wayfinding system that can adapt directions based on the user’s context. Rather than simply providing the shortest route like in existing navigation systems, the developed system searches for routes with fewer turns and generates easy-to-understand verbal explanations. A key feature is prioritizing the inclusion of landmark references at potential points of confusion to provide appropriate contextual cues. Base functions for route searching and verbal direction generation were implemented in this study. A subjective experiment was then conducted to evaluate how user age and the number of provided landmark references impact wayfinding performance when using the system’s verbal directions. Results revealed an age-dependent interaction: increasing landmarks did not universally improve wayfinding success. The optimal number of landmarks varied across age groups. The findings highlight the need to adapt the level of spatial cue details in verbal directions based on the user’s context.
Rin Takahira, Carlos Toshinori Ishi, Takenao Ohkawa
RO-MAN3
2024 Speech-Driven Gesture Generation Using Transformer-Based Denoising Diffusion Probabilistic Models
abstract
While it is crucial for human-like avatars to perform co-speech gestures, existing approaches struggle to generate natural and realistic movements. In the present study, a novel transformer-based denoising diffusion model is proposed to generate co-speech gestures. Moreover, we introduce a practical sampling trick for diffusion models to maintain the continuity between the generated motion segments while improving the within-segment motion likelihood and naturalness. Our model can be used for online generation since it generates gestures for a short segment of speech, e.g., 2 s. We evaluate our model on two large-scale speech-gesture datasets with finger movements using objective measurements and a user study, showing that our model outperforms all other baselines. Our user study is based on the Metahuman platform in the Unreal Engine, a popular tool for creating human-like avatars and motions.
Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro
IEEE Trans. Hum. Mach. Syst.3
2023 QUICKVC: A Lightweight VITS-Based Any-to-Many Voice Conversion Model using ISTFT for Faster Conversion
abstract
With the development of automatic speech recognition and text-to-speech technology, high-quality voice conversion can be achieved by extracting source content information and target speaker information to reconstruct waveforms. However, current methods still require improvement in terms of inference speed. In this study, we propose a lightweight VITS-based voice conversion model that uses the HuBERT-Soft model to extract content information features. Unlike the original VITS model, we use the inverse short-time Fourier transform to replace the most computationally expensive part. Through subjective and objective experiments on synthesized speech, the proposed model is capable of natural speech generation and it is very efficient at inference time. Experimental results show that our model can generate samples at over 5000 KHz on the 3090 GPU and over 250 KHz on the i9-10900K CPU, achieving faster speed in comparison to baseline methods using the same hardware configuration.
Houjian Guo, Carlos Toshinori Ishi, Hiroshi Ishiguro
ASRU3
2023 Using Joint Training Speaker Encoder With Consistency Loss to Achieve Cross-Lingual Voice Conversion and Expressive Voice Conversion
abstract
Voice conversion systems have made significant advancements in terms of naturalness and similarity in common voice conversion tasks. However, their performance in more complex tasks such as cross-lingual voice conversion and expressive voice conversion remains imperfect. In this study, we propose a novel approach that combines a joint training speaker encoder and content features extracted from the cross-lingual speech recognition model Whisper to achieve high-quality cross-lingual voice conversion. Additionally, we introduce a speaker consistency loss to the joint encoder, which improves the similarity between the converted speech and the reference speech. To further explore the capabilities of the joint speaker encoder, we use the Phonetic posteriorgram as the content feature, which enables the model to effectively reproduce both the speaker characteristics and the emotional aspects of the reference speech.
Houjian Guo, Carlos Toshinori Ishi, Hiroshi Ishiguro
ASRU3
2023 HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in Conversation
abstract
The prediction of dialogue acts (DA) labels on utterance-level in conversations can be treated as a sequence labeling problem, which requires context- and speaker-aware semantic comprehension, especially for Japanese. In this study, we pro-posed a hierarchical attention with the graph neural network (HAG) to consider the contextual interconnections as well as the semantics carried by the sentence itself. Concretely, the model use long-short term memory networks (LSTMs) to perform a context-aware encoding within a dialogue window. Then, we construct the context graph by aggregating the neighboring utterances. Subsequently, a speaker feature transformation is executed with a graph attention network (GAT) to calculate the interconnections, while a context-level feature selection is performed with a gated graph convolutional network (GatedGCN) to select the salient utterances that contribute to the DA classification. Finally, we merge the representations of different levels and conduct a classification with two dense layers. We evaluate the proposed model on Japanese dialogue act dataset (JPS-DA). The experimental results show that our method outperforms the baselines.
Changzeng Fu, Zhenghan Chen, Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro
ICASSP6
2023 Recognizing Real-World Intentions using A Multimodal Deep Learning Approach with Spatial-Temporal Graph Convolutional Networks
abstract
Identifying intentions is a critical task for comprehending the actions of others, anticipating their future behavior, and making informed decisions. However, it is challenging to recognize intentions due to the uncertainty of future human activities and the complex influence factors. In this work, we explore the method of recognizing intentions alluded under human behaviors in the real world, aiming to boost intelligent systems' ability to recognize potential intentions and understand human behaviors. We collect data containing real-world human behaviors before using a hand dispenser and a temperature scanner at the building entrance. These data are processed and labeled into intention categories. A questionnaire is conducted to survey the human ability in inferring the intentions of others. Skeleton data and image features are extracted inspired by the answer to the questionnaire. For skeleton-based intention recognition, we propose a spatial-temporal graph convolutional network that performs graph convolutions on both part-based graphs and adaptive graphs, which achieves the best performance compared with baseline models in the same task. A deep-learning-based method using multimodal features is proposed to automatically infer intentions, which is demonstrated to accurately predict intentions based on past behaviors in the experiment, significantly outperforming humans.
Carlos Toshinori Ishi, Bowen Wu 0002, Hiroshi Ishiguro
IROS3
2023 An Adversarial Training Based Speech Emotion Classifier With Isolated Gaussian Regularization
abstract
Speaker individual bias may cause emotion-related features to form clusters with irregular borders (non-Gaussian distributions), making the model sensitive to local irregularities of pattern distributions, resulting in the model over-fit of the in-domain dataset. This problem may cause a decrease in the validation scores in cross-domain (i.e., speaker-independent, channel-variant) implementation. To mitigate this problem, in this paper, we propose an adversarial training-based classifier to regularize the distribution of latent representations to further smooth the boundaries among different categories. In the regularization phase, the representations are mapped into Gaussian distributions in an unsupervised manner to improve the discriminative ability of the latent representations. A single Gaussian distribution is used for mapping the latent representations in our previous study. In this presented work, we adopt a mixture of isolated Gaussian distributions. Moreover, multi-instance learning was adopted by dividing speech into a bag of segments to capture the most salient part of presenting an emotion. The model was evaluated on the IEMOCAP and MELD datasets with in-corpus speaker-independent sittings. In addition, we investigated the accuracy of cross-corpus sittings in simulating speaker-independent and channel-variants. In the experiment, the proposed model was compared not only with baseline models but also with different configurations of our model. The results show that the proposed model is competitive with respect to the baseline, as demonstrated both by in-corpus and cross-corpus validation.
Changzeng Fu, Carlos Toshinori Ishi, Hiroshi Ishiguro
IEEE Trans. Affect. Comput.3
2022 Butsukusa: A Conversational Mobile Robot Describing Its Own Observations and Internal States
abstract
This paper presents an autonomous conversational mobile robot Butsukusa that can describe its own observations and internal states during patrolling tasks. The proposed robot can observe the surrounding environment using the recognition module for objects, humans, environment, localization, and speech and then move autonomously around an indoor living space. Interaction skills via language are required for the robot to perform in such human-centered spaces. To investigate a better communication protocol with users, we evaluate various language generation patterns based on different observations and interaction patterns. The evaluation results indicate that the importance of describing the robot's observation results and internal states, as well as the necessity of an appropriate description, depends on the situation.
Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino, Carlos Toshinori Ishi, Yasutomo Kawanishi, Yutaka Nakamura, Takashi Minato, Yasuki Saito, Michihiko Minoh
HRI4
2022 Controlling the Impression of Robots via GAN-based Gesture Generation
abstract
As a type of body language, gestures can largely affect the impressions of human-like robots perceived by users. Recent data-driven approaches to the generation of co-speech gestures have successfully promoted the naturalness of produced gestures. These approaches also possess greater generalizability to work under various contexts than rule-based methods. However, most have no direct control over the human impressions of robots. The main obstacle is that creating a dataset that covers various impression labels is not trivial. In this study, based on previous findings in cognitive science on robot impressions, we present a heuristic method to control them without manual labeling, and demonstrate its effectiveness on a virtual agent and partially on a humanoid robot through subjective experiments with 50 participants.
Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro
IROS4
2022 Expression of Personality by Gaze Movements of an Android Robot in Multi-Party Dialogues*
abstract
In this study, we describe an improved version of our proposed model to generate gaze movements (eye and head movements) of a dialogue robot in multi-party dialogue situations, and investigated how the impressions change for models created by data of speakers with different personalities. For that purpose, we used a multimodal three-party dialogue data, and first analyzed the distributions of (1) the gaze target (towards dialogue partners or gaze aversion), (2) the gaze duration, and (3) the eyeball direction during gaze aversion. We then generated gaze behaviors in an android robot (Nikola) with the data of two people who were found to have distinctive personalities, and conducted subjective evaluation experiments. Results showed that a significant difference was found in the perceived personalities between the motions generated by the two models.
Taiken Shintani, Carlos Toshinori Ishi, Hiroshi Ishiguro
RO-MAN2
2022 An improved CycleGAN-based emotional voice conversion model by augmenting temporal dependency with a transformer
Changzeng Fu, Carlos Toshinori Ishi, Hiroshi Ishiguro
Speech Commun.3
2021 Analysis of Role-Based Gaze Behaviors and Gaze Aversions, and Implementation of Robot's Gaze Control for Multi-party Dialogue
abstract
In a multi-person face-to-face dialogue, people naturally gaze towards others or avert their gazes, according to their dialogue roles and mental states. The goal of this research is to develop a robot/agent that can generate human-like eye movements, in order to achieve smoother and more engaged dialogue interactions with multiple users. In this study, we analyze the gaze behaviors in three-party dialogue data, accounting for turn-taking, dialogue roles and gaze aversions during the dialogue interactions. Based on the analysis results, we implemented gaze models on a humanoid robot. Subjective evaluation experiments showed that natural behaviors are achieved by our proposed gaze control system, which accounts for dialogue roles and eyeball movement control.
Taiken Shintani, Carlos Toshinori Ishi, Hiroshi Ishiguro
HAI2
2021 MAEC: Multi-Instance Learning with an Adversarial Auto-Encoder-Based Classifier for Speech Emotion Recognition
abstract
In this paper, we propose an adversarial auto-encoder-based classifier, which can regularize the distribution of latent representation to smooth the boundaries among categories. Moreover, we adopt multi-instance learning by dividing speech into a bag of segments to capture the most salient moments for presenting an emotion. The proposed model was trained on the IEMOCAP dataset and evaluated on the in-corpus validation set (IEMOCAP) and the cross-corpus validation set (MELD). The experiment results show that our model outperforms the baseline on in-corpus validation and increases the scores on cross-corpus validation with regularization.
Changzeng Fu, Carlos Toshinori Ishi, Hiroshi Ishiguro
ICASSP3
2021 Analysis of Eye Gaze Reasons and Gaze Aversions During Three-Party Conversations
Carlos Toshinori Ishi, Taiken Shintani
Interspeech1
2019 A Neural Turn-Taking Model without RNN
Carlos Toshinori Ishi, Hiroshi Ishiguro
INTERSPEECH2
2019 Analysis of factors influencing the impression of speaker individuality in android robots
abstract
Humans use not only verbal information but also non-verbal information in daily communication. Among the non-verbal information, we have proposed methods for automatically generating hand gestures in android robots, with the purpose of generating natural human-like motion. In this study, we investigate the effects of hand gesture models trained/designed for different speakers on the impression of the individuality through android robots. We consider that it is possible to express individuality in the robot, by creating hand motion that are unique to that individual. Three factors were taken into account: the appearance of the robot, the voice, and the hand motion. Subjective evaluation experiments were conducted by comparing motions generated in two android robots, two speaker voices, and two motion types, to evaluate how each modality affects the impression of the speaker individuality. Evaluation results indicated that all these three factors affect the impression of speaker individuality, while different trends were found depending on whether or not the android is copy of an existent person.
Ryusuke Mikata, Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro
RO-MAN2
2017 Prosodic Analysis of Attention-Drawing Speech
Carlos Toshinori Ishi, Jun Arai, Norihiro Hagita
INTERSPEECH1
2017 Motion Analysis in Vocalized Surprise Expressions
Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro
INTERSPEECH1
2017 Turn-Taking Estimation Model Based on Joint Embedding of Lexical and Prosodic Contents
Carlos Toshinori Ishi, Hiroshi Ishiguro
INTERSPEECH2
2017 Probabilistic nod generation model based on estimated utterance categories
abstract
We propose a probabilistic model that generates nod motions based on utterance categories estimated from the speech input. The model comprises two main blocks. In the first block, dialogue act-related categories are estimated from the input speech. Considering the correlations between dialogue acts and head motions, the utterances are classified into three categories having distinct nod distributions. Linguistic information extracted from the input speech is fed to a cluster of classifiers which are combined to estimate the utterance categories. In the second block, nod motion parameters are generated based on the categories estimated by the classifiers. The nod motion parameters are represented as probability distribution functions (PDFs) inferred from human motion data. By using speech energy features, the parameters are sampled from the PDFs belonging to the estimated categories. Subjective experiment results indicate that the motions generated by our proposed approach are considered more natural than those of a previous model using fixed nod shapes and hand-labeled utterance categories.
Carlos Toshinori Ishi, Hiroshi Ishiguro
IROS2
2017 Probabilistic 3-D Mapping of Sound-Emitting Structures Based on Acoustic Ray Casting
abstract
This paper presents a two-step framework for creating the three-dimensional (3-D) sound map of an environment with a mobile robot. The first step is the creation of a map that describes the geometry of the environment. The second step is the addition of the acoustic information to the geometric map. The result is a sound map that shows the probability of emitting sound for all the structures in the environment. To build the sound map, a mobile robot equipped with a microphone array drives through the mapped environment. During this drive, the acoustic information gathered by the microphone array is accumulated in a probabilistic manner. First, the likelihood of sound source presence in a set of directions is evaluated from the acoustic power received from these directions. Then, using an estimate of the robot's pose, an acoustic ray casting procedure transfers this likelihood to the structures in the geometric map. Finally, the probability that these structures emit sound is updated accordingly to the likelihood. Experimental results show that the sound maps are: accurate as it was possible to localize sound sources in 3-D and practical as different types of environments were mapped.
Jani Even, Jonas Furrer, Luis Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita
IEEE Trans. Robotics4
2016 Motion generation in android robots during laughing speech
abstract
We are dealing with the problem of generating natural human-like motions during speech in android robots, which have human-like appearances. So far, automatic generation methods have been proposed for lip and head motions of tele-presence robots, based on the speech signal of the tele-operator. In the present study, we aim for extending the speech-driven motion generation methods for laughing speech, since laughter often occurs in natural dialogue interactions and may cause miscommunication if there is mismatch between audio and visual modalities. Based on analysis results of human behaviors during laughing speech, we proposed a motion generation method given the speech signal and the laughing speech intervals. Subjective experiments were conducted using our android robot by generating five different motion types, considering several modalities. Evaluation results show the effectiveness of controlling different parts of the face, head and upper body (eyelid narrowing, lip corner/cheek raising, eye blinking, head motion and upper body motion control).
Carlos Toshinori Ishi, Tomo Funayama, Takashi Minato, Hiroshi Ishiguro
IROS1
2016 Hearing support system using environment sensor network
abstract
In order to solve the problems of current hearing aid devices, we make use of environment sensor network, and propose a hearing support system, where individual target and anti-target sound sources in the environment can be selected, and spatial information of the target sound sources is reconstructed. The performance of the selective sound separation module was evaluated for different noise conditions. Results showed that signal-to-noise ratios of around 15dB could be achieved by the proposed system for a 65dB babble noise plus directional music noise condition. In the same noise condition, subjective intelligibility tests were conducted, and an improvement of 65 to 90% word intelligibility rates could be achieved by using the proposed hearing support system.
Carlos Toshinori Ishi, Jani Even, Norihiro Hagita
IROS1
2016 ERICA: The ERATO Intelligent Conversational Android
abstract
The development of an android with convincingly lifelike appearance and behavior has been a long-standing goal in robotics, and recent years have seen great progress in many of the technologies needed to create such androids. However, it is necessary to actually integrate these technologies into a robot system in order to assess the progress that has been made towards this goal and to identify important areas for future work. To this end, we are developing ERICA, an autonomous android system capable of conversational interaction, featuring advanced sensing and speech synthesis technologies, and arguably the most humanlike android built to date. Although the project is ongoing, initial development of the basic android platform has been completed. In this paper we present an overview of the requirements and design of the platform, describe the development process of an interactive application, report on ERICA's first autonomous public demonstration, and discuss the main technical challenges that remain to be addressed in order to create humanlike, autonomous androids.
Dylan F. Glas, Takashi Minato, Carlos Toshinori Ishi, Tatsuya Kawahara, Hiroshi Ishiguro
RO-MAN3
2016 Speech driven trunk motion generating system based on physical constraint
abstract
We developed a method to automatically generate humanlike trunk motions (neck and waist motions) of a conversational android from its speech in real time. It is based on a spring-damper dynamical model to simulate a human's trunk movement involved in speech. Differing from the existing methods based on machine learning, our system can easily modulate the motions generated due to speech patterns since the parameters in the model correspond to muscle stiffness. The experimental result showed that the android motions generated by our model could be perceived as more natural and motivate participants to talk with the android more, compared with simple copying of human motions.
Kurima Sakai, Takashi Minato, Carlos Toshinori Ishi, Hiroshi Ishiguro
RO-MAN3
2015 Bringing the Scene Back to the Tele-operator: Auditory Scene Manipulation for Tele-presence Systems
abstract
In a tele-operated robot system, the reproduction of auditory scenes, conveying 3D spatial information of sound sources in the remote robot environment, is important for the transmission of remote presence to the tele-operator. We proposed a tele-presence system which is able to reproduce and manipulate the auditory scenes of a remote robot environment, based on the spatial information of human voices around the robot, matched with the operator's head orientation. In the robot side, voice sources are localized and separated by using multiple microphone arrays and human tracking technologies, while in the operator side, the operator's head movement is tracked and used to relocate the spatial positions of the separated sources. Interaction experiments with humans in the robot environment indicated that the proposed system had significantly higher accuracy rates for perceived direction of sounds, and higher subjective scores for sense of presence and listenability, compared to a baseline system using stereo binaural sounds obtained by two microphones located at the humanoid robot's ears. We also proposed three different user interfaces for augmented auditory scene control. Evaluation results indicated higher subjective scores for sense of presence and usability in two of the interfaces (control of voice amplitudes based on virtual robot positioning, and amplification of voices in the frontal direction).
Carlos Toshinori Ishi, Hiroshi Ishiguro
HRI2
2015 Audio augmented point clouds for applications in robotics
abstract
This paper presents a method for representing acoustic information with point clouds by tying it to geometrical features. The motivation is to create a representation of this information that is well suited for mobile robotic applications. In particular, the proposed approach is designed to take advantage of the use of multiple coordinate frames. As an illustrative example, we present a way to create an audio augmented point cloud by adding estimated audio power to the point cloud created by a RGB-D camera. A few applications of this method are presented.
Jani Even, Florent Ferreri, Atsushi Watanabe, Luis Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita
IROS5
2015 Speech activity detection and face orientation estimation using multiple microphone arrays and human position information
abstract
We developed a system for detecting the speech activity intervals of multiple speakers by combining multiple microphone arrays and human tracking technologies. We also proposed a method for estimating the face orientation of the detected speakers. The developed system was evaluated in two steps: individual utterances in different positions and orientations; and simultaneous dialogues by multiple speakers. Evaluation results revealed that the proposed system could detect speech activity intervals with more than 90% of accuracy, and face orientations with standard deviations within 30 degrees, in situations excluding the cases where all arrays are in the opposite direction to the speaker's face orientation.
Carlos Toshinori Ishi, Jani Even, Norihiro Hagita
IROS1
2015 Robot-assisted acoustic inspection of infrastructures - cooperative hammer sounding inspection
abstract
This work presents a human-robot cooperative approach for infrastructure inspection. The goal is to create a robot that assists the human inspector during hammer sounding inspections. Hammer sounding is a frequently used inspection technique that detects invisible defects under the surface of concrete by striking the surface with a hammer and listening the resulting sound. The conventional hammer sounding inspection is time-consuming, and there is no convenient way to represent exhaustively the test results. The proposed approach solves these two problems by having an assistant robot following the inspector, and always being able to look at the hammer impact position. The assistant robot accurately estimates the position of the impact in real-time and creates a detailed representation of the test results. Experimental results show the process for creating the detailed inspection report. The accuracy of the human-robot cooperative approach is evaluated for a real world application. The average error of the impact point estimation was 32 millimeters and the standard deviation was 30 millimeters.
Atsushi Watanabe, Jani Even, Luis Yoichi Morales Saiki, Carlos Toshinori Ishi
IROS4
2015 Online speech-driven head motion generating system and evaluation on a tele-operated robot
abstract
We developed a tele-operated robot system where the head motions of the robot are controlled by combining those of the operator with the ones which are automatically generated from the operator's voice. The head motion generation is based on dialogue act functions which are estimated from linguistic and prosodic information extracted from the speech signal. The proposed system was evaluated through an experiment where participants interact with a tele-operated robot. Subjective scores indicated the effectiveness of the proposed head motion generation system, even under limitations in the dialogue act estimation.
Kurima Sakai, Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro
RO-MAN2
2014 Mapping sound emitting structures in 3D
abstract
This paper presents a framework for creating a 3D map of an environment that contains the probability of a geometric feature to emit a sound. The goal is to provide an automated tool for condition monitoring of plants. The map is created by a mobile platform equipped with a microphone array and laser range sensors. The microphone array is used to estimate the sound power received from different directions whereas the laser range sensors are used for estimating the platform pose in the environment. During navigation, a ray casting method projects the audio measurements made onboard the mobile platform to the map of the environment. Experimental results show that the created map is an efficient tool for sound source localization.
Jani Even, Luis Yoichi Morales Saiki, Nagasrikanth Kallakuri, Jonas Furrer, Carlos Toshinori Ishi, Norihiro Hagita
ICRA5
2014 Analysis of laughter events in real science classes by using multiple environment sensor data
Carlos Toshinori Ishi, Hiroaki Hatano, Norihiro Hagita
INTERSPEECH1
2014 Audio ray tracing for position estimation of entities in blind regions
abstract
This paper presents a framework for making a mobile robot aware of an entity in the blind region of its laser range finders when that entity emits sound. First in a mapping stage, a 3D description of the environment that contains information about acoustic reflection is created. Then during operation, the robot combines estimated directions of arrival of sound with this 3D description to detect entities that are not visible by line of sight sensors but could be heard because of sound reflections. Using this approach, it is possible to restrict the hypothesis about the position of a sound emitting entity in the blind region to a small set of candidate depth values.
Jani Even, Luis Yoichi Morales Saiki, Nagasrikanth Kallakuri, Carlos Toshinori Ishi, Norihiro Hagita
IROS4
2014 Analysis of relationship between head motion events and speech in dialogue conversations
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
Speech Commun.1
2013 Probabilistic approach for building auditory maps with a mobile microphone array
abstract
This paper presents a multi-modal sensor approach for mapping sound sources using an omni-directional microphone array on an autonomous mobile robot. A fusion of audio data (from the microphone array), odometry information and the laser range scan data (from the robot) was used to precisely localize and map the audio sources in an environment. An audio map is created while the robot is autonomously navigating through the environment by continuously generating audio scans with a steered response power (SRP) algorithm. Using the poses of the robot, rays are cast in the map in all directions given by the SRP. Then each occupied cell in the geometric map hit by a ray is assigned a likelihood of containing a sound source. This likelihood is derived from the SRP at that particular instant. Since the localization of the robot is probabilistic, the uncertainty in the pose of the robot in the geometric map is propagated to the occupied cells hit during the ray casting. This process is repeated while the robot is in motion and the map is updated after every audio scan. The generated sound maps were reused and the changes in the audio environment were updated by the robot as it identifies these changes.
Nagasrikanth Kallakuri, Jani Even, Luis Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita
ICRA4
2013 Analysis of factors involved in the choice of rising or non-rising intonation in question utterances appearing in conversational speech
Hiroaki Hatano, Miyako Kiso, Carlos Toshinori Ishi
INTERSPEECH3
2013 Creation of radiated sound intensity maps using multi-modal measurements onboard an autonomous mobile platform
abstract
This paper presents a method for mapping the radiated sound intensity of an environment using an autonomous mobile platform. The sound intensities radiated by the objects are estimated by combining the sound intensity at the platform's position (estimated with a steered response power algorithm) and the distances to the objects (estimated using laser range finders). By combining the estimated sound intensity at the platform's position with the platform's pose obtained from a particle filter based localization algorithm, the sound intensity radiated from the objects is registered in the cells of a grid map covering the environment. This procedure creates a map of the radiated sound intensity that contains information about the sound directivity. To illustrate the effectiveness of the proposed method, a map of radiated sound intensity is created for a test environment. Then the position and the directivity of the sound sources in the test environment are estimated from this map.
Jani Even, Nagasrikanth Kallakuri, Luis Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita
IROS4
2013 Using multiple microphone arrays and reflections for 3D localization of sound sources
abstract
We proposed a method for estimating sound source locations in a 3D space by integrating sound directions estimated by multiple microphone arrays and taking advantage of reflection information. Two types of sources with different directivity properties (human speech and loudspeaker speech) were evaluated for different positions and orientations. Experimental results showed the effectiveness of using reflection information, depending on the position and orientation of the sound sources relative to the array, walls, and the source type. The use of reflection information increased the source position detection rates by 10% on average and up to 60% for the best case.
Carlos Toshinori Ishi, Jani Even, Norihiro Hagita
IROS1
2013 Using sound reflections to detect moving entities out of the field of view
abstract
This paper presents a method for detecting moving entities that are in the robot's path but not in the field of view of sensors like laser scanners, cameras or ultrasonic sensors. The proposed system makes use of passive acoustic localization methods which receive information from occluded regions (at intersections or corners) because of the multipath nature of sound propagation. Contrary to the conventional sensors, this method does not require line of sight. In particular, specular reflections in the environment make it possible to detect moving entities that emit sound such as a walking person or a rolling cart. This idea was exploited for safe navigation of a mobile platform at intersections. The passive acoustic localization output is combined with a 3D geometric map of the environment that is precise enough to estimate sound propagation and reflection using ray casting methods. This gives the robot the ability to detect a moving entity out of the field of view of the sensors that require line of sight. Then the robot is able to recalculate its path and waits until the detected entity is out of its path so that it is safe to move to its destination. To illustrate the performance of the proposed method, a comparison of the robot's navigation with and without the audio sensing is provided for several intersection scenarios.
Nagasrikanth Kallakuri, Jani Even, Luis Yoichi Morales Saiki, Carlos Toshinori Ishi, Norihiro Hagita
IROS4
2013 Analysis of the visual Lombard effect and automatic recognition experiments
Panikos Heracleous, Carlos Toshinori Ishi, Miki Sato, Hiroshi Ishiguro, Norihiro Hagita
Comput. Speech Lang.2
2012 The role of the Lombard reflex in parkinson's disease
abstract
Parkinson's disease (PD) is a severe disease with many symptoms, including speech disorders. Although many methods exist to treat some of PD's symptoms, therapies for speech impairment are not effective and satisfactory, resulting in an open area of research. The current project aims at taking advantage of the Lombard reflex to improve the speech loudness of PD patients. As a first step, the experience of the Lombard reflex by Japanese PD people was confirmed, and the perception of PD patients' speech was evaluated by several subjects. In a following step, methods based on masking sound will be used for intensive training and for self-training of PD patients. However, after intensive training, PD patients may be able to talk louder even without masking noise. In addition, the design and the development of a device based on masking sound that can be used by PD patients while using phone is under consideration.
Panikos Heracleous, Jani Even, Carlos Toshinori Ishi, Takahiro Miyashita, Norihiro Hagita, Masaki Kondo, Kyoko Takanohara
BIBE3
2012 Generation of nodding, head tilting and eye gazing for human-robot dialogue interaction
abstract
Head motion occurs naturally and in synchrony with speech during human dialogue communication, and may carry paralinguistic information, such as intentions, attitudes and emotions. Therefore, natural-looking head motion by a robot is important for smooth human-robot interaction. Based on rules inferred from analyses of the relationship between head motion and dialogue acts, this paper proposes a model for generating head tilting and nodding, and evaluates the model using three types of humanoid robot (a very human-like android, "Geminoid F", a typical humanoid robot with less facial degrees of freedom, "Robovie R2", and a robot with a 3-axis rotatable neck and movable lips, "Telenoid R2"). Analysis of subjective scores shows that the proposed model including head tilting and nodding can generate head motion with increased naturalness compared to nodding only or directly mapping people's original motions without gaze information. We also find that an upwards motion of a robot's face can be used by robots which do not have a mouth in order to provide the appearance that utterance is taking place. Finally, we conduct an experiment in which participants act as visitors to an information desk attended by robots. As a consequence, we verify that our generation model performs equally to directly mapping people's original motions with gaze information in terms of perceived naturalness.
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
HRI2
2012 Fusion of standard and alternative acoustic sensors for robust automatic speech recognition
abstract
This paper focuses on the problem of environmental noises in human-human communication and in automatic speech recognition. To deal with this problem, the use of alternative acoustic sensors -which are attached to the talker and receive the uttered speech through skin or bones- is investigated. In the current study, throat microphones and ear bone microphones are integrated with standard microphones using several fusion methods. The results obtained show that the recognition rates in noisy environments are drastically increased when these sensors are integrated with standard microphones. Moreover, the system does not show any recognition degradations in clean environments. In fact, recognition rates also increase slightly in clean environments. Using late fusion to integrate a throat microphone, an ear bone microphone, and a standard microphone, we achieved a 44% relative improvement in recognition rate in a noisy environment and a 24% relative improvement in recognition rate in a clean environment.
Panikos Heracleous, Jani Even, Carlos Toshinori Ishi, Takahiro Miyashita, Norihiro Hagita
ICASSP3
2012 Evaluation of a formant-based speech-driven lip motion generation
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
INTERSPEECH1
2012 Combining laser range finders and local steered response power for audio monitoring
abstract
This paper presents an audio monitoring system for detecting and identifying people engaged in a conversation. The proposed method is hands-free as it uses a microphone array to acquire the sound. A particularity of the approach is the use of a laser range finder based human tracker system. The human tracker monitors the locations of people then local steered response power is used to detect the people speaking and localize precisely their mouths. Then an audio stream is created for each person and used to perform speaker identification. Experimental results show that the use of the human tracker has several benefits compared to an audio only approach.
Jani Even, Carlos Toshinori Ishi, Panikos Heracleous, Takahiro Miyashita, Norihiro Hagita
IROS2
2012 Evaluation of formant-based lip motion generation in tele-operated humanoid robots
abstract
Generating natural motion in robots is important for improving human-robot interaction. We developed a tele-operation system where the lip motion of a remote humanoid robot is automatically controlled from the operator's voice. In the present work, we introduce an improved version of our proposed speech-driven lip motion generation method, where lip height and width degrees are estimated based on vowel formant information. The method requires the calibration of only one parameter for speaker normalization. Lip height control is evaluated in two types of humanoid robots (Telenoid-R2 and Geminoid-F). Subjective evaluation indicated that the proposed audio-based method can generate lip motion with naturalness superior to vision-based and motion capture-based approaches. Partial lip width control was shown to improve lip motion naturalness in Geminoid-F, which also has an actuator for stretching the lip corners. Issues regarding online real-time processing are also discussed.
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
IROS1
2012 Body-conductive acoustic sensors in human-robot communication
Panikos Heracleous, Carlos Toshinori Ishi, Takahiro Miyashita, Norihiro Hagita
LREC2
2011 Range Based Multi Microphone Array Fusion for Speaker Activity Detection in Small Meetings
Jani Even, Panikos Heracleous, Carlos Toshinori Ishi, Norihiro Hagita
INTERSPEECH3
2011 Improved Acoustic Characterization of Breathy and Whispery Voices
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
INTERSPEECH1
2011 Analysis of Acoustic-Prosodic Features Related to Paralinguistic Information Carried by Interjections in Dialogue Speech
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
INTERSPEECH1
2011 Multi-modal front-end for speaker activity detection in small meetings
abstract
Small informal meetings of two to four participants are very common in work environments. For this reason, a convenient way for recording and archiving these meetings is of great interest. In order to efficiently archive such meetings, an important task to address is to keep trace of “who talked when” during a meeting. This paper proposes a new multi-modal approach to tackle this speaker activity detection problem. One of the novelty of the proposed approach is that it uses a human tracker that relies on scanning laser range finders (LRFs) to localize the participants. This choice is especially relevant for robotic applications as robots are often equipped with LRFs for navigation purpose. In the proposed system, a table top microphone array in the center of the meeting room acquires the audio data while the LRF based human tracker monitors the movement of the participants. Then the speaker activity detection is performed using Gaussian mixture models that were trained before hand. An experiment reproducing a meeting configuration demonstrates the performance of the system for speaker activity detection. In particular, the proposed hands free system maintains an good level of performance compared to the use of close talking microphone while participants are simultaneously speaking.
Jani Even, Panikos Heracleous, Carlos Toshinori Ishi, Norihiro Hagita
IROS3
2011 The effects of microphone array processing on pitch extraction in real noisy environments
abstract
Pitch extraction is important for communication robots, since pitch may carry information about intention, attitude or emotion expression from the user's speech. However, current pitch extraction methods are not robust enough in real noisy environments. In the present work, we propose pitch extraction methods by combining microphone array and auditory scene analysis technologies, and evaluate pitch extraction of multiple speakers in real noisy environments. Evaluation results show that the proposed ML-PSACF (maximum likelihood adaptive beamformer with peak-pruned summary autocorrelation function) contributes to reduce the effects of interference and noise, leading to improvements of 23%±5% on pitch estimation rates, in comparison to the baseline of not using array processing.
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
IROS1
2010 Head motions during dialogue speech and nod timing control in humanoid robots
abstract
Head motion naturally occurs in synchrony with speech and may carry paralinguistic information, such as intention, attitude and emotion, in dialogue communication. With the aim of verifying the relationship between head motion and the dialogue acts carried by speech, analyses were conducted on motion-captured data for several speakers during natural dialogues. The analysis results first confirmed the trends of our previous work, showing that regardless of the speaker, nods frequently occur during speech utterances, not only for expressing dialogue acts such as agreement and affirmation, but also appearing at the last syllable of the phrase, in strong phrase boundaries, especially when the speaker is talking confidently, or expressing interest in the interlocutor's talk. Inter-speaker variability indicated that the frequency of head motion may vary according to the speaker's age or status, while intra-speaker variability indicated that the frequency of head motion also differs depending on the inter-personal relationship with the interlocutor. A simple model for generating nods based on rules inferred from the analysis results was proposed and evaluated in two types of humanoid robots. Subjective scores showed that the proposed model could generate head motions with naturalness comparable to the original motions.
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
HRI1
2010 Close speaker cancellation for suppression of non-stationary background noise for hands-free speech interface
abstract
This paper presents a noise cancellation method based on the ability to efficiently cancel a close target speaker contribution from the signals observed at a microphone array. The proposed method exploits this specificity in the case of the hands-free speech interface. This method is in particular able to deal with non-stationary noise. The method can be divided in three steps. First, the steering vector pointing at the target user is estimated from the covariance of the observed signals. Then the noise estimate is obtained by cancelling the user's contribution. During this step the speech pauses are also estimated. Finally a post-filter is used to suppress this estimated noise from the observed signals. The post-filter strength is controlled by using the estimated noise during the speech pauses as reference. A 20k-words dictation task in presence of non-stationary diffuse background noise at different SNR levels illustrates the effectiveness of the proposed method.
Jani Even, Carlos Toshinori Ishi, Hiroshi Saruwatari, Norihiro Hagita
INTERSPEECH2
2010 Sound interval detection of multiple sources based on sound directivity
abstract
Utterance interval detection is a bottleneck for the current speech recognition performance in robots embedded in real noisy environments. In the present work, we make use of sound localization technology using a microphone array, not only for localizing, but also for detecting sound intervals of multiple sound sources. In our previous work we have implemented and evaluated sound localization in the 3D-space using the MUSIC (MUltiple SIgnal Classification) method. In the present work, we proposed a method for detecting sound intervals based on the sound directivity information inferred from the dynamics of the MUSIC spectrogram. The proposed method achieved high sound interval detection accuracies and low insertion rates compared with the previous sound localization results.
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
IROS1
2009 Evaluation of a MUSIC-based real-time sound localization of multiple sound sources in real noisy environments
abstract
With the goal of improving human-robot speech communication, the localization of multiple sound sources in the 3D-space based on the MUSIC algorithm was implemented and evaluated in a humanoid robot embedded in real noisy environments. The effects of several parameters related to the MUSIC algorithm on sound source localization and real-time performances were evaluated, for recordings in different environments. Real-time processing could be achieved by reducing the frame size to 4 ms, without degrading the sound localization performance. A method was also proposed for determination of the number of sources, which is an important parameter that influences the performance of the MUSIC algorithm. The proposed method achieved localization accuracies and insertion rates comparable with the case where the ideal number of sources is given.
Carlos Toshinori Ishi, Olivier Chatot, Hiroshi Ishiguro, Norihiro Hagita
IROS1
2008 A semi-autonomous communication robot: a field trial at a train station
abstract
This paper reports an initial field trial with a prototype of a semiautonomous communication robot at a train station. We developed an operator-requesting mechanism to achieve semiautonomous operation for a communication robot functioning in real environments. The operator-requesting mechanism autonomously detects situations that the robot cannot handle by itself; a human operator helps by assuming control of the robot.
Masahiro Shiomi, Daisuke Sakamoto, Takayuki Kanda 0001, Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
HRI4
2008 The meanings carried by interjections in spontaneous speech
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
INTERSPEECH1
2008 Automatic extraction of paralinguistic information using prosodic features related to F
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
Speech Commun.1
2008 A Method for Automatic Detection of Vocal Fry
abstract
Vocal fry (also called creak, creaky voice, and pulse register phonation) is a voice quality that carries important linguistic or paralinguistic information, depending on the language. We propose a set of acoustic measures and a method for automatically detecting vocal fry segments in speech utterances. A glottal pulse-synchronized method is proposed to deal with the very low fundamental frequency properties of vocal fry segments, which cause problems in the classic short-term analysis methods. The proposed acoustic measures characterize power, aperiodicity, and similarity properties of vocal fry signals. The basic idea of the proposed method is to scan for local power peaks in a ldquovery short-termrdquo power contour for obtaining glottal pulse candidates, check for periodicity properties, and evaluate a similarity measure between neighboring glottal pulse candidates for deciding the possibility of being vocal fry pulses. In the periodicity analysis, autocorrelation peak properties are taken into account for avoiding misdetection of periodicity in vocal fry segments. Evaluation of the proposed acoustic measures in the automatic detection resulted in 74% correct detection, with an insertion error rate of 13%.
Carlos Toshinori Ishi, Ken-Ichi Sakakibara, Hiroshi Ishiguro, Norihiro Hagita
IEEE Trans. Speech Audio Process.1
2008 A Robust Speech Recognition System for Communication Robots in Noisy Environments
abstract
The application range of communication robots could be widely expanded by the use of automatic speech recognition (ASR) systems with improved robustness for noise and for speakers of different ages. In past researches, several modules have been proposed and evaluated for improving the robustness of ASR systems in noisy environments. However, this performance might be degraded when applied to robots, due to problems caused by distant speech and the robot's own noise. In this paper, we implemented the individual modules in a humanoid robot, and evaluated the ASR performance in a real-world noisy environment for adults' and children's speech. The performance of each module was verified by adding different levels of real environment noise recorded in a cafeteria. Experimental results indicated that our ASR system could achieve over 80% word accuracy in 70-dBA noise. Further evaluation of adult speech recorded in a real noisy environment resulted in 73% word accuracy.
Carlos Toshinori Ishi, Shigeki Matsuda, Takayuki Kanda 0001, Takatoshi Jitsuhiro, Hiroshi Ishiguro, Satoshi Nakamura 0001, Norihiro Hagita
IEEE Trans. Robotics1
2007 Analysis of head motions and speech in spoken dialogue
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
INTERSPEECH1
2007 Analysis of head motions and speech, and head motion control in an android
abstract
With the aim of automatically generating head motions during speech utterances, analyses are conducted for verifying the relations between head motions and linguistic and paralinguistic information carried by speech utterances. Motion captured data are recorded during natural dialogue, and the rotation angles are estimated from the head marker data. Analysis results showed that nods frequently occur during speech utterances, not only for expressing specific dialog acts such as agreement and affirmation, but also as indicative of syntactic or semantic units, appearing at the last syllable of the phrases, in strong phrase boundaries. Analyses are also conducted on the dependence on linguistic, prosodic and voice quality information of other head motions, like shakes and tilts, and discuss about the potentiality for their use in automatic generation of head motions. The paper also proposes a method for controlling the head actuators of an android based on the rotation angles, and evaluates the mapping from the human head motions.
Carlos Toshinori Ishi, Judith Haas, Freerk Pieter Wilbers, Hiroshi Ishiguro, Norihiro Hagita
IROS1
2007 A blendshape model for mapping facial motions to an android
abstract
An important part of natural, and therefore effective, communication is facial motion. The android Repliee Q2 should therefore display realistic facial motion. In computer graphics animation, such motion is created by mapping human motion to the animated character. This paper proposes a method for mapping human facial motion to the android. This is done using a linear model of the android, based on blendshape models used in computer graphics. The model is derived from motion capture of the android and therefore also models the android's physical limitations. The paper shows that the blendshape method can be successfully applied to the android. Also, it is shown that a linear model is sufficient for representing android facial motion, which means control can be very straightforward. Measurements of the produced motion identify the physical limitations of the android and allow identifying the main areas for improvement of the android design.
Freerk Pieter Wilbers, Carlos Toshinori Ishi, Hiroshi Ishiguro
IROS2
2006 Analysis of prosodic and linguistic cues of phrase finals for turn-taking and dialog acts
abstract
This paper presents an analysis on the functions carried by phrase final tones in turn-taking and dialog acts, taking into account linguistic information about the part of speech (particles and auxiliary verbs) attributed to the morphemes at phrase finals. Natural conversational speech data are segmented in inter-pause units, and each utterance unit is arranged according to the phrase final morphemes. Turntaking functions are annotated, and tones of each phrase final are described by acoustic-prosodic features. Analysis results show a relationship between tones and turn-taking functions in most of the morphemes, while no clear relationship is found in some classes of morphemes which are final particles.
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
INTERSPEECH1
2006 Evaluation of Prosodic and Voice Quality Features on Automatic Extraction of Paralinguistic Information
abstract
Aiming to realize a non-verbal communication between humans and robots, the use of acoustic parameters related with voice quality features, besides classical prosodic features, is proposed and evaluated for automatic extraction of paralinguistic information (intentions, attitudes, and emotions) in dialog speech. Experimental results indicated that prosodic features were effective for detecting groups of paralinguistic information expressing specific functions (such as affirmation, denial, and asking for repetition), accounting for 61% of the global identification rate. Voice quality features were effective for detecting part of the paralinguistic information expressing emotions or attitudes (such as surprise, disgust and admiration), leading to 12 % improvement in the global identification rate
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
IROS1
2005 Proposal of acoustic measures for automatic detection of vocal fry
abstract
Vocal fry is a voice quality that often appears in relaxed voices indicating low tension, or in pressed voices expressing attitudes/feelings of surprise, admiration and suffering. We propose a set of acoustic measures for automatically detecting vocal fry segments in speech utterances. In order to deal with vocal fry utterances with very low fundamental frequencies, where classic short-term analysis methods become problematic, a glottal pulse synchronized method is proposed. The acoustic measures are based on power, periodicity and similarity properties of vocal fry signals. The basic idea is to scan for power peaks in a “very short-term ” power contour, check for periodicity properties and evaluate a similarity measure between power peaks for deciding the possibility of vocal fry pulses. Sub-harmonic properties are also taken into account in the periodicity analysis. Evaluation of the proposed measures in automatic detection resulted in 73.3 % correct detection, with an insertion error rate of 3.9 %. 1.
Carlos Toshinori Ishi, Hiroshi Ishiguro, Norihiro Hagita
INTERSPEECH1
2004 A new acoustic measure for aspiration noise detection
abstract
In this paper, we propose a new acoustic measure for detecting aspiration noise in vowels. The measure is an index of synchronization between frequency bands around the first and third formants. The measure is based on the principle that the vocal tract responses to the glottal excitation are synchronized between these frequency bands when aspiration noise is absent, and uncorrelated otherwise. Evaluation results show that the proposed measure can be used together with spectral slope measures for automatic detection of aspiration noise. 1.
Carlos Toshinori Ishi
INTERSPEECH1
2003 Perceptually-related acoustic-prosodic features of phrase finals in spontaneous speech
abstract
With the aim of automatically categorizing phrase final tones, investigations are conducted on the relationship between acoustic-prosodic parameters and perceptual tone categories. Three types of acoustic parameters are proposed: one related to pitch movement within the phrase final, one related to pitch reset prior to the phrase final, and one related to the length of the phrase final. A classification tree is used to evaluate automatic categorization of phrase final tone types, resulting in 76 % correct classification for the best combination among the proposed acoustic parameters. Experiments are also conducted to verify the perceived degree of pitch change within a phrase final, and the perceived degree of pitch reset. While a good relationship is found between the perceptual scores and some of the acoustic parameters, our results also advocate a continuous rather than a categorical relationship between some of the phrase final tone-types considered. 1.
Carlos Toshinori Ishi, Parham Mokhtari, Nick Campbell 0001
INTERSPEECH1
2003 Mora F0 representation for accent type identification in continuous speech and considerations on its relation with perceived pitch values
Carlos Toshinori Ishi, Keikichi Hirose, Nobuaki Minematsu
Speech Commun.1
2001 Identification of accent and intonation in sentences for CALL systems
abstract
In order to construct a CALL (Computer Aided Language Learning) system that can teach learners accent and intonation of Japanese, it's necessary to automatically identify accent types and intonation types in sentence utterances.For this purpose, several acoustic (prosodic) features of speech were investigated taking their effects on human perception into account.For the accent type identification method, the use of average values of F0 in mora and target values of F0 in mora final was evaluated in CV and VC units.Average values of VC units and target values of CV units showed better performance in the identification task.As for the intonation identification, several acoustic features were investigated to represent 6 types of sentence final tones, each conveying different information of intention and perceptual impression.The proposed acoustic features for relative duration and sentence final pitch change showed good correspondence to perceptual features.
Carlos Toshinori Ishi, Nobuaki Minematsu, Ryuji Nishide, Keikichi Hirose
INTERSPEECH1
2000 Identification of Japanese double-mora phonemes considering speaking rate for the use in CALL systems
abstract
Fig.1. Two utterances “Sorewa oQtodesu ” (fast speech) and “Sorewa otodesu ” (slow speech) of a same speaker. The duration of the long phone /Qt / in the upper utterance is shorter than the short phone /t / in the lower utterance, showing the danger of using absolute values.
Carlos Toshinori Ishi, Keikichi Hirose, Nobuaki Minematsu
INTERSPEECH1
2000 The distribution of fillers in lectures in the Japanese language
Michiko Watanabe, Carlos Toshinori Ishi
INTERSPEECH2
1999 A system for learning the pronunciation of Japanese pitch accent
Goh Kawai, Carlos Toshinori Ishi
EUROSPEECH2