Komei Sugiura

dblp:77/2654 · DBLP profile ↗
← Back
46ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0002-0261-0510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 13 since 2021Systems, architecture and hardware · 10 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author
YearPublicationVenuePosition
2026 LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
abstract
We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas they do not always demonstrate high performance. To address these issues, we propose Pearl, an LLM-free supervised metric for image captioning, which is applicable to both reference-based and reference-free settings. We introduce a novel mechanism that learns the representations of image--caption and caption--caption similarities. Furthermore, we construct a human-annotated dataset for image captioning metrics that comprises approximately 333k human judgments collected from 2,360 annotators across over 75k images. Pearl outperformed other existing LLM-free metrics on the Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and FOIL datasets in both reference-based and reference-free settings.
Shinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki, Komei Sugiura
AAAI5
2026 ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
Daichi Yashima, Shuhei Kurita, Yusuke Oda, Shuntaro Suzuki, Seitaro Otsuki, Komei Sugiura
ICPR (3)6
2025 VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
abstract
In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs).Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions.Moreover, recent LLM-as-a-Judge approaches suffer from slow inference due to their reliance on autoregressive inference and early fusion of visual information.To address these limitations, we propose VELA, an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework.Furthermore, we propose LongCap-Arena, a benchmark specifically designed for evaluating metrics for long captions.This benchmark comprises 7,805 images, the corresponding human-provided long reference captions and long candidate captions, and 32,246 human judgments from three distinct perspectives: Descriptiveness, Relevance, and Fluency.We demonstrated that VELA outperformed existing metrics and achieved superhuman performance on LongCap-Arena.Our code and dataset are available at https://vela.kinsta.page/.
Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki, Komei Sugiura
EMNLP5
2025 Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos
abstract
Understanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstrations. However, automatic action sequence generation is not always perfect due to the distribution gap between the training and test environments. To bridge this gap, human intervention could be very effective, such as telling the robot agent what should be done. Motivated by this, we propose an error-correction-based action replanning approach that regenerates better action sequences using (1) automatically generated actions from a pretrained action generator and (2) human error-correction in natural language. We collected singlearm robot action sequences aligned to human action instruction for the cooking video dataset YouCook2. We trained the proposed errorcorrection-based action replanning model using a pre-trained multimodal LLM model (AVBLIP-2), generating a pair of (a) single-arm robot micro-step action sequences and (b) action descriptions in natural language simultaneously. To assess the performance of error correction, we collected human feedback on correcting errors in the automatically generated robot actions. Experiments show that our proposed interactive replanning model trained in a multitask manner using action sequence and description outperformed the baseline model in all types of scores.
Chiori Hori, Motonari Kambara, Komei Sugiura, Kei Ota, Sameer Khurana, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux
ICASSP3
2025 Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images
abstract
Accurate, reliable solar flare prediction is crucial for mitigating potential disruptions to critical infrastructure, while predicting solar flares remains a significant challenge. Existing methods based on heuristic physical features often lack representation learning from solar images. On the other hand, end-to-end learning approaches struggle to model long-range temporal dependencies in solar images. In this study, we propose Deep Space Weather Model (Deep SWM), which is based on multiple deep state space models for handling both ten-channel solar images and long-range spatio-temporal dependencies. Deep SWM also features a sparse masked autoencoder, a novel pretraining strategy that employs a two-phase masking approach to preserve crucial regions such as sunspots while compressing spatial information. Furthermore, we built FlareBench, a new public benchmark for solar flare prediction covering a full 11-year solar activity cycle, to validate our method. Our method outperformed baseline methods and even human expert performance on standard metrics in terms of performance and reliability. The project page can be found at https://keio-smilab25.github.io/DeepSWM.
Shunya Nagashima, Komei Sugiura
ICCV2
2024 Deneb: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
Kazuki Matsuda, Yuiga Wada, Komei Sugiura
ACCV (3)3
2024 Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
abstract
Establishing an automatic evaluation metric that closely aligns with human judgments is essential for effectively developing image captioning models. Recent data-driven metrics have demonstrated a stronger correlation with human judgments than classic metrics such as CIDEr; however they lack sufficient capabilities to handle hallucinations and generalize across diverse images and texts partially because they compute scalar similarities merely using embeddings learned from tasks unrelated to image captioning evaluation. In this study, we propose Polos, a supervised automatic evaluation metric for image captioning models. Polos computes scores from multimodal inputs, using a parallel feature extraction mechanism that leverages embeddings trained through large-scale contrastive learning. To train Polos, we introduce Multimodal Metric Learning from Human Feedback (M2LHF), a framework for developing metrics based on human feedback. We constructed the Polaris dataset, which comprises 131K human judgments from 550 evaluators, which is approximately ten times larger than standard datasets. Our approach achieved state-of-the-art performance on Composite, Flickr8K-Expert, Flickr8K-CF, PASCAL-50S, FOIL, and the Polaris dataset, thereby demonstrating its effectiveness and robustness.
Yuiga Wada, Kanta Kaneda, Daichi Saito, Komei Sugiura
CVPR4
2024 Layer-Wise Relevance Propagation with Conservation Property for ResNet
Seitaro Otsuki, Tsumugi Iida, Félix Doublet, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura
ECCV (43)7
2024 Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models
abstract
We consider the task of generating segmentation masks for the target object from an object manipulation instruction, which allows users to give open vocabulary instructions to domestic service robots. Conventional segmentation generation approaches often fail to account for objects outside the camera’s field of view and cases in which the order of vertices differs but still represents the same polygon, which leads to erroneous mask generation. In this study, we propose a novel method that generates segmentation masks from open vocabulary instructions. We implement a novel loss function using optimal transport to prevent significant loss where the order of vertices differs but still represents the same polygon. To evaluate our approach, we constructed a new dataset based on the REVERIE dataset and Matterport3D dataset. The results demonstrated the effectiveness of the proposed method compared with existing mask generation methods. Remarkably, our best model achieved a +16.32% improvement on the dataset compared with a representative polygon-based method.
Takayuki Nishimura, Katsuyuki Kuyo, Motonari Kambara, Komei Sugiura
IROS4
2023 JaSPICE: Automatic Evaluation Metric Using Predicate-Argument Structures for Image Captioning Models
abstract
Image captioning studies heavily rely on automatic evaluation metrics such as BLEU and METEOR.However, such n-gram-based metrics have been shown to correlate poorly with human evaluation, leading to the proposal of alternative metrics such as SPICE for English; however, no equivalent metrics have been established for other languages.Therefore, in this study, we propose an automatic evaluation metric called JaSPICE, which evaluates Japanese captions based on scene graphs.The proposed method generates a scene graph from dependencies and the predicate-argument structure, and extends the graph using synonyms.We conducted experiments employing 10 image captioning models trained on STAIR Captions and PFN-PIC and constructed the Shichimi dataset, which contains 103,170 human evaluations.The results showed that our metric outperformed the baseline metrics for the correlation coefficient with the human evaluation.
Yuiga Wada, Kanta Kaneda, Komei Sugiura
CoNLL3
2023 Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation Instructions
abstract
In this study, we aim to develop a model that comprehends a natural language instruction (e.g., “Go to the living room and get the nearest pillow to the radio art on the wall”) and generates a segmentation mask for the target everyday object. The task is challenging because it requires (1) the understanding of the referring expressions for multiple objects in the instruction, (2) the prediction of the target phrase of the sentence among the multiple phrases, and (3) the generation of pixel-wise segmentation masks rather than bounding boxes. Studies have been conducted on language-based segmentation methods; however, they sometimes mask irrelevant regions for complex sentences. In this paper, we propose the Multimodal Diffusion Segmentation Model (MDSM), which generates a mask in the first stage and refines it in the second stage. We introduce a crossmodal parallel feature extraction mechanism and extend diffusion probabilistic models to handle crossmodal features. To validate our model, we built a new dataset based on the well-known Matterport3D and REVERIE datasets. This dataset consists of instructions with complex referring expressions accompanied by real indoor environmental images that feature various target objects, in addition to pixel-wise segmentation masks. The performance of MDSM surpassed that of the baseline method by a large margin of +10.13 mean IoU.
Yui Iioka, Yu Yoshida, Yuiga Wada, Shumpei Hatanaka, Komei Sugiura
IROS5
2023 Switching Head-Tail Funnel UNITER for Dual Referring Expression Comprehension with Fetch-and-Carry Tasks
abstract
This paper describes a domestic service robot (DSR) that fetches everyday objects and carries them to specified destinations according to free-form natural language instructions. Given an instruction such as “Move the bottle on the left side of the plate to the empty chair,” the DSR is expected to identify the bottle and the chair from multiple candidates in the environment and carry the target object to the destination. Most of the existing multimodal language understanding methods are impractical in terms of computational complexity because they require inferences for all combinations of target object candidates and destination candidates. We propose Switching Head-Tail Funnel UNITER, which solves the task by predicting the target object and the destination individually using a single model. Our method is validated on a dataset based on a standard dataset for Vision-and-Language Navigation with object manipulation tasks. The results show that our method outperforms the baseline method in terms of language comprehension accuracy. Furthermore, we conduct physical experiments in which a DSR delivers standardized everyday objects in a standardized domestic environment as requested by instructions with referring expressions. The experimental results show that the object grasping and placing actions are achieved with success rates of more than 90 %.
Ryosuke Korekata, Motonari Kambara, Yu Yoshida, Shintaro Ishikawa, Yosuke Kawasaki, Masaki Takahashi 0001, Komei Sugiura
IROS7
2023 Prototypical Contrastive Transfer Learning for Multimodal Language Understanding
abstract
Although domestic service robots are expected to assist individuals who require support, they cannot currently interact smoothly with people through natural language. For example, given the instruction “Bring me a bottle from the kitchen,” it is difficult for such robots to specify the bottle in an indoor environment. Most conventional models have been trained on real-world datasets that are labor-intensive to collect, and they have not fully leveraged simulation data through a transfer learning framework. In this study, we propose a novel transfer learning approach for multimodal language understanding called Prototypical Contrastive Transfer Learning (PCTL), which uses a new contrastive loss called Dual ProtoNCE. We introduce PCTL to the task of identifying target objects in domestic environments according to free-form natural language instructions. To validate PCTL, we built new real-world and simulation datasets. Our experiment demonstrated that PCTL outperformed existing methods. Specifically, PCTL achieved an accuracy of 78.1 %, whereas simple fine-tuning achieved an accuracy of 73.4 %.
Seitaro Otsuki, Shintaro Ishikawa, Komei Sugiura
IROS3
2022 Visual Explanation Generation Based on Lambda Attention Branch Networks
Tsumugi Iida, Takumi Komatsu, Kanta Kaneda, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura
ACCV (2)7
2022 Flare Transformer: Solar Flare Prediction Using Magnetograms and Sunspot Physical Features
Kanta Kaneda, Yuiga Wada, Tsumugi Iida, Naoto Nishizuka, Yûki Kubo, Komei Sugiura
ACCV (2)6
2022 Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass Estimation
abstract
For human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-based geometric algorithm to estimate 3D models of containers for the estimation of their capacity and dimensions. We further use these estimations to estimate their mass in a Convolutional Neural Network model. Experiments show that our Transformer model produced encouraging results in both estimations. While challenges remain in our mask-based algorithm and Convolutional Neural Network model, their results revealed several ways for improvement.
Tomoya Matsubara, Seitaro Otsuki, Yuiga Wada, Haruka Matsuo, Takumi Komatsu, Yui Iioka, Komei Sugiura, Hideo Saito 0001
ICASSP7
2022 Relational Future Captioning Model for Explaining Likely Collisions in Daily Tasks
abstract
Domestic service robots that support daily tasks are a promising solution for elderly or disabled people. It is crucial for domestic service robots to explain the collision risk before they perform actions. In this paper, our aim is to generate a caption about a future event. We propose the Relational Future Captioning Model (RFCM), a crossmodal language generation model for the future captioning task. The RFCM has the Relational Self-Attention Encoder to extract the relationships between events more effectively than the conventional self-attention in transformers. We conducted comparison experiments, and the results show the RFCM outperforms a baseline method on two datasets.
Motonari Kambara, Komei Sugiura
ICIP2
2022 Moment-based Adversarial Training for Embodied Language Comprehension
abstract
In this paper, we focus on a vision-and-language task in which a robot is instructed to execute household tasks. Given an instruction such as "Rinse off a mug and place it in the coffee maker," the robot is required to locate the mug, wash it, and put it in the coffee maker. This is challenging because the robot needs to break down the instruction sentences into subgoals and execute them in the correct order. On the ALFRED benchmark, the performance of state-of-the-art methods is still far lower than that of humans. This is partially because existing methods sometimes fail to infer subgoals that are not explicitly specified in the instruction sentences. We propose Moment-based Adversarial Training (MAT), which uses two types of moments for perturbation updates in adversarial training. We introduce MAT to the embedding spaces of the instruction, subgoals, and state representations to handle their varieties. We validated our method on the ALFRED benchmark, and the results demonstrated that our method outperformed the baseline method for all the metrics on the benchmark.
Shintaro Ishikawa, Komei Sugiura
ICPR2
2021 Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual Dialogue
abstract
Building an interactive artificial intelligence that can ask questions about the real world is one of the biggest challenges for vision and language problems. In particular, goal-oriented visual dialogue, where the aim of the agent is to seek information by asking questions during a turn-taking dialogue, has been gaining scholarly attention recently. While several existing models based on the GuessWhat?! dataset [10] have been proposed, the Questioner typically asks simple category-based questions or absolute spatial questions. This might be problematic for complex scenes where the objects share attributes, or in cases where descriptive questions are required to distinguish objects. In this paper, we propose a novel Questioner architecture, called Unified Questioner Transformer (UniQer), for descriptive question generation with referring expressions. In addition, we build a goal-oriented visual dialogue task called CLEVR Ask. It synthesizes complex scenes that require the Questioner to generate descriptive questions. We train our model with two variants of CLEVR Ask datasets. The results of the quantitative and qualitative evaluations show that UniQer outperforms the baseline.
Shoya Matsumori, Kosuke Shingyouchi, Yuki Abe 0002, Yosuke Fukuchi, Komei Sugiura, Michita Imai
ICCV5
2021 Visual Explanation using Attention Mechanism in Actor-Critic-based Deep Reinforcement Learning
abstract
Deep reinforcement learning (DRL) has great potential for acquiring the optimal action in complex environments such as games and robot control. However, it is difficult to analyze the decision-making of the agent, i.e., the reasons it selects the action acquired by learning. In this work, we propose Mask-Attention A3C (Mask A3C), which introduces an attention mechanism into Asynchronous Advantage Actor-Critic (A3C), which is an actor-critic-based DRL method, and can analyze the decision-making of an agent in DRL. A3C consists of a feature extractor that extracts features from an image, a policy branch that outputs the policy, and a value branch that outputs the state value. In this method, we focus on the policy and value branches and introduce an attention mechanism into them. The attention mechanism applies a mask processing to the feature maps of each branch using mask-attention that expresses the judgment reason for the policy and state value with a heat map. We visualized mask-attention maps for games on the Atari 2600 and found we could easily analyze the reasons behind an agent's decision-making in various game tasks. Furthermore, experimental results showed that the agent could achieve a higher performance by introducing the attention mechanism.
Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura
IJCNN5
2017 Grounded language understanding for manipulation instructions using GAN-based classification
abstract
The target task of this study is grounded language understanding for domestic service robots (DSRs). In particular, we focus on instruction understanding for short sentences where verbs are missing. This task is of critical importance to build communicative DSRs because manipulation is essential for DSRs. Existing instruction understanding methods usually estimate missing information only from non-grounded knowledge; therefore, whether the predicted action is physically executable or not was unclear. In this paper, we present a grounded instruction understanding method to estimate appropriate objects given an instruction and situation. We extend the Generative Adversarial Nets (GAN) and build a GAN-based classifier using latent representations. To quantitatively evaluate the proposed method, we have developed a data set based on the standard data set used for visual question answering (VQA). Experimental results have shown that the proposed method gives the better result than baseline methods.
Komei Sugiura, Hisashi Kawai
ASRU1
2017 Sentence Selection Based on Extended Entropy Using Phonetic and Prosodic Contexts for Statistical Parametric Speech Synthesis
abstract
This paper proposes a sentence selection technique for constructing phonetically and prosodically balanced compact recording scripts for speech synthesis. In the conventional corpus design of speech synthesis, a greedy algorithm that maximizes phonetic coverage is often used. However, for statistical parametric speech synthesis, balances of multiple phonetic and prosodic contextual factors are important as well as the coverage. To take account of both of the phonetic and prosodic contextual balances in sentence selection, we introduce an extended entropy of phonetic and prosodic contexts, such as biphone/triphone, accent/stress/tone, and sentence length. For detailed investigation, conventional and proposed techniques are evaluated using Japanese, English, and Chinese corpora. The objective experimental results show that the proposed technique achieves better coverage and balance of contexts. In addition, speech synthesis experiments based on hidden Markov models reveal that the generated speech parameters become closer to those of the natural speech compared with other conventional sentence selection techniques. Subjective evaluations show that the proposed sentence selection based on the extended entropy improves the naturalness of the synthetic speech while maintaining the similarity to the original sample.
Takashi Nose, Yusuke Arao, Takao Kobayashi, Komei Sugiura, Yoshinori Shiga
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Analysis of Long-Term and Large-Scale Experiments on Robot Dialogues Using a Cloud Robotics Platform
abstract
To build conversational robots, roboticists are required to have deep knowledge of both robotics and spoken dialogue systems. Unlike using stand-alone speech recognition/ synthesis toolkits, a cloud robotics platform for human-robot communication enables high-quality speech recognition and synthesis that is optimized to human-robot interactions. This is challenging because we need to build a wide variety of functionalities ranging from a stable cloud platform to high-quality multilingual speech recognition and synthesis engines. From this background, we constructed rospeex [1], which is a cloud robotics platform for multilingual spoken dialogues with robots. Over 20,000 unique users have used rospeex in the two years since it was launched. In this paper, we propose a method to reduce the response time in rospeex; and analyze its effectiveness. We also analyze the server logs of rospeex that we have collected.
Komei Sugiura, Koji Zettsu
HRI1
2016 Space-time multiple regression model for grid-based population estimation in urban areas
abstract
We can collect, store, and analyze a huge amount of information about human mobility and social interaction activities due to the emergence of information and communication technologies and location-enabled mobile devices under cyber physical system frameworks. The high spatial resolution of population data on a multi-temporal scale is required by transport planners, human geographers, social scientists, and emergency management teams. In this study, we build a space-time multiple regression model to estimate grid-based (500 m × 500 m) spatial resolution at multi-temporal scale (30-min intervals) population data based on the space-time relationship among geospatially enabled person trip (PT) survey data and incorporate both mobile call (MC) and geotagged Twitter (GT) data. Since using geospatially enabled PT survey data as dependent variables enables us to acquire actual population amounts, which strongly depend on MCs and social interaction activities. Although many grids have a strong correlation between PT and MC/GT, some show fewer correlation results, especially where the grids have factories, schools, and workshops in which fewer MCs are found but a large population is presented. Although GT data are sparser than MCs, people from amusement and tourist areas can be detected by GT data. The space-time multiple regression model can also estimate the different amounts of populations based on human travel behavior that changes over space and time. According to accuracy assessments, the night-time estimated results, especially between 00:00 and 06:30, strongly correlate with national census data except in places where the grids have railway and subway stations.
KoKo Lwin, Komei Sugiura, Koji Zettsu
Int. J. Geogr. Inf. Sci.2
2016 Dynamically pre-trained deep recurrent neural networks using environmental monitoring data for predicting PM2.5
abstract
Fine particulate matter ([Formula: see text]) has a considerable impact on human health, the environment and climate change. It is estimated that with better predictions, US$9 billion can be saved over a 10-year period in the USA (State of the science fact sheet air quality. http://www.noaa.gov/factsheets/new, 2012). Therefore, it is crucial to keep developing models and systems that can accurately predict the concentration of major air pollutants. In this paper, our target is to predict [Formula: see text] concentration in Japan using environmental monitoring data obtained from physical sensors with improved accuracy over the currently employed prediction models. To do so, we propose a deep recurrent neural network (DRNN) that is enhanced with a novel pre-training method using auto-encoder especially designed for time series prediction. Additionally, sensors selection is performed within DRNN without harming the accuracy of the predictions by taking advantage of the sparsity found in the network. The numerical experiments show that DRNN with our proposed pre-training method is superior than when using a canonical and a state-of-the-art auto-encoder training method when applied to time series prediction. The experiments confirm that when compared against the [Formula: see text] prediction system VENUS (National Institute for Environmental Studies. Visual Atmospheric Environment Utility System. http://envgis5.nies.go.jp/osenyosoku/, 2014), our technique improves the accuracy of [Formula: see text] concentration level predictions that are being reported in Japan.
Bun Theang Ong, Komei Sugiura, Koji Zettsu
Neural Comput. Appl.2
2015 Constrained region selection method based on configuration space for visualization in scientific dataset search
abstract
We consider constrained label placement problem considering touch interface such as smartphone or tablet. For scientific dataset search, the search results are shown on the global map based on spatial information. There spatial region are often unevenly distributed and most of them are overlapped each other. To select non-overlapped regions from overlapped regions can be considered as a combinational optimization problem which is known as NP-hard. Also, the applications for touch interface should be designed considering finger pad size. In this paper, we propose rectangular label placement method in configuration space for scientific dataset search with touch interface. The idea of configuration is often used for path planning for robots to avoid obstacles. The proposed method apply the idea of configuration space to find the un-overlapped regions to show the selectable regions for touch interface. Furthermore, the relevance between search query and each datasets are considered, so that the user can select more relevant datasets from original search results. This method can be applied not only for scientific dataset search system but also any other applications which shows their results in global map. The experimental evaluations show that the proposed method achieves superior performance to the compared methods.
Shin'ichi Takeuchi, Komei Sugiura, Yuhei Akahoshi, Koji Zettsu
IEEE BigData2
2015 Entropy-based sentence selection for speech synthesis using phonetic and prosodic contexts
Takashi Nose, Yusuke Arao, Takao Kobayashi, Komei Sugiura, Yoshinori Shiga, Akinori Ito
INTERSPEECH4
2015 Rospeex: A cloud robotics platform for human-robot spoken dialogues
abstract
To build conversational robots, roboticists are required to have deep knowledge of both robotics and spoken dialogue systems. Although they can use existing cloud services that were built for other services, e.g., voice search, it will be difficult to share robotics-specific speech corpora obtained as server logs, because they will get buried in non-robotics-related logs. Building a cloud platform especially for the robotics community will benefit not only individual robot developers but also the robotics community since we can share the log corpus collected by it. This is challenging because we need to build a wide variety of functionalities ranging from a stable cloud platform to high-quality multilingual speech recognition and synthesis engines. In this paper, we propose “rospeex,” which is a cloud robotics platform for multilingual spoken dialogues with robots. We analyze the logs we have collected by operating rospeex for more than a year. Our key contribution lies in building a cloud robotics platform and allowing the robotics community to use it without payment or authentication.
Komei Sugiura, Koji Zettsu
IROS1
2015 RoboCup@Home: Analysis and results of evolving competitions for domestic and service robots
abstract
Scientific competitions are becoming more common in many research areas of artificial intelligence and robotics, since they provide a shared testbed for comparing different solutions and enable the exchange of research results. Moreover, they are interesting for general audiences and industries. Currently, many major research areas in artificial intelligence and robotics are organizing multiple-year competitions that are typically associated with scientific conferences. One important aspect of such competitions is that they are organized for many years. This introduces a temporal evolution that is interesting to analyze. However, the problem of evaluating a competition over many years remains unaddressed. We believe that this issue is critical to properly fuel changes over the years and measure the results of these decisions. Therefore, this article focuses on the analysis and the results of evolving competitions. In this article, we present the [email protected] competition, which is the largest worldwide competition for domestic service robots, and evaluate its progress over the past seven years. We show how the definition of a proper scoring system allows for desired functionalities to be related to tasks and how the resulting analysis fuels subsequent changes to achieve general and robust solutions implemented by the teams. Our results show not only the steadily increasing complexity of the tasks that [email protected] robots can solve but also the increased performance for all of the functionalities addressed in the competition. We believe that the methodology used in [email protected] for evaluating competition advances and for stimulating changes can be applied and extended to other robotic competitions as well as to multi-year research projects involving Artificial Intelligence and Robotics.
Luca Iocchi, Dirk Holz, Javier Ruiz-del-Solar, Komei Sugiura, Tijn van der Zant
Artif. Intell.4
2014 Dynamic pre-training of Deep Recurrent Neural Networks for predicting environmental monitoring data
abstract
In this paper, we introduce a Deep Recurrent Neural Network (DRNN) that is trained using a novel autoencoder pre-training method especially designed for the task of time series prediction. Our main objective is to perform predictions of environmental monitoring data using open sensors with improved accuracy over the currently employed methods. The numerical experiments show that our proposed pre-training method is superior that a canonical and a state-of-the-art auto-encoder training method when applied to time series prediction. On the specific case of fine particulate matter (PM2.5) forecasting in Japan, the experiments confirm that when compared against the PM2.5prediction system VENUS employed by the Japanese Government, our technique improves the accuracy of PM2.5concentration level predictions that are being reported in Japan.
Bun Theang Ong, Komei Sugiura, Koji Zettsu
IEEE BigData2
2014 A new dimension for RoboCup @home: human-robot interaction between virtual and real worlds
abstract
This work proposes a new approach to realize embodied and multimodal HRI between virtual robot and real world human for HRI challenges in RoboCup @Home.
Jeffrey Too Chuan Tan, Tetsunari Inamura, Yoshinobu Hagiwara, Komei Sugiura, Takayuki Nagai
HRI4
2014 Non-monologue HMM-based speech synthesis for service robots: A cloud robotics approach
abstract
Robot utterances generally sound monotonous, unnatural, and unfriendly because their Text-to-Speech (TTS) systems are not optimized for communication but for text-reading. Here we present a non-monologue speech synthesis for robots. We collected a speech corpus in a non-monologue style in which two professional voice talents read scripted dialogues. Hidden Markov models (HMMs) were then trained with the corpus and used for speech synthesis. We conducted experiments in which the proposed method was evaluated by 24 subjects in three scenarios: text-reading, dialogue, and domestic service robot (DSR) scenarios. In the DSR scenario, we used a physical robot and compared our proposed method with a baseline method using the standard Mean Opinion Score (MOS) criterion. Our experimental results showed that our proposed method's performance was (1) at the same level as the baseline method in the text-reading scenario and (2) exceeded it in the DSR scenario. We deployed our proposed system as a cloud-based speech synthesis service so that it can be used without any cost.
Komei Sugiura, Yoshinori Shiga, Hisashi Kawai, Teruhisa Misu, Chiori Hori
ICRA1
2014 On RoboCup@Home - Past, Present and Future of a Scientific Competition for Service Robots
Dirk Holz, Javier Ruiz-del-Solar, Komei Sugiura, Sven Wachsmuth
RoboCup3
2013 Complementary Integration of Heterogeneous Crowd-Sourced Datasets for Enhanced Social Analytics
abstract
On behalf of the rapidly and widely disseminated smartphone technology into the public, lots of social network sites and location-based social applications are accumulating a huge volume of massive crowd's daily experiences and thoughts in an unprecedented scale. We can regard them as novel data sources for accomplishing various social analytics, which have usually required lots of efforts to collect crowds' opinion and behavioral data. Thus, we can take advantages of abundant social datasets by integrating them appropriately. However, when we integrate disparate sources to derive a comprehensive view for a survey, it is necessary to know intrinsic exclusive values of each data source compared to others in an intuitive and succinct way. In fact, lots of efforts and time are wasted to overview various datasets consequently to confidently choose a dataset to be integrated in a final result. In this paper, we propose a complementarity index, which can estimate the exclusive usefulness of data sources in terms of spatial and topical coverage when selecting data sources for social analytics purposes. We conducted an experiment about complementarity measurement with two real social datasets from Twitter and VoiceTra; the latter is a speech-to-speech translation app, with which we can additionally obtain crowds' verbal translation logs. With the proposed complementarity index, we can measure the capability of a dataset comparing to others before integrating datasets, thus enabling analysts to examine much more datasets from as many related data sources as possible by focusing on exclusive coverage and relative strength of relevant topics.
Ryong Lee, Kyoung-Sook Kim 0001, Komei Sugiura, Koji Zettsu, Yutaka Kidawara
MDM (2)3
2013 Utterance Classification Using Linguistic and Non-linguistic Information for Network-Based Speech-to-Speech Translation Systems
abstract
Network-based mobile services, such as speech-to-speech translation and voice search, enable the construction of large-scale log database including speech. We have developed a smartphone application called VoiceTra for speech-to-speech translation and have collected 10,000,000 utterances so far. This huge corpus is unique in size and spatio-temporal information; it contains information on anonymized user locations. This spatiotemporal corpus can be used for improving the accuracy of its speech recognition and machine translation, and it will open the door for the study of the location dependency of vocabulary and new applications for location-based services. This paper first analyzes the corpus and then presents a novel method for classifying utterances using linguistic and non-linguistic information. L2-regularized Logistic Regression is used for utterance classification. Our experiments performed on the VoiceTra log corpus revealed that our proposed method outperformed baseline methods in terms of F measure.
Komei Sugiura, Ryong Lee, Hideki Kashioka, Koji Zettsu, Yutaka Kidawara
MDM (2)1
2013 Development of RoboCup@Home Simulation towards Long-term Large Scale HRI
Tetsunari Inamura, Jeffrey Too Chuan Tan, Komei Sugiura, Takayuki Nagai
RoboCup3
2011 Motion generation by reference-point-dependent trajectory HMMs
abstract
This paper presents an imitation learning method for object manipulation such as rotating an object or placing one object on another. In the proposed method, motions are learned using reference-point-dependent probabilistic models. Trajectory hidden Markov models (HMMs) are used as the probabilistic models so that smooth trajectories can be generated from the HMMs. The method was evaluated in physical experiments in terms of motion generation. In the experiments, a robot learned motions from observation, and it generated motions under different object placement. Experimental results showed that appropriate motions were generated even when the object placement was changed.
Komei Sugiura, Naoto Iwahashi, Hideki Kashioka
IROS1
2010 Robot-directed speech detection using Multimodal Semantic Confidence based on speech, image, and motion
abstract
In this paper, we propose a novel method to detect robot-directed (RD) speech that adopts the Multimodal Semantic Confidence (MSC) measure. The MSC measure is used to decide whether the speech can be interpreted as a feasible action under the current physical situation in an object manipulation task. This measure is calculated by integrating speech, image, and motion confidence measures with weightings that are optimized by logistic regression. Experimental results show that, compared with a baseline method that uses speech confidence only, MSC achieved an absolute increase of 5% for clean speech and 12% for noisy speech in terms of average maximum F-measure.
Xiang Zuo, Naoto Iwahashi, Ryo Taguchi, Shigeki Matsuda, Komei Sugiura, Kotaro Funakoshi, Mikio Nakano, Natsuki Oka
ICASSP5
2010 Learning novel objects using out-of-vocabulary word segmentation and object extraction for home assistant robots
abstract
This paper presents a method for learning novel objects from audio-visual input. Objects are learned using out-of-vocabulary word segmentation and object extraction. The latter half of this paper is devoted to evaluations. We propose the use of a task adopted from the RoboCup@Home league as a standard evaluation for real world applications. We have implemented proposed method on a real humanoid robot and evaluated it through a task called “Supermarket”. The results reveal that our integrated system works well in the real application. In fact, our robot outperformed the maximum score obtained in RoboCup@Home 2009 competitions.
Muhammad Attamimi, Attamini Mizutani, Tomoaki Nakamura, Komei Sugiura, Takayuki Nagai, Naoto Iwahashi, Takashi Omori
ICRA4
2010 Active learning of confidence measure function in robot language acquisition framework
abstract
In an object manipulation dialogue, a robot may misunderstand an ambiguous command from a user, such as “Place the cup down (on the table),” potentially resulting in an accident. Although making confirmation questions before all motion will decrease the risk of this failure, the user will find it more convenient if confirmation questions are not made under trivial situations. This paper proposes a method for estimating ambiguity in the commands by introducing an active learning framework with Bayesian logistic regression to human-robot spoken dialogue. We conducted physical experiments in which a user and a manipulator-based robot communicated in spoken language to manipulate toys.
Komei Sugiura, Naoto Iwahashi, Hideki Kashioka, Satoshi Nakamura 0001
IROS1
2010 Detecting robot-directed speech by situated understanding in object manipulation tasks
abstract
In this paper, we propose a novel method for a robot to detect robot-directed speech, that is, to distinguish speech that users speak to a robot from speech that users speak to other people or to themselves. The originality of this work is the introduction of a multimodal semantic confidence (MSC) measure, which is used for domain classification of input speech based on the decision on whether the speech can be interpreted as a feasible action under the current physical situation in an object manipulation task. This measure is calculated by integrating speech, object, and motion confidence with weightings that are optimized by logistic regression. Then we integrate this measure with gaze tracking and conduct experiments under conditions of natural human-robot interaction. Experimental results show that the proposed method achieves a high performance of 94% and 96% in average recall and precision rates, respectively, for robot-directed speech detection.
Xiang Zuo, Naoto Iwahashi, Ryo Taguchi, Kotaro Funakoshi, Mikio Nakano, Shigeki Matsuda, Komei Sugiura, Natsuki Oka
RO-MAN7
2010 Modeling Spoken Decision Making Dialogue and Optimization of its Dialogue Strategy
Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001
SIGDIAL Conference2
2010 Dialogue strategy optimization to assist user's decision for spoken consulting dialogue systems
abstract
This paper addresses a user model and dialogue state definition in spoken consulting dialogue systems that help users in making decision. When selecting from a set of alternatives, users have various decision criteria for making decision. Users often do not have a definite goal or criteria for selection, and thus they may find not only what kind of information the system can provide but their own preference or factors that they should emphasize. In this paper, we model such consulting dialogue as partially observable Markov decision process (POMDP). We then present an optimization of dialogue strategy to help users make better decisions.
Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001
SLT2
2009 Bayesian learning of confidence measure function for generation of utterances and motions in object manipulation dialogue task
abstract
This paper proposes a method that generates motions and utterances in an object manipulation dialogue task. The proposed method integrates belief modules for speech, vision, and motions into a probabilistic framework so that a user’s utterances can be understood based on multimodal information. Responses to the utterances are optimized based on an integrated confidence measure function for the integrated belief modules. Bayesian logistic regression is used for the learning of the confidence measure function. The experimental results revealed that the proposed method reduced the failure rate from 12% down to 2.6% while the rejection rate was less than 24%. Index Terms: multimodal spoken dialogue system, robot language acquisition, confidence, Bayesian logistic regression
Komei Sugiura, Naoto Iwahashi, Hideki Kashioka, Satoshi Nakamura 0001
INTERSPEECH1
2008 Motion recognition and generation by combining reference-point-dependent probabilistic models
abstract
This paper presents a method to recognize and generate sequential motions for object manipulation such as placing one object on another or rotating it. Motions are learned using reference-point-dependent probabilistic models, which are then transformed to the same coordinate system and combined for motion recognition/generation. We conducted physical experiments in which a user demonstrated the manipulation of puppets and toys, and obtained a recognition accuracy of 63% for the sequential motions. Furthermore, the results of motion generation experiments performed with a robot arm are presented.
Komei Sugiura, Naoto Iwahashi
IROS1
2005 Exploiting interaction between sensory morphology and learning
abstract
This paper proposes a system that automatically designs the sensory morphology of an autonomous robot. This system uses two kinds of adaptation, ontogenetic adaptation and phylogenetic adaptation, to optimize the sensory morphology of the robot, in ontogenetic adaptation, individuals with many different sensory morphologies use reinforcement learning to adapt to a task. In phylogenetic adaptation, a genetic algorithm is used to select morphologies with which the robot can learn the task fasten We made the system design a line-following robot, and carried out experiments to compare the design solution with a hand-coded design. The results have shown that the designed robot outperforms the hand-coded design in terms of line-following accuracy and learning speed, although it has fewer sensors than hand-coded robots. The paper also shows the effective use of sensory morphology obtained by our system.
Komei Sugiura, Makoto Akahane, Takayuki Shiose, Katsunori Shimohara, Osamu Katai
SMC1