VLDB 2026 Research / reviewers in the wild / expert
Kotaro Funakoshi
dblp:21/3705
· DBLP profile ↗
75ranked-venue papers
12as first author
24since 2021 · last 2026
0000-0002-4529-4634ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 10 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 19 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9Systems, architecture and hardware · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causal2Vec: Improving Decoder-only LLMs as Embedding Models through a Contextual TokenabstractDecoder-only large language models (LLMs) have been increasingly adopted to build embedding models for diverse tasks.To overcome the inherent limitations of causal attention in representation learning, many existing methods modify the attention mechanism to be bidirectional, potentially undermining LLMs' ability to extract semantic information acquired during pretraining.Meanwhile, leading unidirectional approaches often rely on extra input text to generate contextualized embeddings, inevitably increasing computational costs.In this work, we propose Causal2Vec, a general-purpose embedding model tailored to enhance the performance of decoder-only LLMs without altering their original architectures or introducing significant computational overhead.Specifically, we first employ a lightweight BERT-style model to preencode the input text into a single Contextual token, which is then prepended to the LLM's input sequence, allowing each token to capture contextualized information even without attending to future tokens.Furthermore, to mitigate the recency bias introduced by last-token pooling, we concatenate the last hidden states of Contextual and EOS tokens as the final text embedding.In practice, Causal2Vec achieves a new state-of-the-art performance on the MTEB benchmark among models trained solely on publicly available retrieval datasets. Ailiang Lin, Zhuoyun Li, Yusong Wang 0003, Kotaro Funakoshi, Manabu Okumura |
ACL (1) | 4 |
| 2026 | Retrieving Responses Useful as References in Counseling via LLM-Generated Structurally Analogous DialoguesabstractRetrieving past similar dialogues to assist response generation in the present context is a promising approach to addressing the shortage of highly experienced professionals in various domains, such as user support and mental health counseling. However, unlike traditional retrieval, this task requires finding dialogue histories that lead to useful subsequent responses—a property we define as referenceability—rather than relying solely on superficial topic matching. Learning this property directly is difficult due to the impracticality of manually annotating large-scale data. To address this, we propose a framework that leverages large language models to generate pseudo-similar counseling dialogues with strict message-level correspondence to original sessions. We then use these generated dialogues as weak supervision for contrastive learning of an embedding model. By pairing original dialogue histories with their generated counterparts, the model learns to pull together histories that share referenceable subsequent counselor responses, even if their explicit situations differ. Experimental results show that our fine-tuned model significantly improves retrieval performance, increasing Hit@1 from 0.167 to 0.500 and MRR from 0.232 to 0.576. Human evaluations also confirm that the model retrieves more referenceable responses, improving the average score from 2.22 to 2.66. Yu Nakagawa, Nozomu Ikeda, Kotaro Funakoshi, Manabu Okumura |
SIGDIAL | 3 |
| 2025 | Unveiling the Power of Source: Source-based Minimum Bayes Risk Decoding for Neural Machine TranslationabstractMaximum a posteriori decoding, a commonly used method for neural machine translation (NMT), aims to maximize the estimated posterior probability.However, high estimated probability does not always lead to high translation quality.Minimum Bayes Risk (MBR) decoding (Kumar and Byrne, 2004) offers an alternative by seeking hypotheses with the highest expected utility.Inspired by Quality Estimation (QE) reranking which uses the QE model as a ranker (Fernandes et al., 2022), we propose source-based MBR (sMBR) decoding, a novel approach that utilizes quasi-sources (generated via paraphrasing or back-translation) as "support hypotheses" and a reference-free quality estimation metric as the utility function, marking the first work to solely use sources in MBR decoding.Experiments show that sMBR outperforms QE reranking and the standard MBR decoding.Our findings suggest that sMBR is a promising approach for NMT decoding.1 NMT x h 0 Boxuan Lyu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura |
ACL (1) | 3 |
| 2025 | S2N: A Synthetic Data-Driven Approach for Speaker-To-Dialogue Attribution in NovelsabstractThe Speaker-To-Dialogue Attribution (SDA) task targets attributing the speaker for each utterance in the dialogue context from literature (novel) texts. Although previous studies have attempted to optimize model architectures and employ efficient data usage, SDA still faces the severe challenge of data deficiency due to its high annotation cost. To address this issue, following the extractive machine reading comprehension formulation (MRC) of SDA for realistic concerns, we propose the first synthetic data-driven SDA framework, which automatically synthesizes large-scale SDA training samples using drama script texts. Given the inherent differences between novel texts and scripts, we employ an MLM-Cloze synthesis approach to transform script fragments into novel-style paragraphs. Additionally, to mitigate risks introduced by synthetic data-driven training, we also propose two anti-risk strategies, namely the Joint Contrastive Learning Machine Reader framework and the Dynamic Train-Val Data Influence metric. Experiments on existing benchmarks demonstrate the effectiveness of our methodology. Our code and data are released at https://github.com/huangyiqian-h1kk/S2N. Yiqian Huang 0007, Kotaro Funakoshi, Manabu Okumura, Yang Cao 0011 |
IJCNN | 3 |
| 2025 | Key Challenges in Multimodal Task-Oriented Dialogue Systems: Insights from a Large Competition-Based DatasetabstractChallenges in multimodal task-oriented dialogue between humans and systems, particularly those involving audio and visual interactions, have not been sufficiently explored or shared, forcing researchers to define improvement directions individually without a clearly shared roadmap. To address these challenges, we organized a competition for multimodal task-oriented dialogue systems and constructed a large competition-based dataset of 1,865 minutes of Japanese task-oriented dialogues. This dataset includes audio and visual interactions between diverse systems and human participants. After analyzing system behaviors identified as problematic by the human participants in questionnaire surveys and notable methods employed by the participating teams, we identified key challenges in multimodal task-oriented dialogue systems and discussed potential directions for overcoming these challenges. Shiki Sato, Shinji Iwata, Asahi Hentona, Yuta Sasaki, Takato Yamazaki, Shoji Moriya, Masaya Ohagi, Hirofumi Kikuchi, Zhiyang Qi, Takashi Kodama, Akinobu Lee, Masato Komuro, Hiroyuki Nishikawa, Ryosaku Makino, Takashi Minato, Kurima Sakai, Tomo Funayama, Kotaro Funakoshi, Mayumi Usami, Michimasa Inaba, Tetsuro Takahashi, Ryuichiro Higashinaka |
SIGDIAL | 19 |
| 2025 | Analyzing Dialogue System Behavior in a Specific Situation Requiring Interpersonal ConsiderationabstractIn human-human conversation, interpersonal consideration for the interlocutor is essential, and similar expectations are increasingly placed on dialogue systems. This study examines the behavior of dialogue systems in a specific interpersonal scenario where a user vents frustrations and seeks emotional support from a long-time friend represented by a dialogue system. We conducted a human evaluation and qualitative analysis of 15 dialogue systems under this setting. These systems implemented diverse strategies, such as structuring dialogue into distinct phases, modeling interpersonal relationships, and incorporating cognitive behavioral therapy techniques. Our analysis reveals that these approaches contributed to improved perceived empathy, coherence, and appropriateness, highlighting the importance of design choices in socially sensitive dialogue. Tetsuro Takahashi, Hirofumi Kikuchi, Hiroyuki Nishikawa, Masato Komuro, Ryosaku Makino, Shiki Sato, Yuta Sasaki, Shinji Iwata, Asahi Hentona, Takato Yamazaki, Shoji Moriya, Masaya Ohagi, Zhiyang Qi, Takashi Kodama, Akinobu Lee, Takashi Minato, Kurima Sakai, Tomo Funayama, Kotaro Funakoshi, Mayumi Usami, Michimasa Inaba, Ryuichiro Higashinaka |
SIGDIAL | 20 |
| 2024 | myMediCon: End-to-End Burmese Automatic Speech Recognition for Medical ConversationsabstractEnd-to-End Automatic Speech Recognition (ASR) models have significantly advanced the field of speech processing by streamlining traditionally complex ASR system pipelines, promising enhanced accuracy and efficiency. Despite these advancements, there is a notable absence of freely available medical conversation speech corpora for Burmese, which is one of the low-resource languages. Addressing this gap, we present a manually curated Burmese Medical Speech Conversations (myMediCon) corpus, encapsulating conversations among medical doctors, nurses, and patients. Utilizing the ESPnet speech processing toolkit, we explore End-to-End ASR models for the Burmese language, focus on Transformer and Recurrent Neural Network (RNN) architectures. Our corpus comprises 12 speakers, including three males and nine females, with a total speech duration of nearly 11 hours within the medical domain. To assess the ASR performance, we applied word and syllable segmentation to the text corpus. ASR models were evaluated using Character Error Rate (CER), Word Error Rate (WER), and Translation Error Rate (TER). The experimental results indicate that the RNN-based Burmese speech recognition with syllable-level segmentation achieved the best performance, yielding a CER of 9.7%. Moreover, the RNN approach significantly outperformed the Transformer model. Hay Man Htun, Ye Kyaw Thu, Hutchatai Chanlekha, Kotaro Funakoshi, Thepchai Supnithi |
LREC/COLING | 4 |
| 2024 | Can Respiration Make Spoken Interactions Better?abstractWith the growing use of agents such as chatbots in daily life, enhancing human-agent spoken interactions has become a critical area of research. This study explores the potential of respiratory information to enhance these interactions. We implemented two functions: speech collision avoidance based on speech onset prediction using human respiratory information, and pseudo-respiration presentation to humans by the vertical movements of a robot. We developed a conversational robot incorporating these functions and conducted experiments with 26 human participants. The experimental results confirmed these functions contributed to the smoothness of interactions. Takao Obi, Kotaro Funakoshi |
HAI | 2 |
| 2024 | Enhancing Image Clustering with Captions
Yuanyuan Cai, Satoshi Kosugi, Kotaro Funakoshi, Manabu Okumura |
PACLIC | 3 |
| 2024 | LPLS: A Selection Strategy Based on Pseudo-Labeling Status for Semi-Supervised Active Learning in Text Classification
Chun-Fang Chuang, Dongyuan Li, Satoshi Kosugi, Kotaro Funakoshi, Manabu Okumura |
PACLIC | 4 |
| 2024 | FINE-LMT: Fine-Grained Feature Learning for Multi-modal Machine Translation
Ying Zhang 0065, Dongyuan Li, Jialun Shen, Mingkun Xu, Kotaro Funakoshi, Manabu Okumura |
PRICAI (2) | 7 |
| 2024 | Using Respiration for Enhancing Human-Robot DialogueabstractThis paper presents the development and capabilities of a spoken dialogue robot that uses respiration to enhance human-robot dialogue.By employing a respiratory estimation technique that uses video input, the dialogue robot captures user respiratory information during dialogue.This information is then used to prevent speech collisions between the user and the robot and to present synchronized pseudorespiration with the user, thereby enhancing the smoothness and engagement of human-robot dialogue. Takao Obi, Kotaro Funakoshi |
SIGDIAL | 2 |
| 2023 | After: Active Learning Based Fine-Tuning Framework for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) has drawn increasing attention for its applications in human-machine interaction. However, existing SER methods ignore the information gap between the pre-training speech recognition task and the downstream SER task, leading to sub-optimal performance. Moreover, they require much time to fine-tune on each specific speech dataset, restricting their effectiveness in real-world scenes with large-scale noisy data. To address these issues, we propose an active learning (AL) based Fine-Tuning framework for SER that leverages task adaptation pre-training (TAPT) and AL methods to enhance performance and efficiency. Specifically, we first use TAPT to minimize the information gap between the pre-training and the downstream task. Then, AL methods are used to iteratively select a subset of the most informative and diverse samples for fine-tuning, reducing time consumption. Experiments demonstrate that using only 20% pt. samples improves 8.45% pt. accuracy and reduces 79% pt. time consumption. Dongyuan Li, Kotaro Funakoshi, Manabu Okumura |
ASRU | 3 |
| 2023 | Temporal and Topological Augmentation-based Cross-view Contrastive Learning Model for Temporal Link PredictionabstractWith the booming development of social media, temporal link prediction (TLP), as a core technology, has been receiving increasing attention. However, current methods are based on graph neural networks, which suffer from the over-smoothing issue and easily yield indistinguishable node representations, degrading the prediction accuracy. Besides, they lack the ability to eliminate noisy temporal information and ignore the importance of high-order neighbor information for measuring the link probability between nodes. To solve these issues, we design a cross-view graph contrastive learning (GCL) framework for TLP, called Tacl. We first design two augmented views for GCL by enhancing the temporal and topological information to obtain distinguishable node representations. Then, we learn the evolution rule of temporal networks to help constrain consistency of node representations and eliminate noise. Finally, we incorporate the high-order neighbor information to measure the link probability between nodes. Extensive experiments demonstrate the effectiveness and robustness of Tacl. Dongyuan Li, Shiyin Tan, Yusong Wang 0003, Kotaro Funakoshi, Manabu Okumura |
CIKM | 4 |
| 2023 | Generative Replay Inspired by Hippocampal Memory Indexing for Continual Language LearningabstractContinual learning aims to accumulate knowledge to solve new tasks without catastrophic forgetting for previously learned tasks.Research on continual learning has led to the development of generative replay, which prevents catastrophic forgetting by generating pseudosamples for previous tasks and learning them together with new tasks.Inspired by the biological brain, we propose the hippocampal memory indexing to enhance the generative replay by controlling sample generation using compressed features of previous training samples.It enables the generation of a specific training sample from previous tasks, thus improving the balance and quality of generated replay samples.Experimental results indicate that our method effectively controls the sample generation and consistently outperforms the performance of current generative replay methods. 1 Aru Maekawa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura |
EACL | 3 |
| 2023 | Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionabstractMultimodal emotion recognition aims to recognize emotions for each utterance of multiple modalities, which has received increasing attention for its application in human-machine interaction.Current graph-based methods fail to simultaneously depict global contextual features and local diverse uni-modal features in a dialogue.Furthermore, with the number of graph layers increasing, they easily fall into over-smoothing.In this paper, we propose a method for joint modality fusion and graph contrastive learning for multimodal emotion recognition (JOYFUL), where multimodality fusion, contrastive learning, and emotion recognition are jointly optimized.Specifically, we first design a new multimodal fusion mechanism that can provide deep interaction and fusion between the global contextual and uni-modal specific features.Then, we introduce a graph contrastive learning framework with inter-view and intra-view contrastive losses to learn more distinguishable representations for samples with different sentiments.Extensive experiments on three benchmark datasets indicate that JOYFUL achieved state-of-the-art (SOTA) performance compared to all baselines. Dongyuan Li, Kotaro Funakoshi, Manabu Okumura |
EMNLP | 3 |
| 2023 | EMP: Emotion-guided Multi-modal Fusion and Contrastive Learning for Personality Traits RecognitionabstractMulti-modal personality traits recognition aims to recognize personality traits precisely by utilizing different modality information, which has received increasing attention for its potential applications in human-computer interaction. Current methods almost fail to extract distinguishable features, remove noise, and align features from different modalities, which dramatically affects the accuracy of personality traits recognition. To deal with these issues, we propose an emotion-guided multi-modal fusion and contrastive learning framework for personality traits recognition. Specifically, we first use supervised contrastive learning to extract deeper and more distinguishable features from different modalities. After that, considering the close correlation between emotions and personalities, we use an emotion-guided multi-modal fusion mechanism to guide the feature fusion, which eliminates the noise and aligns the features from different modalities. Finally, we use an auto-fusion structure to enhance the interaction between different modalities to further extract essential features for final personality traits recognition. Extensive experiments on two benchmark datasets indicate that our method achieves state-of-the-art performance and robustness. Yusong Wang 0003, Dongyuan Li, Kotaro Funakoshi, Manabu Okumura |
ICMR | 3 |
| 2023 | A Follow-up Study on Evaluation Metrics Using Follow-up Utterances
Toshiki Kawamoto, Yuki Okano, Takato Yamazaki, Kotaro Funakoshi, Manabu Okumura |
PACLIC | 5 |
| 2022 | A-TIP: Attribute-aware Text Infilling via Pre-trained Language ModelabstractText infilling aims to restore incomplete texts by filling in blanks, which has attracted more attention recently because of its wide application in ancient text restoration and text rewriting. However, attribute- aware text infilling is yet to be explored, and existing methods seldom focus on the infilling length of each blank or the number/location of blanks. In this paper, we propose an Attribute-aware Text Infilling method via a Pre-trained language model (A-TIP), which contains a text infilling component and a plug- and-play discriminator. Specifically, we first design a unified text infilling component with modified attention mechanisms and intra- and inter-blank positional encoding to better perceive the number of blanks and the infilling length for each blank. Then, we propose a plug-and-play discriminator to guide generation towards the direction of improving attribute relevance without decreasing text fluency. Finally, automatic and human evaluations on three open-source datasets indicate that A-TIP achieves state-of- the-art performance compared with all baselines. Dongyuan Li, Jingyi You, Kotaro Funakoshi, Manabu Okumura |
COLING | 3 |
| 2022 | Generating Repetitions with Appropriate Repeated WordsabstractToshiki Kawamoto, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Toshiki Kawamoto, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura |
NAACL-HLT | 3 |
| 2022 | Joint Learning-based Heterogeneous Graph Attention Network for Timeline SummarizationabstractJingyi You, Dongyuan Li, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jingyi You, Dongyuan Li, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura |
NAACL-HLT | 4 |
| 2021 | Towards Table-to-Text Generation with Numerical ReasoningabstractLya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, Hiroya Takamura. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, Hiroya Takamura |
ACL/IJCNLP (1) | 3 |
| 2021 | Robust Dynamic Clustering for Temporal NetworksabstractDynamic community detection (or graph clustering) in temporal networks has attracted much attention because it is promising for revealing the underlying mechanism of complex real-world systems. Current methods are criticized for the independence of graph representation learning and graph clustering, considerable noise during temporal information smoothing, and high time complexity. We propose a R obust T emporal S moothing C lustering method (RTSC), which involves joint graph representation learning and graph clustering, to solve these problems. RTSC can be formulated as a constrained multi-objective optimization problem. Specifically, three-order successive snapshots are first projected into the same subspace via graph embedding. We then use the embedding matrices to learn a common low-rank block-diagonal matrix that contains current clustering information and specific noise matrices with a sparse constraint to remove noise at each time step. To efficiently solve the challenging optimization problem, we also propose an optimization procedure based on the augmented Lagrangian multiplier (ALM) scheme. Experimental results on six artificial datasets and four real-world dynamic network datasets indicate that RTSC performs better than six state-of-the-art algorithms for dynamic clustering in temporal networks. Jingyi You, Chenlong Hu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura |
CIKM | 4 |
| 2021 | Generating Weather Comments from Meteorological SimulationsabstractSoichiro Murakami, Sora Tanaka, Masatsugu Hangyo, Hidetaka Kamigaito, Kotaro Funakoshi, Hiroya Takamura, Manabu Okumura. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Soichiro Murakami, Sora Tanaka, Masatsugu Hangyo, Hidetaka Kamigaito, Kotaro Funakoshi, Hiroya Takamura, Manabu Okumura |
EACL | 5 |
| 2020 | Who Speaks Next? Turn Change and Next Speaker Prediction in Multimodal Multiparty InteractionabstractTurn change prediction and next speaker prediction are two important tasks in multimodal, multiparty human-agent interaction. Predicting a change of dialogue turn and the most probable next speaker can help an agent to decide whether he should contribute to the discussion or wait for someone else to speak. In this research, we propose a machine learning-based approach for both turn change and next speaker prediction. Individual as well as combined models are explored to tackle these tasks. Results show that the proposed models outperform baselines. An ablation study is also performed to measure the importance of different features. Usman Malik, Julien Saunier, Kotaro Funakoshi, Alexandre Pauchet |
ICTAI | 3 |
| 2019 | Personal Partner Agents for Cooperative IntelligenceabstractWe advocate cooperative intelligence (CI) that achieves its goals in cooperating with other agents, particularly human beings, with limited resources but in complex and dynamic environments. CI is important because it delivers better performances in achieving a broad range of tasks; furthermore, cooperativeness is key to human intelligence, and the processes of cooperation can contribute to help people gain several life values. This paper discusses elements in CI and our research approach to CI. We identify the four aspects of CI: adaptive intelligence, collective intelligence, coordinative intelligence, and collaborative intelligence. We also take an approach that focuses on the implementation of coordinative intelligence in the form of personal partner agents (PPAs) and consider the design of our robotic research platform to physically realize PPAs. Kotaro Funakoshi, Hideaki Shimazaki, Takatsune Kumada, Hiroshi Tsujino |
HRI | 1 |
| 2019 | Audio-Visual SLAM towards Human Tracking and Human-Robot Interaction in Indoor EnvironmentsabstractWe propose a novel audio-visual simultaneous and localization (SLAM) framework that exploits human pose and acoustic speech of human sound sources to allow a robot equipped with a microphone array and a monocular camera to track, map, and interact with human partners in an indoor environment. Since human interaction is characterized by features perceived in not only the visual modality, but the acoustic modality as well, SLAM systems must utilize information from both modalities. Using a state-of-the-art beamforming technique, we obtain sound components correspondent to speech and noise; and estimate the Direction-of-Arrival (DoA) estimates of active sound sources as useful representations of observed features in the acoustic modality. Through estimated human pose by a monocular camera, we obtain the relative positions of humans as representation of observed features in the visual modality. Using these techniques, we attempt to eliminate restrictions imposed by intermittent speech, noisy periods, reverberant periods, triangulation of sound-source range, and limited visual field-of-views; and subsequently perform early fusion on these representations. We develop a system that allows for complimentary action between audio-visual sensor modalities in the simultaneous mapping of multiple human sound sources and the localization of observer position. Aaron Chau, Kouhei Sekiguchi, Aditya Arie Nugraha, Kazuyoshi Yoshii, Kotaro Funakoshi |
RO-MAN | 5 |
| 2018 | Vibrational Artificial Subtle Expressions: Conveying System's Confidence Level to Users by Means of Smartphone VibrationabstractArtificial subtle expressions (ASEs) are machine-like expressions used to convey a system's confidence level to users intuitively. So far, auditory ASEs using beep sounds, visual ASEs using LEDs, and motion ASEs using robot movements have been implemented and shown to be effective. In this paper, we propose a novel type of ASE that uses vibration (vibrational ASEs). We implemented the vibrational ASEs on a smartphone and conducted experiments to confirm whether they can convey a system's confidence level to users in the same way as the other types of ASEs. The results clearly showed that vibrational ASEs were able to accurately and intuitively convey the designed confidence level to participants, demonstrating that ASEs can be applied in a variety of applications in real environments. Takanori Komatsu, Kazuki Kobayashi, Seiji Yamada, Kotaro Funakoshi, Mikio Nakano |
CHI | 4 |
| 2018 | Interaction Modeling Based on Segmenting Two Persons Motions Using Coupled GP-HSMMabstractHumans interact with one another daily and learn from their experiences via observation and interaction. To create robots that can coexist with humans, it is important that they learn how to appropriately interact in the human community. In this paper, we propose a novel model, the coupled Gaussian process hidden semi-Markov model (GP-HSMM), which enables robots to learn rules of interaction between two persons by observing them in an unsupervised manner. The continuous motions of the persons are segmented into discrete actions based on GP-HSMM, and the relationships between the actions are extracted. Moreover, all corresponding actions are not simultaneously conducted by two persons during actual interaction. Thus, the coupled GP-HSMM accounts for these lags. We conduct experiments using the motion data of interaction games. Experimental results showed that the coupled GP-HSMM can estimate actions, lags between them and their relationships. Satoru Oshikawa, Tomoaki Nakamura, Takayuki Nagai, Masahide Kaneko, Kotaro Funakoshi, Naoto Iwahashi, Mikio Nakano |
RO-MAN | 5 |
| 2017 | Response Times when Interpreting Artificial Subtle Expressions are Shorter than with Human-like Speech SoundsabstractArtificial subtle expressions (ASEs) are machine-like expressions used to convey a system's confidence level to users intuitively. In this paper, we focus on the cognitive loads of users in interpreting ASEs in this study. Specifically, we assume that a shorter response time indicates less cognitive load, and we hypothesize that users will show a shorter response time when interpreting ASEs compared with speech sounds. We succeeded in verifying our hypothesis in a web-based investigation done to comprehend participants' cognitive loads by measuring their response times in interpreting ASEs and speeches. Takanori Komatsu, Kazuki Kobayashi, Seiji Yamada, Kotaro Funakoshi, Mikio Nakano |
CHI | 4 |
| 2017 | An Attention-based Regression Model for Grounding Textual Phrases in ImagesabstractGrounding, or localizing, a textual phrase in an image is a challenging problem that is integral to visual language understanding. Previous approaches to this task typically make use of candidate region proposals, where end performance depends on that of the region proposal method and additional computational costs are incurred. In this paper, we treat grounding as a regression problem and propose a method to directly identify the region referred to by a textual phrase, eliminating the need for external candidate region prediction. Our approach uses deep neural networks to combine image and text representations and refines the target region with attention models over both image subregions and words in the textual phrase. Despite the challenging nature of this task and sparsity of available data, in evaluation on the ReferIt dataset, our proposed method achieves a new state-of-the-art in performance of 37.26% accuracy, surpassing the previously reported best by over 5 percentage points. We find that combining image and text attention models and an image attention area-sensitive loss function contribute to substantial improvements. Ko Endo, Masaki Aono, Eric Nichols, Kotaro Funakoshi |
IJCAI | 4 |
| 2017 | Boredom Recognition Based on Users' Spontaneous Behaviors in Multiparty Human-Robot Interactions
Yasuhiro Shibasaki, Kotaro Funakoshi, Koichi Shinoda |
MMM (1) | 2 |
| 2016 | Nonparametric Bayesian Models for Spoken Language UnderstandingabstractIn this paper, we propose a new generative approach for semantic slot filling task in spoken language understanding using a nonparametric Bayesian formalism.Slot filling is typically formulated as a sequential labeling problem, which does not directly deal with the posterior distribution of possible slot values.We present a nonparametric Bayesian model involving the generation of arbitrary natural language phrases, which allows an explicit calculation of the distribution over an infinite set of slot values.We demonstrate that this approach significantly improves slot estimation accuracy compared to the existing sequential labeling algorithm. Kei Wakabayashi, Johane Takeuchi, Kotaro Funakoshi, Mikio Nakano |
EMNLP | 3 |
| 2016 | The dialogue breakdown detection challenge: Task description, datasets, and evaluation metrics
Ryuichiro Higashinaka, Kotaro Funakoshi, Yuka Kobayashi, Michimasa Inaba |
LREC | 2 |
| 2015 | Investigating Ways of Interpretations of Artificial Subtle Expressions Among Different Languages: A Case of Comparison Among Japanese, German, Portuguese and Mandarin Chinese
Takanori Komatsu, Rui Prada, Kazuki Kobayashi, Seiji Yamada, Kotaro Funakoshi, Mikio Nakano |
CogSci | 5 |
| 2015 | Fatal or not? Finding errors that lead to dialogue breakdowns in chat-oriented dialogue systemsabstractRyuichiro Higashinaka, Masahiro Mizukami, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Ryuichiro Higashinaka, Masahiro Mizukami, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi |
EMNLP | 3 |
| 2015 | Towards Taxonomy of Errors in Chat-oriented Dialogue SystemsabstractRyuichiro Higashinaka, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi, Masahiro Mizukami. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015. Ryuichiro Higashinaka, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi, Masahiro Mizukami |
SIGDIAL Conference | 2 |
| 2014 | Probabilistic multiparty dialogue management for a game master robotabstractWe present our ongoing research on multiparty dialogue management for a game master robot which engages multiple human participants to play a quiz game. The robot invites passing people to join the game, instructs participants on the rules of the game, and leads them in the game. The robot has to manage people leaving and coming at arbitrary times. Our approach maintains a dialogue manager for each participant, and a module takes a final action with each decision cycle; responsible to decide "what/whom/when to say". We have implemented the dialogue manager with a probabilistic rules approach [4] and made preliminary evaluations with our multiparty human-robot game dialogue data that was collected in a WoZ fashion. Casey Kennington, Kotaro Funakoshi, Yuki Takahashi, Mikio Nakano |
HRI | 2 |
| 2014 | Integration of various concepts and grounding of word meanings using multi-layered multimodal LDA for sentence generationabstractIn the field of intelligent robotics, object handling by robots can be achieved by capturing not only the object concept through object categorization, but also other concepts (e.g., the movement while using the object), as well as the relationship between concepts. Moreover, capturing the concepts of places and people is also necessary to enable the robot to gain real-world understanding. In this study, we propose multi-layered multimodal latent Dirichlet allocation (mMLDA) to realize the formation of various concepts, and the integration of those concepts, by robots. Because concept formation and integration can be conducted by mMLDA, the formation of each concept affects others, resulting in a more appropriate formation. Another issue to be addressed in this paper is the language acquisition by the robots. We propose a method to infer which words are originally connected to a concept using mutual information between words and concepts. Moreover, the order of concepts in teaching utterances can be learned using a simple Markov model, which corresponds to grammar. This grammar can be used to generate sentences that represent the observed information. We report the results of experiments to evaluate the effectiveness of the proposed method. Muhammad Attamimi, Muhammad Fadlil, Kasumi Abe, Tomoaki Nakamura, Kotaro Funakoshi, Takayuki Nagai |
IROS | 5 |
| 2014 | Mutual learning of an object concept and language model based on MLDA and NPYLMabstractHumans develop their concept of an object by classifying it into a category, and acquire language by interacting with others at the same time. Thus, the meaning of a word can be learnt by connecting the recognized word and concept. We consider such an ability to be important in allowing robots to flexibly develop their knowledge of language and concepts. Accordingly, we propose a method that enables robots to acquire such knowledge. The object concept is formed by classifying multimodal information acquired from objects, and the language model is acquired from human speech describing object features. We propose a stochastic model of language and concepts, and knowledge is learnt by estimating the model parameters. The important point is that language and concepts are interdependent. There is a high probability that the same words will be uttered to objects in the same category. Similarly, objects to which the same words are uttered are highly likely to have the same features. Using this relation, the accuracy of both speech recognition and object classification can be improved by the proposed method. However, it is difficult to directly estimate the parameters of the proposed model, because there are many parameters that are required. Therefore, we approximate the proposed model, and estimate its parameters using a nested Pitman-Yor language model and multimodal latent Dirichlet allocation to acquire the language and concept, respectively. Tomoaki Nakamura, Takayuki Nagai, Kotaro Funakoshi, Shogo Nagasaka, Tadahiro Taniguchi, Naoto Iwahashi |
IROS | 3 |
| 2013 | A Robotic Agent in a Virtual Environment that Performs Situated Incremental Understanding of Navigational Utterances
Takashi Yamauchi, Mikio Nakano, Kotaro Funakoshi |
SIGDIAL Conference | 3 |
| 2013 | Correcting phoneme recognition errors in learning word pronunciation through speech interaction
Xiang Zuo, Taisuke Sumii, Naoto Iwahashi, Mikio Nakano, Kotaro Funakoshi, Natsuki Oka |
Speech Commun. | 5 |
| 2012 | How Can We Live with Overconfident or Unconfident Systems?: A Comparison of Artificial Subtle Expressions with Human-like Expression
Takanori Komatsu, Kazuki Kobayashi, Seiji Yamada, Kotaro Funakoshi, Mikio Nakano |
CogSci | 4 |
| 2012 | Impressions made by blinking light used to create artificial subtle expressions and by robot appearance in human-robot speech interactionabstractThe impressions made by a blinking light used to create artificial subtle expressions (ASEs) and by a robot's appearance on users were investigated. The blinking light, which shows the user that the robot is performing speech recognition and thereby prevents utterance collisions, was separated from the robot by embedding it in a pedestal unit. In an evaluation experiment, participants performed five tasks with a spoken dialogue system coupled to a robot placed on the pedestal. The participants' impressions of the dialogue interactions and of the robot were obtained under four conditions (w/ light blinking or w/o blinking; humanoid or cuboid robot). The cuboid robot created a stronger impression of comfort and excitement for the interactions while the blinking light did not create a strong impression of anything. The robot's appearance and the blinking did not create a strong impression of anything for the robot. This suggests that the blinking light in the pedestal unit is a factor that is independent of robot appearance, meaning that the pedestal unit can be applied to robots with various appearances. Kazuki Kobayashi, Kotaro Funakoshi, Seiji Yamada, Mikio Nakano, Takanori Komatsu, Yasunori Saito |
RO-MAN | 2 |
| 2012 | A Unified Probabilistic Approach to Referring Expressions
Kotaro Funakoshi, Mikio Nakano, Takenobu Tokunaga, Ryu Iida |
SIGDIAL Conference | 1 |
| 2011 | Interpretations of Artificial Subtle Expressions (ASEs) in Terms of Different Types of Artifact: A Comparison of an on-screen Artifact with A Robot
Takanori Komatsu, Seiji Yamada, Kazuki Kobayashi, Kotaro Funakoshi, Mikio Nakano |
ACII (2) | 4 |
| 2011 | The chanty bear: a new application for hri researchabstractThis paper presents yet another English-teaching robot, while putting emphasis on the merits which are offered by second language education to human robot interaction (HRI) research. The chanty bear, our prototype robot based on a rhythmic teaching method of English called Jazz Chants is introduced. Kotaro Funakoshi, Tomoya Mizumoto, Ryo Nagata, Mikio Nakano |
HRI | 1 |
| 2011 | Learning Place-Names from Spoken Utterances and Localization Results by Mobile RobotabstractThis paper proposes a method for the unsupervised learning of place-names from pairs of a spoken utterance and a localization result, which represents a current location of a mobile robot, without any priori linguistic knowledge other than a phoneme acoustic model.In previous work, we have proposed a lexical learning method based on statistical model selection.This method can learn the words that represent a single object, such as proper nouns, but cannot learn the words that represent classes of objects, such as general nouns.This paper describes improvements of the method for learning both a phoneme sequence of each word and a distribution of objects that the word represents. Ryo Taguchi, Yuji Yamada, Koosuke Hattori, Taizo Umezaki, Masahiro Hoguro, Naoto Iwahashi, Kotaro Funakoshi, Mikio Nakano |
INTERSPEECH | 7 |
| 2011 | Autonomous acquisition of multimodal information for online object concept formation by robotsabstractThis paper proposes a robot that acquires multi-modal information, i.e. auditory, visual, and haptic information, fully autonomous way using its embodiment. We also propose an online algorithm of multimodal categorization based on the acquired multimodal information and words, which are partially given by human users. The proposed framework makes it possible for the robot to learn object concepts naturally in everyday operation in conjunction with a small amount of linguistic information from human users. In order to obtain multimodal information, the robot detects an object on a fla surface. Then the robot grasps and shakes it for gaining haptic and auditory information. For obtaining visual information, the robot uses a hand held small observation table, so that the robot can control the viewpoints for observing the object. As for the multimodal concept formation, the multimodal LDA using Gibbs sampling is extended to the online version in this paper. The proposed algorithms are implemented on a real robot and tested using real everyday objects in order to show validity of the proposed system. Takaya Araki, Tomoaki Nakamura, Takayuki Nagai, Kotaro Funakoshi, Mikio Nakano, Naoto Iwahashi |
IROS | 4 |
| 2011 | Blinking light patterns as artificial subtle expressions in human-robot speech interactionabstractUsers' impressions of blinking light expressions used as artificial subtle expressions have been investigated. In a preliminary experiment, thirteen blinking patterns were used for investigating participants' impressions of their agreeableness. The highest and lowest valued blinking patterns were identified and used for a speech interaction experiment. In this experiment, 52 participants tried to reserve hotel rooms with a spoken dialogue system coupled with an interface robot using a blinking light expression. A sine wave, a random wave, a rectangular wave, and a no-blinking condition were used as artificial subtle expressions to express a robot's internal state of “processing” or “recognizing”. The results of a questionnaire showed the conditions did not significantly differ in terms of agreeableness, but the sine wave and the rectangular wave were evaluated as “more useful” than the no-blinking condition. Results of factor analyses suggested that the rectangular wave provides a comfortable impression of the dialogue. Kazuki Kobayashi, Kotaro Funakoshi, Seiji Yamada, Mikio Nakano, Takanori Komatsu, Yasunori Saito |
RO-MAN | 2 |
| 2011 | A Two-Stage Domain Selection Framework for Extensible Multi-Domain Spoken Dialogue Systems
Mikio Nakano, Kazunori Komatani, Kyoko Matsuyama, Kotaro Funakoshi, Hiroshi G. Okuno |
SIGDIAL Conference | 5 |
| 2011 | A multi-expert model for dialogue and behavior control of conversational robots and agents
Mikio Nakano, Yuji Hasegawa, Kotaro Funakoshi, Johane Takeuchi, Toyotaka Torii, Kazuhiro Nakadai, Naoyuki Kanda, Kazunori Komatani, Hiroshi G. Okuno, Hiroshi Tsujino |
Knowl. Based Syst. | 3 |
| 2010 | Artificial subtle expressions: intuitive notification methodology of artifactsabstractWe describe artificial subtle expressions (ASEs) as intuitive notification methodology for artifacts' internal states for users. We prepared two types of audio ASEs; one was a flat artificial sound (flat ASE), and the other was a sound that decreased in pitch (decreasing ASE). These two ASEs were played after a robot made a suggestion to the users. Specifically, we expected that the decreasing ASE would inform users of the robot's lower level of confidence about the suggestions. We then conducted a simple experiment to observe whether the participants accepted or rejected the robot's suggestion in terms of the ASEs. The results showed that they accepted the robot's suggestion when the flat ASE was used, whereas they rejected it when the decreasing ASE was used. Therefore, we found that the ASEs succeeded in conveying the robot's internal state to the users accurately and intuitively. Takanori Komatsu, Seiji Yamada, Kazuki Kobayashi, Kotaro Funakoshi, Mikio Nakano |
CHI | 4 |
| 2010 | Similarities and differences in users' interaction with a humanoid and a pet robotabstractIn this paper, we compare user behavior towards the humanoid robot ASIMO and the dog-shaped robot AIBO in a simple task, in which the users has to teach commands and feedback to the robot. Anja Austermann, Seiji Yamada, Kotaro Funakoshi, Mikio Nakano |
HRI | 3 |
| 2010 | Robot-directed speech detection using Multimodal Semantic Confidence based on speech, image, and motionabstractIn this paper, we propose a novel method to detect robot-directed (RD) speech that adopts the Multimodal Semantic Confidence (MSC) measure. The MSC measure is used to decide whether the speech can be interpreted as a feasible action under the current physical situation in an object manipulation task. This measure is calculated by integrating speech, image, and motion confidence measures with weightings that are optimized by logistic regression. Experimental results show that, compared with a baseline method that uses speech confidence only, MSC achieved an absolute increase of 5% for clean speech and 12% for noisy speech in terms of average maximum F-measure. Xiang Zuo, Naoto Iwahashi, Ryo Taguchi, Shigeki Matsuda, Komei Sugiura, Kotaro Funakoshi, Mikio Nakano, Natsuki Oka |
ICASSP | 6 |
| 2010 | Learning naturally spoken commands for a robot
Anja Austermann, Seiji Yamada, Kotaro Funakoshi, Mikio Nakano |
INTERSPEECH | 3 |
| 2010 | Real-time 3D visual sensor for robust object recognitionabstractThis paper presents a novel 3D measurement system, which yields both depth and color information in real time, by calibrating a time-of-flight and two CCD cameras. The problem of occlusions is solved by the proposed fast occluded-pixel detection algorithm. Since the system uses two CCD cameras, missing color information of occluded pixels is covered by one another. We also propose a robust object recognition using the 3D visual sensor. Multiple cues, such as color, texture and 3D (depth) information, are integrated in order to recognize various types of objects under varying lighting conditions. We have implemented the system on our autonomous robot and made the robot do recognition tasks (object learning, detection, and recognition) in various environments. The results revealed that the proposed recognition system provides far better performance than the previous system that is based only on color and texture information. Muhammad Attamimi, Akira Mizutani, Tomoaki Nakamura, Takayuki Nagai, Kotaro Funakoshi, Mikio Nakano |
IROS | 5 |
| 2010 | Does the appearance of a robot affect users' ways of giving Commands and feedback?abstractOur study compares users' interaction with a humanoid robot and a dog-shaped pet-robot. We conducted a user study in which the participants had to teach object names as well as simple commands to either the humanoid or the pet-robot and give feedback to the robot for correct and incorrect performance. While we found, that the way of uttering commands rather depends on personal preference than on the robots' appearance, the way of giving positive and negative feedback differed significantly between both robots: We found that for the pet-robot users gave reward in a similar way as giving reward to a real dog by touching it and commenting on its performance by uttering feedback like “well done” or “that was right”. For the humanoid, users typically did not use touch as a reward and rather used personal expressions like “thank you” to praise the robot. Our findings suggest that users actually rely to some degree on the appearance of a robot as a cue for deciding how to interact with it. Anja Austermann, Seiji Yamada, Kotaro Funakoshi, Mikio Nakano |
RO-MAN | 3 |
| 2010 | Detecting robot-directed speech by situated understanding in object manipulation tasksabstractIn this paper, we propose a novel method for a robot to detect robot-directed speech, that is, to distinguish speech that users speak to a robot from speech that users speak to other people or to themselves. The originality of this work is the introduction of a multimodal semantic confidence (MSC) measure, which is used for domain classification of input speech based on the decision on whether the speech can be interpreted as a feasible action under the current physical situation in an object manipulation task. This measure is calculated by integrating speech, object, and motion confidence with weightings that are optimized by logistic regression. Then we integrate this measure with gaze tracking and conduct experiments under conditions of natural human-robot interaction. Experimental results show that the proposed method achieves a high performance of 94% and 96% in average recall and precision rates, respectively, for robot-directed speech detection. Xiang Zuo, Naoto Iwahashi, Ryo Taguchi, Kotaro Funakoshi, Mikio Nakano, Shigeki Matsuda, Komei Sugiura, Natsuki Oka |
RO-MAN | 4 |
| 2010 | Non-humanlike Spoken Dialogue: A Design Perspective
Kotaro Funakoshi, Mikio Nakano, Kazuki Kobayashi, Takanori Komatsu, Seiji Yamada |
SIGDIAL Conference | 1 |
| 2010 | Correction of phoneme recognition errors in word learning through speech interactionabstractThis paper describes a novel method that enables users to teach systems the phoneme sequences of new words through speech interaction. Using the method, users can correct mis-recognized phoneme sequences incrementally by making corrective utterances. Each corrective utterance may include the whole or a segment of the word. During the interaction, if the correction using the utterance results in a better phoneme sequence than the previous one, a user can stop the interaction or make a corrective utterance again. Otherwise the user can reject the utterance. The originalities of this method are 1) interactive correction by speech, 2) the use of spoken word segments for locating mis-recognized phonemes and, 3) the use of generalized posterior probability (GPP) as a measure of correcting mis-recognized phonemes. The experimental results show that the proposed method achieved 96.8% in phoneme accuracy and 79.1% in word accuracy, with less than seven corrective utterances. Xiang Zuo, Taisuke Sumii, Naoto Iwahashi, Kotaro Funakoshi, Mikio Nakano, Natsuki Oka |
SLT | 4 |
| 2009 | Improving speech understanding accuracy with limited training data using multiple language models and multiple understanding modelsabstractWe aim to improve a speech understanding module with a small amount of training data. A speech understanding module uses a language model (LM) and a language understanding model (LUM). A lot of training data are needed to improve the models. Such data collection is, however, difficult in an actual process of development. We therefore design and develop a new framework that uses multiple LMs and LUMs to improve speech understanding accuracy under various amounts of training data. Even if the amount of available training data is small, each LM and each LUM can deal well with different types of utterances and more utterances are understood by using multiple LM and LUM. As one implementation of the framework, we develop a method for selecting the most appropriate speech understanding result from several candidates. The selection is based on probabilities of correctness calculated by logistic regressions. We evaluate our framework with various amounts of training data. Index Terms: speech understanding, multiple language models and language understanding models, limited training data Masaki Katsumaru, Mikio Nakano, Kazunori Komatani, Kotaro Funakoshi, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 4 |
| 2009 | Learning lexicons from spoken utterances based on statistical model selection
Ryo Taguchi, Naoto Iwahashi, Takashi Nose, Kotaro Funakoshi, Mikio Nakano |
INTERSPEECH | 4 |
| 2008 | Smoothing human-robot speech interactions by using a blinking-light as subtle expressionabstractSpeech overlaps, undesired collisions of utterances between systems and users, harm smooth communication and degrade the usability of systems. We propose a method to enable smooth speech interactions between a user and a robot, which enables subtle expressions by the robot in the form of a blinking LED attached to its chest. In concrete terms, we show that, by blinking an LED from the end of the user's speech until the robot's speech, the number of undesirable repetitions, which are responsible for speech overlaps, decreases, while that of desirable repetitions increases. In experiments, participants played a last-and-first game with the robot. The experimental results suggest that the blinking-light can prevent speech overlaps between a user and a robot, speed up dialogues, and improve user's impressions. Kotaro Funakoshi, Kazuki Kobayashi, Mikio Nakano, Seiji Yamada, Yasuhiko Kitamura, Hiroshi Tsujino |
ICMI | 1 |
| 2008 | Rapid Prototyping of Robust Language Understanding Modules for Spoken Dialogue Systems
Yuichiro Fukubayashi, Kazunori Komatani, Mikio Nakano, Kotaro Funakoshi, Hiroshi Tsujino, Tetsuya Ogata, Hiroshi G. Okuno |
IJCNLP | 4 |
| 2008 | Smoothing human-robot speech interaction with blinking-light expressionsabstractWe propose a method to enable smooth speech interactions between a user and a robot. Our method is based on subtle expression whereby a robot blinks a small LED attached to its chest. We performed experiments in which participants played a last-and-first games and counted the number of repetitions made by the participants and analyzed their impression of the game and the robot. The experimental results suggested that the blinking-light could prevent utterance collisions between a user and a robot and could create familiar and attentive impressions about the game on users. Kazuki Kobayashi, Kotaro Funakoshi, Seiji Yamada, Mikio Nakano, Yasuhiko Kitamura, Hiroshi Tsujino |
RO-MAN | 2 |
| 2007 | Robust acquisition and recognition of spoken location names by domestic robotsabstractThis paper presents a method that enables a conversational domestic robot to learn location names through speech interaction. Each acquired name is associated with a point on the map coordinate system of the robot. Both for acquisition and recognition of location names, a bag- of-words-based categorization technique is used. Namely, the robot acquires a location name as a frequency pattern of words, and recognizes a spoken location name by computing similarity between the patterns. This makes the robot robust not only against speech recognition errors but also against out- of-vocabulary names. We designed a dialogue and behavior management subsystem that learns location names by using our proposed method and navigates to indicated locations, and implemented the subsystem on an omnidirectional cart robot. The result of a preliminary evaluation of the implemented robot with human subjects suggested this approach is promising. Kotaro Funakoshi, Mikio Nakano, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Noriyuki Kimura, Naoto Iwahashi |
IROS | 1 |
| 2007 | A markup language for describing interactive humanoid robot presentationsabstractThis paper presents a multi-modal presentation markup language for humanoid robots, MPML-HR ver. 3.0, which is able to describe presentation contents including speech-based interactions with audiences. Previous versions of MPML-HR do not feature any interaction functionality which dynamically changes the presentation according to the utterances by audiences, although such interaction makes the presentation more effective and understandable. Since MPML-HR ver. 3.0 inherits simple descriptions of previous versions of MPML-HR, the content designer can describe interactive presentations without configuring conventional complicated multi-modal interactive systems. Yoshitaka Nishimura, Shinichiro Minotsu, Hiroshi Dohi, Mitsuru Ishizuka, Mikio Nakano, Kotaro Funakoshi, Johane Takeuchi, Yuji Hasegawa, Hiroshi Tsujino |
IUI | 6 |
| 2006 | Identifying Repair Targets in Action Control Dialogue
Kotaro Funakoshi, Takenobu Tokunaga |
EACL | 1 |
| 2006 | Group-Based Generation of Referring Expressions
Kotaro Funakoshi, Satoru Watanabe, Takenobu Tokunaga |
INLG | 1 |
| 2005 | Understanding Referring Expressions Involving Perceptual GroupingabstractThis paper deals with understanding referring expressions involving perceptual grouping. The ability to use referring expressions is important for conversational agents aimed at real-world interaction. We conducted a psychological experiment to collect referring expressions involving perceptual grouping. A set of methods to identify referents based on the collected data are presented. We were able to identify 78.8% of the referents in the collected expressions Kotaro Funakoshi, Satoru Watanabe, Takenobu Tokunaga, Naoko Kuriyama |
CW | 1 |
| 2004 | Generation of Relative Referring Expressions based on Perceptual Grouping
Kotaro Funakoshi, Satoru Watanabe, Naoko Kuriyama, Takenobu Tokunaga |
COLING | 1 |
| 2004 | Generating Referring Expressions Using Perceptual Groups
Kotaro Funakoshi, Satoru Watanabe, Naoko Kuriyama, Takenobu Tokunaga |
INLG | 1 |
| 2004 | K2: Animated Agents that Understand Speech Commands and Perform Actions
Takenobu Tokugana, Kotaro Funakoshi, Hozumi Tanaka |
PRICAI | 2 |
| 2002 | Processing Japanese Self-correction in Speech Dialog Systems
Kotaro Funakoshi, Takenobu Tokunaga, Hozumi Tanaka |
COLING | 1 |