Seiya Kawano

dblp:249/4687 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
17since 2021 · last 2025
0000-0002-5830-8169ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression
abstract
To improve user engagement during conversations with dialogue systems, we must improve individual dialogue responses and dialogue impressions such as consistency, personality, and empathy throughout the entire dialogue. While such dialogue systems have been developing rapidly with the help of large language models (LLMs), reinforcement learning from AI feedback (RLAIF) has attracted attention to align LLM-based dialogue models for such dialogue impressions. In RLAIF, a reward model based on another LLM is used to create a training signal for an LLM-based dialogue model using zero-shot/few-shot prompting techniques. However, evaluating an entire dialogue only by prompting LLMs is challenging. In this study, the supervised fine-tuning (SFT) of LLMs prepared reward models corresponding to 12 metrics related to the impression of the entire dialogue for evaluating dialogue responses. We tuned our dialogue models using the reward model signals as feedback to improve the impression of the system. The results of automatic and human evaluations showed that tuning the dialogue model using our reward model corresponding to dialogue impression improved the evaluation of individual metrics and the naturalness of the dialogue response.
Kai Yoshida, Masahiro Mizukami, Seiya Kawano, Canasai Kruengkrai, Hiroaki Sugiyama, Koichiro Yoshino
ICASSP3
2025 Rapport-Building Dialogue Strategies for Deeper Connection: Integrating Proactive Behavior, Personalization, and Aizuchi Backchannels
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
INTERSPEECH3
2025 Co-Speech Motion for Virtual Agents in Dialogue Using LLM-Driven Primitive Action Selection
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
INTERSPEECH3
2025 Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
Kiyotada Mori, Seiya Kawano, Angel F. Garcia Contreras, Koichiro Yoshino
INTERSPEECH2
2025 What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
Kiyotada Mori, Seiya Kawano, Carlos Toshinori Ishi, Angel F. Garcia Contreras, Koichiro Yoshino
INTERSPEECH2
2025 Using Language Models to Generate and Forget the Narrative Memories of an Assistive Robot
Angel F. Garcia Contreras, Wen-Yu Chang, Seiya Kawano, Yun-Nung Chen, Koichiro Yoshino
MMM (5)3
2025 RoboDJ: Live Commentary Robots System Driven by Physical- and Cyber-World Observations
Yasutomo Kawanishi, Yutaka Nakamura, Taiken Shintani, Carlos Toshinori Ishi, Seiya Kawano, Koichiro Yoshino, Takashi Minato, Michihiko Minoh
MMM (5)5
2025 What Should Autonomous Robots Verbalize and What Should They Not?
Daichi Yoshihara, Akishige Yuguchi, Seiya Kawano, Takamasa Iio, Koichiro Yoshino
MMM (5)3
2025 LLM-Driven Approach for Motion Control in Human-Robot Dialogue for Elevating Engagement
abstract
Non-verbal behaviors, such as body movements, play a crucial role in enhancing a robot’s speech to elevate engagement in human-robot dialogue. Many existing approach based on rules offered natural and engaging motions aligned with the robot’s utterances but required significant resources to maintain. Recent methods leveraging large language models (LLMs) offer a promising alternative to reduce these costs. However, there is a trade-off between flexibility and safety when determining whether the language model should generate motions based on joint angle parameters or action primitives. In this study, we evaluated two LLM-based motion control models: one for motion generation based on joint angle parameters (LLM-GJA) and the other for motion generation based on primitive actions (LLM-GPA). Our human evaluations indicated that directly generating joint angles outperformed generating action primitives in naturalness, timing consistency, and overall engagement, even achieving performance comparable to rule-based systems. This work highlights the potential of LLMs in generating expressive and contextually appropriate robot motions at the joint angle level.
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
RO-MAN3
2025 Multi-step or Direct: A Proactive Home-Assistant System Based on Commonsense Reasoning
abstract
There is a growing expectation for the realization of proactive home-assistant robots that can assist users in their daily lives. It is essential to develop a framework that closely observes the user’s surrounding context, selectively extracts relevant information, and infers the user’s needs to proactively propose appropriate assistance. In this study, we first extend the Do-I-Demand dataset to define expected proactive assistance actions in domestic situations, where users make ambiguous utterances. These behaviors were defined based on common patterns of support that a majority of users would expect from a robot. We subsequently constructed a framework that infers users’ expected assistance actions from ambiguous utterances through commonsense reasoning. We explored two approaches: (1) multi-step reasoning using COMET as a commonsense reasoning engine, and (2) direct reasoning using large language models. Our experimental results suggest that both the multi-step and direct reasoning methods can successfully derive necessary assistance actions even when dealing with ambiguous user utterances.
Konosuke Yamasaki, Shohei Tanaka, Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino
SIGDIAL4
2024 ClaimBrush: A Novel Framework for Automated Patent Claim Refinement Based on Large Language Models
abstract
Automatic refinement of patent claims in patent applications is crucial from the perspective of intellectual property strategy. In this paper, we propose "ClaimBrush," a novel framework for automated patent claim refinement that includes a dataset and a rewriting model. We constructed a dataset for training and evaluating patent claim rewriting models by collecting a large number of actual patent claim rewriting cases from the patent examination process. Using the constructed dataset, we built an automatic patent claim rewriting model by fine-tuning a large language model. Furthermore, we enhanced the performance of the automatic patent claim rewriting model by applying preference optimization based on a prediction model of patent examiners’ Office Actions. The experimental results showed that our proposed rewriting model outperformed heuristic baselines and zero-shot learning in state-of-the-art large language models. Moreover, preference optimization based on patent examiners’ preferences boosted the performance of patent claim refinement.
Seiya Kawano, Hirofumi Nonaka, Koichiro Yoshino
IEEE Big Data1
2024 A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
abstract
Situated conversations, which refer to visual information as visual question answering (VQA), often contain ambiguities caused by reliance on directive information. This problem is exacerbated because some languages, such as Japanese, often omit subjective or objective terms. Such ambiguities in questions are often clarified by the contexts in conversational situations, such as joint attention with a user or user gaze information. In this study, we propose the Gaze-grounded VQA dataset (GazeVQA) that clarifies ambiguous questions using gaze information by focusing on a clarification process complemented by gaze information. We also propose a method that utilizes gaze target estimation results to improve the accuracy of GazeVQA tasks. Our experimental results showed that the proposed method improved the performance in some cases of a VQA system on GazeVQA and identified some typical problems of GazeVQA tasks that need to be improved.
Shun Inadumi, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino
LREC/COLING2
2024 J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution
abstract
Understanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution, a system must ground the verbal information that appears in user interactions to the visual information observed in egocentric views. To this end, we propose a multimodal reference resolution task and construct a Japanese Conversation dataset for Real-world Reference Resolution (J-CRe3). Our dataset contains egocentric video and dialogue audio of real-world conversations between two people acting as a master and an assistant robot at home. The dataset is annotated with crossmodal tags between phrases in the utterances and the object bounding boxes in the video frames. These tags include indirect reference relations, such as predicate-argument structures and bridging references as well as direct reference relations. We also constructed an experimental model and clarified the challenges in multimodal reference resolution tasks.
Nobuhiro Ueda, Hideko Habe, Akishige Yuguchi, Seiya Kawano, Yasutomo Kawanishi, Sadao Kurohashi, Koichiro Yoshino
LREC/COLING4
2024 Rapport-Driven Virtual Agent: Rapport Building Dialogue Strategy for Improving User Experience at First Meeting
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
INTERSPEECH3
2023 Operative Action Captioning for Estimating System Actions
abstract
Human-assistive systems, such as robots, need to correctly understand the surrounding situation based on obser-vations and output the required support actions for humans. Language is one of the important channels to communicate with humans, and robots are required to have the ability to express their understanding and action-planning results. In this study, we propose a new task of operative action captioning that estimates and verbalizes the actions to be taken by the system in a human-assisting domain. We constructed a system that outputs a verbal description of a possible operative action that changes the current state to the given target state. We collected a dataset consisting of two images as observations, which express the current state and the state changed by actions and a caption that describes the actions that change the current state to the target state, by crowdsourcing in daily life situations. Then we constructed a system that estimates an operative action by a caption. Since the operative action's caption is expected to contain some state-changing actions, we use scene graph prediction as an auxiliary task because the events written in the scene graphs correspond to the state changes. Experimental results showed that our system successfully described the operative actions that should be conducted between the current and target states. The auxiliary tasks that predict the scene graphs improved the quality of the estimation results.
Taiki Nakamura, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino
ICRA2
2022 Butsukusa: A Conversational Mobile Robot Describing Its Own Observations and Internal States
abstract
This paper presents an autonomous conversational mobile robot Butsukusa that can describe its own observations and internal states during patrolling tasks. The proposed robot can observe the surrounding environment using the recognition module for objects, humans, environment, localization, and speech and then move autonomously around an indoor living space. Interaction skills via language are required for the robot to perform in such human-centered spaces. To investigate a better communication protocol with users, we evaluate various language generation patterns based on different observations and interaction patterns. The evaluation results indicate that the importance of describing the robot's observation results and internal states, as well as the necessity of an appropriate description, depends on the situation.
Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino, Carlos Toshinori Ishi, Yasutomo Kawanishi, Yutaka Nakamura, Takashi Minato, Yasuki Saito, Michihiko Minoh
HRI2
2022 Multimodal Persuasive Dialogue Corpus using Teleoperated Android
Seiya Kawano, Muteki Arioka, Akishige Yuguchi, Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara, Satoshi Nakamura 0001, Koichiro Yoshino
INTERSPEECH1
2019 Neural Conversation Model Controllable by Given Dialogue Act Based on Adversarial Learning and Label-aware Objective
abstract
Building a controllable neural conversation model (NCM) is an important task.In this paper, we focus on controlling the responses of NCMs by using dialogue act labels of responses as conditions.We introduce an adversarial learning framework for the task of generating conditional responses with a new objective to a discriminator, which explicitly distinguishes sentences by using labels.This change strongly encourages the generation of label-conditioned sentences.We compared the proposed method with some existing methods for generating conditional responses.The experimental results show that our proposed method has higher controllability for dialogue acts even though it has higher or comparable naturalness to existing methods.
Seiya Kawano, Koichiro Yoshino, Satoshi Nakamura 0001
INLG1