Kiyosu Maeda

dblp:257/8023 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-3270-1974ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
abstract
A quintessential feature of human intelligence is the ability to create ad hoc conventions over time to achieve shared goals efficiently. We investigate how communication strategies evolve through repeated collaboration as people coordinate on shared procedural abstractions. To this end, we conducted an online unimodal study (n = 98) using natural language to probe abstraction hierarchies. In a follow-up lab study (n = 40), we examined how multimodal communication (speech and gestures) changed during physical collaboration. Pairs used augmented reality to isolate their partner’s hand and voice; one participant viewed a 3D virtual tower and sent instructions to the other, who built the physical tower. Participants became faster and more accurate by establishing linguistic and gestural abstractions and using cross-modal redundancy to emphasize key changes from previous interactions. Based on these findings, we extend probabilistic models of convention formation to multimodal settings, capturing shifts in modality preferences. Our findings and model provide building blocks for designing convention-aware intelligent agents situated in the physical world.
Kiyosu Maeda, William P. McCarthy, Ching-Yi Tsai, Jeffrey Mu, Robert D. Hawkins, Judith E. Fan, Parastoo Abtahi
CHI1
2025 Using Gesture and Language to Establish Multimodal Conventions in Collaborative Physical Tasks
Kiyosu Maeda, Ching-Yi Tsai, Judith E. Fan, Parastoo Abtahi
CogSci1
2025 ConverSearch: Supporting Experts in Human Behavior Analysis of Conversational Videos with a Multimodal Scene Search Tool
abstract
Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a lack of comprehensive, user-friendly tools that streamline the processing of diverse multimodal queries impedes efficiency and objectivity. To address this gap, we developed ConverSearch , a visual-programming-based tool based on insights for effective interface and implementation design derived from a formative study with experts. The tool allows experts to integrate various machine learning algorithms to capture human behavioral cues without the need for coding. Our user study, employing the System Usability Scale (SUS) and satisfaction metrics, demonstrated high user preference, reflecting the tool’s ease of use and effectiveness in supporting scene search tasks. Additionally, through a deployment trial within industrial organizations, we confirmed the tool’s objectivity, reusability, and potential to enhance expert workflows. This suggests the advantages of expert-AI collaboration in domains requiring human contextual understanding and demonstrates how customizable, transparent tools yielding reusable artifacts can support expert-driven tasks in complex, multimodal environments.
Riku Arakawa, Kiyosu Maeda, Hiromu Yakura
ACM Trans. Interact. Intell. Syst.2
2023 BlendMR: A Computational Method to Create Ambient Mixed Reality Interfaces
abstract
Mixed Reality (MR) systems display content freely in space, and present nearly arbitrary amounts of information, enabling ubiquitous access to digital information. This approach, however, introduces clutter and distraction if too much virtual content is shown. We present BlendMR, an optimization-based MR system that blends virtual content onto the physical objects in users’ environments to serve as ambient information displays. Our approach takes existing 2D applications and meshes of physical objects as input. It analyses the geometry of the physical objects and identifies regions that are suitable hosts for virtual elements. Using a novel integer programming formulation, our approach then optimally maps selected contents of the 2D applications onto the object, optimizing for factors such as importance and hierarchy of information, viewing angle, and geometric distortion. We evaluate BlendMR by comparing it to a 2D window baseline. Study results show that BlendMR decreases clutter and distraction, and is preferred by users. We demonstrate the applicability of BlendMR in a series of results and usage scenarios.
Violet Yinuo Han, Hyunsung Cho, Kiyosu Maeda, Alexandra Ion, David Lindlbauer
Proc. ACM Hum. Comput. Interact.3
2022 CalmResponses: Displaying Collective Audience Reactions in Remote Communication
abstract
We propose a system displaying audience eye gaze and nod reactions for enhancing synchronous remote communication. Recently, we have had increasing opportunities to speak to others remotely. In contrast to offline situations, however, speakers often have difficulty observing audience reactions at once in remote communication, which makes them feel more anxious and less confident in their speeches. Recent studies have proposed methods of presenting various audience reactions to speakers. Since these methods require additional devices to measure audience reactions, they are not appropriate for practical situations. Moreover, these methods do not present overall audience reactions. In contrast, we design and develop CalmResponses, a browser-based system which measures audience eye gaze and nod reactions only with a built-in webcam and collectively presents them to speakers. The results of our two user studies indicated that the number of fillers in speaker’s speech decreases when audiences’ eye gaze is presented, and their self-rating score increases when audiences’ nodding is presented. Moreover, comments from audiences suggested benefits of CalmResponses for them in terms of co-presence and privacy concerns.
Kiyosu Maeda, Riku Arakawa, Jun Rekimoto
IMX1
2020 BulkScreen: Saliency-Based Automatic Shape Representation of Digital Images with a Vertical Pin-Array Screen
abstract
Digital images appearing on displays in everyday activities (e.g., photos on a smartphone) are automatically and instantly rendered without manual intervention such that we can seamlessly appreciate them. In contrast, shape displays require manual designs of outputs upon actuation of input images to render 3D shapes. In this work, we aim to achieve automatic and on-the-spot actuation of digital images so that we can seamlessly see 3D physical images. To this end, we developed BulkScreen, an image projection system that can automatically render 3D shapes of input images on a vertical pin-array screen. Our approach is based on a deep-neural-network saliency estimation coupled with our post-processing algorithm. We believe this spontaneous actuation mechanism facilitates applications with shape displays such as real-time picture browsing and display advertisement, building on the benefit of representing physical shapes; tangibility.
Riku Arakawa, Yudai Tanaka, Hiromu Kawarasaki, Kiyosu Maeda
TEI4