Ryo Ueda

dblp:191/3366 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2024 VAD Emotion Control in Visual Art Captioning via Disentangled Multimodal Representation
abstract
Art evokes distinct affective responses, leading to the generation of emotional verbal expressions in response to visual stimuli like visual art images. While previous research has addressed controlling linguistic impressions in terms of basic emotion categories, less attention has been given to their continuous, dimensional nature. This paper aims to continuously modulate emotions across valence, arousal, and dominance (VAD) dimensions by employing multimodal representation learning (MMRL) and crossmodal inference. We utilize a multimodal variational autoencoder for MMRL, encoding visual and textual stimuli into a continuous joint vector representation. We then auxiliarily tune this representation to explicitly include the VAD dimensions in a disentangled manner. This enables nuanced emotional control in image captioning, where visual inputs are converted into vectors and then into text. Trained on the ArtEmis dataset, which includes emotion-evoking captions for visual art, as well as additional datasets annotated with VAD scores for either text or visual art, our model demonstrates that manipulating VAD intensities in the vector representation results in continuous caption variations, reflecting the intended emotional changes.
Ryo Ueda, Hiromi Narimatsu, Yusuke Miyao, Shiro Kumano
ACII1
2024 Emergent Word Order Universals from Cognitively-Motivated Language Models
abstract
Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin
ACL (1)2
2024 Emergent Communication with Stack-Based Agents
Daichi Kato, Ryo Ueda, Jason Naradowsky, Yusuke Miyao
CogSci2
2024 Lewis's Signaling Game as beta-VAE For Natural Word Lengths and Segments
abstract
As a sub-discipline of evolutionary and computational linguistics, emergent communication (EC) studies communication protocols, called emergent languages, arising in simulations where agents communicate. A key goal of EC is to give rise to languages that share statistical properties with natural languages. In this paper, we reinterpret Lewis's signaling game, a frequently used setting in EC, as beta-VAE and reformulate its objective function as ELBO. Consequently, we clarify the existence of prior distributions of emergent languages and show that the choice of the priors can influence their statistical properties. Specifically, we address the properties of word lengths and segmentation, known as Zipf's law of abbreviation (ZLA) and Harris's articulation scheme (HAS), respectively. It has been reported that the emergent languages do not follow them when using the conventional objective. We experimentally demonstrate that by selecting an appropriate prior distribution, more natural segments emerge, while suggesting that the conventional one prevents the languages from following ZLA and HAS.
Ryo Ueda, Tadahiro Taniguchi
ICLR1
2023 Emotion-Controllable Impression Utterance Generation for Visual Art
abstract
The degree of subjectivity and the type of emotions when people express their impressions of an object depend on various factors, including their affective states, psychological traits, goals, and social norms. However, most previous efforts in generating people’s impressions of objects have focused on extreme cases, i.e., fully objective or fully emotional, as represented by the MS-COCO and ArtEmis datasets, respectively. We propose an emotional impression generation method that continuously controls the degree of subjectivity versus objectivity and the type of emotion in a unified manner. In the framework of ConCap, which allows controlling text style by modifying an auxiliary input called prefix, we propose to use two types of prefixes to jointly control the degree of subjectivity and the type of emotion: subjective-style prefix and categorical-emotional-style prefix. The subjective-style prefixes are holistic and attempt to describe the entire emotion felt by a group of people, although different individuals may feel different emotions. The categorical-emotional-style prefix is more individual-oriented and tries to focus on a single or few specific emotions. An experiment using both objectively and subjectively descriptive datasets shows qualitatively and quantitatively that the proposed method provides good gradual control of impression expressions of visual art in terms of both subjectivity and objectivity as well as emotion categories.
Ryo Ueda, Hiromi Narimatsu, Yusuke Miyao, Shiro Kumano
ACII1
2023 Emergent Communication with Attention
Ryokan Ri, Ryo Ueda, Jason Naradowsky
CogSci2
2023 On the Word Boundaries of Emergent Languages Based on Harris's Articulation Scheme
Ryo Ueda, Taiga Ishii, Yusuke Miyao
ICLR1
2022 Cross-Linguistic Study on Affective Impression and Language for Visual Art Using Neural Speaker
abstract
Visual art is one of the ideal targets for affective computing because viewing artworks is a regular experience for many people, and it elicits various affective appraisals, cognitions, and reactions in the viewer. To gain a detailed understanding of the interplay between visual content, its emotional impact, and linguistic explanations of this impact, visual art datasets have been proposed. One example is the ArtEmis dataset, which contains emotion categorization and linguistic expressions by crowds of people in reaction to numerous paintings. Linguistic expressions are influenced both by culture in how to appraise art and by language in how to verbalize one's impression. However, cultural and linguistic differences in this domain have not been fully explored. Therefore, we collected a new dataset (ArtEmis-JP) consisting of 16 k emotion labels and utterances in Japan, one of the countries most frequently compared with Western countries in psychology, while carefully following the procedures of the original ArtEmis study conducted with English speakers. In this paper, we report the commonalities and differences between the original ArtEmis dataset and our Japanese dataset using basic statistical comparisons and performance comparisons for a state-of-the-art neural speaker in utterance generation. Going beyond the original study, we newly examined the impact of expertise on emotional categorization and linguistic expressions by examining the differences between experts and non-experts.
Hiromi Narimatsu, Ryo Ueda, Shiro Kumano
ACII2