Yuya Moroto

dblp:232/2947 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
7since 2021 · last 2023
0000-0003-3962-1712ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Multi-View Variational Recurrent Neural Network for Human Emotion Recognition Using Multi-Modal Biological Signals
abstract
In this paper, the Multi-view Variational Recurrent Neural Network (MvVRNN) is proposed for multi-modal human emotion recognition with gaze and brain activity data while humans view images. For realizing accurate emotion recognition, we focus on the following three characteristics of biological signals: 1) the relationship between implicit and explicit information such as gaze and brain activity data, 2) the temporal changes related to human emotions and 3) the effects of noises that can be included during data acquisition. For treating these characteristics, the proposed MvVRNN has several mechanisms including 1) the integration of multi-modal information including implicit and explicit states of humans, 2) the recurrent module for sequential data and 3) the variational approximation based on the Gaussian distribution. The experimental results show that emotion recognition based on the MvVRNN outperforms several existing methods.
Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
ICIP1
2023 Personalized Content Recommender System via Non-verbal Interaction Using Face Mesh and Facial Expression
abstract
Multimedia content recommendation needs to consider users' preferences for each content. Conventional recommender systems consider them with wearable sensors, however, wearing such sensors can lead to a burden on users. In this paper, we construct a recommender system that can explicitly estimate users' preferences without wearable sensors. Specifically, by constructing lightweight but strong machine learning models suitable for our system, the users' interest levels for contents can be estimated from facial images obtained from a widely used webcam. In addition, through the interaction that the user selects displayed contents, our system finds the tendency of personal preferences for recommending contents with high user satisfaction. Our system is available on https://www.lmd-demo.org/2022/start_eng.html.
Yuya Moroto, Rintaro Yanagi, Naoki Ogawa, Kyohei Kamikawa, Keigo Sakurai, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
ACM Multimedia1
2022 Human Emotion Recognition Using Multi-Modal Biological Signals Based On Time Lag-Considered Correlation Maximization
abstract
A human emotion recognition using multi-modal biological signals based on time lag-considered correlation maximization is presented in this paper. Various multi-modal emotion recognition methods for visual stimuli have been studied and they focus on gaze and brain activity data. The visual stimuli captured by human eyes are sent to the brain by neurotransmitters. Thus, there is a time lag between gaze data, which record where humans gaze at, and brain activity data. However, most of the previous methods only integrate features obtained from each data without considering such a time lag. The proposed method newly introduces the mechanism to consider the time lag into the canonical correlation analysis scheme by assuming that the influence of the visual stimuli on brain activity data follows the Poisson distribution. The contribution of this paper is the construction of a recognition method with considering the time lag for getting truly close to the realization of the occurrence mechanism of human emotions. Experimental results show the effectiveness of considering the time lag between gaze and brain activity data.
Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
ICASSP1
2022 Few-Shot Personalized Saliency Prediction with Similarity of Gaze Tendency Using Object-Based Structural Information
abstract
This paper presents a few-shot personalized saliency prediction method with similarity of gaze tendency using object-based structural information. The personalized saliency maps (PSMs) that represent the individual visual attention can be used for analyzing the heterogeneity among personalized preferences, whereas general saliency maps ignore the individual differences. However, the PSM prediction is a difficult task since the acquisition of eye tracking data, which are needed for obtaining PSMs, gives persons a heavy burden. Then, for realizing PSM prediction with a limited amount of training data, the use of the similarity of gaze tendency between persons can be one effective way. It has been reported that human gazes are related to objects and their relative relationships that are semantic and structural information, and we focus on the integration of PSMs predicted for other persons by using object-based similarities of gaze tendency. The advantage in this paper is that we newly focus on similarities of gaze tendency for visually similar objects for solving the lack of eye tracking data by considering both semantic and structural information, simultaneously. By experimenting with the open dataset, the proposed method outperforms the state-of-the-art methods.
Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
ICIP1
2022 Visual Sentiment Prediction Using Cross-Way Few-Shot Learning Based on Knowledge Distillation
abstract
This paper presents a visual sentiment prediction method using cross-way few-shot learning based on knowledge distillation. Previous studies on visual sentiment prediction methods have focused only on one sentiment dataset although there are several sentiment datasets following different sentiment theories. Originally, sentiments are abstract notions common to humans regardless of the difference between sentiment theories. Thus, the use of knowledge obtained from several sentiment datasets can be the effective way to realize robust visual sentiment prediction. To collaboratively use sentiment datasets, there are the following two concerns: training of the model introducing different sentiment theories and prediction of a newly given sample whose sentiment theory is unknown. Thus, we focus on knowledge distillation, which can improve the generalization ability of several tasks, and effective training becomes feasible. In addition, to deal with the different numbers of sentiment labels in a test phase, we newly introduce the cross-way few-shot learning scheme into knowledge distillation. The main contribution in this paper is to integrate knowledge distillation and the cross-way approach for the visual sentiment prediction, and this is the first work for dealing with the difference of sentiment datasets used in the training and test phases. In the end of this paper, the effectiveness of the proposed method is confirmed through experiments using several open datasets.
Yingrui Ye, Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
ICIP2
2022 Affective Embedding Framework with Semantic Representations from Tweets for Zero-Shot Visual Sentiment Prediction
abstract
This paper presents a zero-shot visual sentiment prediction method using semantic representation features of texts from tweets as the non-visual auxiliary data. Previous studies show that visual sentiment prediction methods can only predict the sentiment labels that are the same as the labels of the sentiment theory used in the training dataset, which means that they cannot predict the new sentiment label used in different sentiment theories. To solve the problem of predicting new labels, zero-shot learning has been proposed. The previous zero-shot visual sentiment prediction method uses Word2vec features and the adjective-noun pair features to obtain the semantical relationship between images and sentiment words to predict unseen sentiments. However, many adjective-noun pairs are not related to sentiments, which makes it difficult to compensate for an affective gap between low-level visual features and high-level sentiment semantics. Thus, to better compensate for the affective gap, it is considered to introduce the new non-visual auxiliary data. As people tend to share their feelings with both images and texts on social networking services, the texts from tweets are effective as the side information of the images in visual sentiment prediction. Thus, we introduce the semantic representations from tweets as the new non-visual auxiliary data to construct an affective embedding space, which makes a more effective zero-shot visual sentiment prediction model. Moreover, we propose a cross-dataset zero-shot task for visual sentiment prediction, which is more consistent with the real situation that the testing and training images may be in different domains. The contributions in this paper are to combine several semantic representation features for zero-shot visual sentiment prediction and the proposal of the cross-dataset zero-shot task for visual sentiment prediction. The experiments on several open datasets show the effectiveness of the proposed method.
Yingrui Ye, Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
MMAsia2
2021 Few-Shot Personalized Saliency Prediction using Person Similarity based on Collaborative Multi-Output Gaussian Process Regression
abstract
A few-shot personalized saliency prediction method using person similarity based on collaborative multi-output Gaussian process regression is presented in this paper. Contrary to prediction of general saliency maps, that of personalized saliency maps (PSMs), which is a focus of attention owing to its heterogeneity among individuals, is a challenging problem since the amount of training gaze data is limited due to the burden on new persons. Thus, the proposed method focuses on the similarity of gaze tendency between persons. In the proposed method, collaborative Gaussian process regression (CoMOGP) is adopted for PSM prediction. CoMOGP enables to represent similarity of gaze tendency between the target person and other persons as weights, and then consider the similarity for each image by using visual features obtained from images as inputs. The contributions of the few-shot PSM prediction based on CoMOGP are two-folds. 1) CoMOGP, which is one of probabilistic methods, can avoid the overfitting to small amount of training data. 2) Similarity for each image can be considered by using visual features as inputs. In the experiment using the open dataset, the proposed method outperforms comparative methods including the state-of-the-art method.
Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
ICIP1
2019 Estimation of Emotion Labels via Tensor-Based Spatiotemporal Visual Attention Analysis
abstract
This paper presents emotion label estimation via tensor-based spatiotemporal visual attention analysis. It has been reported in the fields of psychology and neuroscience that human emotions are related to two elements, their visual attention change and objects included in a target image. Therefore, the proposed method focuses on the spatiotemporal change of visual attention of human gazing at objects in the target image and constructs two neural networks which enable the emotion label estimation considering both of the above two elements. Specifically, the proposed method newly constructs a fourth-order tensor, gaze and image tensor (GIT) whose modes correspond to the width, the height and the color channel of the target image and the time axis of visual attention which is used for representing the time change. Then the first network, which consists of general tensor discriminant analysis (GTDA) and extreme learning machine (ELM), estimates the emotion label from the fourth-order GIT with concerning their visual attention change. Furthermore, the second network, which consists of pre-trained convolutoinal neural network-based feature extraction, GTDA and ELM, enables the estimation from the second-order GIT including visual features obtained from objects focused at each time. Finally, the proposed method estimates emotion labels based on decision fusion of the outputs from the two networks. Experimental results show the effectiveness of the proposed method.
Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama
ICIP1