VLDB 2026 Research / reviewers in the wild / expert
Keisuke Maeda
dblp:134/3015
· DBLP profile ↗
72ranked-venue papers
9as first author
56since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 8 first-author · 49 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adversarial Perturbation Shield: Preventing Concept Bleed-through in Continual Learning of Personalized Generative ModelsabstractPersonalized text-to-image diffusion models have gained increasing attention because they can generate images that contain unique concepts based on limited training data. However, in continual learning scenarios, these models suffer from concept bleed-through, where newly introduced concepts frequently overwrite or interfere with the previously learned concepts. Previous studies have attempted to mitigate this issue at the model adaptation level; however, they failed to fully preserve the distinct semantic representations in the latent space. Thus, this paper proposes an adversarial perturbation-based training strategy to address concept bleed-through in continual learning for personalized diffusion models. The proposed method introduces adversarial perturbations into the training images, which strategically shifts their semantic representations in the latent space to ensure that the newly learned concepts remain distinct and do not interfere with the previously acquired knowledge. Unlike structural modifications to the model, the proposed method operates at the data level, which makes it broadly applicable to existing continual personalization frameworks without increasing model complexity. Experimental results demonstrate that the proposed method significantly improves concept separation while maintaining high image fidelity, offering a solution to enhance the reliability of continual learning in personalized generative models. Ziwen Lan, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 2 |
| 2026 | Generalizing Stylized Motion Generation Method by Introducing Metadata-Independent Learning and Unified Multiple Motion DatasetabstractThis study aims to extend the applicability of stylized motion generation methods to be robust for large and diverse motions akin to those found in real-world data. Specifically, we introduce metadata-independent learning alongside style-focused learning, thereby enabling training from motions absent in motion-style datasets. In addition, we construct a novel motion dataset containing both various motions and stylized motions by unifying the multiple datasets to effectively train the model. Our novel learning method and dataset enable stylized motion generation methods to learn from both various motion knowledge and motion-style relations and improve their generalized performance. In downstream tasks, we address motion style transfer and text-to-stylized-motion, validating the enhancement of generalization abilities for each task. Compared to conventional methods, the proposed method demonstrates superior performance in generating and reflecting style, particularly under conditions featuring larger and more diverse motions. The code for this paper and the dataset will be released. Yuki Era, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
IEEE Trans. Multim. | 3 |
| 2025 | Hyperboloid GPLVM for Discovering Continuous Hierarchies via Nonparametric EstimationabstractDimensionality reduction (DR) offers interpretable representations of complex high-dimensional data, and recent DR methods have leveraged hyperbolic geometry to obtain faithful low-dimensional embeddings of high-dimensional hierarchical relationships. However, existing methods are dependent on neighbor embedding, which frequently ruins the continuous nature of the hierarchical structures. This paper proposes hyperboloid Gaussian process latent variable models (hGP-LVMs) to embed high-dimensional hierarchical data while preserving the implicit continuity via nonparametric estimation. We adopt generative modeling using the GP, which provides effective hierarchical embedding and executes ill-posed hyperparameter tuning. This paper presents three variants of the proposed models that employ original point, sparse point, and Bayesian estimations, and we establish their learning algorithms by incorporating the Riemannian optimization and active approximation scheme of the GP-LVM. In addition, we employ the reparameterization trick for scalable learning of the latent variables in the Bayesian estimation method. The proposed hGP-LVMs were applied to several datasets, and the results demonstrate their ability to represent high-dimensional hierarchies in low-dimensional spaces. Koshi Watanabe, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
AISTATS | 2 |
| 2025 | Gradient-Oriented Clustered Federated Learning With Efficient Knowledge Sharing in Non-IID SettingsabstractWe present a novel Clustered Federated Learning (CFL) approach that efficiently shares knowledge among clusters to address non-Independent and Identically Distributed (non-IID) settings. Although conventional CFL has demonstrated strong performance in non-IID settings, a fundamental challenge, the lack of knowledge sharing among clusters remains as a problem to be solved for improvement of generalization. To overcome this limitation, we utilize personalized layers for specific tasks while retaining globally shared layers after clustering. By separating the layers of the model according to their functions, our method enables the model to be well personalized without sacrificing generalizability. Furthermore, we cluster clients based on the gradients of specific layers. Our efficient CFL method can better characterize the data of each client, reduce the memory demand of the client, and mitigate the computational burden of the FL server. We evaluate the effectiveness of our method in various data distribution settings. Kenta Kubota, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 3 |
| 2025 | Generative Dataset Distillation Based on Self-knowledge DistillationabstractDataset distillation is an effective technique for reducing the cost and complexity of model training while maintaining performance by compressing large datasets into smaller, more efficient versions. In this paper, we present a novel generative dataset distillation method that can improve the accuracy of aligning prediction logits. Our approach integrates self-knowledge distillation to achieve more precise distribution matching between the synthetic and original data, thereby capturing the overall structure and relationships within the data. To further improve the accuracy of alignment, we introduce a standardization step on the logits before performing distribution matching, ensuring consistency in the range of logits. Through extensive experiments, we demonstrate that our method outperforms existing state-of-the-art methods, resulting in superior distillation performance. Longzhen Li, Guang Li 0008, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 4 |
| 2025 | Robust Adversarial Defense Based on Non-Transferability of Attack Across Foundation ModelsabstractThis paper presents a novel adversarial defense method that exploits non-transferability of attack across foundation models. Existing adversarial training methods have insufficient robustness to adversarial examples under powerful adversarial attacks in a white-box setting. We clarify that there is no attack transferability between various foundation models through preliminary experiments and propose a defense method that combines the non-transferability with adversarial training. The proposed method constructs a common embedding space of different models and detects adversarial examples by leveraging the different outputs from each model. Based on the detection of adversarial examples, the proposed method can adaptively select a robust model depending on the input images based on non-transferability of attack between models. Experimental results show that the proposed method improves the accuracy of zero-shot class classification for adversarial examples against each model. Koshiro Toishi, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2025 | Triplet Synthesis for Enhancing Composed Image Retrieval via Counterfactual Image GenerationabstractComposed Image Retrieval (CIR) provides an effective way to manage and access large-scale visual data. Construction of the CIR model utilizes triplets that consist of a reference image, modification text describing desired changes, and a target image that reflects these changes. For effectively training CIR models, extensive manual annotation to construct high-quality training datasets, which can be time-consuming and labor-intensive, is required. To deal with this problem, this paper proposes a novel triplet synthesis method by leveraging counterfactual image generation. By controlling visual feature modifications via counterfactual image generation, our approach automatically generates diverse training triplets without any manual intervention. This approach facilitates the creation of larger and more expressive datasets, leading to the improvement of CIR model’s performance. Kenta Uesugi, Naoki Saito 0006, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 3 |
| 2025 | Out-of-Distribution Sample Selection Generated by Diffusion Model toward Model GeneralizationabstractModel generalization that prevents overfitting to the training data is critical to the robustness and reliability of image classification models. Diffusion models are rapidly developing in generating photorealistic images from text, yet it is difficult for a simple text to generate a variety of images that reflect the training data, and there is potential for model generalization. In this paper, we propose a novel method for model generalization with generated images via diffusion models. We make the diffusion model recognize the training data distribution in both the latent and the image space, and obtain generated images reflecting the distribution. In the latent space, we synthesize new features from the training data distribution and decode them into images using the diffusion model. In the image space, we perform selection of unintentionally generated images (out-of-distribution samples) based on the semantic information of the training data. By considering the training data distribution not only in the latent space but also in the image space, we eliminate out-of-distribution samples generated by the diffusion model. Experiments demonstrate that our generated images are faithful to the training data distribution and enhance model generalization performance. Kaede Hayakawa, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2025 | Enhancing Adversarial Robustness of Foundation Models Without Data CentralizationabstractThis paper investigates the adversarial defense for foundation models without data centralization. To deploy foundation models in safety-critical applications, they must be robust against diverse attacks. In the defense against a single attack, foundation models require numerous adversarial examples, leading to a substantial increase in data and computational demands when addressing multiple attacks. However, traditional adversarial defense must centralize all types of adversarial examples, and there has been insufficient investigation into adversarial defense for foundation models without data centralization. Accordingly, we focus on model merging, which integrates fine-tuned models for a specific task to construct a merged model adaptable to any task. Through model merging, we aim to integrate individual models’ defensive capabilities to enhance robustness against various attacks. The main contribution of this paper is demonstrating that defense methods leveraging model merging can achieve classification accuracy approaching that of the defense method with data centralization under multiple attacks. Koshiro Toishi, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2025 | Context-aware Image-to-Music Generation via Bridging Modalities through Musical CaptionsabstractMusic generation has developed remarkably, and research on music generation from text has progressed. However, there are cases where it is difficult to use text as the query when users want to obtain music that is suitable for images, while image-to-music generation is still unexplored. There is a gap between images, which are visual content, and music, and it has been difficult to associate them directly. To address this problem, we realize an image-to-music generation method through musical captions which describe the characteristics of the music. Musical captions possess an enhanced capability to steer the process of music generation. Therefore, if musical captions effectively convey the intended message of the image, they serve as an excellent intermediary between the images and music. The proposed method connects these two different modalities through the medium of musical captions that describe the specialized content of the music. By generating images through musical captions using multi-modal large language models, we construct an image-musical caption pair dataset. Using query-similar images and their paired musical captions in the dataset, in-context learning on a multi-modal large language model is conducted to generate the musical caption corresponding to the target image. The generated musical caption is then input into a text-to-music generative model, and thus, the proposed method enables high-quality image-to-music generation. We conducted experiments to evaluate the quality of the generated music and the consistency with target images through both subjective and objective metrics. The results confirm the effectiveness of our proposed method. The code can be found at https://github.com/lsllsls/CAI2M. Kyohei Kamikawa, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ACM Multimedia | 3 |
| 2025 | Hyperbolic Dataset DistillationabstractTo address the computational and storage challenges posed by large-scale datasets in deep learning, dataset distillation has been proposed to synthesize a compact dataset that replaces the original while maintaining comparable model performance. Unlike optimization-based approaches that require costly bi-level optimization, distribution matching (DM) methods improve efficiency by aligning the distributions of synthetic and original data, thereby eliminating nested optimization. DM achieves high computational efficiency and has emerged as a promising solution. However, existing DM methods, constrained to Euclidean space, treat data as independent and identically distributed points, overlooking complex geometric and hierarchical relationships. To overcome this limitation, we propose a novel hyperbolic dataset distillation method, termed HDD. Hyperbolic space, characterized by negative curvature and exponential volume growth with distance, naturally models hierarchical and tree-like structures. HDD embeds features extracted by a shallow network into the Lorentz hyperbolic space, where the discrepancy between synthetic and original data is measured by the hyperbolic (geodesic) distance between their centroids. By optimizing this distance, the hierarchical structure is explicitly integrated into the distillation process, guiding synthetic samples to gravitate towards the root-centric regions of the original data distribution while preserving their underlying geometric characteristics. Furthermore, we find that pruning in hyperbolic space requires only 20\% of the distilled core set to retain model performance, while significantly improving training stability. Notably, HDD is seamlessly compatible with most existing DM methods, and extensive experiments on different datasets validate its effectiveness. To the best of our knowledge, this is the first work to incorporate the hyperbolic space into the dataset distillation process. The code is available at https://github.com/Guang000/HDD. Guang Li 0008, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
NeurIPS | 3 |
| 2025 | Cross-domain multi-step thinking: Zero-shot fine-grained traffic sign recognition in the wild
Yaozong Gan, Guang Li 0008, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
Knowl. Based Syst. | 4 |
| 2025 | Linear Structure Analysis of Embeddings for Bias Disparity Reduction in Collaborative FilteringabstractRecommender systems personalize user experiences by filtering large volumes of information, and further shape user behavior. Collaborative filtering (CF), a widely used algorithm in this domain, learns user preferences from user-item interactions. However, recent studies indicate that certain CF algorithms can excessively amplify inherent data biases, manifesting as bias disparity. In this study, we examine the latent factor model (LFM), a state-of-the-art CF method that represents users and items as vectors in a shared latent space. By applying linear dimensionality reduction techniques with strong interpretability (such as principal component analysis) to LFM embeddings trained on real-world data, we identify specific axes that encode biases. We then demonstrate the application of these linear relationships in mitigating bias disparity. Experimental results show that our proposed approach can reduce bias disparity with only a slight decrease in recommendation accuracy-an average of$5-8\%$with principal component analysis and$1-3\%$with independent component analysis. Hiroki Okamura, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | Privacy Preserving Gaze Estimation Via Federated Learning Adapted To Egocentric VideoabstractThis paper presents privacy preserving gaze estimation via federated learning adapted to egocentric videos. Gaze estimation with egocentric video stands out among many applications of wearable cameras and is closely related to many high-tech applications envisioned in the future. However, traditional gaze estimation methods mainly depend on a centralized training model, which has the risk of leakage of private information. In this paper, we propose an innovative transformer-based framework that integrates the principles of federated learning into the gaze estimation process using egocentric video data. By training the model without sharing raw gaze data and only updating parameters, this framework can achieve the dual objectives of enhancing the model’s performance and protecting data privacy. Experimental results demonstrate that our framework not only outperforms other federated learning methods but also achieves performance close to that of methods with data sharing. Yuhu Feng, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2024 | Enhancing Noisy Label Learning Via Unsupervised Contrastive Loss with Label Correction Based on Prior KnowledgeabstractTo alleviate the negative impacts of noisy labels, most of the noisy label learning (NLL) methods dynamically divide the training data into two types, "clean samples" and "noisy samples", in the training process. However, the conventional selection of clean samples heavily depends on the features learned in the early stages of training, making it difficult to guarantee the cleanliness of the selected samples in scenarios where the noise ratio is high. In addition, their optimization processes are based on the supervised loss including noisy labels, and effective representation cannot be obtained in the presence of a large number of noisy labels. To address these problems, we propose an effective method capable of robustly performing NLL under extremely high noise ratios. In the proposed method, by introducing the prior knowledge of a pre-trained vision and language model, we can effectively select clean samples since it does not depend on the learning process of NLL. Moreover, the introduction of an unsupervised contrastive learning approach enables the acquisition of noise-robust feature representations in SSL. Experiments with synthetic label noise on CIFAR-10 and CIFAR-100, the benchmark datasets in NLL, demonstrate that the proposed method significantly outperforms state-of-the-art methods. Masaki Kashiwagi, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2024 | Multi-Object Editing in Personalized Text-To-Image Diffusion Model Via Segmentation GuidanceabstractThis paper presents a personalized text-to-image diffusion model for multiple object editing that can improve visual fidelity of the target image and editing ability with a segmentation-based restriction and continual learning. Multiple personalization tasks face the problem of destabilization, especially when the number of targets increases and the concepts of the targets are similar. The proposed method introduces a segmentation guide into continual learning to improve performance for multiple objects. The segmentation guide helps to separate each concept by restricting the regions of target objects during both training and inference. The proposed method learns these concepts by continual learning with Elastic Weight Consolidation, and achieves the output of multiple target objects with concept separation while maintaining visual fidelity. Experimental results demonstrate that the proposed method successfully maintains visual fidelity for multiple target objects. Haruka Matsuda, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 3 |
| 2024 | Caption Unification for Multi-View Lifelogging Images Based on In-Context Learning with Heterogeneous Semantic ContentsabstractThis paper presents a new task of caption unification and a novel caption unification method for multi-view lifelogging images based on in-context learning with heterogeneous semantic contents. Most of the existing image captioning models target a single image and do not consider the common semantic contents among multiple images. Therefore, they suffer from inconsistent captioning for multi-view lifelogging images taken in the same scene, i.e., the images that have common semantic contents regardless of the viewpoints. The proposed method enables training the common semantic contents of multiple images and generating captions that comprehensively contain the semantic contents of the images based on in-context learning. The proposed method has the following two contributions. First, we identify the problem of existing image captioning models and consider caption unification as a new task for multi-view lifelogging images. Second, we generate unified captions based on in-context learning with heterogeneous semantic contents and evaluate the captions from two perspectives, uniformity and retainability. The experimental results demonstrate that the proposed method realizes successful caption unification. Masaya Sato, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2024 | Cross-Domain Few-Shot In-Context Learning For Enhancing Traffic Sign RecognitionabstractIn this paper, we propose a cross-domain few-shot in-context learning method based on the multimodal large language model (MLLM) for enhancing traffic sign recognition (TSR). We first construct a traffic sign detection network based on Vision Transformer Adapter and an extraction module to extract traffic signs from the original road images. To reduce the dependence on training data and improve the performance stability of cross-country TSR, we introduce a cross-domain few-shot in-context learning method based on the MLLM. To enhance MLLM’s fine-grained recognition ability of traffic signs, the proposed method generates corresponding description texts using template traffic signs. These description texts contain key information about the shape, color, and composition of traffic signs, which can stimulate the ability of MLLM to perceive fine-grained traffic sign categories. By using the description texts, our method reduces the cross-domain differences between template and real traffic signs. Our approach requires only simple and uniform textual indications, without the need for large-scale traffic sign images and labels. We perform comprehensive evaluations on the German traffic sign recognition benchmark dataset, the Belgium traffic sign dataset, and two real-world datasets taken from Japan. The experimental results show that our method significantly enhances the TSR performance. Yaozong Gan, Guang Li 0008, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 4 |
| 2024 | Reinforcing Pre-Trained Models Using Counterfactual ImagesabstractThis paper proposes a novel framework to reinforce classification models using language-guided generated counterfactual images. Deep learning classification models are often trained using datasets that mirror real-world scenarios. In this training process, because learning is based solely on correlations with labels, there is a risk that models may learn spurious relationships, such as an overreliance on features not central to the subject, like background elements in images. However, due to the black-box nature of the decision-making process in deep learning models, identifying and addressing these vulnerabilities has been particularly challenging. We introduce a novel framework for reinforcing the classification models, which consists of a two-stage process. First, we identify model weaknesses by testing the model using the counterfactual image dataset, which is generated by perturbed image captions. Subsequently, we employ the counterfactual images as an augmented dataset to fine-tune and reinforce the classification model. Through extensive experiments on several classification models across various datasets, we revealed that fine-tuning with a small set of counterfactual images effectively strengthens the model. Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 3 |
| 2024 | Flexibly manipulating popularity bias for tackling trade-offs in recommendation
Hiroki Okamura, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
Inf. Process. Manag. | 2 |
| 2023 | Improving Dropout in Graph Convolutional Networks for Recommendation via Contrastive LossabstractWe propose a novel graph convolutional network (GCN)-based recommendation model that incorporates a contrastive loss. Although GCN-based recommendation models achieve high recommendation performance, existing models suffer from over-fitting since they explicitly encode the interactions as a graph. To mitigate this problem, while the models perform dropout which just randomly drops edges in the graph, this can miss the essential information that represents users’ preferences. The proposed method improves and extends the dropout via the contrastive loss and enhances the robustness while preserving the essential information. The contrastive loss can correlate the representations in different perturbed graphs and promote their consistency. Conceptually, our method captures users’ preferences which are invariant even in the perturbed interactions. The experimental results demonstrate that our method improves the recommendation performance in real-world datasets. Hiroki Okamura, Keisuke Maeda, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2023 | Estimation of Visual Contents from Human Brain Signals via VQA Based on Brain-Specific AttentionabstractThis paper presents a method for estimation of visual cognitive contents from human brain signals via a newly derived visual question answering (VQA) model. The proposed method can estimate a wide range of cognitive contents from functional magnetic resonance imaging data when subjects viewed images via the VQA model which can effectively utilize low-level and high-level image features extracted from the viewed images. To make use of multiple image features, we newly introduce brain-specific attention into the VQA model. The brain-specific attention can adaptively determine the significance of image features at each level depending on the complexity of the cognitive contents (e.g., category, pattern, number and color). Therefore, we enable flexible and extensive estimation of various cognitive contents. Experimental results show that the proposed method can significantly improve the performance of the visual content estimation. Ryo Shichida, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 3 |
| 2023 | Learning Graph Laplacian from Intrinsic Patterns via Gaussian ProcessabstractIn this paper, we present a novel scheme to learn a graph topology named Laplacian constrained Gaussian process (LCGP). Previous graph learning methods directly use raw observed signals, which often degrades estimation accuracy due to the noise effects and the sample imbalances. LCGP tackles this problem by introducing a small number of intrinsic patterns, that is, LCGP assumes representative signals on a graph. These signals are derived as Bayesian latent variables, and we estimate the graph topology with objective after marginalizing them. Furthermore, the number of the intrinsic patterns is automatically determined by a Bayesian inference procedure. In the experiment, we compared LCGP with baseline and state-of-the-art methods in graph learning and confirmed the effectiveness of LCGP. Koshi Watanabe, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2023 | Defense Against Black-Box Adversarial Attacks Via Heterogeneous Fusion FeaturesabstractThis paper presents an effective approach for the adversarial defense task named a heterogeneous feature fusion network (HFFN). Inspired by the fact that humans can utilize multimodal information to help themselves perceive objects, we introduce the caption features into the classic convolutional neural networks (CNNs) and fuse them with traditional image features. To reduce the "modal gap" between heterogeneous features, we introduce the hetero-center loss that can constrain the distance between the class centers of different modalities. In addition, we further integrate the guided complement entropy loss into HFFN, which can restrict the prediction probability of the model on non-ground truth classes, so that the adversarial robustness of fused heterogeneous features can be deeply improved. We further adopt three state-of-the-art comparison approaches and design four ablation methods to evaluate the effectiveness and adversarial robustness of HFFN. Moreover, in order to compare the performance of HFFN with the comparison methods and ablation methods, we also utilize a substitute model to generate a variety range of adversarial examples. Extensive experiments based on these adversarial examples for the CIFAR-10 and CIFAR-100 datasets exhaustively and strongly demonstrate the superiority of our method. Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2023 | Video-Music Retrieval with Fine-Grained Cross-Modal AlignmentabstractThis paper presents a novel video-music retrieval method for videos containing humans. Our method constructs a cross-modal common embedding space of video features based on human motion and music features and aligns the embedding space based on the beat and genre information. The beat and genre are distinctive properties of music, and music relates motion features better than features extracted from the whole of videos. Therefore, our method realizes the cross-modal retrieval focusing on the motion-music relation and improves the retrieval performance by utilizing motion-based video features and the music property-based embedding space alignment. This is the first task which relates human motion to music for the video-music retrieval. In the experiments, we compare our method with the state-of-the-art video-music and human motion-music retrieval methods and verify that our method achieves more than twice the retrieval performance of the conventional methods in Recall@1. Yuki Era, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 3 |
| 2023 | Feature Integration via Back-Projection Ordering Multi-Modal Gaussian Process Latent Variable Model for Rating PredictionabstractIn this paper, we present a method of feature integration via back-projection ordering multi-modal Gaussian process latent variable model (BPomGP) for rating prediction. In the proposed method, to extract features reflecting the users' interest, we use the known ratings assigned to the viewed contents and users' behavior information while viewing the contents which is related to the users' interest. BPomGP has two important approaches. Unlike the training phase, where the above two types of heterogeneous information are available, behavior information is not given in the test phase. Thus, the projection that transforms content features into integrated features is trained by using both content and behavior features. Furthermore, to reflect the unique characteristics of ratings, their ordering, we constrain the distance between integrated features in a latent space based on the difference in the degree of known ratings. Finally, experimental results show that the proposed method can calculate features that are useful for rating prediction. Kyohei Kamikawa, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2023 | Multi-View Variational Recurrent Neural Network for Human Emotion Recognition Using Multi-Modal Biological SignalsabstractIn this paper, the Multi-view Variational Recurrent Neural Network (MvVRNN) is proposed for multi-modal human emotion recognition with gaze and brain activity data while humans view images. For realizing accurate emotion recognition, we focus on the following three characteristics of biological signals: 1) the relationship between implicit and explicit information such as gaze and brain activity data, 2) the temporal changes related to human emotions and 3) the effects of noises that can be included during data acquisition. For treating these characteristics, the proposed MvVRNN has several mechanisms including 1) the integration of multi-modal information including implicit and explicit states of humans, 2) the recurrent module for sequential data and 3) the variational approximation based on the Gaussian distribution. The experimental results show that emotion recognition based on the MvVRNN outperforms several existing methods. Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2023 | Text-Guided Facial Image Manipulation for Wild Images via Manipulation Direction-Based LossabstractThis paper proposes a novel text-guided facial image manipulation approach to improve robustness against the diversity of input images. Conventional text-guided facial image manipulation methods have achieved highly accurate image manipulation by restricting the aspect ratio and composition of the input image. In practical face image manipulation, images provided by the user often include regions other than the face, such as the upper body, and thus conventional methods have limitations in their application. To solve these problems, we tackle a new task of text-guided facial image manipulation with no restrictions on the composition and aspect ratio of the input image. As a base structure, we employ the face swapping method, which is robust to the diversity of input images and allows to swap faces. To achieve text-guided image manipulation based on the face swapping method, we also employ the manipulation direction, which focuses on the direction of changes between the image and its corresponding text that occurs before and after image manipulation. Experimental results show the effectiveness of the proposed method on wild images that include regions other than faces and have different aspect ratios. Yuto Watanabe, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 3 |
| 2023 | Personalized Content Recommender System via Non-verbal Interaction Using Face Mesh and Facial ExpressionabstractMultimedia content recommendation needs to consider users' preferences for each content. Conventional recommender systems consider them with wearable sensors, however, wearing such sensors can lead to a burden on users. In this paper, we construct a recommender system that can explicitly estimate users' preferences without wearable sensors. Specifically, by constructing lightweight but strong machine learning models suitable for our system, the users' interest levels for contents can be estimated from facial images obtained from a widely used webcam. In addition, through the interaction that the user selects displayed contents, our system finds the tendency of personal preferences for recommending contents with high user satisfaction. Our system is available on https://www.lmd-demo.org/2022/start_eng.html. Yuya Moroto, Rintaro Yanagi, Naoki Ogawa, Kyohei Kamikawa, Keigo Sakurai, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ACM Multimedia | 7 |
| 2022 | Human Emotion Recognition Using Multi-Modal Biological Signals Based On Time Lag-Considered Correlation MaximizationabstractA human emotion recognition using multi-modal biological signals based on time lag-considered correlation maximization is presented in this paper. Various multi-modal emotion recognition methods for visual stimuli have been studied and they focus on gaze and brain activity data. The visual stimuli captured by human eyes are sent to the brain by neurotransmitters. Thus, there is a time lag between gaze data, which record where humans gaze at, and brain activity data. However, most of the previous methods only integrate features obtained from each data without considering such a time lag. The proposed method newly introduces the mechanism to consider the time lag into the canonical correlation analysis scheme by assuming that the influence of the visual stimuli on brain activity data follows the Poisson distribution. The contribution of this paper is the construction of a recognition method with considering the time lag for getting truly close to the realization of the occurrence mechanism of human emotions. Experimental results show the effectiveness of considering the time lag between gaze and brain activity data. Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2022 | Variational Bayesian Graph Convolutional Network for Robust Collaborative FilteringabstractThis paper presents a variational Bayesian graph convolutional network for robust collaborative filtering (VBGCF). Conventional graph convolutional network (GCN)-based recommendation models fully trust the observed interaction graph. However, the data used in real-world applications (e.g., video streaming services) are often incomplete and unreliable. To deal with this realistic situation, we newly introduce the probabilistic model based on variational Bayesian inference to GCN-based recommendation. VBGCF can use various generated graphs instead of the observed interaction graph to learn users’ preferences. Therefore, VBGCF is not affected by the incompleteness and the unreliability and can provide robust recommendation. The results of experiments conducted under the realistic situation show the effectiveness of VBGCF. Nozomu Onodera, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2022 | Distributed Label Dequantized Gaussian Process Latent Variable Model for Multi-View Data IntegrationabstractIn this paper, we present a novel method for multi-view data analysis, distributed label dequantized Gaussian process latent variable model (DLDGP). DLDGP can integrate multi-view data and class information into a common latent space. In the previous multiview methods, the dimension of label features transformed from the class information is much smaller than those of the other modalities, which causes a dimensionality-limitation problem in the latent space. DLDGP extends the dimension of the label features by a distributed label dequantization scheme. Additionally, DLDGP calculates correlation between different classes by encoding class information into distributed features. DLDGP can correctly capture the relationship between multi-view data and obtain the latent features with high expression ability. Experimental results show the effectiveness of our method by using the open dataset. Koshi Watanabe, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2022 | Generative Adversarial Network Including Referring Image Segmentation For Text-Guided Image ManipulationabstractThis paper proposes a novel generative adversarial network to improve the performance of image manipulation using natural language descriptions that contain desired attributes. Text-guided image manipulation aims to semantically manipulate an image aligned with the text description while preserving text-irrelevant regions. To achieve this, we newly introduce referring image segmentation into the generative adversarial network for image manipulation. The referring image segmentation aims to generate a segmentation mask that extracts the text-relevant region. By utilizing the feature map of the segmentation mask in the network, the proposed method explicitly distinguishes the text-relevant and irrelevant regions and has the following two contributions. First, our model can pay attention only to the text-relevant region and manipulate the region aligned with the text description. Second, our model can achieve an appropriate balance between the generation of accurate attributes in the text-relevant region and the reconstruction in the text-irrelevant regions. Experimental results show that the proposed method can significantly improve the performance of image manipulation. Yuto Watanabe, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 3 |
| 2022 | Human-Centric Image Retrieval with Gaze-Based Image CaptioningabstractThis paper presents human-centric image retrieval with gaze-based image captioning. Although the development of cross-modal embedding techniques has enabled advanced image retrieval, many methods have focused only on the information obtained from the contents such as image and text. For further extending the image retrieval, it is necessary to construct retrieval techniques that directly reflect human intentions. In this paper, we propose a new retrieval approach via image captioning based on gaze information by focusing on the fact that the gaze information obtained from humans contains semantic information. Specifically, we construct a transformer, connect caption and gaze trace (CGT) model that learns the relationship among images, captioning provided by humans and gaze traces. Our CGT model enables transformer-based learning by dividing the gaze traces into several bounding boxes, and thus, gaze-based image captioning becomes feasible. By using the obtained captioning for cross-modal retrieval, we can achieve human-centric image retrieval. The technical contribution of this paper is transforming the gaze trace into the captioning via the transformer-based encoder. In the experiments, by comparing the cross-modal embedding method, the effectiveness of the proposed method is proved. Yuhu Feng, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2022 | GCN-Based Multi-Modal Multi-Label Attribute Classification in Anime Illustration Using Domain-Specific Semantic FeaturesabstractThis paper presents a multi-modal multi-label attribute classification model in anime illustration based on Graph Convolutional Networks (GCN) using domain-specific semantic features. In animation production, since creators often intentionally highlight the subtle characteristics of the characters and objects when creating anime illustrations, we focus on the task of multi-label attribute classification. To capture the relationship between attributes, we construct a multi-modal GCN model that can adopt semantic features specific to anime illustration. To generate the domain-specific semantic features that represent the semantic contents of anime illustrations, we construct a new captioning framework for anime illustration by combining real images and their style transformation. The contributions of the proposed method are two-folds. 1) More comprehensive relationships between attributes are captured by introducing GCN with semantic features into the multi-label attribute classification task of anime illustrations. 2) More accurate image captioning of anime illustrations can be generated by a trainable model by using only real-world images. To our best knowledge, this is the first work dealing with multi-label attribute classification in anime illustration. The experimental results show the effectiveness of the proposed method by comparing it with some existing methods including the state-of-the-art methods. Ziwen Lan, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2022 | Gaussian Distributed Graph Constrained Multi-Modal Gaussian Process Latent Variable Model for Ordinal Labeled DataabstractThis paper proposes a Gaussian distributed graph constrained multi-modal Gaussian process latent variable model for ordinal labeled data. Rating data that are used in various real-world applications such as product recommendation can represent user preferences, but the difference between adjacent ratings is often uncertain due to the user’s ambiguity. In order to capture the relationships among multi-modal data including rating data, consideration of the uncertainty is necessary. Therefore, by applying the Gaussian distribution to the rating data, we calculate distributed labels that implicitly include the uncertainty, and thus, the Gaussian distributed graph based on their similarities can be constructed. By introducing a constraint calculated based on the graph Laplacian of the Gaussian distributed graph into the objective function of the multi-modal Gaussian process latent variable model, we can achieve an effective latent space that can consider a label correlation while accounting for the uncertainty. This is the contribution of this paper. The effectiveness of the proposed method is verified by experiments using some open datasets. Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2022 | Few-Shot Personalized Saliency Prediction with Similarity of Gaze Tendency Using Object-Based Structural InformationabstractThis paper presents a few-shot personalized saliency prediction method with similarity of gaze tendency using object-based structural information. The personalized saliency maps (PSMs) that represent the individual visual attention can be used for analyzing the heterogeneity among personalized preferences, whereas general saliency maps ignore the individual differences. However, the PSM prediction is a difficult task since the acquisition of eye tracking data, which are needed for obtaining PSMs, gives persons a heavy burden. Then, for realizing PSM prediction with a limited amount of training data, the use of the similarity of gaze tendency between persons can be one effective way. It has been reported that human gazes are related to objects and their relative relationships that are semantic and structural information, and we focus on the integration of PSMs predicted for other persons by using object-based similarities of gaze tendency. The advantage in this paper is that we newly focus on similarities of gaze tendency for visually similar objects for solving the lack of eye tracking data by considering both semantic and structural information, simultaneously. By experimenting with the open dataset, the proposed method outperforms the state-of-the-art methods. Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2022 | Assessment of Image Manipulation Using Natural Language Description: Quantification of Manipulation DirectionabstractWe propose a novel assessment approach to the performance of image manipulation using natural language descriptions in this paper. Text-guided image manipulation aims to modify an input image aligned with the text description and has recently become a hot topic for its usefulness. For the assessment, we focus on the similarity between "the change in text features" and "the change in image features" and define the direction of the changes as Manipulation Direction (MD). By capturing the changes in each different modality, MD can take into account the direction of image manipulation and quantify the extent to which the image is manipulated aligned with the text description. To the best of our knowledge, this is the first metric specialized in text-guided image manipulation. Experimental results show that MD can calculate evaluation scores that are correlated with subjective scores toward the manipulated images. Yuto Watanabe, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 3 |
| 2022 | Visual Sentiment Prediction Using Cross-Way Few-Shot Learning Based on Knowledge DistillationabstractThis paper presents a visual sentiment prediction method using cross-way few-shot learning based on knowledge distillation. Previous studies on visual sentiment prediction methods have focused only on one sentiment dataset although there are several sentiment datasets following different sentiment theories. Originally, sentiments are abstract notions common to humans regardless of the difference between sentiment theories. Thus, the use of knowledge obtained from several sentiment datasets can be the effective way to realize robust visual sentiment prediction. To collaboratively use sentiment datasets, there are the following two concerns: training of the model introducing different sentiment theories and prediction of a newly given sample whose sentiment theory is unknown. Thus, we focus on knowledge distillation, which can improve the generalization ability of several tasks, and effective training becomes feasible. In addition, to deal with the different numbers of sentiment labels in a test phase, we newly introduce the cross-way few-shot learning scheme into knowledge distillation. The main contribution in this paper is to integrate knowledge distillation and the cross-way approach for the visual sentiment prediction, and this is the first work for dealing with the difference of sentiment datasets used in the training and test phases. In the end of this paper, the effectiveness of the proposed method is confirmed through experiments using several open datasets. Yingrui Ye, Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 3 |
| 2022 | Popularity-Aware Graph Social Recommendation for Fully Non-Interaction UsersabstractIn this paper, we address a novel social recommendation for users who have no interactions with items (unobserved users). This task can provide many applications such as recommendations for cold-start users after the first sign-up and targeted advertising, thus, it seems to be extremely meaningful. However, existing social recommendation methods are unsuitable for this task since they assume that all users have interactions with items or cannot recommend more effectively than MostPopular recommendation. Towards this end, we propose Unobserved user-oriented Graph Social Recommendation (UGSR), which learns the preferences of unobserved users and provides richer recommendations than MostPopular recommendation. The popularity-aware graph convolutional network, which is carefully designed for this task, simultaneously considers some user-item interactions, social relations, and item popularity for the effective user and item modeling. Nozomu Onodera, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
MMAsia | 2 |
| 2022 | Affective Embedding Framework with Semantic Representations from Tweets for Zero-Shot Visual Sentiment PredictionabstractThis paper presents a zero-shot visual sentiment prediction method using semantic representation features of texts from tweets as the non-visual auxiliary data. Previous studies show that visual sentiment prediction methods can only predict the sentiment labels that are the same as the labels of the sentiment theory used in the training dataset, which means that they cannot predict the new sentiment label used in different sentiment theories. To solve the problem of predicting new labels, zero-shot learning has been proposed. The previous zero-shot visual sentiment prediction method uses Word2vec features and the adjective-noun pair features to obtain the semantical relationship between images and sentiment words to predict unseen sentiments. However, many adjective-noun pairs are not related to sentiments, which makes it difficult to compensate for an affective gap between low-level visual features and high-level sentiment semantics. Thus, to better compensate for the affective gap, it is considered to introduce the new non-visual auxiliary data. As people tend to share their feelings with both images and texts on social networking services, the texts from tweets are effective as the side information of the images in visual sentiment prediction. Thus, we introduce the semantic representations from tweets as the new non-visual auxiliary data to construct an affective embedding space, which makes a more effective zero-shot visual sentiment prediction model. Moreover, we propose a cross-dataset zero-shot task for visual sentiment prediction, which is more consistent with the real situation that the testing and training images may be in different domains. The contributions in this paper are to combine several semantic representation features for zero-shot visual sentiment prediction and the proposal of the cross-dataset zero-shot task for visual sentiment prediction. The experiments on several open datasets show the effectiveness of the proposed method. Yingrui Ye, Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
MMAsia | 3 |
| 2022 | Summarizing Data Structures with Gaussian Process and Robust Neighborhood Preservation
Koshi Watanabe, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ECML/PKDD (5) | 2 |
| 2022 | Chain centre loss: A psychology inspired loss function for image sentiment analysis
Yun Liang 0014, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
Neurocomputing | 2 |
| 2021 | Cross-Domain Semi-Supervised Deep Metric Learning for Image Sentiment AnalysisabstractThis paper presents a novel method on image sentiment analysis called cross-domain semi-supervised deep metric learning (CDSS-DML). The proposed method has two contributions. Firstly, since previous researches on image sentiment analysis suffer from the limit of a small amount of well-labeled data, which occurs a decrease in accuracy of classification, CDSS-DML breaks through the limit by training with unlabeled data based on a teacher-student model. Secondly, the proposed method overcomes the difficulty of distribution shift between well-labeled and unlabeled data by jointing three losses. Especially, the proposed method constructs an effective latent space with the joint loss considering the inter-class and the intra-class correlations for image sentiments. From experimental results, the performance improvement with CDSS-DML is confirmed. Yun Liang 0014, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2021 | Classification of Expert-Novice Level Using Eye Tracking And Motion Data via Conditional Multimodal Variational AutoencoderabstractSensor data from wearable devices have been utilized to analyze differences between experts and novices. Previous studies attempted to classify the expert-novice level from sensor data based on supervised learning methods. However, these approaches need to collect enough training data covering various novices’ sensor patterns. In this paper, we propose a semi-supervised anomaly detection approach that requires only sensor data of experts for training and identifies those of novices as anomalies. Our proposed anomaly detection model named conditional multimodal variational autoencoder (CMVAE) has the following two technical contributions: (i) considering action information of persons and (ii) utilizing multimodal sensor data, i.e., eye tracking data and motion data in this case. The proposed method is evaluated on sensor data measured when expert and novice soccer players were shooting, dribbling, and doing soccer ball juggling. Experimental results show that CMVAE can more accurately classify the expert-novice level than previous supervised learning methods and anomaly detection methods using other VAEs. Yusuke Akamatsu, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2021 | Estimation of Visual Features of Viewed Image From Individual and Shared Brain Information Based on FMRI Data Using Probabilistic Generative ModelabstractThis paper presents a method for estimation of visual features based on brain responses measured when subjects view images. The pro-posed method estimates visual features of viewed images by using both individual and shared brain information from functional magnetic resonance imaging (fMRI) data when subjects view images. To extract an effective latent space shared by multiple subjects from high dimensional fMRI data, a probabilistic generative model that can provide a prior distribution to the space is introduced into the proposed method. Also, the extraction of a robust feature space with respect to noise for the individual information becomes feasible via the proposed probabilistic generative model. This is the first contribution of our method. Furthermore, the proposed method constructs a decoder transforming brain information into visual features based on collaborative use of both estimated spaces for individual and shared brain information. This is the second contribution of our method. Experimental results show that the proposed method improves the estimation accuracy of the visual features of viewed images. Takaaki Higashi, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2021 | Feature Integration via Semi-Supervised Ordinally Multi-Modal Gaussian Process Latent Variable ModelabstractThis paper presents a method of feature integration via semi-supervised ordinally multi-modal Gaussian process latent variable model (Semi-OMGP). The proposed method transforms multimodal features into common latent variables suitable for users' interest level estimation. For dealing with the multi-modal features, the proposed method newly derives Semi-OMGP. Semi-OMGP has two contributions. First, Semi-OMGP is suitable for integration between heterogeneous modalities with different distributions by assuming that the similarity matrices of these modalities as observations are generated from latent variables. Second, Semi-OMGP can efficiently use label information by introducing an operator considering the ordinal grade into the prior distribution of latent variables when obtained label information is partially given. Semi-OMGP can simultaneously realize the above contributions, and successful multi-modal feature integration becomes feasible. Experimental results show the effectiveness of the proposed method. Kyohei Kamikawa, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2021 | Multi-Modal Label Dequantized Gaussian Process Latent Variable Model for Ordinal Label EstimationabstractThis paper presents multi-modal label dequantized Gaussian process latent variable model (mLDGP) for ordinal label estimation. mLDGP is constructed based on a probabilistic generative model via Gaussian process and realizes accurate calculation of common latent space from multi-view features including low-dimensional ordinal label features. Conventional methods have a problem that the dimension of the common latent space was limited to that of the label feature, and an enough expressive latent space cannot be obtained. mLDGP, which is constructed by introducing our novel label de-quantization mechanism into the objective function of multi-modal Gaussian process latent variable model (GPLVM), can increase the dimension of label features. Then mLDGP can calculate the effective latent space. Furthermore, mLDGP can estimate projection transforming unknown features of test samples into the common latent space, which was a problem of the conventional GPLVMs. From experimental results obtained by applying our method to the product rating estimation on the online shopping website, it is confirmed that accuracy improvement using mLDGP becomes feasible compared to various methods. Masanao Matsumoto, Keisuke Maeda, Naoki Saito 0006, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2021 | Deep Metric Network Via Heterogeneous Semantics for Image Sentiment AnalysisabstractThis paper presents a novel method for image sentiment analysis called a deep metric network via heterogeneous semantics (DMNHS). The contribution of the proposed method is introduction of the image captioning into image sentiment analysis to reflect a global impression that cannot be represented by classical visual features extracted from images. In order to consider a sentiment correlation between visual and captioning features, the proposed method newly designs a network to integrate these heterogeneous semantics features (HS features). Furthermore, with consideration of relations among sentiments based on the HS features, the proposed method constructs a sentiment latent space by introducing the center loss concerning relationships between different sentiments and enables the classification of image sentiments. From experimental results, the performance improvement via DMN-HS is confirmed. Yun Liang 0014, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2021 | Segmentation-Aware Text-Guided Image ManipulationabstractWe propose a novel approach that improves text-guided image manipulation performance in this paper. Text-guided image manipulation aims at modifying some parts of an input image in accordance with the user’s text description by semantically associating the regions of the image with the text description. We tackle the conventional methods’ problem of modifying undesired parts caused by differences in representation ability between text descriptions and images. Humans tend to pay attention primarily to objects corresponding to the foreground of images, and text descriptions by humans mostly represent the foreground. Therefore, it is necessary to introduce not only a foreground-aware bias based on text descriptions but also a background-aware bias that the text descriptions do not represent. We introduce an image segmentation network into the generative adversarial network for image manipulation to solve the above problem. Comparative experiments with three state-of-the-art methods show the effectiveness of our method quantitatively and qualitatively. Tomoki Haruyama, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 3 |
| 2021 | Cross-Domain Recommendation Method Based On Multi-Layer Graph Analysis With Visual InformationabstractIn this paper, a cross-domain recommendation (CDR) method based on multi-layer graph analysis with visual information is presented. Although previous graph-based CDR methods have effectively used usersʼ ratings of products and their purchase histories, they lack essential information such as product images that can have an impact on usersʼ decision of whether to purchase products. Then the proposed method newly introduces visual features obtained from product images into the graph-based CDR. For dealing with visual features in multiple domains, we focus on both intra-domain and interdomain relationships through visual features. Specifically, to obtain effective embedding features from users and items in a domain, the proposed method newly introduces visual features into an optimization process of the latest graph neural network that considers only user-item interactions. Consideration of the visual similarity of items within a domain becomes feasible, and embedding features with high representation ability can be estimated. Furthermore, to consider visual information between domains, we construct multi-layer graphs for each domain and introduce visual features into the training of a feature transformer across these graphs. Therefore, consideration of both intra-domain and inter-domain relationships through visual features contributes to the performance improvement. To our best knowledge, this is the first trial to introduce visual features into multi-layer graph-based CDR. The effectiveness of our method is demonstrated by comparing several state-of-the-art methods. Taisei Hirakawa, Keisuke Maeda, Takahiro Ogawa 0001, Satoshi Asamizu, Miki Haseyama |
ICIP | 2 |
| 2021 | Time-Lag Aware Multi-Modal Variational Autoencoder Using Baseball Videos And Tweets For Prediction Of Important ScenesabstractA novel method based on time-lag aware multi-modal variational autoencoder for prediction of important scenes (TI-MVAE-PIS) using baseball videos and tweets posted on Twitter is presented in this paper. This paper has the following two technical contributions. First, to effectively use heterogeneous data for the prediction of important scenes, we transform textual, visual and audio features obtained from tweets and videos to the latent features. Then TI-MVAE-PIS can flexibly express the relationships between them in the constructed latent space. Second, since there are time-lags between tweets and the corresponding multiple previous events, Tl-MVAE-PIS considers such time-lags in their relationship estimation for successfully deriving their latent features. Therefore, these two contributions enable accurate important scene prediction. Results of experiments using actual baseball videos and their corresponding tweets show the effectiveness of TI-MVAE-PIS. Kaito Hirasawa, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2021 | Interest Level Estimation via Multi-Modal Gaussian Process Latent Variable Factorization
Kyohei Kamikawa, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2021 | Few-Shot Personalized Saliency Prediction using Person Similarity based on Collaborative Multi-Output Gaussian Process RegressionabstractA few-shot personalized saliency prediction method using person similarity based on collaborative multi-output Gaussian process regression is presented in this paper. Contrary to prediction of general saliency maps, that of personalized saliency maps (PSMs), which is a focus of attention owing to its heterogeneity among individuals, is a challenging problem since the amount of training gaze data is limited due to the burden on new persons. Thus, the proposed method focuses on the similarity of gaze tendency between persons. In the proposed method, collaborative Gaussian process regression (CoMOGP) is adopted for PSM prediction. CoMOGP enables to represent similarity of gaze tendency between the target person and other persons as weights, and then consider the similarity for each image by using visual features obtained from images as inputs. The contributions of the few-shot PSM prediction based on CoMOGP are two-folds. 1) CoMOGP, which is one of probabilistic methods, can avoid the overfitting to small amount of training data. 2) Similarity for each image can be considered by using visual features as inputs. In the experiment using the open dataset, the proposed method outperforms comparative methods including the state-of-the-art method. Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2021 | Correlation-Aware Attention Branch Network Using Multi-Modal Data For Deterioration Level Estimation Of InfrastructuresabstractThis paper presents a correlation-aware attention branch network (CorABN) using multi-modal data for deterioration level estimation of infrastructures. CorABN can collaboratively use visual features from distress images and text features from text data recorded at the inspection to improve the estimation performance of deterioration levels. Specifically, by maximizing correlation between the visual and text features that provide useful information for the deterioration level estimation, a correlation-aware attention map can be generated. Besides, the text features are also utilized in the final estimation along with visual features improved by the attention mechanism to achieve higher estimation performance. With the losses based on both the estimation performance and the correlation, CorABN can train the entire model in an end-to-end manner. Experiments using distress images and their corresponding text data show the effectiveness of the proposed method. Naoki Ogawa, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2021 | Deterioration level estimation via neural network maximizing category-based ordinally supervised multi-view canonical correlation
Keisuke Maeda, Sho Takahashi, Takahiro Ogawa 0001, Miki Haseyama |
Multim. Tools Appl. | 1 |
| 2020 | Important Scene Detection Of Baseball Videos Via Time-Lag Aware Deep Multiset Canonical Correlation MaximizationabstractThis paper presents a new important scene detection method of baseball videos based on correlation maximization between heterogeneous modalities via time-lag aware deep multiset canonical correlation analysis (Tl-dMCCA). The technical contributions of this paper are twofold. First, textual, visual and audio features calculated from tweets and videos are adopted as multi-view time series features. Since Tl-dMCCA which utilizes these features includes the unsupervised embedding scheme via deep networks, the proposed method can flexibly express the relationship between heterogeneous features. Second, since there is the time-lag between posted tweets and the corresponding multiple previous events, Tl-dMCCA considers the time-lag relationships between them. Specifically, we newly introduce the representation of such time-lags into the derivation of their covariance matrices. By considering time-lags via Tl-dMCCA, the proposed method correctly detects important scenes. Kaito Hirasawa, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2020 | Feature Integration Via Geometrical Supervised Multi-View Multi-Label Canonical Correlation For Incomplete Label AssignmentabstractThis paper presents feature integration via geometrical supervised multi-view multi-label canonical correlation analysis (GSM2CCA) for incomplete label assignment. The problem of incomplete labels is frequently encountered in the multi-label classification problem where the training labels are obtained via crowd-sourcing. In such a situation, consideration of only the label correlation, which is the basic approach, is not suitable for improvement of representation ability of features. For dealing with the incomplete label assignment, GSM2CCA constructs effective feature embedding space providing the discriminant ability by introducing both the multi-label correlation and feature similarity of the original feature space into its objective function. Since novel integrated features with high discriminant ability can be calculated by our GSM2CCA, performance improvement of multi-label classification with the incomplete label assignment is realized. The main contribution of this paper is the realization of the effective feature integration via the adoption of the combination use of label similarity and locality preserving projection of heterogeneous features for solving the problem of the incomplete label assignment. The effectiveness of GSM2CCA by applying GSM2CCA-based feature integration to heterogeneous features calculated from various convolutional neural network models is verified via experimental results. Keisuke Maeda, Sho Takahashi, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2020 | Human-centered image classification via a neural network considering visual and biological features
Kazaha Horii, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
Multim. Tools Appl. | 2 |
| 2019 | Multi-feature Fusion Based on Supervised Multi-view Multi-label Canonical Correlation ProjectionabstractThis paper presents multi-feature fusion based on supervised multi-view multi-label canonical correlation projection (sM2CP). The proposed method applies sM2CP-based feature fusion to multiple features obtained from various convolutional neural networks (CNNs) whose characteristics are different. Since new fused features with high representation ability can be obtained, performance improvement of multi-label classification is realized. Specifically, in order to tackle the multi-label problem, sM2CP introduces a label similarity information of label vectors into the objective function of supervised multi-view canonical correlation analysis. Thus, sM2CP can deal with complex label information such as multi-label annotation. The main contribution of this paper is the realization of feature fusion of multiple CNN features for the multi-label problem by introducing multi-label similarity information into the canonical correlation analysis-based feature fusion approach. Experimental results show the effectiveness of sM2CP, which enables effective fusion of multiple CNN features. Keisuke Maeda, Sho Takahashi, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2019 | Neural Network Maximizing Ordinally Supervised Multi-View Canonical Correlation for Deterioration Level EstimationabstractThis paper presents a neural network maximizing ordinally supervised multi-view canonical correlation for deterioration level estimation. The contributions of this paper are twofold. First, in order to calculate features representing deterioration levels on transmission towers, which is one of the infrastructures, a novel neural network handling multi-modal features is constructed from a small amount of training data. Specifically, in our method, effective transformation to features with high discriminant ability without using many hidden layers is realized by setting projection matrices maximizing correlation between multiple features into hidden layer's weights. Second, since there exists ordinal scale in deterioration levels, the proposed method newly derives ordinally supervised multi-view canonical correlation analysis (OsMVCCA). OsMVCCA enables estimation of the effective projection considering not only label information but also their ordinal scales. Experimental results show that the proposed method realizes accurate deterioration level estimation. Keisuke Maeda, Sho Takahashi, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2019 | Estimation of Emotion Labels via Tensor-Based Spatiotemporal Visual Attention AnalysisabstractThis paper presents emotion label estimation via tensor-based spatiotemporal visual attention analysis. It has been reported in the fields of psychology and neuroscience that human emotions are related to two elements, their visual attention change and objects included in a target image. Therefore, the proposed method focuses on the spatiotemporal change of visual attention of human gazing at objects in the target image and constructs two neural networks which enable the emotion label estimation considering both of the above two elements. Specifically, the proposed method newly constructs a fourth-order tensor, gaze and image tensor (GIT) whose modes correspond to the width, the height and the color channel of the target image and the time axis of visual attention which is used for representing the time change. Then the first network, which consists of general tensor discriminant analysis (GTDA) and extreme learning machine (ELM), estimates the emotion label from the fourth-order GIT with concerning their visual attention change. Furthermore, the second network, which consists of pre-trained convolutoinal neural network-based feature extraction, GTDA and ELM, enables the estimation from the second-order GIT including visual features obtained from objects focused at each time. Finally, the proposed method estimates emotion labels based on decision fusion of the outputs from the two networks. Experimental results show the effectiveness of the proposed method. Yuya Moroto, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2018 | Sectional Review Recommendations based on Learner's Comprehension in Video-based Learning
Yusuke Hayashi, Keisuke Maeda, Toshio Honda, Tsukasa Hirashima |
ICCE | 2 |
| 2018 | Structure-mapping Support for Learning by Analogy with Kit-Build Concept Map
Yusuke Hayashi, Kan Yoshida, Keisuke Maeda, Akira Yamanaka, Tsukasa Hirashima |
ICCE | 3 |
| 2018 | A Human-Centered Neural Network Model with Discriminative Locality Preserving Canonical Correlation Analysis for Image ClassificationabstractThis paper presents a human-centered neural network model with discriminative locality preserving canonical correlation analysis (DLPCCA) for image classification. Although construction of multiple hidden layers adopted in recent deep learning methods is effective for extracting semantic features, a large amount of training images is required. In order to extract effective features for image classification successfully from a small amount of training images, the proposed method transforms visual features by using biological information obtained from image viewers as auxiliary information. The proposed method consists of two hidden layers. By constructing the first hidden layer, which can maximize canonical correlation between visual features and features based on biological information, the effective feature transformation can be realized. Specifically, the proposed method uses DLPCCA, which considers label information and preserves local structures. The second hidden layer constructed based on Extreme Learning Machine (ELM) enables classification. Consequently, since the first hidden layer performs the effective feature transformation, the proposed neural network model realizes accurate image classification from a quite small amount of training images. Kazaha Horii, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2018 | Distress classification of class-imbalanced inspection data via correlation-maximizing weighted extreme learning machine
Keisuke Maeda, Sho Takahashi, Takahiro Ogawa 0001, Miki Haseyama |
Adv. Eng. Informatics | 1 |
| 2017 | Automatic martian dust storm detection via decision level fusion basedondeep extreme learning machineabstractThis paper presents an automatic Martian dust storm detection via decision level fusion (DLF) based on deep extreme learning machine (DELM). Since Martian images are taken in multi-wavelength bands, DLF techniques which output a final classification result by integrating multiple classification results are necessary. Furthermore, since the number of Martian images taken by satellites is different for each region, the number of the classification results to be integrated is different. Thus, we present a new DLF framework based on confidence values of the classification results. Specifically, we generate multiple extreme learning machines with kernel classifiers to obtain their classification results. Moreover, we monitor the classification results as confidence values and select the same number of the classification results with high confidence for each region. Finally, these selected results can be integrated by using a DLF based on DELM, which is a multilayered ELM. This integration framework is the biggest contribution of our method. Experimental results show the effectiveness of the DLF based on DELM. Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2017 | Automatic estimation of deterioration level on transmission towers via deep extreme learning machine based on local receptive fieldabstractThis paper presents an automatic estimation method of deterioration levels on transmission towers via Deep Extreme Learning Machine based on Local Receptive Field (DELM-LRF). Although Convolutional Neural Network (CNN) requires a large number of training images, it is difficult to prepare a sufficient number of training images of transmission towers. Thus, we generate a novel estimation method which enables training from a small number of training images. Specifically, we automatically extract image features based on Local Receptive Field (LRF) which combines convolution and pooling without using hand-craft features and estimate deterioration levels via Deep Extreme Learning Machine (DELM), which is a part of efficient deep learning methods. The derivation of DELM-LRF is the biggest contribution of this paper, and it can be trained from less training images compared to CNN. Experimental results show the effectiveness of DELM-LRF for the estimation of deterioration levels on transmission towers. Consequently, the proposed method makes it possible to approach challenging tasks with high expertise having difficulty in preparing enough images. Keisuke Maeda, Sho Takahashi, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2016 | Learning Task Generation from a Series of Propositions of a Learning Topic: Kit-Building Task of Concept Map and Multiple Choice Task of Fill-in-the-blank QuestionsabstractKit-building task of a concept map is a promising exercise for strengthening and assessing learner's comprehension for a topic that a learner already has learnt. In order to investigate the value of the kit-building task of a concept map, we are comparing it with multiple-choice task of fill-in-the-blank questions. As a step to realize the comparison, in this paper, it is described an implementation to generate learning task from a series of propositions, that is, (1) kit-building task of concept map and (2) multiple choice task of fill-in-the-blank questions. Takuya Kitamura, Akira Yamanaka, Keisuke Maeda, Yusuke Hayashi, Tsukasa Hirashima |
ICCE | 3 |
| 2016 | Comparison between Kit-Building Task of Concept Map and Multiple Choice Task of Fill-in-the-blank Question Generated from the Same Series of Propositions
Takuya Kitamura, Akira Yamanaka, Keisuke Maeda, Yusuke Hayashi, Tsukasa Hirashima |
ICCE | 3 |
| 2015 | Forecasting urban dynamics with mobility logs by bilinear Poisson regressionabstractUnderstanding people flow in a city (urban dynamics) is of great importance in urban planning, emergency management, and commercial activity. With the spread of smart devices, many studies on urban dynamics modeling with mobility logs have been conducted. It is predictive analysis, not analysis of the past, that enables various applications contributing to a more prosperous society. To deal with the non-linear effects on urban dynamics from external factors, such as day of the week, national holiday, or weather, we propose a low-rank bilinear Poisson regression model, for a novel and flexible representation of urban dynamics predictive analysis. The results obtained from an experiment with one year's worth of mobility records suggest the high prediction accuracy of the proposed model. We also introduce the following applications: regional event detection via irregularities, visualization of urban dynamics corresponding to urban demographics, and extraction of urban demographics of unknown point of interests. Masamichi Shimosaka, Keisuke Maeda, Takeshi Tsukiji, Kota Tsubouchi |
UbiComp | 2 |
| 2015 | Automatic detection of martian dust storms from heterogeneous data based on decision level fusionabstractThis paper presents automatic detection of Martian dust storms from heterogeneous data (raw data, reflectance data and background subtraction data of the reflectance data) based on decision level fusion. Specifically, the proposed method first extracts image features from these data and selects optimal features for dust storm detection based on the minimal-Redundancy-Maximal-Relevance algorithm. Second, the selected image features are used to train the Support Vector Machine classifier that is constructed on each data. Furthermore, as a main contribution of this paper, the proposed method combines the multiple detection results obtained from the heterogeneous data based on decision level fusion with considering each classifier's detection performance to obtain accurate final detection results. Consequently, the proposed method realizes automatic and accurate detection of Martian dust storms. Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |