EDBT 2026 Demo / reviewers in the wild / expert
Rintaro Yanagi
dblp:232/2765
· DBLP profile ↗
15ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0003-0110-7208ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion Noise Optimization for Synthetic VLM TrainingabstractRecent advances in image generation models have enabled the production of high-quality images, making synthetic images a promising alternative to real images for dataset construction. However, a critical challenge remains in that the performance of Vision–Language Models (VLMs) tends to degrade as the proportion of synthetic images in a dataset increases in conventional approaches. To alleviate the challenge, we introduce a plug-and-play dataset construction framework that enhances text-to-image diffusion models by optimizing their initial noise. Our method treats the initial noise as a learnable parameter and iteratively updates it to maximize text–image alignment based on multiple embedding models without retraining the generator. Since the initial noise plays a crucial role in determining the quality of the synthetic image, its optimization enables the search for initial conditions that yield semantically faithful and realistic images. By improving FID and text–image alignment compared to conventional latent diffusion model (LDM)-based methods, our approach produces synthetic images better suited for training. When CLIP models were trained on such images, it achieved up to +5.09% higher Average R@1 in zero-shot retrieval, +2.88% higher Average top-1 accuracy in zero-shot classification, and +5.05% higher performance in linear-probing. These results demonstrate that initial noise optimization is an effective and scalable strategy for enabling robust VLM training with synthetic images. Ren Ohkubo, Rintaro Yanagi, Hirokatsu Kataoka, Yutaka Satoh |
WACV | 2 |
| 2025 | Approximate Domain Unlearning for Vision-Language ModelsabstractPre-trained Vision-Language Models (VLMs) exhibit strong generalization capabilities, enabling them to recognize a wide range of objects across diverse domains without additional training. However, they often retain irrelevant information beyond the requirements of specific target downstream tasks, raising concerns about computational efficiency and potential information leakage. This has motivated growing interest in approximate unlearning, which aims to selectively remove unnecessary knowledge while preserving overall model performance. Existing approaches to approximate unlearning have primarily focused on {\em class unlearning}, where a VLM is retrained to fail to recognize specified object classes while maintaining accuracy for others. However, merely forgetting object classes is often insufficient in practical applications. For instance, an autonomous driving system should accurately recognize {\em real} cars, while avoiding misrecognition of {\em illustrated} cars depicted in roadside advertisements as {\em real} cars, which could be hazardous. In this paper, we introduce {\em Approximate Domain Unlearning (ADU)}, a novel problem setting that requires reducing recognition accuracy for images from specified domains (e.g., {\em illustration}) while preserving accuracy for other domains (e.g., {\em real}). ADU presents new technical challenges: due to the strong domain generalization capability of pre-trained VLMs, domain distributions are highly entangled in the feature space, making naive approaches based on penalizing target domains ineffective. To tackle this limitation, we propose a novel approach that explicitly disentangles domain distributions and adaptively captures instance-specific domain information. Extensive experiments on four multi-domain benchmark datasets demonstrate that our approach significantly outperforms strong baselines built upon state-of-the-art VLM tuning techniques, paving the way for practical and fine-grained unlearning in VLMs. Code : https://kodaikawamura.github.io/Domain_Unlearning/. Kodai Kawamura, Yuta Goto, Rintaro Yanagi, Hirokatsu Kataoka, Go Irie |
NeurIPS | 3 |
| 2024 | Learning 3D Point Cloud Registration as a Single Optimization Problem
Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Shusaku Sone, Yoshitaka Ushiku |
ACCV (9) | 1 |
| 2024 | Zero-Shot Composed Image Retrieval Considering Query-Target Relationship Leveraging Masked Image-Text PairsabstractThis paper proposes a novel zero-shot composed image retrieval (CIR) method considering the query-target relationship by masked image-text pairs. The objective of CIR is to retrieve the target image using a query image and a query text. Existing methods use a textual inversion network to convert the query image into a pseudo word to compose the image and text and use a pre-trained visual-language model to realize the retrieval. However, they do not consider the query-target relationship to train the textual inversion network to acquire information for retrieval. In this paper, we propose a novel zero-shot CIR method that is trained end-to-end using masked image-text pairs. By exploiting the abundant image-text pairs that are convenient to obtain with a masking strategy for learning the query-target relationship, it is expected that accurate zero-shot CIR using a retrieval-focused textual inversion network can be realized. Experimental results show the effectiveness of the proposed method. Huaying Zhang, Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2024 | DQG: Database Question Generation for Exact Text-based Image Retrieval
Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ACM Multimedia | 1 |
| 2023 | Personalized Content Recommender System via Non-verbal Interaction Using Face Mesh and Facial ExpressionabstractMultimedia content recommendation needs to consider users' preferences for each content. Conventional recommender systems consider them with wearable sensors, however, wearing such sensors can lead to a burden on users. In this paper, we construct a recommender system that can explicitly estimate users' preferences without wearable sensors. Specifically, by constructing lightweight but strong machine learning models suitable for our system, the users' interest levels for contents can be estimated from facial images obtained from a widely used webcam. In addition, through the interaction that the user selects displayed contents, our system finds the tendency of personal preferences for recommending contents with high user satisfaction. Our system is available on https://www.lmd-demo.org/2022/start_eng.html. Yuya Moroto, Rintaro Yanagi, Naoki Ogawa, Kyohei Kamikawa, Keigo Sakurai, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ACM Multimedia | 2 |
| 2023 | Reference-based Dense Pose Estimation via Partial 3D Point Cloud MatchingabstractInteracting with real-world objects is one of the fundamental tasks in multimedia. Despite its importance, existing object pose estimation targets only rigid objects. This demonstration proposes a novel application for non-rigid object pose estimation. Inspired by human dense pose estimation, we represent a pose of a non-rigid object as an indexed point cloud, where each index corresponds to that in a template. The correspondence is identified by a machine-learning-based 3D point cloud matching. Finding correspondence to the template point cloud enables a dense pose estimation with no object-specific learning processes. In the demonstration, we visualize the correspondence of points in observed depth images and the template. We also provide a demonstration of template point cloud reconstruction. Through these systems, onsite visitors can test our system with objects brought by themselves and have an experience with a state-of-the-art 3D point cloud matching method as well as this novel task. Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Yoshitaka Ushiku |
ACM Multimedia | 1 |
| 2022 | Rubber Material Retrieval System using Electron Microscope Images for Rubber Material DevelopmentabstractFor developing valuable rubber materials, machine learning-based computer-aided analysis systems have been attracting a lot of attention. However, these systems mainly focus on analyzing the table and textual data, and the electron microscope images including the rich material information have not been enough analyzed. By effectively using these electron microscope images, further support for the material discovery is realized. In this paper, we present a material information retrieval system via electron microscope image space. Our system aims to support visually and comprehensively grasping the relationships between various rubber materials and those properties. By effectively using the electron microscope image space for material information retrieval, it is expected that the advances in material development are further accelerated. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
MMAsia | 1 |
| 2022 | Interactive Re-ranking via Object Entropy-Guided Question Answering for Cross-Modal Image RetrievalabstractCross-modal image-retrieval methods retrieve desired images from a query text by learning relationships between texts and images. Such a retrieval approach is one of the most effective ways of achieving the easiness of query preparation. Recent cross-modal image-retrieval methods are convenient and accurate when users input a query text that can be used to uniquely identify the desired image. However, in reality, users frequently input ambiguous query texts, and these ambiguous queries make it difficult to obtain desired images. To overcome these difficulties, in this study, we propose a novel interactive cross-modal image-retrieval method based on question answering. The proposed method analyzes candidate images and asks users questions to obtain information that can narrow down retrieval candidates. By only answering questions generated by the proposed method, users can reach their desired images, even when using an ambiguous query text. Experimental results show the proposed method’s effectiveness. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | IR Questioner: QA-based Interactive Retrieval SystemabstractImage retrieval from a given text query (text-to-image retrieval) is one of the most essential systems, and it is effectively utilized for databases (DBs) on the Web. To make them more versatile and familiar, a retrieval system that is adaptive even for personal DBs such as images in smartphones and lifelogging devices should be considered. In this paper, we present a novel text-to-image retrieval system that is specialized for personal DBs. With the cross-modal scheme and the question-answering scheme, the developed system enables users to obtain the desired image effectively even from personal DBs. Our demo is available at https://sites.google.com/view/ir-questioner/. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICMR | 1 |
| 2021 | Database-adaptive Re-ranking for Enhancing Cross-modal Image RetrievalabstractWe propose an approach that enhances arbitrary existing cross-modal image retrieval performance. Most of the cross-modal image retrieval methods mainly focus on direct computation of similarities between a text query and candidate images in an accurate way. However, their retrieval performance is affected by the ambiguity of text queries and the bias of target databases (DBs). Dealing with ambiguous text queries and DBs with bias will lead to accurate cross-modal image retrieval in real-world applications. A DB-adaptive re-ranking method using modality-driven spaces, which can extend arbitrary cross-modal image retrieval methods for enhancing their performance, is proposed in this paper. The proposed method includes two approaches: "DB-adaptive re-ranking'' and "modality-driven clue information extraction''. Our method estimates clue information that can effectively clarify the desired image from the whole set of a target DB and then receives user's feedback for the estimated information. Furthermore, our method extracts more detailed information of a query text and a target DB by focusing on modality-driven spaces, and it enables more accurate re-ranking. Our method allows users to reach their desired single image by just answering questions. Experimental results using MSCOCO, Visual Genome and newly introduced datasets including images with a particular object show that the proposed method can enhance the performance of state-of-the-art cross-modal image retrieval methods. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ACM Multimedia | 1 |
| 2020 | Image Retrieval With Lingual And Visual Paraphrasing Via Generative ModelsabstractA new approach that improves text-based image retrieval (hereinafter referred to as TBIR) performance is proposed in this paper. TBIR methods aim to retrieve a desired image related to a query text. Especially, recent TBIR methods allow us to retrieve images considering word relationships by using a sentence as a query. In these TBIR methods, it is necessary to uniquely identify a desired image from similar images using a single query sentence. However, the diverse expressive styles for a query sentence make it difficult to uniquely identify a desired image. In this paper, we propose a novel TBIR method with paraphrasing on multiple representation spaces. Specifically, by paraphrasing a query sentence on lingual and visual representation spaces, the proposed method can retrieve a desired image from various perspectives and then it can uniquely identify a desired image from similar images. Comprehensive experimental results show the effectiveness of the proposed method. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2020 | Interactive re-ranking for cross-modal retrieval based on object-wise question answeringabstractCross-modal retrieval methods retrieve desired images from a query text by learning relationships between texts and images. This retrieval approach is one of the most effective ways in the easiness of query preparation. Recent cross-modal retrieval is convenient and accurate when users input a query text that can uniquely identify the desired image. Meanwhile, users frequently input ambiguous query texts, and these ambiguous queries make it difficult to obtain the desired images. To alleviate these difficulties, in this paper, we propose a novel interactive cross-modal retrieval method based on question answering (QA) with users. The proposed method analyses candidate images and asks users about information that can narrow retrieval candidates effectively. By only answering the questions generated by the proposed method, users can reach their desired images even from an ambiguous query text. Experimental results show the effectiveness of the proposed method. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
MMAsia | 1 |
| 2019 | Scene Retrieval for Video Summarization Based on Text-to-Image ganabstractWe present a new scene retrieval method based on text-to-image Generative Adversarial Network (GAN) and its application to query-based video summarization. Text-to-image GAN is a deep learning method that can generate images from their corresponding sentences. In this paper, we reveal a characteristic that deep learning-based visual features extracted from images generated by text-to-image GAN include semantic information sufficiently. By utilizing the generated images as queries, the proposed method achieves higher scene retrieval performance than those of the state-of-the-art methods. In addition, we introduce a novel architecture that can consider order relationship of the input sentences to our method for realizing a target video summarization. Specifically, the proposed method generates multiple images thorough text-to-image GAN from multiple sentences summarizing target videos. Their summarized video can be obtained by performing the retrieval of corresponding scenes from the target videos according to the generated images with considering the order relationship. Experimental results show the effectiveness of the proposed method in the retrieval and summarization performance. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2019 | Scene Retrieval from Multiple Resolution Generated Images Based on Text-to-Image GANabstractText-to-image Generative Adversarial Network (GAN) is a deep learning model that generates an image from an input sentence. It is expressly attracting attentions because of its applicability of the generated images. However, many existing studies have still focused on generation of high-quality images, and there are few studies focusing on application of the generated images since text-to-image GANs still cannot produce visually pleasing images in the complicated tasks. In this paper, we apply a text-to-image GAN as a generator of query images for a scene retrieval task to show availability of the visually non-pleasant images. The proposed method utilizes a low-resolution generated image that focuses on a sentence and a high-resolution generated image that focuses on each word of the sentence to retrieve a desired scene. With this mechanism, the proposed method realizes a high-accuracy scene retrieval from a sentence input. Experimental results show the effectiveness of our method. Rintaro Yanagi, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ISCAS | 1 |