Ichiro Ide

dblp:33/339 · DBLP profile ↗
← Back
12ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0003-3942-9296ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 6Information Retrieval & Web Search · 5 (1 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2025 Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
abstract
Recent advances in video captioning are driven by large-scale pretrained models, which follow the standard “pre-training followed by fine-tuning” paradigm, where the full model is fine-tuned for downstream tasks. Although effective, this approach becomes computationally prohibitive as the model size increases. The Parameter-Efficient Fine-Tuning (PEFT) approach offers a promising alternative, but primarily focuses on the language components of Multimodal Large Language Models (MLLMs). Despite recent progress, PEFT remains underexplored in multimodal tasks and lacks sufficient understanding of visual information during fine-tuning the model. To bridge this gap, we propose Query-Adapter (Q-Adapter), a lightweight visual adapter module designed to enhance MLLMs by enabling efficient fine-tuning for the video captioning task. Q-Adapter introduces learnable query tokens and a gating layer into Vision Encoder, enabling effective extraction of sparse, caption-relevant features without relying on external textual supervision. We evaluate Q-Adapter on two well-known video captioning datasets, MSR-VTT and MSVD, where it achieves state-of-the-art performance among the methods that take the PEFT approach across BLEU@4, METEOR, ROUGE-L, and CIDEr metrics. Q-Adapter also achieves competitive performance compared to methods that take the full fine-tuning approach while requiring only 1.4% of the parameters. We further analyze the impact of key hyperparameters and design choices on fine-tuning effectiveness, providing insights into optimization strategies for adapter-based learning. These results highlight the strong potential of Q-Adapter in balancing caption quality and parameter efficiency, demonstrating its scalability for video–language modeling.
Junan Chen 0004, Trung Thanh Nguyen 0006, Takahiro Komamizu, Ichiro Ide
MMAsia4
2024 R-DiP: Re-ranking Based Diffusion Pre-computation for Image Retrieval
Tatsuya Kato, Takahiro Komamizu, Ichiro Ide
DEXA (2)3
2024 Action Selection Learning for Multi-label Multi-view Action Recognition
abstract
Multi-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused on multi-view action recognition in a narrow area with strong labels available, where the onset and offset of each action are labeled at the frame-level. This study focuses on real-world scenarios where cameras are distributed to capture a wide-range area with only weak labels available at the video-level. We propose the method named Multi-view Action Selection Learning (MultiASL), which leverages action selection learning to enhance view fusion by selecting the most useful information from different viewpoints. The proposed method includes a Multi-view Spatial-Temporal Transformer video encoder to extract spatial and temporal features from multi-viewpoint videos. Action Selection Learning is employed at the frame-level, using pseudo ground-truth obtained from weak labels at the video-level, to identify the most relevant frames for action recognition. Experiments in a real-world office environment using the MM-Office dataset demonstrate the superior performance of the proposed method compared to existing methods. The source code is available at https://github.com/thanhhff/MultiASL/.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMAsia4
2023 RecipeMeta: Metapath-enhanced Recipe Recommendation on Heterogeneous Recipe Network
abstract
Recipe is a set of instructions that describes how to make food. It can help people from the preparation of ingredients, food cooking process, etc. to prepare the food, and increasingly in demand on the Web. To help users find the vast amount of recipes on the Web, we address the task of recipe recommendation. Due to multiple data types and relationships in a recipe, we can treat it as a heterogeneous network to describe its information more accurately. To effectively utilize the heterogeneous network, metapath was proposed to describe the higher-level semantic information between two entities by defining a compound path from peer entities. Therefore, we propose a metapath-enhanced recipe recommendation framework, RecipeMeta, that combines GNN (Graph Neural Network)-based representation learning and specific metapath-based information in a recipe to predict User-Recipe pairs for recommendation. Through extensive experiments, we demonstrate that the proposed model, RecipeMeta, outperforms state-of-the-art methods for recipe recommendation.
Jialiang Shi, Takahiro Komamizu, Keisuke Doman, Haruya Kyutoku, Ichiro Ide
MMAsia5
2021 MMArt-ACM'21: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2021
abstract
The International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and social media. The topics of the accepted papers cover an analytic topic on comic contents understanding to generative topics on image synthesis and conversion. The actual MMArt-ACM'21 Proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3460426.
Min-Chun Hu 0001, Ichiro Ide, Kensuke Tobitani
ICMR2
2020 MMArt-ACM'20: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2020
abstract
The International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and deep learning. Despite the strike of the Covid-19 pandemic, this workshop attracts submissions of diverse topics in these two fields, and the workshop program finally consists of five presented papers. The topics cover image retrieval, image transformation and generation, recommendation system, and image/video summarization. The actual MMArt-ACM'20 Proceedings are available in the ACM DL at: https://dl.acm.org/citation.cfm?id=3379173
Wei-Ta Chu, Ichiro Ide, Naoko Nitta, Norimichi Tsumura, Toshihiko Yamasaki
ICMR2
2020 CEA'20: The 12th Workshop on Multimedia for Cooking and Eating Activities
abstract
The 12th Workshop on Multimedia for Cooking and Eating Activities presents This overview introduces the aim of the CEA'20 workshop and the list of papers presented in the workshop.
Ichiro Ide, Yoko Yamakata, Atsushi Hashimoto 0001
ICMR1
2020 Imageability Estimation using Visual and Language Features
abstract
Imageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability.
Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
ICMR3
2014 Estimation of the Representative Story Transition in a Chronological Semantic Structure of News Topics
abstract
It is important to track the flow of topics to thoroughly understand the contents. Accordingly, a method that structures the chronological semantic relations between news stories, namely a "topic thread structure" has been proposed. It allows the comprehensive understanding of a topic by chronologically tracking stories one by one from the initial story. However, this task imposes a user to watch many stories when it contains various sub-topics. Thus, we propose a method that estimates the representative story transition in a topic thread structure. In the proposed method, features obtained from a story and those from the topic thread structure are used for the estimation. We confirmed the effectiveness of the proposed method by comparing the results obtained from the proposed method to the ground truth obtained from votes in a subjective experiment.
Kosuke Kato, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ICMR2
2011 Low Resolution QR-Code Recognition by Applying Super-Resolution Using the Property of QR-Codes
abstract
This paper proposes a method for low resolution QR-code recognition. A QR-code is a two-dimensional binary symbol that can embed various information such as characters and numbers. To recognize a QR-code correctly and stably, the resolution of an input image should be high. In practice, however, recognition of a QR-code is usually difficult due to low resolution when it is captured from a distance. In this paper, we propose a method to improve the performance of low resolution QR-code recognition by using the super-resolution technique that generates a high resolution image from multiple low-resolution images. Although a QR-code is a binary pattern, it is observed as a grayscale image due to the degradation through the capturing process. Especially the pixels around the borders between white and black regions become ambiguous. To overcome this problem, the proposed method introduces a binary pattern constraint to generate super-resolved images appropriate for recognition. Experimental results showed that a recognition rate of 98% can be achieved by the proposed method, which is a 15.7% improvement in comparison with a method using a conventional super-resolution method.
Yuji Kato, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR4
2009 Low-Resolution Character Recognition by Video-Based Super-Resolution
abstract
In this paper, we propose a method for recognizing low-resolution characters using a super-resolution technique. Although portable digital cameras can be used for camera based character recognition, the captured images contain several types of noises which make the recognition task difficult. We introduce a phase of super-resolution before the recognition to enhance the resolution of images obtained from a video. The proposed method uses the subspace method for the recognition of characters which are integrated from multiple low-resolution characters by the super-resolution technique. Experimental results show that the proposed method improves the recognition accuracy; we confirmed that the recognition rate for the input size of 7 times 7 pixels was 90.35%, and for the input size of 9 times 9 pixels was 99.97%.
Ataru Ohkura, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR4
2008 A Hilbert Warping Algorithm for Recognizing Characters from Moving Camera
abstract
We present a method for recognizing characters from image sequences captured by moving camera. In the proposed method, the sequence of the captured images is compared with those of reference character patterns using the concept of analytic signal. Since the captured image sequence can be nonlinearly warped along the time axis due to the movement of a hand-held camera, phase synchronization of two analytic signals is used for the alignment of two image sequences. Hilbert transform is used to convert all the image sequences into analytic signals whose phases are supposed to be increasing. Experimental results showed the usefulness of the proposed phase-based alignment algorithm.
Hiroyuki Ishida, Ichiro Ide, Hiroshi Murase, Tomokazu Takahashi
Document Analysis Systems2