VLDB 2026 Research / reviewers in the wild / expert
Seitaro Otsuki
dblp:321/6903
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0009-8071-6060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 31% Trustworthy machine learning · 27% Language models and text generation · 23% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval › multimedia analysis › multimedia content description
image captioning |
1.0 | 1 | 2026 | LLM-Free Image Captioning Evaluation in Reference-Flexible Settings · AAAI 2026 |
Multimedia analysis and retrieval
image captioning evaluation |
1.0 | 1 | 2026 | LLM-Free Image Captioning Evaluation in Reference-Flexible Settings · AAAI 2026 |
Computer vision › Vision and language › image captioning
image caption evaluation |
0.9 | 1 | 2025 | VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions · EMNLP 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Layer-Wise Relevance Propagation with Conservation Property for ResNet · ECCV (43) 2024 |
Machine learning › Deep learning architectures and training › convolutional neural network
residual network |
0.8 | 1 | 2024 | Layer-Wise Relevance Propagation with Conservation Property for ResNet · ECCV (43) 2024 |
Computer vision › Vision and language
image captioning |
0.3 | 1 | 2026 | LLM-Free Image Captioning Evaluation in Reference-Flexible Settings · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
supervised metric learning · 2.0image-caption representation learning · 2.0multimodal large language model · 0.9layer-wise relevance propagation · 0.8conservation property · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsabstractWe focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas they do not always demonstrate high performance. To address these issues, we propose Pearl, an LLM-free supervised metric for image captioning, which is applicable to both reference-based and reference-free settings. We introduce a novel mechanism that learns the representations of image--caption and caption--caption similarities. Furthermore, we construct a human-annotated dataset for image captioning metrics that comprises approximately 333k human judgments collected from 2,360 annotators across over 75k images. Pearl outperformed other existing LLM-free metrics on the Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and FOIL datasets in both reference-based and reference-free settings. Shinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki, Komei Sugiura |
AAAI | 4 |
| 2026 | ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
Daichi Yashima, Shuhei Kurita, Yusuke Oda, Shuntaro Suzuki, Seitaro Otsuki, Komei Sugiura |
ICPR (3) | 5 |
| 2025 | VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsabstractIn this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs).Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions.Moreover, recent LLM-as-a-Judge approaches suffer from slow inference due to their reliance on autoregressive inference and early fusion of visual information.To address these limitations, we propose VELA, an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework.Furthermore, we propose LongCap-Arena, a benchmark specifically designed for evaluating metrics for long captions.This benchmark comprises 7,805 images, the corresponding human-provided long reference captions and long candidate captions, and 32,246 human judgments from three distinct perspectives: Descriptiveness, Relevance, and Fluency.We demonstrated that VELA outperformed existing metrics and achieved superhuman performance on LongCap-Arena.Our code and dataset are available at https://vela.kinsta.page/. Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki, Komei Sugiura |
EMNLP | 4 |
| 2024 | Layer-Wise Relevance Propagation with Conservation Property for ResNet
Seitaro Otsuki, Tsumugi Iida, Félix Doublet, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura |
ECCV (43) | 1 |
| 2023 | Prototypical Contrastive Transfer Learning for Multimodal Language UnderstandingabstractAlthough domestic service robots are expected to assist individuals who require support, they cannot currently interact smoothly with people through natural language. For example, given the instruction “Bring me a bottle from the kitchen,” it is difficult for such robots to specify the bottle in an indoor environment. Most conventional models have been trained on real-world datasets that are labor-intensive to collect, and they have not fully leveraged simulation data through a transfer learning framework. In this study, we propose a novel transfer learning approach for multimodal language understanding called Prototypical Contrastive Transfer Learning (PCTL), which uses a new contrastive loss called Dual ProtoNCE. We introduce PCTL to the task of identifying target objects in domestic environments according to free-form natural language instructions. To validate PCTL, we built new real-world and simulation datasets. Our experiment demonstrated that PCTL outperformed existing methods. Specifically, PCTL achieved an accuracy of 78.1 %, whereas simple fine-tuning achieved an accuracy of 73.4 %. Seitaro Otsuki, Shintaro Ishikawa, Komei Sugiura |
IROS | 1 |
| 2022 | Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass EstimationabstractFor human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-based geometric algorithm to estimate 3D models of containers for the estimation of their capacity and dimensions. We further use these estimations to estimate their mass in a Convolutional Neural Network model. Experiments show that our Transformer model produced encouraging results in both estimations. While challenges remain in our mask-based algorithm and Convolutional Neural Network model, their results revealed several ways for improvement. Tomoya Matsubara, Seitaro Otsuki, Yuiga Wada, Haruka Matsuo, Takumi Komatsu, Yui Iioka, Komei Sugiura, Hideo Saito 0001 |
ICASSP | 2 |