VLDB 2026 Research / reviewers in the wild / expert
Yuiga Wada
dblp:321/6947
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-3804-4546ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 72% Language models and text generation · 18% Trustworthy machine learning · 5% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › image captioning
image caption evaluation |
1.6 | 2 | 2025 | VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions · EMNLP 2025 Polos: Multimodal Metric Learning from Human Feedback for Image Captioning · CVPR 2024 |
Computer vision › Vision and language
image captioning |
1.1 | 2 | 2026 | Polos: Multimodal Metric Learning from Human Feedback for Image Captioning · CVPR 2024 LLM-Free Image Captioning Evaluation in Reference-Flexible Settings · AAAI 2026 |
Multimedia analysis and retrieval › multimedia analysis › multimedia content description
image captioning |
1.0 | 1 | 2026 | LLM-Free Image Captioning Evaluation in Reference-Flexible Settings · AAAI 2026 |
Multimedia analysis and retrieval
image captioning evaluation |
1.0 | 1 | 2026 | LLM-Free Image Captioning Evaluation in Reference-Flexible Settings · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions · EMNLP 2025 |
Computer vision › Vision and language
vision-language pretraining |
0.8 | 1 | 2024 | Polos: Multimodal Metric Learning from Human Feedback for Image Captioning · CVPR 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2024 | Polos: Multimodal Metric Learning from Human Feedback for Image Captioning · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
supervised metric learning · 2.0image-caption representation learning · 2.0multimodal large language model · 0.9human feedback · 0.8contrastive learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsabstractWe focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas they do not always demonstrate high performance. To address these issues, we propose Pearl, an LLM-free supervised metric for image captioning, which is applicable to both reference-based and reference-free settings. We introduce a novel mechanism that learns the representations of image--caption and caption--caption similarities. Furthermore, we construct a human-annotated dataset for image captioning metrics that comprises approximately 333k human judgments collected from 2,360 annotators across over 75k images. Pearl outperformed other existing LLM-free metrics on the Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and FOIL datasets in both reference-based and reference-free settings. Shinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki, Komei Sugiura |
AAAI | 2 |
| 2025 | VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsabstractIn this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs).Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions.Moreover, recent LLM-as-a-Judge approaches suffer from slow inference due to their reliance on autoregressive inference and early fusion of visual information.To address these limitations, we propose VELA, an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework.Furthermore, we propose LongCap-Arena, a benchmark specifically designed for evaluating metrics for long captions.This benchmark comprises 7,805 images, the corresponding human-provided long reference captions and long candidate captions, and 32,246 human judgments from three distinct perspectives: Descriptiveness, Relevance, and Fluency.We demonstrated that VELA outperformed existing metrics and achieved superhuman performance on LongCap-Arena.Our code and dataset are available at https://vela.kinsta.page/. Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki, Komei Sugiura |
EMNLP | 2 |
| 2024 | Deneb: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
Kazuki Matsuda, Yuiga Wada, Komei Sugiura |
ACCV (3) | 2 |
| 2024 | Polos: Multimodal Metric Learning from Human Feedback for Image CaptioningabstractEstablishing an automatic evaluation metric that closely aligns with human judgments is essential for effectively developing image captioning models. Recent data-driven metrics have demonstrated a stronger correlation with human judgments than classic metrics such as CIDEr; however they lack sufficient capabilities to handle hallucinations and generalize across diverse images and texts partially because they compute scalar similarities merely using embeddings learned from tasks unrelated to image captioning evaluation. In this study, we propose Polos, a supervised automatic evaluation metric for image captioning models. Polos computes scores from multimodal inputs, using a parallel feature extraction mechanism that leverages embeddings trained through large-scale contrastive learning. To train Polos, we introduce Multimodal Metric Learning from Human Feedback (M2LHF), a framework for developing metrics based on human feedback. We constructed the Polaris dataset, which comprises 131K human judgments from 550 evaluators, which is approximately ten times larger than standard datasets. Our approach achieved state-of-the-art performance on Composite, Flickr8K-Expert, Flickr8K-CF, PASCAL-50S, FOIL, and the Polaris dataset, thereby demonstrating its effectiveness and robustness. Yuiga Wada, Kanta Kaneda, Daichi Saito, Komei Sugiura |
CVPR | 1 |
| 2023 | JaSPICE: Automatic Evaluation Metric Using Predicate-Argument Structures for Image Captioning ModelsabstractImage captioning studies heavily rely on automatic evaluation metrics such as BLEU and METEOR.However, such n-gram-based metrics have been shown to correlate poorly with human evaluation, leading to the proposal of alternative metrics such as SPICE for English; however, no equivalent metrics have been established for other languages.Therefore, in this study, we propose an automatic evaluation metric called JaSPICE, which evaluates Japanese captions based on scene graphs.The proposed method generates a scene graph from dependencies and the predicate-argument structure, and extends the graph using synonyms.We conducted experiments employing 10 image captioning models trained on STAIR Captions and PFN-PIC and constructed the Shichimi dataset, which contains 103,170 human evaluations.The results showed that our metric outperformed the baseline metrics for the correlation coefficient with the human evaluation. Yuiga Wada, Kanta Kaneda, Komei Sugiura |
CoNLL | 1 |
| 2023 | Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation InstructionsabstractIn this study, we aim to develop a model that comprehends a natural language instruction (e.g., “Go to the living room and get the nearest pillow to the radio art on the wall”) and generates a segmentation mask for the target everyday object. The task is challenging because it requires (1) the understanding of the referring expressions for multiple objects in the instruction, (2) the prediction of the target phrase of the sentence among the multiple phrases, and (3) the generation of pixel-wise segmentation masks rather than bounding boxes. Studies have been conducted on language-based segmentation methods; however, they sometimes mask irrelevant regions for complex sentences. In this paper, we propose the Multimodal Diffusion Segmentation Model (MDSM), which generates a mask in the first stage and refines it in the second stage. We introduce a crossmodal parallel feature extraction mechanism and extend diffusion probabilistic models to handle crossmodal features. To validate our model, we built a new dataset based on the well-known Matterport3D and REVERIE datasets. This dataset consists of instructions with complex referring expressions accompanied by real indoor environmental images that feature various target objects, in addition to pixel-wise segmentation masks. The performance of MDSM surpassed that of the baseline method by a large margin of +10.13 mean IoU. Yui Iioka, Yu Yoshida, Yuiga Wada, Shumpei Hatanaka, Komei Sugiura |
IROS | 3 |
| 2022 | Flare Transformer: Solar Flare Prediction Using Magnetograms and Sunspot Physical Features
Kanta Kaneda, Yuiga Wada, Tsumugi Iida, Naoto Nishizuka, Yûki Kubo, Komei Sugiura |
ACCV (2) | 2 |
| 2022 | Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass EstimationabstractFor human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-based geometric algorithm to estimate 3D models of containers for the estimation of their capacity and dimensions. We further use these estimations to estimate their mass in a Convolutional Neural Network model. Experiments show that our Transformer model produced encouraging results in both estimations. While challenges remain in our mask-based algorithm and Convolutional Neural Network model, their results revealed several ways for improvement. Tomoya Matsubara, Seitaro Otsuki, Yuiga Wada, Haruka Matsuo, Takumi Komatsu, Yui Iioka, Komei Sugiura, Hideo Saito 0001 |
ICASSP | 3 |