Mahiro Ukai

dblp:382/2174 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0001-6266-9477ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 80% Image recognition and object detection · 11% Language models and text generation · 9%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › image captioning
image caption evaluation
1.012026
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning · AAAI 2026
Computer vision › Vision and language
image captioning
1.012026
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning · AAAI 2026
Computer vision › Vision and language
vision-language model
1.012026
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning · AAAI 2026
Computer vision › Image recognition and object detection › object recognition
object state recognition
0.912025
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models · ACM Multimedia 2025
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.912025
Referring Expression Comprehension for Small Objects · ICCV 2025
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.912025
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models · ACM Multimedia 2025
Natural language and speech › Language models and text generation › prompting
prompt compression
0.812024
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering · ACM Multimedia 2024
Computer vision › Vision and language
visual question answering
0.812024
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

test-time adaptation · 1.0gaussian prior · 1.0analytical solution · 1.0vision-language model · 0.9progressive-iterative zooming adapter · 0.9parameter-efficient fine-tuning · 0.9fine-tuning · 0.9prompt compression · 0.8large language model · 0.8
YearPublicationVenuePosition
2026 DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
abstract
Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challenging, particularly under domain-shift scenarios. To address this issue, we introduce the Distribution-Aware Score Decoder (DISCODE), a novel finetuning-free method that generates robust evaluation scores better aligned with human judgments across diverse domains. The core idea behind DISCODE lies in its test-time adaptive evaluation approach, which introduces the Adaptive Test-Time (ATT) loss, leveraging a Gaussian prior distribution to improve robustness in evaluation score estimation. This loss is efficiently minimized at test time using an analytical solution that we derive. Furthermore, we introduce the Multi-domain Caption Evaluation (MCEval) benchmark, a new image captioning evaluation benchmark covering six distinct domains, designed to assess the robustness of evaluation metrics. In our experiments, we demonstrate that DISCODE achieves state-of-the-art performance as a reference-free evaluation metric across MCEval and four representative existing benchmarks.
Nakamasa Inoue, Kanoko Goto, Masanari Oi, Martyna Gruszka, Mahiro Ukai, Takumi Hirose, Yusuke Sekikawa
AAAI5
2025 Referring Expression Comprehension for Small Objects
abstract
Referring expression comprehension (REC) aims to localize the target object described by a natural language expression. Recent advances in vision-language learning have led to significant performance improvements in REC tasks. However, localizing extremely small objects remains a considerable challenge despite its importance in real-world applications such as autonomous driving. To address this issue, we introduce a novel dataset and method for REC targeting small objects. First, we present the small object REC (SOREC) dataset, which consists of 100,000 pairs of referring expressions and corresponding bounding boxes for small objects in driving scenarios. Second, we propose the progressive-iterative zooming adapter (PIZA), an adapter module for parameter-efficient fine-tuning that enables models to progressively zoom in and localize small objects. In a series of experiments, we apply PIZA to GroundingDINO and demonstrate a significant improvement in accuracy on the SOREC dataset. Our dataset, codes and pre-trained models are publicly available on the project page.
Kanoko Goto, Takumi Hirose, Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue
ICCV3
2025 STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
abstract
Object state recognition aims to identify the specific condition of objects, such as their positional states (e.g., open or closed) and functional states (e.g., on or off). While recent Vision-Language Models (VLMs) are capable of performing a variety of multimodal tasks, it remains unclear how precisely they can identify object states. To alleviate this issue, we introduce the STAte and Transition UnderStanding Benchmark (STATUS Bench), the first benchmark for rigorously evaluating the ability of VLMs to understand subtle variations in object states in diverse situations. Specifically, STATUS Bench introduces a novel evaluation scheme that requires VLMs to perform three tasks simultaneously: object state identification (OSI), image retrieval (IR), and state change identification (SCI). These tasks are defined over our fully hand-crafted dataset involving image pairs, their corresponding object state descriptions and state change descriptions. Furthermore, we introduce a large-scale training dataset, namely STATUS Train, which consists of 13 million semi-automatically created descriptions. This dataset serves as the largest resource to facilitate further research in this area. In our experiments, we demonstrate that STATUS Bench enables rigorous consistency evaluation and reveal that current state-of-the-art VLMs still significantly struggle to capture subtle object state distinctions. Surprisingly, under the proposed rigorous evaluation scheme, most open-weight VLMs exhibited chance-level zero-shot performance. After fine-tuning on STATUS Train, Qwen2.5-VL achieved performance comparable to Gemini 2.0 Flash. These findings underscore the necessity of STATUS Bench and Train for advancing object state recognition in VLM research.
Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue
ACM Multimedia1
2024 PolarDB: Formula-Driven Dataset for Pre-Training Trajectory Encoders
abstract
Formula-driven supervised learning (FDSL) is a growing research topic for finding simple mathematical formulas that generate synthetic data and labels for pre-training neural networks. The main advantage of FDSL is that there is no risk of generating data with ethical implications such as gender bias and racial bias because it does not rely on real data as discussed in previous studies using fractals and polygons for pre-training image encoders. While FDSL has been proposed for pre-training image encoders, it has not been considered for temporal trajectory data. In this paper, we introduce PolarDB, the first formula-driven dataset for pre-training trajectory encoders with an application to fine-grained cutting-method recognition using hand trajectories. More specifically, we generate 270k trajectories for 432 categories on the basis of polar equations and use them to pre-train a Transformer-based trajectory encoder in an FDSL manner. In the experiments, we show that pre-training on PolarDB improves the accuracy of fine-grained cutting-method recognition on cooking videos of EPIC-KITCHEN and Ego4D datasets, where the pre-trained trajectory encoder is used as a plug-in module for a video recognition network.
Sota Miyamoto, Takuma Yagi, Yuto Makimoto, Mahiro Ukai, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Nakamasa Inoue
ICASSP4
2024 AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
abstract
Visual question answering aims to provide responses to questions given visual input. Recently, visual programmatic models (VPMs), which generate programs to answer questions through large language models (LLMs), have attracted attention. However, they often require long input prompts to provide the LLM with sufficient API usage details to generate relevant code. To address this limitation, we propose AdaCoder, an adaptive prompt compression framework for VPMs. AdaCoder operates in two phases: a compression phase and an inference phase. In the compression phase, given a preprompt that describes all API definitions with example code snippets, a set of compressed preprompts is generated, each depending on a specific question type. In the inference phase, AdaCoder predicts the question type and chooses the appropriate corresponding compressed preprompt to generate code to answer the question. In experiments, we apply AdaCoder to ViperGPT and demonstrate that it reduces token length by 71.1%, while maintaining or even improving the performance of visual question answering.
Mahiro Ukai, Shuhei Kurita, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Nakamasa Inoue
ACM Multimedia1