Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daiki Shimada

dblp:194/7684 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0001-5840-0192ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 45% Video understanding and tracking · 35% Trustworthy machine learning · 21%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
video captioning
1.012026
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026
Computer vision › Vision and language › vision-language model › prompt learning
vision-language model prompting
1.012026
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.912025
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives · ICLR 2025
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.912025
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives · ICLR 2025
Computer vision › Video understanding and tracking › multimodal video understanding › audio-visual video understanding
audio-visual event localization
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Computer vision › Video understanding and tracking › temporal understanding
temporal video understanding
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Computer vision › Vision and language
video-language model
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Computer vision › Video understanding and tracking
video object segmentation
0.312026
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026

Methods — techniques the papers use, named apart from their topics

multimodal LLM fine-tuning · 1.7event-based video clustering · 1.7temporal analysis · 1.0segment anything · 1.0multimodal prompting · 1.0chain-of-thought prompting · 1.0temporal invariance · 0.9modality misalignment · 0.9adversarial training · 0.9
YearPublicationVenuePosition
2026 Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
abstract
In this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event boundary detection and coarse event descriptions, and (iii) an InternVL-2.5 Captioner that, conditioned on spatiotemporal visual prompts and chain-of-thought (CoT) guidance, produces detailed, temporally coherent captions about object attributes, actions, states, interactions, and context. The system supports point, box, and region prompts and maintains temporal sensitivity by tracking object states across segments. In contrast to vanilla video captioning that is overly abstract and dense video captioning that is often terse, CAT-V enables object-level specificity with spatial accuracy and temporal coherence, without additional training data.
Yunlong Tang 0002, Jing Bi 0002, Chao Huang 0033, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Pinxin Liu, Mingqian Feng, Junjia Guo, Luchuan Song, Ali Vosoughi, Jinxi He, Zeliang Zhang 0001, Jiebo Luo 0001, Chenliang Xu
AAAI5
2025 Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains. By fine-tuning multimodal LLMs with temporal annotations from well-annotated datasets, e.g., dense video captioning datasets, their temporal understanding capacity in video-language tasks can be obtained. However, there is a notable lack of untrimmed audio-visual video datasets with precise temporal annotations for events. This deficiency hinders LLMs from learning the alignment between time, audio-visual events, and text tokens, thus impairing their ability to localize audio-visual events in videos temporally. To address this gap, we introduce PU-VALOR, a comprehensive audio-visual dataset comprising over 114,081 pseudo-untrimmed videos with detailed temporal annotations. PU-VALOR is derived from the large-scale but coarse-annotated audio-visual dataset VALOR, through a subtle method involving event-based video clustering, random temporal scaling, and permutation. By fine-tuning a multimodal LLM on PU-VALOR, we developed AVicuna, a model capable of aligning audio-visual events with temporal intervals and corresponding text tokens. AVicuna excels in temporal localization and time-aware dialogue capabilities. Our experiments demonstrate that AVicuna effectively handles temporal understanding in audio-visual videos and achieves state-of-the-art performance on open-ended video QA, audio-visual QA, and audio-visual event dense localization tasks.
Yunlong Tang 0002, Daiki Shimada, Jing Bi 0002, Mingqian Feng, Hang Hua, Chenliang Xu
AAAI2
2025 Multi-Agent Path Finding Using Provisionally Booking Nodes for Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara
ICAART (3)1
2025 Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
abstract
While audio-visual learning equips models with a richer understanding of the real world by leveraging multiple sensory modalities, this integration also introduces new vulnerabilities to adversarial attacks. In this paper, we present a comprehensive study of the adversarial robustness of audio-visual models, considering both temporal and modality-specific vulnerabilities. We propose two powerful adversarial attacks: 1) a temporal invariance attack that exploits the inherent temporal redundancy across consecutive time segments and 2) a modality misalignment attack that introduces incongruence between the audio and visual modalities. These attacks are designed to thoroughly assess the robustness of audio-visual models against diverse threats. Furthermore, to defend against such attacks, we introduce a novel audio-visual adversarial training framework. This framework addresses key challenges in vanilla adversarial training by incorporating efficient adversarial perturbation crafting tailored to multi-modal data and an adversarial curriculum strategy. Extensive experiments in the Kinetics-Sounds dataset demonstrate that our proposed temporal and modality-based attacks in degrading model performance can achieve state-of-the-art performance, while our adversarial training defense largely improves the adversarial robustness as well as the adversarial training efficiency.
Zeliang Zhang 0001, Susan Liang, Daiki Shimada, Chenliang Xu
ICLR3
2024 Path Finding with Flexible Provisional Booking in Multi-agent Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara
PRIMA1
2016 Document classification through image-based character embedding and wildcard training
abstract
Languages such as Chinese and Japanese have a significantly large number (several thousands) of alphabets as compared to other languages, and each of their sentences consists of several concatenated words with wide varieties of inflected forms; thus appropriate word segmentation is quite difficult. Therefore, recently proposed sophisticated language-processing methods designed for languages such as English cannot be applied. In this paper, we address those issues and propose a new and efficient document classification technique for such languages. The proposed method is characterized into a new “image-based character embedding” method and character-level convolutional neural networks method with “wildcard training.” The first method encodes each character based on its pictorial structures and preserves them. Further, the second method treats some of the input characters as wildcards in the classification stage and functions as efficient data augmentation. We confirmed that our proposed method showed superior performance when compared conventional methods for Japanese document classification problems.
Daiki Shimada, Ryunosuke Kotani, Hitoshi Iyatomi
IEEE BigData1