EDBT 2026 Demo / reviewers in the wild / expert
Daiki Shimada
dblp:194/7684
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0001-5840-0192ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 45% Video understanding and tracking · 35% Trustworthy machine learning · 21% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
video captioning |
1.0 | 1 | 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026 |
Computer vision › Vision and language › vision-language model › prompt learning
vision-language model prompting |
1.0 | 1 | 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.9 | 1 | 2025 | Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.9 | 1 | 2025 | Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives · ICLR 2025 |
Computer vision › Video understanding and tracking › multimodal video understanding › audio-visual video understanding
audio-visual event localization |
0.9 | 1 | 2025 | Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025 |
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding |
0.9 | 1 | 2025 | Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025 |
Computer vision › Video understanding and tracking › temporal understanding
temporal video understanding |
0.9 | 1 | 2025 | Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025 |
Computer vision › Vision and language
video-language model |
0.9 | 1 | 2025 | Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025 |
Computer vision › Video understanding and tracking
video object segmentation |
0.3 | 1 | 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
multimodal LLM fine-tuning · 1.7event-based video clustering · 1.7temporal analysis · 1.0segment anything · 1.0multimodal prompting · 1.0chain-of-thought prompting · 1.0temporal invariance · 0.9modality misalignment · 0.9adversarial training · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal PromptingabstractIn this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event boundary detection and coarse event descriptions, and (iii) an InternVL-2.5 Captioner that, conditioned on spatiotemporal visual prompts and chain-of-thought (CoT) guidance, produces detailed, temporally coherent captions about object attributes, actions, states, interactions, and context. The system supports point, box, and region prompts and maintains temporal sensitivity by tracking object states across segments. In contrast to vanilla video captioning that is overly abstract and dense video captioning that is often terse, CAT-V enables object-level specificity with spatial accuracy and temporal coherence, without additional training data. Yunlong Tang 0002, Jing Bi 0002, Chao Huang 0033, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Pinxin Liu, Mingqian Feng, Junjia Guo, Luchuan Song, Ali Vosoughi, Jinxi He, Zeliang Zhang 0001, Jiebo Luo 0001, Chenliang Xu |
AAAI | 5 |
| 2025 | Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal UnderstandingabstractLarge language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains. By fine-tuning multimodal LLMs with temporal annotations from well-annotated datasets, e.g., dense video captioning datasets, their temporal understanding capacity in video-language tasks can be obtained. However, there is a notable lack of untrimmed audio-visual video datasets with precise temporal annotations for events. This deficiency hinders LLMs from learning the alignment between time, audio-visual events, and text tokens, thus impairing their ability to localize audio-visual events in videos temporally. To address this gap, we introduce PU-VALOR, a comprehensive audio-visual dataset comprising over 114,081 pseudo-untrimmed videos with detailed temporal annotations. PU-VALOR is derived from the large-scale but coarse-annotated audio-visual dataset VALOR, through a subtle method involving event-based video clustering, random temporal scaling, and permutation. By fine-tuning a multimodal LLM on PU-VALOR, we developed AVicuna, a model capable of aligning audio-visual events with temporal intervals and corresponding text tokens. AVicuna excels in temporal localization and time-aware dialogue capabilities. Our experiments demonstrate that AVicuna effectively handles temporal understanding in audio-visual videos and achieves state-of-the-art performance on open-ended video QA, audio-visual QA, and audio-visual event dense localization tasks. Yunlong Tang 0002, Daiki Shimada, Jing Bi 0002, Mingqian Feng, Hang Hua, Chenliang Xu |
AAAI | 2 |
| 2025 | Multi-Agent Path Finding Using Provisionally Booking Nodes for Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara |
ICAART (3) | 1 |
| 2025 | Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality PerspectivesabstractWhile audio-visual learning equips models with a richer understanding of the real world by leveraging multiple sensory modalities, this integration also introduces new vulnerabilities to adversarial attacks.
In this paper, we present a comprehensive study of the adversarial robustness of audio-visual models, considering both temporal and modality-specific vulnerabilities. We propose two powerful adversarial attacks: 1) a temporal invariance attack that exploits the inherent temporal redundancy across consecutive time segments and 2) a modality misalignment attack that introduces incongruence between the audio and visual modalities. These attacks are designed to thoroughly assess the robustness of audio-visual models against diverse threats. Furthermore, to defend against such attacks, we introduce a novel audio-visual adversarial training framework. This framework addresses key challenges in vanilla adversarial training by incorporating efficient adversarial perturbation crafting tailored to multi-modal data and an adversarial curriculum strategy. Extensive experiments in the Kinetics-Sounds dataset demonstrate that our proposed temporal and modality-based attacks in degrading model performance can achieve state-of-the-art performance, while our adversarial training defense largely improves the adversarial robustness as well as the adversarial training efficiency. Zeliang Zhang 0001, Susan Liang, Daiki Shimada, Chenliang Xu |
ICLR | 3 |
| 2024 | Path Finding with Flexible Provisional Booking in Multi-agent Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara |
PRIMA | 1 |
| 2016 | Document classification through image-based character embedding and wildcard trainingabstractLanguages such as Chinese and Japanese have a significantly large number (several thousands) of alphabets as compared to other languages, and each of their sentences consists of several concatenated words with wide varieties of inflected forms; thus appropriate word segmentation is quite difficult. Therefore, recently proposed sophisticated language-processing methods designed for languages such as English cannot be applied. In this paper, we address those issues and propose a new and efficient document classification technique for such languages. The proposed method is characterized into a new “image-based character embedding” method and character-level convolutional neural networks method with “wildcard training.” The first method encodes each character based on its pictorial structures and preserves them. Further, the second method treats some of the input characters as wildcards in the classification stage and functions as efficient data augmentation. We confirmed that our proposed method showed superior performance when compared conventional methods for Japanese document classification problems. Daiki Shimada, Ryunosuke Kotani, Hitoshi Iyatomi |
IEEE BigData | 1 |