EDBT 2026 Demo / reviewers in the wild / expert
Aobotao Dai
dblp:346/1029
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 52% 3D vision · 26% Vision and language · 17% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
object tracking |
0.9 | 1 | 2025 | Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment · NeurIPS 2025 |
Computer vision › Video understanding and tracking
video instance segmentation |
0.9 | 1 | 2025 | Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment · NeurIPS 2025 |
Computer vision › Vision and language
visual question answering |
0.6 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Multimedia analysis and retrieval › video understanding
multimodal video understanding |
0.6 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.2 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
multimodal transformer fusion · 1.1answer generation · 1.1OCR token extraction · 1.1large language model · 0.9SAM2 · 0.92d segmentor · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environmentabstract360 video captures the complete surrounding scenes with the ultra-large field of view of 360x180. This makes 360 scene understanding tasks, e.g., segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With the recent emergence of foundation models, the community is, however, impeded by the lack of large-scale, labelled real-world datasets. This is caused by the inherent spherical properties, e.g., severe distortion in polar regions, and content discontinuities, rendering the annotation costly yet complex. This paper introduces Leader360V, the first large-scale (10K+), labeled real-world 360 video datasets for instance segmentation and tracking. Our datasets enjoy high scene diversity, ranging from indoor and urban settings to natural and dynamic outdoor scenes. To automate annotation, we design an automatic labeling pipeline, which subtly coordinates pre-trained 2D segmentors and large language models (LLMs) to facilitate the labeling. The pipeline operates in three novel stages. Specifically, in the Initial Annotation Phase, we introduce a Semantic- and Distortion-aware Refinement (SDR) module, which combines object mask proposals from multiple 2D segmentors with LLM-verified semantic labels. These are then converted into mask prompts to guide SAM2 in generating distortion-aware masks for subsequent frames. In the Auto-Refine Annotation Phase, missing or incomplete regions are corrected either by applying the SDR again or resolving the discontinuities near the horizontal borders. The Manual Revision Phase finally incorporates LLMs and human annotators to further refine and validate the annotations. Extensive user studies and evaluations demonstrate the effectiveness of our labeling pipeline. Meanwhile, experiments confirm that Leader360V significantly enhances model performance for 360 video segmentation and tracking, paving the way for more scalable 360 scene understanding. We release our dataset and code at {https://leader360v.github.io/Leader360V_HomePage/} for better understanding. Dingwen Xiao, Aobotao Dai, Yexin Liu, Tianbo Pan, Shiqi Wen, Lei Chen 0002, Lin Wang 0040 |
NeurIPS | 3 |
| 2022 | Towards Video Text Visual Question Answering: Benchmark and BaselineabstractThere are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring, shopping ads and e-learning videos) where temporal reasoning ability is required. To this end, we propose a new task named Video Text Visual Question Answering (ViteVQA in short) that aims at answering questions by reasoning texts and visual information spatiotemporally in a given video. In particular, on the one hand, we build the first ViteVQA benchmark dataset named M4-ViteVQA --- the abbreviation of Multi-category Multi-frame Multi-resolution Multi-modal benchmark for ViteVQA, which contains 7,620 video clips of 9 categories (i.e., shopping, traveling, driving, vlog, sport, advertisement, movie, game and talking) and 3 kinds of resolutions (i.e., 720p, 1080p and 1176x664), and 25,123 question-answer pairs. On the other hand, we develop a baseline method named T5-ViteVQA for the ViteVQA task. T5-ViteVQA consists of five transformers. It first extracts optical character recognition (OCR) tokens, question features, and video representations via two OCR transformers, one language transformer and one video-language transformer, respectively. Then, a multimodal fusion transformer and an answer generation module are applied to fuse multimodal information and generate the final prediction. Extensive experiments on M4-ViteVQA demonstrate the superiority of T5-ViteVQA to the existing approaches of TextVQA and VQA tasks. The ViteVQA benchmark is available in https://github.com/bytedance/VTVQA. Minyi Zhao, Bingjia Li, Wanqing Li 0007, Shijie Xuyang, Zhihang Yu, Xinkun Yu, Guangze Li, Aobotao Dai, Shuigeng Zhou |
NeurIPS | 11 |