EDBT 2026 Demo / reviewers in the wild / expert
Xinkun Yu
dblp:346/0865
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Vision and language · 77% Image recognition and object detection · 23% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual question answering |
0.6 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Multimedia analysis and retrieval › video understanding
multimodal video understanding |
0.6 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.2 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
multimodal transformer fusion · 1.1answer generation · 1.1OCR token extraction · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Towards Video Text Visual Question Answering: Benchmark and BaselineabstractThere are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring, shopping ads and e-learning videos) where temporal reasoning ability is required. To this end, we propose a new task named Video Text Visual Question Answering (ViteVQA in short) that aims at answering questions by reasoning texts and visual information spatiotemporally in a given video. In particular, on the one hand, we build the first ViteVQA benchmark dataset named M4-ViteVQA --- the abbreviation of Multi-category Multi-frame Multi-resolution Multi-modal benchmark for ViteVQA, which contains 7,620 video clips of 9 categories (i.e., shopping, traveling, driving, vlog, sport, advertisement, movie, game and talking) and 3 kinds of resolutions (i.e., 720p, 1080p and 1176x664), and 25,123 question-answer pairs. On the other hand, we develop a baseline method named T5-ViteVQA for the ViteVQA task. T5-ViteVQA consists of five transformers. It first extracts optical character recognition (OCR) tokens, question features, and video representations via two OCR transformers, one language transformer and one video-language transformer, respectively. Then, a multimodal fusion transformer and an answer generation module are applied to fuse multimodal information and generate the final prediction. Extensive experiments on M4-ViteVQA demonstrate the superiority of T5-ViteVQA to the existing approaches of TextVQA and VQA tasks. The ViteVQA benchmark is available in https://github.com/bytedance/VTVQA. Minyi Zhao, Bingjia Li, Wanqing Li 0007, Shijie Xuyang, Zhihang Yu, Xinkun Yu, Guangze Li, Aobotao Dai, Shuigeng Zhou |
NeurIPS | 9 |