Xinkun Yu

dblp:346/0865 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 77% Image recognition and object detection · 23%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual question answering
0.612022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022
Multimedia analysis and retrieval › video understanding
multimodal video understanding
0.612022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022
Computer vision › Image recognition and object detection › text recognition
optical character recognition
0.212022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

multimodal transformer fusion · 1.1answer generation · 1.1OCR token extraction · 1.1
YearPublicationVenuePosition
2022 Towards Video Text Visual Question Answering: Benchmark and Baseline
abstract
There are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring, shopping ads and e-learning videos) where temporal reasoning ability is required. To this end, we propose a new task named Video Text Visual Question Answering (ViteVQA in short) that aims at answering questions by reasoning texts and visual information spatiotemporally in a given video. In particular, on the one hand, we build the first ViteVQA benchmark dataset named M4-ViteVQA --- the abbreviation of Multi-category Multi-frame Multi-resolution Multi-modal benchmark for ViteVQA, which contains 7,620 video clips of 9 categories (i.e., shopping, traveling, driving, vlog, sport, advertisement, movie, game and talking) and 3 kinds of resolutions (i.e., 720p, 1080p and 1176x664), and 25,123 question-answer pairs. On the other hand, we develop a baseline method named T5-ViteVQA for the ViteVQA task. T5-ViteVQA consists of five transformers. It first extracts optical character recognition (OCR) tokens, question features, and video representations via two OCR transformers, one language transformer and one video-language transformer, respectively. Then, a multimodal fusion transformer and an answer generation module are applied to fuse multimodal information and generate the final prediction. Extensive experiments on M4-ViteVQA demonstrate the superiority of T5-ViteVQA to the existing approaches of TextVQA and VQA tasks. The ViteVQA benchmark is available in https://github.com/bytedance/VTVQA.
Minyi Zhao, Bingjia Li, Wanqing Li 0007, Shijie Xuyang, Zhihang Yu, Xinkun Yu, Guangze Li, Aobotao Dai, Shuigeng Zhou
NeurIPS9