Bingjia Li

dblp:305/2924 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 70% Vision and language · 30%
Computer graphics and multimedia
2 papers
Image and video processing · 67% Multimedia analysis and retrieval · 33%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
scene text recognition
0.612022
C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022
Computer vision › Image recognition and object detection
text recognition
0.612022
C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022
Computer vision › Vision and language
visual question answering
0.612022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022
Multimedia analysis and retrieval › video understanding
multimodal video understanding
0.612022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution
0.612022
C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022
Image and video processing
super-resolution
0.612022
C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022
Computer vision › Image recognition and object detection › text recognition
optical character recognition
0.212022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

recognizer feedback · 1.1multimodal transformer fusion · 1.1cross-modal clue fusion · 1.1character-level language model · 1.1answer generation · 1.1OCR token extraction · 1.1
YearPublicationVenuePosition
2025 HiREN: Towards higher supervision quality for better scene text image super-resolution
Minyi Zhao, Yi Xu 0003, Bingjia Li, Jihong Guan, Shuigeng Zhou
Neurocomputing3
2022 Two-Stage Multimodality Fusion for High-Performance Text-Based Visual Question Answering
Bingjia Li, Minyi Zhao, Shuigeng Zhou
ACCV (4)1
2022 C3-STISR: Scene Text Image Super-resolution with Triple Clues
abstract
Scene text image super-resolution (STISR) has been regarded as an important pre-processing task for text recognition from low-resolution scene text images. Most recent approaches use the recognizer's feedback as clues to guide super-resolution. However, directly using recognition clue has two problems: 1) Compatibility. It is in the form of probability distribution, has an obvious modal gap with STISR - a pixel-level task; 2) Inaccuracy. it usually contains wrong information, thus will mislead the main task and degrade super-resolution performance. In this paper, we present a novel method C3-STISR that jointly exploits the recognizer's feedback, visual and linguistical information as clues to guide super-resolution. Here, visual clue is from the images of texts predicted by the recognizer, which is informative and more compatible with the STISR task; while linguistical clue is generated by a pre-trained character-level language model, which is able to correct the predicted texts. We design effective extraction and fusion mechanisms for the triple cross-modal clues to generate a comprehensive and unified guidance for super-resolution. Extensive experiments on TextZoom show that C3-STISR outperforms the SOTA methods in fidelity and recognition performance. Code is available in https://github.com/zhaominyiz/C3-STISR.
Minyi Zhao, Fan Bai 0001, Bingjia Li, Shuigeng Zhou
IJCAI4
2022 Towards Video Text Visual Question Answering: Benchmark and Baseline
abstract
There are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring, shopping ads and e-learning videos) where temporal reasoning ability is required. To this end, we propose a new task named Video Text Visual Question Answering (ViteVQA in short) that aims at answering questions by reasoning texts and visual information spatiotemporally in a given video. In particular, on the one hand, we build the first ViteVQA benchmark dataset named M4-ViteVQA --- the abbreviation of Multi-category Multi-frame Multi-resolution Multi-modal benchmark for ViteVQA, which contains 7,620 video clips of 9 categories (i.e., shopping, traveling, driving, vlog, sport, advertisement, movie, game and talking) and 3 kinds of resolutions (i.e., 720p, 1080p and 1176x664), and 25,123 question-answer pairs. On the other hand, we develop a baseline method named T5-ViteVQA for the ViteVQA task. T5-ViteVQA consists of five transformers. It first extracts optical character recognition (OCR) tokens, question features, and video representations via two OCR transformers, one language transformer and one video-language transformer, respectively. Then, a multimodal fusion transformer and an answer generation module are applied to fuse multimodal information and generate the final prediction. Extensive experiments on M4-ViteVQA demonstrate the superiority of T5-ViteVQA to the existing approaches of TextVQA and VQA tasks. The ViteVQA benchmark is available in https://github.com/bytedance/VTVQA.
Minyi Zhao, Bingjia Li, Wanqing Li 0007, Shijie Xuyang, Zhihang Yu, Xinkun Yu, Guangze Li, Aobotao Dai, Shuigeng Zhou
NeurIPS2