Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zewen Gao

dblp:317/5092 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 33% Image recognition and object detection · 33% Knowledge representation and reasoning · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.612022
Knowledge Mining with Scene Text for Fine-Grained Recognition · CVPR 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge discovery
0.612022
Knowledge Mining with Scene Text for Fine-Grained Recognition · CVPR 2022
Computer vision › Vision and language › multimodal understanding
scene text understanding
0.612022
Knowledge Mining with Scene Text for Fine-Grained Recognition · CVPR 2022

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 0.6knowbert · 0.6
YearPublicationVenuePosition
2022 Knowledge Mining with Scene Text for Fine-Grained Recognition
abstract
Recently, the semantics of scene text has been proven to be essential in fine-grained image classification. However, the existing methods mainly exploit the literal meaning of scene text for fine-grained recognition, which might be irrelevant when it is not significantly related to objects/scenes. We propose an end-to-end trainable network that mines implicit contextual knowledge behind scene text image and enhance the semantics and correlation to fine-tune the image representation. Unlike the existing methods, our model integrates three modalities: visual feature extraction, text semantics extraction, and correlating background knowledge to fine-grained image classification. Specifically, we employ KnowBert to retrieve relevant knowledge for semantic representation and combine it with image features for fine-grained classification. Experiments on two benchmark datasets, Con-Text, and Drink Bottle, show that our method outperforms the state-of-the-art by 3.72% mAP and 5.39% mAp, respectively. To further validate the effectiveness of the proposed method, we create a new dataset on crowd activity recognition for the evaluation. The source code and new dataset of this work are available at this repository11https://github.com/lanfeng4659/KnowledgeMiningWithSceneText.
Hao Wang 0207, Junchao Liao, Tianheng Cheng, Zewen Gao, Hao Liu 0003, Bo Ren 0002, Xiang Bai, Wenyu Liu 0001
CVPR4