Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dongsheng Xu 0001

dblp:56/7948-1 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0003-2481-5669ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 66% Segmentation and scene understanding · 24% Language models and text generation · 10%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
instance segmentation
0.812024
Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point · AAAI 2024
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.812024
Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point · AAAI 2024
Computer vision › Segmentation and scene understanding
referring image segmentation
0.812024
Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point · AAAI 2024
Computer vision › Vision and language
image captioning
0.712023
Zero-TextCap: Zero-shot Framework for Text-based Image Captioning · ACM Multimedia 2023
Natural language and speech › Language models and text generation
masked language modeling
0.712023
Zero-TextCap: Zero-shot Framework for Text-based Image Captioning · ACM Multimedia 2023
Computer vision › Vision and language
multimodal reasoning
0.712023
Scene-text Oriented Visual Entailment: Task, Dataset and Solution · ACM Multimedia 2023
Computer vision › Vision and language › multimodal understanding
scene text understanding
0.712023
Scene-text Oriented Visual Entailment: Task, Dataset and Solution · ACM Multimedia 2023
Computer vision › Vision and language › image captioning
text-based image captioning
0.712023
Zero-TextCap: Zero-shot Framework for Text-based Image Captioning · ACM Multimedia 2023
Computer vision › Vision and language › visual reasoning
visual entailment
0.712023
Scene-text Oriented Visual Entailment: Task, Dataset and Solution · ACM Multimedia 2023
Computer vision › Vision and language
vision-language dataset
0.212023
Scene-text Oriented Visual Entailment: Task, Dataset and Solution · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

soft ground-truth · 0.8point-based cross-modal comprehension · 0.8binary classification · 0.8multimodal transformer · 0.7masked language model · 0.7hybrid sampling · 0.7OCR · 0.7CLIP-based generation guidance · 0.7
YearPublicationVenuePosition
2025 Visual primitives as words: Alignment and interaction for compositional zero-shot learning
Feng Shuang 0002, Jiahuan Li, Qingbao Huang, Wenye Zhao, Dongsheng Xu 0001, Haonan Cheng
Pattern Recognit.5
2025 DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based image captioning
Dongsheng Xu 0001, Qingbao Huang, Xingmao Zhang, Haonan Cheng, Yi Cai 0001
Pattern Recognit.1
2024 Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point
abstract
As a fundamental and challenging task in the vision and language domain, Referring Expression Comprehension (REC) has shown impressive improvements recently. However, for a complex task that couples the comprehension of abstract concepts and the localization of concrete instances, one-stage approaches are bottlenecked by computing and data resources. To obtain a low-cost solution, the prevailing two-stage approaches decouple REC into localization (region proposal) and comprehension (region-expression matching) at region-level, but the solution based on isolated regions cannot sufficiently utilize the context and is usually limited by the quality of proposals. Therefore, it is necessary to rebuild an efficient two-stage solution system. In this paper, we propose a point-based two-stage framework for REC, in which the two stages are redefined as point-based cross-modal comprehension and point-based instance localization. Specifically, we reconstruct the raw bounding box and segmentation mask into center and mass scores as soft ground-truth for measuring point-level cross-modal correlations. With the soft ground-truth, REC can be approximated as a binary classification problem, which fundamentally avoids the impact of isolated regions on the optimization process. Remarkably, the consistent metrics between center and mass scores allow our system to directly optimize grounding and segmentation by utilizing the same architecture. Experiments on multiple benchmarks show the feasibility and potential of our point-based paradigm. Our code available at https://github.com/VILAN-Lab/PBREC-MT.
Peizhi Zhao, Shiyi Zheng, Wenye Zhao, Dongsheng Xu 0001, Pijian Li, Yi Cai 0001, Qingbao Huang
AAAI4
2023 Scene-text Oriented Visual Entailment: Task, Dataset and Solution
abstract
Visual Entailment (VE) is a fine-grained reasoning task aiming to predict whether the image semantically entails a hypothesis in textual form.Existing studies of VE only focus on basic visual attributes but largely overlook the importance of scene text, which usually entails rich semantic information and crucial clues (e.g., time, place, affiliation, and topic), leading to superficial design of hypothesis or incorrect entailment prediction. To fill this gap, we propose a new task called scene-text oriented Visual Entailment (STOVE), which requires models to predict whether an image semantically entails the corresponding hypothesis designed based on the scene text-centered visual information.STOVE task challenges a model to deeply understand the interplay between language and images containing scene text, requiring aligning hypotheses tokens, scene text, and visual contents.To support the researches on STOVE, we further collect a dataset termed TextVE, consisting of 23,864 images and 47,728 hypotheses related to scene text, which is constructed with the strategy of minimizing biases.Additionally, we present a baseline named MMTVE applying a multimodal transformer to model the spatial, semantic, and visual reasoning relations between multiple scene text tokens, hypotheses, and visual features.Experimental results illustrate that our model is effective in comprehending STOVE and achieves outstanding performance.Our codes are available at https://github.com/VISLANG-Lab/TextVE.
Nan Li 0055, Pijian Li, Dongsheng Xu 0001, Wenye Zhao, Yi Cai 0001, Qingbao Huang
ACM Multimedia3
2023 Zero-TextCap: Zero-shot Framework for Text-based Image Captioning
abstract
Text-based image captioning is a vital but under-explored task, which aims to describe images by captions containing scene text automatically. Recent studies have made encouraging progress, but they are still suffering from two issues. Firstly, current models cannot capture and generate scene text in non-Latin script languages, which severely limits the objectivity and the information completeness of generated captions. Secondly, current models tend to describe images with monotonous and templated style, which greatly limits the diversity of the generated captions. Although the above-mentioned issues can be alleviated through carefully designed annotations, this process is undoubtedly laborious and time-consuming. To address the above issues, we propose a Zero-shot Framework for Text-based Image Captioning (Zero-TextCap). Concretely, to generate candidate sentences starting from the prompt 'Image of' and iteratively refine them to improve the quality and diversity of captions, we introduce a Hybrid-sampling masked language model (H-MLM). To read multi-lingual scene text and model the relationships between them, we introduce a robust OCR system. To ensure that the captions generated by H-MLM contain scene text and are highly relevant to the image, we propose a CLIP-based generation guidance module to insert OCR tokens and filter candidate sentences. Our Zero-TextCap is capable of generalizing captions containing multi-lingual scene text and boosting the diversity of captions. Sufficient experiments demonstrate the effectiveness of our proposed Zero-TextCap. Our codes are available at https://github.com/Gemhuang79/Zero_TextCap.
Dongsheng Xu 0001, Wenye Zhao, Yi Cai 0001, Qingbao Huang
ACM Multimedia1