Zuan Gao

dblp:359/4127 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0009-0001-0587-4167ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Representation and self-supervised learning · 40% Image recognition and object detection · 31% Vision and language · 27%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 72% Multimedia analysis and retrieval · 28%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
scene text detection
1.522025
Masked Text Pre-Training for Scene Text Detection · IEEE Trans. Multim. 2025
Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection · ACM Multimedia 2023
Computer vision › Image recognition and object detection
scene text recognition
1.522024
Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition · IJCAI 2024
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
1.422024
Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition · IJCAI 2024
Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection · ACM Multimedia 2023
Computer vision › Vision and language
image captioning
0.912025
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis · CVPR 2025
Computer vision › Vision and language › multimodal understanding
multimodal table understanding
0.912025
SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis · CVPR 2025
Multimedia analysis and retrieval
multimodal evaluation
0.912025
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.812024
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
style-content disentanglement
0.812024
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024
Visual content generation and editing › visual text generation
scene text editing
0.812024
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024
Visual content generation and editing › image editing
text image editing
0.812024
How Control Information Influences Multilingual Text Image Generation and Editing? · NeurIPS 2024
Visual content generation and editing
visual text generation
0.812024
How Control Information Influences Multilingual Text Image Generation and Editing? · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling
0.712023
Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection · ACM Multimedia 2023
Data integration and cleaning › data generation
synthetic data generation
0.312025
SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis · CVPR 2025
Machine learning › Generative modeling
diffusion model
0.212024
How Control Information Influences Multilingual Text Image Generation and Editing? · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

table image rendering · 1.7question-answer generation · 1.7precision and hit metrics · 1.7QA-based evaluation · 1.7LLM-based synthesis · 1.7image decoder · 1.5controlnet · 1.5alignment loss · 1.5masked pre-training · 0.9two-stage generation · 0.8symmetric superimposition modeling · 0.8fourier analysis · 0.8
YearPublicationVenuePosition
2025 SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis
abstract
Due to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly. To address the above issues, we design a simple yet effective synthesis framework that consists of two independent steps: table image rendering and table question and answer (Q&A) pairs generation. We use table codes (HTML, LaTeX, Markdown) to synthesize images and generate Q&A pairs with large language model (LLM). This approach leverages LLMs high concurrency and low cost to boost annotation efficiency and reduce expenses. By inputting code instead of images, LLMs can directly access the content and structure of the table, reducing hallucinations in table understanding and improving the accuracy of generated Q&A pairs. Finally, we synthesize a large-scale MTU dataset, SynTab, containing 636K images and 1.8M samples costing within $200 in US dollars. We further introduce a generalist tabular multimodal model, SynTab-LLaVA. This model not only effectively extracts local textual content within the table but also enables global modeling of relationships between cells. SynTab-LLaVA achieves SOTA performance on 21 out of 24 in-domain and out-of-domain benchmarks, demonstrating the effectiveness and generalization of our method. The Code is available at SynTab-LLaVA.
Bangbang Zhou, Zuan Gao, Boqiang Zhang, Zhineng Chen, Hongtao Xie 0001
CVPR2
2025 CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
abstract
Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation, they remain limited to vague-view or object-view analyses and incomplete visual element coverage. In this paper, we introduce CAPability, a comprehensive multi-view benchmark for evaluating visual captioning across 12 dimensions spanning six critical views. We curate nearly 11K human-annotated images and videos with visual element annotations to evaluate the generated captions. CAPability stably assesses both the correctness and thoroughness of captions with \textit{precision} and \textit{hit} metrics. By converting annotations to QA pairs, we further introduce a heuristic metric, \textit{know but cannot tell} ($K\bar{T}$), indicating a significant performance gap between QA and caption capabilities. Our work provides a holistic analysis of MLLMs' captioning abilities, as we identify their strengths and weaknesses across various dimensions, guiding future research to enhance specific aspects of their capabilities.
Chen-Wei Xie, Feiwu Yu, Jixuan Chen, Pandeng Li, Boqiang Zhang, Nianzu Yang, Yinglu Li, Zuan Gao, Hongtao Xie 0001
NeurIPS10
2025 Masked Text Pre-Training for Scene Text Detection
Hongtao Xie 0001, Keran Wang, Bangbang Zhou, Yuxin Wang 0002, Weigang Qi, Yadong Qu, Zuan Gao, Dongming Zhang 0004
IEEE Trans. Multim.8
2024 Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
abstract
Scene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly coupled features for all tasks, resulting in sub-optimal performance. We propose a Disentangled Representation Learning framework (DARLING) aimed at disentangling these two types of features for improved adaptability in better addressing various downstream tasks (choose what you really need). Specifically, we synthesize a dataset of image pairs with identical style but different content. Based on the dataset, we decouple the two types of features by the supervision design. Clearly, we directly split the visual representation into style and content features, the content features are supervised by a text recognition loss, while an alignment loss aligns the style features in the image pairs. Then, style features are employed in reconstructing the counterpart image via an image decoder with a prompt that indicates the counterpart's content. Such an operation effectively decouples the features based on their distinctive properties. To the best of our knowledge, this is the first time in the field of scene text that disentangles the inherent properties of the text images. Our method achieves state-of-the-art performance in Scene Text Recognition, Removal, and Editing.
Boqiang Zhang, Hongtao Xie 0001, Zuan Gao
CVPR3
2024 Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
Zuan Gao, Yuxin Wang 0002, Yadong Qu, Boqiang Zhang, Zixiao Wang 0002, Hongtao Xie 0001
IJCAI1
2024 How Control Information Influences Multilingual Text Image Generation and Editing?
abstract
Visual text generation has significantly advanced through diffusion models aimed at producing images with readable and realistic text. Recent works primarily use a ControlNet-based framework, employing standard font text images to control diffusion models. Recognizing the critical role of control information in generating high-quality text, we investigate its influence from three perspectives: input encoding, role at different stages, and output features. Our findings reveal that: 1) Input control information has unique characteristics compared to conventional inputs like Canny edges and depth maps. 2) Control information plays distinct roles at different stages of the denoising process. 3) Output control features significantly differ from the base and skip features of the U-Net decoder in the frequency domain. Based on these insights, we propose TextGen, a novel framework designed to enhance generation quality by optimizing control information. We improve input and output features using Fourier analysis to emphasize relevant information and reduce noise. Additionally, we employ a two-stage generation framework to align the different roles of control information at different stages. Furthermore, we introduce an effective and lightweight dataset for training. Our method achieves state-of-the-art performance in both Chinese and English text generation. The code and dataset are available at https://github.com/CyrilSterling/TextGen.
Boqiang Zhang, Zuan Gao, Yadong Qu, Hongtao Xie 0001
NeurIPS2
2023 Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection
abstract
Scene text detection has made great progress recently with the wide use of pre-training. Nonetheless, existing scene text detection methods still suffer from two problems: 1) Limited annotated real data reduces the feature robustness. 2) Detectors perform poorly on text lacking of visual information. In this paper, we explore the potential of the CLIP model, and propose a novel self-supervised Masked Text Modeling (MTM) pre-training method for scene text detection, which can be trained with unlabeled data and improve the linguistic reasoning ability for text occlusion. Different from previous randomly pixel-level masking methods, MTM performs a targeted text-aware masking process under an unsupervised manner. Specifically, MTM consists of text perception and masked text modeling. In the text perception step, benefiting from the text-friendliness of CLIP, a Text Perception Module is proposed to attend to text area by computing the similarity between the text and image tokens from CLIP model. In the masked text modeling step, a Text-aware Masking Strategy is designed to mask the text area, and the Masked Text Modeling Module is used to reconstruct the masked texts. MTM obtains the ability to reason the linguistic information of masked texts with the reconstruction. This robust feature extraction learned by MTM ensures a more discriminative representation for the text lacking of visual information. Moreover, a new text dataset named OcclusionText is proposed to evaluate the robustness for text occlusion of detection methods. Extensive experiments on public benchmarks demonstrate that our MTM can boost the performance of existing text detectors.
Keran Wang, Hongtao Xie 0001, Yuxin Wang 0002, Dongming Zhang 0004, Yadong Qu, Zuan Gao, Yongdong Zhang 0001
ACM Multimedia6