EDBT 2026 Demo / reviewers in the wild / expert
Zuan Gao
dblp:359/4127
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0009-0001-0587-4167ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Representation and self-supervised learning · 40% Image recognition and object detection · 31% Vision and language · 27% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 72% Multimedia analysis and retrieval · 28% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
scene text detection |
1.5 | 2 | 2025 | Masked Text Pre-Training for Scene Text Detection · IEEE Trans. Multim. 2025 Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection · ACM Multimedia 2023 |
Computer vision › Image recognition and object detection
scene text recognition |
1.5 | 2 | 2024 | Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition · IJCAI 2024 Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
1.4 | 2 | 2024 | Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition · IJCAI 2024 Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection · ACM Multimedia 2023 |
Computer vision › Vision and language
image captioning |
0.9 | 1 | 2025 | CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis · CVPR 2025 |
Computer vision › Vision and language › multimodal understanding
multimodal table understanding |
0.9 | 1 | 2025 | SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis · CVPR 2025 |
Multimedia analysis and retrieval
multimodal evaluation |
0.9 | 1 | 2025 | CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.8 | 1 | 2024 | Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
style-content disentanglement |
0.8 | 1 | 2024 | Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024 |
Visual content generation and editing › visual text generation
scene text editing |
0.8 | 1 | 2024 | Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing · CVPR 2024 |
Visual content generation and editing › image editing
text image editing |
0.8 | 1 | 2024 | How Control Information Influences Multilingual Text Image Generation and Editing? · NeurIPS 2024 |
Visual content generation and editing
visual text generation |
0.8 | 1 | 2024 | How Control Information Influences Multilingual Text Image Generation and Editing? · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling |
0.7 | 1 | 2023 | Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection · ACM Multimedia 2023 |
Data integration and cleaning › data generation
synthetic data generation |
0.3 | 1 | 2025 | SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis · CVPR 2025 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | How Control Information Influences Multilingual Text Image Generation and Editing? · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
table image rendering · 1.7question-answer generation · 1.7precision and hit metrics · 1.7QA-based evaluation · 1.7LLM-based synthesis · 1.7image decoder · 1.5controlnet · 1.5alignment loss · 1.5masked pre-training · 0.9two-stage generation · 0.8symmetric superimposition modeling · 0.8fourier analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled SynthesisabstractDue to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly. To address the above issues, we design a simple yet effective synthesis framework that consists of two independent steps: table image rendering and table question and answer (Q&A) pairs generation. We use table codes (HTML, LaTeX, Markdown) to synthesize images and generate Q&A pairs with large language model (LLM). This approach leverages LLMs high concurrency and low cost to boost annotation efficiency and reduce expenses. By inputting code instead of images, LLMs can directly access the content and structure of the table, reducing hallucinations in table understanding and improving the accuracy of generated Q&A pairs. Finally, we synthesize a large-scale MTU dataset, SynTab, containing 636K images and 1.8M samples costing within $200 in US dollars. We further introduce a generalist tabular multimodal model, SynTab-LLaVA. This model not only effectively extracts local textual content within the table but also enables global modeling of relationships between cells. SynTab-LLaVA achieves SOTA performance on 21 out of 24 in-domain and out-of-domain benchmarks, demonstrating the effectiveness and generalization of our method. The Code is available at SynTab-LLaVA. Bangbang Zhou, Zuan Gao, Boqiang Zhang, Zhineng Chen, Hongtao Xie 0001 |
CVPR | 2 |
| 2025 | CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and ThoroughnessabstractVisual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation, they remain limited to vague-view or object-view analyses and incomplete visual element coverage. In this paper, we introduce CAPability, a comprehensive multi-view benchmark for evaluating visual captioning across 12 dimensions spanning six critical views. We curate nearly 11K human-annotated images and videos with visual element annotations to evaluate the generated captions. CAPability stably assesses both the correctness and thoroughness of captions with \textit{precision} and \textit{hit} metrics. By converting annotations to QA pairs, we further introduce a heuristic metric, \textit{know but cannot tell} ($K\bar{T}$), indicating a significant performance gap between QA and caption capabilities. Our work provides a holistic analysis of MLLMs' captioning abilities, as we identify their strengths and weaknesses across various dimensions, guiding future research to enhance specific aspects of their capabilities. Chen-Wei Xie, Feiwu Yu, Jixuan Chen, Pandeng Li, Boqiang Zhang, Nianzu Yang, Yinglu Li, Zuan Gao, Hongtao Xie 0001 |
NeurIPS | 10 |
| 2025 | Masked Text Pre-Training for Scene Text Detection
Hongtao Xie 0001, Keran Wang, Bangbang Zhou, Yuxin Wang 0002, Weigang Qi, Yadong Qu, Zuan Gao, Dongming Zhang 0004 |
IEEE Trans. Multim. | 8 |
| 2024 | Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and EditingabstractScene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly coupled features for all tasks, resulting in sub-optimal performance. We propose a Disentangled Representation Learning framework (DARLING) aimed at disentangling these two types of features for improved adaptability in better addressing various downstream tasks (choose what you really need). Specifically, we synthesize a dataset of image pairs with identical style but different content. Based on the dataset, we decouple the two types of features by the supervision design. Clearly, we directly split the visual representation into style and content features, the content features are supervised by a text recognition loss, while an alignment loss aligns the style features in the image pairs. Then, style features are employed in reconstructing the counterpart image via an image decoder with a prompt that indicates the counterpart's content. Such an operation effectively decouples the features based on their distinctive properties. To the best of our knowledge, this is the first time in the field of scene text that disentangles the inherent properties of the text images. Our method achieves state-of-the-art performance in Scene Text Recognition, Removal, and Editing. Boqiang Zhang, Hongtao Xie 0001, Zuan Gao |
CVPR | 3 |
| 2024 | Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
Zuan Gao, Yuxin Wang 0002, Yadong Qu, Boqiang Zhang, Zixiao Wang 0002, Hongtao Xie 0001 |
IJCAI | 1 |
| 2024 | How Control Information Influences Multilingual Text Image Generation and Editing?abstractVisual text generation has significantly advanced through diffusion models aimed at producing images with readable and realistic text. Recent works primarily use a ControlNet-based framework, employing standard font text images to control diffusion models. Recognizing the critical role of control information in generating high-quality text, we investigate its influence from three perspectives: input encoding, role at different stages, and output features. Our findings reveal that: 1) Input control information has unique characteristics compared to conventional inputs like Canny edges and depth maps. 2) Control information plays distinct roles at different stages of the denoising process. 3) Output control features significantly differ from the base and skip features of the U-Net decoder in the frequency domain. Based on these insights, we propose TextGen, a novel framework designed to enhance generation quality by optimizing control information. We improve input and output features using Fourier analysis to emphasize relevant information and reduce noise. Additionally, we employ a two-stage generation framework to align the different roles of control information at different stages. Furthermore, we introduce an effective and lightweight dataset for training. Our method achieves state-of-the-art performance in both Chinese and English text generation. The code and dataset are available at https://github.com/CyrilSterling/TextGen. Boqiang Zhang, Zuan Gao, Yadong Qu, Hongtao Xie 0001 |
NeurIPS | 2 |
| 2023 | Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text DetectionabstractScene text detection has made great progress recently with the wide use of pre-training. Nonetheless, existing scene text detection methods still suffer from two problems: 1) Limited annotated real data reduces the feature robustness. 2) Detectors perform poorly on text lacking of visual information. In this paper, we explore the potential of the CLIP model, and propose a novel self-supervised Masked Text Modeling (MTM) pre-training method for scene text detection, which can be trained with unlabeled data and improve the linguistic reasoning ability for text occlusion. Different from previous randomly pixel-level masking methods, MTM performs a targeted text-aware masking process under an unsupervised manner. Specifically, MTM consists of text perception and masked text modeling. In the text perception step, benefiting from the text-friendliness of CLIP, a Text Perception Module is proposed to attend to text area by computing the similarity between the text and image tokens from CLIP model. In the masked text modeling step, a Text-aware Masking Strategy is designed to mask the text area, and the Masked Text Modeling Module is used to reconstruct the masked texts. MTM obtains the ability to reason the linguistic information of masked texts with the reconstruction. This robust feature extraction learned by MTM ensures a more discriminative representation for the text lacking of visual information. Moreover, a new text dataset named OcclusionText is proposed to evaluate the robustness for text occlusion of detection methods. Extensive experiments on public benchmarks demonstrate that our MTM can boost the performance of existing text detectors. Keran Wang, Hongtao Xie 0001, Yuxin Wang 0002, Dongming Zhang 0004, Yadong Qu, Zuan Gao, Yongdong Zhang 0001 |
ACM Multimedia | 6 |