EDBT 2026 Demo / reviewers in the wild / expert
Yadong Qu
dblp:80/1926
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0003-0265-5011ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Igd: Instructional Graphic Design With Multimodal Layer Generatio
Yadong Qu, Hongtao Xie 0001, Yongdong Zhang 0001, Shancheng Fang, Yuxin Wang 0002, Zhineng Chen |
ICCV | 1 |
| 2025 | IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware GenerationabstractMeme creation is a creative process that blends images and text. However, existing methods lack critical components, failing to support intent-driven caption-layout generation and personalized generation, making it difficult to generate high-quality memes. To address this limitation, we propose IterMeme, an end-to-end interactive meme creation framework that utilizes a unified Multimodal Large Language Model (MLLM) to facilitate seamless collaboration among multiple components. To overcome the absence of a caption-layout generation component, we develop a robust layout representation method and construct a large-scale image-caption-layout dataset, MemeCap, which enhances the model’s ability to comprehend emotions and coordinate caption-layout generation effectively. To address the lack of a personalization component, we introduce a parameter-shared dual-LLM architecture that decouples the intricate representations of reference images and text. Furthermore, we incorporate the expert-guided M³OE for fine-grained identity properties (IP) feature extraction and cross-modal fusion. By dynamically injecting features into every layer of the model, we enable adaptive refinement of both visual and semantic information. Experimental results demonstrate that IterMeme significantly advances the field of meme creation by delivering consistently high-quality outcomes. The code, model, and dataset will be open-sourced to the community. Yaqi Cai, Shancheng Fang, Yadong Qu, Meng Shao, Hongtao Xie 0001 |
IJCAI | 3 |
| 2025 | Masked Text Pre-Training for Scene Text Detection
Hongtao Xie 0001, Keran Wang, Bangbang Zhou, Yuxin Wang 0002, Weigang Qi, Yadong Qu, Zuan Gao, Dongming Zhang 0004 |
IEEE Trans. Multim. | 7 |
| 2024 | Leveraging Text Localization for Scene Text Removal via Text-Aware Masked Image Modeling
Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Yadong Qu, Fengjun Guo, Pengwei Liu |
ECCV (66) | 4 |
| 2024 | Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
Zuan Gao, Yuxin Wang 0002, Yadong Qu, Boqiang Zhang, Zixiao Wang 0002, Hongtao Xie 0001 |
IJCAI | 3 |
| 2024 | Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
Bangbang Zhou, Yadong Qu, Zixiao Wang 0002, Boqiang Zhang, Hongtao Xie 0001 |
IJCAI | 2 |
| 2024 | LATextSpotter: Empowering Transformer Decoder with Length Perception AbilityabstractScene text spotting aims to integrate scene text detection and recognition into a unified framework. The existing transformer-based methods lack fine-grained positional information and linguistic information, limiting the convergence and performance of the model. In this paper, we propose a Length-Awear Text Spotter (LATextSpotter) to alleviate this problem by explicitly introducing two types of prior knowledge. First, the location of each character is initialized by coarsely locating the text instance and predicting the length, which provides effective guidance for the subsequent position-sensitive decoder. It is worth noting that the model requires only word-level supervision to achieve decent performance in the absence of expensive character-level annotations. Second, we design a mask prediction strategy based on the length information that masks character information at the feature level, and guides the model to predict the missing part. It empowers the decoder with language modeling capability without introducing extra modules. Additionally, considering the coordination between each module, a multi-stage training strategy is proposed to optimize the convergence process. Quantitative experiments demonstrate that LATextSpotter achieves the optimal end-to-end performance on arbitrary-shaped benchmarks by 76.6% and competitive spotting performance on multi-oriented datasets. Yadong Qu, Hongtao Xie 0001, Yongdong Zhang 0001 |
ISCAS | 2 |
| 2024 | Boosting Semi-Supervised Scene Text Recognition via Viewing and SummarizingabstractExisting scene text recognition (STR) methods struggle to recognize challenging texts, especially for artistic and severely distorted characters. The limitation lies in the insufficient exploration of character morphologies, including the monotonousness of widely used synthetic training data and the sensitivity of the model to character morphologies. To address these issues, inspired by the human learning process of viewing and summarizing, we facilitate the contrastive learning-based STR framework in a self-motivated manner by leveraging synthetic and real unlabeled data without any human cost. In the viewing process, to compensate for the simplicity of synthetic data and enrich character morphology diversity, we propose an Online Generation Strategy to generate background-free samples with diverse character styles. By excluding background noise distractions, the model is encouraged to focus on character morphology and generalize the ability to recognize complex samples when trained with only simple synthetic data. To boost the summarizing process, we theoretically demonstrate the derivation error in the previous character contrastive loss, which mistakenly causes the sparsity in the intra-class distribution and exacerbates ambiguity on challenging samples. Therefore, a new Character Unidirectional Alignment Loss is proposed to correct this error and unify the representation of the same characters in all samples by aligning the character features in the student model with the reference features in the teacher model. Extensive experiment results show that our method achieves SOTA performance (94.7\% and 70.9\% average accuracy on common benchmarks and Union14M-Benchmark). Code will be available. Yadong Qu, Bangbang Zhou, Hongtao Xie 0001, Yongdong Zhang 0001 |
NeurIPS | 1 |
| 2024 | How Control Information Influences Multilingual Text Image Generation and Editing?abstractVisual text generation has significantly advanced through diffusion models aimed at producing images with readable and realistic text. Recent works primarily use a ControlNet-based framework, employing standard font text images to control diffusion models. Recognizing the critical role of control information in generating high-quality text, we investigate its influence from three perspectives: input encoding, role at different stages, and output features. Our findings reveal that: 1) Input control information has unique characteristics compared to conventional inputs like Canny edges and depth maps. 2) Control information plays distinct roles at different stages of the denoising process. 3) Output control features significantly differ from the base and skip features of the U-Net decoder in the frequency domain. Based on these insights, we propose TextGen, a novel framework designed to enhance generation quality by optimizing control information. We improve input and output features using Fourier analysis to emphasize relevant information and reduce noise. Additionally, we employ a two-stage generation framework to align the different roles of control information at different stages. Furthermore, we introduce an effective and lightweight dataset for training. Our method achieves state-of-the-art performance in both Chinese and English text generation. The code and dataset are available at https://github.com/CyrilSterling/TextGen. Boqiang Zhang, Zuan Gao, Yadong Qu, Hongtao Xie 0001 |
NeurIPS | 3 |
| 2023 | Exploring Stroke-Level Modifications for Scene Text EditingabstractScene text editing (STE) aims to replace text with the desired one while preserving background and styles of the original text. However, due to the complicated background textures and various text styles, existing methods fall short in generating clear and legible edited text images. In this study, we attribute the poor editing performance to two problems: 1) Implicit decoupling structure. Previous methods of editing the whole image have to learn different translation rules of background and text regions simultaneously. 2) Domain gap. Due to the lack of edited real scene text images, the network can only be well trained on synthetic pairs and performs poorly on real-world images. To handle the above problems, we propose a novel network by MOdifying Scene Text image at strokE Level (MOSTEL). Firstly, we generate stroke guidance maps to explicitly indicate regions to be edited. Different from the implicit one by directly modifying all the pixels at image level, such explicit instructions filter out the distractions from background and guide the network to focus on editing rules of text regions. Secondly, we propose a Semi-supervised Hybrid Learning to train the network with both labeled synthetic images and unpaired real scene text images. Thus, the STE model is adapted to real-world datasets distributions. Moreover, two new datasets (Tamper-Syn2k and Tamper-Scene) are proposed to fill the blank of public evaluation datasets. Extensive experiments demonstrate that our MOSTEL outperforms previous methods both qualitatively and quantitatively. Datasets and code will be available at https://github.com/qqqyd/MOSTEL. Yadong Qu, Qingfeng Tan, Hongtao Xie 0001, Yuxin Wang 0002, Yongdong Zhang 0001 |
AAAI | 1 |
| 2023 | Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text DetectionabstractScene text detection has made great progress recently with the wide use of pre-training. Nonetheless, existing scene text detection methods still suffer from two problems: 1) Limited annotated real data reduces the feature robustness. 2) Detectors perform poorly on text lacking of visual information. In this paper, we explore the potential of the CLIP model, and propose a novel self-supervised Masked Text Modeling (MTM) pre-training method for scene text detection, which can be trained with unlabeled data and improve the linguistic reasoning ability for text occlusion. Different from previous randomly pixel-level masking methods, MTM performs a targeted text-aware masking process under an unsupervised manner. Specifically, MTM consists of text perception and masked text modeling. In the text perception step, benefiting from the text-friendliness of CLIP, a Text Perception Module is proposed to attend to text area by computing the similarity between the text and image tokens from CLIP model. In the masked text modeling step, a Text-aware Masking Strategy is designed to mask the text area, and the Masked Text Modeling Module is used to reconstruct the masked texts. MTM obtains the ability to reason the linguistic information of masked texts with the reconstruction. This robust feature extraction learned by MTM ensures a more discriminative representation for the text lacking of visual information. Moreover, a new text dataset named OcclusionText is proposed to evaluate the robustness for text occlusion of detection methods. Extensive experiments on public benchmarks demonstrate that our MTM can boost the performance of existing text detectors. Keran Wang, Hongtao Xie 0001, Yuxin Wang 0002, Dongming Zhang 0004, Yadong Qu, Zuan Gao, Yongdong Zhang 0001 |
ACM Multimedia | 5 |
| 2023 | What is the Real Need for Scene Text Removal? Exploring the Background Integrity and Erasure Exhaustivity PropertiesabstractAs a crucial application in privacy protection, scene text removal (STR) has received amounts of attention in recent years. However, existing approaches coarsely erasing texts from images ignore two important properties: the background texture integrity (BI) and the text erasure exhaustivity (EE). These two properties directly determine the erasure performance, and how to maintain them in a single network is the core problem for STR task. In this paper, we attribute the lack of BI and EE properties to the implicit erasure guidance and imbalanced multi-stage erasure respectively. To improve these two properties, we propose a new ProgrEssively Region-based scene Text eraser (PERT). There are three key contributions in our study. First, a novel explicit erasure guidance is proposed to enhance the BI property. Different from implicit erasure guidance modifying all the pixels in the entire image, our explicit one accurately performs stroke-level modification with only bounding-box level annotations. Second, a new balanced multi-stage erasure is constructed to improve the EE property. By balancing the learning difficulty and network structure among progressive stages, each stage takes an equal step towards the text-erased image to ensure the erasure exhaustivity. Third, we propose two new evaluation metrics called BI-metric and EE-metric, which make up the shortcomings of current evaluation tools in analyzing BI and EE properties. Compared with previous methods, PERT outperforms them by a large margin in both BI-metric ( ↑ 6.13 %) and EE-metric ( ↑ 1.9 %), obtaining SOTA results with high speed (71 FPS) and at least 25% lower parameter complexity. Code will be available at https://github.com/wangyuxin87/PERT. Yuxin Wang 0002, Hongtao Xie 0001, Zixiao Wang 0002, Yadong Qu, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | ADNet: Rethinking the Shrunk Polygon-Based Approach in Scene Text DetectionabstractTo localize text regions and separate close instances, the shrunk polygon is widely used in recent scene text detection methods. However, there exist two problems: 1) Existing methods fail to consider the aspect ratio sensitive problem when reconstructing the text instance from shrunk polygon. 2) Texts with extreme aspect ratios will lead to the fracture of shrunk polygons. To handle these two problems, in this paper, we propose a novel Adaptive Dilation Network (ADNet) to focus on the reconstruction process from shrunk polygon, which aims to provide a tight and complete text representation. Firstly, instead of using a fixed dilation factor, ADNet uses an aspect ratio-wise dilation factor to reconstruct the text region from shrunk polygon for each text instance. Such an instance-wise dilation factor considers the scale correlation between the original and shrunk polygon, and thus can guide an adaptive text region reconstruction for texts with large aspect ratio variance. Secondly, to deal with the fracture of detection results, a new Efficient Spatial Relationship Module (ESRM) is devised to capture long-range dependencies with low computation cost. ESRM uses a novel Weighted Pooling to reduce the resolution of feature maps without much information loss. Compared with the existing methods, ADNet further explores the potential of shrunk polygon-based approaches and obtains excellent detection results at an impressive speed. Extensive experiments on several datasets (Total-Text, CTW1500, MSRA-TD500 and ICDAR2015) verify the superiority of our method. Yadong Qu, Hongtao Xie 0001, Shancheng Fang, Yuxin Wang 0002, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2002 | Efficient Tree-Based Multicast in Wormhole-Routed Networks
Jianping Song, Zifeng Hou, Yadong Qu |
HiPC | 3 |