VLDB 2026 Research / reviewers in the wild / expert
Shancheng Fang
dblp:201/0063
· DBLP profile ↗
31ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-3100-3664ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 13 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with CitationsabstractGenerating with citations is crucial for trustworthy Large Language Models (LLMs), yet even advanced LLMs often produce mismatched or irrelevant citations. Existing methods over-optimize citation fidelity while overlooking relevance to the user query, which degrades answer quality and robustness in real-world settings with noisy or irrelevant retrieved content. Moreover, the prevailing single-pass paradigm struggles to deliver optimal answers in long-form generation that requiring multiple citations. To address these limitations, we propose FineRef, a framework based on Fine-grained error Reflection, which explicitly teaches the model to self-identify and correct two key citation errors—mismatch and irrelevance—on a per-citation basis. FineRef follows a two-stage training strategy. The first stage instills an “attempt–reflect–correct” behavioral pattern via supervised fine-tuning, using fine-grained and controllable reflection data constructed by specialized lightweight models. An online self-reflective bootstrapping strategy is designed to improve generalization by iteratively enriching training data with verified, self-improving examples. To further enhance the self-reflection and correction capability, the second stage applies process-level reinforcement learning with a multi-dimensional reward scheme that promotes reflection accuracy, answer quality, and correction gain. Experiments on the ALCE benchmark demonstrate that FineRef significantly improves both citation performance and answer accuracy. Our 7B model outperforms GPT-4 by up to 18% in Citation F1 and 4% in EM Recall, while also surpassing the state-of-the-art model across key evaluation metrics. FineRef also exhibits strong generalization and robustness in domain transfer settings and noisy retrieval scenarios. Yixing Peng, Licheng Zhang 0002, Shancheng Fang, Yi Liu 0148, Peijian Gu, Quan Wang 0002 |
AAAI | 3 |
| 2025 | Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video GenerationabstractSora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask2DiT, a novel approach that establishes fine-grained, one-to-one alignment between video segments and their corresponding text annotations. Specifically, we introduce a symmetric binary mask at each attention layer within the DiT architecture, ensuring that each text annotation applies exclusively to its respective video segment while preserving temporal coherence across visual tokens. This attention mechanism enables precise segment-level textual-to-visual alignment, allowing the DiT architecture to effectively handle video generation tasks with a fixed number of scenes. To further equip the DiT architecture with the ability to generate additional scenes based on existing ones, we incorporate a segment-level conditional mask, which conditions each newly generated segment on the preceding video segments, thereby enabling auto-regressive scene extension. Both qualitative and quantitative experiments confirm that Mask2DiT excels in maintaining visual consistency across segments while ensuring semantic alignment between each segment and its corresponding text description. Our project page is https://tianhao-qi.github.io/Mask2DiTProject/. Tianhao Qi, Jianlong Yuan, Wanquan Feng, Shancheng Fang, Jiawei Liu 0001, SiYu Zhou 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
CVPR | 4 |
| 2025 | Igd: Instructional Graphic Design With Multimodal Layer Generatio
Yadong Qu, Hongtao Xie 0001, Yongdong Zhang 0001, Shancheng Fang, Yuxin Wang 0002, Zhineng Chen |
ICCV | 4 |
| 2025 | IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware GenerationabstractMeme creation is a creative process that blends images and text. However, existing methods lack critical components, failing to support intent-driven caption-layout generation and personalized generation, making it difficult to generate high-quality memes. To address this limitation, we propose IterMeme, an end-to-end interactive meme creation framework that utilizes a unified Multimodal Large Language Model (MLLM) to facilitate seamless collaboration among multiple components. To overcome the absence of a caption-layout generation component, we develop a robust layout representation method and construct a large-scale image-caption-layout dataset, MemeCap, which enhances the model’s ability to comprehend emotions and coordinate caption-layout generation effectively. To address the lack of a personalization component, we introduce a parameter-shared dual-LLM architecture that decouples the intricate representations of reference images and text. Furthermore, we incorporate the expert-guided M³OE for fine-grained identity properties (IP) feature extraction and cross-modal fusion. By dynamically injecting features into every layer of the model, we enable adaptive refinement of both visual and semantic information. Experimental results demonstrate that IterMeme significantly advances the field of meme creation by delivering consistently high-quality outcomes. The code, model, and dataset will be open-sourced to the community. Yaqi Cai, Shancheng Fang, Yadong Qu, Meng Shao, Hongtao Xie 0001 |
IJCAI | 2 |
| 2025 | GRIP: A Graph-Based Reasoning Instruction ProducerabstractLarge-scale, high-quality data is essential for advancing the reasoning capabilities of large language models (LLMs). As publicly available Internet data becomes increasingly scarce, synthetic data has emerged as a crucial research direction. However, existing data synthesis methods often suffer from limited scalability, insufficient sample diversity, and a tendency to overfit to seed data, which constrains their practical utility. In this paper, we present \textit{\textbf{GRIP}}, a \textbf{G}raph-based \textbf{R}easoning \textbf{I}nstruction \textbf{P}roducer that efficiently synthesizes high-quality and diverse reasoning instructions. \textit{GRIP} constructs a knowledge graph by extracting high-level concepts from seed data, and uniquely leverages both explicit and implicit relationships within the graph to drive large-scale and diverse instruction data synthesis, while employing open-source multi-model supervision to ensure data quality. We apply \textit{GRIP} to the critical and challenging domain of mathematical reasoning. Starting from a seed set of 7.5K math reasoning samples, we construct \textbf{GRIP-MATH}, a dataset containing 2.1 million synthesized question-answer pairs. Compared to similar synthetic data methods, \textit{GRIP} achieves greater scalability and diversity while also significantly reducing costs. On mathematical reasoning benchmarks, models trained with GRIP-MATH demonstrate substantial improvements over their base models and significantly outperform previous data synthesis methods. Jiankang Wang, Yuxin Wang 0002, Mengting Xing, Shancheng Fang, Hongtao Xie 0001 |
NeurIPS | 6 |
| 2024 | DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image GenerationabstractWhile large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity and follow the text prompts simultaneously for conditioned input face images and texts. Despite existing encoder-based methods achieving high efficiency and decent face similarity, the generated image often fails to follow the textual prompts. To ease this editability issue, we present DreamIdentity, to learn edit-friendly and accurate face-identity representations in the word embedding space. Specifically, we propose self-augmented editability learning to enhance the editability for projected embedding, which is achieved by constructing paired generated celebrity's face and edited celebrity images for training, aiming at transferring mature editability of off-the-shelf text-to-image models in celebrity to unseen identities. Furthermore, we design a novel dedicated face-identity encoder to learn an accurate representation of human faces, which applies multi-scale ID-aware features followed by a multi-embedding projector to generate the pseudo words in the text embedding space directly. Extensive experiments show that our method can generate more text-coherent and ID-preserved images with negligible time overhead compared to the standard text-to-image generation process. Zhuowei Chen, Shancheng Fang, Mengqi Huang, Zhendong Mao 0001 |
AAAI | 2 |
| 2024 | DEADiff: An Efficient Stylization Diffusion Model with Disentangled RepresentationsabstractThe diffusion-based text-to-image model harbors im-mense potential in transferring reference style. However, current encoder-based approaches significantly impair the text controllability of text-to-image models while transfer-ring styles. In this paper, we introduce DEADiff to address this issue using the following two strategies: 1) a mecha-nism to decouple the style and semantics of reference images. The decoupled feature representations are first extracted by Q-Formers which are instructed by different text descriptions. Then they are injected into mutually exclusive subsets of cross-attention layers for better disentanglement. 2) A non-reconstructive learning method. The Q-Formers are trained using paired images rather than the identical target, in which the reference image and the ground-truth image are with the same style or semantics. We show that DEADiff attains the best visual stylization results and optimal balance between the text controllability inherent in the text-to-image model and style similarity to the reference image, as demonstrated both quantitatively and qualitatively. Our project page is https://tianhao-qi.github.io/DEADiff‘/. Tianhao Qi, Shancheng Fang, Yanze Wu, Hongtao Xie 0001, Jiawei Liu 0001, Lang Chen, Yongdong Zhang 0001 |
CVPR | 2 |
| 2024 | CDistNet: Perceiving Multi-domain Character Distance for Robust Text Recognition
Tianlun Zheng, Zhineng Chen, Shancheng Fang, Hongtao Xie 0001, Yu-Gang Jiang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Sentiment-Oriented Transformer-Based Variational Autoencoder Network for Live Video CommentingabstractAutomatic live video commenting is getting increasing attention due to its significance in narration generation, topic explanation, etc. However, the diverse sentiment consideration of the generated comments is missing from current methods. Sentimental factors are critical in interactive commenting, and there has been lack of research so far. Thus, in this article, we propose a Sentiment-oriented Transformer-based Variational Autoencoder (So-TVAE) network, which consists of a sentiment-oriented diversity encoder module and a batch attention module, to achieve diverse video commenting with multiple sentiments and multiple semantics. Specifically, our sentiment-oriented diversity encoder elegantly combines a VAE and random mask mechanism to achieve semantic diversity under sentiment guidance, which is then fused with cross-modal features to generate live video comments. A batch attention module is also proposed in this article to alleviate the problem of missing sentimental samples, caused by the data imbalance that is common in live videos as the popularity of videos varies. Extensive experiments on Livebot and VideoIC datasets demonstrate that the proposed So-TVAE outperforms the state-of-the-art methods in terms of the quality and diversity of generated comments. Related code is available at https://github.com/fufy1024/So-TVAE . Fengyi Fu, Shancheng Fang, Weidong Chen 0013, Zhendong Mao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Crossing the Gap: Domain Generalization for Image CaptioningabstractExisting image captioning methods are under the assumption that the training and testing data are from the same domain or that the data from the target domain (i.e., the domain that testing data lie in) are accessible. However, this assumption is invalid in real-world applications where the data from the target domain is inaccessible. In this paper, we introduce a new setting called Domain Generalization for Image Captioning (DGIC), where the data from the target domain is unseen in the learning process. We first construct a benchmark dataset for DGIC, which helps us to investigate models' domain generalization (DG) ability on unseen domains. With the support of the new benchmark, we further propose a new framework called language-guided semantic metric learning (LSML) for the DGIC setting. Experiments on multiple datasets demonstrate the challenge of the task and the effectiveness of our newly proposed benchmark and LSML framework. Yuchen Ren 0001, Zhendong Mao 0001, Shancheng Fang, Yan Lu 0001, Tong He 0001, Yongdong Zhang 0001, Wanli Ouyang |
CVPR | 3 |
| 2023 | ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text SpottingabstractScene text spotting is of great importance to the computer vision community due to its wide variety of applications. Recent methods attempt to introduce linguistic knowledge for challenging recognition rather than pure visual classification. However, how to effectively model the linguistic rules in end-to-end deep networks remains a research challenge. In this paper, we argue that the limited capacity of language models comes from 1) implicit language modeling; 2) unidirectional feature representation; and 3) language model with noise input. Correspondingly, we propose an autonomous, bidirectional and iterative ABINet++ for scene text spotting. First, the autonomous suggests enforcing explicitly language modeling by decoupling the recognizer into vision model and language model and blocking gradient flow between both models. Second, a novel bidirectional cloze network (BCN) as the language model is proposed based on bidirectional feature representation. Third, we propose an execution manner of iterative correction for the language model which can effectively alleviate the impact of noise input. Additionally, based on an ensemble of the iterative predictions, a self-training method is developed which can learn from unlabeled images effectively. Finally, to polish ABINet++ in long text recognition, we propose to aggregate horizontal features by embedding Transformer units inside a U-Net, and design a position and content attention module which integrates character order and content to attend to character features precisely. ABINet++ achieves state-of-the-art performance on both scene text recognition and scene text spotting benchmarks, which consistently demonstrates the superiority of our method in various environments especially on low-quality images. Besides, extensive experiments including in English and Chinese also prove that, a text spotter that incorporates our language modeling method can significantly improve its performance both in accuracy and speed compared with commonly used attention-based recognizers. Code is available at https://github.com/FangShancheng/ABINet-PP. Shancheng Fang, Zhendong Mao 0001, Hongtao Xie 0001, Yuxin Wang 0002, Chenggang Yan 0001, Yongdong Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | ADNet: Rethinking the Shrunk Polygon-Based Approach in Scene Text DetectionabstractTo localize text regions and separate close instances, the shrunk polygon is widely used in recent scene text detection methods. However, there exist two problems: 1) Existing methods fail to consider the aspect ratio sensitive problem when reconstructing the text instance from shrunk polygon. 2) Texts with extreme aspect ratios will lead to the fracture of shrunk polygons. To handle these two problems, in this paper, we propose a novel Adaptive Dilation Network (ADNet) to focus on the reconstruction process from shrunk polygon, which aims to provide a tight and complete text representation. Firstly, instead of using a fixed dilation factor, ADNet uses an aspect ratio-wise dilation factor to reconstruct the text region from shrunk polygon for each text instance. Such an instance-wise dilation factor considers the scale correlation between the original and shrunk polygon, and thus can guide an adaptive text region reconstruction for texts with large aspect ratio variance. Secondly, to deal with the fracture of detection results, a new Efficient Spatial Relationship Module (ESRM) is devised to capture long-range dependencies with low computation cost. ESRM uses a novel Weighted Pooling to reduce the resolution of feature maps without much information loss. Compared with the existing methods, ADNet further explores the potential of shrunk polygon-based approaches and obtains excellent detection results at an impressive speed. Extensive experiments on several datasets (Total-Text, CTW1500, MSRA-TD500 and ICDAR2015) verify the superiority of our method. Yadong Qu, Hongtao Xie 0001, Shancheng Fang, Yuxin Wang 0002, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Learning Pixel Affinity Pyramid for Arbitrary-Shaped Text DetectionabstractArbitrary-shaped text detection in natural images is a challenging task due to the complexity of the background and the diversity of text properties. The difficulty lies in two aspects: accurate separation of adjacent texts and sufficient text feature representation. To handle these problems, we consider text detection as instance segmentation and propose a novel text detection framework, which jointly learns semantic segmentation and a pixel affinity pyramid in a unified fully convolutional network. Specifically, the pixel affinity pyramid is proposed to encode multi-scale instance affiliation relationships of pixels, which is not only robust to varying shapes of text but also provides an accurate boundary description for separating closely located texts. In the inference phase, a simple but effective post-processing is presented to reconstruct text instances from the semantic segmentation results under the guidance of the learned pixel affinity pyramid, achieving good accuracy and efficiency. Furthermore, to enhance the representation of text features in the neural network, two modules — the Region Enhancement Module (REM) and Attentional Fusion Module (AFM) — are proposed. The REM models the semantic correlations of regional features to enhance the features from the text area, which effectively suppresses false-positive detection. The AFM adaptively fuses multi-scale textual information through an attention mechanism to obtain abundant text semantic features, which benefits multi-sized text detection. Extensive ablation experiments are conducted demonstrating the effectiveness of the REM and AFM. Evaluation results on standard benchmarks, including Total-Text, ICDAR2015, SCUT-CTW1500, and MSRA-TD500, show that our method surpasses most existing text detectors and achieves state-of-the-art performance, denoting its superior capability in detecting arbitrary-shaped texts. Zilong Fu, Hongtao Xie 0001, Shancheng Fang, Yuxin Wang 0002, Mengting Xing, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | ER-SAN: Enhanced-Adaptive Relation Self-Attention Network for Image CaptioningabstractImage captioning (IC), bringing vision to language, has drawn extensive attention. Precisely describing visual relations between image objects is a key challenge in IC. We argue that the visual relations, that is geometric positions (i.e., distance and size) and semantic interactions (i.e., actions and possessives), indicate the mutual correlations between objects. Existing Transformer-based methods typically resort to geometric positions to enhance the representation of visual relations, yet only using the shallow geometric is unable to precisely cover the complex and actional correlations. In this paper, we propose to enhance the correlations between objects from a comprehensive view that jointly considers explicit semantic and geometric relations, generating plausible captions with accurate relationship predictions. Specifically, we propose a novel Enhanced-Adaptive Relation Self-Attention Network (ER-SAN). We design the direction-sensitive semantic-enhanced attention, which considers content objects to semantic relations and semantic relations to content objects attention to learn explicit semantic-aware relations. Further, we devise an adaptive re-weight relation module that determines how much semantic and geometric attention should be activated to each relation feature. Extensive experiments on MS-COCO dataset demonstrate the effectiveness of our ER-SAN, with improvements of CIDEr from 128.6% to 135.3%, achieving state-of-the-art performance. Codes will be released \url{https://github.com/CrossmodalGroup/ER-SAN}. Zhendong Mao 0001, Shancheng Fang, Hao Li 0189 |
IJCAI | 3 |
| 2022 | Background Layout Generation and Object Knowledge Transfer for Text-to-Image GenerationabstractText-to-Image generation (T2I) aims to generate realistic and semantically consistent images according to the natural language descriptions. Built upon the recent advances in generative adversarial networks (GANs), existing T2I models have made great process. However, a close inspection of their generated images shows two major limitations: 1) the background (e.g., fence, lake) of the generated image with the complicated, real-world scene tends to be unrealistic; 2) the object (e.g., elephant, zebra) in the generated image often presents highly distorted shape or key parts missing. To address these limitations, we propose a two-stage T2I approach, where the first stage redesigns the text-to-layout process to incorporate the background layout with the existing object layout, the second stage transfers the object knowledge from an existing class-to-image model to the layout-to-image process to improve the object fidelity. Specifically, a transformer-based architecture is introduced as the layout generator to learn the mapping from text to layout of object and background, and a Text-attended Layout-aware feature Normalization (TL-Norm) is proposed to adaptively transfer the object knowledge to the image generation. Benefitting from the background layout and transferred object knowledge, the proposed approach significantly surpasses previous state-of-the-art methods in the image quality metric and achieves superior image-text alignment performance. Zhuowei Chen, Zhendong Mao 0001, Shancheng Fang, Bo Hu 0036 |
ACM Multimedia | 3 |
| 2022 | Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningabstractNews Image Captioning aims to generate descriptions for images embedded in news articles, including plentiful real-world concepts, especially about named entities. However, existing methods are limited in the entity-level template. Not only is it labor-intensive to craft the template, but it is error-prone due to local entity-aware, which solely constrains the prediction output at each language model decoding step with corrupted entity relationship. To overcome the problem, we investigate a concise and flexible paradigm to achieve global entity-aware by introducing a prompting mechanism with fine-tuning pre-trained models, named Fine-tuning with Multi-modal Entity Prompts for News Image Captioning (NewsMEP). Firstly, we incorporate two pre-trained models: (i) CLIP, translating the image with open-domain knowledge; (ii) BART, extended to encode article and image simultaneously. Moreover, leveraging the BART architecture, we can easily take the end-to-end fashion. Secondly, we prepend the target caption with two prompts to utilize entity-level lexical cohesion and inherent coherence in the pre-trained language model. Concretely, the visual prompts are obtained by mapping CLIP embeddings, and contextual vectors automatically construct the entity-oriented prompts. Thirdly, we provide an entity chain to control caption generation that focuses on entities of interest. Experiments results on two large-scale publicly available datasets, including detailed ablation studies, show that our NewsMEP not only outperforms state-of-the-art methods in general caption metrics but also achieves significant performance in precision and recall of various named entities. Jingjing Zhang 0007, Shancheng Fang, Zhendong Mao 0001, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Semi-Supervised Text Detection With Accurate Pseudo-LabelsabstractRecent scene text detection methods have made great progress. However, existing methods rely heavily on extensive labeled data, which is very time-consuming and expensive. In this letter, we propose a novel semi-supervised text detection method to alleviate the dependence of text detectors on labeled data by generating accurate pseudo-labels and performing effective data augmentations. Specifically, a dual-threshold pseudo-label generation algorithm is designed to divide the prediction results into background regions, text regions, and uncertain regions. The definition of uncertain regions obviously improves the accuracy of pseudo-labels. To obtain accurate pseudo-labels for text at various scales, we first design a scale-aware loss function to adaptively adjust the loss weight of different scale texts. Then, a multi-scale feature extraction module is proposed to extract multi-scale text features and adaptively weight these features according to the scale of the text. Moreover, effective data augmentations are explored to use unlabeled data to improve the robustness of the model to various texts. Experiments show that our method achieves state-of-the-art performance on several datasets(e.g., outperforms existing methods by 2.0% on TD500). Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Yongdong Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Semantically Similarity-Wise Dual-Branch Network for Scene Graph GenerationabstractScene graph generation aims to detect visual entities and relationships between them from an image. The object-level visual information is of vital importance for predicting accurate relationships. However, most existing methods essentially encode visual information with coarse supervised information, since they regard different relationships as mutually exclusive semantics with equal-distance labels by taking cross-entropy function as the main training loss. Intuitively, different relationship semantics naturally have their own similarity and dissimilarity with different level distances,i.e., the topological information of relationship semantics. It can serve as an inspiring hint to aid learning to grasp the key related visual information. Accordingly, we propose a Semantically Similarity-wise Dual-branch Network (SSDN) which introduces topological information of relationship semantics as extra supervision to aid learning extracting and encoding relationship-related visual information. To avoid possible chaotic feature learning and enable the introduced knowledge to be better absorbed during inference, we design a dual-branch framework consisting of an auxiliary branch and an inference branch. The topological information extracted from the groundtruth is introduced at the front end of the auxiliary branch which then generates a soft embedding to be propagated to the inference branch in a knowledge distillation manner. Extensive experiments show that our model averagely outperforms state-of-the-art approaches on benchmark Visual Genome and VRD significantly, which demonstrates its effectiveness and superiority. Zhendong Mao 0001, Shancheng Fang, Wenyu Zang, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | PETR: Rethinking the Capability of Transformer-Based Language Model in Scene Text RecognitionabstractThe exploration of linguistic information promotes the development of scene text recognition task. Benefiting from the significance in parallel reasoning and global relationship capture, transformer-based language model (TLM) has achieved dominant performance recently. As a decoupled structure from the recognition process, we argue that TLM's capability is limited by the input low-quality visual prediction. To be specific: 1) The visual prediction with low character-wise accuracy increases the correction burden of TLM. 2) The inconsistent word length between visual prediction and original image provides a wrong language modeling guidance in TLM. In this paper, we propose a Progressive scEne Text Recognizer (PETR) to improve the capability of transformer-based language model by handling above two problems. Firstly, a Destruction Learning Module (DLM) is proposed to consider the linguistic information in the visual context. DLM introduces the recognition of destructed images with disordered patches in the training stage. Through guiding the vision model to restore patch orders and make word-level prediction on the destructed images, visual prediction with high character-wise accuracy is obtained by exploring inner relationship between the local visual patches. Secondly, a new Language Rectification Module (LRM) is proposed to optimize the word length for language guidance rectification. Through progressively implementing LRM in different language modeling steps, a novel progressive rectification network is constructed to handle some extremely challenging cases (e.g. distortion, occlusion, etc.). By utilizing DLM and LRM, PETR enhances the capability of transformer-based language model from a more general aspect, that is, focusing on the reduction of correction burden and rectification of language modeling guidance. Compared with parallel transformer-based methods, PETR obtains 1.0% and 0.8% improvement on regular and irregular datasets respectively while introducing only 1.7M additional parameters. The extensive experiments on both English and Chinese benchmarks demonstrate that PETR achieves the state-of-the-art results. Yuxin Wang 0002, Hongtao Xie 0001, Shancheng Fang, Mengting Xing, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Boundary-Aware Arbitrary-Shaped Scene Text Detector With Learnable Embedding NetworkabstractBenefiting from the popularity of deep learning theory, scene text detection algorithms have developed rapidly in recent years. Methods representing text region by text segmentation map are proved to capture arbitrary-shaped text in a more flexible and accurate way. However, such segmentation-based methods are prone to be disturbed by the text-like background patterns (like the fence, grass, etc.), which generally suffer from imprecise boundary detail problem. In this paper, LEMNet is proposed to handle the imprecise boundary problem by guiding the generation of text boundary based on a priori constraint. In the training stage, Boundary Segmentation Branch is firstly constructed to predict coarse boundary mask for each text instance. Then, through mapping pixels into an embedding space, the proposed Pixel Embedding Branch makes the embedding representation of boundary points learn to be more similar, meanwhile enlarging the characteristic distance between background points and boundary points. During inference, noise in the coarse boundary segmentation map can be effectively suppressed by a Noisy Point Suppression Algorithm among pixel embedding vectors. In this way, LEMNet can generate a more precise boundary description of text regions. To further enhance the distinguishability of boundary features, we propose a Context Enhancement Module to capture feature interactions in different representation subspaces, in which features are parallelly performed attention and concatenated to generate enhanced features. Extensive experiments are conducted over four challenging datasets, which demonstrate the effectiveness of LEMNet. Specifically, LEMNet achieves F-measure of 85.2%, 87.6% and 85.2% on CTW1500, Total-Text and MSRA-TD500 respectively, which is the latest SOTA. Mengting Xing, Hongtao Xie 0001, Qingfeng Tan, Shancheng Fang, Yuxin Wang 0002, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text RecognitionabstractLinguistic knowledge is of great benefit to scene text recognition. However, how to effectively model linguistic rules in end-to-end deep networks remains a research challenge. In this paper, we argue that the limited capacity of language models comes from: 1) implicitly language modeling; 2) unidirectional feature representation; and 3) language model with noise input. Correspondingly, we propose an autonomous, bidirectional and iterative ABINet for scene text recognition. Firstly, the autonomous suggests to block gradient flow between vision and language models to enforce explicitly language modeling. Secondly, a novel bidirectional cloze network (BCN) as the language model is proposed based on bidirectional feature representation. Thirdly, we propose an execution manner of iterative correction for language model which can effectively alleviate the impact of noise input. Additionally, based on the ensemble of iterative predictions, we propose a self-training method which can learn from unlabeled images effectively. Extensive experiments indicate that ABINet has superiority on low-quality images and achieves state-of-the-art results on several mainstream benchmarks. Besides, the ABINet trained with ensemble self-training shows promising improvement in realizing human-level recognition. Code is available at https://github.com/FangShancheng/ABINet. Shancheng Fang, Hongtao Xie 0001, Yuxin Wang 0002, Zhendong Mao 0001, Yongdong Zhang 0001 |
CVPR | 1 |
| 2021 | From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkabstractIn this paper, we abandon the dominant complex language model and rethink the linguistic learning process in the scene text recognition. Different from previous methods considering the visual and linguistic information in two separate structures, we propose a Visual Language Modeling Network (VisionLAN), which views the visual and linguistic information as a union by directly enduing the vision model with language capability. Specially, we introduce the text recognition of character-wise occluded feature maps in the training stage. Such operation guides the vision model to use not only the visual texture of characters, but also the linguistic information in visual context for recognition when the visual cues are confused (e.g. occlusion, noise, etc.). As the linguistic information is acquired along with visual features without the need of extra language model, Vision-LAN significantly improves the speed by 39% and adaptively considers the linguistic information to enhance the visual features for accurate recognition. Furthermore, an Occlusion Scene Text (OST) dataset is proposed to evaluate the performance on the case of missing character-wise visual cues. The state of-the-art results on several benchmarks prove our effectiveness. Code and dataset are available at https://github.com/wangyuxin87/VisionLAN. Yuxin Wang 0002, Hongtao Xie 0001, Shancheng Fang, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
ICCV | 3 |
| 2021 | End-to-end Boundary Exploration for Weakly-supervised Semantic SegmentationabstractIt is full of challenges for weakly supervised semantic segmentation (WSSS) acquiring the pixel-level object location with only image-level annotations. Especially, the single-stage methods learn image- and pixel-level labels simultaneously to avoid complicated multi-stage computations and sophisticated training procedures. In this paper, we argue that using a single model to accomplish image- and pixel-level classification will fall into the balance of multi-target and consequently weakens the recognition capability. Because the image-level task tends to learn position-independent features, but the pixel-level task tends to be position-sensitive. Hence, we propose an effective encoder-decoder framework to explore object boundaries and solve the above dilemma. The encoder and decoder learn position-independent and position-sensitive features independently during the end-to-end training. In addition, a global soft pooling is suggested to suppress background pixels' activation for the encoder training and further improve the class activation map (CAM) performance. The edge annotations for the decoder training are synthesized by the high confidence CAMs, which do not requires extra supervision. The extensive experiments on the Pascal VOC12 dataset demonstrate that our method achieves state-of-the-art compared to the end-to-end approaches. It gets 63.6% and 65.7% mIoU scores on val and test sets respectively. Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Yue Hu 0002, Jianlong Tan |
ACM Multimedia | 2 |
| 2021 | TDI TextSpotter: Taking Data Imbalance into Account in Scene Text SpottingabstractRecent scene text spotters that integrate text detection module and recognition module have made significant progress. However, existing methods encounter two problems. 1). The data imbalance issue between text detection module and text recognition module limits the performance of text spotters. 2). The default left-to-right reading direction leads to errors in unconventional text spotting. In this paper, we propose a novel scene text spotter TDI to solve these problems. Firstly, in order to solve the data imbalance problem, a sample generation algorithm is proposed to generate plenty of samples online for training the text recognition module by using character features and character labels. Secondly, a weakly supervised character generation algorithm is designed to generate character-level labels from word-level labels for the sample generation algorithm and the training of the text detection module. Finally, in order to spot arbitrarily arranged text correctly, a direction perception module is proposed to perceive the reading direction of text instance. Experiments on several benchmarks show that these designs can significantly improve the performance of text spotter. Specifically, our method outperforms state-of-the-art methods on three public datasets in both text detection and end-to-end text recognition, which fully proves the effectiveness and robustness of our method. Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Jing Wang 0221, Zhengjun Zha, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2020 | CRNet: A Center-aware Representation for Detecting Text of Arbitrary ShapesabstractExisting scene text detection methods achieve state-of-the-art performance by designing elaborate anchors or complex post-processing. Nonetheless, most methods still face the dilemma of detecting adjacent texts as one instance and long text with large character spacing as multiple fragments. To tackle these problems, we propose an anchor-free scene text detector leveraging Center-aware Representation to achieve accurate arbitrary-shaped scene text detection namely CRNet. Firstly, we propose a center-aware location algorithm to explicitly learn center regions and center points of text instances, which is able to separate adjacent text instances effectively. Then, a multi-scale context extraction module capable of extracting local context, long-range dependencies and global context adaptively is designed to effectively perceive long text with large character spacing. Finally, a low-level features enhancement block is introduced to enhance the geometric information of text. Extensive experiments conducted on several benchmarks including SCUT-CTW1500, Total-Text, ICDAR2015, ICDAR2017 MLT, and MSRA-TD500 demonstrate the effectiveness of our method. Specifically, without any anchor and complicated post-processing, our CRNet achieves 84.2% and 85.1% on CTW1500 and MSRA-TD500 in F-measure, outperforming all state-of-the-art anchor-based and anchor-free methods. Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Yan Li 0068, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2019 | MLTS: A Multi-Language Scene Text SpotterabstractScene text detection and recognition are popular research topics in computer vision due to its various applications such as autonomous driving, blind assistance and text translation. However, many methods currently can only detect or recog-nize the text of one language. In scene text images, we can often see text in multi-language appearing on the same image. However, there is no valid model for multi-language text spotting. In this paper, an end-to-end method for multi-language scene text detection, recognition and script identification is proposed. The method, called MLTS, is an abbreviation of a Multi-Language Scene Text Spotter. By designing a special backbone for text and combining two different kinds of attention. MLTS achieves state-of-the-art performance for both joint localization and script identification in natural images and in cropped word script identification, the precision, recall and F-measure are 0.7145, 0.6583 and 0.6852 respectively, while the corresponding values of the best existing methods are 0.5759, 0.6207, 0.5974 respectively. Additionally, our MLTS achieves comparable performance on ICDAR2013 and ICDAR2015, which proves the effectiveness of the model. Yu Zhou 0016, Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
ICME | 2 |
| 2019 | Learning to Draw Text in Natural Images with Conditional Adversarial NetworksabstractIn this work, we propose an entirely learning-based method to automatically synthesize text sequence in natural images leveraging conditional adversarial networks. As vanilla GANs are clumsy to capture structural text patterns, directly employing GANs for text image synthesis typically results in illegible images. Therefore, we design a two-stage architecture to generate repeated characters in images. Firstly, a character generator attempts to synthesize local character appearance independently, so that the legible characters in sequence can be obtained. To achieve style consistency of characters, we propose a novel style loss based on variance-minimization. Secondly, we design a pixel-manipulation word generator constrained by self-regularization, which learns to convert local characters to plausible word image. Experiments on SVHN dataset and ICDAR, IIIT5K datasets demonstrate our method is able to synthesize visually appealing text images. Besides, we also show the high-quality images synthesized by our method can be used to boost the performance of a scene text recognition algorithm. Shancheng Fang, Hongtao Xie 0001, Jianlong Tan, Yongdong Zhang 0001 |
IJCAI | 1 |
| 2019 | Convolutional Attention Networks for Scene Text RecognitionabstractIn this article, we present Convoluitional Attention Networks (CAN) for unconstrained scene text recognition. Recent dominant approaches for scene text recognition are mainly based on Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN), where the CNN encodes images and the RNN generates character sequences. Our CAN is different from these methods; our CAN is completely built on CNN and includes an attention mechanism. The distinctive characteristics of our method include (i) CAN follows encoder-decoder architecture, in which the encoder is a deep two-dimensional CNN and the decoder is a one-dimensional CNN; (ii) the attention mechanism is applied in every convolutional layer of the decoder, and we propose a novel spatial attention method using average pooling; and (iii) position embeddings are equipped in both a spatial encoder and a sequence decoder to give our networks a sense of location. We conduct experiments on standard datasets for scene text recognition, including Street View Text , IIIT5K, and ICDAR datasets. The experimental results validate the effectiveness of different components and show that our convolutional-based method achieves state-of-the-art or competitive performance over prior works, even without the use of RNN. Hongtao Xie 0001, Shancheng Fang, Zhengjun Zha, Yating Yang, Yan Li 0068, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Deep Convolutional Nets for Pulmonary Nodule Detection and Classification
Nannan Sun, Dongbao Yang, Shancheng Fang, Hongtao Xie 0001 |
KSEM (2) | 3 |
| 2018 | Attention and Language Ensemble for Scene Text Recognition with Convolutional Sequence ModelingabstractRecent dominant approaches for scene text recognition are mainly based on convolutional neural network (CNN) and recurrent neural network (RNN), where the CNN processes images and the RNN generates character sequences. Different from these methods, we propose an attention-based architecture1 which is completely based on CNNs. The distinctive characteristics of our method include: (1) the method follows encoder-decoder architecture, in which the encoder is a two-dimensional residual CNN and the decoder is a deep one-dimensional CNN. (2) An attention module that captures visual cues, and a language module that models linguistic rules are designed equally in the decoder. Therefore the attention and language can be viewed as an ensemble to boost predictions jointly. (3) Instead of using a single loss from language aspect, multiple losses from attention and language are accumulated for training the networks in an end-to-end way. We conduct experiments on standard datasets for scene text recognition, including Street View Text, IIIT5K and ICDAR datasets. The experimental results show our CNN-based method has achieved state-of-the-art performance on several benchmark datasets, even without the use of RNN. Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Nannan Sun, Jianlong Tan, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2017 | Detecting Uyghur text in complex background images with convolutional neural network
Shancheng Fang, Hongtao Xie 0001, Zhineng Chen, Shiai Zhu, Xiaoyan Gu 0001, Xingyu Gao 0001 |
Multim. Tools Appl. | 1 |