Canjie Luo

dblp:209/9580 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
12since 2021 · last 2023
0000-0002-1089-7611ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021
YearPublicationVenuePosition
2023 SLOGAN: Handwriting Style Synthesis for Arbitrary-Length and Out-of-Vocabulary Text
abstract
Large amounts of labeled data are urgently required for the training of robust text recognizers. However, collecting handwriting data of diverse styles, along with an immense lexicon, is considerably expensive. Although data synthesis is a promising way to relieve data hunger, two key issues of handwriting synthesis, namely, style representation and content embedding, remain unsolved. To this end, we propose a novel method that can synthesize parameterized and controllable handwriting S tyles for arbitrary-Length and O ut-of-vocabulary text based on a G enerative A dversarial N etwork (GAN), termed SLOGAN. Specifically, we propose a style bank to parameterize specific handwriting styles as latent vectors, which are input to a generator as style priors to achieve the corresponding handwritten styles. The training of the style bank requires only writer identification of the source images, rather than attribute annotations. Moreover, we embed the text content by providing an easily obtainable printed style image, so that the diversity of the content can be flexibly achieved by changing the input printed image. Finally, the generator is guided by dual discriminators to handle both the handwriting characteristics that appear as separated characters and in a series of cursive joins. Our method can synthesize words that are not included in the training vocabulary and with various new styles. Extensive experiments have shown that high-quality text images with great style diversity and rich vocabulary can be synthesized using our method, thereby enhancing the robustness of the recognizer.
Canjie Luo, Zhe Li 0046, Dezhi Peng
IEEE Trans. Neural Networks Learn. Syst.1
2022 Look Closer to Supervise Better: One-Shot Font Generation via Component-Based Discriminator
abstract
Automatic font generation remains a challenging research issue due to the large amounts of characters with complicated structures. Typically, only a few samples can serve as the style/content reference (termed few-shot learning), which further increases the difficulty to preserve local style patterns or detailed glyph structures. We investigate the drawbacks of previous studies and find that a coarsegrained discriminator is insufficient for supervising a font generator. To this end, we propose a novel Component-Aware Module (CAM), which supervises the generator to decouple content and style at a more fine-grained level, i.e., the component level. Different from previous studies struggling to increase the complexity of generators, we aim to perform more effective supervision for a relatively simple generator to achieve its full potential, which is a brand new perspective for font generation. The whole framework achieves remarkable results by coupling component-level supervision with adversarial learning, hence we call it Component-Guided GAN, shortly CG-GAN. Extensive experiments show that our approach outperforms state-of-the-art one-shot font generation methods. Furthermore, it can be applied to handwritten word synthesis and scene text image editing, suggesting the generalization of our approach.
Yuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu, Shenggao Zhu, Nicholas Jing Yuan
CVPR2
2022 SimAN: Exploring Self-Supervised Representation Learning of Scene Text via Similarity-Aware Normalization
abstract
Recently self-supervised representation learning has drawn considerable attention from the scene text recognition community. Different from previous studies using contrastive learning, we tackle the issue from an alternative perspective, i.e., by formulating the representation learning scheme in a generative manner. Typically, the neighboring image patches among one text line tend to have similar styles, including the strokes, textures, colors, etc. Motivated by this common sense, we augment one image patch and use its neighboring patch as guidance to recover itself. Specifically, we propose a Similarity-Aware Normalization (SimAN) module to identify the different patterns and align the corresponding styles from the guiding patch. In this way, the network gains representation capability for distinguishing complex patterns such as messy strokes and cluttered backgrounds. Experiments show that the proposed SimAN significantly improves the representation quality and achieves promising performance. Moreover, we surprisingly find that our self-supervised generative network has impressive potential for data synthesis, text image editing, and font interpolation, which suggests that the proposed SimAN has a wide range of practical applications.
Canjie Luo, Jingdong Chen
CVPR1
2022 Don't Forget Me: Accurate Background Recovery for Text Removal via Modeling Local-Global Context
Chongyu Liu, Canjie Luo, Bangdong Chen, Fengjun Guo, Kai Ding 0009
ECCV (28)4
2022 ChaCo: Character Contrastive Learning for Handwritten Text Recognition
Jiapeng Wang 0003, Canjie Luo, Yang Xue 0001
ICFHR5
2022 Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild
abstract
Camera-captured document images usually suffer from perspective and geometric deformations. It is of great value to rectify them when considering poor visual aesthetics and the deteriorated performance of OCR systems. Recent learning-based methods intensively focus on the accurately cropped document image. However, this might not be sufficient for overcoming practical challenges, including document images either with large marginal regions or without margins. Due to this impracticality, users struggle to crop documents precisely when they encounter large marginal regions. Simultaneously, dewarping images without margins is still an insurmountable problem. To the best of our knowledge, there is still no complete and effective pipeline for rectifying document images in the wild. To address this issue, we propose a novel approach called Marior (Margin Removal and Iterative Content Rectification). Marior follows a progressive strategy to iteratively improve the dewarping quality and readability in a coarse-to-fine manner. Specifically, we divide the pipeline into two modules: margin removal module (MRM) and iterative content rectification module (ICRM). First, we predict the segmentation mask of the input image to remove the margin, thereby obtaining a preliminary result. Then we refine the image further by producing dense displacement flows to achieve content-aware rectification. We determine the number of refinement iterations adaptively. Experiments demonstrate the state-of-the-art performance of our method on public benchmarks. The resources are available at https://github.com/ZZZHANG-jx/Marior for further comparison.
Jiaxin Zhang 0003, Canjie Luo, Fengjun Guo, Kai Ding 0009
ACM Multimedia2
2022 PageNet: Towards End-to-End Weakly Supervised Page-Level Handwritten Chinese Text Recognition
Dezhi Peng, Canjie Luo, Songxuan Lai
Int. J. Comput. Vis.4
2021 Implicit Feature Alignment: Learn To Convert Text Recognizer to Text Spotter
abstract
Text recognition is a popular research subject with many associated challenges. Despite the considerable progress made in recent years, the text recognition task itself is still constrained to solve the problem of reading cropped line text images and serves as a subtask of optical character recognition (OCR) systems. As a result, the final text recognition result is limited by the performance of the text detector. In this paper, we propose a simple, elegant and effective paradigm called Implicit Feature Alignment (IFA), which can be easily integrated into current text recognizers, resulting in a novel inference mechanism called IFA- inference. This enables an ordinary text recognizer to process multi-line text such that text detection can be completely freed. Specifically, we integrate IFA into the two most prevailing text recognition streams (attention-based and CTC-based) and propose attention-guided dense prediction (ADP) and Extended CTC (ExCTC). Furthermore, the Wasserstein-based Hollow Aggregation Cross-Entropy (WH-ACE) is proposed to suppress negative predictions to assist in training ADP and ExCTC. We experimentally demonstrate that IFA achieves state-of-the-art performance on end-to-end document recognition tasks while maintaining the fastest speed, and ADP and ExCTC complement each other on the perspective of different application scenarios. Code will be available at https://github.com/Wang-Tianwei/Implicit-feature-alignment.
Dezhi Peng, Zhe Li 0046, Mengchao He, Yongpan Wang, Canjie Luo
CVPR8
2021 A Multi-level Progressive Rectification Mechanism for Irregular Scene Text Recognition
Qianying Liao, Qingxiang Lin, Canjie Luo, Jiaxin Zhang 0003, Dezhi Peng
ICDAR (4)4
2021 Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
Tong He 0001, Hao Chen 0041, Xinyu Wang 0010, Canjie Luo, Shuaitao Zhang, Chunhua Shen
Int. J. Comput. Vis.5
2021 Separating Content from Style Using Adversarial Learning for Recognizing Text in the Wild
Canjie Luo, Qingxiang Lin, Chunhua Shen
Int. J. Comput. Vis.1
2021 STAN: A sequential transformation attention-based network for scene text recognition
Qingxiang Lin, Canjie Luo, Songxuan Lai
Pattern Recognit.2
2020 Decoupled Attention Network for Text Recognition
abstract
Text recognition has attracted considerable research interests because of its various applications. The cutting-edge text recognition methods are based on attention mechanisms. However, most of attention methods usually suffer from serious alignment problem due to its recurrency alignment operation, where the alignment relies on historical decoding results. To remedy this issue, we propose a decoupled attention network (DAN), which decouples the alignment operation from using historical decoding results. DAN is an effective, flexible and robust end-to-end text recognizer, which consists of three components: 1) a feature encoder that extracts visual features from the input image; 2) a convolutional alignment module that performs the alignment operation based on visual features from the encoder; and 3) a decoupled text decoder that makes final prediction by jointly using the feature map and attention maps. Experimental results show that DAN achieves state-of-the-art performance on multiple text recognition tasks, including offline handwritten text recognition and regular/irregular scene text recognition. Codes will be released.1
Canjie Luo, Xiaoxue Chen, Yaqiang Wu, Qianying Wang 0002, Mingxiang Cai
AAAI4
2020 On the General Value of Evidence, and Bilingual Scene-Text Visual Question Answering
abstract
Visual Question Answering (VQA) methods have made incredible progress, but suffer from a failure to generalize. This is visible in the fact that they are vulnerable to learning coincidental correlations in the data rather than deeper relations between image content and ideas expressed in language. We present a dataset that takes a step towards addressing this problem in that it contains questions expressed in two languages, and an evaluation process that co-opts a well understood image-based metric to reflect the method’s ability to reason. Measuring reasoning directly encourages generalization by penalizing answers that are coincidentally correct. The dataset reflects the scene-text version of the VQA problem, and the reasoning evaluation can be seen as a text-based version of a referring expression challenge. Experiments and analyses are provided that show the value of the dataset. The dataset is available at www.est-vqa.org.
Xinyu Wang 0010, Chunhua Shen, Chun Chet Ng, Canjie Luo, Chee Seng Chan, Anton van den Hengel
CVPR5
2020 Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition
abstract
Handwritten text and scene text suffer from various shapes and distorted patterns. Thus training a robust recognition model requires a large amount of data to cover diversity as much as possible. In contrast to data collection and annotation, data augmentation is a low cost way. In this paper, we propose a new method for text image augmentation. Different from traditional augmentation methods such as rotation, scaling and perspective transformation, our proposed augmentation method is designed to learn proper and efficient data augmentation which is more effective and specific for training a robust recognizer. By using a set of custom fiducial points, the proposed augmentation method is flexible and controllable. Furthermore, we bridge the gap between the isolated processes of data augmentation and network optimization by joint learning. An agent network learns from the output of the recognition network and controls the fiducial points to generate more proper training samples for the recognition network. Extensive experiments on various benchmarks, including regular scene text, irregular scene text and handwritten text, show that the proposed augmentation and the joint learning methods significantly boost the performance of the recognition networks. A general toolkit for geometric augmentation is available.
Canjie Luo, Yongpan Wang
CVPR1
2020 Adaptive embedding gate for attention-based scene text recognition
abstract
Scene text recognition has attracted particular research interest because it is a very challenging problem and has various applications. The most cutting-edge methods are attentional encoder-decoder frameworks that learn the alignment between the input image and output sequences. In particular, the decoder recurrently outputs predictions, using the prediction of the previous step as a guidance for every time step. In this study, we point out that the inappropriate use of previous predictions in existing attentional decoders restricts the recognition performance and brings instability. To handle this problem, we propose a novel module, namely adaptive embedding gate (AEG). The proposed AEG focuses on introducing high-order character language models to attentional decoders by controlling the information transmission between adjacent characters. AEG is a flexible module and can be easily integrated into the state-of-the-art attentional decoders for scene text recognition. We evaluate its effectiveness as well as robustness on a number of standard benchmarks, including the IIIT5K, SVT, SVT-P, CUTE80, and ICDAR datasets. Experimental results demonstrate that AEG can significantly boost recognition performance and bring better robustness.
Xiaoxue Chen, Canjie Luo
Neurocomputing5
2020 EPAN: Effective parts attention network for scene text recognition
abstract
For most previous attention-based scene text recognition methods, images are transformed into high-level feature vectors that form a feature map with height equal to one. Such vectors may contain unnecessary noise that limits recognition performance. To address this issue, in this paper, we propose the effective parts attention network (EPAN) which can attentively highlight the character region for more precise recognition. EPAN consists of a text image encoder and character effective parts decoder (CEPD), and it is end-to-end trainable. The former separates the high-dimensional feature map into one-dimensional vectors row-by-row, which are connected to a bidirectional long short term memory unit to encode contextual information. Subsequently, the CEPD transforms the vectors using a novel glimpse network at each time step to roughly determine the position of the characters. Then the CEPD uses a refinement network to generate a mask to gradually localize the precise position of important parts of the current character. Experiments were conducted on various benchmarks, including IIIT5K-Words, Street View Text, ICDAR 2003, ICDAR 2013, CUTE80, Street View Text Perspective, and ICDAR 2015, which demonstrated that the proposed EPAN method significantly outperformed or was comparable to existing methods in terms of lexicon-free word accuracy. Additionally, substantial qualitative results further demonstrated the robustness of our method.
Yunlong Huang, Zenghui Sun, Canjie Luo
Neurocomputing4
2020 SaHAN: Scale-aware hierarchical attention network for scene text recognition
Jiaxin Zhang 0003, Canjie Luo, Weiying Zhou
Pattern Recognit. Lett.2
2020 EraseNet: End-to-End Text Removal in the Wild
abstract
Scene text removal has attracted increasing research interests owing to its valuable applications in privacy protection, camera-based virtual reality translation, and image editing. However, existing approaches, which fall short on real applications, are mainly because they were evaluated on synthetic or unrepresentative datasets. To fill this gap and facilitate this research direction, this paper proposes a real-world dataset called SCUT-EnsText that consists of 3,562 diverse images selected from public scene text reading benchmarks, and each image is scrupulously annotated to provide visually plausible erasure targets. With SCUT-EnsText, we design a novel GANbased model termed EraseNet that can automatically remove text located on the natural images. The model is a two-stage network that consists of a coarse-erasure sub-network and a refinement sub-network. The refinement sub-network targets improvement in the feature representation and refinement of the coarse outputs to enhance the removal performance. Additionally, EraseNet contains a segmentation head for text perception and a local-global SN-Patch-GAN with spectral normalization (SN) on both the generator and discriminator for maintaining the training stability and the congruity of the erased regions. A sufficient number of experiments are conducted on both the previous public dataset and the brand-new SCUT-EnsText. Our EraseNet significantly outperforms the existing state-of-the-art methods in terms of all metrics, with remarkably superior higherquality results. The dataset and code will be made available at https://github.com/HCIILAB/SCUT-EnsText.
Chongyu Liu, Shuaitao Zhang, Canjie Luo, Yongpan Wang
IEEE Trans. Image Process.5
2019 Tightness-Aware Evaluation Protocol for Scene Text Detection
abstract
Evaluation protocols play key role in the developmental progress of text detection methods. There are strict requirements to ensure that the evaluation methods are fair, objective and reasonable. However, existing metrics exhibit some obvious drawbacks: 1) They are not goal-oriented; 2) they cannot recognize the tightness of detection methods; 3) existing one-to-many and many-to-one solutions involve inherent loopholes and deficiencies. Therefore, this paper proposes a novel evaluation protocol called Tightness-aware Intersect-over-Union (TIoU) metric that could quantify completeness of ground truth, compactness of detection, and tightness of matching degree. Specifically, instead of merely using the IoU value, two common detection behaviors are properly considered; meanwhile, directly using the score of TIoU to recognize the tightness. In addition, we further propose a straightforward method to address the annotation granularity issue, which can fairly evaluate word and text-line detections simultaneously. By adopting the detection results from published methods and general object detection frameworks, comprehensive experiments on ICDAR 2013 and ICDAR 2015 datasets are conducted to compare recent metrics and the proposed TIoU metric. The comparison demonstrated some promising new prospects, e.g., determining the methods and frameworks for which the detection is tighter and more beneficial to recognize. Our method is extremely simple; however, the novelty is none other than the proposed metric can utilize simplest but reasonable improvements to lead to many interesting and insightful prospects and solving most the issues of the previous metrics. The code is publicly available at https://github.com/Yuliang-Liu/TIoU-metric.
Zecheng Xie, Canjie Luo, Shuaitao Zhang, Lele Xie
CVPR4
2019 ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text - RRC-ArT
abstract
This paper reports the ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text - RRC-ArT that consists of three major challenges: i) scene text detection, ii) scene text recognition, and iii) scene text spotting. A total of 78 submissions from 46 unique teams/individuals were received for this competition. The top performing score of each challenge is as follows: i) T1 - 82.65%, ii) T2.1 - 74.3%, iii) T2.2 - 85.32%, iv) T3.1 - 53.86%, and v) T3.2 - 54.91%. Apart from the results, this paper also details the ArT dataset, tasks description, evaluation metrics and participants' methods. The dataset, the evaluation kit as well as the results are publicly available at the challenge website.
Chee Kheng Chng, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, Chee Seng Chan, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, Junyu Han
ICDAR10
2019 Attention After Attention: Reading Text in the Wild with Cross Attention
abstract
Recent methods mostly regarded scene text recognition as a sequence-to-sequence problem. These methods roughly transform the image into a feature sequence and use the algorithms for sequence-to-sequence problem like CTC or attention to decode the characters. However, text in images is distributed in a two-dimensional (2D) space and roughly converting the features of text into a feature sequence may introduce extra noise, especially if the text is irregular. In this paper, we propose a novel framework named cross attention network, which learns to attend to local features of a 2D feature map corresponding to individual characters. The network contains two 1D attention networks, which operates harmoniously in two directions. Thus, one of the attention modules vertically attends to the features corresponding to the whole text of 2D features and the other horizontal module selects the local features to decode individual characters. Extensive experiments are performed on various regular benchmarks, including SVT, ICDAR2003, ICDAR2013, and IIIT5K-Words, which demonstrate that the proposed model either outperforms or is comparable to all previous methods. Moreover, the model is evaluated on irregular benchmarks including SVT-Perspective, CUTE80 and ICDAR 2015. The performance on irregular benchmarks shows the robustness of our model.
Yunlong Huang, Canjie Luo, Qingxiang Lin, Weiying Zhou
ICDAR2
2019 ICDAR 2019 Competition on Large-Scale Street View Text with Partial Labeling - RRC-LSVT
abstract
Robust text reading from street view images provides valuable information for various applications. Performance improvement of existing methods in such a challenging scenario heavily relies on the amount of fully annotated training data, which is costly and in-efficient to obtain. To scale up the amount of training data while keeping the labeling procedure cost-effective, this competition introduces a new challenge on Large-scale Street View Text with Partial Labeling (LSVT), providing 5,0000 and 400,000 images in full and weak annotations, respectively. This competition aims to explore the abilities of state-of-the-art methods to detect and recognize text instances from large-scale street view images, closing gaps between research benchmarks and real applications. During the competition period, a total number of 41 teams participate in the two tasks with 132 valid submissions, i.e., text detection and end-to-end text spotting. This paper includes dataset descriptions, task definitions, evaluation protocols and results summaries of ICDAR 2019-LSVT challenge.
Yipeng Sun, Dimosthenis Karatzas, Chee Seng Chan, Zihan Ni, Chee Kheng Chng, Canjie Luo, Chun Chet Ng, Junyu Han, Errui Ding, Jingtuo Liu
ICDAR8
2019 Curved scene text detection via transverse and longitudinal sequence connection
abstract
Curved text detection is a difficult problem that has not been addressed sufficiently. To highlight the difficulties in reading curved text in a real environment, we constructed a curved text dataset called CTW1500, which includes over 10,000 text annotations in 1500 images, and used it to formulate a polygon-based curved text detector that can detect curved text without using an empirical combination. With the seamless integration of recurrent transverse and longitudinal offset connection, our method explores context information instead of predicting points independently, resulting in smoother and more accurate detection. Our approach is designed as a universal method, meaning it can be trained using rectangular or quadrilateral bounding boxes, requiring no extra effort. Experimental results on the CTW1500 dataset and Total-text demonstrated that our method with only a light backbone can outperform state-of-the-art methods by a large margin. Our method also achieved state-of-the-art performance on the MSRA-TD500 dataset, demonstrating its promising generalization ability. Code, datasets, and label-tool are available at https://github.com/Yuliang-Liu/Curve-Text-Detector.
Shuaitao Zhang, Canjie Luo, Sheng Zhang 0024
Pattern Recognit.4
2019 MORAN: A Multi-Object Rectified Attention Network for scene text recognition
Canjie Luo, Zenghui Sun
Pattern Recognit.1
2018 Feature Enhancement Network: A Refined Scene Text Detector
abstract
In this paper, we propose a refined scene text detector with a novel Feature Enhancement Network (FEN)for Region Proposal and Text Detection Refinement. Retrospectively, both region proposal with only 3 x 3 sliding-window feature and text detection refinement with single scale high level feature are insufficient, especially for smaller scene text. Therefore, we design a new FEN network with task-specific, low and high level semantic features fusion to improve the performance of text detection. Besides, since unitary position-sensitive RoI pooling in general object detection is unreasonable for variable text regions, an adaptively weighted position-sensitive RoI pooling layer is devised for further enhancing the detecting accuracy. To tackle the sample-imbalance problem during the refinement stage,we also propose an effective positives mining strategy for efficiently training our network. Experiments on ICDAR2011 and 2013 robust text detection benchmarks demonstrate that our method can achieve state-of-the-art results, outperforming all reported methods in terms of F-measure.
Sheng Zhang 0024, Canjie Luo
AAAI4
2018 ICPR2018 Contest on Robust Reading for Multi-Type Web Images
abstract
Electronic commerce has infiltrated every aspect of our daily lives, which offers great convenience for shopping, advertising, etc. Text in the web images is responsible to convey essential information for consumers. Algorithms that read text in these web images can facilitate applications of various types, such as goods surveillance, products classification, and intelligent retrieval or recommendation. Despite of various existing text reading tasks, this contest introduces a novel large-scale dataset named MTWI that contains 20,000 images, which is the first dataset that is mainly constructed by Chinese and English web text. Three tasks (web text recognition, web text detection, and end-to-end web text detection and recognition) were set up for encouraging more research on the web text reading problem. The contest was held from February 2, 2018 to May 26, 2018 with 289 valid submissions from 4,282 registered teams. Throughout this report, we describe the details of this new dataset, the purposes and definitions of the tasks, the evaluation protocols, and the summaries of the results.
Mengchao He, Zhibo Yang 0003, Sheng Zhang 0024, Canjie Luo, Feiyu Gao, Qi Zheng 0002, Yongpan Wang, Xin Zhang 0013
ICPR5