EDBT 2026 Demo / reviewers in the wild / expert
Chengquan Zhang
dblp:163/1795
· DBLP profile ↗
31ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0001-8254-5773ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Recognition-Synergistic Scene Text EditingabstractScene text editing aims to modify text content within scene images while maintaining style consistency. Traditional methods achieve this by explicitly disentangling style and content from the source image and then fusing the style with the target content, while ensuring content consistency using a pre-trained recognition model. Despite notable progress, these methods suffer from complex pipelines, leading to suboptimal performance in complex scenarios. In this work, we introduce Recognition-Synergistic Scene Text Editing (RS-STE), a novel approach that fully exploits the intrinsic synergy of text recognition for editing. Our model seamlessly integrates text recognition with text editing within a unified framework, and leverages the recognition model’s ability to implicitly disentangle style and content while ensuring content consistency. Specifically, our approach employs a multi-modal parallel decoder based on transformer architecture, which predicts both text content and stylized images in parallel. Additionally, our cyclic self-supervised fine-tuning strategy enables effective training on unpaired real-world data without ground truth, enhancing style and content consistency through a twice-cyclic generation process. Built on a relatively simple architecture, RS-STE achieves state-of-the-art performance on both synthetic and real-world benchmarks, and further demonstrates the effectiveness of leveraging the generated hard cases to boost the performance of downstream recognition tasks. Code is available at https://github.com/ZhengyaoFang/RS-STE. Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
CVPR | 4 |
| 2024 | Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents
Mengjun Cheng, Chengquan Zhang, Chang Liu 0047, Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
ECCV (45) | 2 |
| 2024 | WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-Only Supervised Text Spotting
Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Fanglin Chen 0001, Guangming Lu 0002, Wenjie Pei |
ECCV (31) | 4 |
| 2024 | Towards Unified Multi-granularity Text Detection with Interactive AttentionabstractExisting OCR engines or document image analysis systems typically rely on training separate models for text detection in varying scenarios and granularities, leading to significant computational complexity and resource demands. In this paper, we introduce "Detect Any Text" (DAT), an advanced paradigm that seamlessly unifies scene text detection, layout analysis, and document page detection into a cohesive, end-to-end model. This design enables DAT to efficiently manage text instances at different granularities, including word, line, paragraph and page. A pivotal innovation in DAT is the across-granularity interactive attention module, which significantly enhances the representation learning of text instances at varying granularities by correlating structural information across different text queries. As a result, it enables the model to achieve mutually beneficial detection performances across multiple text granularities. Additionally, a prompt-based segmentation module refines detection outcomes for texts of arbitrary curvature and complex layouts, thereby improving DAT’s accuracy and expanding its real-world applicability. Experimental results demonstrate that DAT achieves state-of-the-art performances across a variety of text-related benchmarks, including multi-oriented/arbitrarily-shaped scene text detection, document layout analysis and page detection tasks. Xingyu Wan, Chengquan Zhang, Pengyuan Lv, Sen Fan, Zihan Ni, Errui Ding, Jingdong Wang 0001 |
ICML | 2 |
| 2024 | Beyond Memory Safety: an Empirical Study on Bugs and Fixes of Rust ProgramsabstractRust is a nascent programming language designed to improve memory safety for system programming while maintaining high performance. The Rust language ensures memory safety through its ownership mechanism and by performing compile-time checks on safe code. However, for low-level controls, developers are allowed to bypass these checks by marking their code as unsafe, which in turn introduces memory vulnerabilities. Beyond these memory-related concerns, the existence and nature of other common bugs such as run-time panics have not been thoroughly explored. In this paper, we conduct a comprehensive empirical study to characterize bugs and their fixes beyond memory safety concerns by manually inspecting bug patches in Rust programs. We identify 790 bug fixes from 1100 commits in six widely-used Rust projects and the Rust standard library, and then investigate their root causes and symptoms. Furthermore, we analyze the relationships between these bugs and unsafe code (i.e., whether they are caused by the use of unsafe code and to what extent it impacts them). Our bug study introduces a classification of 15 root causes and 6 symptoms, and categorizes bugs into different groups according to their relationships with safe/unsafe code. We identify 19 major findings and draw broader lessons from them to guide the research community towards future directions in program testing, analysis, fault localization, and repair for Rust language. Chengquan Zhang, Yang Feng 0003, Yaokun Zhang, Yuxuan Dai, Baowen Xu |
QRS | 1 |
| 2024 | Irregular text block recognition via decoupling visual, linguistic, and positional information
Chengquan Zhang, Jiaxin Zhang 0003, Zecheng Xie, Pengyuan Lv |
Pattern Recognit. | 3 |
| 2023 | ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai |
ICDAR (2) | 2 |
| 2023 | StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training
Yuechen Yu, Yulin Li 0004, Chengquan Zhang, Xiaoqiang Zhang 0006, Zengyuan Guo, Xiameng Qin, Junyu Han, Errui Ding, Jingdong Wang 0001 |
ICLR | 3 |
| 2023 | Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document UnderstandingabstractTransformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unable to handle the layout representation in documents, e.g. word, line and paragraph, on different granularity levels and seem hard to achieve a good trade-off between efficiency and performance. To tackle the concerns, we propose Fast-StrucTexT, an efficient multi-modal framework based on the StrucTexT algorithm with an hourglass transformer architecture, for visual document understanding. Specifically, we design a modality-guided dynamic token merging block to make the model learn multi-granularity representation and prunes redundant tokens. Additionally, we present a multi-modal interaction module called Symmetry Cross-Attention (SCA) to consider multi-modal fusion and efficiently guide the token mergence. The SCA allows one modality input as query to calculate cross attention with another modality in a dual phase. Extensive experiments on FUNSD, SROIE, and CORD datasets demonstrate that our model achieves the state-of-the-art performance and almost 1.9x faster inference time than the state-of-the-art methods. Mingliang Zhai, Yulin Li 0004, Xiameng Qin, Qunyi Xie, Chengquan Zhang, Yuwei Wu 0001, Yunde Jia |
IJCAI | 6 |
| 2023 | GridFormer: Towards Accurate Table Structure Recognition via Grid PredictionabstractAll tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex and edge of a grid. First, we propose a flexible table representation in the form of an M X N grid. In this representation, the vertexes and edges of the grid store the localization and adjacency information of the table. Then, we introduce a DETR-style table structure recognizer to efficiently predict this multi-objective information of the grid in a single shot. Specifically, given a set of learned row and column queries, the recognizer directly outputs the vertexes and edges information of the corresponding rows and columns. Extensive experiments on five challenging benchmarks which include wired, wireless, multi-merge-cell, oriented, and distorted tables demonstrate the competitive performance of our model over other methods. Pengyuan Lv, Weihong Ma, Hongyi Wang 0008, Yuechen Yu, Chengquan Zhang, Yang Xue 0001, Jingdong Wang 0001 |
ACM Multimedia | 5 |
| 2023 | Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningabstractDue to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection. Despite great progress, these methods show deficiencies in robustness and still suffer from false positives and instance adhesion. Different from existing methods which integrate multiple-granularity features or multiple outputs, we resort to the perspective of representation learning in which auxiliary tasks are utilized to enable the encoder to jointly learn robust features with the main task of per-pixel classification during optimization. For semantic representation learning, we propose global-dense semantic contrast (GDSC), in which a vector is extracted for global semantic representation, then used to perform element-wise contrast with the dense grid features. To learn instance-aware representation, we propose to combine top-down modeling (TDM) with the bottom-up framework to provide implicit instance-level clues for the encoder. With the proposed GDSC and TDM, the encoder network learns stronger representation without introducing any parameters and computations during inference. Equipped with a very light decoder, the detector can achieve more robust real-time scene text detection. Experimental results on four public datasets show that the proposed method can outperform or be comparable to the state-of-the-art on both accuracy and speed. Specifically, the proposed method achieves 87.2% F-measure with 48.2 FPS on Total-Text and 89.6% F-measure with 36.9 FPS on MSRA-TD500 on a single GeForce RTX 2080 Ti GPU. Xugong Qin, Pengyuan Lv, Chengquan Zhang, Yu Zhou 0015, Peng Zhang 0044, Hailun Lin, Weiping Wang 0005 |
ACM Multimedia | 3 |
| 2022 | Decoupling Recognition from Detection: Single Shot Self-Reliant Scene Text SpotterabstractTypical text spotters follow the two-stage spotting strategy: detect the precise boundary for a text instance first and then perform text recognition within the located text region. While such strategy has achieved substantial progress, there are two underlying limitations. 1) The performance of text recognition depends heavily on the precision of text detection, resulting in the potential error propagation from detection to recognition. 2) The RoI cropping which bridges the detection and recognition brings noise from background and leads to information loss when pooling or interpolating from feature maps. In this work we propose the single shot Self-Reliant Scene Text Spotter (SRSTS), which circumvents these limitations by decoupling recognition from detection. Specifically, we conduct text detection and recognition in parallel and bridge them by the shared positive anchor point. Consequently, our method is able to recognize the text instances correctly even though the precise text boundaries are challenging to detect. Additionally, our method reduces the annotation cost for text detection substantially. Extensive experiments on regular-shaped benchmark and arbitrary-shaped benchmark demonstrate that our SRSTS compares favorably to previous state-of-the-art spotters in terms of both accuracy and efficiency. Pengyuan Lv, Guangming Lu 0002, Chengquan Zhang, Wenjie Pei |
ACM Multimedia | 4 |
| 2021 | PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering NetworkabstractThe reading of arbitrarily-shaped text has received increasing research attention. However, existing text spotters are mostly built on two-stage frameworks or character-based methods, which suffer from either Non-Maximum Suppression (NMS), Region-of-Interest (RoI) operations, or character-level annotations. In this paper, to address the above problems, we propose a novel fully convolutional Point Gathering Network (PGNet) for reading arbitrarily-shaped text in real-time. The PGNet is a single-shot text spotter, where the pixel-level character classification map is learned with proposed PG-CTC loss avoiding the usage of character-level annotations. With a PG-CTC decoder, we gather high-level character classification vectors from two-dimensional space and decode them into text symbols without NMS and RoI operations involved, which guarantees high efficiency. Additionally, reasoning the relations between each character and its neighbors, a graph refinement module (GRM) is proposed to optimize the coarse recognition and improve the end-to-end performance. Experiments prove that the proposed method achieves competitive accuracy, meanwhile significantly improving the running speed. In particular, in Total-Text, it runs at 46.7 FPS, surpassing the previous spotters with a large margin. Chengquan Zhang, Fei Qi 0001, Xiaoqiang Zhang 0006, Pengyuan Lv, Junyu Han, Jingtuo Liu, Errui Ding, Guangming Shi |
AAAI | 2 |
| 2021 | StrucTexT: Structured Text Understanding with Multi-Modal TransformersabstractStructured text understanding on Visually Rich Documents (VRDs) is a crucial part of Document Intelligence. Due to the complexity of content and layout in VRDs, structured text understanding has been a challenging task. Most existing studies decoupled this problem into two sub-tasks: entity labeling and entity linking, which require an entire understanding of the context of documents at both token and segment levels. However, little work has been concerned with the solutions that efficiently extract the structured data from different levels. This paper proposes a unified framework named StrucTexT, which is flexible and effective for handling both sub-tasks. Specifically, based on the transformer, we introduce a segment-token aligned encoder to deal with the entity labeling and entity linking tasks at different levels of granularity. Moreover, we design a novel pre-training strategy with three self-supervised tasks to learn a richer representation. StrucTexT uses the existing Masked Visual Language Modeling task and the new Sentence Length Prediction and Paired Boxes Direction tasks to incorporate the multi-modal information across text, image, and layout. We evaluate our method for structured text understanding at segment-level and token-level and show it outperforms the state-of-the-art counterparts with significantly superior performance on the FUNSD, SROIE, and EPHOIE datasets. Yulin Li 0004, Yuxi Qian, Yuechen Yu, Xiameng Qin, Chengquan Zhang, Junyu Han, Jingtuo Liu, Errui Ding |
ACM Multimedia | 5 |
| 2021 | End-to-end video text detection with online tracking
Hongyuan Yu, Yan Huang 0008, Lihong Pi, Chengquan Zhang, Liang Wang 0001 |
Pattern Recognit. | 4 |
| 2020 | Towards Accurate Scene Text Recognition With Semantic Reasoning NetworksabstractScene text image contains two levels of contents: visual texture and semantic information. Although the previous scene text recognition methods have made great progress over the past few years, the research on mining semantic information to assist text recognition attracts less attention, only RNN-like structures are explored to implicitly model semantic information. However, we observe that RNN based methods have some obvious shortcomings, such as time-dependent decoding manner and one-way serial transmission of semantic context, which greatly limit the help of semantic information and the computation efficiency. To mitigate these limitations, we propose a novel end-to-end trainable framework named semantic reasoning network (SRN) for accurate scene text recognition, where a global semantic reasoning module (GSRM) is introduced to capture global semantic context through multi-way parallel transmission. The state-of-the-art results on 7 public benchmarks, including regular text, irregular text and non-Latin long text, verify the effectiveness and robustness of the proposed method. In addition, the speed of SRN has significant advantages over the RNN based methods, demonstrating its value in practical use. Deli Yu, Chengquan Zhang, Junyu Han, Jingtuo Liu, Errui Ding |
CVPR | 3 |
| 2020 | Learning Global Structure Consistency for Robust Object TrackingabstractFast appearance variations and the distractions of similar objects are two of the most challenging problems in visual object tracking. Unlike many existing trackers that focus on modeling only the target, in this work, we consider the transient variations of the whole scene. The key insight is that the object correspondence and spatial layout of the whole scene are consistent (i.e., global structure consistency) in consecutive frames which helps to disambiguate the target from distractors. Moreover, modeling transient variations enables to localize the target under fast variations. Specifically, we propose an effective and efficient short-term model that learns to exploit the global structure consistency in a short time and thus can handle fast variations and distractors. Since short-term modeling falls short of handling occlusion and out of the views, we adopt the long-short term paradigm and use a long-term model that corrects the short-term model when it drifts away from the target or the target is not present. These two components are carefully combined to achieve the balance of stability and plasticity during tracking. We empirically verify that the proposed tracker can tackle the two challenging scenarios and validate it on large scale benchmarks. Remarkably, our tracker improves state-of-the-art-performance on VOT2018 from 0.440 to 0.460, GOT-10k from 0.611 to 0.640, and NFS from 0.619 to 0.629. Bi Li 0005, Chengquan Zhang, Zhibin Hong, Xu Tang 0007, Jingtuo Liu, Junyu Han, Errui Ding, Wenyu Liu 0001 |
ACM Multimedia | 2 |
| 2019 | Look More Than Once: An Accurate Detector for Text of Arbitrary ShapesabstractPrevious scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challenging text instances, such as extremely long text and arbitrarily shaped text. To address these two problems, we present a novel text detector namely LOMO, which localizes the text progressively for multiple times (or in other word, LOok More than Once). LOMO consists of a direct regressor (DR), an iterative refinement module (IRM) and a shape expression module (SEM). At first, text proposals in the form of quadrangle are generated by DR branch. Next, IRM progressively perceives the entire long text by iterative refinement based on the extracted feature blocks of preliminary proposals. Finally, a SEM is introduced to reconstruct more precise representation of irregular text by considering the geometry properties of text instance, including text region, text center line and border offsets. The state-of-the-art results on several public benchmarks including ICDAR2017-RCTW, SCUT-CTW1500, Total-Text, ICDAR2015 and ICDAR17-MLT confirm the striking robustness and effectiveness of LOMO. Chengquan Zhang, Borong Liang, Zuming Huang, Mengyi En, Junyu Han, Errui Ding, Xinghao Ding |
CVPR | 1 |
| 2019 | An End-to-End Video Text Detector with Online TrackingabstractVideo text detection is considered as one of the most difficult tasks in document analysis due to the following two challenges: 1) the difficulties caused by video scenes, i.e., motion blur, illumination changes, and occlusion; 2) the properties of text including variants of fonts, languages, orientations, and shapes. Most existing methods attempt to enhance the performance of video text detection by cooperating with video text tracking, but treat these two tasks separately. In this work, we propose an end-to-end video text detection model with online tracking to address these two challenges. Specifically, in the detection branch, we adopt ConvLSTM to capture spatial structure information and motion memory. In the tracking branch, we convert the tracking problem to text instance association, and an appearance-geometry descriptor with memory mechanism is proposed to generate robust representation of text instances. By integrating these two branches into one trainable framework, they can promote each other and the computational cost is significantly reduced. Experiments on existing video text benchmarks including ICDAR2013 Video, Minetto and YVT demonstrate that the proposed method significantly outperforms state-of-the-art methods. Our method improves F-score by about 2% on all datasets and it can run realtime with 24.36 fps on TITAN Xp. Hongyuan Yu, Chengquan Zhang, Junyu Han, Errui Ding, Liang Wang 0001 |
ICDAR | 2 |
| 2019 | A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task LearningabstractDetecting scene text of arbitrary shapes has been a challenging task over the past years. In this paper, we propose a novel segmentation-based text detector, namely SAST, which employs a context attended multi-task learning framework based on a Fully Convolutional Network (FCN) to learn various geometric properties for the reconstruction of polygonal representation of text regions. Taking sequential characteristics of text into consideration, a Context Attention Block is introduced to capture long-range dependencies of pixel information to obtain a more reliable segmentation. In post-processing, a Point-to-Quad assignment method is proposed to cluster pixels into text instances by integrating both high-level object knowledge and low-level pixel information in a single shot. Moreover, the polygonal representation of arbitrarily-shaped text can be extracted with the proposed geometric properties much more effectively. Experiments on several benchmarks, including ICDAR2015, ICDAR2017-MLT, SCUT-CTW1500, and Total-Text, demonstrate that SAST achieves better or comparable performance in terms of accuracy. Furthermore, the proposed algorithm runs at 27.63 FPS on SCUT-CTW1500 with a Hmean of 81.0% on a single NVIDIA Titan Xp graphics card, surpassing most of the existing segmentation-based methods. Chengquan Zhang, Fei Qi 0001, Zuming Huang, Mengyi En, Junyu Han, Jingtuo Liu, Errui Ding, Guangming Shi |
ACM Multimedia | 2 |
| 2019 | Editing Text in the WildabstractIn this paper, we are interested in editing text in natural images, which aims to replace or modify a word in the source image with another one while maintaining its realistic look. This task is challenging, as the styles of both background and text need to be preserved so that the edited image is visually indistinguishable from the source image. Specifically, we propose an end-to-end trainable style retention network (SRNet) that consists of three modules: text conversion module, background inpainting module and fusion module. The text conversion module changes the text content of the source image into the target text while keeping the original text style. The background inpainting module erases the original text, and fills the text region with appropriate texture. The fusion module combines the information from the two former modules, and generates the edited text images. To our knowledge, this work is the first attempt to edit text in natural images at the word level. Both visual effects and quantitative results on synthetic and real-world dataset (ICDAR 2013) fully confirm the importance and necessity of modular decomposition. We also conduct extensive experiments to validate the usefulness of our method in various real-world applications such as text image synthesis, augmented reality (AR) translation, information hiding, etc. Chengquan Zhang, Jiaming Liu 0003, Junyu Han, Jingtuo Liu, Errui Ding, Xiang Bai |
ACM Multimedia | 2 |
| 2018 | Detecting Text in the Wild with Deep Character Embedding Network
Jiaming Li 0010, Chengquan Zhang, Yipeng Sun, Junyu Han, Errui Ding |
ACCV (4) | 2 |
| 2018 | TextNet: Irregular Text Reading from Images with an End-to-End Trainable Network
Yipeng Sun, Chengquan Zhang, Zuming Huang, Jiaming Liu 0003, Junyu Han, Errui Ding |
ACCV (3) | 2 |
| 2017 | WordSup: Exploiting Word Annotations for Character Based Text DetectionabstractImagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese, Japanese, mathematical expression and etc. It is natural and convenient to construct a common text detection engine based on character detectors. However, training character detectors requires a vast of location annotated characters, which are expensive to obtain. Actually, the existing real text datasets are mostly annotated in word or line level. To remedy this dilemma, we propose a weakly supervised framework that can utilize word annotations, either in tight quadrangles or the more loose bounding boxes, for character detector training. When applied in scene text detection, we are thus able to train a robust character detector by exploiting word annotations in the rich large-scale real scene text datasets, e.g. ICDAR15 [19] and COCO-text [39]. The character detector acts as a key role in the pipeline of our text detection engine. It achieves the state-of-the-art performance on several challenging scene text detection benchmarks. We also demonstrate the flexibility of our pipeline by various scenarios, including deformed text detection and math expression recognition. Chengquan Zhang, Yuxuan Luo 0002, Junyu Han, Errui Ding |
ICCV | 2 |
| 2017 | Text/non-text image classification in the wild with convolutional neural networks
Xiang Bai, Baoguang Shi, Chengquan Zhang, Xuan Cai |
Pattern Recognit. | 3 |
| 2016 | Multi-oriented Text Detection with Fully Convolutional NetworksabstractIn this paper, we propose a novel approach for text detection in natural images. Both local and global cues are taken into account for localizing text lines in a coarse-to-fine procedure. First, a Fully Convolutional Network (FCN) model is trained to predict the salient map of text regions in a holistic manner. Then, text line hypotheses are estimated by combining the salient map and character components. Finally, another FCN classifier is used to predict the centroid of each character, in order to remove the false hypotheses. The framework is general for handling text in multiple orientations, languages and fonts. The proposed method consistently achieves the state-of-the-art performance on three text detection benchmarks: MSRA-TD500, ICDAR2015 and ICDAR2013. Zheng Zhang 0022, Chengquan Zhang, Wei Shen 0002, Cong Yao, Wenyu Liu 0001, Xiang Bai |
CVPR | 2 |
| 2016 | Distinguishing text/non-text natural images with Multi-Dimensional Recurrent Neural NetworksabstractIn this paper, we focus on the text/non-text classification problem: distinguishing images that contain text from a lot of natural images. To this end, we propose a novel neural network architecture, termed Convolutional Multi-Dimensional Recurrent Neural Network (CMDRNN), which distinguishes text/non-text images by classifying local image blocks, taking both region pixels and dependencies among blocks into account. The network is composed of a Convolutional Neural Network (CNN) and a Multi-Dimensional Recurrent Neural Network (MDRNN). The CNN extracts rich and high-level image representation, while the MDRNN analyzes dependencies along multiple directions and produces block-level predictions. By evaluating CMDRNN on a public dataset, we observe improvements over prior arts in terms of both speed and accuracy. Pengyuan Lv, Baoguang Shi, Chengquan Zhang, Xiang Bai |
ICPR | 3 |
| 2016 | Symmetry-based object proposal for text detectionabstractScene text detection and recognition have become active research topics in computer vision. In this paper, we focus on the detection of text proposal from wild images. Text proposals attempt to generate a relatively small set of bounding box proposals that are most likely to contain text. Different from previous methods that merge similar region based on property of individual region, we assumed that text word bare strong symmetry property. We propose a new algorithm that exploit the symmetry property to directly generate word-level proposals. Proposals generation process using the region features, and rank process making use of the symmetry structures in text groups. Experiments on two standard datasets demonstrate that the proposed algorithm has achieve the state-of-the-art performance, especially in the case of smaller proposal number. Xuelei Zhang, Zheng Zhang 0022, Chengquan Zhang, Xiang Bai |
ICPR | 3 |
| 2016 | Traffic sign detection and recognition using fully convolutional network guided proposals
Yingying Zhu 0005, Chengquan Zhang, Duoyou Zhou, Xinggang Wang, Xiang Bai, Wenyu Liu 0001 |
Neurocomputing | 2 |
| 2015 | Automatic script identification in the wildabstractWith the rapid increase of transnational communication and cooperation, people frequently encounter multilingual scenarios in various situations. In this paper, we are concerned with a relatively new problem: script identification at word or line levels in natural scenes. A large-scale dataset with a great quantity of natural images and 10 types of widely-used languages is constructed and released. In allusion to the challenges in script identification in real-world scenarios, a deep learning based algorithm is proposed. The experiments on the proposed dataset demonstrate that our algorithm achieves superior performance, compared with conventional image classification or script identification methods, including as the original CNN architecture, LLC and GLCM. Baoguang Shi, Cong Yao, Chengquan Zhang, Feiyue Huang, Xiang Bai |
ICDAR | 3 |
| 2015 | Automatic discrimination of text and non-text natural imagesabstractWith the rapid growth of image and video data, there comes an interesting yet challenging problem: How to organize and utilize such large volume of data? Textual content in images and videos is an important source of information, which can be of great usefulness and assistance. Therefore, we investigate in this paper the problem of text image discrimination, which aims at distinguishing natural images with text from those without text. To tackle this problem, we propose a method that combines three mature techniques in this area, namely: MSER, CNN and BoW. To better evaluate the proposed algorithm, we also construct a large benchmark for text image discrimination, which includes natural images in a variety of scenarios. This algorithm has proven to be both effective and efficient, thus it can serve as a tool for mining valuable textual information from huge amount of image and video data. Chengquan Zhang, Cong Yao, Baoguang Shi, Xiang Bai |
ICDAR | 1 |