VLDB 2026 Research / reviewers in the wild / expert
Cong Yao
dblp:82/11201
· DBLP profile ↗
71ranked-venue papers
4as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 3 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 48 · 4 first-author · 22 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Granularity Prediction with Learnable Fusion for Scene Text Recognition
Cheng Da, Cong Yao |
Int. J. Comput. Vis. | 3 |
| 2025 | ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction DataabstractRecently, large language models (LLMs) and multimodal large language models (MLLMs) have demonstrated promising results on document visual question answering (VQA) task, particularly after training on document instruction datasets. An effective evaluation method for document instruction data is crucial in constructing instruction data with high efficacy, which, in turn, facilitates the training of LLMs and MLLMs for document VQA. However, most existing evaluation methods for instruction data are limited to the textual content of the instructions themselves, thereby hindering the effective assessment of document instruction datasets and constraining their construction. In this paper, we propose ProcTag, a data-oriented method that assesses the efficacy of document instruction data. ProcTag innovatively performs tagging on the execution process of instructions rather than the instruction text itself. By leveraging the diversity and complexity of these tags to assess the efficacy of the given dataset, ProcTag enables selective sampling or filtering of document instructions. Furthermore, DocLayPrompt, a novel semi-structured layout-aware document prompting strategy, is proposed for effectively representing documents. Experiments demonstrate that sampling existing open-sourced and generated document VQA/instruction datasets with ProcTag significantly outperforms current methods for evaluating instruction data. Impressively, with ProcTag-based sampling in the generated document datasets, only 30.5 percent of the document instructions are required to achieve 100 percent efficacy compared to the complete dataset. Yufan Shen, Chuwei Luo, Zhaoqing Zhu, Qi Zheng 0002, Jiajun Bu, Cong Yao |
AAAI | 8 |
| 2025 | Physics-Guided Multimodal Neural Networks for Big Data - Driven Magnetic Component Design
Jin Zhang 0018, Cong Yao, Wengen Li, Qiyou Xie, Qiuzhen Wan, Chunye Gong |
IEEE Big Data | 2 |
| 2025 | Bi-VLDoc: bidirectional vision-language modeling for visually-rich document understanding
Chuwei Luo, Guozhi Tang, Qi Zheng 0002, Cong Yao, Yang Xue 0001, Luo Si |
Int. J. Document Anal. Recognit. | 4 |
| 2025 | LORE++: Logical location regression network for table structure recognition with pre-training
Rujiao Long, Hangdi Xing, Zhibo Yang 0003, Qi Zheng 0002, Fei Huang 0002, Cong Yao |
Pattern Recognit. | 7 |
| 2025 | Generative compositor for few-shot visual information extractionabstractVisual Information Extraction (VIE), aiming at extracting structured information from visually rich document images , plays a pivotal role in document processing. Considering various layouts, semantic scopes, and languages, VIE encompasses an extensive range of types, potentially numbering in the thousands. However, many of these types suffer from a lack of training data , which poses significant challenges. In this paper, we propose a novel generative model , named Generative Compositor, to address the challenge of few-shot VIE. The Generative Compositor is a hybrid pointer-generator network that emulates the operations of a compositor by retrieving words from the source text and assembling them based on the provided prompts. Furthermore, three pre-training strategies are employed to enhance the model’s perception of spatial context information. Besides, a prompt-aware resampler is specially designed to enable efficient matching by leveraging the entity-semantic prior contained in prompts. The introduction of the prompt-based retrieval mechanism and the pre-training strategies enable the model to acquire more effective spatial and semantic clues with limited training samples . Experiments demonstrate that the proposed method achieves highly competitive results in the full-sample training, while notably outperforms the baseline in the 1-shot, 5-shot, and 10-shot settings. Zhibo Yang 0003, Wei Hua 0005, Sibo Song, Cong Yao, Yingying Zhu 0005, Wenqing Cheng, Xiang Bai |
Pattern Recognit. | 4 |
| 2025 | HierCode: A lightweight hierarchical codebook for zero-shot Chinese text recognition
Yuyi Zhang 0002, Dezhi Peng, Peirong Zhang 0001, Zhenhua Yang, Zhibo Yang 0003, Cong Yao |
Pattern Recognit. | 7 |
| 2024 | FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningabstractAutomatic font generation is an imitation task, which aims to create a font library that mimics the style of reference images while preserving the content from source images. Although existing font generation methods have achieved satisfactory performance, they still struggle with complex characters and large style variations. To address these issues, we propose FontDiffuser, a diffusion-based image-to-image one-shot font generation method, which innovatively models the font imitation task as a noise-to-denoise paradigm. In our method, we introduce a Multi-scale Content Aggregation (MCA) block, which effectively combines global and local content cues across different scales, leading to enhanced preservation of intricate strokes of complex characters. Moreover, to better manage the large variations in style transfer, we propose a Style Contrastive Refinement (SCR) module, which is a novel structure for style representation learning. It utilizes a style extractor to disentangle styles from images, subsequently supervising the diffusion model via a meticulously designed style contrastive loss. Extensive experiments demonstrate FontDiffuser's state-of-the-art performance in generating diverse characters and styles. It consistently excels on complex characters and large style changes compared to previous methods. The code is available at https://github.com/yeungchenwa/FontDiffuser. Zhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang 0002, Cong Yao |
AAAI | 5 |
| 2024 | LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingabstractRecently, leveraging large language models (LLMs) or multimodal large language models (MLLMs) for document understanding has been proven very promising. However, previous works that employ LLMs/MLLMs for document understanding have not fully explored and utilized the document layout information, which is vital for precise document understanding. In this paper, we propose LayoutLLM, an LLM/MLLM based method for document understanding. The core of LayoutLLM is a layout instruction tuning strategy, which is specially designed to enhance the comprehension and utilization of document layouts. The proposed layout instruction tuning strategy consists of two components: Layout-aware Pre-training and Layout-aware Supervised Fine-tuning. To capture the characteristics of document layout in Layout-aware Pre-training, three groups of pretraining tasks, corresponding to document-level, region-level and segment-level information, are introduced. Furthermore, a novel module called layout chain-of-thought (LayoutCoT) is devised to enable LayoutLLM to focus on regions relevant to the question and generate accurate answers. LayoutCoT is effective for boosting the performance of document understanding. Meanwhile, it brings a certain degree of interpretability, which could facilitate manual inspection and correction. Experiments on standard benchmarks show that the proposed LayoutLLM significantly outperforms existing methods that adopt open-source 7B LLMs/MLLMs for document understanding. Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng 0002, Cong Yao |
CVPR | 6 |
| 2024 | OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionabstractRecently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capable of processing document-based questions. Various methods have been proposed to address the challenging problem of VsTP. However, due to the diversified targets and heterogeneous schemas, previous works usually design task-specific architectures and objectives for individual tasks, which in- advertently leads to modal isolation and complex workflow. In this paper, we propose a unified paradigm for parsing visually-situated text across diverse scenarios. Specifically, we devise a universal model, called OmniParser, which can simultaneously handle three typical visually-situated text parsing tasks: text spotting, key information extraction, and table recognition. In OmniParser, all tasks share the unified encoder-decoder architecture, the unified objective: point- conditioned text generation, and the unified input&output representation: prompt & structured sequences. Extensive experiments demonstrate that the proposed OmniParser achieves state-of-the-art (SOTA) or highly competitive performances on 7 datasets for the three visually-situated text parsing tasks, despite its unified, concise design. The code is available at AdvancedLiterateMachinery. Jianqiang Wan, Sibo Song, Wenwen Yu, Wenqing Cheng, Fei Huang 0002, Xiang Bai, Cong Yao, Zhibo Yang 0003 |
CVPR | 8 |
| 2024 | WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
Zirui Shao, Feiyu Gao, Hangdi Xing, Zepeng Zhu, Jiajun Bu, Qi Zheng 0002, Cong Yao |
ECCV (8) | 8 |
| 2024 | Platypus: A Generalized Specialist Model for Reading Text in Various Forms
Peng Wang 0028, Zhaohai Li, Jun Tang 0008, Humen Zhong, Fei Huang 0002, Zhibo Yang 0003, Cong Yao |
ECCV (35) | 7 |
| 2024 | Visual Text Generation in the Wild
Jiawei Liu 0006, Feiyu Gao, Wenyu Liu 0001, Xinggang Wang, Peng Wang 0028, Fei Huang 0002, Cong Yao, Zhibo Yang 0003 |
ECCV (53) | 8 |
| 2024 | DocHieNet: A Large and Diverse Dataset for Document Hierarchy ParsingabstractParsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding.However, previously the research on this topic has been largely hindered since most existing datasets are small-scale, or contain documents of only a single type, which are characterized by a lack of document diversity.Moreover, there is a significant discrepancy in the annotation standards across datasets.In this paper, we introduce a large and diverse document hierarchy parsing (DHP) dataset to compensate for the data scarcity and inconsistency problem.We aim to set a new standard as a more practical, long-standing benchmark.Meanwhile, we present a new DHP framework designed to grasp both fine-grained text content and coarsegrained pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling the multi-page and multi-level challenges in DHP.Through exhaustive experiments, we validate the effectiveness of our proposed dataset and method 1 . Hangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao, Jiajun Bu, Qi Zheng 0002, Cong Yao |
EMNLP | 8 |
| 2024 | VL-Reader: Vision and Language Reconstructor is an Effective Scene Text RecognizerabstractText recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towards advanced text recognition, there are three key challenges: (1) an encoder capable of representing the visual and semantic distributions; (2) a decoder that ensures the alignment between vision and semantics; and (3) consistency in the framework during pre-training, if it exists, and fine-tuning. Inspired by masked autoencoding, a successful pre-training strategy in both vision and language, we propose an innovative scene text recognition approach, named VL-Reader. The novelty of the VL-Reader lies in the pervasive interplay between vision and language throughout the entire process. Concretely, we first introduce a Masked Visual-Linguistic Reconstruction (MVLR) objective, which aims at simultaneously modeling visual and linguistic information. Then, we design a Masked Visual-Linguistic Decoder (MVLD) to further leverage masked vision-language context and achieve bi-modal feature interaction. The architecture of VL-Reader maintains consistency from pre-training to fine-tuning. In the pre-training stage, VL-Reader reconstructs both masked visual and text tokens, while in the fine-tuning stage, the network degrades to reconstruct all characters from an image without any masked regions. VL-reader achieves an average accuracy of 97.1% on six typical datasets, surpassing the SOTA by 1.1%. The improvement was even more significant on challenging datasets. The results demonstrate that vision and language reconstructor can serve as an effective scene text recognizer. Humen Zhong, Zhibo Yang 0003, Zhaohai Li, Peng Wang 0028, Jun Tang 0008, Wenqing Cheng, Cong Yao |
ACM Multimedia | 7 |
| 2024 | Will the GDPR Restrain Health Data Access Bodies Under the European Health Data Space (EHDS)?abstractThe plans for a European Health Data Space (EHDS) envisage an ambitious and radical platform that will inter alia make the sharing of secondary health data easier. It will encourage the systematic sharing of health data and provide a legal framework for it to be shared by Health Data Access Bodies (HDABs) based in each of the Member States. Whilst this promises to bring about major benefits for research and innovation, it also raises serious questions given the intrinsic sensitivity of health data. Fears concerning privacy harms on the individual level and detrimental effects on the societal level have been raised. This article discusses two of the main protective pillars designed to allay such concerns. The first is that the proposal clearly outlines several contexts for which a Health Data Access Permit (HDAP) should and should not be granted. The second is that a request for an HDAP must also be compliant with the GDPR (inter alia requiring a valid legal basis and respecting data processing principles such as ‘minimization’ and ‘storage limitation’). As this article discusses, in some instances the need to have a valid legal basis under the GDPR may make it difficult to obtain a data access permit, in particular for some of the commercially orientated grounds outlined within the EHDS proposal. A further important issue concerns the ability of HDABs to analyse the compatibility permit requests under the GDPR and relevant national law at both speed and scale. Paul Quinn, Erika Ellyne, Cong Yao |
Comput. Law Secur. Rev. | 3 |
| 2024 | HMS-RRT: A novel hybrid multi-strategy rapidly-exploring random tree algorithm for multi-robot collaborative exploration in unknown environments
Yuming Ning, Tuanjie Li, Cong Yao, Wenqian Du 0001 |
Expert Syst. Appl. | 3 |
| 2023 | LORE: Logical Location Regression Network for Table Structure RecognitionabstractTable structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they either count on additional heuristic rules to recover the table structures, or require a huge amount of training data and time-consuming sequential decoders. In this paper, we propose an alternative paradigm. We model TSR as a logical location regression problem and propose a new TSR framework called LORE, standing for LOgical location REgression network, which for the first time combines logical location regression together with spatial location regression of table cells. Our proposed LORE is conceptually simpler, easier to train and more accurate than previous TSR models of other paradigms. Experiments on standard benchmarks demonstrate that LORE consistently outperforms prior arts. Code is available at https:// github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/DocumentUnderstanding/LORE-TSR. Hangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu, Qi Zheng 0002, Liangcheng Li, Cong Yao |
AAAI | 7 |
| 2023 | Building A Mobile Text Recognizer via Truncated SVD-based Knowledge Distillation-Guided NAS
Weifeng Lin, Canyu Xie, Dezhi Peng, Cong Yao, Mengchao He |
BMVC | 7 |
| 2023 | GeoLayoutLM: Geometric Pre-training for Visual Information ExtractionabstractVisual information extraction (VIE) plays an important role in Document Intelligence. Generally, it is divided into two tasks: semantic entity recognition (SER) and relation extraction (RE). Recently, pre-trained models for documents have achieved substantial progress in VIE, particularly in SER. However, most of the existing models learn the geometric representation in an implicit way, which has been found insufficient for the RE task since geometric information is especially crucial for RE. Moreover, we reveal another factor that limits the performance of RE lies in the objective gap between the pre-training phase and the finetuning phase for RE. To tackle these issues, we propose in this paper a multi-modal framework, named GeoLayoutLM, for VIE. GeoLayoutLM explicitly models the geometric relations in pre-training, which we call geometric pre-training. Geometric pre-training is achieved by three specially designed geometry-related pre-training tasks. Additionally, novel relation heads, which are pre-trained by the geometric pre-training tasks and fine-tuned for RE, are elaborately designed to enrich and enhance the feature representation. According to extensive experiments on standard VIE benchmarks, GeoLayoutLM achieves highly competitive scores in the SER task and significantly outperforms the previous state-of-the-arts for RE (e.g., the F1 score of RE on FUNSD is boosted from 80.35% to 89.45%)11https://github.com/AlibabaResearch/AdvancedLiterateMachinery. Chuwei Luo, Changxu Cheng, Qi Zheng 0002, Cong Yao |
CVPR | 4 |
| 2023 | Modeling Entities as Semantic Points for Visual Information Extraction in the WildabstractRecently, Visual Information Extraction (VIE) has been becoming increasingly important in both the academia and industry, due to the wide range of real-world applications. Previously, numerous works have been proposed to tackle this problem. However, the benchmarks used to assess these methods are relatively plain, i.e., scenarios with real-world complexity are not fully represented in these benchmarks. As the first contribution of this work, we curate and release a new dataset for VIE, in which the document images are much more challenging in that they are taken from real applications, and difficulties such as blur, partial occlusion, and printing shift are quite common. All these factors may lead to failures in information extraction. Therefore, as the second contribution, we explore an alternative approach to precisely and robustly extract key information from document images under such tough conditions. Specifically, in contrast to previous methods, which usually either incorporate visual information into a multi-modal architecture or train text spotting and information extraction in an end-to-end fashion, we explicitly model entities as semantic points, i.e., center points of entities are enriched with semantic information describing the attributes and relationships of different entities, which could largely benefit entity labeling and linking. Extensive experiments on standard benchmarks in this field as well as the proposed dataset demonstrate that the proposed method can achieve significantly enhanced performance on entity labeling and linking, compared with previous state-of-the-art models. Dataset is available at https://www.modelscope.cn/datasets/damo/SIBR/summary. Zhibo Yang 0003, Rujiao Long, Sibo Song, Humen Zhong, Wenqing Cheng, Xiang Bai, Cong Yao |
CVPR | 8 |
| 2023 | Conditional Text Image Generation with Diffusion ModelsabstractCurrent text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and annotating enough real text images. In this paper, we explore the problem of text image generation, by taking advantage of the powerful abilities of Diffusion Models in generating photo-realistic and diverse image samples with given conditions, and propose a method called Conditional Text Image Generation with Diffusion Models (CTIG-DM for short). To conform to the characteristics of text images, we devise three conditions: image condition, text condition, and style condition, which can be used to control the attributes, contents, and styles of the samples in the image generation process. Specifically, four text image generation modes, namely: (1) synthesis mode, (2) augmentation mode, (3) recovery mode, and (4) imitation mode, can be derived by combining and configuring these three conditions. Extensive experiments on both handwritten and scene text demonstrate that the proposed CTIG-DM is able to produce image samples that simulate real-world complexity and diversity, and thus can boost the performance of existing text recognizers. Besides, CTIG-DM shows its appealing potential in domain adaptation and generating images containing Out-Of-Vocabulary (OOV) words. Zhaohai Li, Mengchao He, Cong Yao |
CVPR | 5 |
| 2023 | LISTER: Neighbor Decoding for Length-Insensitive Scene Text RecognitionabstractThe diversity in length constitutes a significant characteristic of text. Due to the long-tail distribution of text lengths, most existing methods for scene text recognition (STR) only work well on short or seen-length text, lacking the capability of recognizing longer text or performing length extrapolation. This is a crucial issue, since the lengths of the text to be recognized are usually not given in advance in real-world applications, but it has not been adequately investigated in previous works. Therefore, we propose in this paper a method called Length-Insensitive Scene TExt Recognizer (LISTER), which remedies the limitation regarding the robustness to various text lengths. Specifically, a Neighbor Decoder is proposed to obtain accurate character attention maps with the assistance of a novel neighbor matrix regardless of the text lengths. Besides, a Feature Enhancement Module is devised to model the long-range dependency with low computation cost, which is able to perform iterations with the neighbor decoder to enhance the feature map progressively. To the best of our knowledge, we are the first to achieve effective length-insensitive scene text recognition. Extensive experiments demonstrate that the proposed LISTER algorithm exhibits obvious superiority on long text recognition and the ability for length extrapolation, while comparing favourably with the previous state-of-the-art methods on standard benchmarks for STR (mainly short text)1. Changxu Cheng, Peng Wang 0103, Cheng Da, Qi Zheng 0002, Cong Yao |
ICCV | 5 |
| 2023 | Vision Grid Transformer for Document Layout AnalysisabstractDocument pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those pretrained in a multi-modal fashion, usually rely on either textual features or visual features. Grid-based models for DLA are multi-modality but largely neglect the effect of pre-training. To fully leverage multi-modal information and exploit pre-training techniques to learn better representation for DLA, in this paper, we present VGT, a two-stream Vision Grid Transformer, in which Grid Transformer (GiT) is proposed and pre-trained for 2D token-level and segment-level semantic understanding. Furthermore, a new dataset named D4LA, which is so far the most diverse and detailed manually-annotated benchmark for document layout analysis, is curated and released. Experiment results have illustrated that the proposed VGT model achieves new state-of-the-art results on DLA tasks, e.g. PubLayNet (95.7%→96.2%), DocBank (79.6%→84.1%), and D4LA (67.7%→68.8%). The code and models as well as the D4LA dataset will be made publicly available1. Cheng Da, Chuwei Luo, Qi Zheng 0002, Cong Yao |
ICCV | 4 |
| 2023 | ICDAR 2023 Competition on Born Digital Video Text Question Answering
Zhibo Yang 0003, Xiaoge Song, Sibo Song, Tong Lu 0002, Xiang Bai, Cheng-Lin Liu 0001, Fei Huang 0002, Cong Yao |
ICDAR (2) | 8 |
| 2023 | Read Ten Lines at One Glance: Line-Aware Semi-Autoregressive Transformer for Multi-Line Handwritten Mathematical Expression RecognitionabstractHandwritten Mathematical Expression Recognition (HMER) plays a critical role in various applications, such as digitized education and scientific research. Although existing methods have achieved promising performance on publicly available datasets, they still struggle to recognize multi-line mathematical expressions (MEs), suffering from complex structures and slow inference speed. To address these issues, we propose a Line-Aware Semi-autoregressive Transformer (LAST) that treats multi-line mathematical expression sequences as two-dimensional dual-end structures. The proposed LAST utilizes a line-wise dual-end decoding strategy to decode multi-line mathematical expressions in parallel and perform dual-end decoding within each line. Specifically, we introduce a line-aware positional encoding module and a line-partitioned dual-end mask to endow LAST with line order awareness and directionality. Additionally, we adopt a shared-task optimization strategy to train LAST in both autoregressive and semi-autoregressive tasks. To evaluate the effectiveness of our approach in real-world scenarios, we have built a new Multi-line Mathematical Expression dataset (M2E), which, to the best of our knowledge, is the first of its kind and boasts with the largest character category, the largest samples of characters, and the longest average sequence length, compared to existing ME datasets. Experimental results on both the M2E dataset and publicly available datasets demonstrate the effectiveness of our proposed method. Notably, our semi-autoregressive decoding approach achieves significantly faster decoding speeds while still achieving state-of-the-art performance compared to the existing methods. Wentao Yang 0003, Zhe Li 0046, Dezhi Peng, Mengchao He, Cong Yao |
ACM Multimedia | 6 |
| 2023 | Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale FusionabstractRecently, segmentation-based scene text detection methods have drawn extensive attention in the scene text detection field, because of their superiority in detecting the text instances of arbitrary shapes and extreme aspect ratios, profiting from the pixel-level descriptions. However, the vast majority of the existing segmentation-based approaches are limited to their complex post-processing algorithms and the scale robustness of their segmentation models, where the post-processing algorithms are not only isolated to the model optimization but also time-consuming and the scale robustness is usually strengthened by fusing multi-scale feature maps directly. In this paper, we propose a Differentiable Binarization (DB) module that integrates the binarization process, one of the most important steps in the post-processing procedure, into a segmentation network. Optimized along with the proposed DB module, the segmentation network can produce more accurate results, which enhances the accuracy of text detection with a simple pipeline. Furthermore, an efficient Adaptive Scale Fusion (ASF) module is proposed to improve the scale robustness by fusing features of different scales adaptively. By incorporating the proposed DB and ASF with the segmentation network, our proposed scene text detector consistently achieves state-of-the-art results, in terms of both detection accuracy and speed, on five standard benchmarks. Minghui Liao, Zhisheng Zou, Zhaoyi Wan, Cong Yao, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Revisiting Document Image Dewarping by Grid RegularizationabstractThis paper addresses the problem of document image dewarping, which aims at eliminating the geometric distortion in document images for document digitization. Instead of designing a better neural network to approximate the optical flow fields between the inputs and outputs, we pursue the best readability by taking the text lines and the document boundaries into account from a constrained optimization perspective. Specifically, our proposed method first learns the boundary points and the pixels in the text lines and then follows the most simple observation that the boundaries and text lines in both horizontal and vertical directions should be kept after dewarping to introduce a novel grid regularization scheme. To obtain the final forward mapping for dewarping, we solve an optimization problem with our proposed grid regularization. The experiments comprehensively demonstrate that our proposed approach outperforms the prior arts by large margins in terms of readability (with the metrics of Character Errors Rate and the Edit Distance) while maintaining the best image quality on the publicly-available DocUNet benchmark. Xiangwei Jiang, Rujiao Long, Nan Xue 0001, Zhibo Yang 0003, Cong Yao, Gui-Song Xia |
CVPR | 5 |
| 2022 | Vision-Language Pre-Training for Boosting Scene Text DetectorsabstractRecently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning for scene text detection, a task that intrinsically involves cross-modal interaction between the two modalities: vision and language, since text is the written form of language. Concretely, we propose to learn contextualized, joint representations through vision-language pretraining, for the sake of enhancing the performance of scene text detectors. Towards this end, we devise a pre-training architecture with an image encoder, a text encoder and a cross-modal encoder, as well as three pretext tasks: image-text contrastive learning (ITC), masked language modeling (MLM) and word-in-image prediction (WIP). The pretrained model is able to produce more informative representations with richer semantics, which could readily benefit existing scene text detectors (such as EAST and PSENet) in the down-stream text detection task. Extensive experiments on standard benchmarks demonstrate that the proposed paradigm can significantly improve the performance of various representative text detectors, outperforming previous pre-training approaches. The code and pre-trained models will be publicly released. Sibo Song, Jianqiang Wan, Zhibo Yang 0003, Jun Tang 0008, Wenqing Cheng, Xiang Bai, Cong Yao |
CVPR | 7 |
| 2022 | Levenshtein OCR
Cheng Da, Peng Wang 0103, Cong Yao |
ECCV (28) | 3 |
| 2022 | Multi-granularity Prediction for Scene Text Recognition
Peng Wang 0103, Cheng Da, Cong Yao |
ECCV (28) | 3 |
| 2022 | Facial Attribute Transformers for Precise and Robust Makeup TransferabstractIn this paper, we address the problem of makeup transfer, which aims at transplanting the makeup from the reference face to the source face while preserving the identity of the source. Existing makeup transfer methods have made notable progress in generating realistic makeup faces, but do not perform well in terms of color fidelity and spatial transformation. To tackle these issues, we propose a novel Facial Attribute Transformer (FAT) and its variant Spatial FAT for high-quality makeup transfer. Drawing inspirations from the Transformer in NLP, FAT is able to model the semantic correspondences and interactions between the source face and reference face, and then precisely estimate and transfer the facial attributes. To further facilitate shape deformation and transformation of facial parts, we also integrate thin plate splines (TPS) into FAT, thus creating Spatial FAT, which is the first method that can transfer geometric attributes in addition to color and texture. Extensive qualitative and quantitative experiments demonstrate the effectiveness and superiority of our proposed FATs in the following aspects: (1) ensuring high-fidelity color transfer; (2) allowing for geometric transformation of facial parts; (3) handling facial variations (such as poses and shadows) and (4) supporting high-resolution face generation. Zhaoyi Wan, Jie An 0002, Cong Yao, Jiebo Luo 0001 |
WACV | 5 |
| 2021 | MOST: A Multi-Oriented Scene Text Detector With Localization RefinementabstractOver the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they might still fall short when handling text instances of extreme aspect ratios and varying scales. To tackle such difficulties, we propose in this paper a new algorithm for scene text detection, which puts forward a set of strategies to significantly improve the quality of text localization. Specifically, a Text Feature Alignment Module (TFAM) is proposed to dynamically adjust the receptive fields of features based on initial raw detections; a Position-Aware Non-Maximum Suppression (PA-NMS) module is devised to selectively concentrate on reliable raw detections and exclude unreliable ones; besides, we propose an Instance-wise IoU loss for balanced training to deal with text instances of different scales. An extensive ablation study demonstrates the effectiveness and superiority of the proposed strategies. The resulting text detection system, which integrates the proposed strategies with a leading scene text detector EAST, achieves state-of-the-art or competitive performance on various standard benchmarks for text detection while keeping a fast running speed. Minghang He, Minghui Liao, Zhibo Yang 0003, Humen Zhong, Jun Tang 0008, Wenqing Cheng, Cong Yao, Yongpan Wang, Xiang Bai |
CVPR | 7 |
| 2021 | Scene Text Detection and Recognition: The Deep Learning Era
Shangbang Long, Cong Yao |
Int. J. Comput. Vis. | 3 |
| 2021 | Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary ShapesabstractUnifying text detection and text recognition in an end-to-end training fashion has become a new trend for reading text in the wild, as these two tasks are highly relevant and complementary. In this paper, we investigate the problem of scene text spotting, which aims at simultaneous text detection and recognition in natural images. An end-to-end trainable neural network named as Mask TextSpotter is presented. Different from the previous text spotters that follow the pipeline consisting of a proposal generation network and a sequence-to-sequence recognition network, Mask TextSpotter enjoys a simple and smooth end-to-end learning procedure, in which both detection and recognition can be achieved directly from two-dimensional space via semantic segmentation. Further, a spatial attention module is proposed to enhance the performance and universality. Benefiting from the proposed two-dimensional representation on both detection and recognition, it easily handles text instances of irregular shapes, for instance, curved text. We evaluate it on four English datasets and one multi-language dataset, achieving consistently superior performance over state-of-the-art methods in both detection and end-to-end text recognition tasks. Moreover, we further investigate the recognition module of our method separately, which significantly outperforms state-of-the-art methods on both regular and irregular text datasets for scene text recognition. Minghui Liao, Pengyuan Lv, Minghang He, Cong Yao, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Real-Time Scene Text Detection with Differentiable BinarizationabstractRecently, segmentation-based methods are quite popular in scene text detection, as the segmentation results can more accurately describe scene text of various shapes such as curve text. However, the post-processing of binarization is essential for segmentation-based detection, which converts probability maps produced by a segmentation method into bounding boxes/regions of text. In this paper, we propose a module named Differentiable Binarization (DB), which can perform the binarization process in a segmentation network. Optimized along with a DB module, a segmentation network can adaptively set the thresholds for binarization, which not only simplifies the post-processing but also enhances the performance of text detection. Based on a simple segmentation network, we validate the performance improvements of DB on five benchmark datasets, which consistently achieves state-of-the-art results, in terms of both detection accuracy and speed. In particular, with a light-weight backbone, the performance improvements by DB are significant so that we can look for an ideal tradeoff between detection accuracy and efficiency. Specifically, with a backbone of ResNet-18, our detector achieves an F-measure of 82.8, running at 62 FPS, on the MSRA-TD500 dataset. Code is available at: https://github.com/MhLiao/DB. Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 0006, Xiang Bai |
AAAI | 3 |
| 2020 | TextScanner: Reading Characters in Order for Robust Scene Text RecognitionabstractDriven by deep learning and a large volume of data, scene text recognition has evolved rapidly in recent years. Formerly, RNN-attention-based methods have dominated this field, but suffer from the problem of attention drift in certain situations. Lately, semantic segmentation based algorithms have proven effective at recognizing text of different forms (horizontal, oriented and curved). However, these methods may produce spurious characters or miss genuine characters, as they rely heavily on a thresholding procedure operated on segmentation maps. To tackle these challenges, we propose in this paper an alternative approach, called TextScanner, for scene text recognition. TextScanner bears three characteristics: (1) Basically, it belongs to the semantic segmentation family, as it generates pixel-wise, multi-channel segmentation maps for character class, position and order; (2) Meanwhile, akin to RNN-attention-based methods, it also adopts RNN for context modeling; (3) Moreover, it performs paralleled prediction for character position and class, and ensures that characters are transcripted in the correct order. The experiments on standard benchmark datasets demonstrate that TextScanner outperforms the state-of-the-art methods. Moreover, TextScanner shows its superiority in recognizing more difficult text such as Chinese transcripts and aligning with target characters. Zhaoyi Wan, Minghang He, Xiang Bai, Cong Yao |
AAAI | 5 |
| 2020 | On Vocabulary Reliance in Scene Text RecognitionabstractThe pursuit of high performance on public benchmarks has been the driving force for research in scene text recognition, and notable progresses have been achieved. However, a close investigation reveals a startling fact that the state-of-the-art methods perform well on images with words within vocabulary but generalize poorly to images with words outside vocabulary. We call this phenomenon ``vocabulary reliance''. In this paper, we establish an analytical framework, in which different datasets, metrics and module combinations for quantitative comparisons are devised, to conduct an in-depth study on the problem of vocabulary reliance in scene text recognition. Key findings include: (1) Vocabulary reliance is ubiquitous, i.e., all existing algorithms more or less exhibit such characteristic; (2) Attention-based decoders prove weak in generalizing to words outside vocabulary and segmentation-based decoders perform well in utilizing visual features; (3) Context modeling is highly coupled with the prediction layers. These findings provide new insights and can benefit future research in scene text recognition. Furthermore, we propose a simple yet effective mutual learning strategy to allow models of two families (attention-based and segmentation-based) to learn collaboratively. This remedy alleviates the problem of vocabulary reliance and significantly improves the overall scene text recognition performance. Zhaoyi Wan, Jielei Zhang, Jiebo Luo 0001, Cong Yao |
CVPR | 5 |
| 2020 | Differentiable Feature Aggregation Search for Knowledge Distillation
Yushuo Guan, Bingxuan Wang, Yuanxing Zhang, Cong Yao, Kaigui Bian, Jian Tang 0008 |
ECCV (17) | 5 |
| 2020 | A New Perspective for Flexible Feature Gathering in Scene Text Recognition Via Character Anchor PoolingabstractIrregular scene text recognition has attracted much attention from the research community, mainly due to the complexity of shapes of text in natural scene. However, recent methods either rely on shape-sensitive modules such as bounding box regression, or discard sequence learning. To tackle these issues, we propose a pair of coupling modules, termed as Character Anchoring Module (CAM) and Anchor Pooling Module (APM), to extract high-level semantics from two-dimensional space to form feature sequences. The proposed CAM localizes the text in a shape-insensitive way by design by anchoring characters individually. APM then interpolates and gathers features flexibly along the character anchors which enables sequence learning. The complementary modules realize a harmonic unification of spatial information and sequence learning. With the proposed modules, our recognition system surpasses previous state-of-the-art scores on irregular and perspective text datasets, including, ICDAR 2015, CUTE, and Total-Text, while paralleling state-of-the-art performance on regular text datasets. Shangbang Long, Yushuo Guan, Kaigui Bian, Cong Yao |
ICASSP | 4 |
| 2020 | SynthText3D: synthesizing scene text images from 3D virtual worlds
Minghui Liao, Boyu Song, Shangbang Long, Minghang He, Cong Yao, Xiang Bai |
Sci. China Inf. Sci. | 5 |
| 2020 | Self-Similarity Action ProposalabstractTemporal action proposal generation, which aims to locate temporal segments that may contain actions, is a key prepositive step of various video analysis tasks, like temporal action detection. In this letter, we present Self-Similarity Action Proposal (SSAP), a simple method that generates action proposals using the self-similarity of videos. Specifically, a basic low-level index, structural similarity, is adopted to measure the similarity between adjacent frames. Potential action boundaries are located by thresholding the similarity values and candidate action segments are successively generated by grouping the boundaries. A segment evaluation module (SEM) is further employed to score and refine the segments. The framework achieves state-of-the-art performance on THUMOS14 and competitive results on ActivityNet v1.3. Notably, on THUMOS14, it achieves over 4% improvement on the average recall at 50 proposals and 3.3% gain in [email protected] when combined with an existing action classifier for temporal action detection. Yuchao Sun, Jianghu Lu, Cong Yao, Yu Zhou 0016 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Scene Text Recognition from Two-Dimensional PerspectiveabstractInspired by speech recognition, recent state-of-the-art algorithms mostly consider scene text recognition as a sequence prediction problem. Though achieving excellent performance, these methods usually neglect an important fact that text in images are actually distributed in two-dimensional space. It is a nature quite different from that of speech, which is essentially a one-dimensional signal. In principle, directly compressing features of text into a one-dimensional form may lose useful information and introduce extra noise. In this paper, we approach scene text recognition from a two-dimensional perspective. A simple yet effective model, called Character Attention Fully Convolutional Network (CA-FCN), is devised for recognizing the text of arbitrary shapes. Scene text recognition is realized with a semantic segmentation network, where an attention mechanism for characters is adopted. Combined with a word formation module, CA-FCN can simultaneously recognize the script and predict the position of each character. Experiments demonstrate that the proposed algorithm outperforms previous methods on both regular and irregular text datasets. Moreover, it is proven to be more robust to imprecise localizations in the text detection phase, which are very common in practice. Minghui Liao, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lv, Cong Yao, Xiang Bai |
AAAI | 7 |
| 2019 | Scene Text Detection with Supervised Pyramid Context NetworkabstractScene text detection methods based on deep learning have achieved remarkable results over the past years. However, due to the high diversity and complexity of natural scenes, previous state-of-the-art text detection methods may still produce a considerable amount of false positives, when applied to images captured in real-world environments. To tackle this issue, mainly inspired by Mask R-CNN, we propose in this paper an effective model for scene text detection, which is based on Feature Pyramid Network (FPN) and instance segmentation. We propose a supervised pyramid context network (SPCNET) to precisely locate text regions while suppressing false positives.Benefited from the guidance of semantic information and sharing FPN, SPCNET obtains significantly enhanced performance while introducing marginal extra computation. Experiments on standard datasets demonstrate that our SPCNET clearly outperforms start-of-the-art methods. Specifically, it achieves an F-measure of 92.1% on ICDAR2013, 87.2% on ICDAR2015, 74.1% on ICDAR2017 MLT and 82.9% on Enze Xie, Yuhang Zang, Shuai Shao 0005, Gang Yu 0002, Cong Yao |
AAAI | 5 |
| 2019 | Symmetry-Constrained Rectification Network for Scene Text RecognitionabstractReading text in the wild is a very challenging task due to the diversity of text instances and the complexity of natural scenes. Recently, the community has paid increasing attention to the problem of recognizing text instances with irregular shapes. One intuitive and effective way to handle this problem is to rectify irregular text to a canonical form before recognition. However, these methods might struggle when dealing with highly curved or distorted text instances. To tackle this issue, we propose in this paper a Symmetry-constrained Rectification Network (ScRN) based on local attributes of text instances, such as center line, scale and orientation. Such constraints with an accurate description of text shape enable ScRN to generate better rectification results than existing methods and thus lead to higher recognition accuracy. Our method achieves state-of-the-art performance on text with both regular and irregular shapes. Specifically, the system outperforms existing algorithms by a large margin on datasets that contain quite a proportion of irregular text instances, e.g., ICDAR 2015, SVT-Perspective and CUTE80. Yushuo Guan, Minghui Liao, Kaigui Bian, Song Bai 0001, Cong Yao, Xiang Bai |
ICCV | 7 |
| 2019 | ASTER: An Attentional Scene Text Recognizer with Flexible RectificationabstractA challenging aspect of scene text recognition is to handle text with distortions or irregular layout. In particular, perspective text and curved text are common in natural scenes and are difficult to recognize. In this work, we introduce ASTER, an end-to-end neural network model that comprises a rectification network and a recognition network. The rectification network adaptively transforms an input image into a new one, rectifying the text in it. It is powered by a flexible Thin-Plate Spline transformation which handles a variety of text irregularities and is trained without human annotations. The recognition network is an attentional sequence-to-sequence model that predicts a character sequence directly from the rectified image. The whole model is trained end to end, requiring only images and their groundtruth text. Through extensive experiments, we verify the effectiveness of the rectification and demonstrate the state-of-the-art recognition performance of ASTER. Furthermore, we demonstrate that ASTER is a powerful component in end-to-end recognition systems, for its ability to enhance the detector. Baoguang Shi, Xinggang Wang, Pengyuan Lv, Cong Yao, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Multi-Oriented Scene Text Detection via Corner Localization and Region SegmentationabstractPrevious deep learning based state-of-the-art scene text detection methods can be roughly classified into two categories. The first category treats scene text as a type of general objects and follows general object detection paradigm to localize scene text by regressing the text box locations, but troubled by the arbitrary-orientation and large aspect ratios of scene text. The second one segments text regions directly, but mostly needs complex post processing. In this paper, we present a method that combines the ideas of the two types of methods while avoiding their shortcomings. We propose to detect scene text by localizing corner points of text bounding boxes and segmenting text regions in relative positions. In inference stage, candidate boxes are generated by sampling and grouping corner points, which are further scored by segmentation maps and suppressed by NMS. Compared with previous methods, our method can handle long oriented text naturally and doesn't need complex post processing. The experiments on ICDAR2013, ICDAR2015, MSRA-TD500, MLT and COCO-Text demonstrate that the proposed algorithm achieves better or comparable results in both accuracy and efficiency. Based on VGG16, it achieves an F-measure of 84.3% on ICDAR2015 and 81.5% on MSRA-TD500. Pengyuan Lv, Cong Yao, Shuicheng Yan, Xiang Bai |
CVPR | 2 |
| 2018 | TextSnake: A Flexible Representation for Detecting Text of Arbitrary Shapes
Shangbang Long, Jiaqiang Ruan, Cong Yao |
ECCV (2) | 6 |
| 2018 | Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
Pengyuan Lv, Minghui Liao, Cong Yao, Xiang Bai |
ECCV (14) | 3 |
| 2017 | EAST: An Efficient and Accurate Scene Text DetectorabstractPrevious approaches for scene text detection have already achieved promising performances across various benchmarks. However, they usually fall short when dealing with challenging scenarios, even when equipped with deep neural network models, because the overall performance is determined by the interplay of multiple stages and components in the pipelines. In this work, we propose a simple yet powerful pipeline that yields fast and accurate text detection in natural scenes. The pipeline directly predicts words or text lines of arbitrary orientations and quadrilateral shapes in full images, eliminating unnecessary intermediate steps (e.g., candidate aggregation and word partitioning), with a single neural network. The simplicity of our pipeline allows concentrating efforts on designing loss functions and neural network architecture. Experiments on standard datasets including ICDAR 2015, COCO-Text and MSRA-TD500 demonstrate that the proposed algorithm significantly outperforms state-of-the-art methods in terms of both accuracy and efficiency. On the ICDAR 2015 dataset, the proposed algorithm achieves an F-score of 0.7820 at 13.2fps at 720p resolution. Xinyu Zhou 0004, Cong Yao, Yuzhi Wang, Shuchang Zhou 0001, Weiran He, Jiajun Liang |
CVPR | 2 |
| 2017 | Auto-Encoder Guided GAN for Chinese Calligraphy SynthesisabstractIn this paper, we investigate the Chinese calligraphy synthesis problem: synthesizing Chinese calligraphy images with specified style from standard font(eg. Hei font) images (Fig. 1(a)). Recent works mostly follow the stroke extraction and assemble pipeline which is complex in the process and limited by the effect of stroke extraction. In this work we treat the calligraphy synthesis problem as an image-to-image translation problem and propose a deep neural network based model which can generate calligraphy images from standard font images directly. Besides, we also construct a large scale benchmark that contains various styles for Chinese calligraphy synthesis. We evaluate our method as well as some baseline methods on the proposed dataset, and the experimental results demonstrate the effectiveness of our proposed model. Pengyuan Lv, Xiang Bai, Cong Yao, Zhen Zhu 0006, Tengteng Huang, Wenyu Liu 0001 |
ICDAR | 3 |
| 2017 | ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)abstractChinese is the most widely used language in the world. Algorithms that read Chinese text in natural images facilitate applications of various kinds. Despite the large potential value, datasets and competitions in the past primarily focus on English, which bares very different characteristics than Chinese. This report introduces RCTW, a new competition that focuses on Chinese text reading. The competition features a large-scale dataset with over 12,000 annotated images. Two tasks, namely text localization and end-to-end recognition, are set up. The competition took place from January 20 to May 31, 2017. 23 valid submissions were received from 19 teams. This report includes dataset description, task definitions, evaluation protocols, and results summaries and analysis. Through this competition, we call for more future research on the Chinese text reading problem. Baoguang Shi, Cong Yao, Minghui Liao, Pei Xu 0006, Linyan Cui, Serge J. Belongie, Shijian Lu, Xiang Bai |
ICDAR | 2 |
| 2017 | An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text RecognitionabstractImage-based sequence recognition has been a long-standing research topic in computer vision. In this paper, we investigate the problem of scene text recognition, which is among the most important and challenging tasks in image-based sequence recognition. A novel neural network architecture, which integrates feature extraction, sequence modeling and transcription into a unified framework, is proposed. Compared with previous systems for scene text recognition, the proposed architecture possesses four distinctive properties: (1) It is end-to-end trainable, in contrast to most of the existing algorithms whose components are separately trained and tuned. (2) It naturally handles sequences in arbitrary lengths, involving no character segmentation or horizontal scale normalization. (3) It is not confined to any predefined lexicon and achieves remarkable performances in both lexicon-free and lexicon-based scene text recognition tasks. (4) It generates an effective yet much smaller model, which is more practical for real-world application scenarios. The experiments on standard benchmarks, including the IIIT-5K, Street View Text and ICDAR datasets, demonstrate the superiority of the proposed algorithm over the prior arts. Moreover, the proposed algorithm performs well in the task of image-based music score recognition, which evidently verifies the generality of it. Baoguang Shi, Xiang Bai, Cong Yao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Robust Scene Text Recognition with Automatic RectificationabstractRecognizing text in natural images is a challenging task with many unsolved problems. Different from those in documents, words in natural images often possess irregular shapes, which are caused by perspective distortion, curved character placement, etc. We propose RARE (Robust text recognizer with Automatic REctification), a recognition model that is robust to irregular text. RARE is a speciallydesigned deep neural network, which consists of a Spatial Transformer Network (STN) and a Sequence Recognition Network (SRN). In testing, an image is firstly rectified via a predicted Thin-Plate-Spline (TPS) transformation, into a more "readable" image for the following SRN, which recognizes text through a sequence recognition approach. We show that the model is able to recognize several types of irregular text, including perspective text and curved text. RARE is end-to-end trainable, requiring only images and associated text labels, making it convenient to train and deploy the model in practical systems. State-of-the-art or highly-competitive performance achieved on several benchmarks well demonstrates the effectiveness of the proposed model. Baoguang Shi, Xinggang Wang, Pengyuan Lv, Cong Yao, Xiang Bai |
CVPR | 4 |
| 2016 | Multi-oriented Text Detection with Fully Convolutional NetworksabstractIn this paper, we propose a novel approach for text detection in natural images. Both local and global cues are taken into account for localizing text lines in a coarse-to-fine procedure. First, a Fully Convolutional Network (FCN) model is trained to predict the salient map of text regions in a holistic manner. Then, text line hypotheses are estimated by combining the salient map and character components. Finally, another FCN classifier is used to predict the centroid of each character, in order to remove the false hypotheses. The framework is general for handling text in multiple orientations, languages and fonts. The proposed method consistently achieves the state-of-the-art performance on three text detection benchmarks: MSRA-TD500, ICDAR2015 and ICDAR2013. Zheng Zhang 0022, Chengquan Zhang, Wei Shen 0002, Cong Yao, Wenyu Liu 0001, Xiang Bai |
CVPR | 4 |
| 2016 | Scene text detection and recognition: recent advances and future trends
Yingying Zhu 0005, Cong Yao, Xiang Bai |
Frontiers Comput. Sci. | 2 |
| 2016 | Deep Learning Representation using Autoencoder for 3D Shape Retrieval
Zhuotun Zhu, Xinggang Wang, Song Bai 0001, Cong Yao, Xiang Bai |
Neurocomputing | 4 |
| 2016 | Script identification in the wild via discriminative convolutional neural network
Baoguang Shi, Xiang Bai, Cong Yao |
Pattern Recognit. | 3 |
| 2016 | Strokelets: A Learned Multi-Scale Mid-Level Representation for Scene Text RecognitionabstractIn this paper, we are concerned with the problem of automatic scene text recognition, which involves localizing and reading characters in natural images. We investigate this problem from the perspective of representation and propose a novel multi-scale representation, which leads to accurate, robust character identification and recognition. This representation consists of a set of mid-level primitives, termed strokelets, which capture the underlying substructures of characters at different granularities. The Strokelets possess four distinctive advantages: 1) usability: automatically learned from character level annotations; 2) robustness: insensitive to interference factors; 3) generality: applicable to variant languages; and 4) expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of the strokelets and demonstrate the effectiveness of the text recognition algorithm built upon the strokelets. Moreover, we show the method to incorporate the strokelets to improve the performance of scene text detection. Xiang Bai, Cong Yao, Wenyu Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Symmetry-based text line detection in natural scenesabstractRecently, a variety of real-world applications have triggered huge demand for techniques that can extract textual information from natural scenes. Therefore, scene text detection and recognition have become active research topics in computer vision. In this work, we investigate the problem of scene text detection from an alternative perspective and propose a novel algorithm for it. Different from traditional methods, which mainly make use of the properties of single characters or strokes, the proposed algorithm exploits the symmetry property of character groups and allows for direct extraction of text lines from natural images. The experiments on the latest ICDAR benchmarks demonstrate that the proposed algorithm achieves state-of-the-art performance. Moreover, compared to conventional approaches, the proposed algorithm shows stronger adaptability to texts in challenging scenarios. Zheng Zhang 0022, Wei Shen 0002, Cong Yao, Xiang Bai |
CVPR | 3 |
| 2015 | Relaxed Multiple-Instance SVM with Application to Object DiscoveryabstractMultiple-instance learning (MIL) has served as an important tool for a wide range of vision applications, for instance, image classification, object detection, and visual tracking. In this paper, we propose a novel method to solve the classical MIL problem, named relaxed multiple-instance SVM (RMI-SVM). We treat the positiveness of instance as a continuous variable, use Noisy-OR model to enforce the MIL constraints, and optimize them jointly in a unified framework. The optimization problem can be efficiently solved using stochastic gradient decent. The extensive experiments demonstrate that RMI-SVM consistently achieves superior performance on various benchmarks for MIL. Moreover, we simply applied RMI-SVM to a challenging vision task, common object discovery. The state-of-the arts results of object discovery on PASCAL VOC datasets further confirm the advantages of the proposed method. Xinggang Wang, Zhuotun Zhu, Cong Yao, Xiang Bai |
ICCV | 3 |
| 2015 | Automatic script identification in the wildabstractWith the rapid increase of transnational communication and cooperation, people frequently encounter multilingual scenarios in various situations. In this paper, we are concerned with a relatively new problem: script identification at word or line levels in natural scenes. A large-scale dataset with a great quantity of natural images and 10 types of widely-used languages is constructed and released. In allusion to the challenges in script identification in real-world scenarios, a deep learning based algorithm is proposed. The experiments on the proposed dataset demonstrate that our algorithm achieves superior performance, compared with conventional image classification or script identification methods, including as the original CNN architecture, LLC and GLCM. Baoguang Shi, Cong Yao, Chengquan Zhang, Feiyue Huang, Xiang Bai |
ICDAR | 2 |
| 2015 | Automatic discrimination of text and non-text natural imagesabstractWith the rapid growth of image and video data, there comes an interesting yet challenging problem: How to organize and utilize such large volume of data? Textual content in images and videos is an important source of information, which can be of great usefulness and assistance. Therefore, we investigate in this paper the problem of text image discrimination, which aims at distinguishing natural images with text from those without text. To tackle this problem, we propose a method that combines three mature techniques in this area, namely: MSER, CNN and BoW. To better evaluate the proposed algorithm, we also construct a large benchmark for text image discrimination, which includes natural images in a variety of scenarios. This algorithm has proven to be both effective and efficient, thus it can serve as a tool for mining valuable textual information from huge amount of image and video data. Chengquan Zhang, Cong Yao, Baoguang Shi, Xiang Bai |
ICDAR | 2 |
| 2014 | Multiple Stage Residual Model for Accurate Image Classification
Song Bai 0001, Xinggang Wang, Cong Yao, Xiang Bai |
ACCV (1) | 3 |
| 2014 | Strokelets: A Learned Multi-scale Representation for Scene Text RecognitionabstractDriven by the wide range of applications, scene text detection and recognition have become active research topics in computer vision. Though extensively studied, localizing and reading text in uncontrolled environments remain extremely challenging, due to various interference factors. In this paper, we propose a novel multi-scale representation for scene text recognition. This representation consists of a set of detectable primitives, termed as strokelets, which capture the essential substructures of characters at different granularities. Strokelets possess four distinctive advantages: (1) Usability: automatically learned from bounding box labels, (2) Robustness: insensitive to interference factors, (3) Generality: applicable to variant languages, and (4) Expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of strokelets and demonstrate the effectiveness of the proposed algorithm for text recognition. Cong Yao, Xiang Bai, Baoguang Shi, Wenyu Liu 0001 |
CVPR | 1 |
| 2014 | Human Detection Using Learned Part Alphabet and Pose Dictionary
Cong Yao, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki |
ECCV (5) | 1 |
| 2014 | A Unified Framework for Multioriented Text Detection and RecognitionabstractHigh level semantics embodied in scene texts are both rich and clear and thus can serve as important cues for a wide range of vision applications, for instance, image understanding, image indexing, video search, geolocation, and automatic navigation. In this paper, we present a unified framework for text detection and recognition in natural images. The contributions of this paper are threefold: 1) text detection and recognition are accomplished concurrently using exactly the same features and classification scheme; 2) in contrast to methods in the literature, which mainly focus on horizontal or near-horizontal texts, the proposed system is capable of localizing and reading texts of varying orientations; and 3) a new dictionary search method is proposed, to correct the recognition errors usually caused by confusions among similar yet different characters. As an additional contribution, a novel image database with texts of different scales, colors, fonts, and orientations in diverse real-world scenarios, is generated and released. Extensive experiments on standard benchmarks as well as the proposed database demonstrate that the proposed system achieves highly competitive performance, especially on multioriented texts. Cong Yao, Xiang Bai, Wenyu Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Traffic sign classification using two-layer image representationabstractThis paper makes use of locality-constrained linear coding (LLC) in a two-layer image representation framework for traffic sign recognition. As a multi-category classification problem with unbalanced frequencies and variations, many machine learning approaches have been adopted with some low level features for traffic sign recognition. To the best of our knowledge, this is the first method using coding features for traffic sign recognition. First, we extract features(dense SIFT features, HOG features and LBP features) and encode them with a k-means generated codebook and LLC. Second, each traffic sign image is represented by the features generated by spatial pyramid matching (SPM). Then, all the image representations from each kind of features are concatenated together as the final image representation. Finally, we show that a linear SVM classifier trained with this image representation can achieve the state-of-the-art recognition rate of 99.67% on the well-known German Traffic Sign Recognition Benchmark. Yingying Zhu 0005, Xinggang Wang, Cong Yao, Xiang Bai |
ICIP | 3 |
| 2012 | Detecting texts of arbitrary orientations in natural imagesabstractWith the increasing popularity of practical vision systems and smart phones, text detection in natural scenes becomes a critical yet challenging task. Most existing methods have focused on detecting horizontal or near-horizontal texts. In this paper, we propose a system which detects texts of arbitrary orientations in natural images. Our algorithm is equipped with a two-level classification scheme and two sets of features specially designed for capturing both the intrinsic characteristics of texts. To better evaluate our algorithm and compare it with other competing algorithms, we generate a new dataset, which includes various texts in diverse real-world scenarios; we also propose a protocol for performance evaluation. Experiments on benchmark datasets and the proposed dataset demonstrate that our algorithm compares favorably with the state-of-the-art algorithms when handling horizontal texts and achieves significantly enhanced performance on texts of arbitrary orientations in complex natural scenes. Cong Yao, Xiang Bai, Wenyu Liu 0001, Yi Ma 0001, Zhuowen Tu |
CVPR | 1 |
| 2012 | Online Random Ferns for robust visual tracking
Cong Rao, Cong Yao, Xiang Bai, Weichao Qiu, Wenyu Liu 0001 |
ICPR | 2 |
| 2012 | Co-Transduction for Shape RetrievalabstractIn this paper, we propose a new shape/object retrieval algorithm, namely, co-transduction. The performance of a retrieval system is critically decided by the accuracy of adopted similarity measures (distances or metrics). In shape/object retrieval, ideally, intraclass objects should have smaller distances than interclass objects. However, it is a difficult task to design an ideal metric to account for the large intraclass variation. Different types of measures may focus on different aspects of the objects: for example, measures computed based on contours and skeletons are often complementary to each other. Our goal is to develop an algorithm to fuse different similarity measures for robust shape retrieval through a semisupervised learning framework. We name our method co-transduction, which is inspired by the co-training algorithm. Given two similarity measures and a query shape, the algorithm iteratively retrieves the most similar shapes using one measure and assigns them to a pool for the other measure to do a re-ranking, and vice versa. Using co-transduction, we achieved an improved result of 97.72% (bull's-eye measure) on the MPEG-7 data set over the state-of-the-art performance. We also present an algorithm called tri-transduction to fuse multiple-input similarities, and it achieved 99.06% on the MPEG-7 data set. Our algorithm is general, and it can be directly applied on input similarity measures/metrics; it is not limited to object shape retrieval and can be applied to other tasks for ranking/retrieval. Xiang Bai, Bo Wang 0044, Cong Yao, Wenyu Liu 0001, Zhuowen Tu |
IEEE Trans. Image Process. | 3 |