Pengyuan Lv

dblp:168/4701 · also Pengyuan Lyu · DBLP profile ↗
← Back
33ranked-venue papers
16as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 11 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Recognition-Synergistic Scene Text Editing
abstract
Scene text editing aims to modify text content within scene images while maintaining style consistency. Traditional methods achieve this by explicitly disentangling style and content from the source image and then fusing the style with the target content, while ensuring content consistency using a pre-trained recognition model. Despite notable progress, these methods suffer from complex pipelines, leading to suboptimal performance in complex scenarios. In this work, we introduce Recognition-Synergistic Scene Text Editing (RS-STE), a novel approach that fully exploits the intrinsic synergy of text recognition for editing. Our model seamlessly integrates text recognition with text editing within a unified framework, and leverages the recognition model’s ability to implicitly disentangle style and content while ensuring content consistency. Specifically, our approach employs a multi-modal parallel decoder based on transformer architecture, which predicts both text content and stylized images in parallel. Additionally, our cyclic self-supervised fine-tuning strategy enables effective training on unpaired real-world data without ground truth, enhancing style and content consistency through a twice-cyclic generation process. Built on a relatively simple architecture, RS-STE achieves state-of-the-art performance on both synthetic and real-world benchmarks, and further demonstrates the effectiveness of leveraging the generated hard cases to boost the performance of downstream recognition tasks. Code is available at https://github.com/ZhengyaoFang/RS-STE.
Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei
CVPR2
2025 FreMamba: A Frequency-Domain Mamba Model for Hyperspectral Image Change Detection
abstract
The task of hyperspectral image change detection (HSI-CD) is to identify the subtle category changes of land surfaces by utilizing the rich spectral information of bi-temporal hyperspectral images (HSIs). Recent advanced deep learning methods have improved the performance of HSI-CD. However, HSIs have the mixed-pixel phenomenon, which affects the ability of HSI-CD models to discriminate land-cover changes. In this paper, a frequency domain Mamba (FreMamba) model is proposed to precisely capture the details of the spatial and spectral changes in bi-temporal HSIs, with consideration of the abovementioned problem. The proposed FreMamba model utilizes the capability of the Mamba model to adaptively filter the critical change information in long-range feature sequences. This is combined with frequency domain information to enhance the low-frequency global information and high-frequency spatial detail representation, based on a dual-branch network structure. A low-frequency spectral-guided attention (LSGA) module is proposed for the low-frequency Mamba branch, where it is embedded within the Mamba block via residual connections. Global low-frequency features are aggregated by the LSGA module with noise suppression, while the spectral change discriminability of ground objects is adaptively enhanced via the subsequent Mamba blocks. In the high-frequency Mamba branch, a high-frequency spatial geometric enhancement (HSGE) module is proposed that preserves edge details by spatial frequency domain decomposition to extract high-frequency geometric features. Experiments on three public HSI-CD datasets demonstrate that the proposed FreMamba model can accurately capture detailed change information and can outperform the existing state-of-the-art HSI-CD methods.
Pengyuan Lv, Feilong Shan, Yune Cao, Yanfei Zhong
IEEE Trans. Geosci. Remote. Sens.1
2025 QSCDNet: A Hybrid Quantum Spectral Change Detection Network for Hyperspectral Image Change Detection
abstract
Hyperspectral image change detection (HSI-CD) is an important remote sensing technique for identifying fine-grained land-cover change. Deep learning methods such as convolutional neural networks (CNNs) and transformers have achieved good performance in HSI-CD. However, due to the pseudo-changes caused by imaging conditions, the spectral change characteristics within each pixel often exhibit uncertainties. In this study, differing from the traditional deep learning methods, we aimed to relate the abovementioned spectral change uncertainty to the perspective of the quantum state and built a dual-branch hybrid quantum neural network for HSI-CD (QSCDNet). The quantum branch consists of several 2-D quantum spectral change convolutional blocks (QSCCBs). These blocks provide a wider variety of expression forms of the spectral change, independent of the pseudo-change impact, through a parameterized quantum circuit (PQC), to better extract the change information of the spectral features. The CNN branch is designed based on a channel attention network architecture to provide traditional network change features. The output features of the quantum network branch and the CNN branch are then fused based on a cross-domain change feature fusion module (CDFM). The impact of the number of QSCCBs was also analyzed to verify the effectiveness of increasing the depth of the quantum network structure. The proposed method was tested on three public HSI-CD datasets and compared with the state-of-the-art methods to validate its potential in the field of HSI-CD.
Pengyuan Lv, Ye Gao 0006, Heng Hu, Yanfei Zhong
IEEE Trans. Geosci. Remote. Sens.1
2025 SiamS²F: Satellite Video Single-Object Tracking Based on a Siamese Spectral-Spatial-Frame Correlation Network
abstract
In recent years, satellite video object tracking has received widespread attention as a research hotspot, but is also faced with many challenges. In satellite video, moving objects usually consist of only a few pixels, which makes it more difficult for the tracker to distinguish the target from the background. Furthermore, occlusion factors such as clouds, trees, and bridges can bring challenges when relocating the target from adjacent frames. In this article, to solve the above-mentioned problems, we propose a Siamese spectral-spatial-frame (SiamS2F) correlation network, which combines deep spectral, spatial, and frame features to enhance the interaction between video frames, to better focus on the target continuity. First, the deep features of the video are extracted with a Siamese backbone, along with a channel-spatial attention module. Second, a spectral-spatial-frame (SSF) module is proposed, where the output features of the backbone are fused based on the graph attention layer, and then the intraframe and interframe information of the fused features is enhanced by a newly designed multiframe interactive attention (MFIA) mechanism. Third, to solve the problem of similar objects, an additional center loss function is proposed in the classification regression head (CRH), where the search area is limited to a local range by adding an inbox identifier, to reduce the impact of the similar objects around the target. The proposed method was tested on two challenging benchmark remote sensing datasets—SatSOT and SV248S—where SiamS2F outperformed the related state-of-the-art trackers.
Pengyuan Lv, Xianyan Gao, Yanfei Zhong
IEEE Trans. Geosci. Remote. Sens.1
2024 WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-Only Supervised Text Spotting
Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Fanglin Chen 0001, Guangming Lu 0002, Wenjie Pei
ECCV (31)3
2024 Towards Unified Multi-granularity Text Detection with Interactive Attention
abstract
Existing OCR engines or document image analysis systems typically rely on training separate models for text detection in varying scenarios and granularities, leading to significant computational complexity and resource demands. In this paper, we introduce "Detect Any Text" (DAT), an advanced paradigm that seamlessly unifies scene text detection, layout analysis, and document page detection into a cohesive, end-to-end model. This design enables DAT to efficiently manage text instances at different granularities, including word, line, paragraph and page. A pivotal innovation in DAT is the across-granularity interactive attention module, which significantly enhances the representation learning of text instances at varying granularities by correlating structural information across different text queries. As a result, it enables the model to achieve mutually beneficial detection performances across multiple text granularities. Additionally, a prompt-based segmentation module refines detection outcomes for texts of arbitrary curvature and complex layouts, thereby improving DAT’s accuracy and expanding its real-world applicability. Experimental results demonstrate that DAT achieves state-of-the-art performances across a variety of text-related benchmarks, including multi-oriented/arbitrarily-shaped scene text detection, document layout analysis and page detection tasks.
Xingyu Wan, Chengquan Zhang, Pengyuan Lv, Sen Fan, Zihan Ni, Errui Ding, Jingdong Wang 0001
ICML3
2024 A Semi-Supervised Pyramid Cross-Temporal Attention Transformer for Change Detection in High-Resolution Remote Sensing Images
abstract
The vision transformer (ViT) model has the advantage of being able to model the long-range dependencies in the imagery and has been studied for the task of remote sensing image change detection (CD). However, the performance of the existing transformer-based CD methods is not satisfactory in the case of limited labeled data. The original self-attention mechanism cannot effectively extract the change information, and the large number of parameters in the ViT model makes the model difficult to train. To solve the above-mentioned problems, a semi-supervised pyramid cross-temporal attention transformer for change detection (CT2RCDSS) is proposed in this letter. The CT2RCDSS method follows an encoder-decoder structure. The encoder utilizes a dual-branch structure, containing the combination of the proposed cross-temporal attention (PCTA) and pyramid self-attention (PSA) mechanisms, which is designed to consider the interaction of the features from different time phases and enhance the changes at different scales. In the decoder, a series of deconvolutional layers with skip connections are utilized, and a Softmax layer follows to acquire the final binary change map. In addition, a semi-supervised training strategy, which reduces the errors in the pseudo-labels generated from the models initialized with different parameters, is used to improve the model stability while using unlabeled data. The experiments showed that the proposed method can achieve a superior F1-score and intersection over union (IoU), which indicates the potential of the proposed method.
Pengyuan Lv, Mengchen Li, Yanfei Zhong
IEEE Geosci. Remote. Sens. Lett.1
2024 QETR: A Query-Enhanced Transformer for Remote Sensing Image Object Detection
abstract
Recently, transformer models have been introduced into the field of remote sensing image object detection, benefiting from their ability to model long-term information. However, the existing transformer-based object detection methods mainly consider the global interaction of local elements and have a limited ability to enhance the local information, which can bring some difficulties in distinguishing real objects and a complex background. In this letter, a query-enhanced transformer (QETR) model is proposed to solve the above problems. The proposed model consists of three main parts: an encoder, a decoder, and a detection head. A Swin transformer is used to extract deep features in the encoder. In the decoder, the object and anchor queries are initialized and the feature and position information of the objects is learned by the multi-head self-attention and cross-attention mechanisms, respectively. Furthermore, a query align module along with a scale controller are proposed to enhance the object information around the local queries by limiting the attention to a certain range without losing important information. Finally, the boundaries and types of the objects are acquired from the detection head based on bipartite matching. To verify the effectiveness of the proposed method, comparative experiments were carried out with other state-of-the-art methodologies on two public datasets: the High-Resolution Remote Sensing Detection (HRRSD) dataset and the object detection in optical remote sensing images (DIOR) dataset. The experimental results confirm the effectiveness and superiority of the QETR model, which achieved 71.5% and 91.1% mAP values on the DIOR and HRRSD datasets, respectively.
Pengyuan Lv, Yanfei Zhong
IEEE Geosci. Remote. Sens. Lett.2
2024 Irregular text block recognition via decoupling visual, linguistic, and positional information
Chengquan Zhang, Jiaxin Zhang 0003, Zecheng Xie, Pengyuan Lv
Pattern Recognit.6
2024 A Semi-Supervised Semantic and Spatial Change Detail Retention Network for Semantic Change Detection in Remote Sensing Images
abstract
The task of semantic change detection (SCD) in remote sensing images (RSIs) is aimed at identifying the multiple “from-to” changes in land cover. However, the prevailing multitask SCD network structure does not directly model the semantic change information, leading to potential inaccuracies in extracting multiple detailed changes. In this article, a semi-supervised semantic and spatial change detail retention network (S3CDRNet) is proposed to precisely model the details of spatial changes and the distribution of multiple change categories for RSIs. The encoder of the proposed S3CDRNet is based on a Siamese structure, where a shallow convolutional neural network (CNN) followed by a transformer are used to extract the local-global change features. A precise semantic change perception module (PSCPM) based on large kernel convolution is then introduced to enhance the weak semantic changes. The decoder consists of several deconvolution layers to restore the original resolution, and a pixelwise semantic change map is acquired. To further distinguish the inherent imbalanced change types, an adaptive category-balanced semi-supervised learning (ACBSS) strategy is developed to better use the abundant class distribution information from the unlabeled image pairs. In the experiments conducted in this study, we made an in-depth study of multiple SCD based on three public datasets: the SECOND dataset, the Landsat-SCD dataset, and the Hi-UCD mini dataset. The results show the potential of the proposed method under scenarios with various change classes and an imbalanced change class problem, compared with the state-of-the-art SCD methods.
Pengyuan Lv, Yanfei Zhong
IEEE Trans. Geosci. Remote. Sens.1
2023 ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai
ICDAR (2)14
2023 GridFormer: Towards Accurate Table Structure Recognition via Grid Prediction
abstract
All tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex and edge of a grid. First, we propose a flexible table representation in the form of an M X N grid. In this representation, the vertexes and edges of the grid store the localization and adjacency information of the table. Then, we introduce a DETR-style table structure recognizer to efficiently predict this multi-objective information of the grid in a single shot. Specifically, given a set of learned row and column queries, the recognizer directly outputs the vertexes and edges information of the corresponding rows and columns. Extensive experiments on five challenging benchmarks which include wired, wireless, multi-merge-cell, oriented, and distorted tables demonstrate the competitive performance of our model over other methods.
Pengyuan Lv, Weihong Ma, Hongyi Wang 0008, Yuechen Yu, Chengquan Zhang, Yang Xue 0001, Jingdong Wang 0001
ACM Multimedia1
2023 Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation Learning
abstract
Due to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection. Despite great progress, these methods show deficiencies in robustness and still suffer from false positives and instance adhesion. Different from existing methods which integrate multiple-granularity features or multiple outputs, we resort to the perspective of representation learning in which auxiliary tasks are utilized to enable the encoder to jointly learn robust features with the main task of per-pixel classification during optimization. For semantic representation learning, we propose global-dense semantic contrast (GDSC), in which a vector is extracted for global semantic representation, then used to perform element-wise contrast with the dense grid features. To learn instance-aware representation, we propose to combine top-down modeling (TDM) with the bottom-up framework to provide implicit instance-level clues for the encoder. With the proposed GDSC and TDM, the encoder network learns stronger representation without introducing any parameters and computations during inference. Equipped with a very light decoder, the detector can achieve more robust real-time scene text detection. Experimental results on four public datasets show that the proposed method can outperform or be comparable to the state-of-the-art on both accuracy and speed. Specifically, the proposed method achieves 87.2% F-measure with 48.2 FPS on Total-Text and 89.6% F-measure with 36.9 FPS on MSRA-TD500 on a single GeForce RTX 2080 Ti GPU.
Xugong Qin, Pengyuan Lv, Chengquan Zhang, Yu Zhou 0015, Peng Zhang 0044, Hailun Lin, Weiping Wang 0005
ACM Multimedia2
2022 High Resolution Remote Sensing Image Semantic Segmentation Based on Ultra-Lightweight Fully Convolution Neural Network
abstract
In recent years, fully convolutional neural networks (FCNs) have been widely used in the field of remote sensing image semantic segmentation. However, these networks have huge amount of parameters and cost much computational efficiency. In this paper, an ultra-lightweight network (ULN) is proposed to overcome this problem. The proposed ULN model uses the encoder-decoder architecture to acquire the pixelwise result. In ULN, the efficient spatial pyramid network (ESPNet) is used to extract deep semantic features with fewer parameters. Considering the dilated convolutions will lose some semantic information in the encoding process, the feature enhancement block (FEB) is proposed. The recurrent criss-cross attention module is added at the end of skip connection to acquire the global contextual information. The proposed ULN is tested on the ISPRS Vaihingen dataset, the results show that our network achieves competitive results with fewer parameters(1.5M).
Pengyuan Lv, Yanfei Zhong, Liangpei Zhang 0001
IGARSS2
2022 Review of Vision Transformer Models for Remote Sensing Image Scene Classification
abstract
As an important semantic understanding method of remote sensing images, scene classification has received much attention in recent years. Convolutional neural network (CNN) is the representative deep learning method for scene classification which has powerful ability in feature extraction. However, the multilevel features in CNN are acquired by hierarchical convolutional layers which have difficulty in considering the interaction of different objects in the scene. Vision transformer (ViT) model provides a new way to understand the image by directly modeling the contextual information of local patches. This paper makes a review of recent progress of ViT models in the field of computer vision and remote sensing. The major contributions are as follows: 1) A brief review of the traditional scene classification methods is made; 2) ViT based models for scene classification are introduced and compared with CNN models; 3) Experiments of recent ViT models are performed and analyzed on UCM and NWPU datasets.
Pengyuan Lv, Yanfei Zhong, Liangpei Zhang 0001
IGARSS1
2022 Decoupling Recognition from Detection: Single Shot Self-Reliant Scene Text Spotter
abstract
Typical text spotters follow the two-stage spotting strategy: detect the precise boundary for a text instance first and then perform text recognition within the located text region. While such strategy has achieved substantial progress, there are two underlying limitations. 1) The performance of text recognition depends heavily on the precision of text detection, resulting in the potential error propagation from detection to recognition. 2) The RoI cropping which bridges the detection and recognition brings noise from background and leads to information loss when pooling or interpolating from feature maps. In this work we propose the single shot Self-Reliant Scene Text Spotter (SRSTS), which circumvents these limitations by decoupling recognition from detection. Specifically, we conduct text detection and recognition in parallel and bridge them by the shared positive anchor point. Consequently, our method is able to recognize the text instances correctly even though the precise text boundaries are challenging to detect. Additionally, our method reduces the annotation cost for text detection substantially. Extensive experiments on regular-shaped benchmark and arbitrary-shaped benchmark demonstrate that our SRSTS compares favorably to previous state-of-the-art spotters in terms of both accuracy and efficiency.
Pengyuan Lv, Guangming Lu 0002, Chengquan Zhang, Wenjie Pei
ACM Multimedia2
2022 SCViT: A Spatial-Channel Feature Preserving Vision Transformer for Remote Sensing Image Scene Classification
abstract
Convolutional neural network (CNN)-based methods are widely used in remote sensing image scene classification and can obtain excellent performances. However, the stacked receptive fields in the CNN-based methods have limitations in modeling the long-range dependencies of local features. The vision transformer (ViT) model provides a good solution as it directly considers the global interactions of local patches by the self-attention mechanism. However, the vanilla ViT model, which simply splits images into fixed-size patches treated as tokens, mainly considers the global information in the spatial domain. In this article, a spatial-channel feature preserving ViT (SCViT) model is proposed, which considers both the detailed geometric information of the high-spatial-resolution (HSR) imagery and the contribution of the different channels contained in the classification token. First, in the proposed method, tokens are generated by progressively aggregating the neighboring overlapping patches to extract the local structural features of the imagery. Second, a multihead self-attention (MSA) mechanism is used to model the global interactions of the tokens in the encoder. A lightweight channel attention (LCA) module is then introduced to consider the importance of the different channels in the classification token. Finally, a multilayer perceptron (MLP) is used to acquire the final results. Compared with the state-of-the-art scene classification methods, the experimental results confirm the potential of using ViT models in remote sensing image scene classification.
Pengyuan Lv, Yanfei Zhong, Fang Du, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Land-Use/Land-Cover Change Detection Based on Class-Prior Object-Oriented Conditional Random Field Framework for High Spatial Resolution Remote Sensing Imagery
abstract
High spatial resolution (HSR) remote sensing images can reflect more subtle changes and more specific types of land use and land cover (LULC) due to the abundant spatial geometric information. In this article, a class-prior object-oriented conditional random field (COCRF) framework consisting of a binary change detection (CD) task and a multiclass CD task is proposed to fill the application gap. In the proposed framework, the class-prior knowledge is used to improve the construction of the unary potential in both the binary and multiclass CD tasks, to reduce the influence of spectral variability. The binary CD result provides a constraint to the multiclass CD result. As a result, both parts have effective interaction. The class posterior probability images of two dates can be obtained automatically with the class-prior knowledge by sample migration. Furthermore, an object constraint described by the class dispersion within the objects is added to improve the smoothness in local objects, while the pairwise potential improves the smoothness of the whole area by using the eight-neighborhood spectral information of the center pixel. By integrating the above approaches, the problems of error accumulation and the manual intervention required in the traditional multiclass CD methods can be relieved. An adaptive parameter estimation strategy is also adopted in the proposed framework, to save the time required for manual parameter setting. The proposed COCRF framework was validated on two HSR remote sensing image data sets, where it achieved a better performance than the other state-of-the-art CD methods.
Sunan Shi, Yanfei Zhong, Ji Zhao 0006, Pengyuan Lv, Yinhe Liu, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering Network
abstract
The reading of arbitrarily-shaped text has received increasing research attention. However, existing text spotters are mostly built on two-stage frameworks or character-based methods, which suffer from either Non-Maximum Suppression (NMS), Region-of-Interest (RoI) operations, or character-level annotations. In this paper, to address the above problems, we propose a novel fully convolutional Point Gathering Network (PGNet) for reading arbitrarily-shaped text in real-time. The PGNet is a single-shot text spotter, where the pixel-level character classification map is learned with proposed PG-CTC loss avoiding the usage of character-level annotations. With a PG-CTC decoder, we gather high-level character classification vectors from two-dimensional space and decode them into text symbols without NMS and RoI operations involved, which guarantees high efficiency. Additionally, reasoning the relations between each character and its neighbors, a graph refinement module (GRM) is proposed to optimize the coarse recognition and improve the end-to-end performance. Experiments prove that the proposed method achieves competitive accuracy, meanwhile significantly improving the running speed. In particular, in Total-Text, it runs at 46.7 FPS, surpassing the previous spotters with a large margin.
Chengquan Zhang, Fei Qi 0001, Xiaoqiang Zhang 0006, Pengyuan Lv, Junyu Han, Jingtuo Liu, Errui Ding, Guangming Shi
AAAI6
2021 Open Set Face Anti-Spoofing in Unseen Attacks
abstract
In this paper, we propose an end-to-end open set face anti-spoofing (OSFA) approach for unseen attack recognition. Previous domain generalization approaches aim to align multiple domains beyond one common subspace, leading to performance degradation due to the discrepancy of different domains. To address this issue, our approach formulates face anti-spoofing (FAS) in an open set recognition framework, which learns compact representation for each known class in parallel to recognizing unseen attack examples. To this end, we introduce the statistical extreme value theory incorporated in our objective under the multi-task framework. Moreover, we develop an identity-aware contrastive learning method, preventing us from confusion in unseen attack examples versus hard examples. Experimental results on four datasets demonstrate the robustness of our proposed OSFA, especially under diverse categories of unseen attacks.
Hao Liu 0019, Pengyuan Lv, Zekuan Yu
ACM Multimedia4
2021 Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
abstract
Unifying text detection and text recognition in an end-to-end training fashion has become a new trend for reading text in the wild, as these two tasks are highly relevant and complementary. In this paper, we investigate the problem of scene text spotting, which aims at simultaneous text detection and recognition in natural images. An end-to-end trainable neural network named as Mask TextSpotter is presented. Different from the previous text spotters that follow the pipeline consisting of a proposal generation network and a sequence-to-sequence recognition network, Mask TextSpotter enjoys a simple and smooth end-to-end learning procedure, in which both detection and recognition can be achieved directly from two-dimensional space via semantic segmentation. Further, a spatial attention module is proposed to enhance the performance and universality. Benefiting from the proposed two-dimensional representation on both detection and recognition, it easily handles text instances of irregular shapes, for instance, curved text. We evaluate it on four English datasets and one multi-language dataset, achieving consistently superior performance over state-of-the-art methods in both detection and end-to-end text recognition tasks. Moreover, we further investigate the recognition module of our method separately, which significantly outperforms state-of-the-art methods on both regular and irregular text datasets for scene text recognition.
Minghui Liao, Pengyuan Lv, Minghang He, Cong Yao, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Scene Text Recognition from Two-Dimensional Perspective
abstract
Inspired by speech recognition, recent state-of-the-art algorithms mostly consider scene text recognition as a sequence prediction problem. Though achieving excellent performance, these methods usually neglect an important fact that text in images are actually distributed in two-dimensional space. It is a nature quite different from that of speech, which is essentially a one-dimensional signal. In principle, directly compressing features of text into a one-dimensional form may lose useful information and introduce extra noise. In this paper, we approach scene text recognition from a two-dimensional perspective. A simple yet effective model, called Character Attention Fully Convolutional Network (CA-FCN), is devised for recognizing the text of arbitrary shapes. Scene text recognition is realized with a semantic segmentation network, where an attention mechanism for characters is adopted. Combined with a word formation module, CA-FCN can simultaneously recognize the script and predict the position of each character. Experiments demonstrate that the proposed algorithm outperforms previous methods on both regular and irregular text datasets. Moreover, it is proven to be more robust to imprecise localizations in the text detection phase, which are very common in practice.
Minghui Liao, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lv, Cong Yao, Xiang Bai
AAAI6
2019 Learning Shape-Aware Embedding for Scene Text Detection
abstract
We address the problem of detecting scene text in arbitrary shapes, which is a challenging task due to the high variety and complexity of the scene. Specifically, we treat text detection as instance segmentation and propose a segmentation-based framework, which extracts each text instance as an independent connected component. To distinguish different text instances, our method maps pixels onto an embedding space where pixels belonging to the same text are encouraged to appear closer to each other and vise versa. In addition, we introduce a Shape-Aware Loss to make training adaptively accommodate various aspect ratios of text instances and the tiny gaps among them, and a new post-processing pipeline to yield precise bounding box predictions. Experimental results on three challenging datasets (ICDAR15, MSRA-TD500 and CTW1500) demonstrate the effectiveness of our work.
Zhuotao Tian, Michelle Shu, Pengyuan Lv, Ruiyu Li, Chao Zhou 0001, Xiaoyong Shen, Jiaya Jia
CVPR3
2019 ASTER: An Attentional Scene Text Recognizer with Flexible Rectification
abstract
A challenging aspect of scene text recognition is to handle text with distortions or irregular layout. In particular, perspective text and curved text are common in natural scenes and are difficult to recognize. In this work, we introduce ASTER, an end-to-end neural network model that comprises a rectification network and a recognition network. The rectification network adaptively transforms an input image into a new one, rectifying the text in it. It is powered by a flexible Thin-Plate Spline transformation which handles a variety of text irregularities and is trained without human annotations. The recognition network is an attentional sequence-to-sequence model that predicts a character sequence directly from the rectified image. The whole model is trained end to end, requiring only images and their groundtruth text. Through extensive experiments, we verify the effectiveness of the rectification and demonstrate the state-of-the-art recognition performance of ASTER. Furthermore, we demonstrate that ASTER is a powerful component in end-to-end recognition systems, for its ability to enhance the detector.
Baoguang Shi, Xinggang Wang, Pengyuan Lv, Cong Yao, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Multi-Oriented Scene Text Detection via Corner Localization and Region Segmentation
abstract
Previous deep learning based state-of-the-art scene text detection methods can be roughly classified into two categories. The first category treats scene text as a type of general objects and follows general object detection paradigm to localize scene text by regressing the text box locations, but troubled by the arbitrary-orientation and large aspect ratios of scene text. The second one segments text regions directly, but mostly needs complex post processing. In this paper, we present a method that combines the ideas of the two types of methods while avoiding their shortcomings. We propose to detect scene text by localizing corner points of text bounding boxes and segmenting text regions in relative positions. In inference stage, candidate boxes are generated by sampling and grouping corner points, which are further scored by segmentation maps and suppressed by NMS. Compared with previous methods, our method can handle long oriented text naturally and doesn't need complex post processing. The experiments on ICDAR2013, ICDAR2015, MSRA-TD500, MLT and COCO-Text demonstrate that the proposed algorithm achieves better or comparable results in both accuracy and efficiency. Based on VGG16, it achieves an F-measure of 84.3% on ICDAR2015 and 81.5% on MSRA-TD500.
Pengyuan Lv, Cong Yao, Shuicheng Yan, Xiang Bai
CVPR1
2018 Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
Pengyuan Lv, Minghui Liao, Cong Yao, Xiang Bai
ECCV (14)1
2018 Unsupervised Change Detection Based on Hybrid Conditional Random Field Model for High Spatial Resolution Remote Sensing Imagery
abstract
High spatial resolution (HSR) remote sensing images provide detailed geometric information about land cover. As a result, it is possible to detect more subtle changes with the help of HSR images. However, due to the increased spatial resolution and the limited spectral information, it is difficult to identify the real changes only through the spectral feature of the image. To fully explore the spectral–spatial information and improve the change detection performance for HSR images, this paper proposes the hybrid conditional random field (HCRF) model, which combines the traditional random field method with an object-based technique. In the proposed method, the spectral discriminative information of a single pixel is extracted by the unary potential, which is modeled using a soft clustering method to make an initial separation of changed and unchanged pixels. The pairwise potential then considers the contextual information of adjacent pixels to favor spatial smoothing. An object term is also introduced in the HCRF model to keep the homogeneity of changed objects. By the use of these approaches, the oversmoothing problem of the random field-based methods and the detection error caused by the segmentation strategy in the object-based methods can be relieved. The proposed method was tested on three HSR image data sets and outperformed the compared state-of-the-art techniques.
Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 Auto-Encoder Guided GAN for Chinese Calligraphy Synthesis
abstract
In this paper, we investigate the Chinese calligraphy synthesis problem: synthesizing Chinese calligraphy images with specified style from standard font(eg. Hei font) images (Fig. 1(a)). Recent works mostly follow the stroke extraction and assemble pipeline which is complex in the process and limited by the effect of stroke extraction. In this work we treat the calligraphy synthesis problem as an image-to-image translation problem and propose a deep neural network based model which can generate calligraphy images from standard font images directly. Besides, we also construct a large scale benchmark that contains various styles for Chinese calligraphy synthesis. We evaluate our method as well as some baseline methods on the proposed dataset, and the experimental results demonstrate the effectiveness of our proposed model.
Pengyuan Lv, Xiang Bai, Cong Yao, Zhen Zhu 0006, Tengteng Huang, Wenyu Liu 0001
ICDAR1
2017 Change detection based on structural conditional random field framework for high spatial resolution remote sensing imagery
abstract
In this paper, a structural conditional random field framework (SCRF) is proposed to detect the detailed change information from high spatial resolution (HSR) remote sensing imagery. Traditional random field based methods encounter the over-smoothing problem when deal with HSR images and the boundary of changed objects cannot be preserved well. To solve this problem, in SCRF, fuzzy c means (FCM) is used to model the unary potential while avoiding the independent assumption. Pairwise potentials with different shapes are selected as the structural set to model the spatial features of land cover such as buildings and roads. Based on SCRF, a set of change belief maps are generated to describe the observed image from different aspects. An object based fusion strategy is then followed to combine the belief maps to get the refined result. The results of the proposed method on two HSR data sets outperform some state-of-art algorithms.
Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Ailong Ma, Liangpei Zhang 0001
IGARSS1
2016 Robust Scene Text Recognition with Automatic Rectification
abstract
Recognizing text in natural images is a challenging task with many unsolved problems. Different from those in documents, words in natural images often possess irregular shapes, which are caused by perspective distortion, curved character placement, etc. We propose RARE (Robust text recognizer with Automatic REctification), a recognition model that is robust to irregular text. RARE is a speciallydesigned deep neural network, which consists of a Spatial Transformer Network (STN) and a Sequence Recognition Network (SRN). In testing, an image is firstly rectified via a predicted Thin-Plate-Spline (TPS) transformation, into a more "readable" image for the following SRN, which recognizes text through a sequence recognition approach. We show that the model is able to recognize several types of irregular text, including perspective text and curved text. RARE is end-to-end trainable, requiring only images and associated text labels, making it convenient to train and deploy the model in practical systems. State-of-the-art or highly-competitive performance achieved on several benchmarks well demonstrates the effectiveness of the proposed model.
Baoguang Shi, Xinggang Wang, Pengyuan Lv, Cong Yao, Xiang Bai
CVPR3
2016 Distinguishing text/non-text natural images with Multi-Dimensional Recurrent Neural Networks
abstract
In this paper, we focus on the text/non-text classification problem: distinguishing images that contain text from a lot of natural images. To this end, we propose a novel neural network architecture, termed Convolutional Multi-Dimensional Recurrent Neural Network (CMDRNN), which distinguishes text/non-text images by classifying local image blocks, taking both region pixels and dependencies among blocks into account. The network is composed of a Convolutional Neural Network (CNN) and a Multi-Dimensional Recurrent Neural Network (MDRNN). The CNN extracts rich and high-level image representation, while the MDRNN analyzes dependencies along multiple directions and produces block-level predictions. By evaluating CMDRNN on a public dataset, we observe improvements over prior arts in terms of both speed and accuracy.
Pengyuan Lv, Baoguang Shi, Chengquan Zhang, Xiang Bai
ICPR1
2016 Unsupervised change detection model based on hybrid conditional random field for high spatial resolution remote sensing imagery
abstract
In this paper, an unsupervised change detection model based on hybrid conditional random field model (HCRF) is proposed for high spatial resolution (HSR) remote sensing imagery. Traditional random field based algorithms are mainly based on the analysis of the difference image which ignores the spatial-temporal change information of ground objects which is important in dealing with HSR imagery. Thus in HCRF, a new graph structure is designed to explore the correlation of corresponding ground objects from different times to get a better result. The unary potential is selected as the probabilistic result of change vector analysis (CVA), the pairwise potential is modeled to consider the contextual information of difference image and the similarity between objects from bi-temporal original images is considered using an object term. The proposed method is tested on two HSR data sets (IKONOS and QuickBird) and out performs some state-of-art algorithms.
Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Liangpei Zhang 0001
IGARSS1
2016 Change Detection Based on a Multifeature Probabilistic Ensemble Conditional Random Field Model for High Spatial Resolution Remote Sensing Imagery
abstract
In this letter, a multifeature probabilistic ensemble conditional random field (MFPECRF) model is proposed to perform the task of change detection for high spatial resolution (HSR) remote sensing imagery. MFPECRF not only considers the spectral feature of single pixels but also the interaction between neighborhood pixels and the structural property of the ground objects in HSR imagery to give a higher detection accuracy than the traditional random field methods, which only utilize spectral and label information. In the unary potential, the spectral and morphological features of the difference image are combined using a probabilistic ensemble strategy, and the pairwise potential considers the contextual information of the observed field. The parameters of MFPECRF are estimated using a piecewise strategy, and the final result is obtained by the use of the loopy belief propagation algorithm. The experimental results of two groups of HSR multispectral images confirm the potential of the proposed method in improving the detection accuracy for HSR imagery.
Pengyuan Lv, Yanfei Zhong, Ji Zhao 0006, Hongzan Jiao, Liangpei Zhang 0001
IEEE Geosci. Remote. Sens. Lett.1