VLDB 2026 Research / reviewers in the wild / expert
Peirui Cheng
dblp:220/7625
· DBLP profile ↗
20ranked-venue papers
5as first author
17since 2021 · last 2025
0000-0002-4993-6753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldabstractExisting vision-based 3D occupancy prediction methods are inherently limited in accuracy due to their exclusive reliance on street-view imagery, neglecting the potential benefits of incorporating satellite views. We propose SA-Occ, the first Satellite-Assisted 3D occupancy prediction model, which leverages GPS & IMU to integrate historical yet readily available satellite imagery into real-time applications, effectively mitigating limitations of ego-vehicle perceptions, involving occlusions and degraded performance in distant regions. To address the core challenges of cross-view perception, we propose: 1) Dynamic-Decoupling Fusion, which resolves inconsistencies in dynamic regions caused by the temporal asynchrony between satellite and street views; 2) 3D-Proj Guidance, a module that enhances 3D feature extraction from inherently 2D satellite imagery; and 3) Uniform Sampling Alignment, which aligns the sampling density between street and satellite views. Evaluated on Occ3D-nuScenes, SA-Occ achieves state-of-the-art performance, especially among single-frame methods, with a 39.05% mIoU (a 6.97% improvement), while incurring only 6.93 ms of additional latency per frame. Our code and newly curated dataset are available at https://github.com/chenchen235/SA-Occ. Chen Chen 0036, Zhirui Wang 0003, Taowei Sheng, Yundu Li, Peirui Cheng, Luning Zhang, Kaiqiang Chen, Yanfeng Hu, Xue Yang 0005, Xian Sun 0001 |
ICCV | 6 |
| 2025 | Physics-Guided Deep Learning 3-D Inversion Based on Magnetic DataabstractThe 3-D inversion of magnetic data based on deep learning has achieved great success. This method relies on neural networks to extract features from a large amount of data and then generate structures, with high prediction accuracy and fast speed. However, this data-driven inversion method lacks a closed-form analytical expression and is likened to a “black box.” This feature makes it lack theoretical guidance and has poor interpretability. In addition, the training data often cannot fully cover all aspects of the target, so data-driven methods face the challenge of weak generalization ability in practical applications. Therefore, this letter proposes to integrate physical knowledge into the inversion method based on deep learning. The architecture and loss function of the deep learning model are designed based on the forward modeling of the magnetic field. This can not only improve the interpretability of the algorithm and provide a deeper understanding of the model but also help to make up for the lack of generalization capabilities of deep learning methods. Xiaoqing Shi, Zhirui Wang 0003, Xue Lu, Peirui Cheng, Luning Zhang, Liangjin Zhao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | UCDNet: Multi-UAV Collaborative 3-D Object Detection Network by Reliable Feature MappingabstractMulti-unmanned aerial vehicle (UAV) collaborative 3-D object detection can comprehend complex environments by integrating complementary information, with applications encompassing traffic monitoring, delivery services, and agricultural management. However, the extremely broad observations in aerial remote sensing and significant perspective differences across multiple UAVs make it challenging to achieve precise and consistent feature mapping from 2-D images to 3-D space in multi-UAV collaborative 3-D object detection paradigm. To address the problem, we propose an unparalleled camera-based multi-UAV collaborative 3-D object detection paradigm called UCDNet. Specifically, the depth information from the UAVs to the ground is explicitly utilized as a strong prior to provide a reference for more accurate and generalizable feature mapping. Additionally, we design a homologous point geometric consistency loss as an auxiliary self-supervision, which directly influences the feature mapping module, thereby strengthening the global consistency of multiview perception. Experiments on AeroCollab3D and CoPerception-UAVs datasets show that our method increases 4.7% and 10% mean Average Precision (mAP) respectively compared to the baseline, which demonstrates the superiority of UCDNet. Pengju Tian, Zhirui Wang 0003, Peirui Cheng, Zhechao Wang, Liangjin Zhao, Menglong Yan, Xue Yang 0005, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | AVCPNet: An AAV-Vehicle Collaborative Perception Network for 3-D Object DetectionabstractWith the advancement of collaborative perception, the role of autonomous aerial vehicle (AAV)–vehicle collaborative perception has become increasingly significant. The demand for collaborative perception from various perspectives to construct comprehensive perceptual information is rising. However, challenges emerge due to differences in the field of view (FOV) between cross-domain agents and their varying sensitivities to image information. Furthermore, accurate depth information is essential for collaboration to transform image features into bird’s eye view (BEV) features. To address these challenges, we propose a framework specifically designed for aerial-ground collaboration. First, to address the deficiency of datasets for aerial-ground collaboration, we have developed a virtual dataset named V2U-COO for our research. Second, we design a cross-domain cross-adaptation (CDCA) module to align the target information obtained from different domains, thereby achieving more accurate perception results. Finally, we introduce a collaborative depth optimization (CDO) module to obtain more precise depth estimation results, leading to more accurate perception results. We conduct extensive experiments on both our virtual dataset and a public dataset to validate the effectiveness of our framework. Our method resolves the feature fusion issue under significant height differences, a challenge that previous BEV generation methods struggled to address effectively. Our experiments on the V2U-COO and DAIR-V2X datasets demonstrate improvements in detection accuracy of 6.1% and 2.7%, respectively. Our code will be released athttps://github.com/wyccoo/uvcp. Zhirui Wang 0003, Peirui Cheng, Pengju Tian, Ziyang Yuan, Liangjin Zhao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | RingMo-Galaxy: A Remote Sensing Distributed Foundation Model for Diverse Downstream TasksabstractRemote sensing lightweight foundation models have successfully achieved online perception, providing real-time intelligent interpretation. However, their capabilities are restricted to inferences solely based on their respective observations and models, thus lacking a comprehensive understanding of large-scale remote sensing scenarios. To address this limitation, we propose RingMo-Galaxy, a remote sensing distributed foundation model based on generalized information mapping and interaction. RingMo-Galaxy can realize online collaborative perception across multiple platforms and diverse downstream tasks by mapping observations into a unified space and implementing a task-agnostic information interaction strategy. Specifically, we leverage the ground-based geometric prior of remote sensing oblique observations to change feature mapping from absolute to relative depth estimation, thereby enhancing the model’s ability to extract generalized features across diverse heights and perspectives. In addition, we present a dual-branch information compression module to decouple high-frequency and low-frequency features, achieving feature-level compression while preserving critical task-agnostic details. To support our research, we collect a multitask simulation dataset named AirCo-MultiTasks, specifically designed for multi-unmanned aerial vehicle (UAV) collaborative observation. We also conduct extensive experiments, including 3-D object detection, instance segmentation, and trajectory prediction. The numerous results demonstrate that our proposed RingMo-Galaxy achieves state-of-the-art performance across various downstream tasks. Zhechao Wang, Zhirui Wang 0003, Peirui Cheng, Liangjin Zhao, Pengju Tian, Mingxin Chen, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | TeCCo: A Terminal-Cloud Cross-Domain Collaborative Framework for Remote Sensing Image ClassificationabstractThe terminal-cloud collaborative framework boosts precision and efficiency by integrating cloud computing power with low-latency terminal responsiveness, offering a suitable solution for the growing demands of multiplatform remote sensing (RS) image interpretation. However, the significant differences in data distribution across various RS platforms present a great challenge in balancing the cloud’s centralized processing capabilities with the local interpretation abilities of different terminals. To address this challenge, we propose a terminal-cloud cross-domain collaborative (TeCCo) framework that inherits the efficiency advantages of multiple platforms while ensuring high-accuracy interpretation of diverse data distributions from different terminals. First, the dual classifier co-learning (DCCL) module is designed to enhance cloud robustness. By combining a multilayer perceptron for instance-level classification and a graph convolutional network (GCN) for feature-level aggregation, it achieves mutual supervision and improves feature alignment across different data distributions. Second, the hypernetwork personalization (HNP) module is introduced to generate personalized classifier parameters for each terminal with little fine-tuning cost, allowing terminals to maintain their uniqueness while benefiting from the generalization advantages of collaborative training. Finally, a data-assisted progressive inference mechanism is proposed to enhance accuracy by jointly clustering the features transmitted from terminals and the features of supervised data in the cloud. Extensive experiments demonstrate that TeCCo effectively addresses data distribution challenges, enhancing both the generalization of the cloud model and the personalization of terminal models, achieving state-of-the-art (SOTA) performance in cross-domain and multiplatform RS image classification. Peirui Cheng, Liangjin Zhao, Zhirui Wang 0003, Lingyu Kong, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Drones Help Drones: A Collaborative Framework for Multi-Drone Object Trajectory Prediction and BeyondabstractCollaborative trajectory prediction can comprehensively forecast the future motion of objects through multi-view complementary information. However, it encounters two main challenges in multi-drone collaboration settings. The expansive aerial observations make it difficult to generate precise Bird's Eye View (BEV) representations. Besides, excessive interactions can not meet real-time prediction requirements within the constrained drone-based communication bandwidth. To address these problems, we propose a novel framework named "Drones Help Drones" (DHD). Firstly, we incorporate the ground priors provided by the drone's inclined observation to estimate the distance between objects and drones, leading to more precise BEV generation. Secondly, we design a selective mechanism based on the local feature discrepancy to prioritize the critical information contributing to prediction tasks during inter-drone interactions. Additionally, we create the first dataset for multi-drone collaborative prediction, named "Air-Co-Pred", and conduct quantitative and qualitative experiments to validate the effectiveness of our DHD framework. The results demonstrate that compared to state-of-the-art approaches, DHD reduces position deviation in BEV representations by over 20\% and requires only a quarter of the transmission ratio for interactions while achieving comparable prediction performance. Moreover, DHD also shows promising generalization to the collaborative 3D object detection in CoPerception-UAVs. Zhechao Wang, Peirui Cheng, Minxing Chen, Pengju Tian, Zhirui Wang 0003, Xue Yang 0005, Xian Sun 0001 |
NeurIPS | 2 |
| 2024 | FS-DCL: Distributed Collaborative Learning for Few-Shot Remote Sensing Image ClassificationabstractWith the development of on-orbit hardware and distributed multiplatform observation systems in satellite remote sensing (RS) scenario, on-orbit collaborative model updating has become a promising trend. Due to restrictions of imaging conditions and storage resources, on-orbit updating is usually carried out with limited samples. However, existing collaborative learning methods rarely consider the few-shot problem. To address this issue, this letter innovatively proposes a distributed collaborative learning method for few-shot RS image classification (FS-DCL), which encourages the collaboration between satellites with similar data distribution to supplement useful information for each satellite, and design on-orbit models to extract more discriminative features. Specifically, a personalized parameter aggregation strategy (PPAS) is proposed to generate personalized parameters for each satellite based on information from satellites with similar data distributions, providing information gain to alleviate problems of insufficient samples. Besides, a feature enhancement method (FEM) is applied to on-orbit models to enhance the feature representation and produce a more discriminative feature space, thus improving the accuracy of few-shot metric classification. Extensive experiments on two RS datasets demonstrate the superiority of FS-DCL. Peirui Cheng, Yuelei Wang, Zhirui Wang 0003, Kaiqiang Chen, Xian Sun 0001, Daobing Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | MDCNet: A Multiplatform Distributed Collaborative Network for Object Detection in Remote Sensing ImageryabstractWith the recent development of remote sensing (RS) technology, the amount of RS platforms has witnessed a substantial increase, and the capacity of Earth observation has been greatly enhanced. The interpretation of RS images has also gradually evolved from traditional centralized ground processing to on-orbit processing. However, the traditional single-platform on-orbit processing is limited to a single source of information, which results in the underutilization of the advantages of multiplatform observation in the current RS field, and restricts the accuracy of inference tasks. To tackle the aforementioned problem, we propose a multiplatform distributed collaborative inference network, which can combine the intermediate features from multiple platforms to improve the accuracy of inference tasks. First, we proposed the collaboration map generator, which generates the collaboration map for optimal collaborator selection autonomously. Second, a spatial feature compression (SFC) module is designed to compress the interplatform transmission features, adapting spatially sparse distribution characteristics of RS objects. Finally, a feature fusion module containing spatial priors is proposed to fuse the features collected from multiple platforms to obtain more precise inference results. We conducted extensive experiments on three public datasets and verified the effectiveness of the proposed framework. On the NWPU VHR-10 dataset, for example, the proposed method improves the detection accuracy by 13.7% and 10.3% under two experimental settings compared with a single platform and compresses the intermediate data transmission between platforms by more than 80%. Shujing Duan, Peirui Cheng, Zhechao Wang, Zhirui Wang 0003, Kaiqiang Chen, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | FCIL-MSN: A Federated Class-Incremental Learning Method for Multisatellite NetworksabstractMulti-satellite networks have become the prevalent mode for remote sensing intelligent interpretation, with the onboard models requiring class-incremental updates to accommodate the new categories emerging in evolving data and tasks. Traditional model updating methods, which involve uploading models separately after ground-based updating, are inefficient due to limited uplink bandwidth and cumbersome ground update processes while underutilizing potential computing resources on satellites. To address the aforementioned problems, this paper innovatively proposes a collaborative in-orbit incremental update method termed FCIL-MSN, which leverages observational information and computing resources from multi-satellite networks. Firstly, FCIL-MSN achieves collaborative onboard model updates by introducing federated class-incremental learning into multi-satellite networks. Secondly, a bias calibration-guided relationship distillation module constructs a pseudo-feature set by collaborative multi-satellite networks, which alleviates the model bias caused by class imbalance from a global perspective, thereby enhancing model performance. Finally, a gradient information aggregation module is designed to facilitate the exclusion of unfavorable local updates by measuring the contribution of each terminal, thereby accelerating the convergence while obtaining the global model. We conduct extensive experiments on two datasets for scene classification tasks to verify the effectiveness of our proposed method. Experimental results demonstrate that FCIL-MSN outperforms existing general FCIL methods, improving average classification accuracy by 1.45% and decreasing the performance degradation rate by 6.40%. Ziqing Niu, Peirui Cheng, Zhirui Wang 0003, Liangjin Zhao, Xian Sun 0001, Zhi Guo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | RingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid FrameworkabstractIn recent years, remote sensing (RS) vision foundation models such as RingMo have emerged and achieved excellent performance in various downstream tasks. However, the high demand for computing resources limits the application of these models on edge devices. It is necessary to design a more lightweight foundation model to support on-orbit RS image interpretation. Existing methods face challenges in achieving lightweight solutions while retaining generalization in RS image interpretation. This is due to the complex high and low-frequency spectral components in RS images, which make traditional single CNN or Vision Transformer methods unsuitable for the task. Therefore, this paper proposes RingMo-lite, a RS lightweight network with a CNN-Transformer hybrid framework, which effectively exploits the frequency-domain properties of RS to optimize the interpretation process on several tasks like classification, object detection, semantic segmentation, and change detection. It is combined by the Transformer module as a low-pass filter to extract global features of RS images through a dual-branch structure, and the CNN module as a stacked high-pass filter to extract fine-grained details effectively. Furthermore, a novelty-designed frequency-domain masked image modeling (FD-MIM) is employed during the pretraining stage for self-supervised learning, which combines the high-frequency and low-frequency characteristics of each image patch. This approach effectively captures the latent feature representation in RS data. As shown in Fig. 1, compared with RingMo, the proposed RingMo-lite reduces the parameters over 60% in various RS image interpretation tasks, the average accuracy drops by less than 2% in most of the scenes and achieves SOTA performance compared to models of the similar size. In addition, our work will be integrated into the MindSpore computing platform in the near future. Yuelei Wang, Liangjin Zhao, Zhechao Wang, Ziqing Niu, Peirui Cheng, Kaiqiang Chen, Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | A Data-Related Patch Proposal for Semantic Segmentation of Aerial ImagesabstractLarge-size images cannot be directly put into GPU for training and need to be cropped to patches due to GPU memory limitation. The commonly used cropping methods before are random cropping and sequential cropping, which are crude and fatally inefficient. Firstly, categories of datasets are often imbalanced, and just simple cropping misses an excellent opportunity to make the data distribution balanced. Secondly, the training needs to crop a large number of patches to cover all patterns, which greatly increases the training time. This problem is of great practical hazards but is often overlooked by previous works. The optimal solution is to generate valuable patches. Valuable patches refer to the value to network training, i.e., the value of this patch for the convergence of the network, and the improvement of the accuracy. To this end, we propose a data-related patch proposal strategy to sample high valuable patches. The core idea is to score each patch according to the accuracy of each category, so as to perform balanced sampling. Compared with random cropping or sequential cropping, our method can improve the segmentation accuracy and accelerate the training vastly. Moreover, our method also shows great advantages over the loss-based balanced approaches. Experiments on Deepglobe and Potsdam show the excellent effect of our method. Lianlei Shan, Guiqin Zhao, Jun Xie 0003, Peirui Cheng, Xiaobin Li 0006, Zhepeng Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Detect Arbitrary-Shaped Text via Adaptive Thresholding and Localization Quality EstimationabstractThe two-stage scene text detection algorithms based on Mask R-CNN have achieved good performances on multiple challenging benchmarks. However, their effectiveness is degraded due to artificially setting constant thresholds and low localization quality of candidate boxes. In this paper, we present a novel scene text detection method based on Mask R-CNN and the proposed method, named LOAD, proposes adaptive threshold module and localization quality estimation module to address the above two problems. We propose two kinds of adaptive thresholds which are used for the filtering of candidate boxes and the binarization of pixels respectively. We introduce the self-attention mechanism to obtain the global information for generating the adaptive thresholds. Besides, we introduce the localization quality estimation into our model to obtain more accurate candidate boxes for subsequent segmentation. Comparative experiments are conducted on five benchmarks(ICDAR 2015, ICDAR 2017, MSRA-TD500, Total-Text and CTW1500), and the results demonstrate that the proposed method achieves the state-of-the-art performance with an F-measure of 91.0%, 78.7%, 87.4%, 90.6% and 86.0%. We also provide adequate ablation experiments to demonstrate the effectiveness of the proposed components. Peirui Cheng, Yuzhong Zhao, Weiqiang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | DCM: A Distributed Collaborative Training Method for the Remote Sensing Image ClassificationabstractAs the number of aero and space remote sensing platforms increases, distributed observation and real-time terminal processing become mainstream in the future. However, most of the training methods for the multi-platform are still limited to centralized structures or independent training based on a single platform, which is inefficient or limited in accuracy. In order to solve this problem, we innovatively propose a distributed collaborative method (DCM) for remote sensing image classification training in this article. First, the proposed training method, which is based on one cloud and several terminals, can aggregate different parameters of the terminal network to the cloud to improve global accuracy. Second, a sample proximity network is designed to process the problem of data heterogeneity on different terminal networks, which further improves the accuracy during the model fusion on the cloud. Third, a multi-layer grouped concatenation module is applied after the model fusion to extract hierarchical features with different categories of remote sensing images. Experimental results on the challenging remote sensing image classification dataset FAIR1M show that the proposed training method has better collaborative learning ability than the centralized-based model or terminal-trained lightweight network under the heterogeneous data. Yuelei Wang, Zhirui Wang 0003, Peirui Cheng, Xuan Zeng 0004, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | DCNNet: A Distributed Convolutional Neural Network for Remote Sensing Image ClassificationabstractWith the development of information technology, multiplatform collaborative collection and processing of remote sensing (RS) images has become a significant trend. However, the existing models are challenging to achieve accurate and efficient image interpretation on RS multiplatform systems. To solve this problem, we propose a novel distributed convolutional neural network (DCNNet) and demonstrate the superiority of our method in RS image classification. First, a progressive inference mechanism is introduced to support most images to be classified in advance with satisfactory accuracy, which minimizes redundant cloud transmission and achieves higher inference acceleration. Meanwhile, a distributed self-distillation paradigm is designed to integrate and refine in-depth features, performing efficient knowledge transfer between the terminals and the cloud network. Second, a multiscale feature fusion (MSFF) module is presented to extract valid receptive fields and assign weights to crucial channel dimension features. Finally, a sampling augmentation (SA) attention is proposed to enhance the effective feature representation of RS images through a bottom-up and top-down feedforward structure. We conducted extensive experiments and visual analyses on three benchmark scene classification datasets and one fine-grained dataset. Compared with the existing methods, DCNNet consolidates several advantages in terms of accuracy, computation, transmission, and processing efficiency into a single framework for multiplatform RS image classification. Zhirui Wang 0003, Peirui Cheng, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Direct regression scene text detection with accuracy scoring
Peirui Cheng, Yuzhong Zhao, Yuanqiang Cai, Weiqiang Wang 0001 |
Neurocomputing | 1 |
| 2021 | Scale-Residual Learning Network for Scene Text DetectionabstractDetecting incidentally captured text in the wild remains an open problem due to challenging factors including unconstrained scenarios and large scale variation. In this paper, we establish a large-scale scene text detection dataset (LS-Text), containing 36, 000 images and 270, 783 text instances with various scales and complex scenarios, to promote the research of text detection. We propose a Scale-residual Learning Network (SLN) to deal with the scale variation problem in a progressive optimization manner. Specifically, we integrate both learnable feature concatenation and feature up-sampling operator. It can effectively eliminate the residuals between the outputs of SLN and ground-truth text instances by processing both the Feature Fusion Residuals (FFR) and the Scale Transformation Residuals (STR), simultaneously. By stacking multi-scale feature maps in a deep-to-shallow manner, SLN continuously optimizes feature representation by accumulating strong semantic information and rich texture details in a scale-residual learning way. Extensive experimental results on five challenging datasets demonstrate the state-of-the-art performance of the proposed SLN model, and the challenging aspects related to real-world scenarios of the proposed LS-Text dataset. Both the source code of SLN and the LS-Text dataset are available athttps://github.com/SLN-Text-Detection. Yuanqiang Cai, Chang Liu 0047, Peirui Cheng, Dawei Du, Libo Zhang 0001, Weiqiang Wang 0001, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | A Direct Regression Scene Text Detector With Position-Sensitive SegmentationabstractDirect regression methods have demonstrated their success on various multi-oriented benchmarks for scene text detection due to the high recall rate for small targets and the direct regression for text boxes. However, too many false positive candidates and inaccurate position regression still limit the performance of these methods. In this paper, we propose an end-to-end method by introducing position-sensitive segmentation into the direct regression method to overcome these shortcomings. We generate the ground truth of position-sensitive segmentation maps based on the information of text boxes so that the position-sensitive segmentation module can be trained synchronously with the direct regression module. Besides, more information about the relative position of text is provided for the network through the training of position-sensitive segmentation maps, which improves the expressiveness of the network. We also introduce spatial pyramid of position-sensitive segmentation into the proposed method considering the huge differences in sizes and aspect ratios of scene texts and we propose position-sensitive COI(Corner area of Interest) pooling into the proposed method to speed up the inference. Experiments on datasets ICDAR2015, MLT-17 and COCO-Text demonstrate that the proposed method has a comparable performance with state-of-the-art methods while it is more efficient. We also provide abundant ablation experiments to demonstrate the effectiveness of these improvements in our proposed method. Peirui Cheng, Yuanqiang Cai, Weiqiang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Multi-scale Scene Text Detection via Resolution TransformabstractScene text detection is a challenging task because there are many small text targets in the natural scene and the size of scene text varies greatly. For the current popular scene text detection methods, such as EAST and Textboxes, it is difficult to detect small text and text with large differences in size well at the same time. To solve the problem, we propose a novel multi-scale scene text detection method based on EAST. The proposed method extracts high-resolution feature maps at multiple scales via resolution transform and detects text on these feature maps. Through the resolution transform module, the proposed method can detect both kinds of text well. Besides, we use aggregated feature pyramid module to efficiently pass both low-level and high-level information to feature maps at each scale. Experiments on datasets ICDAR2015 and COCO-Text demonstrate that the proposed method has a comparable performance with state-of-the-art methods and it is more efficient. For ICDAR2015 and COCO-Text datasets, the proposed method achieves an F-score of 0.84 and 0.42 respectively. Peirui Cheng, Weiqiang Wang 0001, Yuanqiang Cai |
ICME | 1 |
| 2018 | A Multi-Oriented Scene Text Detector with Position-Sensitive SegmentationabstractScene text detection has been studied for a long time and lots of approaches have achieved promising performances. Most approaches regard text as a specific object and utilize the popular frameworks of object detection to detect scene text. However, scene text is different from general objects in terms of orientations, sizes and aspect ratios. In this paper, we present an end-to-end multi-oriented scene text detection approach, which combines the object detection framework with the position-sensitive segmentation. For a given image, features are extracted through a fully convolutional network. Then they are input into text detection branch and position-sensitive segmentation branch simultaneously, where text detection branch is used for generating candidates and position-sensitive segmentation branch is used for generating segmentation maps. Finally the candidates generated by text detection branch are projected onto the position-sensitive segmentation maps for filtering. The proposed approach utilizes the merits of position-sensitive segmentation to improve the expressiveness of the proposed network. Additionally, the approach uses position-sensitive segmentation maps to further filter the candidates so as to highly improve the precision rate. Experiments on datasets ICDAR2015 and COCO-Text demonstrate that the proposed method outperforms previous state-of-the-art methods. For ICDAR2015 dataset, the proposed method achieves an F-score of 0.83 and a precision rate of 0.87. Peirui Cheng, Weiqiang Wang 0001 |
ICMR | 1 |