Rui Zhang 0056

dblp:60/2536-56 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0002-7386-2694ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2023 MUGS: A Multiple Granularity Semi-supervised Method for Text Recognition
Qianyi Jiang, Lingling Zhao, Rui Zhang 0056
ICDAR (5)5
2021 Heterogeneous Network Based Semi-supervised Learning for Scene Text Recognition
Qianyi Jiang, Nan Li 0071, Rui Zhang 0056, Xiaolin Wei
ICDAR (4)4
2021 Scene Text Detection with Scribble Line
Yang Qiu 0002, Minghui Liao, Rui Zhang 0056, Xiaolin Wei, Xiang Bai
ICDAR (4)4
2020 ALEC: An Accurate, Light and Efficient Network for CAPTCHA Recognition
Nan Li 0071, Qianyi Jiang, Rui Zhang 0056, Xiaolin Wei
DAS4
2020 A Method for Scene Text Style Transfer
Gaojing Zhou, Yongsheng Zhou, Rui Zhang 0056, Xiaolin Wei
DAS5
2020 An Improved Convolutional Block Attention Module for Chinese Character Recognition
Yongsheng Zhou, Rui Zhang 0056, Xiaolin Wei
DAS3
2020 ReADS: A Rectified Attentional Double Supervised Network for Scene Text Recognition
abstract
In recent years, scene text recognition is always regarded as a sequence-to-sequence problem. Connectionist Temporal Classification (CTC) and Attentional sequence recognition (Attn) are two very prevailing approaches to tackle this problem while they may fail in some scenarios respectively. CTC concentrates more on every individual character but is weak in text semantic dependency modeling. Attn based methods have better context semantic modeling ability while tends to overfit on limited training data. In this paper, we elaborately design a Rectified Attentional Double Supervised Network (ReADS) for general scene text recognition. To overcome the weakness of CTC and Attn, both of them are applied in our method but with different modules in two supervised branches which can make a complementary to each other. Moreover, effective spatial and channel attention mechanisms are introduced to eliminate background noise and extract valid foreground information. Finally, a simple rectified network is implemented to rectify irregular text. The ReADS can be trained end-to-end and only word-level annotations are required. Extensive experiments on various benchmarks verify the effectiveness of ReADS which achieves state-of-the-art performance.
Qianyi Jiang, Nan Li 0071, Rui Zhang 0056, Xiaolin Wei
ICPR4
2020 Robust Lexicon-Free Confidence Prediction for Text Recognition
abstract
Benefiting from the success of deep learning, Optical Character Recognition (OCR) is booming in recent years. As we all know, the text recognition results are vulnerable to slight perturbation in input images, thus a method for measuring how reliable the results are is crucial. In this paper, we present a novel method for confidence measurement given a text recognition result, which can be embedded in any text recognizer with little overheads. Our method consists of two stages with a coarse-to-fine style. The first stage generates multiple candidates for voting coarse scores by a Single-Input Multi-Output network (SIMO). The second stage calculates a refined confidence score referred by the voting result and the conditional probabilities of the Top-1 probable recognition sequence. Highly competitive performance is achieved on several standard benchmarks which validate the efficiency and effectiveness of the proposed method. Moreover, it can be adopted in both Latin and non-Latin languages.
Qianyi Jiang, Rui Zhang 0056, Xiaolin Wei
ICPR3
2019 Scene Text Detection with Feature Pyramid Network and Linking Segments
abstract
Scene text detection is one of the most challenging problems in computer vision and has attracted great interest. Different from generic object detection, scene text detection mainly suffers from the large variance of scale, aspect ratio, and orientation in scene text. In this paper, we propose an effective and efficient model (SEG-FPN) for scene text detection, which is based on Feature Pyramid Network (FPN) and Linking Segments (SegLink). We incorporate feature pyramid mechanism with Single Shot Detector (SSD) framework to deal with different scale texts, and link locally detectable elements to detect texts of different orientations and aspect ratios. Moreover, compared with SSD, we enlarge the feature map of deep layers to better localize the large texts and recognize the small texts accurately. Experiments on ICDAR2015 and ICDAR2013 datasets demonstrate that our method can achieve comparable performance in terms of both accuracy and time. Specifically, SEG-FPN achieves an f-measure of 0.820 at 10.3 fps for 1280*768 ICDAR 2015 Incidental text images, and an f-measure of 0.879 at 19.2 fps for 512*512 ICDAR 2013 focused scene text images.
Rui Zhang 0056, Yongsheng Zhou, Dong Wang 0004
ICDAR2
2019 ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
abstract
Chinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinese characters can be arranged in various layouts with numerous fonts. The Chinese signboards in street view are a good choice for Chinese scene text images since they have different backgrounds, fonts and layouts. We organized a competition called ICDAR2019-ReCTS, which mainly focuses on reading Chinese text on signboard. This report presents the final results of the competition. A large-scale dataset of 25,000 annotated signboard images, in which all the text lines and characters are annotated with locations and transcriptions, were released. Four tasks, namely character recognition, text line recognition, text line detection and end-to-end recognition were set up. Besides, considering the Chinese text ambiguity issue, we proposed a multi ground truth (multi-GT) evaluation method to make evaluation fairer. The competition started on March 1, 2019 and ended on April 30, 2019. 262 submissions from 46 teams are received. Most of the participants come from universities, research institutes, and tech companies in China. There are also some participants from the United States, Australia, Singapore, and Korea. 21 teams submit results for Task 1, 23 teams submit results for Task 2, 24 teams submit results for Task 3, and 13 teams submit results for Task 4. The official website for the competition is http://rrc.cvc.uab.es/?ch=12.
Rui Zhang 0056, Xiang Bai, Baoguang Shi, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar, Yongsheng Zhou, Qianyi Jiang, Nan Li 0071, Dong Wang 0004, Minghui Liao
ICDAR1