VLDB 2026 Research / reviewers in the wild / expert
Rui Zhang 0056
dblp:60/2536-56
· DBLP profile ↗
10ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0002-7386-2694ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | MUGS: A Multiple Granularity Semi-supervised Method for Text Recognition
Qianyi Jiang, Lingling Zhao, Rui Zhang 0056 |
ICDAR (5) | 5 |
| 2021 | Heterogeneous Network Based Semi-supervised Learning for Scene Text Recognition
Qianyi Jiang, Nan Li 0071, Rui Zhang 0056, Xiaolin Wei |
ICDAR (4) | 4 |
| 2021 | Scene Text Detection with Scribble Line
Yang Qiu 0002, Minghui Liao, Rui Zhang 0056, Xiaolin Wei, Xiang Bai |
ICDAR (4) | 4 |
| 2020 | ALEC: An Accurate, Light and Efficient Network for CAPTCHA Recognition
Nan Li 0071, Qianyi Jiang, Rui Zhang 0056, Xiaolin Wei |
DAS | 4 |
| 2020 | A Method for Scene Text Style Transfer
Gaojing Zhou, Yongsheng Zhou, Rui Zhang 0056, Xiaolin Wei |
DAS | 5 |
| 2020 | An Improved Convolutional Block Attention Module for Chinese Character Recognition
Yongsheng Zhou, Rui Zhang 0056, Xiaolin Wei |
DAS | 3 |
| 2020 | ReADS: A Rectified Attentional Double Supervised Network for Scene Text RecognitionabstractIn recent years, scene text recognition is always regarded as a sequence-to-sequence problem. Connectionist Temporal Classification (CTC) and Attentional sequence recognition (Attn) are two very prevailing approaches to tackle this problem while they may fail in some scenarios respectively. CTC concentrates more on every individual character but is weak in text semantic dependency modeling. Attn based methods have better context semantic modeling ability while tends to overfit on limited training data. In this paper, we elaborately design a Rectified Attentional Double Supervised Network (ReADS) for general scene text recognition. To overcome the weakness of CTC and Attn, both of them are applied in our method but with different modules in two supervised branches which can make a complementary to each other. Moreover, effective spatial and channel attention mechanisms are introduced to eliminate background noise and extract valid foreground information. Finally, a simple rectified network is implemented to rectify irregular text. The ReADS can be trained end-to-end and only word-level annotations are required. Extensive experiments on various benchmarks verify the effectiveness of ReADS which achieves state-of-the-art performance. Qianyi Jiang, Nan Li 0071, Rui Zhang 0056, Xiaolin Wei |
ICPR | 4 |
| 2020 | Robust Lexicon-Free Confidence Prediction for Text RecognitionabstractBenefiting from the success of deep learning, Optical Character Recognition (OCR) is booming in recent years. As we all know, the text recognition results are vulnerable to slight perturbation in input images, thus a method for measuring how reliable the results are is crucial. In this paper, we present a novel method for confidence measurement given a text recognition result, which can be embedded in any text recognizer with little overheads. Our method consists of two stages with a coarse-to-fine style. The first stage generates multiple candidates for voting coarse scores by a Single-Input Multi-Output network (SIMO). The second stage calculates a refined confidence score referred by the voting result and the conditional probabilities of the Top-1 probable recognition sequence. Highly competitive performance is achieved on several standard benchmarks which validate the efficiency and effectiveness of the proposed method. Moreover, it can be adopted in both Latin and non-Latin languages. Qianyi Jiang, Rui Zhang 0056, Xiaolin Wei |
ICPR | 3 |
| 2019 | Scene Text Detection with Feature Pyramid Network and Linking SegmentsabstractScene text detection is one of the most challenging problems in computer vision and has attracted great interest. Different from generic object detection, scene text detection mainly suffers from the large variance of scale, aspect ratio, and orientation in scene text. In this paper, we propose an effective and efficient model (SEG-FPN) for scene text detection, which is based on Feature Pyramid Network (FPN) and Linking Segments (SegLink). We incorporate feature pyramid mechanism with Single Shot Detector (SSD) framework to deal with different scale texts, and link locally detectable elements to detect texts of different orientations and aspect ratios. Moreover, compared with SSD, we enlarge the feature map of deep layers to better localize the large texts and recognize the small texts accurately. Experiments on ICDAR2015 and ICDAR2013 datasets demonstrate that our method can achieve comparable performance in terms of both accuracy and time. Specifically, SEG-FPN achieves an f-measure of 0.820 at 10.3 fps for 1280*768 ICDAR 2015 Incidental text images, and an f-measure of 0.879 at 19.2 fps for 512*512 ICDAR 2013 focused scene text images. Rui Zhang 0056, Yongsheng Zhou, Dong Wang 0004 |
ICDAR | 2 |
| 2019 | ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on SignboardabstractChinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinese characters can be arranged in various layouts with numerous fonts. The Chinese signboards in street view are a good choice for Chinese scene text images since they have different backgrounds, fonts and layouts. We organized a competition called ICDAR2019-ReCTS, which mainly focuses on reading Chinese text on signboard. This report presents the final results of the competition. A large-scale dataset of 25,000 annotated signboard images, in which all the text lines and characters are annotated with locations and transcriptions, were released. Four tasks, namely character recognition, text line recognition, text line detection and end-to-end recognition were set up. Besides, considering the Chinese text ambiguity issue, we proposed a multi ground truth (multi-GT) evaluation method to make evaluation fairer. The competition started on March 1, 2019 and ended on April 30, 2019. 262 submissions from 46 teams are received. Most of the participants come from universities, research institutes, and tech companies in China. There are also some participants from the United States, Australia, Singapore, and Korea. 21 teams submit results for Task 1, 23 teams submit results for Task 2, 24 teams submit results for Task 3, and 13 teams submit results for Task 4. The official website for the competition is http://rrc.cvc.uab.es/?ch=12. Rui Zhang 0056, Xiang Bai, Baoguang Shi, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar, Yongsheng Zhou, Qianyi Jiang, Nan Li 0071, Dong Wang 0004, Minghui Liao |
ICDAR | 1 |