VLDB 2026 Research / reviewers in the wild / expert
Shuyang Feng
dblp:276/3624
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Image and video processing · 100% | |
| Artificial intelligence
1 paper |
Image recognition and object detection · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution |
1.4 | 2 | 2025 | Scene Text Image Super-Resolution Via Semantic Distillation and Text Perceptual Loss · IEEE Trans. Multim. 2025 Scene Text Image Super-Resolution via Parallelly Contextual Attention Network · ACM Multimedia 2021 |
Computer vision › Image recognition and object detection
scene text recognition |
0.9 | 1 | 2025 | Scene Text Image Super-Resolution Via Semantic Distillation and Text Perceptual Loss · IEEE Trans. Multim. 2025 |
Image and video processing
super-resolution |
0.9 | 1 | 2025 | Scene Text Image Super-Resolution Via Semantic Distillation and Text Perceptual Loss · IEEE Trans. Multim. 2025 |
Image and video processing
image restoration |
0.5 | 1 | 2021 | Scene Text Image Super-Resolution via Parallelly Contextual Attention Network · ACM Multimedia 2021 |
Image and video processing › super-resolution
image super-resolution |
0.5 | 1 | 2021 | Scene Text Image Super-Resolution via Parallelly Contextual Attention Network · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
adversarial training · 2.2semantic distillation · 1.7perceptual loss · 1.7contextual attention · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scene Text Image Super-Resolution Via Semantic Distillation and Text Perceptual LossabstractText Super-Resolution (SR) technology aims to recover lost information in low-resolution text images. With the proposal of TextZoom, which is the first dataset aiming at text super-resolution in real scenes, more and more scene text super-resolution models have been presented on the basis of it. Although these methods have achieved excellent performance, they do not consider how to make full and efficient use of semantic information. Out of this consideration, a Semantic-aware Trident Network (STNet) for Scene Text Image Super-Resolution is proposed. Specifically, pre-trained text recognition model ASTER (Attentional Scene Text Recognizer) is utilized to assist this process in two ways. Firstly, a novel basic block named Semantic-aware Trident Block (STB) is designed to build the STNet, which incorporates an added branch for semantic distillation to learn semantic information of pre-trained recognition model. Secondly, we expand our model in an adversarial training manner and propose new text perceptual loss based on ASTER to further enhance semantic information in SR images. Extensive experiments on TextZoom dataset show that compared with directly recognizing bicubic images, the proposed STNet boosts the recognition accuracy of ASTER, MORAN (Multi-Object Rectified Attention Network), and CRNN (Convolutional Recurrent Neural Network) by 17.4%, 18.2%, and 24.3%, respectively, which is higher than the performance of several existing state-of-the-art (SOTA) SR network models. Besides, experiments in real scenes (on ICDAR 2015 dataset) and in restricted scenarios (defense against adversarial attacks) validate that addition of semantic information enables the proposed method to achieve promising cross-dataset performance. Since the proposed method is trained on cropped images, when applied to real-world scenarios, locations of text in natural images are firstly localized through scene text detection methods, and then cropped text images are obtained based on detected text positions. Cairong Zhao, Shuyang Feng, Xuekuan Wang |
IEEE Trans. Multim. | 3 |
| 2023 | Text-Enhanced Scene Image Super-Resolution via Stroke Mask and Orthogonal AttentionabstractLow-resolution text images are very commonplace in real life and their information is hard to be extracted by using existing text recognition methods only. Although this problem can be solved by introducing super-resolution (SR) techniques, most existing SR methods fail to process stroke regions and background regions of input text images distinctively. In this paper, we propose a text-specific super-resolution network named Text Enhanced Attention Network (TEAN) to solve this problem. First of all, we compensate for disadvantages of traditional thresholding mask operation proposed in Text Super-Resolution Network (TSRN) by utilizing deep-learning based semantic segmentation method to get correct masks as prior semantic information and propose a Text-Segmented-Contextual-Attention (TSCA) branch on the basis of them. Besides, we design an Orthogonal Contextual Attention Module (OCAM) working with TSCA to implicitly enhance stroke regions of LR images. Secondly, to effectively fuse shallow features and deep features of SR model, we propose a convolutional structure named Weight Balanced Fusion Module (WBFM) to improve traditional feature fusion methods of SR network. Finally, extensive experiments on TextZoom dataset demonstrate that the proposed network can improve the recognition accuracy of text images on existing text recognition models. Using TEAN to process low-resolution text images improves the recognition accuracy by 25.4% on CRNN, by 17.4% on ASTER, by 17.3% on MORAN, by 20.7% on NRTR, by 17.3% on SAR and by 15.9% on MASTER compared with directly recognizing them, which attains competitive performances against state-of-the-art methods. Furthermore, cross-dataset experiments on IC15_2077 demonstrate that TEAN is helpful for scene text recognition task, especially for low-resolution images even with the cross-domain issue. Cairong Zhao, Shuyang Feng, Duoqian Miao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkabstractOptical degradation blurs text shapes and edges, so existing scene text recognition methods have difficulties in achieving desirable results on low-resolution (LR) scene text images acquired in real-world environments. The above problem can be solved by efficiently extracting sequential information to reconstruct super-resolution (SR) text images, which remains a challenging task. In this paper, we propose a Parallelly Contextual Attention Network (PCAN), which effectively learns sequence-dependent features and focuses more on high-frequency information of the reconstruction in text images. Firstly, we explore the importance of sequence-dependent features in horizontal and vertical directions parallelly for text SR, and then design a parallelly contextual attention block to adaptively select the key information in the text sequence that contributes to image super-resolution. Secondly, we propose a hierarchically orthogonal texture-aware attention module and an edge guidance loss function, which can help to reconstruct high-frequency information in text images. Finally, we conduct extensive experiments on TextZoom dataset, and the results can be easily incorporated into mainstream text recognition algorithms to further improve their performance in LR image recognition. Besides, our approach exhibits great robustness in defending against adversarial attacks on seven mainstream scene text recognition datasets, which means it can also improve the security of the text recognition pipeline. Compared with directly recognizing LR images, our method can respectively improve the recognition accuracy of ASTER, MORAN, and CRNN by 14.9%, 14.0%, and 20.1%. Our method outperforms eleven state-of-the-art (SOTA) SR methods in terms of boosting text recognition performance. Most importantly, it outperforms the current optimal text-orient SR method TSRN by 3.2%, 3.7%, and 6.0% on the recognition accuracy of ASTER, MORAN, and CRNN respectively. Cairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding, Jun Wu 0006, Fumin Shen, Heng Tao Shen |
ACM Multimedia | 2 |
| 2020 | Path Aggregation and Dual Supervision Network for Scene Text Detection
Shuyang Feng, Cairong Zhao |
PRCV (3) | 1 |