Wei Feng 0016

dblp:17/1152-16 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0003-3711-6333ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2024 SignParser: An End-to-End Framework for Traffic Sign Understanding
Wei Feng 0016, Cheng-Lin Liu 0001
Int. J. Comput. Vis.2
2024 Video Text Detection With Robust Feature Representation
abstract
Existing video text detection methods mostly track texts with appearance feature only, thus are easily influenced by the change of perspective and illumination. In this paper, we propose an end-to-end video text detector that tracks texts based on robust feature representation fusing multiple descriptors. First, we introduce a character center segmentation branch to extract semantic feature, which encodes the category and position information of characters. And for extracting the topology feature of each text instance, we propose a relative position awareness branch to encode the relative position information among texts. Then, an adaptive feature fusion network is proposed to dynamically fuse multiple descriptors to generate a robust feature representation for more robust tracking. In addition, to promote the research and evaluation in this field, we also construct a large Bilingual Road scene Video Text dataset, named BiRViT-1K, which contains 1000 videos of Chinese and English texts. Experimental results show the proposed semantic and topology features are beneficial to the text detection and tracking performance, and the proposed method achieves state-of-the-art performance on four public video text benchmarks ICDAR 2015 Video, YVT, RT-1K and BOVText, and two Chinese scene text benchmarks CASIA10K and MSRA-TD500.
Wei Feng 0016, Mengbiao Zhao, Xu-Yao Zhang, Cheng-Lin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Social Relation Reasoning Based on Triangular Constraints
abstract
Social networks are essentially in a graph structure where persons act as nodes and the edges connecting nodes denote social relations. The prediction of social relations, therefore, relies on the context in graphs to model the higher-order constraints among relations, which has not been exploited sufficiently by previous works, however. In this paper, we formulate the paradigm of the higher-order constraints in social relations into triangular relational closed-loop structures, i.e., triangular constraints, and further introduce the triangular reasoning graph attention network (TRGAT). Our TRGAT employs the attention mechanism to aggregate features with triangular constraints in the graph, thereby exploiting the higher-order context to reason social relations iteratively. Besides, to acquire better feature representations of persons, we introduce node contrastive learning into relation reasoning. Experimental results show that our method outperforms existing approaches significantly, with higher accuracy and better consistency in generating social relation graphs.
Wei Feng 0016, Shuqi Mei, Cheng-Lin Liu 0001
AAAI3
2023 Texts as points: Scene text detection with point supervision
Mengbiao Zhao, Wei Feng 0016, Cheng-Lin Liu 0001
Pattern Recognit. Lett.2
2022 Mixed-Supervised Scene Text Detection With Expectation-Maximization Algorithm
abstract
Scene text detection is an important and challenging task in computer vision. For detecting arbitrarily-shaped texts, most existing methods require heavy data labeling efforts to produce polygon-level text region labels for supervised training. In order to reduce the cost in data labeling, we study mixed-supervised arbitrarily-shaped text detection by combining various weak supervision forms (e.g., image-level tags, coarse, loose and tight bounding boxes), which are far easier to annotate. Whereas the existing weakly-supervised learning methods (such as multiple instance learning) do not promote full object coverage, to approximate the performance of fully-supervised detection, we propose an Expectation-Maximization (EM) based mixed-supervised learning framework to train scene text detector using only a small amount of polygon-level annotated data combined with a large amount of weakly annotated data. The polygon-level labels are treated as latent variables and recovered from the weak labels by the EM algorithm. A new contour-based scene text detector is also proposed to facilitate the use of weak labels in our mixed-supervised learning framework. Extensive experiments on six scene text benchmarks show that (1) using only 10% strongly annotated data and 90% weakly annotated data, our method yields comparable performance to that of fully supervised methods, (2) with 100% strongly annotated data, our method achieves state-of-the-art performance on five scene text benchmarks (CTW1500, Total-Text, ICDAR-ArT, MSRA-TD500, and C-SVT), and competitive results on the ICDAR2015 Dataset. We will make our weakly annotated datasets publicly available.
Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001
IEEE Trans. Image Process.2
2021 Semantic-Aware Video Text Detection
abstract
Most existing video text detection methods track texts with appearance features, which are easily influenced by the change of perspective and illumination. Compared with appearance features, semantic features are more robust cues for matching text instances. In this paper, we propose an end-to-end trainable video text detector that tracks texts based on semantic features. First, we introduce a new character center segmentation branch to extract semantic features, which encode the category and position of characters. Then we propose a novel appearance-semantic-geometry descriptor to track text instances, in which se-mantic features can improve the robustness against appearance changes. To overcome the lack of character-level an-notations, we propose a novel weakly-supervised character center detection module, which only uses word-level annotated real images to generate character-level labels. The proposed method achieves state-of-the-art performance on three video text benchmarks ICDAR 2013 Video, Minetto and RT-1K, and two Chinese scene text benchmarks CA-SIA10K and MSRA-TD500.
Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001
CVPR1
2021 Learning to Understand Traffic Signs
abstract
One of the intelligent transportation system's critical tasks is to understand traffic signs and convey traffic information to humans. However, most related works are focused on the detection and recognition of traffic sign texts or symbols, which is not sufficient for understanding. Besides, there has been no public dataset for traffic sign understanding research. Our work takes the first step towards addressing this problem. First, we propose a "CASIA-Tencent Chinese Traffic Sign Understanding Dataset" (CTSU Dataset), which contains 5000 images of traffic signs with rich semantic descriptions. Second, we introduce a novel multi-task learning architecture that extracts text and symbol information from traffic signs, reasons the relationship between texts and symbols, classifies signs into different categories, and finally, composes the descriptions of the signs. Experiments show that the task of traffic sign understanding is achievable, and our architecture demonstrates state-of-the-art and superior performance. The CTSU Dataset is available at http://www.nlpr.ia.ac.cn/databases/CASIA-Tencent%20CTSU/index.html.
Wei Feng 0016, Shuqi Mei, Cheng-Lin Liu 0001
ACM Multimedia2
2021 Residual Dual Scale Scene Text Spotting by Fusing Bottom-Up and Top-Down Processing
Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001
Int. J. Comput. Vis.1
2020 Mutually Guided Dual-Task Network for Scene Text Detection
abstract
Scene text detection has been studied extensively. Existing methods detect either words or text lines and use either word-level or line-level annotated data for training. In this paper, we propose a dual-task network that can perform word-level and line-level text detection simultaneously and use training data of both levels of annotation to boost the performance. The dual-task network has two detection heads for word-level and line-level text detection, respectively. Then we propose a mutual guidance scheme for the joint training of the two tasks with two modules: line filtering module utilizes the output feature map of the text line detector to filter out the non-text regions for the word detector, and word enhancing module provides prior positions of words for the text line detector depending on the output feature map of the word detector. Experimental results of word-level and line-level text detection demonstrate the effectiveness of the proposed dual-task network and mutual guidance scheme, and the results of our method are competitive with state-of-the-art methods.
Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICPR2
2019 Turbo Learning Framework for Human-Object Interactions Recognition and Human Pose Estimation
abstract
Human-object interactions (HOI) recognition and pose estimation are two closely related tasks. Human pose is an essential cue for recognizing actions and localizing the interacted objects. Meanwhile, human action and their interacted objects’ localizations provide guidance for pose estimation. In this paper, we propose a turbo learning framework to perform HOI recognition and pose estimation simultaneously. First, two modules are designed to enforce message passing between the tasks, i.e. pose aware HOI recognition module and HOI guided pose estimation module. Then, these two modules form a closed loop to utilize the complementary information iteratively, which can be trained in an end-to-end manner. The proposed method achieves the state-of-the-art performance on two public benchmarks including Verbs in COCO (V-COCO) and HICO-DET datasets.
Wei Feng 0016, Wentao Liu 0002, Chen Qian 0006, Xiaolin Hu 0001
AAAI1
2019 TextDragon: An End-to-End Framework for Arbitrary Shaped Text Spotting
abstract
Most existing text spotting methods either focus on horizontal/oriented texts or perform arbitrary shaped text spotting with character-level annotations. In this paper, we propose a novel text spotting framework to detect and recognize text of arbitrary shapes in an end-to-end manner, using only word/line-level annotations for training. Motivated from the name of TextSnake, which is only a detection model, we call the proposed text spotting framework TextDragon. In TextDragon, a text detector is designed to describe the shape of text with a series of quadrangles, which can handle text of arbitrary shapes. To extract arbitrary text regions from feature maps, we propose a new differentiable operator named RoISlide, which is the key to connect arbitrary shaped text detection and recognition. Based on the extracted features through RoISlide, a CNN and CTC based text recognizer is introduced to make the framework free from labeling the location of characters. The proposed method achieves state-of-the-art performance on two curved text benchmarks CTW1500 and Total-Text, and competitive results on the ICDAR 2015 Dataset.
Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICCV1
2018 Scene Text Detection with Recurrent Instance Segmentation
abstract
Convolutional Neural Network (CNN) based scene text detection methods mostly employ the semantic segmentation (text/non-text classification) task to localize the regions of texts. However, they cannot distinguish different text-lines like instance segmentation. In this paper, we propose a novel framework based on Fully Convolutional Networks (FCN) and Recurrent Neural Network (RNN) to achieve both scene text detection and instance segmentation. The FCN is used to classify text and non-text regions, and the RNN utilizes the features extracted by FCN to simultaneously detect and segment one text instance at each time step. Meanwhile, it also extracts bounding boxes by a much simpler way than the non-maximum suppression (NMS) method. The proposed method achieves competitive results on two public benchmarks including ICDAR 2015 Incidental Scene Text Dataset and ICDAR 2013 Focused Scene Text Dataset. Moreover, the benefits of adding regression task in the RNN module are manifested.
Wei Feng 0016, Cheng-Lin Liu 0001
ICPR1