Feng Chen 0040

dblp:21/3047-40 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
25since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
YearPublicationVenuePosition
2025 SPAN: A Salient Patch-Clue Aware Network for Cross-Domain Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) plays a critical role in ensuring the security of face recognition system from different kinds of presentation attacks. Most existing FAS research faces several limitations: 1) insufficient consideration of the role of local fine-grained information, 2) the assumption that spoofing patterns are uniformly distributed across the entire image, neglecting the uneven distribution of spoofing clues, and 3) an overemphasis on intra-domain scenarios, leading to limited generalization capabilities for unseen domains. In this paper, we propose a Salient Patch-Clue Aware Network (SPAN) for cross-domain face anti-spoofing to tackle the aforementioned issues. Specifically, we use all patches cropped from the complete image as input to the FAS network, enabling the network to focus on local information while avoiding information loss. Additionally, we propose a patch perception mechanism to extract key regions containing salient spoofing clues, such as reflections and edges, thereby reducing interference from irrelevant information. Furthermore, we introduce a pixel perception mechanism to capture finer-grained details. Based on these two mechanisms, we design a Salient Clue Perception Module (SCPM). We conduct cross-domain experiments on CASIA-FASD, Idiap Replay-Attack, MSU-MFSD, and OULU-NPU datasets. Our method achieves state-of-the-art HTER on seven protocols, especially excelling on M&I to C and M&I to O, surpassing the second place by 9.79% and 6.19%, showcasing strong generalization capability. The codes are available at https://github.com/SPAN2025/SPAN.
Liangfeng Zhang, Lei Chen 0069, Jinhui Lin, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
IJCNN7
2024 Multi-task Learning for License Plate Recognition in Unconstrained Scenarios
Zhen-Lun Mo, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICDAR (1)4
2024 HQOD: Harmonious Quantization for Object Detection
abstract
Task inharmony problem commonly occurs in modern object detectors, leading to inconsistent qualities between classification and regression tasks. The predicted boxes with high classification scores but poor localization positions or low classification scores but accurate localization positions will worsen the performance of detectors after Non-Maximum Suppression. Furthermore, when object detectors collaborate with Quantization- Aware Training (QAT), we observe that the task inharmony problem will be further exacerbated, which is considered one of the main causes of the performance degradation of quantized detectors. To tackle this issue, we propose the Harmonious Quantization for Object Detection (HQOD) framework, which consists of two components. Firstly, we propose a task-correlated loss to encourage detectors to focus on improving samples with lower task harmony quality during QAT. Secondly, a harmonious Intersection over Union (IoU) loss is incorporated to balance the optimization of the regression branch across different IoU levels. The proposed HQOD can be easily integrated into different QAT algorithms and detectors. Remarkably, on the MS COCO dataset, our 4-bit ATSS with ResNet-50 backbone achieves a state-of-the- art mAP of 39.6%, even surpassing the full-precision one. Codes are available at https://github.com/Menace-Dragon/VP-QOD.
Zhiwei Dong, Song-Lu Chen, Ruiyao Zhang, Shutong Ti, Feng Chen 0040, Xu-Cheng Yin
ICME6
2024 Towards Low-resource License Plate Recognition via Feature Shuffling
abstract
Manual annotation is costly and limits the availability of sufficient annotated license plates for training recognition models. Small-scale license plate datasets (i.e., low-resource) often exhibit a long-tailed distribution in character classes at some character positions, primarily due to their limited variation in character permutations. Previous methods tend to prioritize head classes with high occurrence probability when applied to small-scale datasets. To solve this problem, we propose feature shuffling to balance the occurrence distribution across various character classes, thereby improving the recognition of tail classes with low occurrence probability. Moreover, we introduce global perception to holistically understand the overall character layout for effective feature shuffling. Extensive experiments on the small-scale UFPR and SSIG-SegPlate datasets demonstrate that our method achieves state-of-the-art results, with an average improvement of 43.70% over the baseline. Experiments on RodoSol and CCPD prove our method achieves state-of-the-art performance on large-scale datasets, verifying its generality.
Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICME4
2024 Transmitted and Aggregated Self-Attention for Automatic Speech Recognition
Tian-Hao Zhang, Xinyuan Qian 0001, Feng Chen 0040, Xu-Cheng Yin
INTERSPEECH3
2024 Improving Small License Plate Detection with Bidirectional Vehicle-Plate Relation
Songkang Dai, Song-Lu Chen, Qi Liu 0041, Chao Zhu 0003, Feng Chen 0040, Xu-Cheng Yin
MMM (2)6
2024 Irregular License Plate Recognition via Global Information Integration
Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
MMM (2)4
2024 Integrated Recognition of Arbitrary-Oriented Multi-line Billet Number
Zhongjie Hu, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
PRCV (7)5
2024 Improving Multi-Type License Plate Recognition via Learning Globally and Contrastively
abstract
Previous license plate recognition (LPR) methods have achieved impressive performance on single-type license plates. However, multi-type license plate recognition is still challenging due to various character layouts and fonts. There are two main problems: one is that recognition models are prone to incorrectly perceive the location of characters due to diverse character layouts, and the other is that characters of different categories may have similar glyphs due to various fonts, causing character misidentification. Therefore, to solve the above problems, we propose two plug-and-play modules based on an attention-based framework for multi-type license plate recognition. First, we propose a global modeling module to integrate character layout information to precisely perceive the location of characters, thus generating accurate predictions. Second, a position-aware contrastive learning module is proposed to enhance the robustness and discriminability of features to alleviate character misidentification of similar glyphs. Finally, to verify the effectiveness and generality, we apply the proposed modules to six baseline models, and the results demonstrate that the proposed method can achieve state-of-the-art performance on three multi-type license plate datasets. Moreover, extensive experiments prove that our proposed modules can significantly improve performance by 6.8% on RODOSOL-ALPR with a small parameter increase.
Qi Liu 0041, Song-Lu Chen, Tian-Hao Zhang, Feng Chen 0040, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.5
2023 Self-Convolution for Automatic Speech Recognition
abstract
Self-attention plays a significant role in recent automatic speech recognition (ASR) models with promising results. However, it suffers from high computational complexity and weak capability in modeling local information. In contrast, the convolutional neural network (CNN) is computationally effective and superior in learning local information. Whereas it fails in self-interaction and capturing long-range dependence among input tokens. Accordingly, we take their complementary advantages and propose a new module, namely self-convolution, to compensate for each individual limitations. Specifically, self-convolution generates convolution kernels at each token (to model local information) which are then used to convolve itself (for self-interaction). Moreover, we bring in global information during the generation of convolution kernel to enhance the learning of long-range dependencies. In this way, the advantages of self-attention and CNN are both utilized. We conduct rigorous experiments on LibriSpeech, Tedlium2, and AIShell1 datasets and demonstrate that our proposed self-convolution can achieve superior ASR performance than self-attention with less computational cost.
Qi Liu 0041, Xinyuan Qian 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICASSP5
2023 End-to-End Multi-line License Plate Recognition with Cascaded Perception
Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
ICDAR (5)3
2023 Complex Glyph Enhancement for License Plate Generation
Yu-Xiang Chen, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICIG (1)6
2023 InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu 0041, Xinyuan Qian 0001, Li-Fang Wei, Feng Chen 0040, Song-Lu Chen, Xu-Cheng Yin
INTERSPEECH6
2023 Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding
Tian-Hao Zhang, Haibo Qin, Zhi-Hao Lai, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xinyuan Qian 0001, Xu-Cheng Yin
INTERSPEECH6
2023 LiteHandNet: A Lightweight Hand Pose Estimation Network via Structural Feature Enhancement
Zhi-Yong Huang, Song-Lu Chen, Qi Liu 0041, Chong-Jian Zhang, Feng Chen 0040, Xu-Cheng Yin
MMM (1)5
2023 Feature Enhancement and Reconstruction for Small Object Detection
Chong-Jian Zhang, Song-Lu Chen, Qi Liu 0041, Zhi-Yong Huang, Feng Chen 0040, Xu-Cheng Yin
MMM (1)5
2023 Hypersphere guided embedding for masked face recognition
Xiaobin Zhu 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin, Lei Chen 0069
Pattern Recognit. Lett.4
2022 Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition
abstract
Non-autoregressive (NAR) transformer models have been studied intensively in automatic speech recognition (ASR), and many NAR transformer models is to use the causal mask to limit token dependencies. However, the causal mask is designed for the left-to-right decoding process of the non-parallel autoregressive (AR) transformer, which is inappropriate for the parallel NAR transformer since it ignores the right-to-left contexts. Some methods are proposed to utilize right-to-left contexts with an extra decoder, but these methods increase the model complexity. To tackle the above problems, we propose a new non-autoregressive transformer with a unified bidirectional decoder (NAT-UBD), which can simultaneously utilize left-to-right and right-to-left contexts for ASR. However, direct use of bidirectional contexts will cause information leakage, which means the decoder output can be affected by the character information of the input in the same position. To avoid information leakage, we propose a novel attention mask and modify vanilla queries, keys, and values matrices for NAT-UBD. Experimental results verify that NAT-UBD can achieve character error rates (CERs) of 5.0%/5.5% on the Aishell-1 dev/test sets, outperforming all previous NAR transformer models. Moreover, NAT-UBD can run 49.8× faster than the AR transformer baseline when decoding in a single step.
Chuan-Fei Zhang, Tian-Hao Zhang, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICASSP5
2022 Adaptive Rounding Compensation for Post-training Quantization
Jinhui Lin, Song-Lu Chen, Ruiyao Zhang, Zhiwei Dong, Feng Chen 0040, Xu-Cheng Yin
ICONIP (5)7
2022 Semi-Supervised Fine-Grained Classification with Web Data via Noisy Sample Selection
abstract
For fine-grained classification, it is extremely difficult and costly to acquire the annotated data. Hence, some studies propose to use web data for fine-grained classification. However, the web data contains tremendous noisy labels, which can affect the classification results. Although many previous studies propose to discard noisy data via sample selection, they also discard some valid data. The valid data denotes hard or mislabeled samples that can enhance the robustness of the model. To solve the above problems, we propose a novel method to discard irrelevant noisy data from web data while keeping valid data for fine-grained classification. Specifically, we divide the web data into clean and noisy samples and then distinguish the noisy samples into open-set and close-set noises. Finally, the model is constructed in a semi-supervised manner, where the clean samples are used as the labeled set, and the close-set noises are used as the unlabeled set. Extensive experiments verify that our method can improve the classification performance by an average of 1.89% on three fine-grained benchmark datasets compared with the current methods. The experimental results prove the effectiveness of the combination of sample selection and semi-supervised training strategy.
Meng-Xuan Li, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICPR5
2022 DANet: Dynamic Attention to Spoof Patterns for Face Anti-Spoofing
abstract
Face anti-spoofing is a vital part to protect the security of face recognition systems. Many existing face anti-spoofing methods rely on convolutional neural networks (CNNs) and achieve competitive performance. However, due to the power of CNNs, these methods will extract information that is irrelevant to spoof patterns, such as acquisition equipment and environmental characteristics, which makes the network vulnerable to changes of the illumination or camera. In this work, we propose a plug-and-play module called DyAttention, which can improve the robustness against environmental changes. Moreover, we build a network named DANet with DyAttention, which can accurately capture the spoof patterns from coarse to fine. DANet can dynamically capture the texture differences between live and spoof samples in the facial area. Specifically, we use the spatial attention mechanism to generate a mask of the facial area. Then, we extract the intrinsic texture patterns and piecewise enhance them via dynamic activation for clean representation, where the texture patterns are not affected by the environmental and domain factors. Through experiments on three benchmark datasets, our DANet achieves state-of-the-art intra-dataset accuracy on CASIA-MFSD, Replay-Attack, and OULU-NPU. Meanwhile, DANet can enhance the cross-dataset performance between CASIA-MFSD and Replay-Attack, improving the average HTER by 1.3%.
Chun-Yu Sun, Song-Lu Chen, Xinjie Li 0002, Feng Chen 0040, Xu-Cheng Yin
ICPR4
2022 Anchor-Free Location Refinement Network for Small License Plate Detection
Zhen-Jia Li, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
PRCV (4)4
2022 Depth-Guided Progressive Network for Object Detection
abstract
Multi-scale object detection in natural scenes is still challenging. To enhance the multi-scale perception capability, some algorithms combine the lower-level and higher-level information via multi-scale feature fusion strategies. However, the inherent spatial properties among instances and relations between foreground and background are ignored. In addition, the human-defined “center-based” regression quality evaluation strategy, predicting a high-to-low score based on a linear relationship with the distance to the center of ground-truth box, is not robust to scale-variant objects. In this work, we propose a Depth-Guided Progressive Network (DGPNet) for multi-scale object detection. Specifically, besides the prediction of classification and localization, the depth is estimated and used to guide the image features in a weighted manner to obtain a better spatial representation. Therefore, depth estimation and 2D object detection are simultaneously learned via a unified network, where the depth features are merged as auxiliary information into the detection branch to enhance the discrimination among multi-scale objects. Moreover, to overcome the difficulty of empirically fitting the localization quality function, high-quality predicted boxes on scale-variant objects are more adaptively obtained by an IoU-aware progressive sampling strategy. We divide the sampling process into two stages, i.e., “statistical-aware” and “IoU-aware”. The former selects thresholds for positive samples based on statistical characteristics of multi-scale instances, and the latter further selects high-quality samples by IoU on the basis of the former. Therefore, the final ranking scores better reflect the quality of localization. Experiments verify that our method outperforms state-of-the-art methods on the KINS and Cityscapes dataset.
Jia-Wei Ma, Song-Lu Chen, Feng Chen 0040, Shu Tian, Jingyan Qin, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.4
2021 Fast Recognition for Multidirectional and Multi-type License Plates with 2D Spatial Attention
Qi Liu 0041, Song-Lu Chen, Zhen-Jia Li, Feng Chen 0040, Xu-Cheng Yin
ICDAR (4)5
2021 End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
Song-Lu Chen, Shu Tian, Jia-Wei Ma, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
Neurocomputing6
2020 Simultaneous End-to-End Vehicle and License Plate Detection With Multi-Branch Attention Neural Network
abstract
Vehicle and license plate detection plays an important role in intelligent transportation systems and is still a challenging task in real applications, such as on-road scenarios. Recently, Convolutional Neural Network (CNN)-based detectors achieve the state-of-the-art performance. However, it is difficult to efficiently detect the vehicle and license plate simultaneously in most cases. With a single network, the vehicle can affect the detection of the license plate due to the inclusion relation. In this paper, we propose an end-to-end deep neural network for detecting the vehicle and the license plate simultaneously in a given image, where two separate branches with different convolutional layers are designed for vehicle detection and license plate detection, respectively. In consideration of the license plate's small size and fairly obvious features as well as the vehicle's various size and rather complex features, the license plates are detected with low-level features and the vehicles are localized with multi-level features in corresponding convolutional layers. Moreover, a task-specific anchor design strategy is employed to obtain better predictions. Besides, the attention mechanisms and feature-fusion strategies are utilized to improve the detection performance of small-scale objects. A variety of experiments on real datasets and public datasets verify that our proposed method has fairly high accuracy and efficiency.
Song-Lu Chen, Jia-Wei Ma, Feng Chen 0040, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.4