VLDB 2026 Research / reviewers in the wild / expert
Song-Lu Chen
dblp:187/9111 · also Songlu Chen
· DBLP profile ↗
37ranked-venue papers
3as first author
35since 2021 · last 2026
0000-0002-0780-658XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 22 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 16 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Optical Flow: Latent Micro-Motion as Visual Evidence in UAV VideoabstractOptical flow has long dominated motion representation in video by focusing on explicit pixel displacement caused by object or camera movement. In this work, we argue that video also contains a largely overlooked form of motion information, namely latent micro-motion, which arises from subtle, structure-constrained responses of rigid components to physical interaction with the environment. We study this phenomenon in UAV video, a physically grounded setting where onboard structures are continuously exposed to aerodynamic forces. Although such micro-motions are low in amplitude and are often treated as noise or residual vibration, we show that they form a consistent visual signal that becomes observable through structure-aware and temporally aggregated analysis, even when using simple segmentation and coarse motion descriptors. Through an exploratory analysis, we demonstrate that micro-motion patterns exhibit clear structure and respond systematically to changes in wind conditions, with particularly strong sensitivity to wind direction and weaker dependence on wind magnitude in the examined scenarios. These observations suggest that micro-motion constitutes a distinct regime of motion information in video, complementary to explicit displacement, and motivate a broader reconsideration of how motion is represented and exploited in physically grounded multimedia scenarios. Bowen Zhang 0011, Song-Lu Chen, Xiaobin Zhu 0001, Xu-Cheng Yin |
ICMR | 3 |
| 2026 | Character recognition on continuous casting slabs via rotated detection and corner regression
Zhongjie Hu, Song-Lu Chen, Xiuxin Ge, Xu-Cheng Yin |
Pattern Recognit. Lett. | 4 |
| 2025 | AtomNet: Designing Tiny Models from Operators Under Extreme MCU ConstraintsabstractTiny machine learning (TinyML) has attracted heightened attention for its ability to provide low-cost and instantaneous performance on edge devices. Particularly, the commonly used microcontroller unit (MCU) imposes extreme constraints on peak memory (SRAM) and storage (Flash). Existing TinyML methods often rely on a customized and hard-to-obtain inference libraries, as well as necessitate a time-consuming search for a deployable architecture using advanced Neural Architecture Search (NAS) algorithms. To solve these problems, we fully exploit the resources on MCU and deduce hardware-oriented guidelines for designing models under extreme MCU constraints. In detail, we delve into thorough information about the atom operators by collecting the runtime data of Flash, SRAM, and latency to build a dataset named AtomDB. Based on AtomDB, several critical operator guidelines are established to fully utilize limited Flash and SRAM, while minimizing latency. By transferring the guidelines to analyze blocks, we propose a hybrid pattern that organizes appropriate blocks at different network stages to form the AtomNet, a more hardware-oriented architecture, to handle the former SRAM bottleneck and the latter Flash bottleneck. Extensive experiments demonstrate the effectiveness of the exploitation of the hardware characteristics. Remarkably, AtomNet pioneeringly achieve 3.5% accuracy enhancement and more than 15% latency reduction on 320KB MCU using readily available official inference libraries for ImageNet tasks, surpassing the current state-of-the-art method. Zhiwei Dong, Mingzhu Shen, Shihao Bai, Xiuying Wei, Jinyang Guo 0002, Ruihao Gong, Song-Lu Chen, Xianglong Liu 0001, Xu-Cheng Yin |
AAAI | 7 |
| 2025 | Tool Playgrounds: A Comprehensive and Analyzable Benchmark for LLM Tool InvocationabstractThe rapid advancement of large language models (LLMs) has paved the way for their use in solving real-world problems, which in turn has significantly driven the development of tool-assisted LLMs. This progress necessitates thorough evaluation methods. However, existing benchmarks typically only provide end-to-end scores but lack in-depth analysis and often suffer from issues such as instability. To address this gap, we have meticulously designed the Tool Playgrounds framework, a comprehensive, analyzable, and extensible benchmark. This framework evaluates boundary dimensions such as parameter missing interaction, parameter correction, tool failover, and leveraging internal knowledge. Our findings indicate that even the most advanced commercial models frequently overlook these essential aspects and face challenges in managing complex tool usage. To foster further research and development, we have made our code, dataset, and leaderboard publicly available on https://github.com/zhiwei-dong/ToolPlaygrounds. Zhiwei Dong, Ruihao Gong, Yang Yong, Yongqiang Yao, Song-Lu Chen, Xu-Cheng Yin |
ICASSP | 6 |
| 2025 | Data-Free Post-Training Quantization with Block-wise Enhanced Sample GenerationabstractData-free quantization is known for quantizing a pre-trained deep neural network without access to any training data, which applies to many real-world scenarios in that the training data is unavailable due to security, user privacy, or proprietary concerns. Most of the existing data-free quantization methods adopt a generator-quantization framework, which generator network to synthesize fake samples and Quantization-Aware Training (QAT) to quantize model. While the combination of the generator network and QAT can result in good accuracy for quantized models, the diversity of generated samples is lacking and quantizing a single model may take over 10 hours, which contrasts with Post-Training Quantization (PTQ)’s time-saving potential but poor accuracy. In order to address these issues, we have made improvements to the data generation and quantization process. In detail, 1) We propose Generator Exploration Enhancement (GEE) for utilizing the batch normalization statistics and adversarial sample exploration to enhance the quality and diversity of synthetic samples; 2) We introduce Block-wise Sample Generation (BSG) to collectively optimize individual blocks and the generator, leveraging PTQ as a foundation to boost workflow efficiency. Experiment results show that our proposed method improves both the performance and the efficiency of the data-free quantization compared to that of existing methods. Significantly, BSG achieves an 18% accuracy improvement and reduces quantization time by over 50% for 3-bit ResNet-18 in ImageNet tasks, surpassing the current state-of-the-art QAT method. Ruiyao Zhang, Zhiwei Dong, Shutong Ti, Song-Lu Chen, Xu-Cheng Yin |
ICASSP | 5 |
| 2025 | SPAN: A Salient Patch-Clue Aware Network for Cross-Domain Face Anti-SpoofingabstractFace anti-spoofing (FAS) plays a critical role in ensuring the security of face recognition system from different kinds of presentation attacks. Most existing FAS research faces several limitations: 1) insufficient consideration of the role of local fine-grained information, 2) the assumption that spoofing patterns are uniformly distributed across the entire image, neglecting the uneven distribution of spoofing clues, and 3) an overemphasis on intra-domain scenarios, leading to limited generalization capabilities for unseen domains. In this paper, we propose a Salient Patch-Clue Aware Network (SPAN) for cross-domain face anti-spoofing to tackle the aforementioned issues. Specifically, we use all patches cropped from the complete image as input to the FAS network, enabling the network to focus on local information while avoiding information loss. Additionally, we propose a patch perception mechanism to extract key regions containing salient spoofing clues, such as reflections and edges, thereby reducing interference from irrelevant information. Furthermore, we introduce a pixel perception mechanism to capture finer-grained details. Based on these two mechanisms, we design a Salient Clue Perception Module (SCPM). We conduct cross-domain experiments on CASIA-FASD, Idiap Replay-Attack, MSU-MFSD, and OULU-NPU datasets. Our method achieves state-of-the-art HTER on seven protocols, especially excelling on M&I to C and M&I to O, surpassing the second place by 9.79% and 6.19%, showcasing strong generalization capability. The codes are available at https://github.com/SPAN2025/SPAN. Liangfeng Zhang, Lei Chen 0069, Jinhui Lin, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
IJCNN | 6 |
| 2025 | Decoupling and Interaction: task coordination in single-stage object detection
Jia-Wei Ma, Shu Tian, Haixia Man, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin |
Multim. Tools Appl. | 4 |
| 2024 | Multi-task Learning for License Plate Recognition in Unconstrained Scenarios
Zhen-Lun Mo, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin |
ICDAR (1) | 2 |
| 2024 | HQOD: Harmonious Quantization for Object DetectionabstractTask inharmony problem commonly occurs in modern object detectors, leading to inconsistent qualities between classification and regression tasks. The predicted boxes with high classification scores but poor localization positions or low classification scores but accurate localization positions will worsen the performance of detectors after Non-Maximum Suppression. Furthermore, when object detectors collaborate with Quantization- Aware Training (QAT), we observe that the task inharmony problem will be further exacerbated, which is considered one of the main causes of the performance degradation of quantized detectors. To tackle this issue, we propose the Harmonious Quantization for Object Detection (HQOD) framework, which consists of two components. Firstly, we propose a task-correlated loss to encourage detectors to focus on improving samples with lower task harmony quality during QAT. Secondly, a harmonious Intersection over Union (IoU) loss is incorporated to balance the optimization of the regression branch across different IoU levels. The proposed HQOD can be easily integrated into different QAT algorithms and detectors. Remarkably, on the MS COCO dataset, our 4-bit ATSS with ResNet-50 backbone achieves a state-of-the- art mAP of 39.6%, even surpassing the full-precision one. Codes are available at https://github.com/Menace-Dragon/VP-QOD. Zhiwei Dong, Song-Lu Chen, Ruiyao Zhang, Shutong Ti, Feng Chen 0040, Xu-Cheng Yin |
ICME | 3 |
| 2024 | Towards Low-resource License Plate Recognition via Feature ShufflingabstractManual annotation is costly and limits the availability of sufficient annotated license plates for training recognition models. Small-scale license plate datasets (i.e., low-resource) often exhibit a long-tailed distribution in character classes at some character positions, primarily due to their limited variation in character permutations. Previous methods tend to prioritize head classes with high occurrence probability when applied to small-scale datasets. To solve this problem, we propose feature shuffling to balance the occurrence distribution across various character classes, thereby improving the recognition of tail classes with low occurrence probability. Moreover, we introduce global perception to holistically understand the overall character layout for effective feature shuffling. Extensive experiments on the small-scale UFPR and SSIG-SegPlate datasets demonstrate that our method achieves state-of-the-art results, with an average improvement of 43.70% over the baseline. Experiments on RodoSol and CCPD prove our method achieves state-of-the-art performance on large-scale datasets, verifying its generality. Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin |
ICME | 2 |
| 2024 | Improving Small License Plate Detection with Bidirectional Vehicle-Plate Relation
Songkang Dai, Song-Lu Chen, Qi Liu 0041, Chao Zhu 0003, Feng Chen 0040, Xu-Cheng Yin |
MMM (2) | 2 |
| 2024 | Irregular License Plate Recognition via Global Information Integration
Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
MMM (2) | 3 |
| 2024 | Integrated Recognition of Arbitrary-Oriented Multi-line Billet Number
Zhongjie Hu, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
PRCV (7) | 3 |
| 2024 | Improving license plate recognition via diverse stylistic plate generation
Qi Liu 0041, Song-Lu Chen, Yu-Xiang Chen, Xu-Cheng Yin |
Pattern Recognit. Lett. | 2 |
| 2024 | M3TTS: Multi-modal text-to-speech of multi-scale style control for dubbing
Li-Fang Wei, Xinyuan Qian 0001, Tian-Hao Zhang, Song-Lu Chen, Xu-Cheng Yin |
Pattern Recognit. Lett. | 5 |
| 2024 | Improving Multi-Type License Plate Recognition via Learning Globally and ContrastivelyabstractPrevious license plate recognition (LPR) methods have achieved impressive performance on single-type license plates. However, multi-type license plate recognition is still challenging due to various character layouts and fonts. There are two main problems: one is that recognition models are prone to incorrectly perceive the location of characters due to diverse character layouts, and the other is that characters of different categories may have similar glyphs due to various fonts, causing character misidentification. Therefore, to solve the above problems, we propose two plug-and-play modules based on an attention-based framework for multi-type license plate recognition. First, we propose a global modeling module to integrate character layout information to precisely perceive the location of characters, thus generating accurate predictions. Second, a position-aware contrastive learning module is proposed to enhance the robustness and discriminability of features to alleviate character misidentification of similar glyphs. Finally, to verify the effectiveness and generality, we apply the proposed modules to six baseline models, and the results demonstrate that the proposed method can achieve state-of-the-art performance on three multi-type license plate datasets. Moreover, extensive experiments prove that our proposed modules can significantly improve performance by 6.8% on RODOSOL-ALPR with a small parameter increase. Qi Liu 0041, Song-Lu Chen, Tian-Hao Zhang, Feng Chen 0040, Xu-Cheng Yin |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Sample Weighting with Hierarchical Equalization Loss for Dense Object DetectionabstractLabel assignment (LA) is one of the essential phases in the object detection paradigm and aims to classify samples as foreground or background. Current LA strategies generally discriminate samples by explicit thresholds and then calculate weighted losses based on their significances. However, existing methods mostly neglect to consider the importance of samples comprehensively due to the uneven distribution of objects and the limitations of detector structures. In this paper, we propose a hierarchical equalization loss (HEL) by reconsidering the underlying factors affecting sample weights. First, we mitigate sample imbalance at three progressive levels. (1) Task level. We propose task-reconciled weights (TRW) to overcome the effects caused by inter-task inconsistencies (i.e., the inherent differences of classification and localization). (2) Instance level. We propose instance-aware normalization (IAN) for reconstructing the distribution of sample weights within an instance to suppress environmental noise. (3) Pyramid level. We propose hierarchical modulation (HM) to alleviate the unbalanced distribution of multi-scale objects on feature pyramids. Then, we stack the above three mechanisms and formulate the effective weighted loss. Moreover, we propose a staggered candidate bag construction (SCBC) mechanism to further improve the robustness of our method. Without adding any extra overhead, HEL can improve the performance of representative detectors by an impressive margin. Equipped with HEL, a single “ResNet-50+FPN+Head” detector can achieve a performance of 41.9 AP on COCO under 1× schedule, outperforming other existing LA methods. Extensive experiments conducted on multiple backbones and datasets demonstrate the effectiveness of our method. Jia-Wei Ma, Lei Chen 0069, Shu Tian, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin |
IEEE Trans. Multim. | 5 |
| 2023 | Self-Convolution for Automatic Speech RecognitionabstractSelf-attention plays a significant role in recent automatic speech recognition (ASR) models with promising results. However, it suffers from high computational complexity and weak capability in modeling local information. In contrast, the convolutional neural network (CNN) is computationally effective and superior in learning local information. Whereas it fails in self-interaction and capturing long-range dependence among input tokens. Accordingly, we take their complementary advantages and propose a new module, namely self-convolution, to compensate for each individual limitations. Specifically, self-convolution generates convolution kernels at each token (to model local information) which are then used to convolve itself (for self-interaction). Moreover, we bring in global information during the generation of convolution kernel to enhance the learning of long-range dependencies. In this way, the advantages of self-attention and CNN are both utilized. We conduct rigorous experiments on LibriSpeech, Tedlium2, and AIShell1 datasets and demonstrate that our proposed self-convolution can achieve superior ASR performance than self-attention with less computational cost. Qi Liu 0041, Xinyuan Qian 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
ICASSP | 4 |
| 2023 | End-to-End Multi-line License Plate Recognition with Cascaded Perception
Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin |
ICDAR (5) | 1 |
| 2023 | Complex Glyph Enhancement for License Plate Generation
Yu-Xiang Chen, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
ICIG (1) | 3 |
| 2023 | InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu 0041, Xinyuan Qian 0001, Li-Fang Wei, Feng Chen 0040, Song-Lu Chen, Xu-Cheng Yin |
INTERSPEECH | 7 |
| 2023 | Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding
Tian-Hao Zhang, Haibo Qin, Zhi-Hao Lai, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xinyuan Qian 0001, Xu-Cheng Yin |
INTERSPEECH | 4 |
| 2023 | LiteHandNet: A Lightweight Hand Pose Estimation Network via Structural Feature Enhancement
Zhi-Yong Huang, Song-Lu Chen, Qi Liu 0041, Chong-Jian Zhang, Feng Chen 0040, Xu-Cheng Yin |
MMM (1) | 2 |
| 2023 | Feature Enhancement and Reconstruction for Small Object Detection
Chong-Jian Zhang, Song-Lu Chen, Qi Liu 0041, Zhi-Yong Huang, Feng Chen 0040, Xu-Cheng Yin |
MMM (1) | 2 |
| 2023 | Hypersphere guided embedding for masked face recognition
Xiaobin Zhu 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin, Lei Chen 0069 |
Pattern Recognit. Lett. | 3 |
| 2023 | Self-supervised contrastive speaker verification with nearest neighbor positive instances
Li-Fang Wei, Chuan-Fei Zhang, Tian-Hao Zhang, Song-Lu Chen, Xu-Cheng Yin |
Pattern Recognit. Lett. | 5 |
| 2022 | Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech RecognitionabstractNon-autoregressive (NAR) transformer models have been studied intensively in automatic speech recognition (ASR), and many NAR transformer models is to use the causal mask to limit token dependencies. However, the causal mask is designed for the left-to-right decoding process of the non-parallel autoregressive (AR) transformer, which is inappropriate for the parallel NAR transformer since it ignores the right-to-left contexts. Some methods are proposed to utilize right-to-left contexts with an extra decoder, but these methods increase the model complexity. To tackle the above problems, we propose a new non-autoregressive transformer with a unified bidirectional decoder (NAT-UBD), which can simultaneously utilize left-to-right and right-to-left contexts for ASR. However, direct use of bidirectional contexts will cause information leakage, which means the decoder output can be affected by the character information of the input in the same position. To avoid information leakage, we propose a novel attention mask and modify vanilla queries, keys, and values matrices for NAT-UBD. Experimental results verify that NAT-UBD can achieve character error rates (CERs) of 5.0%/5.5% on the Aishell-1 dev/test sets, outperforming all previous NAR transformer models. Moreover, NAT-UBD can run 49.8× faster than the AR transformer baseline when decoding in a single step. Chuan-Fei Zhang, Tian-Hao Zhang, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
ICASSP | 4 |
| 2022 | Adaptive Rounding Compensation for Post-training Quantization
Jinhui Lin, Song-Lu Chen, Ruiyao Zhang, Zhiwei Dong, Feng Chen 0040, Xu-Cheng Yin |
ICONIP (5) | 4 |
| 2022 | Semi-Supervised Fine-Grained Classification with Web Data via Noisy Sample SelectionabstractFor fine-grained classification, it is extremely difficult and costly to acquire the annotated data. Hence, some studies propose to use web data for fine-grained classification. However, the web data contains tremendous noisy labels, which can affect the classification results. Although many previous studies propose to discard noisy data via sample selection, they also discard some valid data. The valid data denotes hard or mislabeled samples that can enhance the robustness of the model. To solve the above problems, we propose a novel method to discard irrelevant noisy data from web data while keeping valid data for fine-grained classification. Specifically, we divide the web data into clean and noisy samples and then distinguish the noisy samples into open-set and close-set noises. Finally, the model is constructed in a semi-supervised manner, where the clean samples are used as the labeled set, and the close-set noises are used as the unlabeled set. Extensive experiments verify that our method can improve the classification performance by an average of 1.89% on three fine-grained benchmark datasets compared with the current methods. The experimental results prove the effectiveness of the combination of sample selection and semi-supervised training strategy. Meng-Xuan Li, Qi Liu 0041, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
ICPR | 4 |
| 2022 | DANet: Dynamic Attention to Spoof Patterns for Face Anti-SpoofingabstractFace anti-spoofing is a vital part to protect the security of face recognition systems. Many existing face anti-spoofing methods rely on convolutional neural networks (CNNs) and achieve competitive performance. However, due to the power of CNNs, these methods will extract information that is irrelevant to spoof patterns, such as acquisition equipment and environmental characteristics, which makes the network vulnerable to changes of the illumination or camera. In this work, we propose a plug-and-play module called DyAttention, which can improve the robustness against environmental changes. Moreover, we build a network named DANet with DyAttention, which can accurately capture the spoof patterns from coarse to fine. DANet can dynamically capture the texture differences between live and spoof samples in the facial area. Specifically, we use the spatial attention mechanism to generate a mask of the facial area. Then, we extract the intrinsic texture patterns and piecewise enhance them via dynamic activation for clean representation, where the texture patterns are not affected by the environmental and domain factors. Through experiments on three benchmark datasets, our DANet achieves state-of-the-art intra-dataset accuracy on CASIA-MFSD, Replay-Attack, and OULU-NPU. Meanwhile, DANet can enhance the cross-dataset performance between CASIA-MFSD and Replay-Attack, improving the average HTER by 1.3%. Chun-Yu Sun, Song-Lu Chen, Xinjie Li 0002, Feng Chen 0040, Xu-Cheng Yin |
ICPR | 2 |
| 2022 | Anchor-Free Location Refinement Network for Small License Plate Detection
Zhen-Jia Li, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin |
PRCV (4) | 2 |
| 2022 | Depth-Guided Progressive Network for Object DetectionabstractMulti-scale object detection in natural scenes is still challenging. To enhance the multi-scale perception capability, some algorithms combine the lower-level and higher-level information via multi-scale feature fusion strategies. However, the inherent spatial properties among instances and relations between foreground and background are ignored. In addition, the human-defined “center-based” regression quality evaluation strategy, predicting a high-to-low score based on a linear relationship with the distance to the center of ground-truth box, is not robust to scale-variant objects. In this work, we propose a Depth-Guided Progressive Network (DGPNet) for multi-scale object detection. Specifically, besides the prediction of classification and localization, the depth is estimated and used to guide the image features in a weighted manner to obtain a better spatial representation. Therefore, depth estimation and 2D object detection are simultaneously learned via a unified network, where the depth features are merged as auxiliary information into the detection branch to enhance the discrimination among multi-scale objects. Moreover, to overcome the difficulty of empirically fitting the localization quality function, high-quality predicted boxes on scale-variant objects are more adaptively obtained by an IoU-aware progressive sampling strategy. We divide the sampling process into two stages, i.e., “statistical-aware” and “IoU-aware”. The former selects thresholds for positive samples based on statistical characteristics of multi-scale instances, and the latter further selects high-quality samples by IoU on the basis of the former. Therefore, the final ranking scores better reflect the quality of localization. Experiments verify that our method outperforms state-of-the-art methods on the KINS and Cityscapes dataset. Jia-Wei Ma, Song-Lu Chen, Feng Chen 0040, Shu Tian, Jingyan Qin, Xu-Cheng Yin |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Fast Recognition for Multidirectional and Multi-type License Plates with 2D Spatial Attention
Qi Liu 0041, Song-Lu Chen, Zhen-Jia Li, Feng Chen 0040, Xu-Cheng Yin |
ICDAR (4) | 2 |
| 2021 | Robust Chinese License Plate Generation via Foreground Text and Background Separation
Qi Liu 0041, Song-Lu Chen, Xu-Cheng Yin |
ICIG (3) | 3 |
| 2021 | End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
Song-Lu Chen, Shu Tian, Jia-Wei Ma, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin |
Neurocomputing | 1 |
| 2020 | Semantic Bilinear Pooling for Fine-Grained RecognitionabstractNaturally, fine-grained recognition, e.g., vehicle identification or bird classification, has specific hierarchical labels, where fine categories are always harder to be classified than coarse categories. However, most of the recent deep learning based methods neglect the semantic structure of fine-grained objects and do not take advantage of the traditional fine-grained recognition techniques (e.g. coarse-to-fine classification). In this paper, we propose a novel framework with a two-branch network (coarse branch and fine branch), i.e., semantic bilinear pooling, for fine-grained recognition with a hierarchical label tree. This framework can adaptively learn the semantic information from the hierarchical levels. Specifically, we design a generalized cross-entropy loss for the training of the proposed framework to fully exploit the semantic priors via considering the relevance between adjacent levels and enlarge the distance between samples of different coarse classes. Furthermore, our method leverages only the fine branch when testing so that it adds no overhead to the testing time. Experimental results show that our proposed method achieves state-of-the-art performance on four public datasets. Xinjie Li 0002, Song-Lu Chen, Chao Zhu 0003, Xu-Cheng Yin |
ICPR | 3 |
| 2020 | Simultaneous End-to-End Vehicle and License Plate Detection With Multi-Branch Attention Neural NetworkabstractVehicle and license plate detection plays an important role in intelligent transportation systems and is still a challenging task in real applications, such as on-road scenarios. Recently, Convolutional Neural Network (CNN)-based detectors achieve the state-of-the-art performance. However, it is difficult to efficiently detect the vehicle and license plate simultaneously in most cases. With a single network, the vehicle can affect the detection of the license plate due to the inclusion relation. In this paper, we propose an end-to-end deep neural network for detecting the vehicle and the license plate simultaneously in a given image, where two separate branches with different convolutional layers are designed for vehicle detection and license plate detection, respectively. In consideration of the license plate's small size and fairly obvious features as well as the vehicle's various size and rather complex features, the license plates are detected with low-level features and the vehicles are localized with multi-level features in corresponding convolutional layers. Moreover, a task-specific anchor design strategy is employed to obtain better predictions. Besides, the attention mechanisms and feature-fusion strategies are utilized to improve the detection performance of small-scale objects. A variety of experiments on real datasets and public datasets verify that our proposed method has fairly high accuracy and efficiency. Song-Lu Chen, Jia-Wei Ma, Feng Chen 0040, Xu-Cheng Yin |
IEEE Trans. Intell. Transp. Syst. | 1 |