Qi Wang 0061

dblp:19/1924-61 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0003-0445-5603ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Receptive field weighted representation and context enhancement for SAR ship detection
Cheng Zha, Weidong Min, Qi Wang 0061, Di Gai, Hongyue Xiang
Expert Syst. Appl.4
2026 CLIP-based partial-wise prompt learning for unsupervised vehicle re-identification
Qi Wang 0061, Xin Xiong 0016
Expert Syst. Appl.3
2026 Residual Mamba-Driven Multiscale Attentive Network With Boundary Enhancement for IoT-Enabled Medical Image Segmentation
abstract
In IoT-enabled intelligent healthcare systems, medical images are frequently acquired in real-time from heterogeneous imaging sensors such as dermoscopic devices and MRI scanners. In resource-limited or edge-deployed settings, achieving precise and rapid image segmentation plays a crucial role in facilitating early diagnosis and supporting clinical decision-making. This paper proposes a medical image segmentation method based on a residual Mamba backbone network, combining a multi-scale gated attention (MGA) module with a boundary enhancement (BE) module to effectively enhance the model’s feature representation and boundary localization capabilities. Specifically, the R-Mamba backbone network combines the advantages of statespace modeling and convolutional feature extraction, achieving efficient fusion of global context and local details. The MGA module dynamically captures multi-scale semantic information through dilated convolutions and gating mechanisms, enhancing the model’s adaptability to targets of different scales and shapes. The BE module significantly strengthens boundary representation and fine-grained structural segmentation through multi-scale convolutions and channel-spatial dual attention mechanisms. Additionally, this paper designs a multi-loss function joint optimization strategy to comprehensively constrain region overlap, pixel classification, and structural consistency. Experimental validation on ISIC skin lesion and LGG brain tumor datasets shows competitive performance compared to several mainstream models under the tested conditions.
Guoqiang Ren, Qi Wang 0061, Jieying Tu, Pengxiang Su, Hengrui Liu, Di Gai, Peng Luo 0005, Shuxiao Li
IEEE Internet Things J.2
2026 Dual-Student Adversarial Framework With Discriminator and Consistency-Driven Learning for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised medical image segmentation is essential for alleviating the cost of manual annotation in clinical applications. However, existing methods often suffer from unreliable pseudo-labels and confirmation bias in consistency-based training, which can lead to unstable optimization and degraded performance. To address these issues, a novel method named dual-Student adversarial framework with discriminator and consistency-driven learning for semi-supervised medical image segmentation is proposed. Specifically, an adversarial learning-based segmentation refinement (ALSR) module is designed to encourage prediction diversity between two student networks and leverage a shared discriminator for adversarial refinement of pseudo-labels. To further stabilize the consistency process, a residual exponential moving average (R-EMA) is applied in the uncertainty estimation with inter-instance consistency measurement (UIM) module to construct a robust teacher model, while noisy voxel predictions are selectively filtered based on uncertainty estimation. In addition, a Contrastive Representation Stabilization (CRS) module is developed to enhance voxel-level semantic alignment by performing contrastive learning only on confident regions, improving feature discriminability and structural consistency. Extensive experiments on benchmark datasets demonstrate that our method consistently outperforms prior state-of-the-art approaches.
Haifan Wu, Yuhan Geng, Di Gai, Jieying Tu, Xin Xiong 0016, Qi Wang 0061
IEEE J. Biomed. Health Informatics6
2026 Outliers Adaptation Exploration and Centroids Matching Label Refinement for Unsupervised Person Re-Identification
abstract
Existing unsupervised person re-identification (Re-ID) methods obtain pseudo-labels mainly by clustering to optimize the model. Most methods only use clustered instances to provide supervised information for model training without considering the possible value information of un-clustered outliers. Some methods use un-clustered outliers for training and show promising results, but their strategies for using un-clustered outliers are limited in applicability. Furthermore, unsatisfactory feature embedding and imperfect clustering cannot guarantee that instances within clusters have the same identity, resulting in generated pseudo-labels containing noise. To solve the above problems, outliers adaptation exploration (OAE) and centroids matching label refinement (CMLR) are proposed in this paper. First, OAE is designed to explore the value information of un-clustered outliers. OAE implements the nearest clustered instance neighborhood constraint and the maximum centroid distance constraint on the un-clustered outliers based on the feature distribution to update the clusters, thereby achieving a healthy balance between the number of training samples and model accuracy. Second, CMLR is proposed to alleviate the inherent noise in pseudo-labels. CMLR considers the similarity relationship between clustered instances and clusters in accordance with the cluster distribution to refine pseudo-labels, prompting features to learn from probability distributions with cluster distribution similarity relationship information. Extensive experiments demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art performance in unsupervised learning and unsupervised domain adaptation settings.
Weidong Min, Qi Wang 0061, Ziyang Deng
IEEE Trans. Multim.4
2025 Vehiclemae: View-Asymmetry Mutual Learning for Vehicle Re-Identification Pre-Training Via Masked Autoencoders
Qi Wang 0061, Dong Wang 0080, Di Gai, Xin Xiong 0016, Jiyang Xu, Ruihua Zhou
ICCV1
2025 Threefold Encoder Interaction: Hierarchical Multi-Grained Semantic Alignment for Cross-Modal Food Retrieval
abstract
Current cross-modal food retrieval approaches focus mainly on the global visual appearance of food without explicitly considering multi-grained information. Additionally, direct calculation of the global similarity of image-recipe pairs is not particularly effective in terms of latent alignment, which suffers from mismatch during the mutual image-recipe retrieval process. This paper proposes a threefold encoder interaction (TEI) cross-modal food retrieval framework to maintain the multi-granularity of food images and the multi-levels of textual recipes to address the aforementioned challenges. The TEI framework comprises an image encoder, a recipe encoder, and a multi-grained interaction encoder. We simultaneously propose a multi-grained relation-aware attention (MRA) embedded in the multi-grained interaction encoder to capture multi-grained food visual features. The multi-grained interaction similarity scores are calculated to better establish the multi-grained correlation between recipe and image entities based on the extracted hierarchical textual and multi-grained visual features. Finally, a hierarchical multi-grained semantic alignment loss is designed to supervise the whole process of cross-modal training using the multi-grained interaction similarity scores. Extensive qualitative and quantitative experiments on the Recipe1M dataset have demonstrated that the proposed TEI framework achieves multi-grained semantic alignment between image and text modalities and is superior to other state-of-the-art methods in cross-modal food retrieval tasks.
Qi Wang 0061, Dong Wang 0080, Weidong Min, Di Gai, Cheng Zha, Yuling Zhong
IEEE Trans. Multim.1
2024 Joint training with local soft attention and dual cross-neighbor label smoothing for unsupervised person re-identification
abstract
Existing unsupervised person re-identification approaches fail to fully capture the fine-grained features of local regions, which can result in people with similar appearances and different identities being assigned the same label after clustering. The identity-independent information contained in different local regions leads to different levels of local noise. To address these challenges, joint training with local soft attention and dual cross-neighbor label smoothing (DCLS) is proposed in this study. First, the joint training is divided into global and local parts, whereby a soft attention mechanism is proposed for the local branch to accurately capture the subtle differences in local regions, which improves the ability of the re-identification model in identifying a person’s local significant features. Second, DCLS is designed to progressively mitigate label noise in different local regions. The DCLS uses global and local similarity metrics to semantically align the global and local regions of the person and further determines the proximity association between local regions through the cross information of neighboring regions, thereby achieving label smoothing of the global and local regions throughout the training process. In extensive experiments, the proposed method outperformed existing methods under unsupervised settings on several standard person re-identification datasets.
Weidong Min, Qi Wang 0061, Qingpeng Zeng, Shimiao Cui, Jiongjin Chen
Comput. Vis. Media4
2024 SAM-driven MAE pre-training and background-aware meta-learning for unsupervised vehicle re-identification
abstract
Distinguishing identity-unrelated background information from discriminative identity information poses a challenge in unsupervised vehicle re-identification (Re-ID). Re-ID models suffer from varying degrees of background interference caused by continuous scene variations. The recently proposed segment anything model (SAM) has demonstrated exceptional performance in zero-shot segmentation tasks. The combination of SAM and vehicle Re-ID models can achieve efficient separation of vehicle identity and background information. This paper proposes a method that combines SAM-driven mask autoencoder (MAE) pre-training and background-aware meta-learning for unsupervised vehicle Re-ID. The method consists of three sub-modules. First, the segmentation capacity of SAM is utilized to separate the vehicle identity region from the background. SAM cannot be robustly employed in exceptional situations, such as those with ambiguity or occlusion. Thus, in vehicle Re-ID downstream tasks, a spatially-constrained vehicle background segmentation method is presented to obtain accurate background segmentation results. Second, SAM-driven MAE pre-training utilizes the aforementioned segmentation results to select patches belonging to the vehicle and to mask other patches, allowing MAE to learn identity-sensitive features in a self-supervised manner. Finally, we present a background-aware meta-learning method to fit varying degrees of background interference in different scenarios by combining different background region ratios. Our experiments demonstrate that the proposed method has state-of-the-art performance in reducing background interference variations.
Dong Wang 0080, Qi Wang 0061, Weidong Min, Di Gai, Yuhan Geng
Comput. Vis. Media2
2024 Vision-language constraint graph representation learning for unsupervised vehicle re-identification
Dong Wang 0080, Qi Wang 0061, Zhiwei Tu, Weidong Min, Xin Xiong 0016, Yuling Zhong, Di Gai
Expert Syst. Appl.2
2024 Feature ensemble network for medical image segmentation with multi-scale atrous transformer
abstract
Abstract Recent years have witnessed notable advancements in medical image segmentation through deep convolutional neural networks. However, a notable limitation lies in the local operation of convolution, which hinders the ability to fully exploit global semantic information. To overcome the challenges prevalent in medical image segmentation, the feature ensemble network with multi‐scale atrous transformer is proposed. At the core of the approach lies the multi‐scale contextual integration module, which is based on the multi‐scale atrous transformer and facilitates contextual integration of multi‐level features. To extract discriminative fine‐grained features of the target region, a hybrid attention mechanism that synergistically combines spatial and channel attention, thereby sharpening the model's focus on crucial target information within high‐level features, is incorporated. Additionally, the channel‐aware feature reconstruction module is introduced as an innovative component engineered to tackle feature similarity issues across different categories. This module performs feature reconstruction based on channel perception, effectively widening the feature gap between categories and enhancing the segmentation capability. It is worth mentioning that our approach surpasses the state‐of‐the‐art method using three benchmark datasets in medical image segmentation.
Di Gai, Yuhan Geng, Xin Xiong 0016, Ruihua Zhou, Qi Wang 0061
IET Image Process.7
2024 Semi-supervised medical image classification based on class prototype matching for soft pseudo labels with consistent regularization
Di Gai, Ruonan Xiong, Weidong Min, Qi Wang 0061, Xin Xiong 0016, Chunjiang Peng
Multim. Tools Appl.5
2023 SAR ship localization method with denoising and feature refinement
abstract
Synthetic Aperture Radar (SAR) ship detection is greatly important to marine transportation monitoring and fishery resource management. To improve the detection accuracy of small ships, an SAR ship localization method with Denoising and Feature Refinement (DFR) is proposed in this paper. It consists of three parts. The first part is the denoising module, which uses non-local mean to suppress the speckle noise of the SAR image . The second part is Hierarchical Feature Fusion (HFF) module. It can integrate more low-level features by adding skip connections. This prevents the low-level spatial position information of the fused features from being diluted by high-level semantic information, therefore it is beneficial to the detection of small ships. The third part is a center-based ship predictor with Feature Refinement (FR). The FR module is proposed to refine the features and reduce the background interference, which is conducive to locate ships more accurately. Extensive experiments are conducted. The experimental results show that after adding the denoising and FR modules, the value of AP 0.5 is increased by 1.7% and 2.3%, respectively, which proves the effectiveness of these two modules. In inshore and offshore scenarios, the AP 0.5 values of DFR are 0.884 and 0.966, respectively, achieving the best results. The proposed method can also be generalized to mark lesion locations in medical images and detect offshore oil production platforms .
Cheng Zha, Weidong Min, Wei Li 0151, Xin Xiong 0016, Qi Wang 0061
Eng. Appl. Artif. Intell.6
2023 Scene-adaptive crowd counting method based on meta learning with dual-input network DMNet
Weidong Min, Qi Wang 0061, Qiyan Fu
Frontiers Comput. Sci.4
2023 Memory-efficient document layout analysis method using LD-net
Weidong Min, Qi Wang 0061, Zitai Wei
Multim. Tools Appl.3
2023 Dual similarity pre-training and domain difference encouragement learning for vehicle re-identification in the wild
Qi Wang 0061, Yuling Zhong, Weidong Min, Di Gai
Pattern Recognit.1
2023 Spatiotemporal Learning Transformer for Video-Based Human Pose Estimation
abstract
Multi-frame human pose estimation has long been an appealing and fundamental issue in visual perception. Owing to the frequent rapid motion and pose occlusion in videos, this task is extremely challenging. Current state-of-the-art methods seek to model spatiotemporal features by equally fusing each frame in the local sequence, which weakens the target frame information. In addition, existing approaches usually emphasize more on deep features while ignoring the detailed information implied in the shallow feature maps, resulting in the dropping of crucial features. To address the above problems, we propose an effective framework, namely spatiotemporal learning transformer for video-based human pose estimation (SLT-Pose), which consists of a Personalized Feature Extraction Module (PFEM), Self-feature Refinement Module (SRM), Cross-frame Temporal Learning Module (CTLM) and Disentangled Keypoint Detector (DKD). To be specific, we propose PFEM which extracts and modulates the individual frame features to adapt to the varying human shape, and integrates single-frame features to obtain the spatiotemporal features. We further present SRM to establish global correlation spatial cues on the target frame to attain the refinement feature. Then, a CTLM is designed to search for the information most closely related to the target frame from the spatiotemporal features to intensify the interaction between the target frame and the local sequence, using both the shallow detailed and the deep semantic representations. Finally, we employ DKD to extract the disentangled characteristics of each joint and encode the articulated joint pairs in the human body, promoting the model to reasonably and accurately predict the keypoint heatmaps. Extensive experiments on three huamn motion benchmarks, including PoseTrack2017, PoseTrack2018, and Sub-JHMDB dataset, demonstrate that SLT-Pose plays favorably against state-of-the-art approaches in terms of both objective evaluation and subjective visual performance.
Di Gai, Runyang Feng, Weidong Min, Xiaosong Yang, Pengxiang Su, Qi Wang 0061
IEEE Trans. Circuits Syst. Video Technol.6
2023 Human Skeleton Feature Optimizer and Adaptive Structure Enhancement Graph Convolution Network for Action Recognition
abstract
Human action recognition based on the graph convolution network (GCN) is a hot topic in computer vision. Existing GCN-based methods fail to capture internal implicit information when extracting action features, thereby leading to over-smoothing in the training stage. These issues result in poor performance and inaccurate extraction of action features. To address these problems, a new GCN is constructed. In this paper, a human skeleton feature optimizer (SFO) and adaptive structure enhancement graph convolution network (ASE-GCN) for action recognition are proposed in an end-to-end manner. To obtain discriminative features, the SFO is proposed to construct a new skeleton representation for action recognition through the connection criterion, which extracts the internal implicit information of action. The action feature of the joint coordinates is extracted by graph structure mask (GSM), directed graph mapping (DGM), and adaptive pooling operation (APO) in the proposed ASE-GCN network. The GSM acts as the regularizer of skeleton structure information to strengthen the representation of the graph structure. The DGM correlates the directed graph with human motion information through kinematic principle, and the APO strengthens the global high-frequency features to alleviate over-smoothing. The proposed method achieves comparable or superior results over state-of-the-art methods when used in experiments on two large public-scale datasets, NTU-RGB+D and Kinetics.
Xin Xiong 0016, Weidong Min, Qi Wang 0061, Cheng Zha
IEEE Trans. Circuits Syst. Video Technol.3
2023 Need Only One More Point (NOOMP): Perspective Adaptation Crowd Counting in Complex Scenes
abstract
Recently, solving the crowd counting problem under occlusion and complex perspective is a hot but difficult topic. Existing methods mainly constructed counters in parallel perspective, but when facing complex perspective, such as the influences of height difference and heavy occlusions, they fail to get good accuracy. To alleviate these problems, this work proposes a novel and interesting framework NOOMP (Need Only One More Point) for perspective adaptation crowd counting task in complex nature scenes. Firstly, this work considers that the common scenes in our daily life usually have the height difference, which brings complex perspective to crowd counting. So, a new labeled method, Absolute-geometry Gaussian Generation is proposed, which only needs one more point for each person in image and gets better accuracy. Secondly, the NOOMP framework consists of meta-learning structure and uses the few-shot way to train the counting model, which can implement the perspective adaptation effective and solve the problem of high label cost. Thirdly, for fitting the characteristic of few-shot learning, this work proposes a new Multi-head Parallel Network (MPNet) for NOOMP. The feature of crowd is extracted by MPNet, which is a hybrid structure composed of shallow network and deep network. This network can save the features of shallow network and the deeper network effectively, which makes MPNet performs well in NOOMP. In addition, this work collects a new dataset, named Multiple Height Differences in Mall (MHDM) for NOOMP, which contains images of different views and height differences from shopping malls and supermarkets. Experiments based on MHDM and other benchmarks show that the NOOMP has good performances in model accuracy and works well for solving perspective change problem.
Qi Wang 0061, Guowei Zhan, Weidong Min, Shimiao Cui
IEEE Trans. Multim.2
2023 Trade-off background joint learning for unsupervised vehicle re-identification
Qi Wang 0061, Weidong Min, Di Gai, Haowen Luo
Vis. Comput.2
2022 ECNFP: Edge-constrained network using a feature pyramid for image inpainting
Zitai Wei, Weidong Min, Qi Wang 0061
Expert Syst. Appl.3
2022 Multiple Granularity Spatiotemporal Network for Sea Surface Temperature Prediction
abstract
Sea surface temperature (SST) prediction has an important practical value in marine disaster prevention and mitigation. Most current methods only use the temporal correlation of SST during prediction, but the spatial correlation is not considered, resulting in low prediction accuracies. In addition, the changing trend of SST as reflected by the single granularity feature is unreliable, and the degrees of dependence between historical SST and future SST tend to vary. In order to overcome these issues, the multiple granularity spatiotemporal network (MGSN) is proposed for SST prediction. The proposed method consists of three parts. First, a multibranch network structure is constructed to extract different temporal features of different granularities. Second, a temporal dependence representation module is developed to represent the different degrees of dependence between historical SST and predicted SST in the temporal dimension. Third, the spatiotemporal fusion prediction module is used to achieve a spatiotemporal prediction of the SST and fuse the prediction results of different granular features. Comparative experiments have been conducted. The experimental results show that the root-mean-square error (RMSE) of the proposed method is reduced by 0.1360, 0.1608, and 0.1448 compared with the RMSE of convolutional LSTM (ConvLSTM), when predicting SST for the next one day, three days, and seven days, respectively. Our method has strong spatiotemporal feature modeling capabilities and is suitable for regional SST prediction.
Cheng Zha, Weidong Min, Xin Xiong 0016, Qi Wang 0061
IEEE Geosci. Remote. Sens. Lett.5
2022 Traffic Sign Recognition Based on Semantic Scene Understanding and Structural Traffic Sign Location
abstract
Traffic sign recognition (TSR) plays an important role in driving assistance system and traffic safety insurance. However, existing methods focus on extracting features of traffic signs and ignore the constraints of spatial positional relationships between traffic signs and other objects in the scene. This way results in incorrectly detecting other similar objects as traffic signs and failing to detect very small traffic signs. A TSR method based on semantic scene understanding and structural traffic sign location is proposed in this study to solve the aforementioned problems. A scene structure model based on the constraints of spatial positional relationships between traffic signs and other objects is proposed to establish trusted search regions. An improved Light-weight RefineNet is used to analyze and understand a scene semantically and accurately and then segment objects in complicated environments precisely. A new network multiscale densely connected object detector (MDCOD) based on densely connected style, multiscale feature fusion, and improved K-means++ algorithms is proposed to recognize very small traffic signs. The trusted traffic signs are found by filtering false candidates outside the scene structure model. The proposed method is tested on Tsinghua-Tencent 100K and German Traffic Sign Detection Benchmark datasets and achieves accuracies of 92.8% and 99.90%, respectively, outperforming the existing methods.
Weidong Min, Ruikang Liu, Daojing He, Qingting Wei, Qi Wang 0061
IEEE Trans. Intell. Transp. Syst.6
2022 Inter-Domain Adaptation Label for Data Augmentation in Vehicle Re-Identification
abstract
Vehicle re-identification (Re-ID) methods often fail to achieve robust performance due to insufficient training data and domain diversities. Although state-of-the-art methods apply image-to-image translation or web data to achieve data augmentation, the construct of new datasets will not only introduce noise, but also undergo a mismatch issue with the source domain. Moreover, the label noise of cross-domain data in existing label distribution technologies cannot be alleviated. In this paper, a multi-domain joint learning with inter-domain adaptation label smoothing regularization (IALSR) is proposed using a semi-supervised learning framework. The overall framework consists of two parts. In one part, a multi-domain joint network (MJNet) is proposed to learn multiple vehicle attributes simultaneously. The output of the training model is employed to group several inter-domain subsets, which are regarded as different domains. To adapt to domain diversities, style transfer models are learned for each pair of subsets to generate free and rich data as a novel data augmentation approach. In the other part, IALSR, which preserves self-similarity and domain-transitivity, is designed to smooth the noise of style-transferred data. Upon our basis, we further introduce the web data to verify the superiority of the IALSR. The results of extensive experimental on two large-scale vehicle Re-ID datasets demonstrate that the proposed approach is superior to other state-of-the-art ones.
Qi Wang 0061, Weidong Min, Cheng Zha, Zitai Wei
IEEE Trans. Multim.1
2022 3D Skeleton and Two Streams Approach to Person Re-identification Using Optimized Region Matching
abstract
Person re-identification (Re-ID) is a challenging and arduous task due to non-overlapping views, complex background, and uncontrollable occlusion in video surveillance. An existing method for capturing pedestrian local region information is to divide person regions into horizontal stripes, which may lead to invalid features and erroneous learning. To solve this problem, this paper proposes a 3D skeleton and a two-stream approach to person Re-ID. The first stream of the method uses the 3D skeleton for background filtering and region segmentation. The second stream uses Siamese net to extract the global descriptor. The features of the two streams are fused to preserve the integrity of the person. An optimized region matching method for metric learning is designed. Extensive comparing experiments were conducted with state-of-the-art Re-ID methods on the Market-1501, CUHK03, and DukeMTMC-reID datasets. Experimental results show that the proposed method outperforms the existing methods in recognition accuracy.
Weidong Min, Tiemei Huang, Deyu Lin, Qi Wang 0061
ACM Trans. Multim. Comput. Commun. Appl.6
2021 MSR-FAN: Multi-scale residual feature-aware network for crowd counting
abstract
Abstract Crowd counting aims to count the number of people in crowded scenes, which is important to the security systems, traffic control and so on. The existing methods typically using local features cannot properly handle the perspective distortion and the varying scales in congested scene images, and henceforth perform wrong people counting. To alleviate this issue, this study proposes a multi‐scale residual feature‐aware network (MSR‐FAN) that combines multi‐scale features using multiple receptive field sizes and learns the feature‐aware information on each image. The MSR‐FAN is trained end‐to‐end to generate high‐quality density map and evaluate the crowd number. The method consists of three parts. To handle the perspective changes problem, the first part, the direction‐based feature‐enhanced network, is designed to encode the perspective information in four directions based on the initial image feature. The second part, the proposed multi‐scale residual block module, gets the global information to handle the represent the regional feature better. This module explores features of different scales as well as reinforce the global feature. The third part, the feature‐aware block, is designed to extract the feature hidden in the different channels. Experiment results based on benchmark datasets show that the proposed approach outperforms the existing state‐of‐the‐art methods.
Weidong Min, Xin Wei 0002, Qi Wang 0061, Qiyan Fu, Zitai Wei
IET Image Process.4
2021 PFLU and FPFLU: Two novel non-monotonic activation functions in convolutional neural networks
Weidong Min, Qi Wang 0061, Song Zou, Xinhao Chen
Neurocomputing3
2021 Viewpoint adaptation learning with cross-view distance metric for robust vehicle re-identification
Qi Wang 0061, Weidong Min, Ziyuan Yang 0001, Xin Xiong 0016
Inf. Sci.1
2020 Discriminative fine-grained network for vehicle re-identification using two-stage re-ranking
Qi Wang 0061, Weidong Min, Daojing He, Song Zou, Tiemei Huang, Ruikang Liu
Sci. China Inf. Sci.1
2019 New approach to vehicle license plate location based on new model YOLO-L and plate pre-identification
abstract
Currently, the conventional license plate location method fails to detect the license plate under complex road environments such as severe weather conditions and viewpoint changes. Besides, it is difficult for license plate location method based on machine learning to precisely locate the area of license plate. Moreover, license plate location method may incorrectly detect similar objects such as billboards and road signs as license plates. To alleviate these problems, this article proposes a new approach to vehicle license plate location based on new model YOLO‐L and plate pre‐identification. The new model improves in two aspects to precisely locate the area of license plate. First, it uses k‐means++ clustering algorithm to select the best number and size of plate candidate boxes. Second, it modifies the structure and depth of YOLOv2 model. Plate pre‐identification algorithm can effectively distinguish license plates from similar objects. The experimental results show that authors’ proposed method not only achieves a precision of 98.86% and a recall of 98.86%, which outperforms the existing methods, but also has high efficiency in real time.
Weidong Min, Qi Wang 0061, Qingpeng Zeng, Yanqiu Liao
IET Image Process.3