EDBT 2026 Demo / reviewers in the wild / expert
Weidong Min
dblp:29/4943
· DBLP profile ↗
78ranked-venue papers
12as first author
57since 2021 · last 2026
0000-0003-2526-2181ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 4 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 6 since 2021Computer networks · 4 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Receptive field weighted representation and context enhancement for SAR ship detection
Cheng Zha, Weidong Min, Qi Wang 0061, Di Gai, Hongyue Xiang |
Expert Syst. Appl. | 2 |
| 2026 | Occlusion-robust 3D human pose estimation via Weighted Joint Trajectory Topology and Guided Semantic Enhancement
Weidong Min, Jian-Sheng Wu, Ziyang Deng |
Image Vis. Comput. | 2 |
| 2026 | Dual contrastive learning with graph masking: A self-supervised framework for multi-view clustering
Jian-Sheng Wu, Wen-Ting Li, Jun-Yun Wu, Weidong Min |
Neural Networks | 4 |
| 2026 | Dual-Domain Balanced Channel-Spatial Mixed Attention in Convolutional Neural NetworksabstractAbstract Channel-Spatial mixed attention methods have been proven to enhance the representation capability of convolutional neural networks. Most existing channel-spatial mixed attention methods decouple the channel domain and spatial domain, investing different orders of computational complexity in calculating attention for the two domains. However, both the channel domain and spatial domain are equally important and deserve comparable levels of complexity in attention computation. To solve this problem, we propose a novel channel-spatial mixed attention method, which is called D ual-domain B alanced C hannel- S patial M ixed A ttention (DBCSMA). The input tensor is first decomposed into channel and spatial domains using global pooling. Then, the same order of complexity is dedicated to calculating attention in both the channel domain and the spatial domain, and a mixed attention tensor is generated using multiplicative aggregation. Finally, input channel information and spatial information are simultaneously recalibrated using multiplication of this mixed attention tensor and the input tensor. Experimental results demonstrate that DBCSMA achieves superior performance compared to most existing channel-spatial mixed attention methods while using lower complexity. Weidong Min, Shimiao Cui, Lixin Zhan, Xiangrong Zeng, Jiongjin Chen |
Neural Process. Lett. | 2 |
| 2026 | Multi-modal graph-aware pre-training for disease classification and cross-modal retrieval in chest radiology
Weidong Min, Haifan Wu |
Pattern Recognit. | 2 |
| 2026 | Dual-level data-anchor association learning via k-partite graph factorization for multi-view clustering
Jian-Sheng Wu, Yuan-Tong Cheng, Weidong Min, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2026 | High-order Aligned Deep Complementary and View-Specific Similarity Graphs for Unsupervised Multi-View Feature Selection
Jian-Sheng Wu, Jia-Tao Yu, Jun-Yun Wu, Weidong Min, Wei-Shi Zheng 0001 |
Pattern Recognit. | 4 |
| 2026 | Micro-Expression Analysis Based on Self-Adaptive Pseudo-Labeling and Residual Connected Channel Attention MechanismsabstractMicro-expressions can reveal genuine emotions that are not easily concealed, making them invaluable in fields such as psychotherapy and criminal interrogation. However, existing pseudo-labeling-based methods for micro-expression analysis have two major limitations. First, pseudo-labels generated by the sliding window do not account for the actual proportion of micro-expressions in the video, which leads to inaccurate labeling. Second, they predominantly focus on overall features, thereby neglecting subtle features. In this paper, we propose a micro-expression analysis method called Spot-Then-Recognize Method (STRM), which integrates spotting and recognition tasks. To address the first limitation, we propose a Self-Adaptive Pseudo-labeling Method (SAPM) that dynamically assigns pseudo-labels to micro-expression frames according to their actual proportion in the video sequence, thereby improving labeling accuracy. To address second limitation, we design a Multi-Scale Residual Channel Attention Network (MSRCAN) to effectively extract subtle micro-expression features. The MSRCAN comprises three modules: Multi-Scale Shared Network (MSSN), Spotting Network, and Recognition Network. The MSSN initially extracts micro-expression features by performing multi-scale feature extraction with Residual Connected Channel Attention Modules (RCCAM), which are then refined in the spotting and recognition networks. We conducted comprehensive experiments on three short video datasets (CASME II, SMIC-E-HS, SMIC-E-NIR) and two long video datasets (CAS(ME)2, SAMMLV). Experimental results show that our proposed method significantly outperforms existing methods, achieving an overall performance of 58.24%, a 19.62% improvement, and a $1.51\times $ gain over the baseline in terms of micro-expression analysis. Weidong Min |
IEEE Trans. Image Process. | 2 |
| 2026 | Outliers Adaptation Exploration and Centroids Matching Label Refinement for Unsupervised Person Re-IdentificationabstractExisting unsupervised person re-identification (Re-ID) methods obtain pseudo-labels mainly by clustering to optimize the model. Most methods only use clustered instances to provide supervised information for model training without considering the possible value information of un-clustered outliers. Some methods use un-clustered outliers for training and show promising results, but their strategies for using un-clustered outliers are limited in applicability. Furthermore, unsatisfactory feature embedding and imperfect clustering cannot guarantee that instances within clusters have the same identity, resulting in generated pseudo-labels containing noise. To solve the above problems, outliers adaptation exploration (OAE) and centroids matching label refinement (CMLR) are proposed in this paper. First, OAE is designed to explore the value information of un-clustered outliers. OAE implements the nearest clustered instance neighborhood constraint and the maximum centroid distance constraint on the un-clustered outliers based on the feature distribution to update the clusters, thereby achieving a healthy balance between the number of training samples and model accuracy. Second, CMLR is proposed to alleviate the inherent noise in pseudo-labels. CMLR considers the similarity relationship between clustered instances and clusters in accordance with the cluster distribution to refine pseudo-labels, prompting features to learn from probability distributions with cluster distribution similarity relationship information. Extensive experiments demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art performance in unsupervised learning and unsupervised domain adaptation settings. Weidong Min, Qi Wang 0061, Ziyang Deng |
IEEE Trans. Multim. | 3 |
| 2025 | Learning missing instances in intact and projection spaces for incomplete multi-view unsupervised feature selection
Jian-Sheng Wu, Hong-Wei Yu, Yanlan Li, Weidong Min |
Appl. Intell. | 4 |
| 2025 | Multi-view deep reciprocal nonnegative matrix factorization
Jun-Yun Wu, Jian-Sheng Wu, Weidong Min |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Confident local similarity graphs for unsupervised feature selection on incomplete multi-view data
Hong-Wei Yu, Jun-Yun Wu, Jian-Sheng Wu, Weidong Min |
Knowl. Based Syst. | 4 |
| 2025 | Constructing a smoothed Leaky ReLU using a linear combination of the smoothed ReLU and identity function
Weidong Min, Ziyang Deng |
Neural Comput. Appl. | 2 |
| 2025 | Facial expression transformation for anime-style image based on decoder control and attention mask
Xinhao Rao, Weidong Min, Ziyang Deng |
Signal Process. Image Commun. | 2 |
| 2025 | Threefold Encoder Interaction: Hierarchical Multi-Grained Semantic Alignment for Cross-Modal Food RetrievalabstractCurrent cross-modal food retrieval approaches focus mainly on the global visual appearance of food without explicitly considering multi-grained information. Additionally, direct calculation of the global similarity of image-recipe pairs is not particularly effective in terms of latent alignment, which suffers from mismatch during the mutual image-recipe retrieval process. This paper proposes a threefold encoder interaction (TEI) cross-modal food retrieval framework to maintain the multi-granularity of food images and the multi-levels of textual recipes to address the aforementioned challenges. The TEI framework comprises an image encoder, a recipe encoder, and a multi-grained interaction encoder. We simultaneously propose a multi-grained relation-aware attention (MRA) embedded in the multi-grained interaction encoder to capture multi-grained food visual features. The multi-grained interaction similarity scores are calculated to better establish the multi-grained correlation between recipe and image entities based on the extracted hierarchical textual and multi-grained visual features. Finally, a hierarchical multi-grained semantic alignment loss is designed to supervise the whole process of cross-modal training using the multi-grained interaction similarity scores. Extensive qualitative and quantitative experiments on the Recipe1M dataset have demonstrated that the proposed TEI framework achieves multi-grained semantic alignment between image and text modalities and is superior to other state-of-the-art methods in cross-modal food retrieval tasks. Qi Wang 0061, Dong Wang 0080, Weidong Min, Di Gai, Cheng Zha, Yuling Zhong |
IEEE Trans. Multim. | 3 |
| 2025 | Enhanced virtual gesture generation through joint contact representation and contact area alignment
Weidong Min, Ziyang Deng |
Vis. Comput. | 2 |
| 2024 | DSU-GAN: A robust frontal face recognition approach based on generative adversarial network
Deyu Lin, Huanxin Wang, Weidong Min, Chenguang Yao, Yong Liang Guan 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Joint training with local soft attention and dual cross-neighbor label smoothing for unsupervised person re-identificationabstractExisting unsupervised person re-identification approaches fail to fully capture the fine-grained features of local regions, which can result in people with similar appearances and different identities being assigned the same label after clustering. The identity-independent information contained in different local regions leads to different levels of local noise. To address these challenges, joint training with local soft attention and dual cross-neighbor label smoothing (DCLS) is proposed in this study. First, the joint training is divided into global and local parts, whereby a soft attention mechanism is proposed for the local branch to accurately capture the subtle differences in local regions, which improves the ability of the re-identification model in identifying a person’s local significant features. Second, DCLS is designed to progressively mitigate label noise in different local regions. The DCLS uses global and local similarity metrics to semantically align the global and local regions of the person and further determines the proximity association between local regions through the cross information of neighboring regions, thereby achieving label smoothing of the global and local regions throughout the training process. In extensive experiments, the proposed method outperformed existing methods under unsupervised settings on several standard person re-identification datasets. Weidong Min, Qi Wang 0061, Qingpeng Zeng, Shimiao Cui, Jiongjin Chen |
Comput. Vis. Media | 3 |
| 2024 | SAM-driven MAE pre-training and background-aware meta-learning for unsupervised vehicle re-identificationabstractDistinguishing identity-unrelated background information from discriminative identity information poses a challenge in unsupervised vehicle re-identification (Re-ID). Re-ID models suffer from varying degrees of background interference caused by continuous scene variations. The recently proposed segment anything model (SAM) has demonstrated exceptional performance in zero-shot segmentation tasks. The combination of SAM and vehicle Re-ID models can achieve efficient separation of vehicle identity and background information. This paper proposes a method that combines SAM-driven mask autoencoder (MAE) pre-training and background-aware meta-learning for unsupervised vehicle Re-ID. The method consists of three sub-modules. First, the segmentation capacity of SAM is utilized to separate the vehicle identity region from the background. SAM cannot be robustly employed in exceptional situations, such as those with ambiguity or occlusion. Thus, in vehicle Re-ID downstream tasks, a spatially-constrained vehicle background segmentation method is presented to obtain accurate background segmentation results. Second, SAM-driven MAE pre-training utilizes the aforementioned segmentation results to select patches belonging to the vehicle and to mask other patches, allowing MAE to learn identity-sensitive features in a self-supervised manner. Finally, we present a background-aware meta-learning method to fit varying degrees of background interference in different scenarios by combining different background region ratios. Our experiments demonstrate that the proposed method has state-of-the-art performance in reducing background interference variations. Dong Wang 0080, Qi Wang 0061, Weidong Min, Di Gai, Yuhan Geng |
Comput. Vis. Media | 3 |
| 2024 | Collaborative and Discriminative Subspace Learning for unsupervised multi-view feature selection
Jian-Sheng Wu, Yanlan Li, Jun-Xiao Gong, Weidong Min |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Vision-language constraint graph representation learning for unsupervised vehicle re-identification
Dong Wang 0080, Qi Wang 0061, Zhiwei Tu, Weidong Min, Xin Xiong 0016, Yuling Zhong, Di Gai |
Expert Syst. Appl. | 4 |
| 2024 | MSCNet: Dense vehicle counting method based on multi-scale dilated convolution channel-aware deep network
Qiyan Fu, Weidong Min, Chunbo Li |
GeoInformatica | 2 |
| 2024 | Image inpainting network based on multi-level attention mechanismabstractAbstract Image inpainting networks based on deep learning techniques have been widely used in many important fields. However, most inpainting networks fail to generate desirable repaired images. This may be due to their failure to extract effective features and accurately assign high weights to the undamaged regions. To alleviate these problems, an image inpainting network based on gated convolution and multi‐level attention mechanism (IIN‐GCMAM) is proposed in this paper. This network follows encoder–decoder architecture, consisting of the gated convolution encoder (GC‐encoder) and the multi‐level attention mechanism decoder (MAM‐decoder). The GC‐encoder weighs the extracted features with gated convolutions, which reduces the interference caused by the damaged regions. The multi‐level attention mechanism employed in the MAM‐decoder uses multi‐scale feature maps spatially and channel‐wise to improve the consistency in global structure and the fineness of repaired results. Extensive experiments are conducted on the common datasets, Paris StreetView and CelebA. Experimental results indicate that the proposed IIN‐GCMAM can achieve a good performance on the common evaluation metrics and visual effects. It can achieve 0.0408, 0.720, and 22.27 in MAE, SSIM, and PSNR at the mask ratio of 50%–60%, respectively. Hongyue Xiang, Weidong Min, Zitai Wei, Ziyang Deng |
IET Image Process. | 2 |
| 2024 | Spatial Decomposition and Aggregation for Attention in Convolutional Neural NetworksabstractChannel attention has been shown to improve the performance of deep convolutional neural networks efficiently. Channel attention adaptively recalibrates the importance of each channel, determining what to attend to. However, channel attention only encodes inter-channel information but neglects the importance of positional information. Positional information is crucial in determining where to attend to. To address this issue, we propose a novel channel-spatial attention method named Spatial-Decomposition-Aggregation Attention (SDAA) method. First, a high-axis spatial direction is decomposed into multiple low-axis spatial directions. Then, a shared transformation sub-unit establishes attention in each low-axis space direction. Next, all the low-axis attention masks are aggregated into a high-axis attention mask. Finally, the generated high-axis attention mask is fused into the input features, thus enhancing the input features. Essentially, our method is a divide-and-conquer process. Experimental results demonstrate that our SDAA method outperforms the existing channel-spatial attention methods. Weidong Min, Hongyue Xiang, Cheng Zha, Qiyan Fu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | A Novel High-Precision and Low-Latency Abandoned Object Detection Method Under the Hybrid Cloud-Fog Computing ArchitectureabstractAbandoned object Detection (Aod) is of critical importance in the field of public safety. However, the demand on detection accuracy and latency hinders the development of ubiquitous Aod in safety protection, especially for some surveillance devices with relatively low-computational capacity. To this end, a novel high-precision and low-latency Aod method under the hybrid cloud-fog computing architecture is proposed in this article. To be specific, a YOLO-various hidden (YOLO-VH) Aod network model, which is integrated with an efficient dynamic convolution-based ghost module and a Haar wavelet-based downsampling convolution module, is presented to improve the detection accuracy of Aod. In addition, a flexible task offloading strategy is proposed to offload some of the Aod tasks based on the expectation cursor, which is designed to determine the local optimal offloading amount at different times. Finally, extensive experiments are conducted to verify the performance of our proposal through simulations. Our proposal exhibits a reduction of approximately 5.31 million parameters and 30.4 GFLOPs in computation compared with YOLOv9, while demonstrating performance improvements of 25.0% and 38.8% relative to cloud and fog computing, respectively. Furthermore, the total latency for image acquisition, task offloading, and task processing has been observed to be approximately 60% and 15% lower than cloud and fog computing, respectively. Deyu Lin, Junhao Zhao, Fuxin Yu, Weidong Min, Yong Liang Guan 0001 |
IEEE Internet Things J. | 4 |
| 2024 | Semi-supervised medical image classification based on class prototype matching for soft pseudo labels with consistent regularization
Di Gai, Ruonan Xiong, Weidong Min, Qi Wang 0061, Xin Xiong 0016, Chunjiang Peng |
Multim. Tools Appl. | 3 |
| 2024 | Sign language recognition based on skeleton and SK3D-Residual network
Zhanlu Huangfu, Weidong Min, Tianqi Ding, Yanqiu Liao |
Multim. Tools Appl. | 3 |
| 2024 | ShuffleNeMt: modern lightweight convolutional neural network architecture
Weidong Min, Guowei Zhan, Qiyan Fu |
Pattern Anal. Appl. | 2 |
| 2024 | Improved channel attention methods via hierarchical pooling and reducing information loss
Weidong Min, Junwei Han 0001, Shimiao Cui |
Pattern Recognit. | 2 |
| 2024 | DCSG: data complement pseudo-label refinement and self-guided pre-training for unsupervised person re-identification
Jiongjin Chen, Weidong Min, Lixin Zhan |
Vis. Comput. | 3 |
| 2023 | Feature Distribution Fitting with Direction-Driven Weighting for Few-Shot Images ClassificationabstractFew-shot learning has received increasing attention and witnessed significant advances in recent years. However, most of the few-shot learning methods focus on the optimization of training process, and the learning of metric and sample generating networks. They ignore the importance of learning the ground-truth feature distributions of few-shot classes. This paper proposes a direction-driven weighting method to make the feature distributions of few-shot classes precisely fit the ground-truth distributions. The learned feature distributions can generate an unlimited number of training samples for the few-shot classes to avoid overfitting. Specifically, the proposed method consists of two optimization strategies. The direction-driven strategy is for capturing more complete direction information that can describe the feature distributions. The similarity-weighting strategy is proposed to estimate the impact of different classes in the fitting procedure and assign corresponding weights. Our method outperforms the current state-of-the-art performance by an average of 3% for 1-shot on standard few-shot learning benchmarks like miniImageNet, CIFAR-FS, and CUB. The excellent performance and compelling visualization show that our method can more accurately estimate the ground-truth distributions. Xin Wei 0002, Huan Wan, Weidong Min |
AAAI | 4 |
| 2023 | SAR ship localization method with denoising and feature refinementabstractSynthetic Aperture Radar (SAR) ship detection is greatly important to marine transportation monitoring and fishery resource management. To improve the detection accuracy of small ships, an SAR ship localization method with Denoising and Feature Refinement (DFR) is proposed in this paper. It consists of three parts. The first part is the denoising module, which uses non-local mean to suppress the speckle noise of the SAR image . The second part is Hierarchical Feature Fusion (HFF) module. It can integrate more low-level features by adding skip connections. This prevents the low-level spatial position information of the fused features from being diluted by high-level semantic information, therefore it is beneficial to the detection of small ships. The third part is a center-based ship predictor with Feature Refinement (FR). The FR module is proposed to refine the features and reduce the background interference, which is conducive to locate ships more accurately. Extensive experiments are conducted. The experimental results show that after adding the denoising and FR modules, the value of AP 0.5 is increased by 1.7% and 2.3%, respectively, which proves the effectiveness of these two modules. In inshore and offshore scenarios, the AP 0.5 values of DFR are 0.884 and 0.966, respectively, achieving the best results. The proposed method can also be generalized to mark lesion locations in medical images and detect offshore oil production platforms . Cheng Zha, Weidong Min, Wei Li 0151, Xin Xiong 0016, Qi Wang 0061 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Scene-adaptive crowd counting method based on meta learning with dual-input network DMNet
Weidong Min, Qi Wang 0061, Qiyan Fu |
Frontiers Comput. Sci. | 2 |
| 2023 | Multi-level correlation learning for multi-view unsupervised feature selection
Jian-Sheng Wu, Jun-Xiao Gong, Weidong Min |
Knowl. Based Syst. | 4 |
| 2023 | Semantic Segmentation of Point Cloud With Novel Neural Radiation Field ConvolutionabstractPoint cloud semantic segmentation predicts the semantic class of each point, which can help AI machines perceive the real 3-D world. Recently, precision for large-scale point cloud is limited by complex scenarios, data occlusion, and massive data, which remains. This letter proposes neural radiance field convolution (NeRFConv) for large-scale point cloud semantic analysis. First, to conquer some existing methods represent the point cloud local space as the relative position without considering the rotation property of the point cloud. A new spatial direction representation called neural radiance field 7-D is constructed to denote the local point cloud space, which is unaffected by the rotation of the coordinate axes. Then, to be robust against the varying density of the point cloud, an adaptive convolution method is proposed to aggregate the local point cloud features and obtain more expressive local features. Finally, the proposed method achieves state-of-the-art performance on the test dataset. In particular, mIoU of 58.1% and 58.3% were achieved the SensatUrban and SementicKITTI datasets, respectively. Wei Li 0151, Lixin Zhan, Weidong Min, Chenglu Wen |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Downsampling in uniformly-spaced windows for coding-based Palmprint recognition
Ziyuan Yang 0001, Lu Leng, Weidong Min |
Multim. Tools Appl. | 3 |
| 2023 | Memory-efficient document layout analysis method using LD-net
Weidong Min, Qi Wang 0061, Zitai Wei |
Multim. Tools Appl. | 2 |
| 2023 | Dual similarity pre-training and domain difference encouragement learning for vehicle re-identification in the wild
Qi Wang 0061, Yuling Zhong, Weidong Min, Di Gai |
Pattern Recognit. | 3 |
| 2023 | Spatiotemporal Learning Transformer for Video-Based Human Pose EstimationabstractMulti-frame human pose estimation has long been an appealing and fundamental issue in visual perception. Owing to the frequent rapid motion and pose occlusion in videos, this task is extremely challenging. Current state-of-the-art methods seek to model spatiotemporal features by equally fusing each frame in the local sequence, which weakens the target frame information. In addition, existing approaches usually emphasize more on deep features while ignoring the detailed information implied in the shallow feature maps, resulting in the dropping of crucial features. To address the above problems, we propose an effective framework, namely spatiotemporal learning transformer for video-based human pose estimation (SLT-Pose), which consists of a Personalized Feature Extraction Module (PFEM), Self-feature Refinement Module (SRM), Cross-frame Temporal Learning Module (CTLM) and Disentangled Keypoint Detector (DKD). To be specific, we propose PFEM which extracts and modulates the individual frame features to adapt to the varying human shape, and integrates single-frame features to obtain the spatiotemporal features. We further present SRM to establish global correlation spatial cues on the target frame to attain the refinement feature. Then, a CTLM is designed to search for the information most closely related to the target frame from the spatiotemporal features to intensify the interaction between the target frame and the local sequence, using both the shallow detailed and the deep semantic representations. Finally, we employ DKD to extract the disentangled characteristics of each joint and encode the articulated joint pairs in the human body, promoting the model to reasonably and accurately predict the keypoint heatmaps. Extensive experiments on three huamn motion benchmarks, including PoseTrack2017, PoseTrack2018, and Sub-JHMDB dataset, demonstrate that SLT-Pose plays favorably against state-of-the-art approaches in terms of both objective evaluation and subjective visual performance. Di Gai, Runyang Feng, Weidong Min, Xiaosong Yang, Pengxiang Su, Qi Wang 0061 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Human Skeleton Feature Optimizer and Adaptive Structure Enhancement Graph Convolution Network for Action RecognitionabstractHuman action recognition based on the graph convolution network (GCN) is a hot topic in computer vision. Existing GCN-based methods fail to capture internal implicit information when extracting action features, thereby leading to over-smoothing in the training stage. These issues result in poor performance and inaccurate extraction of action features. To address these problems, a new GCN is constructed. In this paper, a human skeleton feature optimizer (SFO) and adaptive structure enhancement graph convolution network (ASE-GCN) for action recognition are proposed in an end-to-end manner. To obtain discriminative features, the SFO is proposed to construct a new skeleton representation for action recognition through the connection criterion, which extracts the internal implicit information of action. The action feature of the joint coordinates is extracted by graph structure mask (GSM), directed graph mapping (DGM), and adaptive pooling operation (APO) in the proposed ASE-GCN network. The GSM acts as the regularizer of skeleton structure information to strengthen the representation of the graph structure. The DGM correlates the directed graph with human motion information through kinematic principle, and the APO strengthens the global high-frequency features to alleviate over-smoothing. The proposed method achieves comparable or superior results over state-of-the-art methods when used in experiments on two large public-scale datasets, NTU-RGB+D and Kinetics. Xin Xiong 0016, Weidong Min, Qi Wang 0061, Cheng Zha |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Need Only One More Point (NOOMP): Perspective Adaptation Crowd Counting in Complex ScenesabstractRecently, solving the crowd counting problem under occlusion and complex perspective is a hot but difficult topic. Existing methods mainly constructed counters in parallel perspective, but when facing complex perspective, such as the influences of height difference and heavy occlusions, they fail to get good accuracy. To alleviate these problems, this work proposes a novel and interesting framework NOOMP (Need Only One More Point) for perspective adaptation crowd counting task in complex nature scenes. Firstly, this work considers that the common scenes in our daily life usually have the height difference, which brings complex perspective to crowd counting. So, a new labeled method, Absolute-geometry Gaussian Generation is proposed, which only needs one more point for each person in image and gets better accuracy. Secondly, the NOOMP framework consists of meta-learning structure and uses the few-shot way to train the counting model, which can implement the perspective adaptation effective and solve the problem of high label cost. Thirdly, for fitting the characteristic of few-shot learning, this work proposes a new Multi-head Parallel Network (MPNet) for NOOMP. The feature of crowd is extracted by MPNet, which is a hybrid structure composed of shallow network and deep network. This network can save the features of shallow network and the deeper network effectively, which makes MPNet performs well in NOOMP. In addition, this work collects a new dataset, named Multiple Height Differences in Mall (MHDM) for NOOMP, which contains images of different views and height differences from shopping malls and supermarkets. Experiments based on MHDM and other benchmarks show that the NOOMP has good performances in model accuracy and works well for solving perspective change problem. Qi Wang 0061, Guowei Zhan, Weidong Min, Shimiao Cui |
IEEE Trans. Multim. | 4 |
| 2023 | Trade-off background joint learning for unsupervised vehicle re-identification
Qi Wang 0061, Weidong Min, Di Gai, Haowen Luo |
Vis. Comput. | 3 |
| 2022 | ECNFP: Edge-constrained network using a feature pyramid for image inpainting
Zitai Wei, Weidong Min, Qi Wang 0061 |
Expert Syst. Appl. | 2 |
| 2022 | Non-Watertight Polygonal Surface Reconstruction From Building Point Cloud via Connection and Data FitabstractPolygonal planes are used to simplify the modeling of the building point cloud, which is widely applied for city planning, 3-D cadastral management, and model rendering. Much of the current research mainly focuses on watertight polygon model reconstruction from the 3-D model data with few missing values, which is highly hypothetical and restricted in application. Therefore, this letter proposes a non-watertight PolyFit (NW-PolyFit) algorithm based on the fit degree of connection and data fitting to reconstruct a non-watertight polygon model from the data with missing values. First, to refine the planar structure in buildings, the refined supporting planes are generated from the detected point cloud primitives. Second, to retain the sharp features, the$\delta $expansion planes are generated from their corresponding supporting planes. We get the relationship between primitives by intersecting those$\delta $expansion planes. Finally, to eliminate the influence of missing data, a fit-and-remove strategy is proposed to filter the generated candidate’s faces, which achieves non-watertight modeling. Experiments show that the NW-PolyFit achieved similar modeling effects for completed data compared with the state-of-the-art methods. The NW-PolyFit achieves non-watertight modeling from the point cloud data with massive missing values while other methods do not. Yifei Jiang, Weidong Min, Wei Li 0151 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Content-Based Remote Sensing Image Retrieval Based on Fuzzy Rules and a Fuzzy DistanceabstractThe methods in remote sensing image retrieval (RSIR) usually search the whole retrieval data set in the retrieval process, which takes much time and is unnecessary. To reduce the overall search time, this letter proposes a new retrieval scheme based on fuzzy rules. The proposed method calculates the fuzzy class membership of images using two ways. The first way predicts the fuzzy class membership by convolutional neural network (CNN). The other uses the image-to-class distance that is a distance between an image and each class on the training data set. The two fuzzy class memberships are used to measure the classification confidence, and a query image is classified into three fuzzy sets, namely, “low classification confidence,” “medium classification confidence,” and “high classification confidence,” based on the classification confidence. The fuzzy rules are built according to fuzzy classification to choose the search space for each fuzzy set. The final search space is determined by the two search spaces obtained by fuzzy rules. Moreover, the fuzzy distance between a query image and a retrieved image is used to improve the retrieval performance, which is calculated according to their fuzzy class memberships and the Euclidean distance between the two images. The experimental results on University of California, Merced data set (UCMD) and PatternNet databases show that our proposed method can not only enhance the retrieval performance but also reduce the search time in comparison to other state-of-the-art techniques. Famao Ye, Meng Dong, Dajun Li, Weidong Min |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Multiple Granularity Spatiotemporal Network for Sea Surface Temperature PredictionabstractSea surface temperature (SST) prediction has an important practical value in marine disaster prevention and mitigation. Most current methods only use the temporal correlation of SST during prediction, but the spatial correlation is not considered, resulting in low prediction accuracies. In addition, the changing trend of SST as reflected by the single granularity feature is unreliable, and the degrees of dependence between historical SST and future SST tend to vary. In order to overcome these issues, the multiple granularity spatiotemporal network (MGSN) is proposed for SST prediction. The proposed method consists of three parts. First, a multibranch network structure is constructed to extract different temporal features of different granularities. Second, a temporal dependence representation module is developed to represent the different degrees of dependence between historical SST and predicted SST in the temporal dimension. Third, the spatiotemporal fusion prediction module is used to achieve a spatiotemporal prediction of the SST and fuse the prediction results of different granular features. Comparative experiments have been conducted. The experimental results show that the root-mean-square error (RMSE) of the proposed method is reduced by 0.1360, 0.1608, and 0.1448 compared with the RMSE of convolutional LSTM (ConvLSTM), when predicting SST for the next one day, three days, and seven days, respectively. Our method has strong spatiotemporal feature modeling capabilities and is suitable for regional SST prediction. Cheng Zha, Weidong Min, Xin Xiong 0016, Qi Wang 0061 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Traffic Sign Recognition Based on Semantic Scene Understanding and Structural Traffic Sign LocationabstractTraffic sign recognition (TSR) plays an important role in driving assistance system and traffic safety insurance. However, existing methods focus on extracting features of traffic signs and ignore the constraints of spatial positional relationships between traffic signs and other objects in the scene. This way results in incorrectly detecting other similar objects as traffic signs and failing to detect very small traffic signs. A TSR method based on semantic scene understanding and structural traffic sign location is proposed in this study to solve the aforementioned problems. A scene structure model based on the constraints of spatial positional relationships between traffic signs and other objects is proposed to establish trusted search regions. An improved Light-weight RefineNet is used to analyze and understand a scene semantically and accurately and then segment objects in complicated environments precisely. A new network multiscale densely connected object detector (MDCOD) based on densely connected style, multiscale feature fusion, and improved K-means++ algorithms is proposed to recognize very small traffic signs. The trusted traffic signs are found by filtering false candidates outside the scene structure model. The proposed method is tested on Tsinghua-Tencent 100K and German Traffic Sign Detection Benchmark datasets and achieves accuracies of 92.8% and 99.90%, respectively, outperforming the existing methods. Weidong Min, Ruikang Liu, Daojing He, Qingting Wei, Qi Wang 0061 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Inter-Domain Adaptation Label for Data Augmentation in Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) methods often fail to achieve robust performance due to insufficient training data and domain diversities. Although state-of-the-art methods apply image-to-image translation or web data to achieve data augmentation, the construct of new datasets will not only introduce noise, but also undergo a mismatch issue with the source domain. Moreover, the label noise of cross-domain data in existing label distribution technologies cannot be alleviated. In this paper, a multi-domain joint learning with inter-domain adaptation label smoothing regularization (IALSR) is proposed using a semi-supervised learning framework. The overall framework consists of two parts. In one part, a multi-domain joint network (MJNet) is proposed to learn multiple vehicle attributes simultaneously. The output of the training model is employed to group several inter-domain subsets, which are regarded as different domains. To adapt to domain diversities, style transfer models are learned for each pair of subsets to generate free and rich data as a novel data augmentation approach. In the other part, IALSR, which preserves self-similarity and domain-transitivity, is designed to smooth the noise of style-transferred data. Upon our basis, we further introduce the web data to verify the superiority of the IALSR. The results of extensive experimental on two large-scale vehicle Re-ID datasets demonstrate that the proposed approach is superior to other state-of-the-art ones. Qi Wang 0061, Weidong Min, Cheng Zha, Zitai Wei |
IEEE Trans. Multim. | 2 |
| 2022 | 3D Skeleton and Two Streams Approach to Person Re-identification Using Optimized Region MatchingabstractPerson re-identification (Re-ID) is a challenging and arduous task due to non-overlapping views, complex background, and uncontrollable occlusion in video surveillance. An existing method for capturing pedestrian local region information is to divide person regions into horizontal stripes, which may lead to invalid features and erroneous learning. To solve this problem, this paper proposes a 3D skeleton and a two-stream approach to person Re-ID. The first stream of the method uses the 3D skeleton for background filtering and region segmentation. The second stream uses Siamese net to extract the global descriptor. The features of the two streams are fused to preserve the integrity of the person. An optimized region matching method for metric learning is designed. Extensive comparing experiments were conducted with state-of-the-art Re-ID methods on the Market-1501, CUHK03, and DukeMTMC-reID datasets. Experimental results show that the proposed method outperforms the existing methods in recognition accuracy. Weidong Min, Tiemei Huang, Deyu Lin, Qi Wang 0061 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Multi-view face generation via unpaired images
Yanni Zou, Weidong Min, Xin Xiong 0016 |
Vis. Comput. | 3 |
| 2021 | Illumination-Enhanced Crowd Counting Based on IC-Net in Low Lighting Conditions
Weidong Min |
ICIG (1) | 2 |
| 2021 | Multimodal graph inference network for scene graph generation
Jingwen Duan, Weidong Min, Deyu Lin, Xin Xiong 0016 |
Appl. Intell. | 2 |
| 2021 | MSR-FAN: Multi-scale residual feature-aware network for crowd countingabstractAbstract Crowd counting aims to count the number of people in crowded scenes, which is important to the security systems, traffic control and so on. The existing methods typically using local features cannot properly handle the perspective distortion and the varying scales in congested scene images, and henceforth perform wrong people counting. To alleviate this issue, this study proposes a multi‐scale residual feature‐aware network (MSR‐FAN) that combines multi‐scale features using multiple receptive field sizes and learns the feature‐aware information on each image. The MSR‐FAN is trained end‐to‐end to generate high‐quality density map and evaluate the crowd number. The method consists of three parts. To handle the perspective changes problem, the first part, the direction‐based feature‐enhanced network, is designed to encode the perspective information in four directions based on the initial image feature. The second part, the proposed multi‐scale residual block module, gets the global information to handle the represent the regional feature better. This module explores features of different scales as well as reinforce the global feature. The third part, the feature‐aware block, is designed to extract the feature hidden in the different channels. Experiment results based on benchmark datasets show that the proposed approach outperforms the existing state‐of‐the‐art methods. Weidong Min, Xin Wei 0002, Qi Wang 0061, Qiyan Fu, Zitai Wei |
IET Image Process. | 2 |
| 2021 | PFLU and FPFLU: Two novel non-monotonic activation functions in convolutional neural networks
Weidong Min, Qi Wang 0061, Song Zou, Xinhao Chen |
Neurocomputing | 2 |
| 2021 | Viewpoint adaptation learning with cross-view distance metric for robust vehicle re-identification
Qi Wang 0061, Weidong Min, Ziyuan Yang 0001, Xin Xiong 0016 |
Inf. Sci. | 2 |
| 2021 | Joint adaptive manifold and embedding learning for unsupervised feature selection
Jian-Sheng Wu, Meng-Xiao Song, Weidong Min, Jian-Huang Lai, Wei-Shi Zheng 0001 |
Pattern Recognit. | 3 |
| 2021 | Driver Yawning Detection Based on Subtle Facial Action RecognitionabstractVarious investigations have shown that driver fatigue is the main cause of traffic accidents. Research on the use of computer vision techniques to detect signs of fatigue from facial actions, such as yawning, has demonstrated good potential. However, accurate and robust detection of yawning is difficult because of the complicated facial actions and expressions of drivers in the real driving environment. Several facial actions and expressions have the same mouth deformation as yawning. Thus, a novel approach to detecting yawning based on subtle facial action recognition is proposed in this study to alleviate the abovementioned problems. A 3D deep learning network with a low time sampling characteristic is proposed for subtle facial action recognition. This network uses 3D convolutional and bidirectional long short-term memory networks for spatiotemporal feature extraction and adopts SoftMax for classification. A keyframe selection algorithm is designed to select the most representative frame sequence from subtle facial actions. This algorithm rapidly eliminates redundant frames using image histograms with low computation cost and detects outliers by median absolute deviation. A series of experiments are also conducted on YawDD benchmark and self-collected datasets. Compared with several state-of-the-art methods, the proposed method has high yawning detection rates and can effectively distinguish yawning from similar facial actions. Hao Yang 0027, Li Liu 0010, Weidong Min, Xiaosong Yang, Xin Xiong 0016 |
IEEE Trans. Multim. | 3 |
| 2020 | S3D-CNN: skeleton-based 3D consecutive-low-pooling neural network for fall detection
Xin Xiong 0016, Weidong Min, Wei-Shi Zheng 0001, Pin Liao, Hao Yang 0027 |
Appl. Intell. | 2 |
| 2020 | Discriminative fine-grained network for vehicle re-identification using two-stage re-ranking
Qi Wang 0061, Weidong Min, Daojing He, Song Zou, Tiemei Huang, Ruikang Liu |
Sci. China Inf. Sci. | 2 |
| 2020 | An Energy-Saving Routing Integrated Economic Theory With Compressive Sensing to Extend the Lifespan of WSNsabstractA novel intercluster routing which simultaneously takes the energy efficiency in both intracluster and intercluster phases into account is proposed in this article, with the aim of extending the lifespan of the wireless sensor networks (WSNs). In the intracluster phase, the data are acquired based on the compressive sensing (CS) theory to cut down extra energy consumption resulted from spatial-temporal correlation. As for the intercluster phase, the economic welfare theory is applied to balance the energy depletion among different clusters. To this end, a novel concept of energy efficiency welfare (E2W) is proposed to promote energy equilibrium during the process of intercluster routing decision making. Subsequently, an energy-saving intercluster routing integrated economic theory with CS (EIREC) is presented and detailed. Finally, extensive experiments are designed and conducted to evaluate its energy efficiency. Comparisons with the existing clustering and CS-based strategies have verified its effectiveness in improving energy efficiency and extending network lifespan. Deyu Lin, Weidong Min |
IEEE Internet Things J. | 2 |
| 2020 | A Survey on Energy-Efficient Strategies in Static Wireless Sensor NetworksabstractA comprehensive analysis on the energy-efficient strategy in static Wireless Sensor Networks (WSNs) that are not equipped with any energy harvesting modules is conducted in this article. First, a novel generic mathematical definition of Energy Efficiency (EE) is proposed, which takes the acquisition rate of valid data, the total energy consumption, and the network lifetime of WSNs into consideration simultaneously. To the best of our knowledge, this is the first time that the EE of WSNs is mathematically defined. The energy consumption characteristics of each individual sensor node and the whole network are expounded at length. Accordingly, the concepts concerning EE, namely the Energy-Efficient Means, the Energy-Efficient Tier, and the Energy-Efficient Perspective, are proposed. Subsequently, the relevant energy-efficient strategies proposed from 2002 to 2019 are tracked and reviewed. Specifically, they respectively are classified into five categories: the Energy-Efficient Media Access Control protocol, the Mobile Node Assistance Scheme, the Energy-Efficient Clustering Scheme, the Energy-Efficient Routing Scheme, and the Compressive Sensing--based Scheme. A detailed elaboration on both of the basic principle and the evolution of them is made. Finally, further analysis on the categories is made and the related conclusion is drawn. To be specific, the interdependence among them, the relationships between each of them, and the Energy-Efficient Means, the Energy-Efficient Tier, and the Energy-Efficient Perspective are analyzed in detail. In addition, the specific applicable scenarios for each of them and the relevant statistical analysis are detailed. The proportion and the number of citations for each category are illustrated by the statistical chart. In addition, the existing opportunities and challenges facing WSNs in the context of the new computing paradigm and the feasible direction concerning EE in the future are pointed out. Deyu Lin, Quan Wang 0006, Weidong Min, Zhiqiang Zhang 0001 |
ACM Trans. Sens. Networks | 3 |
| 2019 | A Large-Scale Repository of Deterministic Regular Expression Patterns and Its Applications
Haiming Chen 0001, Yeting Li, Chunmei Dong, Xinyu Chu, Xiaoying Mou, Weidong Min |
PAKDD (3) | 6 |
| 2019 | Real-time face recognition based on pre-identification and multi-scale classificationabstractIn face recognition, searching a person's face in the whole picture is generally too time‐consuming to ensure high‐detection accuracy. Objects similar to the human face or multi‐view faces in low‐resolution images may result in the failure of face recognition. To alleviate the above problems, a real‐time face recognition method based on pre‐identification and multi‐scale classification is proposed in this study. The face area is segmented based on the proportion of human faces in the pedestrian area to reduce the search range, and faces can be robustly detected in complicated scenarios such as heads moving frequently or with large angles. To accurately recognise small‐scale faces, the authors propose the multi‐scale and multi‐channel shallow convolution network, which combines a multi‐scale mechanism on the feature map with a multi‐channel convolution network for real‐time face recognition. It performs face matching only in the pre‐identified face areas instead of the whole image, therefore it is more efficient. Experimental results showed that the proposed real‐time face recognition method detects and recognises faces correctly, and outperforms the existing methods in terms of effectiveness and efficiency. Weidong Min, Mengdan Fan, Jing Li 0027 |
IET Comput. Vis. | 1 |
| 2019 | New approach to vehicle license plate location based on new model YOLO-L and plate pre-identificationabstractCurrently, the conventional license plate location method fails to detect the license plate under complex road environments such as severe weather conditions and viewpoint changes. Besides, it is difficult for license plate location method based on machine learning to precisely locate the area of license plate. Moreover, license plate location method may incorrectly detect similar objects such as billboards and road signs as license plates. To alleviate these problems, this article proposes a new approach to vehicle license plate location based on new model YOLO‐L and plate pre‐identification. The new model improves in two aspects to precisely locate the area of license plate. First, it uses k‐means++ clustering algorithm to select the best number and size of plate candidate boxes. Second, it modifies the structure and depth of YOLOv2 model. Plate pre‐identification algorithm can effectively distinguish license plates from similar objects. The experimental results show that authors’ proposed method not only achieves a precision of 98.86% and a recall of 98.86%, which outperforms the existing methods, but also has high efficiency in real time. Weidong Min, Qi Wang 0061, Qingpeng Zeng, Yanqiu Liao |
IET Image Process. | 1 |
| 2019 | SAR Image Retrieval Based on Unsupervised Domain Adaptation and ClusteringabstractEfficiently retrieving synthetic aperture radar (SAR) image is an important yet challenging task in the remote sensing field. Due to the shortage of labeled SAR images for fine-tuning convolutional neural network (CNN) models, this letter presents an unsupervised domain adaptation model based on CNN to learn the domain-invariant feature between SAR images and optical aerial images for SAR image retrieving, which can alleviate the burden of manual labeling. We extend a deep CNN to a novel adversarial network by adding the domain discriminator and the pseudolabel predictor. We improve the adaptation capacity of the adversarial network by utilizing the class information of SAR training images, which is obtained by clustering. Compared with the other related methods, the proposed method can enhance retrieval performance with our SAR data set. Famao Ye, Meng Dong, Hailin He, Weidong Min |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2019 | Human fall detection using normalized shape aspect ratio
Weidong Min, Song Zou, Jing Li 0027 |
Multim. Tools Appl. | 1 |
| 2018 | Support vector machine approach to fall recognition based on simplified expression of human skeleton action and fast detection of start key frame using torso angleabstractFalls sustained by subjects can have severe consequences, especially for elderly persons living alone. A fall detection method for indoor environments based on the Kinect sensor and analysis of three‐dimensional skeleton joints information is proposed. Compared with state‐of‐the‐art methods, the authors’ method provides two major improvements. First, possible fall activity is quantified and represented by a one‐dimensional float array with only 32 items, followed by fall recognition using a support vector machine (SVM). Unlike typical deep learning methods, the input parameters of their method are dramatically reduced. Hence, videos are trained and recognised by an SVM with a low time cost. Second, the torso angle is imported to detect the start key frame of a possible fall, which is much more efficient than using a sliding window. Their approach is evaluated on the telecommunication systems team (TST) fall detection dataset v2. The results show that their approach achieves an accuracy of 92.05%, better than other typical methods. According to the characters of machine learning, when more samples are imported, their method is expected to achieve a higher accuracy and stronger capability of fall‐like discrimination. It can be used in real‐time video surveillance because of its time efficiency and robustness. Weidong Min, Leiyue Yao, Zhenrong Lin, Li Liu 0010 |
IET Comput. Vis. | 1 |
| 2018 | Remote Sensing Image Registration Using Convolutional Neural Network FeaturesabstractSuccessful remote sensing image registration is an important step for many remote sensing applications. The scale-invariant feature transform (SIFT) is a well-known method for remote sensing image registration, with many variants of SIFT proposed. However, it only uses local low-level information, and loses much middle- or high-level information to register. Image features extracted by a convolutional neural network (CNN) have achieved the state-of-the-art performance for image classification and retrieval problems, and can provide much middle- and high-level information for remote sensing image registration. Hence, in this letter, we investigate how to calculate the CNN feature, and study the way to fuse SIFT and CNN features for remote sensing image registration. The experimental results demonstrate that the proposed method yields a better registration performance in terms of both the aligning accuracy and the number of correct correspondences. Famao Ye, Yanfei Su, Xuqing Zhao, Weidong Min |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Remote Sensing Image Retrieval Using Convolutional Neural Network Features and Weighted DistanceabstractRemote sensing image retrieval (RSIR) is a fundamental task in remote sensing. Most content-based RSIR approaches take a simple distance as similarity criteria. A retrieval method based on weighted distance and basic features of convolutional neural network (CNN) is proposed in this letter. The method contains two stages. First, in offline stage, the pretrained CNN is fine-tuned by some labeled images from the target data set, then used to extract CNN features, and labeled the images in the retrieval data set. Second, in online stage, we use the fine-tuned CNN model to extract the CNN feature of the query image and calculate the weight of each image class and apply them to calculate the distance between the query image and the retrieved images. Experiments are conducted on two RSIR data sets. Compared with the state-of-the-art methods, the proposed method is simplified but efficient, significantly improving retrieval performance. Famao Ye, Xuqing Zhao, Meng Dong, Weidong Min |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2018 | Recognition of pedestrian activity based on dropped-object detection
Weidong Min, Jing Li 0027, Shaoping Xu |
Signal Process. | 1 |
| 2018 | A New Approach to Track Multiple Vehicles With the Combination of Robust Detection and Two ClassifiersabstractIt plays an important role to accurately track multiple vehicles in intelligent transportation, especially in intelligent vehicles. Due to complicated traffic environments it is difficult to track multiple vehicles accurately and robustly, especially when there are occlusions among vehicles. To alleviate these problems, a new approach is proposed to track multiple vehicles with the combination of robust detection and two classifiers. An improved ViBe algorithm is proposed for robust and accurate detection of multiple vehicles. It uses the gray-scale spatial information to build dictionary of pixel life length to make ghost shadows and object's residual shadows quickly blended into the samples of the background. The improved algorithm takes good post-processing method to restrain dynamic noise. In this paper, we also design a method using two classifiers to further attack the problem of failure to track vehicles with occlusions and interference. It classifies tracking rectangles with confidence values between two thresholds through combining local binary pattern with support vector machine (SVM) classifier and then using a convolutional neural network (CNN) classifier for the second time to remove the interference areas between vehicles and other moving objects. The two classifiers method has both time efficiency advantage of SVM and high accuracy advantage of CNN. Comparing with several existing methods, the qualitative and quantitative analysis of our experiment results showed that the proposed method not only effectively removed the ghost shadows, and improved the detection accuracy and real-time performance, but also was robust to deal with the occlusion of multiple vehicles in various traffic scenes. Weidong Min, Mengdan Fan, Xiaoguang Guo |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | A New Kinect Approach to Judge Unhealthy Sitting Posture Based on Neck Angle and Torso Angle
Leiyue Yao, Weidong Min |
ICIG (1) | 2 |
| 2013 | Network Management Based on Domain Partition for Mobile Agents
Weidong Min |
IDEAL | 2 |
| 2013 | A matrix grammar approach for automatic distributed network resource management
Weidong Min, Yongzhen Ke |
Frontiers Comput. Sci. | 1 |
| 1996 | Automatic mesh generation for multiply connected planar regions based on mesh grading propagation
Weidong Min, Zesheng Tang, Minzhi Wang |
Comput. Aided Des. | 1 |
| 1995 | A new approach to fully automatic mesh generation
Weidong Min, Zesheng Tang, Minzhi Wang |
J. Comput. Sci. Technol. | 1 |
| 1993 | Recognition of dimensions in engineering drawings based on arrowheadabstractAn algorithm for recognizing and understanding dimensions in engineering drawings is presented. The dimensions are classified into 27 patterns and 48 subpatterns. The concept and algorithm of arrowhead-match is presented. A DFA is used to detect and recognize dimensions and rectify their rationality.> Weidong Min, Zesheng Tang |
ICDAR | 1 |
| 1993 | Using web grammar to recognize dimensions in engineering drawings
Weidong Min, Zesheng Tang |
Pattern Recognit. | 1 |