Mingxin Yu

dblp:177/2318 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RTDM: Real-time denoising mamba with progressive self-distillation
Yuchen Bai 0003, Mingxin Yu, Lidan Lu, Xiaoping Lou, Mingli Dong, Zidong Wang 0001, Lianqing Zhu
Knowl. Based Syst.2
2025 ABSOD: Attention based small object detector for outfalls inspection in aerial images
abstract
Strengthening the inspection of outfalls into rivers and oceans can help monitor pollutant emissions to the natural environment. Unmanned aerial vehicle (UAV) with high spatial resolution imagery has become a more efficient method for outfall surveys. At present, outfalls retrieval from UAV images relies on visual interpretation by skilled experts. However, long periods of concentration on detecting outfalls in high-resolution images for an expert easily increase mental load and stress, resulting in missing and false detection. Therefore, we develop a deep learning model, called Attention Based Small Object Detector (ABSOD), to perform outfalls detection in aerial images. In this model, an adaptive spatial correlation pyramid attention (ASCPA) network is proposed to establish long-distance region-to-region relationships between the outfall and its surrounding information more effectively. This network is mainly composed of SPE (Spatial Pyramid Extractor) and SCFM (Spatial Correlation Fusion Module). The purpose of the SPE is to extract multi-scale spatial information on the feature map. The SCFM is used to perform spatial correlation feature recalibration to selectively emphasized informative features. Experimental results show that the proposed network outperforms the state-of-the-art small object detection model in detecting outfalls, and reaches 45.9%, 92.8%, 86.5% and 34.4% in the four metrics of Precision, Recall, AP 0.5 , and AP 0.5:0.95 , respectively. To show the superiority of the ASCPA network, we compared our results with other attention mechanisms, all of them show that the ASCPA network has a competitive performance for outfalls detection. Moreover, based on visualization analysis, the ASCPA network is able to pay more attention on true outfall objects with respect to other attention mechanisms. These promising results demonstrate that the deep learning algorithm can be a feasible solution to assist experts in detecting outfalls with UAV imagery. The model and code are available at https://github.com/ISCLab-Bistu/ASCPA-Attention .
Zhenjia Li, Shengjun Liang, Mingxin Yu
Intell. Data Anal.3
2025 BotLGT: Social bot detection based on LLM and graph transformer
Liyan Shen, Qinglei Guo, Mingxin Yu
Neurocomputing7
2025 Attention-enhanced controllable disentanglement for cloth-changing person re-identification
Yiyuan Ge, Mingxin Yu, Zhihao Chen 0014, Wenshuai Lu, Yuxiang Dai, Huiyu Shi
Vis. Comput.2
2024 Efficient Motion Planning for Manipulators with Control Barrier Function-Induced Neural Controller
abstract
Sampling-based motion planning methods for manipulators in crowded environments often suffer from expensive collision checking and high sampling complexity, which make them difficult to use in real time. To address this issue, we propose a new generalizable control barrier function (CBF)based steering controller to reduce the number of samples needed in a sampling-based motion planner RRT. Our method combines the strength of CBF for real-time collision-avoidance control and RRT for long-horizon motion planning, by using CBF-induced neural controller (CBF-INC) to generate control signals that steer the system towards sampled configurations by RRT. CBF-INC is learned as Neural Networks and has two variants handling different inputs, respectively: state (signed distance) input and point-cloud input from LiDAR. In the latter case, we also study two different settings: fully and partially observed environmental information. Compared to manually crafted CBF which suffers from over-approximating robot geometry, CBF-INC can balance safety and goal-reaching better without being over-conservative. Given state-based input, our neural CBF-induced neural controller-enhanced RRT (CBFINC-RRT) can increase the success rate by 14% while reducing the number of nodes explored by 30%, compared with vanilla RRT on hard test cases. Given LiDAR input where vanilla RRT is not directly applicable, we demonstrate that our CBF-INCRRT can improve the success rate by 10%, compared with planning with other steering controllers. Our project page with supplementary material is at https://mit-realm.github.io/CBFINC-RRT-website/.
Mingxin Yu, Chenning Yu, M.-Mahdi Naddaf-Sh, Devesh Upadhyay, Sicun Gao, Chuchu Fan
ICRA1
2024 Multiple-local feature and attention fused person re-identification method
abstract
Person re-identification (ReID) is widely used in intelligent security, monitoring, criminal investigation and other fields. Aiming at the problems of local occlusion, scale misalignment and attitude change of pedestrian images in actual scenes, we propose a Multi-local Feature and Attention fused network (MFA) used for person re-identification task. Firstly, Channel Point Affinity Attention module (CPAA) is embedded in the backbone network to enhance the ability of the network for extracting local details. The feature map output from the backbone network is horizontally segmented into four local feature maps, and further four branch networks are concatenated to the feature map of the backbone network. The four local feature maps are used to guide the four branch networks to pay more attention on different areas of pedestrians through Global Local Aligned loss (GLA) function. Finally, the pedestrian feature vector containing multi-local features is obtained. The mAP of the network on Market-1501, DukeMTMC-reID,CUHK03 and MSMT17 datasets were 88.6%, 81.4%, 79.5% and 64.7%, and the Rank-1 was 95.8%, 90.1%, 81.2% and 84.1% respectively. In addition, the model also obtained 73.2% and 68.1% of Rank-1 on partial dataset Patial-REID and Patial-iLIDS, respectively. Recently, The MFA model parameter is 28.3M and the inference efficiency is approximately 32 fps to an image with a resulation of 256 × 128. Compared with other ReID methods, our proposed methods achieved a competitive performance for ReID task. The code was available at github:[email protected]:ISCLab-Bistu/MFA.git.
Mingxin Yu, Rui You, Xinglong Ji, Wenshuai Lu
Intell. Data Anal.1
2024 A novel dual-granularity lightweight transformer for vision tasks
abstract
Transformer-based networks have revolutionized visual tasks with their continuous innovation, leading to significant progress. However, the widespread adoption of Vision Transformers (ViT) is limited due to their high computational and parameter requirements, making them less feasible for resource-constrained mobile and edge computing devices. Moreover, existing lightweight ViTs exhibit limitations in capturing different granular features, extracting local features efficiently, and incorporating the inductive bias inherent in convolutional neural networks. These limitations somewhat impact the overall performance. To address these limitations, we propose an efficient ViT called Dual-Granularity Former (DGFormer). DGFormer mitigates these limitations by introducing two innovative modules: Dual-Granularity Attention (DG Attention) and Efficient Feed-Forward Network (Efficient FFN). In our experiments, on the image recognition task of ImageNet, DGFormer surpasses lightweight models such as PVTv2-B0 and Swin Transformer by 2.3% in terms of Top1 accuracy. On the object detection task of COCO, under RetinaNet detection framework, DGFormer outperforms PVTv2-B0 and Swin Transformer with increase of 0.5% and 2.4% in average precision (AP), respectively. Similarly, under Mask R-CNN detection framework, DGFormer exhibits improvement of 0.4% and 1.8% in AP compared to PVTv2-B0 and Swin Transformer, respectively. On the semantic segmentation task on the ADE20K, DGFormer achieves a substantial improvement of 2.0% and 2.5% in mean Intersection over Union (mIoU) over PVTv2-B0 and Swin Transformer, respectively. The code is open-source and available at: https://github.com/ISCLab-Bistu/DGFormer.git.
Mingxin Yu, Wenshuai Lu, Yuxiang Dai, Huiyu Shi, Rui You
Intell. Data Anal.2
2024 MambaTSR: You only need 90k parameters for traffic sign recognition
Yiyuan Ge, Zhihao Chen 0014, Mingxin Yu, Qing Yue, Rui You, Lianqing Zhu
Neurocomputing3
2023 A lightweight vision transformer with symmetric modules for vision tasks
abstract
Transformer-based networks have demonstrated their powerful performance in various vision tasks. However, these transformer-based networks are heavyweight and cannot be applied to edge computing (mobile) devices. Despite that the lightweight transformer network has emerged, several problems remain, i.e., weak feature extraction ability, feature redundancy, and lack of convolutional inductive bias. To address these three problems, we propose a lightweight visual transformer (Symmetric Former, SFormer), which contains two novel modules (Symmetric Block and Symmetric FFN). Specifically, we design Symmetric Block to expand feature capacity inside the module and enhance the long-range modeling capability of attention mechanism. To increase the compactness of the model and introduce inductive bias, we introduce convolutional cheap operations to design Symmetric FFN. We compared the SFormer with existing lightweight transformers on several vision tasks. Remarkably, on the image recognition task of ImageNet [13], SFormer gains 1.2% and 1.6% accuracy improvements compared to PVTv2-b0 and Swin Transformer, respectively. On the semantic segmentation task of ADE20K [64], SFormer delivers performance improvements of 0.2% and 0.7% compared to PVTv2-b0 and Swin Transformer, respectively. On the cityscapes dataset [11], SFormer delivers performance improvements of 2.5% and 4.2% compared to PVTv2-b0 and Swin Transformer, respectively. The code is open-source and available at: https://github.com/ISCLab-Bistu/Symmetric_Former.git.
Shengjun Liang, Mingxin Yu, Wenshuai Lu, Xinglong Ji, Xiongxin Tang, Rui You
Intell. Data Anal.2
2022 OOD-CV: A Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images
Bingchen Zhao, Shaozuo Yu, Wufei Ma, Mingxin Yu, Shenxiao Mei, Angtian Wang, Ju He, Alan L. Yuille, Adam Kortylewski
ECCV (8)4
2022 Rotation-aware dynamic temporal consistency with spatial sparsity correlation tracking
Mingxin Yu, Changlong Wang 0002, Yuhua Zhang, Zhilong Lin
Image Vis. Comput.1
2020 EEG-based tonic cold pain assessment using extreme learning machine
abstract
The purpose of this study is to present a novel method which can objectively identify the subjective perception of tonic pain. To achieve this goal, scalp EEG data are recorded from 16 subjects under the cold stimuli condition. The proposed method is capable of classifying four classes of tonic pai n states, which include No pain, Minor Pain, Moderate Pain, and Severe Pain. Due to multi-class problem of our research an extended Common Spatial Pattern (ECSP) method is first proposed for accurately extracting features of tonic pain from captured EEG data. Then, a single-hidden-layer feedforward network is used as a classifier for pain identification. With the aid of extreme learning machine (ELM) algorithm, the classifier is trained here. The advantages of ELM-based classifier can obtain an optimal and generalized solution for multi-class tonic cold pain. Experimental results demonstrate that the proposed method discriminates the tonic pain successfully. Additionally, to show the superiority for the ELM-based classifier, compared results with the well-known support vector machine (SVM) method show the ELM-based classifier outperform than the SVM-based classifier. These findings may pay the way for providing a direct and objective measure of the subjective perception of tonic pain.
Mingxin Yu, Yingzi Lin, Lianqing Zhu, Guangkai Sun, Yikang Guo
Intell. Data Anal.1
2020 Diverse frequency band-based convolutional neural networks for tonic cold pain assessment using EEG
Mingxin Yu, Bofei Zhu, Lianqing Zhu, Yingzi Lin, Yikang Guo, Guangkai Sun, Mingli Dong
Neurocomputing1
2018 An eye detection method based on convolutional neural networks and support vector machines
abstract
Eye detection plays an important role in many fields, because eyes provide prominent facial feature information. However, changes in face pose, illumination variation, with glasses, and eye occlusions can make it difficult to detect eyes well from facial images. This paper proposes a hybrid model f or eye detection. The model is an integration of two classifiers: Convolutional Neural Networks (CNN) and Support Vector Machines (SVM). In order to improve the speed of detection in the system, an eye variance filter (EVF) is constructed for eliminating most of noneye images to keep less candidate eye images. The CNN then works as a trainable feature extractor to explicitly extract various latent eye features. Finally, the trained SVM classifier is employed for eye verification instead of using the CNN classification function. Experiments applying the model have been conducted on the BioID, IMM, FERET and ORL face databases. Comparisons with other methods on the same databases indicate that this hybrid model has achieved a higher detection accuracy. Extensive experiments demonstrate the robustness and efficiency of our method by testing it on different facial images with varying eye conditions.
Mingxin Yu, Yingzi Lin, David Schmidt 0002, Xiangzhou Wang, Yikang Guo, Bo Liang 0011
Intell. Data Anal.1
2018 Diesel engine modeling based on recurrent neural networks for a hardware-in-the-loop simulation system of diesel generator sets
Mingxin Yu, Xiaoying Tang 0003, Yingzi Lin, Xiangzhou Wang
Neurocomputing1
2016 A spatial-temporal trajectory clustering algorithm for eye fixations identification
abstract
Eye movements mainly consist of fixations and saccades. The identification of eye fixations plays an important role in the process of eye-movement data research. At present, there is no standard method for identifying eye fixations. In this paper, eye movements are regarded as spatial-temporal traj ectories. Hence, we present a spatial-temporal trajectory clustering algorithm for eye fixations identification. The main idea of the algorithm is based on Density-Based Spatial Clustering Algorithm with Noise (DBSCAN), which is commonly used in spatial clustering data. In order to apply DBSCAN to our spatial-temporal clustering data, we modified its original concept and algorithm. In addition, the optimum dispersion threshold (Eps) is derived automatically from the data sets with the aid of the `gap statistic' theory. Using the confusion matrix measurement method, we compared the classification results obtained by our algorithm with four other expert algorithms for eye fixations identification show the proposed algorithm demonstrated an equal or better performance. Also, the robustness of our algorithm to additional noise in Points of Gaze (PoGs) data and changes in sampling rate has been verified.
Mingxin Yu, Yingzi Lin, Jeffrey Breugelmans, Xiangzhou Wang, Guanglai Gao
Intell. Data Anal.1