Ming Ma 0006

dblp:20/1028-6 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-2060-2266ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ICMF-Net: Interactive Cross-Modal Fusion with Attention and Selection Network for Remote Sensing Object Detection
abstract
Remote sensing object detection plays a critical role in applications such as ecological monitoring, mining management, and urban planning, yet remains challenging due to significant scale variations, densely distributed small objects, and complex background interference. Previous works primarily rely on RGB-only or multispectral imagery, often neglecting the DSM data that encodes valuable height information. To address these limitations, we propose ICMF-Net, one of the first studies to leverage both RGB imagery and Digital Surface Model (DSM) data for remote sensing object detection. ICMF-Net consists of three key components: CMASF (Cross-Modal Attention and Selection Fusion), which enhances modality complementarity via adaptive cross-modal attention, feature selection, and height-aware spatial enhancement; FEMR (Feature Extraction and Multi-scale Refinement), which strengthens hierarchical feature representation through multi-scale attention and redundancy-aware convolution; and SPD-Conv (Space-to-Depth Convolution), which replaces standard strided convolutions in the backbone and neck to enable information-preserving downsampling and improve sensitivity to small and densely distributed targets. Extensive experiments on the ISPRS Potsdam and MMOD benchmarks demonstrate that ICMF-Net improves mAP50: 95 by 5.2% and 4.3% over the baseline, respectively, establishing it as a robust and effective RGB–DSM remote sensing object detection framework.
Ming Ma 0006
ICMR4
2025 RS-YOLO: An Optimized YOLOv8 Framework for Mining Area Traffic Safety Monitoring
abstract
The driving conditions in mining regions are intricate, marked by limited visibility and significant disturbances. Current object detection models often encounter high rates of missed detections and false positives when applied to monitoring driver behavior and road conditions for safety. To tackle these challenges, we introduce RS-YOLO (Residual CSP YOLO), an enhanced YOLO framework. Firstly, we develop a novel lightweight feature extraction module that employs expanded feature maps and multi-layer convolutions within an inverted residual architecture. This approach enhances feature extraction efficiency and reduces both missed and false detections. Secondly, we incorporate a pre-activation mechanism to optimize the network, which improves gradient flow, ensures training stability, and accelerates convergence speed. Finally, we integrate the PIOU loss function to more accurately address the alignment errors between predicted and ground truth bounding boxes in complex scenarios. Experimental results demonstrate that RS-YOLO achieves relative improvements of 3.59% compared to YOLOv8n and 1.63% over YOLOv8s on the mining dataset, while also reducing the number of parameters and computational complexity by 50% relative to YOLOv8s. Additionally, on the COCO dataset, RS-YOLO achieves a 27% accuracy improvement over YOLOv8n. These outcomes exhibit RS-YOLO's superior detection capabilities, lower computational demands, and strong generalization performance in challenging environments.
Jingyi Jia, Ming Ma 0006
CSCWD3
2025 DSDN-Net: An Effective Network for Semantic Segmentation in Open-Pit Coal Mining Areas for Land Cover Recognition
abstract
Currently, existing methods in land cover recognition in open-pit coal mining areas face the issue of insufficient accuracy due to multiscale and blurred boundaries when processing remote sensing images. This paper introduces a remote sensing image semantic segmentation network, DSDN-Net, to tackle the issues. DSDN-Net adopts MobileNetV2 as the backbone, and a spatial pyramid pooling structure, Dynamic Snake Dense-ASPP (DSDN-ASPP), which enhances the model’s receptive field based on cross-layer connections is designed to increase the model’s receptive field by introducing dynamic snake convolution and utilizing depthwise convolutions with different kernel sizes, allowing the model to focus more on spatial features in remote sensing images. To handle blurred boundaries in remote sensing image, the decoder of DSDN-Net incorporates a Convolutional Block Attention Module (CBAM) to enhance precision. A dataset containing 5440 remote sensing images from the open-pit coal mining area is constructed using high-resolution remote sensing image. Experimental results on the open-pit coal mining area dataset demonstrate that the proposed DSDN-Net outperforms existing methods in multiple performance metrics.
Ming Ma 0006
ICASSP2
2025 Interactive and Balanced Multimodal Learning via Cross Attention and Gradient Modulation for Compressed Video Action Recognition
abstract
Compressed video action recognition is a crucial task in video processing. Compared with traditional methods, it directly processes RGB (I-frames) and motion (motion vectors and residuals) modalities, which effectively alleviates computational burdens. However, this task suffers from dynamic noise and insufficient interaction between modalities. To address these issues, we propose the Cross-Modal Fusion Modulation Network (CFM-Net), which consists of an RGB stream and a motion stream. The motion stream incorporates a Pseudo Optical Flow Generator (POFG) that reduces motion noise through adversarial learning and optical flow supervision. The interaction between streams is strengthened by the Lightweight Cross-Modal Fusion (LCF) block and the Adaptive Dynamic Gradient Modulation (ADGM) strategy. The LCF facilitates the fusion of RGB and motion modalities through cross-attention, while the ADGM enhances their interaction by balancing multimodal optimization. We evaluate the performance of CFM-Net on the UCF-101 and HMDB-51 benchmarks, demonstrating the effectiveness and accuracy of our approach.
Shaojie Li 0001, Ming Ma 0006
ICASSP3
2025 YOLO-MAOD: An Algorithm for Ground Object Detection in Open-Pit Mine Based on Remote Sensing Images
Ming Ma 0006
ICIC (21)4
2025 YOLO-DYN: An Improved YOLOv8 Model for Multi-Scale Object Detection
abstract
In mine transportation, road and driver behavior detection face challenges in multi-scale and small target recognition. Conventional feature fusion strategies often fail to adapt to the complex, unstructured nature of mining environments, while standard convolution and pooling tend to cause the loss of critical small-target features. To address this, we propose YOLO-DYN, a model designed to simultaneously detect the driver’s operational context and the surrounding road environment. It introduces Adapt-C2f, an adaptive module combining deformable convolution—enhanced by a tailored multi-path coordinate attention mechanism—with a dilated convolution block, effectively preserving key features. A dynamic fusion strategy further adjusts the convolution kernel shape and orientation while integrating multi-level spatial information. Experiments on the public PASCAL VOC dataset and our constructed mining-area dataset show that YOLO-DYN improves mean Average Precision (mAP) by 6.1% and 2.2%, respectively.
Ming Ma 0006
IJCNN4
2025 DC-SNet: Efficient Spatio-Temporal Prediction Network Based On Dual-Domain Collaborative Spatiotemporal Network
abstract
As the most critical safety hazard in open-pit mining, slope deformation requires effective and accurate detection, which holds significant importance for the safety and ecological protection of the mining area. However, as the demand for real-time and precise prediction grows more acute, traditional models relying on single-domain features face significant challenges. Moreover, due to the intrinsic spatial correlations among monitoring points, the prediction accuracy of models incapable of capturing dynamic spatiotemporal features is highly likely to be constrained.To better address these challenges, we propose a novel network framework, the Efficient Spatio-Temporal Prediction Network Based on Dual-Domain Collaborative Spatiotemporal Network (DC-SNet), which innovatively integrates feature extraction mechanisms from both time and frequency domains. We adopt a dual-domain collaborative architecture that enables complementary integration of features from the time and frequency domains, and employ the SpaceT Block module -- specifically designed to extract deeper sequential features from radar data of open-pit mine slopes -- to further enhance the spatial correlations among monitoring points. Furthermore, we introduce a Full Embedding module to effectively capture the dynamic spatiotemporal features within slope displacement time series. Experimental results using open-pit slope radar data and public datasets demonstrate that our model significantly outperforms current mainstream time series prediction models on both datasets.
Ming Ma 0006
ICMR3
2025 MFCA-UNet:A Multi-source Feature Fusion Cross-Attention Enhancement Network for 12 Lead Electrocardiogram Classification
abstract
The intelligent monitoring and classification of electrocardiogram (ECG) signals plays a crucial role in the early diagnosis of cardiovascular diseases. Despite significant progress in deep learning-based ECG analysis, existing models still struggle to effectively capture key information from critical signal regions while neglecting the impact of patient-specific attributes on ECG signals, thereby limiting their accuracy and generalization. To address these issues, we propose MFCA-UNet, a multi-source feature fusion cross-attention enhancement network. Specifically, we integrate a multi-scale feature attention module into the UNet backbone to improve the model’s sensitivity to crucial features. Meanwhile, the lead fusion module is optimized to facilitate effective interaction between local and global information across different leads. Additionally, we design a cross-fusion encoder to represent and extract patient-specific demographic features through feature mapping, which are then cross-fused with ECG signal data. Extensive experiments on the public datasets PTB-XL and Chapman show that the proposed method performs better than current advanced electrocardiogram classification models.
Ming Ma 0006, Meiju Yu
SMC2
2024 YOLO-BS: A Better Object Detection Model for Real-Time Driver Behavior Detection
Yang Xi, Jinxin Guo, Ming Ma 0006
ICIC (12)3
2024 PP-YOLO-CS: A Novel Approach for Real-Time Fire and Smoke Detection in Industrial Park Environments
abstract
Real-time detection of fire and smoke in Industrial park environments is crucial for ensuring safety. Current models face significant challenges in accurately detecting fire and smoke in scenarios with small-sized targets and substantial shape variations, and there is an urgent need to enhance the detection capabilities of the models. Addressing these challenges, this paper proposes the novel PP-YOLO-CS model, based on the PP-YOLOE, integrated CA-ConvNeXt, SP-ASPP, and SP-CSP modules. The model introduces CA-ConvNeXt, which integrates Channel Attention with ConvNeXt, focusing on critical information channels. This improves feature extraction efficiency and accuracy in complex industrial park fire scenarios and addresses the shortcomings of current models in handling complex and diverse data. We propose the Split-based Convolution with Atrous Spatial Pyramid Pooling Structure (SP-ASPP) that captures multi-scale information through varied dilation rates and channel-splitting techniques, effectively addressing target scale and shape variation challenges, enhancing detection accuracy across diverse sizes while maintaining computational efficiency. We present the Split-based Convolution with Cross Stage Partial (SP-CSP) framework, a distinctive approach that excels in precision. The framework integratesthe Atrous Spatial Pyramid Pooling (ASPP) module and Split-based Convolution, effectively enlarging the receptive field to capture finer details, especially in low-resolution images, while simplifying the feature fusion process, reducing information redundancy, and improving the accuracy of the system. To facilitate the development of a fire smoke detection model for industrial park environments, we establish the IndustrialParkData dataset, comprising 7141 images depicting fire and smoke in factory environments. Experimental results demonstrate that CA-ConvNeXt and SP-ASPP improves the accuracy of the PP-YOLOE model by 1.73%, with a simultaneous speed increase of 6.39FPS. Furthermore, CA-ConvNeXt and SP-CSP maintains the same prediction speed as PP-YOLOE while increasing the Average Precision (mAPs) by 2.05%. These results demonstrate the strong performance of the PP-YOLO-CS model in detecting fire and smoke.
Ming Ma 0006, Mengkun Guo
IJCNN1
2023 SDGC-YOLOv5: A More Accurate Model for Small Object Detection
Zhiming Lu, Shaojie Li 0001, Ming Ma 0006
ICANN (7)4
2023 MTFD: Multi-Teacher Fusion Distillation for Compressed Video Action Recognition
abstract
As an important work in computer vision, some recent representative works such as Two-stream networks, 3D ConvNets, and Transformer-based networks have achieved outstanding performance. However, due to the high computational cost, the explosion of computation time and parameters, they cannot meet the needs of real-time applications. The current work utilizes the keyframes, residual and motion information retained by compressed video for computation, which greatly reduces the computational effort but still cannot satisfy real-time applications. Therefore, we propose a multi-teacher fusion distillation framework for compressed video action recognition (MTFD). Unlike the traditional method of transferring the knowledge of single or multiple teachers directly into the student model, we also perform knowledge transfer between teachers. MTFD achieves better knowledge distillation through mutual guidance and information fusion between teachers. Furthermore, we improve the network’s ability to extract motion information, which ultimately reduces the computational effort while maintaining high accuracy.
Jinxin Guo, Shaojie Li 0001, Ming Ma 0006
ICASSP5
2023 META: Motion Excitation With Temporal Attention for Compressed Video Action Recognition
abstract
Compressed video action recognition has gained significant attention recently due to its ability to replace the raw video with I-frames and compressed motion clues, such as motion vectors and residuals. This results in substantial reductions in storage and computation costs. However, this task suffers from coarse and a lack of structures that can capture long-range spatiotemporal dependencies. To address these issues, this paper proposes a novel module called Motion Excitation with Temporal Attention (META) and utilizes network structures that can capture long-range dependencies. The META module stimulates motion information between I-frames and enhances the motion representation of the motion vectors. It first assigns different weights to feature-level frames, and then calculates the feature-level temporal differences from spatiotemporal features. Finally, it utilizes these differences to excite the motion-sensitive channels of the features. In addition, for compressed video action recognition, we have introduced a new network structure that combines CNN and Transformer. It can seamlessly integrate the merits of convolution and self-attention in a concise transformer format. The proposed method is evaluated on the challenging HMDB-51 and UCF-101 datasets. The extensive comparison results and ablation studies demonstrate the effectiveness and strength of the proposed method.
Shaojie Li 0001, Jinxin Guo, Ming Ma 0006
ICPADS4
2023 LAE-Net: Light and Efficient Network for Compressed Video Action Recognition
Jinxin Guo, Ming Ma 0006
MMM (2)4
2023 IFF-Net: I-Frame Fusion Network for Compressed Video Action Recognition
abstract
Compressed video action recognition has received significant attention due to its potential for reducing storage and computational costs. However, the current methods typically only capture a few RGBs and compressed motion cues (e.g., motion vectors and residuals), which are insufficient for modeling actions at their full temporal extent. To address this issue, we propose a Time Domain Fusion (TDF) Module that can extract both low-frequency and high-frequency components from the video and integrate them seamlessly, resulting in the effective integration of abundant motion information into a single frame. More importantly, by using the TDF module, we introduced a new network called I-Frame Fusion Network (IFF-Net). The IFF -Net interacts with the original network (I-frame, motion vector, and residual) in two ways: explicit and implicit. Explicit interaction involves extracting the new representation and the original compressed representation information separately and then performing a later fusion. In contrast, implicit interaction uses the distillation approach, with the IFF-Net acting as the teacher to guide the I-frame network to learn full temporal expressions. Our approach performs better than state-of-the-art methods on the UCF-101 and HMDB-51 datasets for compressed video action recognition.
Shaojie Li 0001, Jinxin Guo, Ming Ma 0006
SMC5
2022 Multi-Knowledge Attention Transfer Framework for Action Recognition
Jinxin Guo, Ming Ma 0006
ICANN (1)3
2019 Effective moving object detection in H.264/AVC compressed domain for video surveillance
Ming Ma 0006, Houbing Song
Multim. Tools Appl.1
2017 Distribution of primary additional errors in fractal encoding method
Shuai Liu 0002, Weina Fu, Liqiang He, Jiantao Zhou 0002, Ming Ma 0006
Multim. Tools Appl.5
2016 Differential trajectory tracking with automatic learning of background reconstruction
Weina Fu, Jiantao Zhou 0002, Shuai Liu 0002, Ming Ma 0006
Multim. Tools Appl.4
2016 A fractal image encoding method based on statistical loss used in agricultural image compression
Shuai Liu 0002, Lingyun Qi, Ming Ma 0006
Multim. Tools Appl.4