Maoxun Yuan

dblp:305/3385 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-7463-7328ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Image recognition and object detection · 59% Segmentation and scene understanding · 21% Vision and language · 14%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
infrared object detection
1.012026
Knowledge-Guided Adversarial Training for Infrared Object Detection via Thermal Radiation Modeling · Int. J. Comput. Vis. 2026
Computer vision › Image recognition and object detection › object detection
multimodal object detection
0.912025
Rethinking Multi-Modal Object Detection From the Perspective of Mono-Modality Feature Learning · ICCV 2025
Computer vision › Vision and language
cross-modal alignment
0.612022
Translation, Scale and Rotation: Cross-Modal Alignment Meets RGB-Infrared Vehicle Detection · ECCV (9) 2022
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter tuning
0.312025
UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

thermal radiation modeling · 1.0adversarial training · 1.0unified framework · 0.9multimodal feature learning · 0.9adapter tuning · 0.9cross-modal alignment · 0.6
YearPublicationVenuePosition
2026 Knowledge-Guided Adversarial Training for Infrared Object Detection via Thermal Radiation Modeling
Shukun Xiong, Maoxun Yuan, Ranjie Duan, Qing Guo 0005, Haibin Duan, Xingxing Wei 0001
Int. J. Comput. Vis.3
2026 Removal Then Selection: A Coarse-to-Fine Fusion Perspective for RGB-Infrared Object Detection
abstract
In recent years, object detection utilizing both visible (RGB) and thermal infrared (IR) imagery has garnered extensive attention and has been widely implemented across a diverse array of fields. By leveraging the complementary properties between RGB and IR images, the object detection task can achieve reliable and robust object localization across a variety of lighting conditions, from daytime to nighttime environments. While RGB-IR multi-modal data input generally enhances overall detection performance, most existing multi-modal object detection methods fail to fully exploit the complementary potential of these two modalities. We believe that this issue arises not only from the challenges associated with effectively integrating multi-modal information but also from the presence of redundant features in both the RGB and IR modalities. The redundant information of each modality will exacerbate the fusion imprecision problems during propagation. To address this issue, we draw inspiration from the human cognitive mechanisms for processing multi-modal information and propose a novel coarse-to-fine perspective to purify and fuse features from both modalities. Specifically, following this perspective, we design a Redundant Spectrum Removal module to remove interfering information within each modality coarsely and a Dynamic Feature Selection module to finely select the desired features for feature fusion. To verify the effectiveness of the coarse-to-fine fusion strategy, we construct a new object detector called the Removal then Selection Detector (RSDet). Extensive experiments on five RGB-IR object detection datasets verify the superior performance of our method. The source code and results are available athttps://github.com/Zhao-Tian-yi/RSDet.git
Tianyi Zhao 0003, Maoxun Yuan, Feng Jiang 0014, Nan Wang 0014, Xingxing Wei 0001
IEEE Trans. Intell. Transp. Syst.2
2025 Rethinking Multi-Modal Object Detection From the Perspective of Mono-Modality Feature Learning
Tianyi Zhao 0003, Yanglei Gao, Maoxun Yuan, Xingxing Wei 0001
ICCV5
2025 UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning
Maoxun Yuan, Tianyi Zhao 0003, Shan Fu, Xue Yang 0005, Xingxing Wei 0001
ACM Multimedia1
2024 SFDFusion: An Efficient Spatial-Frequency Domain Fusion Network for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to utilize the complementary information from two modalities to generate fused images with prominent targets and rich texture details. Most existing algorithms only perform pixel-level or feature-level fusion from different modalities in the spatial domain. They usually overlook the information in the frequency domain, and some of them suffer from inefficiency due to excessively complex structures. To tackle these challenges, this paper proposes an efficient Spatial-Frequency Domain Fusion (SFDFusion) network for infrared and visible image fusion. First, we propose a Dual-Modality Refinement Module (DMRM) to extract complementary information. This module extracts useful information from both the infrared and visible modalities in the spatial domain and enhances fine-grained spatial details. Next, to introduce frequency domain information, we construct a Frequency Domain Fusion Module (FDFM) that transforms the spatial domain to the frequency domain through Fast Fourier Transform (FFT) and then integrates frequency domain information. Additionally, we design a frequency domain fusion loss to provide guidance for the fusion process. Extensive experiments on public datasets demonstrate that our method produces fused images with significant advantages in various fusion metrics and visual effects. Furthermore, our method demonstrates high efficiency in image fusion and good performance on downstream detection tasks, thereby satisfying the real-time demands of advanced visual tasks. The code is available at https://github.com/lqz2/SFDFusion.
Qingle Zhang, Maoxun Yuan
ECAI3
2024 C²Former: Calibrated and Complementary Transformer for RGB-Infrared Object Detection
abstract
Object detection on visible (RGB) and infrared (IR) images, as an emerging solution to facilitate robust detection for around-the-clock applications, has received extensive attention in recent years. With the help of IR images, object detectors have been more reliable and robust in practical applications by using RGB-IR combined information. However, existing methods still suffer from modality miscalibration and fusion imprecision problems. Since transformer has the powerful capability to model the pairwise correlations between different features, in this paper, we propose a novel Calibrated and Complementary Transformer called C2Former to address these two problems simultaneously. In C2Former, we design an Inter-modality Cross-Attention (ICA) module to obtain the calibrated and complementary features by learning the cross-attention relationship between the RGB and IR modality. To reduce the computational cost caused by computing the global attention in ICA, an Adaptive Feature Sampling (AFS) module is introduced to decrease the dimension of feature maps. Because C2Former performs in the feature domain, it can be embedded into existed RGB-IR object detectors via the backbone network. Thus, one single-stage and one two-stage object detector both incorporating our C2Former are constructed to evaluate its effectiveness and versatility. With extensive experiments on the DroneVehicle and KAIST RGB-IR datasets, we verify that our method can fully utilize the RGB-IR complementary information and achieve robust detection results. The code is available at https://github.com/yuanmaoxun/C2Former.git.
Maoxun Yuan, Xingxing Wei 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Adversarial pan-sharpening attacks for object detection in remote sensing
Maoxun Yuan
Pattern Recognit.2
2022 Translation, Scale and Rotation: Cross-Modal Alignment Meets RGB-Infrared Vehicle Detection
Maoxun Yuan, Yinyan Wang, Xingxing Wei 0001
ECCV (9)1