Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chen Chen 0152

dblp:65/4423-152 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0005-7923-9018ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Image recognition and object detection · 30% Face, body and person analysis · 30% Representation and self-supervised learning · 17%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
multimodal object detection
1.622025
Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition Cues · ICCV 2025
Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object Detection · CVPR 2024
Computer vision › Face, body and person analysis › person re-identification
vehicle re-identification
1.622025
UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identification · CVPR 2025
Relation-Aware Weight Sharing in Decoupling Feature Learning Network for UAV RGB-Infrared Vehicle Re-Identification · IEEE Trans. Multim. 2024
Computer vision › 3D vision
depth estimation
0.912025
Dive into Aerial Remote Sensing Underwater Depth Estimation with Hyperspectral Imagery · AAAI 2025
Computer vision › Image recognition and object detection › object detection › aerial object detection
UAV object detection
0.912025
Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition Cues · ICCV 2025
Computer vision › 3D vision › depth estimation › scene depth estimation
underwater depth estimation
0.912025
Dive into Aerial Remote Sensing Underwater Depth Estimation with Hyperspectral Imagery · AAAI 2025
Machine learning › Representation and self-supervised learning › representation matching › feature alignment › modality alignment
cross-modality feature alignment
0.812024
Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object Detection · CVPR 2024
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.812024
Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object Detection · CVPR 2024
Computer vision › Face, body and person analysis
person re-identification
0.812024
Relation-Aware Weight Sharing in Decoupling Feature Learning Network for UAV RGB-Infrared Vehicle Re-Identification · IEEE Trans. Multim. 2024
Computer vision › Image recognition and object detection
RGB-thermal fusion
0.812024
Relation-Aware Weight Sharing in Decoupling Feature Learning Network for UAV RGB-Infrared Vehicle Re-Identification · IEEE Trans. Multim. 2024
Computer vision › Face, body and person analysis › person re-identification › multi-modal person re-identification
visible-infrared person re-identification
0.812024
Relation-Aware Weight Sharing in Decoupling Feature Learning Network for UAV RGB-Infrared Vehicle Re-Identification · IEEE Trans. Multim. 2024
Computer vision › Vision and language › vision-language model
prompt learning
0.312025
Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition Cues · ICCV 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.312025
UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identification · CVPR 2025
Machine learning › Deep learning architectures and training
state space model
0.312025
Dive into Aerial Remote Sensing Underwater Depth Estimation with Hyperspectral Imagery · AAAI 2025
Robotics › Legged, aerial and field robots › aerial robots
unmanned aerial vehicle
0.212024
Relation-Aware Weight Sharing in Decoupling Feature Learning Network for UAV RGB-Infrared Vehicle Re-Identification · IEEE Trans. Multim. 2024

Methods — techniques the papers use, named apart from their topics

soft gating · 0.9prompt-guided condition-aware dynamic fusion · 0.9physical imaging model · 0.9patch-mixed image reconstruction · 0.9multimodal fusion · 0.9modality-augmented contrasting cluster · 0.9modality discrimination adversarial learning · 0.9hyperspectral imaging · 0.9offset modeling · 0.8deformable alignment · 0.8
YearPublicationVenuePosition
2026 Separate to generalization: Two-branch feature separation framework for generalized underwater image restoration
Jiahao Qi, Chen Chen 0152, Kangcheng Bin, Ping Zhong 0001
Pattern Recognit.3
2025 Dive into Aerial Remote Sensing Underwater Depth Estimation with Hyperspectral Imagery
abstract
Visible spectrum images capture limited information from just three discrete bands, often resulting in suboptimal performance in underwater depth estimation (UDE) due to significant information loss from water absorption. In contrast, HSIs, which include hundreds of continuous bands, provide abundant spectral information that offers greater resilience against the adverse effects of water absorption. In this paper, we conduct a comprehensive study to investigate how spectral information can enhance remote sensing UDE through two key aspects: the benchmark dataset and the general framework. For the benchmark dataset, we construct a real-world hyperspectral UDE (HUDE) dataset ATR-HUDE, comprising approximately 500 synchronized hyperspectral and LiDAR data pairs collected from diverse coastal scenes and flight altitudes. Regarding the general framework, we integrate recent advances in state space models and physical imaging models to design a novel HUDE framework named HUDEMamba that estimates underwater depth using both model-driven and data-driven approaches. Experimental results on the constructed benchmark dataset validate the potential of HUDE and the effectiveness of HUDEMamba.
Jiahao Qi, Chen Chen 0152, Dehui Zhu, Kangcheng Bin, Ping Zhong 0001
AAAI3
2025 UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identification
abstract
Cross-Modality Re-Identification (VT-ReID) aims to achieve around-the-clock target matching, benefiting from the strengths of both RGB and infrared (IR) modalities. However, the field is hindered by limited datasets, particularly for vehicle VT-ReID, and by challenges such as modality bias training (MBT), stemming from biased pre-training on ImageNet. To tackle the above issues, this paper introduces an dataset benchmark, named UCM-VeID V2, for vehicle VT-ReID, and proposes a new self-supervised pre-training method, Cross-Modality Patch-Mixed Self-Supervised Learning (PMSL). UCM-VeID V2 dataset features a significant increase in data volume, along with enhancements in multiple aspects. PMSL addresses MBT by learning modality-invariant features through Patch-Mixed Image Reconstruction (PMIR) and Modality Discrimination Adversarial Learning (MDAL), and enhances discriminability with Modality-Augmented Contrasting Cluster (MACC). Comprehensive experiments are carried out to validate the proposed method.
Jiahao Qi, Chen Chen 0152, Kangcheng Bin, Ping Zhong 0001
CVPR3
2025 Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition Cues
abstract
Unmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of high-quality dataset. However, the existing dataset struggles to fully capture real-world complexity for limited imaging conditions. To this end, we introduce a high-diversity dataset ATR-UMOD covering varying scenarios, spanning altitudes from 80m to 300m, angles from 0° to 75°, and all-day, all-year time variations in rich weather and illumination conditions. Moreover, each RGB-IR image pair is annotated with 6 condition attributes, offering valuable high-level contextual information. To meet the challenge raised by such diverse conditions, we propose a novel prompt-guided condition-aware dynamic fusion (PCDF) to adaptively reassign multimodal contributions by leveraging annotated condition cues. By encoding imaging conditions as text prompts, PCDF effectively models the relationship between conditions and multimodal contributions through a task-specific soft-gating transformation. A prompt-guided condition-decoupling module further ensures the availability in practice without condition annotations. Experiments on ATR-UMOD dataset reveal the effectiveness of PCDF.
Chen Chen 0152, Kangcheng Bin, Jiahao Qi, Tianpeng Liu, Zhen Liu 0004, Yongxiang Liu, Ping Zhong 0001
ICCV1
2025 Discriminative Latent-Space Learning for Fine-Grained Object Detection in Remote Sensing Images
abstract
Abstract—Fine-grained object detection (FOD) is essential in many remote sensing image interpretation tasks. Existing FOD methods have achieved remarkable progress in modeling discriminative features for FOD in remote sensing images. However, they receive unsatisfactory recognition accuracy due to the curse of dimensionality (CoD) problem. In this paper, we propose an orthogonal constraint-based discriminative latent-space learning (DLL) method to address the CoD problem. We first optimize a sparse optimization paradigm with convex relaxation to extract shared features between fine-grained objects into a low-dimensional latent-space. Then, we solve the orthogonal space of the latent-space for extracting irrelevant features related to common features from input features, i.e., discriminative features. The sparse optimization paradigm reprojects the underlying trends hidden in high-dimensional feature space into a low-dimensional latent space and hence addresses the CoD problem. We use neural parameters with latent-space and orthogonal constraints to approximately solve the proposed DLL, which can be efficiently optimized under convex programming tools. We theoretically prove the effectiveness of the proposed DLL. By adding our DLL to the existing deep learning-based object detection method, extensive experiments conducted on two datasets demonstrate that our method achieves superior performance compared with other state-of-the-art methods.
Xikun Hu, Wenlin Liu, Xiangsheng Wang, Chen Chen 0152, Ya Jiang, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Physics-Informed Curriculum Learning Framework for Hyperspectral Underwater Target Characterization
abstract
Hyperspectral imaging (HSI) provides fine-grained spectral information essential for material identification and target detection, particularly in complex environments such as underwater scenarios. However, hyperspectral underwater target detection (HUTD) remains challenging due to severe spectral distortions and variability introduced by wavelength-dependent absorption and the dynamic nature of aquatic environments. Existing separation-based and characterization-based methods are often constrained by weak signal responses or a heavy reliance on accurate environmental parameter estimation, which is difficult to achieve in practice. To overcome these limitations, we propose PCL-HUTD, a novel physics-informed curriculum learning framework for robust underwater target characterization without requiring explicit environmental modeling. PCL-HUTD integrates a physics-guided target construction module with a hard-sample aware contrastive learning strategy, enhanced by unsupervised clustering and a perturbation-consistency based sample selection mechanism. Furthermore, a closed-loop curriculum learning paradigm is introduced to progressively refine target representations throughout training. Extensive experiments on three real-world HUTD datasets demonstrate that PCL-HUTD achieves state-of-the-art performance in both detection accuracy and robustness, particularly under challenging conditions with strong background interference. These results validate the effectiveness of our parameter-free, physics-informed approach for underwater hyperspectral target detection.
Jiahao Qi, Chen Chen 0152, Dehui Zhu, Kangcheng Bin, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Masked Spatial-Spectral Autoencoders Are Excellent Hyperspectral Defenders
abstract
Deep learning (DL) methodology contributes a lot to the development of hyperspectral image (HSI) analysis community. However, it also makes HSI analysis systems vulnerable to adversarial attacks. To this end, we propose a masked spatial-spectral autoencoder (MSSA) in this article under self-supervised learning theory, for enhancing the robustness of HSI analysis systems. First, a masked sequence attention learning (MSAL) module is conducted to promote the inherent robustness of HSI analysis systems along spectral channel. Then, we develop a graph convolutional network (GCN) with learnable graph structure to establish global pixel-wise combinations. In this way, the attack effect would be dispersed by all the related pixels among each combination, and a better defense performance is achievable in spatial aspect. Finally, to improve the defense transferability and address the problem of limited labeled samples, MSSA employs spectra reconstruction as a pretext task and fits the datasets in a self-supervised manner. Comprehensive experiments over three benchmarks verify the effectiveness of MSSA in comparison with the state-of-the-art hyperspectral classification methods and representative adversarial defense strategies.
Jiahao Qi, Zhiqiang Gong, Chen Chen 0152, Ping Zhong 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object Detection
abstract
Visible-infrared (RGB-IR) image fusion has shown great potentials in object detection based on unmanned aerial ve-hicles (UAVs). However, the weakly misalignment problem between multimodal image pairs limits its performance in object detection. Most existing methods often ignore the modality gap and emphasize a strict alignment, resulting in an upper bound of alignment quality and an increase of implementation costs. To address these challenges, we propose a novel method named Offset-guided Adaptive Feature Alignment (OAFA), which could adaptively adjust the relative positions between multimodal features. Considering the impact of modality gap on the cross-modality spa-tial matching, a Cross-modality Spatial Offset Modeling (CSOM) module is designed to establish a common sub-space to estimate the precise feature-level offsets. Then, an Offset-guided Deformable Alignment and Fusion (ODAF) module is utilized to implicitly capture optimal fusion po-sitions for detection task rather than conducting a strict alignment. Comprehensive experiments demonstrate that our method not only achieves state-of-the-art performance in the UAVs-based object detection task but also shows strong robustness to the weakly misalignment problem.
Chen Chen 0152, Jiahao Qi, Kangcheng Bin, Ruigang Fu, Xikun Hu, Ping Zhong 0001
CVPR1
2024 Relation-Aware Weight Sharing in Decoupling Feature Learning Network for UAV RGB-Infrared Vehicle Re-Identification
abstract
Owing to the capacity of performing full-time target searches, cross-modality vehicle re-identification based on unmanned aerial vehicles (UAV) is gaining more attention in both video surveillance and public security. However, this promising and innovative research has not been studied sufficiently due to the issue of data inadequacy. Meanwhile, the cross-modality discrepancy and orientation discrepancy challenges further aggravate the difficulty of this task. To this end, we pioneer a cross-modality vehicle Re-ID benchmark named UAV Cross-Modality Vehicle Re-ID (UCM-VeID), containing 753 identities with16015RGB and13913infrared images. Moreover, to meet cross-modality discrepancy and orientation discrepancy challenges, we present a hybrid weights decoupling network (HWDNet) to learn the shared discriminative orientation-invariant features. For the first challenge, we proposed a hybrid weights siamese network with a well-designed weight restrainer and its corresponding objective function to learn both modality-specific and modality shared information. In terms of the second challenge, three effective decoupling structures with two pretext tasks are investigated to flexibly conduct orientation-invariant feature separation task. Comprehensive experiments are carried out to validate the effectiveness of the proposed method.
Jiahao Qi, Chen Chen 0152, Kangcheng Bin, Ping Zhong 0001
IEEE Trans. Multim.3