Haichen Yu

dblp:164/2358 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0001-8123-417XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 29% Deep learning architectures and training · 25% 3D vision · 15%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
expert pruning
1.012026
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision › remote sensing
remote sensing image analysis
1.012026
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training
foundation model
0.912025
RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training › foundation model
remote sensing foundation model
0.912025
RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.0mixture of experts · 1.0dynamic expert pruning · 1.0parameter-efficient fine-tuning · 0.9frequency-enhanced multi-head self-attention · 0.9affine transformation contrastive learning · 0.9
YearPublicationVenuePosition
2026 RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation
abstract
The rapid advancement of foundation models has revolutionized visual representation learning in a self-supervised manner. However, their application in remote sensing (RS) remains constrained by a fundamental gap: existing models predominantly handle single or limited modalities, overlooking the inherently multi-modal nature of RS observations. Optical, synthetic aperture radar (SAR), and multi-spectral data offer complementary insights that significantly reduce the inherent ambiguity and uncertainty in single-source analysis. To bridge this gap, we introduce RingMoE, a unified multi-modal RS foundation model with 14.7 billion parameters, pre-trained on 400 million multi-modal RS images from nine satellites. RingMoE incorporates three key innovations: 1) A hierarchical Mixture-of-Experts (MoE) architecture comprising modal-specialized, collaborative, and shared experts, effectively modeling intra-modal knowledge while capturing cross-modal dependencies to mitigate conflicts between modal representations; 2) Physics-informed self-supervised learning, explicitly embedding sensor-specific radiometric characteristics into the pre-training objectives; 3) Dynamic expert pruning, enabling adaptive model compression from 14.7B to 1B parameters while maintaining performance, facilitating efficient deployment in Earth observation applications. Evaluated across 23 benchmarks spanning six key RS tasks (i.e., classification, detection, segmentation, tracking, change detection, and depth estimation), RingMoE outperforms existing foundation models and sets new SOTAs, demonstrating remarkable adaptability from single-modal to multi-modal scenarios. Beyond theoretical progress, it has been deployed and trialed in multiple sectors, including emergency response, land management, marine sciences, and urban planning.
Hanbo Bi, Yingchao Feng, Boyuan Tong, Haichen Yu, Yongqiang Mao, Wenhui Diao, Peijin Wang, Yue Yu 0001, Hanyang Peng, Yehong Zhang, Kun Fu 0001, Xian Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning
abstract
Aerial Remote Sensing (ARS) vision tasks present significant challenges due to the unique viewing angle characteristics. Existing research has primarily focused on algorithms for specific tasks, which have limited applicability in a broad range of ARS vision applications. This paper proposes RingMo-Aerial, aiming to fill the gap in foundation model research in the field of ARS vision. A Frequency-Enhanced Multi-Head Self-Attention (FE-MSA) mechanism is introduced to strengthen the model's capacity for small-object representation. Complementarily, an affine transformation-based contrastive learning method improves its adaptability to the tilted viewing angles inherent in ARS tasks. Furthermore, the ARS-Adapter, an efficient parameter fine-tuning method, is proposed to improve the model's adaptability and performance in various ARS vision tasks. Experimental results demonstrate that RingMo-Aerial achieves SOTA performance on multiple downstream tasks. This indicates the practicality and efficacy of RingMo-Aerial in enhancing the performance of ARS vision tasks.
Wenhui Diao, Haichen Yu, Kaiyue Kang, Tong Ling, Yingchao Feng, Hanbo Bi, Libo Ren, Xuexue Li, Yongqiang Mao, Xian Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Injecting Linguistic Into Visual Backbone: Query-Aware Multimodal Fusion Network for Remote Sensing Visual Grounding
abstract
The remote sensing visual grounding (RSVG) task focuses on accurately identifying and localizing specific targets in remote sensing (RS) images using descriptive query expressions. Existing methods independently extract visual and textual features, ignoring early complementary information between image and text. This leads to information loss and misalignment, limiting the model’s ability to distinguish similar targets. To address this challenge, we propose the query-aware multimodal fusion network (QAMFN), which introduces an innovative query-guided visual attention (QGVA) mechanism in the early stages of the visual encoder. This mechanism integrates textual information during the early visual feature extraction process, thereby resolving the issue of missing image-text complementary information. QGVA ensures that the visual backbone accurately focuses on local features highly relevant to the query by injecting textual information into the visual encoding process. Additionally, to enhance the model’s ability to integrate multimodal information and adapt to more complex RS images, we introduce the text-semantic attention-guided masking (TAM) module. TAM aggregates multimodal features processed by the backbones and filters out redundant information, producing high-quality fused features. Experiments demonstrate that our approach sets a new record on the DIOR-RSVG dataset, improving accuracy to 81.67% (an absolute increase of 4.98%).
Wenkai Zhang 0002, Hanbo Bi, Shuoke Li, Haichen Yu, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.6