EDBT 2026 Demo / reviewers in the wild / expert
Bin Zhang 0046
dblp:13/5236-46
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-9545-2760ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SkySplat: Generalizable 3D Gaussian Splatting from Multi-Temporal Sparse Satellite ImagesabstractThree-dimensional scene reconstruction from sparse-view satellite images is a long-standing and challenging task. While 3D Gaussian Splatting (3DGS) and its variants have recently attracted attention for its high efficiency, existing methods remain unsuitable for satellite images due to incompatibility with rational polynomial coefficient (RPC) models and limited generalization capability. Recent advances in generalizable 3DGS approaches show potential, but they perform poorly on multi-temporal sparse satellite images due to limited geometric constraints, transient objects, and radiometric inconsistencies. To address these limitations, we propose SkySplat, a novel self-supervised framework that integrates the RPC model into the generalizable 3DGS pipeline, enabling more effective use of sparse geometric cues for improved reconstruction. SkySplat relies only on RGB images and radiometric-robust relative height supervision, thereby eliminating the need for ground-truth height maps. Key components include a Cross-Self Consistency Module (CSCM), which mitigates transient object interference via consistency-based masking, and a multi-view consistency aggregation strategy that refines reconstruction results. Compared to per-scene optimization methods, SkySplat achieves an 86 times speedup over EOGS with higher accuracy. It also outperforms generalizable 3DGS baselines, reducing MAE from 13.18 m to 1.80 m on the DFC19 dataset significantly, and demonstrates strong cross-dataset generalization on the MVS3D benchmark. Xuejun Huang, Xinyi Liu 0002, Yi Wan 0001, Bin Zhang 0046, Mingtao Xiong, Yingying Pei, Yongjun Zhang 0002 |
AAAI | 5 |
| 2026 | Travel Like a Train: Detection and Reconstruction of Rail Tracks From Aerial Images
Yongjun Zhang 0002, Guangshuai Wang, Bin Zhang 0046, Jiwei Deng, Ziqian Huang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Confidence-Aware Superpixel Self-Supervised Learning for Cross-Modal Point Cloud RecognitionabstractThe high cost of 3D annotation severely limits the availability of high-quality labeled data, posing a critical bottleneck to point cloud recognition tasks. In contrast, acquiring and annotating natural images is far more cost-effective, making 2D image-based methods more accessible. With the rise of large-scale pre-trained 2D feature extraction models, which have demonstrated remarkable embedding capabilities across diverse domains, an opportunity emerges to leverage them for 3D point cloud tasks. By transferring knowledge from well-established 2D models, reliance on 3D annotation can be significantly reduced while improving feature extraction efficiency. However, adapting these 2D models, originally trained on natural images, to remote sensing data introduces several challenges. In particularly, the generalization ability of 2D models is limited, and widely used point-pixel constraint methods suffer from high computational complexity and low error tolerance. To address these issues, this letter proposes the confidence-aware superpixel self-supervised learning method (CS-SSP). By incorporating confidence estimation, CS-SSP enhances the reliability of cross-domain feature embeddings, ensuring accurate feature transfer from natural images to remote sensing data. Additionally, a superpixel-based constraint is introduced to reduce computational complexity. Moreover, CS-SSP employs a coarse-to-fine training strategy, progressively refining constraints to mitigate training difficulties. The experiments demonstrate that CS-SSP outperforms baseline methods across all metrics. Notably, when fine-tuning with extremely limited labeled samples, CS-SSP achieves the state-of-the-art accuracy across all key metrics, highlighting its strong advantage in low-data scenarios. Yameng Wang, Bin Zhang 0046, Yi Wan 0001, Zhaoxi Yue, Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | MVSR3D: An End-to-End Framework for Semantic 3-D Reconstruction Using Multiview Satellite ImageryabstractSemantic 3D reconstruction from multi-view images is essential for applications such as 3D city modeling and robot navigation. However, existing methods treat semantic segmentation and height estimation as separate tasks, leading to suboptimal reconstruction results. To bridge this gap, we introduce MVSR3D, the first end-to-end framework for semantic 3D reconstruction using multi-view satellite images. MVSR3D employs a dual-stream architecture, consisting of the segmentation branch (MVSAM) based on Segment Anything Model (SAM) and the height estimation branch based on multi-view stereo. To enhance multi-view feature fusion, we propose the Epipolar Cross Attention (ECA) module in the MVSAM branch, which integrates image embeddings primarily along epipolar line to exploit complementary multi-view information. Unlike conventional multi-task learning approaches, we design dedicated interaction modules—the SAM Feature-Guided (SAM-FG) module and the Elevation-Guided Sparse Prompts Generator (EGSPG)—to facilitate multi-task interaction and feature fusion. Extensive evaluations on the DFC19 and SpaceNet4 datasets demonstrate that MVSR3D significantly outperforms the state-of-the-art multi-view multi-task learning method, improving the mIoU3 metric at a 2.5-meter threshold by 37.09%–45.11%. Xuejun Huang, Xinyi Liu 0002, Yi Wan 0001, Bin Zhang 0046, Yameng Wang, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | CloudViT: A Lightweight Vision Transformer Network for Remote Sensing Cloud DetectionabstractClouds inevitably exist in satellite images, which limit the processing and application of satellite images to a certain extent. Therefore, cloud detection is a preprocessing task in satellite image extraction and analysis processing. However, the existing methods are difficult to mine robust features, and the number of parameters and computation are large, which is not conducive to the deployment of the model. In this letter, cloud vision transformer (CloudViT), a lightweight vision transformer network for cloud detection from satellite imagery, is proposed. In detail, to utilize dark channel priors in multispectral imagery to guide the network to learn features, a multiscale dark channel extractor is used to first predict dark channels, and then, the dark channel features and image features are input to the attention mechanism-based dark channel-guided context aggregation module to enhance image features, which in turn makes cloud detection results more accurate. At the same time, to enhance the transfer ability of the network between different satellite sensors, a plug-and-play channel adaptive module is proposed to deal with the inconsistency of the number of different satellite sensor bands. The experimental results on the Landsat7 dataset show that our network CloudViT outperforms the state-of-the-art methods while keeping the number of parameters and computation small. At the same time, the experimental results on transfer to three other datasets show that using the channel adaptation module can greatly improve the transfer ability of the model. Bin Zhang 0046, Yongjun Zhang 0002, Yansheng Li 0001, Yi Wan 0001, Yongxiang Yao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | ELSR: Efficient Line Segment Reconstruction with Planes and Points GuidanceabstractThree-dimensional (3D) line segments are helpful for scene reconstruction. Most of the existing 3D-line-segment reconstruction algorithms deal with two views or dozens of small-size images; while in practice there are usually hundreds or thousands of large-size images. In this paper, we propose an efficient line segment reconstruction method called ELSR11Available at https://skyearth.org/publication/project/ELSR. ELSR exploits scene planes that are commonly seen in city scenes and sparse 3D points that can be acquired easily from the structure-from-motion (SfM) approach. For two views, ELSR efficiently finds the local scene plane to guide the line matching and exploits sparse 3D points to accelerate and constrain the matching. To reconstruct a 3D line segment with multiple views, ELSR utilizes an efficient abstraction approach that selects representative 3D lines based on their spatial consistence. Our experiments demonstrated that ELSR had a higher accuracy and efficiency than the existing methods. Moreover, our results showed that ELSR could reconstruct 3D lines efficiently for large and complex scenes that contain thousands of large-size images. Yi Wan 0001, Yongjun Zhang 0002, Xinyi Liu 0002, Bin Zhang 0046, Xiqi Wang |
CVPR | 5 |
| 2022 | SCTRANS: A TRANSFORMER NETWORK BASED ON THE SPATIAL AND CHANNEL ATTENTION FOR CLOUD DETECTIONabstractCloud detection is an important preprocessing step for remote sensing image processing and analysis. The current deep-learning-based cloud detection methods are mostly based on Convolutional Neural Network (CNN) which pay more attention to local information. To make more use of the global information, in this article, we propose a transformer-based cloud detection method (SCTrans) based on the spatial and channel attention mechanism. The experiment results show that when using only three-band images on the Landsat7 dataset, the mIoU of the validation set reaches 85.92% and the mIoU of the test set reaches 87.86%. The experimental results show that the proposed network has a higher mIoU and F1 score than Fmask and other networks. Wenke Jiao, Yongjun Zhang 0002, Bin Zhang 0046, Yi Wan 0001 |
IGARSS | 3 |
| 2022 | A Cascaded Cross-Modal Network for Semantic Segmentation from High-Resolution Aerial Imagery and RAW Lidar DataabstractAs various sensors appear, extracting information from multimodal data becomes a prominent topic. Current multimodal approaches for image and LiDAR normally discard the point-to-point topology relationship of the latter to keep the dimension matched. To tackle this task, we propose a cascaded cross-modal network (CCMN) to extract the joint-features from high-resolution aerial imagery and LiDAR point directly, instead of their abridged derivatives. Firstly, point-wise features are extract from raw LiDAR data by a forepart 3D extractor. Subsequently, the LiDAR-derived features are executed spatial reference conversion to project and align to the imagery coordinate space. Finally, the cross-modal compounds containing the obtained feature maps and the corresponding images are placed into a U-shape structure to generate segmentation results. The experiment results indicate that our strategy surpasses the popular multimodal method by 6% on mIoU. Yameng Wang, Bin Zhang 0046, Yi Wan 0001, Yongjun Zhang 0002 |
IGARSS | 2 |
| 2020 | Registration of Multimodal Remote Sensing Images Using Transfer OptimizationabstractMultimodal image registration is critical yet challenging for remote sensing image processing. Due to the large nonlinear intensity differences between the multimodal images, conventional search algorithms tend to get trapped into local optima when optimizing the transformation parameters by maximizing mutual information (MI). To address this problem, inspired by transfer learning, we propose a novel search algorithm named transfer optimization (TO), which can be applied to any optimizer. In TO, an optimizer transfers its better individuals to the other optimizer in each iteration. Thus, TO can share information between two optimizers and take advantage of their search mechanisms, which is helpful to avoid the local optima. Then, the registration of the multimodal remote sensing images using TO is presented. We compare the proposed algorithm with several state-of-the-art algorithms on real and simulated image pairs. Experimental results demonstrate the superiority of our algorithm in terms of registration accuracy. Xiaohu Yan, Yongjun Zhang 0002, Dejun Zhang, Neng Hou, Bin Zhang 0046 |
IEEE Geosci. Remote. Sens. Lett. | 5 |