Bingze Song

dblp:305/0880 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-3968-4529ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
abstract
Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a dynamic test-time scaling law strategy inspired by imagery that adaptively adjusts the inference search space and reward guided by prompts, effectively enhancing generation quality in imaginative scenarios. Furthermore, we introduce LDT-Bench, the first benchmark targeting long-distance semantic prompts, designed to evaluate the creativity of video generation models. It comprises 2,839 challenging concept pairs from diverse recognition datasets and incorporates an automatic evaluation protocol to assess creative capacity. Extensive experiments on LDT-Bench demonstrate that our approach consistently outperforms general generation models and test-time scaling approaches. Additionally, ImagerySearch achieves strong performance on VBench, confirming its effectiveness in improving video generation quality under diverse conditions.
Meiqi Wu, Jiashu Zhu, Xiaokun Feng, Chubin Chen, Bingze Song, Fangyuan Mao, Jiahong Wu 0005, Xiangxiang Chu, Kaiqi Huang
AAAI6
2025 Refining Remote Sensing Image Segmentation Results Based on Vision Foundation Models
abstract
Pixel-level classification of remote sensing images is a fundamental task in Earth science-related research. However, current automated models inevitably produce some segmentation errors. In practical applications, manual review and correction are often required. To address the inevitable segmentation errors in automated remote sensing image segmentation, this letter proposes a remote sensing segmentation correction model. This model is based on the segment anything (SAM) model and requires training only a small number of parameters to enable the model to perceive click prompt. The model consists of two main components: one part is responsible for automated segmentation and the other part is dedicated to refine the results of the automated segmentation. Considering that click-based correction is difficult to learn during training, we designed a two-stage training process for the network. Through experiments on two widely used datasets, it was found that the proposed remote sensing image correction network significantly improves mIOU after incorporating the correction process. After applying fewer than six corrective clicks per category, the mIOU is improved by 12.99% in ISPRS Vaihingen dataset and 6.93% in ISPRS Postam dataset.
Bingze Song, Dongbo Wang, Peng Liu 0024, Gaoliang Xie
IEEE Geosci. Remote. Sens. Lett.1
2025 Diff-HRNet: A Diffusion Model-Based High-Resolution Network for Remote Sensing Semantic Segmentation
abstract
The semantic segmentation methods based on deep neural networks predominantly employ supervised learning, relying heavily on the quantity and quality of annotated samples. Due to the complexity of high-resolution remote sensing imagery, obtaining sufficient and precise pixel-level labeled data is highly challenging. This letter introduces a novel self-supervised learning method using a pretrained denoising diffusion probabilistic model (DDPM) to leverage semantic information from large-scale unlabeled remote sensing imageries. Building on this, a multistage fusion scheme between pretrained features and high-resolution features is proposed, enabling the network to learn more effective strategies to leverage prior information provided by the pretrained model while preserving the rich semantic details of high-resolution images. Experimental results on two remote sensing semantic segmentation datasets show that the proposed Diff-HRNet outperforms all compared methods, demonstrating the potential of pretrained diffusion models in extracting crucial feature representations for semantic segmentation tasks.
Chang Liu 0053, Bingze Song, Huaxin Pei, Pinjie Li, Mengshuo Chen
IEEE Geosci. Remote. Sens. Lett.3
2025 Click Prompt Learning With Feature Encoding for Segmentation of Remote Sensing Images
abstract
Pixel-level annotation tasks are important in the intelligent processing of remote sensing images. For these tasks, Interactive Image Segmentation (IIS) models using click prompts are developing fast in the field of natural images. However, most interactive segmentation models using click prompts are unsuitable for remote sensing images with their current design of click prompts and their interaction schemes with image information. Based on the situation, we used a DETR-like model as the basic framework and redesigned the pixel decoder and the transformer decoder to better suit the task of IIS for remote sensing images. In the pixel decoder, we designed a click prompt with feature encoding to learn click information and a composite attention structure to facilitate interaction between click and image information, allowing the image feature at the click locations to more easily dominate annotation masks. In the transformer decoder, we utilized deformable attention, using only a single initialized query to obtain annotation masks and IoU prediction. In this paper, we trained our model on a composite remote sensing dataset and evaluated its performance on external datasets. The results showcased the model’s adaptability, achieving superior performance compared to existing methods. The code will be available at https://github.com/songbingze/ClickPromptRSIIS.
Bingze Song, Peng Liu 0024, Lingjun Zhao, Lajiao Chen, Mengzhen Xu, Yi Zeng 0002
IEEE Trans. Geosci. Remote. Sens.1
2024 Reconstruction of Large-Scale Missing Data in Remote Sensing Images Using Extend-GAN
abstract
Numerous studies have been conducted on missing data recovery in remote sensing images, such as cloud removal and dead pixels restoration. Nevertheless, reconstructing continuous, extensive, and complete missing areas still poses a significant challenge. In this letter, we propose a new architecture named Extend-generative adversarial network (GAN), which leverages only a low-resolution image with relaxed requirements on spatial resolution and acquisition time as a condition to reconstruct a high-resolution image with large-scale missing areas. We equip Extend-GAN with learnable adaptive region normalization (LARN) to adjust the intensity distribution of pixels to reduce color distortion. We also introduce a new loss function into the training process of Extend-GAN, namely the structural similarity (SSIM)-based triplet loss, which helps to preserve the between missing parts and known regions. Gaofen-2 and Landsat-9 image pairs are used to validate the proposed method. Extend-GAN performs better when comprehensively evaluated on visual effect, quantitative metrics, processing speed, etc. Code is available athttps://github.com/yc-cui/Extend-GAN.
Yongchuan Cui, Peng Liu 0024, Bingze Song, Lingjun Zhao, Yan Ma 0001, Lajiao Chen
IEEE Geosci. Remote. Sens. Lett.3
2022 MLFF-GAN: A Multilevel Feature Fusion With GAN for Spatiotemporal Remote Sensing Images
abstract
Due to the limitation of technology and budget, it is often difficult for sensors of a single remote sensing satellite to have both high temporal resolution and high spatial (HTHS) resolution at the same time. In this paper, we proposed a new Multi-level Feature Fusion with Generative Adversarial Network (MLFF-GAN) for generating fusion HTHS images. MLFF-GAN mainly uses U-net-like architecture and its generator is composed of three stages: feature extraction, feature fusion, and image reconstruction. In feature extraction and reconstruction stage, the generator employs the encoding and decoding structure to extract three groups of multi-level features, which can cope with the huge difference of resolution between high-resolution images and low-resolution images. In the feature fusion stage, Adaptive Instance Normalization (AdaIN) block is designed to learn the global distribution relationship between multi-temporal images, and an attention module (AM) is used to learn the local information weights for the change of small areas. The proposed MLFF-GAN was tested on two Landsat and MODIS datasets. Some state-of-the-art algorithms are comprehensively compared with MLFF-GAN. We also carried on the ablation experiment to test the effectiveness of different sub-module in MLFF-GAN. The experiment results and ablation analysis show the better performances of the proposed method when compared with other methods. The code is available at https://github.com/songbingze/MLFF-GAN.
Bingze Song, Peng Liu 0024, Jun Li 0009, Lizhe Wang 0001, Guojin He, Lajiao Chen
IEEE Trans. Geosci. Remote. Sens.1