Binbin Song

dblp:218/2718 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RouteWinFormer: A Mid-Range Route-Window Transformer With Structure Regularization for Image Restoration
abstract
Window-based Transformers have substantially improved the practicality of image restoration techniques by reducing computational costs, but their local attention mechanism introduces a bottleneck that limits further performance gains. Expanding the attention range holds promise for better image restoration quality. However, the expansion faces two fundamental issues: (1) prohibitive quadratic growth in computational complexity, and (2) increased attention noise in feature aggregation caused by the inclusion of more irrelevant pixels. In this paper, we propose RouteWinFormer, a novel window-based Transformer that enables wider-range attention in image restoration. RouteWinFormer incorporates the Route-Window Attention Module (RWAM) to efficiently extend attention to mid-range regions, which are identified as critical for image restoration according to an analysis of normalized average attention distances. RWAM adopts a router that dynamically selects a subset of nearby windows for aggregation, rather than including all neighboring windows, alleviating the growth in computational complexity. The selection leverages regional similarity based on a window-level description, reducing the inclusion of irrelevant pixels and attention noise during aggregation. In addition, we introduce Multi-Scale Structure Regularization (MSR) for RWAM to learn a better structure prior. MSR encourages the sub-scale branches of a U-shaped network to extract structural information, which serves as a condition for the original-scale branch to learn degradation patterns. Experimental results show that RouteWinFormer outperforms state-of-the-art methods on 8 benchmarks across diverse image restoration tasks, including defocus deblurring, desnowing, dehazing, deraining, and motion deblurring. The source code is available at: https://github.com/QifanLi-vilab/RouteWinFormer.
Qifan Li, Binbin Song, Jisheng Chu, Xiaopeng Fan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Single-Image Reflection Removal via Iterative Prompt Learning of Reflection Level
abstract
Single-image reflection removal (SIRR) aims to restore the latent background layer from a reflection-contaminated image. Despite the promising progress achieved by deep learning-based methods, the roles of negative training samples and descriptive prompts for the reflection severity are underexplored in most existing deep SIRR approaches, limiting their reflection removal performance and generalization capability. In this work, we introduce a novel training framework that synergistically leverages learnable prompts and image data to optimize the restoration network. To this end, we define reflection levels corresponding to varying degrees of reflection interference on the background content and learn reflection-level prompts to supervise the SIRR process. We propose an Iterative Reflection Level Reduction (IRLR) framework composed of a Restoration Network Training Module (RNTM) and a Reflection Level Learning Module (RLLM). Specifically, RNTM predicts the background layer under the guidance of prompts learned by RLLM, while RLLM in turn refines these prompts using outputs from RNTM. The two modules are trained iteratively to progressively reduce the reflection levels of estimated background layers. To initialize the prompts, we construct a dedicated reflection-level dataset for pretraining. For adaptively supervising RNTM, we design a new reflection-level-aware strategy to address the challenge of directly aligning the output background with the minimal reflection level. Comprehensive experimental results demonstrate that the proposed method significantly outperforms state-of-the-art methods on average performance across several released datasets, improving PSNR by 0.82 dB and SSIM by 0.0120, respectively. The source code and dataset are available at https://github.com/NamecantbeNULL/IRLR_SIRR.
Binbin Song, Jiantao Zhou 0001, Shuning Xu, Xina Liu, Haiwei Wu, Xiaopeng Fan 0001, Bihan Wen
IEEE Trans. Image Process.1
2025 Modeling Scattering Effect for Under-Display Camera Image Restoration
Binbin Song, Jiantao Zhou 0001, Xiangyu Chen 0006, Shuning Xu
Int. J. Comput. Vis.1
2024 Direction-Aware Video Demoiréing with Temporal-Guided Bilateral Learning
abstract
Moiré patterns occur when capturing images or videos on screens, severely degrading the quality of the captured images or videos. Despite the recent progresses, existing video demoiréing methods neglect the physical characteristics and formation process of moiré patterns, significantly limiting the effectiveness of video recovery. This paper presents a unified framework, DTNet, a direction-aware and temporal-guided bilateral learning network for video demoiréing. DTNet effectively incorporates the process of moiré pattern removal, alignment, color correction, and detail refinement. Our proposed DTNet comprises two primary stages: Frame-level Direction-aware Demoiréing and Alignment (FDDA) and Tone and Detail Refinement (TDR). In FDDA, we employ multiple directional DCT modes to perform the moiré pattern removal process in the frequency domain, effectively detecting the prominent moiré edges. Then, the coarse and fine-grained alignment is applied on the demoiréd features for facilitating the utilization of neighboring information. In TDR, we propose a temporal-guided bilateral learning pipeline to mitigate the degradation of color and details caused by the moiré patterns while preserving the restored frequency information in FDDA. Guided by the aligned temporal features from FDDA, the affine transformations for the recovery of the ultimate clean frames are learned in TDR. Extensive experiments demonstrate that our video demoiréing method outperforms state-of-the-art approaches by 2.3 dB in PSNR, and also delivers a superior visual experience.
Shuning Xu, Binbin Song, Xiangyu Chen 0006, Jiantao Zhou 0001
AAAI2
2024 Image Demoiréing in RAW and sRGB Domains
Shuning Xu, Binbin Song, Xiangyu Chen 0006, Xina Liu, Jiantao Zhou 0001
ECCV (6)2
2024 CNN Injected transformer for image exposure correction
Shuning Xu, Xiangyu Chen 0006, Binbin Song, Caishi Huang, Jiantao Zhou 0001
Neurocomputing3
2024 A Multispectral Remote Sensing Crop Segmentation Method Based on Segment Anything Model Using Multistage Adaptation Fine-Tuning
abstract
Multispectral information is crucial for remote sensing crop monitoring, but current methods struggle with inadequate feature extraction, leading to poor generalization and incomplete segmentation. The segment anything model (SAM) shows significant potential for generalization across fields, offering a promising solution for crop monitoring. This article introduces a crop segmentation method based on SAM using multistage adaptation fine-tuning, namely MAF-SAM, effectively utilizing information from multispectral remote sensing and the transfer and generalization abilities of SAM. In its first stage, MAF-SAM employs a prefix adapter to extract primary low-level multispectral features and compresses them into three channels to meet the requirements of subsequent stages. The second stage introduces a low-rank adaptation (LoRA) fine-tuning strategy to inject crop-specific knowledge into the image encoder, enhancing MAF-SAM’s adaptability in particular crop segmentation tasks. In the third stage, it utilizes a mask decoder with no-prompt embedding to automatically generate masks with accurate class information. MAF-SAM achieves F1 scores and Intersection over Union (IoU) for soybean and corn of 0.9294, 0.8680, 0.8760, and 0.7723, respectively, along with a Kappa coefficient of 0.9543. It demonstrates superior temporal and spatial transfer capabilities relative to five other advanced segmentation methods in our study area.
Binbin Song, Hui Yang 0017, Yanlan Wu, Peng Zhang 0059, Biao Wang 0005, Guichao Han
IEEE Trans. Geosci. Remote. Sens.1
2023 Under-Display Camera Image Restoration with Scattering Effect
abstract
The under-display camera (UDC) provides consumers with a full-screen visual experience without any obstruction due to notches or punched holes. However, the semitransparent nature of the display inevitably introduces the severe degradation into UDC images. In this work, we address the UDC image restoration problem with the specific consideration of the scattering effect caused by the display. We explicitly model the scattering effect by treating the display as a piece of homogeneous scattering medium. With the physical model of the scattering effect, we improve the image formation pipeline for the image synthesis to construct a realistic UDC dataset with ground truths. To suppress the scattering effect for the eventual UDC image recovery, a two-branch restoration network is designed. More specifically, the scattering branch leverages global modeling capabilities of the channel-wise self-attention to estimate parameters of the scattering effect from degraded images. While the image branch exploits the local representation advantage of CNN to recover clear scenes, implicitly guided by the scattering branch. Extensive experiments are conducted on both real-world and synthesized data, demonstrating the superiority of the proposed method over the state-of-the-art UDC restoration techniques. The source code and dataset are available at https://github.com/NamecantbeNULL/SRUDC.
Binbin Song, Xiangyu Chen 0006, Shuning Xu, Jiantao Zhou 0001
ICCV1
2023 Real-Scene Reflection Removal With RAW-RGB Image Pairs
abstract
Most brands of modern consumer digital cameras nowadays are able to provide RAW-RGB image pairs conveniently, even in the automatic mode. RAW images store pixel intensities linearly related to the radiance, which could be beneficial for the image reflection removal (IRR) task. However, existing IRR solutions, usually directly restoring the background in the non-linear RGB domain, severely overlook the valuable information conveyed by readily-available RAW images. Such a negligence may limit the performance of IRR methods on real-scene images. To mitigate this deficiency, we propose a Cascaded RAW and RGB Restoration Network (CR3Net) by leveraging both the RGB images and their paired RAW versions. Specifically, we firstly separate background and reflection layers in the linear RAW domain, and then restore the two layers in the non-linear RGB format by converting RAW features into the RGB domain. A novel RAW-to-RGB module (RRM) is devised to upsample these features and mimic pointwise mappings in the camera image signal processor (ISP). In addition, we collect the first real-world dataset that contains paired RAW and RGB images for IRR. Compared with state-of-the-art approaches, our method achieves a significant performance gain of about 2.07dB in PSNR, 0.028 in SSIM, and 0.0123 in LPIPS tested on the captured dataset. The source code and dataset are available athttps://github.com/NamecantbeNULL/RAW_RGB_RR.
Binbin Song, Jiantao Zhou 0001, Xiangyu Chen 0006, Shile Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2022 Multistage Curvature-Guided Network for Progressive Single Image Reflection Removal
abstract
Thanks to the powerful learning capability, deep neural networks (DNNs) have acquired broad applications in single image reflection removal. The DNN-based algorithms relax the constraints of specific priors and learn to generate visually pleasant background layers from massive training data. However, most of them employ a single network structure to recover both the semantic information and local details of the background, which may lead to obvious reflection residue or even failure. To mitigate this deficiency, in this work, we propose a Multi-stage Curvature-guided De-Reflection Network (MCDRNet), which combines multiple network architectures in a unified framework to progressively reconstruct the background layer and refine the fine-grained details. Our framework consists of three stages, where the encoder-decoders are exploited in the first two stages to recover the semantic components of background layers with lower scales and a variant ResNet is applied in the last stage to refine the background details with the original input resolution. In the first two stages, to introduce the structural guidance for the reflection removal, we cascade another decoder branch to restore the curvature map of the background. In addition, at the end of the first two stages, instead of directly passing the intermediate estimates to the next stage, we propose a Non-local Attention Module (NAM) to augment and transmit the features from decoders. Extensive experimental results on several public datasets demonstrate that the proposed MCDRNet outperforms the state-of-the-art methods quantitatively and generates visually better reflection removal results. The source code and pre-trained models are available athttps://github.com/NamecantbeNULL/MCDRNet.
Binbin Song, Jiantao Zhou 0001, Haiwei Wu
IEEE Trans. Circuits Syst. Video Technol.1
2019 Improving Person Search by Adaptive Feature Pyramid-based Multi-Scale Matching
abstract
Within person search problem, existing pedestrian detectors usually generate person bounding boxes with different sizes, which will degrade the performance of person re-identification in matching process. The built-in feature pyramid hierarchy to address the multi-scale matching problem adopted in previous studies, is prone to be effective for reducing the semantic gaps among different level features. Therefore, in this paper we propose an Adaptive Feature Pyramid (AFP)-based multi-scale matching method for person search. The network adaptively learns a set of adjusting vectors to make the low-level feature distributions approximate the high-level. Moreover, we integrate the AFP structure into a ResNet-50 based person search framework to search a target person from a gallery of whole scene images. Extensive experimental results show that our approach performs superiority over the state-of-the-art person search methods on CUHK-SYSU dataset.
Binbin Song
VCIP1
2018 Host load prediction with long short-term memory in cloud computing
Binbin Song, Yao Yu 0001, Yu Zhou 0007, Sidan Du
J. Supercomput.1