EDBT 2026 Demo / reviewers in the wild / expert
Yuanbo Wen 0002
dblp:262/3144-2
· DBLP profile ↗
23ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0001-7599-5645ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unpaired iterative prompt learning for real-world image deraining
Yuanbo Wen 0002, Tao Gao 0001, Shan Liang 0002, Ting Chen 0003 |
Expert Syst. Appl. | 1 |
| 2026 | Prior-oriented specific and triple-view general prompts for multi-weather degraded image restoration
Yuanbo Wen 0002, Tao Gao 0001, Shan Liang 0002, Ting Chen 0003 |
Expert Syst. Appl. | 1 |
| 2026 | Multi-perspective prompt and assimilated self-modulation transformer for adverse weather removal
Yuanbo Wen 0002, Tao Gao 0001, Shan Liang 0002, Zixiang Liu, Ting Chen 0003 |
Expert Syst. Appl. | 1 |
| 2026 | A wavelet-guided and physics-aware network for remote sensing image dehazing
Qianxi Zhang, Ting Chen 0003, Tao Gao 0001, Yuanbo Wen 0002, Shan Liang 0002 |
Expert Syst. Appl. | 4 |
| 2026 | Dual-attention cooperative and multi-view gated transformer for adverse weather removal
Shan Liang 0002, Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003 |
Pattern Recognit. | 2 |
| 2026 | Prompt-oriented and frequency-regularized schrödinger bridge for unpaired rain streaks and raindrops removal
Yuanbo Wen 0002, Ting Chen 0003, Tao Gao 0001 |
Pattern Recognit. | 1 |
| 2026 | DynStaticNet: A biological vision-inspired dual-branch all-in-one network for video weather removal
Qianxi Zhang, Tao Gao 0001, Ting Chen 0003, Yuanbo Wen 0002, Tao Lei 0003 |
Pattern Recognit. | 4 |
| 2026 | Structure-Preserving Frequency-Regularized Text-Guided Optimal Transport for Unpaired Rain Streaks and Raindrops RemovalabstractThe removal of rain streaks and raindrops is crucial for enhancing the image visibility and mitigating the weather degradations. However, most existing approaches rely on the paired rainy and clean images, which are challenging to obtain in real-world scenarios. To this end, we propose a novel structure-preserving frequency-regularized text-guided optimal transport (SFTOT) framework, which formulates the unpaired rain streaks and raindrops removal as an optimal transport problem. Specifically, we introduce a structure-preserving transport cost, incorporating the structural similarity constraint to minimize the duality gap between the primal and dual formulations, while preserving the structural details of reconstructed images. Furthermore, by embedding the inherent frequency sparsity of rain streaks and raindrops into the transport cost, we derive a frequency-regularized optimal transport objective, ensuring consistency in frequency distributions between the generated and clean images. Additionally, we employ a pre-trained one-step stable diffusion model as the restoration network, which is fine-tuned using the low-rank adaptation (LoRA) adapters and zero convolutional layers, while integrating the domain-specific text prompts for both degraded and clean images to guide the generation process. Extensive experiments demonstrate that our method surpasses the existing well-performing unpaired learning approaches, achieving notable improvements in both the fidelity and photo-realism. Yuanbo Wen 0002, Tao Gao 0001, Qianxi Zhang, Jing Zhang 0052, Ting Chen 0003, Lidong Liu |
IEEE Trans. Multim. | 1 |
| 2025 | Multi-axis Prompt and Multi-dimension Fusion Network for All-in-one Weather-degraded Image RestorationabstractExisting approaches aiming to remove adverse weather degradations compromise the image quality and incur the long processing time. To this end, we introduce a multi-axis prompt and multi-dimension fusion network (MPMF-Net). Specifically, we develop a multi-axis prompts learning block (MPLB), which learns the prompts along three separate axis planes, requiring fewer parameters and achieving superior performance. Moreover, we present a multi-dimension feature interaction block (MFIB), which optimizes intra-scale feature fusion by segregating features along height, width and channel dimensions. This strategy enables more accurate mutual attention and adaptive weight determination. Additionally, we propose the coarse-scale degradation-free implicit neural representations (CDINR) to normalize the degradation levels of different weather conditions. Extensive experiments demonstrate the significant improvements of our model over the recent well-performing approaches in both reconstruction fidelity and inference time. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
AAAI | 1 |
| 2025 | Brain-Inspired Spiking Neural Networks for Energy-Efficient Object DetectionabstractBrain-inspired spiking neural networks (SNNs) have the capability of energy-efficient processing of temporal information. However, leveraging the rich dynamic characteristics of SNNs and prior works in artificial neural networks (ANNs) to construct an effective object detection model for visual tasks remains an open question for further exploration. To develop a directly-trained , low energy consumption and high-performance multi-scale SNN model, we propose a novel interpretable object detection framework Multi-scale Spiking Detector (MSD). Initially, we propose a spiking convolutional neuron as a core component of the Optic Nerve Nucleus Block (ONNB), designed to significantly enhance the deep feature extraction capabilities of SNNs. ONNB enables direct training with improved energy efficiency, demonstrating superior performance compared to state-of-the-art ANN-to-SNN conversion and SNN techniques. In addition, we propose a Multi-scale Spiking Detection Framework to emulate the biological response and comprehension of stimuli from different objects. Wherein, spiking multi-scale fusion and the spiking detector are employed to integrate features across different depths and to detect response outcomes, respectively. Our method outperforms state-of-the-art ANN detectors, with only 7.8 M parameters and 6.43 mJ energy consumption. MSD obtains the mean average precision (mAP) of 62.0% and 66.3% on COCO and Gen1 datasets, respectively. Tao Gao 0001, Yisheng An, Ting Chen 0003, Jing Zhang 0052, Yuanbo Wen 0002, Mengkun Liu, Qianxi Zhang |
CVPR | 6 |
| 2025 | Cross-Level Interaction and Intralevel Fusion Network for Remote Sensing Image DehazingabstractExisting approaches have significantly advanced remote sensing image dehazing. However, they often rely on conventional encoder-decoder architectures, leading to prolonged inference times. To this end, we propose a novel cross-level interaction and intra-level fusion network for remote sensing image dehazing (CINet), which shifts the focus from encoder-decoder dependencies to an innovative hierarchical architecture centered on skip connections, leading to competitive dehazing performance with decreased calculating complexity. Furthermore, we introduce a cross-level multi-view interaction module (CMIM) to facilitate effective interactions between features across hierarchical levels, mitigating the information loss commonly caused by repeated down-sampling operations. Meanwhile, we develop an intra-level dual-dimension fusion module (IDFM), which leverages height-wise and width-wise self-attention to capture rich spatial-aware information, enabling robust and efficient intra-level feature fusion. Additionally, we propose a multi-view progressive extraction block (MPEB), which decomposes features into four distinct components and applies convolutions with diverse kernel sizes, groups, and dilation factors. This design promotes progressive feature learning while significantly reducing computational overhead. Extensive experiments conducted on nine publicly available datasets validate the effectiveness and superiority of our proposed model. Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003, Mengkun Liu, Lidong Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | All-in-One Weather-Degraded Image Restoration Via Adaptive Degradation-Aware Self-Prompting ModelabstractExisting approaches for all-in-one weather-degraded image restoration suffer from inefficiencies in leveraging degradation-aware priors, resulting in sub-optimal performance in adapting to different weather conditions. To this end, we develop an adaptive degradation-aware self-prompting model (ADSM) for all-in-one weather-degraded image restoration. Specifically, our model employs the contrastive language-image pre-training model (CLIP) to facilitate the training of our proposed latent prompt generators (LPGs), which represent three types of latent prompts to characterize the degradation type, degradation property and image caption. Moreover, we integrate the acquired degradation-aware prompts into the time embedding of diffusion model to improve degradation perception. Meanwhile, we employ the latent caption prompt to guide the reverse sampling process using the cross-attention mechanism, thereby guiding the accurate image reconstruction. Furthermore, to accelerate the reverse sampling procedure of diffusion model and address the limitations of frequency perception, we introduce a wavelet-oriented noise estimating network (WNE-Net). Extensive experiments conducted on eight publicly available datasets demonstrate the effectiveness of our proposed approach in both task-specific and all-in-one applications. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003 |
IEEE Trans. Multim. | 1 |
| 2024 | Multi-Dimension Queried and Interacting Network for Stereo Image DerainingabstractEliminating the rain degradation in stereo images poses a formidable challenge, which necessitates the efficient exploitation of mutual information present between the dual views. To this end, we devise MQINet, which employs multi-dimension queries and interactions for stereo image deraining. More specifically, our approach incorporates a context-aware dimension-wise queried block (CDQB). This module leverages dimension-wise queries that are independent of the input features and employs global context-aware attention (GCA) to capture essential features while avoiding the entanglement of redundant or irrelevant information. Meanwhile, we introduce an intra-view physics-aware attention (IPA) based on the inverse physical model of rainy images. IPA extracts shallow features that are sensitive to the physics of rain degradation, facilitating the reduction of rain-related artifacts during the early learning period. Furthermore, we integrate a cross-view multi-dimension interacting attention mechanism (CMIA) to foster comprehensive feature interaction between the two views across multiple dimensions. Extensive experimental evaluations demonstrate the superiority of our model over EPRRNet and StereoIRR, achieving respective improvements of 4.18 dB and 0.45 dB in PSNR. Code and models are available at https://github.com/chdwyb/MQINet. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
ICASSP | 1 |
| 2024 | Encoder-Minimal and Decoder-Minimal Framework for Remote Sensing Image DehazingabstractHaze obscures remote sensing images, hindering valuable information extraction. To this end, we propose RSHazeNet, an encoder-minimal and decoder-minimal framework for efficient remote sensing image dehazing. Specifically, regarding the process of merging features within the same level, we develop an innovative module called intra-level transposed fusion module (ITFM). This module employs adaptive transposed self-attention to capture comprehensive context-aware information, facilitating the robust context-aware feature fusion. Meanwhile, we present a cross-level multi-view interaction module (CMIM) to enable effective interactions between features from various levels, mitigating the loss of information due to the repeated sampling operations. In addition, we propose a multi-view progressive extraction block (MPEB) that partitions the features into four distinct components and employs convolution with varying kernel sizes, groups, and dilation factors to facilitate view-progressive feature learning. Extensive experiments demonstrate the superiority of our proposed RSHazeNet. We release the source code and all pre-trained models at https://github.com/chdwyb/RSHazeNet. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
ICASSP | 1 |
| 2024 | Unpaired Photo-realistic Image Deraining with Energy-informed Diffusion ModelabstractExisting unpaired image deraining approaches face challenges in accurately capture the distinguishing characteristics between the rainy and clean domains, resulting in residual degradation and color distortion within the reconstructed images. To this end, we propose an energy-informed diffusion model for unpaired photo-realistic image deraining (UPID-EDM). Initially, we delve into the intricate visual-language priors embedded within the contrastive language-image pre-training model (CLIP), and demonstrate that the CLIP priors aid in the discrimination of rainy and clean images. Furthermore, we introduce a dual-consistent energy function (DEF) that retains the rain-irrelevant characteristics while eliminating the rain-relevant features. This energy function is trained by the non-corresponding rainy and clean images. In addition, we employ the rain-relevance discarding energy function (RDEF) and the rain-irrelevance preserving energy function (RPEF) to direct the reverse sampling procedure of a pre-trained diffusion model, effectively removing the rain streaks while preserving the image contents. Extensive experiments demonstrate that our energy-informed model surpasses the existing unpaired learning approaches in terms of both supervised and no-reference metrics. Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003 |
ACM Multimedia | 1 |
| 2024 | A novel dual-stage progressive enhancement network for single image deraining
Tao Gao 0001, Yuanbo Wen 0002, Jing Zhang 0052, Ting Chen 0003 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Neural Schrödinger bridge for unpaired real-world image deraining
Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003 |
Inf. Sci. | 1 |
| 2024 | A Self-Supplementary and Revised Network for Remote Sensing Object DetectionabstractObject detection is an essential and crucial task in interpretation of optical remote sensing images (RSIs). However, its performance is usually limited due to the complex background and multiscale characteristics of targets. To overcome these limitations, a self-supplementary and revised anchor-free detector is proposed. First, to reduce the computational cost of detection, a partial bottleneck (PBottleneck) structure is designed to efficiently extract multiscale feature information in a lightweight manner. Second, pure spatial feature pyramid network (PSFPN) attaches importance to description of distance and suppresses environmental disturbance by a devised multidirectional distance attention (MDDA) mechanism. In addition, pure fusion strategy (PFS) is created to boost information with no occlusion between various features. Third, toward the multiscale objects issue, self-learning supplementary and revised module (SSRM) is explored to generate more abundant and balanced expression by adaptively incorporating the supplementary and corrected information from adjacent features. Finally, comprehensive experiments are conducted on several publicly available datasets, demonstrating effectiveness of our proposed detector, leading to a new benchmark. Tao Gao 0001, Zixiang Liu, Guiping Wu, Yuanbo Wen 0002, Lidong Liu, Ting Chen 0003, Jing Zhang 0052 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Restoring vision in rain-by-snow weather with simple attention-based sampling cross-hierarchy Transformer
Yuanbo Wen 0002, Tao Gao 0001, Kaihao Zhang, Peng Cheng 0002, Ting Chen 0003 |
Pattern Recognit. | 1 |
| 2024 | From heavy rain removal to detail restoration: A faster and better network
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003 |
Pattern Recognit. | 1 |
| 2024 | Frequency-Oriented Efficient Transformer for All-in-One Weather-Degraded Image RestorationabstractAdverse weather conditions, such as rain, raindrop, snow and haze, consistently degrade images in an unpredictable manner, thereby rendering existing task-specific and task-aligned methods inadequate in addressing this formidable problem. To this end, we investigate the application of Transformer in image restoration and introduce an efficient frequency-oriented method called AIRFormer, which is designed to restore weather-degraded images comprehensively and holistically. Specifically, we identify that the initial self-attention mechanism exhibits distinctive properties akin to a low-pass filter. Therefore, we construct a frequency-guided Transformer encoder by incorporating wavelet-based prior information to guide the extraction of image features. Additionally, considering the non-specific frequency characteristics of self-attention in the later stages, we develop a frequency-refined Transformer decoder that incorporates learnable task-specific queries across spatial dimensions, channel dimensions, and wavelet domains. To facilitate the training of our proposed method, we curate a comprehensive benchmark dataset named AIR40K that, encompasses a wide range of challenging scenarios. Extensive experimental evaluations demonstrate the superiority of our AIRFormer over both task-aligned and all-in-one methods across 15 publicly available datasets. Notably, AIRFormer achieves the best trade-off between the inference time and quality of reconstructed image, comparing with existing methods such as TransWeather and Restormer. The source code, dataset and pre-trained models will be available at https://github.com/chdwyb/AIRFormer. Tao Gao 0001, Yuanbo Wen 0002, Kaihao Zhang, Jing Zhang 0052, Ting Chen 0003, Lidong Liu, Wenhan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Attention-Free Global Multiscale Fusion Network for Remote Sensing Object DetectionabstractRemote sensing object detection (RSOD) encounters challenges in complex backgrounds and small object detection, which are interconnected and unable to address separately. To this end, we propose an attention-free global multiscale fusion network (AGMF-Net). Initially, we present a spatial bias module (SBM) to obtain long-range dependencies as a part of our proposal global information extraction module (GIEM). GIEM efficiently captures the global information, overcoming challenges posed by complex backgrounds. Moreover, we propose multitask enhanced structure (MES) and multitask feature pretreatment (MFP) to enhance the feature representation of multiscale targets, while eliminating the interference from complex backgrounds. In addition, an efficient context decoupled detector (ECDD) is presented to provide distinct features for regression and classification tasks, aiming to improve the efficiency of RSOD. Extensive experiments demonstrate that our proposed method achieves superior performance compared with the state-of-the-art detectors. Specifically, AGMF-Net obtains the mean average precision (mAP) of 73.2%, 92.03%, 95.21%, and 94.30% on detection in optical remote sensing images (DIOR), high resolution remote sensing detection (HRRSD), Northwestern Polytechnical University Very High Resolution-10 (NWPU VHR-10), and RSOD datasets, respectively. Tao Gao 0001, Yuanbo Wen 0002, Ting Chen 0003, Qianqian Niu, Zixiang Liu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Encoder-Free Multiaxis Physics-Aware Fusion Network for Remote Sensing Image DehazingabstractCurrent methods for remote sensing image dehazing confront noteworthy computational intricacies and yield suboptimal dehazed outputs, thereby circumscribing their pragmatic applicability. To this end, we propose EMPF-Net, a novel encoder-free multi-axis physics-aware fusion network that exhibits both light-weighted characteristics and computational efficiency. In our pipeline, we contend that conventional u-shaped networks allocate substantial computational resources to encode haze-degraded features, which play a subordinate role in the reconstruction process. Consequently, our encoder stages solely incorporate down-sampling operations. To improve the representation efficiency and enhance the generalization capabilities, we devise a multi-axis partial queried learning block (MPQLB) that primarily concentrates on learning dimension-wise queries, instead of relying solely on strictly-correlated content of the input features. Furthermore, we augment the reconstruction procedure by incorporating ground truth supervision into each stage via a supervised cross-scale transposed attention module (SCTAM). It calculates attention maps under the guidance of clean images, thereby suppressing less informative features to propagate to the subsequent level. In addition, to address the challenge of ineffective intral-level feature fusion, which result in insufficient elimination of haze-degraded information and negatively impact the quality of reconstructed images, we introduce a physics-aware intra-level fusion module (PIFM). This module harnesses a physical inversion model to facilitate the intra-level feature interaction and alleviate the interference of dehazing-irrelevant information. Our proposed EMPF-Net is evaluated on 12 publicly available datasets, and the experimental results substantiate our superiority in terms of both metrical scores and visual quality, despite being equipped with a modest parameter count of 300 K. Our approach is readily accessible at https://github.com/chdwyb/EMPF-Net. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |