Wenxuan Fang 0001

dblp:320/0402-1 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0004-3435-3426ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Curriculum adaptation for one-stream RGB-T tracking
Xiantao Hu, Fansheng Zeng, Bineng Zhong 0001, Zhangyong Tang, Wenxuan Fang 0001, Jun Li 0027, Ying Tai, Jian Yang 0003
Pattern Recognit.5
2026 WeatherCycle: Unpaired Multi-Weather Restoration via Color Space Decoupled Cycle Learning
abstract
Unsupervised image restoration under multi-weather conditions remains a fundamental yet underexplored challenge. While existing methods often rely on task-specific physical priors, their narrow focus limits scalability and generalization to diverse real-world weather scenarios. In this work, we propose WeatherCycle, a unified unpaired framework that reformulates weather restoration as a bidirectional degradation-content translation cycle, guided by degradation-aware curriculum regularization. At its core, WeatherCycle employs alumina-chroma decompositionstrategy to decouple degradation from content without modeling complex weather, enabling domain conversion between degraded and clean images. To model diverse and complex degradations, we propose aLumina Degradation Guidance Module(LDGM), which learns luminance degradation priors from a degraded image pool and injects them into clean images via frequency-domain amplitude modulation, enabling controllable and realistic degradation modeling. Additionally, we incorporate aDifficulty-Aware Contrastive Regularization(DACR) module that identifies hard samples via a CLIP-based classifier and enforces contrastive alignment between hard samples and restored features to enhance semantic consistency and robustness. Extensive experiments across serve multi-weather datasets, demonstrate that our method achieves state-of-the-art performance among unsupervised approaches, with strong generalization to complex weather degradations.
Wenxuan Fang 0001, Jiangwei Weng, Jianjun Qian, Jian Yang 0003, Jun Li 0027
IEEE Trans. Circuits Syst. Video Technol.1
2026 DC2MNet: Lightweight and Efficient Discrete Cosine Channel Modulation Network for Image Restoration
abstract
Image restoration aims to remove degradation factors (such as blur, snow e.g.) from the damaged image and reconstruct a clean image. Although some methods seek solutions from the frequency domain and are proven to be effective, they are still faced two challenges: (i) Degradation blurs cannot be removed well, and (ii) Inverse transform in frequency domain is computationally expensive. To this end, we propose a lightweight and efficient Discrete Cosine Channel Modulation Network (DC2MNet) for recovering images of multiple degraded conditions from the frequency and spatial perspectives. Specifically, we propose a Discrete Cosine Channel Modulation (DCCM) module to extract the most informative lowest-frequency components of features, and subsequently utilize the channel modulation to reconstruct the global structure of the corresponding feature, avoiding inverse transform in high-dimensional spaces. Furthermore, to effectively remove degradation, we propose a Spatial Mask Modulation (SMM) module to suppress degradation blurs in high-frequency features and emphasize local details that are beneficial to image restoration via pixel-level spatial attention. Finally, we embed the DCCM module and SMM module into the Channel Spatial Modulation Block (CSMB) to form the basic component of DC2MNet, which achieves SOTA performance on various restoration tasks through extensive experiments, including image dehazing, deraining, desnowing and multi-weather restoration. The code and pre-trained models will be open source in this repository.
Guoqing Zhang 0002, Wenxuan Fang 0001, Yupeng Shang, Yuhui Zheng, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.2
2026 DVDPEC: Driving-Video Dehazing via Position Embedding-Based Codebook
abstract
Despite significant progress in real-world image dehazing, efficiently generating high-fidelity, haze-free videos (especially in driving scenarios) remains challenging. Existing methods generally extend image dehazing techniques to videos by employing pre-trained single image dehazing models for preprocessing followed by refinement stages. However, this disjointed two-stage process often leads to unrealistic textures and loss of detail, as it fails to leverage large amounts of high-quality images for prior learning and the subsequent refinement struggles to correct temporal inconsistencies across frames introduced in the first stage. To address these issues, we propose DVDPEC: a Driving Video Dehazing framework utilizing a Position Embedding-based (PE-based) Codebook and a novel Flow Selective Block (FSB). The PE-based codebook stores fine-grained, spatially aware textural information specific to driving videos and leverages implicit positional embeddings for precise, position-aware codebook matching. This enables accurate prior retrieval and improves dehazing results. The FSB aggregates information from adjacent frames by dynamically combining both image flow and prior flow, effectively mitigating flow estimation ambiguities caused by haze. It enhances information fusion across frames, leading to more coherent and visually appealing dehazed videos. Extensive experiments demonstrate that DVDPEC achieves state-of-the-art performance on real-world driving video dehazing tasks, significantly enhancing texture preservation and visual fidelity.
Yu Zheng 0036, Wenxuan Fang 0001, Xiantao Hu, Junkai Fan, Jiangwei Weng, Jun Li 0027, Kai Zhang 0008, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2025 Guided Real Image Dehazing Using YCbCr Color Space
abstract
Image dehazing, particularly with learning-based methods, has gained significant attention due to its importance in real-world applications. However, relying solely on the RGB color space often fall short, frequently leaving residual haze. This arises from two main issues: the difficulty in obtaining clear textural features from hazy RGB images and the complexity of acquiring real haze/clean image pairs outside controlled environments like smoke-filled scenes. To address these issues, we first propose a novel Structure Guided Dehazing Network (SGDN) that leverages the superior structural properties of YCbCr features over RGB. It comprises two key modules: Bi-Color Guidance Bridge (BGB) and Color Enhancement Module (CEM). BGB integrates a phase integration module and an interactive attention module, utilizing the rich texture features of the YCbCr space to guide the RGB space, thereby recovering clearer features in both frequency and spatial domains. To maintain tonal consistency, CEM further enhances the color perception of RGB features by aggregating YCbCr channel information. Furthermore, for effective supervised learning, we introduce a Real-World Well-Aligned Haze dataset, which includes a diverse range of scenes from various geographical regions and climate conditions. Experimental results demonstrate that our method surpasses existing state-of-the-art methods across multiple real-world smoke/haze datasets.
Wenxuan Fang 0001, Junkai Fan, Yu Zheng 0036, Jiangwei Weng, Ying Tai, Jun Li 0027
AAAI1
2025 Cross-modal Gaussian Localization Distillation for Optical Information guided SAR Object Detection
abstract
Synthetic Aperture Radar (SAR) images contain a dense clutter of objects that can be better characterized using bounding boxes with angles. However, accurately detecting the angles of objects remains challenging due to the imaging mechanism of SAR. To address this issue, we propose a novel knowledge distillation method called cross-modal Gaussian Localization Distillation (GaLD). It aims to improve SAR object detection performance by utilizing the angle information from optical images. Specifically, we convert the oriented bounding box into a Gaussian distribution and design a Gaussian Angle Distillation (GAD) loss function to align the angle information between optical and SAR images. In addition, to mitigate the negative impact of low-quality angle information on the network, we design an Adaptive Weighting Strategy (AWS) to guide the student network to prioritize high-quality angle information. The lack of high-quality oriented labels and objects in the OGSOD-1.0 dataset has hindered the progress in related fields. Therefore, we have added high-quality oriented labels and images to the OGSOD-1.0 dataset to construct a new dataset. Extensive experiments demonstrate the effectiveness and superiority of our proposed GaLD over existing methods. The dataset and code are available at: https://github.com/wchao0601/GaLD.
Lei Luo 0001, Wenxuan Fang 0001, Jian Yang 0003
ICASSP3
2025 MSOD: A Large-Scale Multiscene Dataset and a Novel Diagonal-Geometry Loss for SAR Object Detection
abstract
Synthetic Aperture Radar (SAR) has attracted significant attention due to its excellent all-weather imaging capabilities. However, SAR image object detection methods face two major challenges: 1) Most existing datasets are small in volume and single in category and scene. 2) Existing IoU-based loss functions cannot fully capture the relationship between prediction and target bounding boxes. To further advance the development of the SAR object detection method, we construct a large-scale multi-scene SAR object detection dataset called MSOD. It comprises three distinct scenarios, containing 40K images and about 1M instances of interest classified into six categories. In addition, we propose a novel diagonal-based similarity loss, Diagonal-Geometry IoU (DGIoU), to optimize the performance of SAR object detection by measuring the similarity between the diagonal of the prediction and target boxes. Specifically, we equivalently represent a rectangular box as a diagonal, and then define DGIoU based on the similarity of a set of sampling points between the diagonals of the predicted box and the target box. DGIoU effectively characterizes the difference between the predicted box and the target box, particularly in box inclusion and separation cases, resulting in improved localization accuracy. Numerous experimental results demonstrate that MSOD is closer to practical application and more challenging than existing SAR image datasets, and serves as a strong benchmark for evaluating the effectiveness of various IoU loss functions. The dataset and code are available at: https://github.com/wchao0601/MSOD-DGIoU.
Wenxuan Fang 0001, Xiang Li 0041, Jian Yang 0003, Lei Luo 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Rethinking general time series analysis from a frequency domain perspective
Jili Fan, Jiayu Fang, Wenxuan Fang 0001, Min Xia 0002
Knowl. Based Syst.4
2024 SDBAD-Net: A Spatial Dual-Branch Attention Dehazing Network Based on Meta-Former Paradigm
abstract
Image dehazing is an emblematical low-level vision task that aims at restoring haze-free images from haze images. Recently, some methods adopts deep learning techniques to rebuild haze-free images. However, in real-world scenarios, complex degradation of captured images and non-uniform spatial distributions of haze will significantly weaken the generalization ability of these models. Accordingly, we propose a novel Spatial Dual-Branch Attention Dehazing network (SDBAD-Net) based on the Meta-Former paradigm for end-to-end dehazing. Specifically, we firstly design a robust Spatial Dual-Branch Attention (SDBA) module to filter the haze distribution features from different densities, which is suitable for both uniform and non-uniform situations. Secondly, we introduce a Structural Features Supplementary (SFS) module to dynamically fuse the contextual structural features in a nonlinear manner, so as to correct the image distortion caused by the lack of structural details. Finally, the quantitative and qualitative experiments are carried out on two challenging datasets, and the results show that our method outperforms most of state-of-the-art algorithms with fewer parameters and faster speed, especially surpassing FFA-Net with only 50% parameters and 7% computational costs. In addition, we ulteriorly explore its performance on object detection in foggy weather with our model on the challenging Real-world Task-driven Testing Set (RTTS), and the surprising results further prove the robustness and wide-applicability of our method.
Guoqing Zhang 0002, Wenxuan Fang 0001, Yuhui Zheng, Ruili Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Optimizing Attention in a Transformer for Multihorizon, Multienergy Load Forecasting in Integrated Energy Systems
abstract
Accurate forecasting of multienergy loads is essential for designing, operating, scheduling, and managing integrated energy systems (IESs). Recent research suggests that transformer models have the potential to improve long-sequence predictions. However, existing transformer models often emphasize capturing temporal dependencies while neglecting crucial dependencies among different variables necessary for multienergy load forecasting. Moreover, transformer models encounter challenges related to quadratic time complexity and significant memory usage, which hinder their direct applicability to tasks involving long-sequence, multienergy load forecasting. To tackle these challenges, we propose a model called DTformer and apply it to the task of multihorizon, multienergy load forecasting in IES. Within DTformer, we employ patch embedding to convert the input multienergy load sequences into a 3-D vector array, preserving both temporal and variable information. Subsequently, we propose the temporal top windowed attention (TWA) module and the dual variable attention module to handle extended temporal dependencies and intervariable dependencies. Importantly, the computational complexity and memory requirements of the TWA model are regulated at a level of$O(N^\frac{4}{3})$. Through extensive experimentation, we found that our DTformer surpasses baseline models in terms of performance using the IES dataset sourced from Arizona State University's Tempe campus.
Jili Fan, Min Xia 0002, Wenxuan Fang 0001, Jun Liu 0100
IEEE Trans. Ind. Informatics4