Fangwei Hao

dblp:302/7568 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Du-CIPT: Dual Cross-Modal Interactive Pyramid Transformer for RGB-Thermal salient object detection and segmentation
Jiesheng Wu, Ji Du, Fangwei Hao, Jiankang Hong
Signal Process. Image Commun.3
2026 RA-COD: Retrieval-Augmented Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) is pivotal for segmenting objects that seamlessly blend into their surroundings. While prior endeavors demonstrate impressive performance through training on predefined labels, they heavily rely on labor-intensive data annotation and struggle to adapt to open-world scenarios. In this light, we propose RA-COD, a training-free paradigm that enables COD by retrieving the most similar samples from the prototype repository. The efficacy of RA-COD hinges on 1) capturing the nuanced resemblance between objects and their environments and 2) excelling in dense prediction tasks. To achieve (1), the crux lies in ensuring diversity and discriminability within the prototype repository. In this context, we propose GenPro, an automated pipeline for crafting Generative Prototypes. GenPro integrates a range of foundation models, including the Diffusion Model, Vision-Language Model, Segment Anything Model (SAM), and DINOv2, in a complementary manner that synergistically generates diverse and distinguishable prototype samples. To achieve (2), we propose C2F to retrieve camouflaged objects in a Coarse-to-Fine regime. We commence with pixel-level retrieval in the feature space, which generates a coarse mask that effectively captures class discrimination and object localization. Further refinement is achieved by extracting bounding boxes from this coarse mask to prompt SAM in generating mask proposals for region-level retrieval. Evaluations on four benchmarks showcase that RA-COD achieves state-of-the-art performance compared to existing training-free methods.
Ji Du, Jiesheng Wu, Desheng Kong, Fangwei Hao, Jing Xu 0008, Ping Li 0016
IEEE Trans. Image Process.4
2025 Shift the Lens: Environment-Aware Unsupervised Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) seeks to distinguish objects from their highly similar backgrounds. Existing work has essentially focused on isolating camouflaged objects from the environment, demonstrating ever-improving performance but at the cost of extensive annotations and complex optimizations. In this paper, we diverge from this paradigm and shift the lens to isolating the salient environment from the camouflaged object. We introduce EASE, an Environment-Aware unSupErvised COD framework that identifies the environment by referencing an environment prototype library and detects camouflaged objects by inverting the retrieved environmental features. Specifically, our approach (DiffPro) uses large multimodal models, diffusion models, and vision-foundation models to construct the environment prototype library. To retrieve environments from the library and refrain from confusing foreground and background, we incorporate three retrieval schemes: Kernel Density Estimation-based Adaptive Threshold (KDE-AT), Global-to-Local pixel-level retrieval (G2L), and Self-Retrieval (SR). Our experiments demonstrate significant improvements over current unsupervised methods, with EASE achieving an average gain of over 10% on the COD10K dataset. When integrated with SAM, EASE surpasses prompt-based segmentation approaches and performs competitively with state-of-the-art fully-supervised methods. Code is available at https://github.com/xiaohainku/EASE.
Ji Du, Fangwei Hao, Mingyang Yu 0001, Desheng Kong, Jiesheng Wu, Jing Xu 0008, Ping Li 0016
CVPR2
2025 Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
abstract
At the core of Camouflaged Object Detection (COD) lies segmenting objects from their highly similar surroundings. Previous efforts navigate this challenge primarily through image-level modeling or annotation-based optimization. Despite advancing considerably, this commonplace practice hardly taps valuable dataset-level contextual information or relies on laborious annotations. In this paper, we propose RISE, a RetrIeval SElf-augmented paradigm that exploits the entire training dataset to generate pseudo-labels for single images, which could be used to train COD models. RISE begins by constructing prototype libraries for environments and camouflaged objects using training images (without ground truth), followed by K-Nearest Neighbor (KNN) retrieval to generate pseudo-masks for each image based on these libraries. It is important to recognize that using only training images without annotations exerts a pronounced challenge in crafting high-quality prototype libraries. In this light, we introduce a Clustering-then-Retrieval (CR) strategy, where coarse masks are first generated through clustering, facilitating subsequent histogram-based image filtering and cross-category retrieval to produce high-confidence prototypes. In the KNN retrieval stage, to alleviate the effect of artifacts in feature maps, we propose Multi-View KNN Retrieval (MVKR), which integrates retrieval results from diverse views to produce more robust and precise pseudo-masks. Extensive experiments demonstrate that RISE outperforms state-of-the-art unsupervised and prompt-based methods. Code is available at https://github.com/xiaohainku/RISE.
Ji Du, Xin Wang 0118, Fangwei Hao, Mingyang Yu 0001, Chunyuan Chen, Jiesheng Wu, Jing Xu 0008, Ping Li 0016
ICCV3
2025 Large coordinate attention network for lightweight image super-resolution
Fangwei Hao, Jiesheng Wu, Haotian Lu 0003, Ji Du, Jing Xu 0008, Xiaoxuan Xu
Eng. Appl. Artif. Intell.1
2025 Information sparsity guided transformer for multi-modal medical image super-resolution
Haotian Lu 0003, Jie Mei 0004, Fangwei Hao, Jing Xu 0008
Expert Syst. Appl.5
2025 Towards context-aware convolutional network for image restoration
Fangwei Hao, Ji Du, Weiyun Liang, Jing Xu 0008, Xiaoxuan Xu
Knowl. Based Syst.1
2025 UpGen: Unleashing Potential of Foundation Models for Training-Free Camouflage Detection via Generative Models
abstract
Camouflaged Object Detection (COD) aims to segment objects resembling their environment. To address the challenges of extensive annotations and complex optimizations in supervised learning, recent prompt-based segmentation methods excavate insightful prompts from Large Vision-Language Models (LVLMs) and refine them using various foundation models. These are subsequently fed into the Segment Anything Model (SAM) for segmentation. However, due to the hallucinations of LVLMs and insufficient image-prompt interactions during the refinement stage, these prompts often struggle to capture well-established class differentiation and localization of camouflaged objects, resulting in performance degradation. To provide SAM with more informative prompts, we present UpGen, a pipeline that prompts SAM with generative prompts without requiring training, marking a novel integration of generative models with LVLMs. Specifically, we propose the Multi-Student-Single-Teacher (MSST) knowledge integration framework to alleviate hallucinations of LVLMs. This framework integrates insights from multiple sources to enhance the classification of camouflaged objects. To enhance interactions during the prompt refinement stage, we are the first to leverage generative models on real camouflage images to produce SAM-style prompts without fine-tuning. By capitalizing on the unique learning mechanism and structure of generative models, we effectively enable image-prompt interactions and generate highly informative prompts for SAM. Our extensive experiments demonstrate that UpGen outperforms weakly-supervised models and its SAM-based counterparts. We also integrate our framework into existing weakly-supervised methods to generate pseudo-labels, resulting in consistent performance gains. Moreover, with minor adjustments, UpGen shows promising results in open-vocabulary COD, referring COD, salient object detection, marine animal segmentation, and transparent object segmentation.
Ji Du, Jiesheng Wu, Desheng Kong, Weiyun Liang, Fangwei Hao, Jing Xu 0008, Grace Guiling Wang, Ping Li 0016
IEEE Trans. Image Process.5
2025 Boosting Foreground-Background Disentanglement for Camouflaged Object Detection
abstract
In nature, certain objects exhibit patterns that closely resemble their backgrounds, a phenomenon commonly referred to as Camouflaged Object Detection (COD). We argue that existing COD approaches often suffer from insufficient discriminability for these objects, which we attribute to a lack of effective disentangling of foreground and background representations. To address this, we propose a novel Foreground-Background Disentanglement Network (FBD-Net) that enhances foreground-background disentanglement learning to improve discriminability. Specifically, we design an Edge-guided Foreground-Background Decoupling (EFBD) module, which facilitates the separated learning of foreground and background representations. Additionally, we introduce the Foreground-Background Representation Disentangling Head (DisHead) to further boost the discriminative power of the model. The DisHead consists of two objectives: the Edge Objective and the FoBa Objective. Furthermore, we propose three complementary modules: the Context Aggregation Module (CAM) for initial coarse object detection, the Scale-Interaction Enhanced Pyramid (SIEP) for multi-scale information extraction, and the Cross-Stage Adaptive Fusion (CSAF) module for subtle clue accumulation. Extensive experiments demonstrate that both our CNN-based and Transformer-based FBD-Nets outperform 26 state-of-the-art COD methods across four public datasets. Codes will be released on https://github.com/TomorrowJW/FBD-Net-COD .
Jiesheng Wu, Fangwei Hao, Jing Xu 0008
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Lightweight blueprint residual network for single image super-resolution
Fangwei Hao, Jiesheng Wu, Weiyun Liang, Jing Xu 0008, Ping Li 0016
Expert Syst. Appl.1
2024 Transformer Fusion and Pixel-Level Contrastive Learning for RGB-D Salient Object Detection
abstract
Current RGB-D salient object detection (RGB-D SOD) methods mainly develop a generalizable model trained by binary cross-entropy (BCE) loss based on convolutional or Transformer backbones. However, they usually exploit convolutional modules to fuse multi-modality features, with little attention paid to capturing the long-range multi-modality interactions for feature fusion. Furthermore, BCE loss does not explicitly explore intra- and inter-pixel relationships in a joint embedding space. To address these issues, we propose a cross-modality interaction parallel-transformer (CIPT) module, which better captures the long-range multi-modality interactions, generating more comprehensive fusion features. Besides, we propose a pixel-level contrastive learning (PCL) method that improves inter-pixel discrimination and intra-pixel compactness, resulting in a well-structured embedding space and a better saliency detector. Specifically, we propose an asymmetric network (TPCL) for RGB-D SOD, which consists of a Swin V2 Transformer-based backbone and a designed lightweight backbone (LDNet). Moreover, an edge-guided module and a feature enhancement (FE) module are proposed to refine the learned fusion features. Extensive experiments demonstrate that our method achieves excellent performance against 15 state-of-the-art methods on seven public datasets. We expect our work to facilitate the exploration of applying Transformer and contrastive learning for RGB-D SOD tasks.
Jiesheng Wu, Fangwei Hao, Weiyun Liang, Jing Xu 0008
IEEE Trans. Multim.2
2023 Mask-and-Edge Co-Guided Separable Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) involves segmenting objects that share similar patterns, such as color and texture, with their surroundings. Current methods typically employ multiple well-designed modules or rely on edge cues to learn object feature representations for COD. However, these methods still struggle to capture the discriminative semantics between camouflaged objects (foreground) and background, possibly generating blurry prediction maps. To address these limitations, we propose a novel mask-and-edge co-guided separable network (MECS-Net) for COD that leverages both edge and mask cues to learn more discriminative representations and improve detection performance. Specifically, we design a mask-and-edge co-guided separable attention (MECSA) module, which consists of three flows for separately capturing edge, foreground, and background semantics. In addition, we propose a multi-scale enhancement fusion (MEF) module to aggregate multi-scale features of objects. The predictions are decoded in a top-down manner. Extensive experiments and visualizations demonstrate that our CNN-based and Transformer-based MECS-Net outperform 13 state-of-the-art methods on four popular COD datasets. Codes and results are availablehttps://github.com/TomorrowJW/MECS-Net-COD$\ast$.
Jiesheng Wu, Weiyun Liang, Fangwei Hao, Jing Xu 0008
IEEE Signal Process. Lett.3
2022 Efficient residual attention network for single image super-resolution
Fangwei Hao, Taiping Zhang, Linchang Zhao, Yuan Yan Tang
Appl. Intell.1
2021 Channel Hourglass Residual Network For Single Image Super-Resolution
abstract
Deep convolutional neural networks (CNNs) for Super-Resolution (SR) from low-resolution (LR) images have achieved remarkable reconstruction performance with the utilization of residual networks and visual attention mechanism. However, the existing single image super-resolution (SISR) methods with deeper or wider network architectures encounter module representation bottleneck and neglect module efficiency in real-world applications. To solve these issues, in this paper, we design channel hourglass residual structure (CHRS) consisted of several nested residual modules for reducing parameters and extracting more representational features. Furthermore, we integrate channel attention (CA) mechanism into CHRS to generate channel hourglass residual block (CHRB) which can be easily extended to other methods for improving performance. We also propose channel hourglass residual network (CHRN) which not only pays attention to network learning efficiency but also learns more discriminative expressions. Extensive experiments demonstrate the effectiveness of our CHRN and the generalization ability of our CHRB.
Fangwei Hao, XinDi Ma, Taiping Zhang, Yuan Yan Tang
IJCNN1