Ji Du

dblp:320/2099 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0001-9388-8146ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Boosting semi-supervised camouflaged object detection with representative samples and better labels
Chunyuan Chen, Weiyun Liang, Ji Du, Xinjian Wei, Jing Xu 0008, Frank Jiang 0001
Knowl. Based Syst.3
2026 Du-CIPT: Dual Cross-Modal Interactive Pyramid Transformer for RGB-Thermal salient object detection and segmentation
Jiesheng Wu, Ji Du, Fangwei Hao, Jiankang Hong
Signal Process. Image Commun.2
2026 Learn From Examples: In-Context Learning for Camouflaged Object Detection
abstract
Recently, new paradigms of camouflaged object detection (COD), such as referring COD (Ref-COD) and collaborative COD (Co-COD), have been proposed to enhance task performance. However, there remains a lack of in-depth exploration of how to utilize reference information more effectively. In this paper, we introduce in-context learning camouflaged object detection (ICL-COD) as a novel paradigm of COD, which leverages camouflaged image samples and their corresponding annotations as visual examples to guide the model in better perceiving camouflage and recognizing camouflaged objects. We propose the ICL-Camo network, with the design of a context mining module (CMM) to mine fine-grained contextual information contained in the visual examples, and a context guiding module (CGM) that utilizes the contextual information mined from the examples as guidance to shift the attention of the target image features on potential camouflaged regions, thus enhancing its perception of camouflaged objects. Extensive experiments conducted on the COD benchmarks and other relevant tasks demonstrate the effectiveness of our proposed ICL-COD paradigm and ICL-Camo network. Code and results are available at: https://github.com/h0t-zer0/ICL-Camo.
Chunyuan Chen, Weiyun Liang, Ji Du, Jing Xu 0008, Ping Li 0016, Grace Guiling Wang
IEEE Trans. Image Process.3
2026 RA-COD: Retrieval-Augmented Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) is pivotal for segmenting objects that seamlessly blend into their surroundings. While prior endeavors demonstrate impressive performance through training on predefined labels, they heavily rely on labor-intensive data annotation and struggle to adapt to open-world scenarios. In this light, we propose RA-COD, a training-free paradigm that enables COD by retrieving the most similar samples from the prototype repository. The efficacy of RA-COD hinges on 1) capturing the nuanced resemblance between objects and their environments and 2) excelling in dense prediction tasks. To achieve (1), the crux lies in ensuring diversity and discriminability within the prototype repository. In this context, we propose GenPro, an automated pipeline for crafting Generative Prototypes. GenPro integrates a range of foundation models, including the Diffusion Model, Vision-Language Model, Segment Anything Model (SAM), and DINOv2, in a complementary manner that synergistically generates diverse and distinguishable prototype samples. To achieve (2), we propose C2F to retrieve camouflaged objects in a Coarse-to-Fine regime. We commence with pixel-level retrieval in the feature space, which generates a coarse mask that effectively captures class discrimination and object localization. Further refinement is achieved by extracting bounding boxes from this coarse mask to prompt SAM in generating mask proposals for region-level retrieval. Evaluations on four benchmarks showcase that RA-COD achieves state-of-the-art performance compared to existing training-free methods.
Ji Du, Jiesheng Wu, Desheng Kong, Fangwei Hao, Jing Xu 0008, Ping Li 0016
IEEE Trans. Image Process.1
2025 Shift the Lens: Environment-Aware Unsupervised Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) seeks to distinguish objects from their highly similar backgrounds. Existing work has essentially focused on isolating camouflaged objects from the environment, demonstrating ever-improving performance but at the cost of extensive annotations and complex optimizations. In this paper, we diverge from this paradigm and shift the lens to isolating the salient environment from the camouflaged object. We introduce EASE, an Environment-Aware unSupErvised COD framework that identifies the environment by referencing an environment prototype library and detects camouflaged objects by inverting the retrieved environmental features. Specifically, our approach (DiffPro) uses large multimodal models, diffusion models, and vision-foundation models to construct the environment prototype library. To retrieve environments from the library and refrain from confusing foreground and background, we incorporate three retrieval schemes: Kernel Density Estimation-based Adaptive Threshold (KDE-AT), Global-to-Local pixel-level retrieval (G2L), and Self-Retrieval (SR). Our experiments demonstrate significant improvements over current unsupervised methods, with EASE achieving an average gain of over 10% on the COD10K dataset. When integrated with SAM, EASE surpasses prompt-based segmentation approaches and performs competitively with state-of-the-art fully-supervised methods. Code is available at https://github.com/xiaohainku/EASE.
Ji Du, Fangwei Hao, Mingyang Yu 0001, Desheng Kong, Jiesheng Wu, Jing Xu 0008, Ping Li 0016
CVPR1
2025 Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
abstract
At the core of Camouflaged Object Detection (COD) lies segmenting objects from their highly similar surroundings. Previous efforts navigate this challenge primarily through image-level modeling or annotation-based optimization. Despite advancing considerably, this commonplace practice hardly taps valuable dataset-level contextual information or relies on laborious annotations. In this paper, we propose RISE, a RetrIeval SElf-augmented paradigm that exploits the entire training dataset to generate pseudo-labels for single images, which could be used to train COD models. RISE begins by constructing prototype libraries for environments and camouflaged objects using training images (without ground truth), followed by K-Nearest Neighbor (KNN) retrieval to generate pseudo-masks for each image based on these libraries. It is important to recognize that using only training images without annotations exerts a pronounced challenge in crafting high-quality prototype libraries. In this light, we introduce a Clustering-then-Retrieval (CR) strategy, where coarse masks are first generated through clustering, facilitating subsequent histogram-based image filtering and cross-category retrieval to produce high-confidence prototypes. In the KNN retrieval stage, to alleviate the effect of artifacts in feature maps, we propose Multi-View KNN Retrieval (MVKR), which integrates retrieval results from diverse views to produce more robust and precise pseudo-masks. Extensive experiments demonstrate that RISE outperforms state-of-the-art unsupervised and prompt-based methods. Code is available at https://github.com/xiaohainku/RISE.
Ji Du, Xin Wang 0118, Fangwei Hao, Mingyang Yu 0001, Chunyuan Chen, Jiesheng Wu, Jing Xu 0008, Ping Li 0016
ICCV1
2025 MambaCOD: Cross-Modal Mamba Fusion Network with Adapter Tuning for RGB-D Camouflaged Object Detection
Jiesheng Wu, Lizheng Zhang, Fuyu Zhang, Biao Jie, Ji Du
PRCV (16)7
2025 Large coordinate attention network for lightweight image super-resolution
Fangwei Hao, Jiesheng Wu, Haotian Lu 0003, Ji Du, Jing Xu 0008, Xiaoxuan Xu
Eng. Appl. Artif. Intell.4
2025 Towards context-aware convolutional network for image restoration
Fangwei Hao, Ji Du, Weiyun Liang, Jing Xu 0008, Xiaoxuan Xu
Knowl. Based Syst.2
2025 UpGen: Unleashing Potential of Foundation Models for Training-Free Camouflage Detection via Generative Models
abstract
Camouflaged Object Detection (COD) aims to segment objects resembling their environment. To address the challenges of extensive annotations and complex optimizations in supervised learning, recent prompt-based segmentation methods excavate insightful prompts from Large Vision-Language Models (LVLMs) and refine them using various foundation models. These are subsequently fed into the Segment Anything Model (SAM) for segmentation. However, due to the hallucinations of LVLMs and insufficient image-prompt interactions during the refinement stage, these prompts often struggle to capture well-established class differentiation and localization of camouflaged objects, resulting in performance degradation. To provide SAM with more informative prompts, we present UpGen, a pipeline that prompts SAM with generative prompts without requiring training, marking a novel integration of generative models with LVLMs. Specifically, we propose the Multi-Student-Single-Teacher (MSST) knowledge integration framework to alleviate hallucinations of LVLMs. This framework integrates insights from multiple sources to enhance the classification of camouflaged objects. To enhance interactions during the prompt refinement stage, we are the first to leverage generative models on real camouflage images to produce SAM-style prompts without fine-tuning. By capitalizing on the unique learning mechanism and structure of generative models, we effectively enable image-prompt interactions and generate highly informative prompts for SAM. Our extensive experiments demonstrate that UpGen outperforms weakly-supervised models and its SAM-based counterparts. We also integrate our framework into existing weakly-supervised methods to generate pseudo-labels, resulting in consistent performance gains. Moreover, with minor adjustments, UpGen shows promising results in open-vocabulary COD, referring COD, salient object detection, marine animal segmentation, and transparent object segmentation.
Ji Du, Jiesheng Wu, Desheng Kong, Weiyun Liang, Fangwei Hao, Jing Xu 0008, Grace Guiling Wang, Ping Li 0016
IEEE Trans. Image Process.1
2025 HRC-Net: Learning Visual Hypothesis, Representative, and Collaboration for Multi-Domain Image Inpainting
abstract
Multi-domain image inpainting utilizes complementary contextual information from auxiliary domain images to restore corrupted regions. While existing methods reconstruct auxiliary images to provide additional guidance, they face fundamental limitations: recovered pixels with complex patterns often lack representative details, while oversimplified patterns offer insufficient contextual information. To address these challenges, we propose HRC-Net, a novel framework incorporating three generative sub-networks for the comprehensive image inpainting task. Our architecture consists of: (1) A Hypothesis Sub-network that enables robust samplings of pixel-wise hypotheses from multi-domain inputs; (2) A Representative Sub-network that learns to score hypothesis quality based on contextual relevance; and (3) a Collaboration Sub-network that optimizes adaptive fusion kernels to integrate the most pertinent details. Together, these components model the joint distribution of representative scores and convolutional kernels, fostering a precise interaction between auxiliary hypotheses and target image corruption to meticulously repair the target image. Extensive evaluations across multiple benchmark datasets demonstrate HRC-Net's superior performance, significantly outperforming state-of-the-art methods in both quantitative metrics and visual quality.
Xin Wang 0118, Di Lin 0002, Wanchao Su, Ji Du, Jie Zhang 0090, Haotian Dong, Ke Xu 0010, Qing Guo 0005, Ping Li 0016
ACM Trans. Graph.4