EDBT 2026 Demo / reviewers in the wild / expert
Jie Wang 0095
dblp:29/5259-95
· DBLP profile ↗
16ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-7266-8179ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhanced visual prompt meets low-light saliency detection
Nana Yu, Jie Wang 0095, Yahong Han |
Pattern Recognit. | 2 |
| 2026 | Low-Light Salient Object Detection via Representation DecouplingabstractIn low-light scenes, images often suffer low contrast, poor distinction between the target and the background, and loss of regional information. This makes it challenging for Salient Object Detection (SOD) algorithms to identify and locate the targets. Most existing methods address low-light SOD by first enhancing low-light images and then performing saliency detection. However, splitting these into two sub-tasks may result in negative transfer effects. Additionally, some methods attempt to integrate these two sub-tasks into an end-to-end framework. However, conflicts may arise during the training process due to the differing features required by low-light enhancement and saliency detection. To address this conflict, We propose a decoupled representation network called DRNet, whose core is the construction of a Visual Center Decoupler (VCD). The VCD decouples the learned representations into enhancement-specific and SOD-specific embeddings. This module provides an implicit way to balance the specific requirements of the two subtasks. On the one hand, the enhancement-specific embeddings use RGB illumination constraints and Local Binary Patterns (LBP) feature aggregation to constrain illumination and maintain texture feature stability. On the other hand, the SOD-specific embeddings utilize dynamic multi-scale convolutions to integrate fine-grained details required at multiple scales. Finally, to validate the performance advantages of DRNet, we conduct extensive comparative experiments between the proposed DRNet and existing single-modal methods. Comparative experiments with some representative bi-modal methods further demonstrate the merits of our proposed method. Nana Yu, Jie Wang 0095, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Visual Consensus Prompting for Co-Salient Object DetectionabstractExisting co-salient object detection (CoSOD) methods generally employ a three-stage architecture (i.e., encoding, consensus extraction & dispersion, and prediction) along with a typical full fine-tuning paradigm. Although they yield certain benefits, they exhibit two notable limitations: 1) This architecture relies on encoded features to facilitate consensus extraction, but the meticulously extracted consensus does not provide timely guidance to the encoding stage. 2) This paradigm involves globally updating all parameters of the model, which is parameter-inefficient and hinders the effective representation of knowledge within the foundation model for this task. Therefore, in this paper, we propose an interaction-effective and parameter-efficient concise architecture for the CoSOD task, addressing two key limitations. It introduces, for the first time, a parameter-efficient prompt tuning paradigm and seamlessly embeds consensus into the prompts to formulate task-specific Visual Consensus Prompts (VCP). Our VCP aims to induce the frozen foundation model to perform better on CoSOD tasks by formulating task-specific visual consensus prompts with minimized tunable parameters. Concretely, the primary insight of the purposeful Consensus Prompt Generator (CPG) is to enforce limited tunable parameters to focus on co-salient representations and generate consensus prompts. The formulated Consensus Prompt Disperser (CPD) leverages consensus prompts to form task-specific visual consensus prompts, thereby arousing the powerful potential of pre-trained models in addressing CoSOD tasks. Extensive experiments demonstrate that our concise VCP outperforms 13 cutting-edge full fine-tuning models, achieving the new state of the art (with 6.8% improvement in Fmmetrics on the most challenging CoCA dataset). Source code has been available at https://github.com/WJ-CV/VCP. Jie Wang 0095, Nana Yu, Yahong Han |
CVPR | 1 |
| 2025 | Progressive expansion for semi-supervised bi-modal salient object detection
Jie Wang 0095, Nana Yu, Yahong Han |
Pattern Recognit. | 1 |
| 2025 | Explicitly Disentangling and Exclusively Fusing for Semi-Supervised Bi-Modal Salient Object DetectionabstractBi-modal (RGB-T and RGB-D) salient object detection (SOD) aims to enhance detection performance by leveraging the complementary information between modalities. While significant progress has been made, two major limitations persist. Firstly, mainstream fully supervised methods come with a substantial burden of manual annotation, while weakly supervised or unsupervised methods struggle to achieve satisfactory performance. Secondly, the indiscriminate modeling of local detailed information (object edge) and global contextual information (object body) often results in predicted objects with incomplete edges or inconsistent internal representations. In this work, we propose a novel paradigm to effectively alleviate the above limitations. Specifically, we first enhance the consistency regularization strategy to build a basic semi-supervised architecture for the bi-modal SOD task, which ensures that the model can benefit from massive unlabeled samples while effectively alleviating the annotation burden. Secondly, to ensure detection performance (i.e., complete edges and consistent bodies), we disentangle the SOD task into two parallel sub-tasks: edge integrity fusion prediction and body consistency fusion prediction. Achieving these tasks involves two key steps: 1) the explicitly disentangling scheme decouples salient object features into edge and body features, and 2) the exclusively fusing scheme performs exclusive integrity or consistency fusion for each of them. Eventually, our approach demonstrates significant competitiveness compared to 26 fully supervised methods, while effectively alleviating 90% of the annotation burden. Furthermore, it holds a substantial advantage over 15 non-fully supervised methods. Jie Wang 0095, Xiangji Kong, Nana Yu, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Intra-Modality Self-Enhancement Mirror Network for RGB-T Salient Object DetectionabstractThe inherent imaging properties of sensors result in two distinct differences between the data from the two modalities in RGB-T Salient Object Detection (SOD) tasks. Namely, differences in imaging effectiveness due to varying sensitivities to specific scenes and fundamental domain differences resulting from differences in reflecting scene characteristics. Existing methods primarily focus on pursuing unique cross-modal fusion designs to enhance model performance. However, not only do direct cross-modal fusion modes fail to improve the effectiveness of original features, but intricate cross-modal fusion designs also increase the domain differences between modalities, thereby resulting in suboptimal performance. Therefore, in this paper, we no longer insist on pursuing unique cross-modal fusion designs but instead contemplate how to enhance the effectiveness of original features within modalities (mitigating differences in imaging effectiveness) and utilize a concise cross-modal fusion mechanism (alleviating the impact of domain differences) to achieve satisfactory performance. In this spirit, we propose the Intra-modality Self-enhancement Mirror Network (ISMNet) for RGB-T salient object detection. The core of ISMNet is the proposed Intra-modality Cross-scale Self-enhancement Module (ICSM). The main insight of ICSM is to exploit saliency clues by modeling the correlation between intra-modality cross-scale features (which exhibit strong correlations and small domain differences), thereby enhancing the effectiveness of original multi-scale features within modalities. We employ the proposed novel paradigm to mirror-expand existing typical paradigms to obtain a more robust model architecture. Extensive experiments demonstrate that our proposed new architecture and the introduced universal Intra-modality Cross-scale Self-enhancement Module effectively improve the effectiveness of original features and promote the achievement of state-of-the-art performance. Jie Wang 0095, Jinwen Xi, Jie Shi 0011, Xueying Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Single-Group Generalized RGB and RGB-D Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) aims to segment the co-occurring salient objects in a given group of relevant images. Existing methods typically rely on extensive group training data to enhance the model’s CoSOD capabilities. However, fitting prior knowledge of the extensive group results in a significant performance gap between the seen and out-of-sample image groups. Relaxing such a fitting with fewer prior groups may improve the generalization ability of CoSOD while alleviating the annotation burdens. Hence, it is essential to explore the use of fewer groups during the training phase, such as using only single group, to pursue a highly generalized CoSOD model. We term this new setting as Sg-CoSOD, which aims to train a model using only a single group and effectively apply it to any unseen RGB and RGB-D CoSOD test groups. Towards Sg-CoSOD, it is important to ensure detection performance with limited data and release class dependency with only a single-group. Thus, we present a method, i.e., cross-excitation between saliency and ‘Co’, which decouples the CoSOD task into two parallel branches: ‘Co’ To Saliency (CTS) and Saliency To ‘Co’ (STC). The CTS branch focuses on mining group consensus to guide image co-saliency predictions, while the STC branch is dedicated to using saliency priors to motivate group consensus mining. Furthermore, we propose a Class-Agnostic Triplet (CAT) loss to constrain intra-group consensus while suppressing the model from acquiring class prior knowledge. Extensive experiments on RGB and RGB-D CoSOD tasks with multiple unknown groups show that our model has higher generalization capabilities (e.g., for large-scale datasets CoSOD3k and CoSal1k with multiple generalized groups, we obtain a gain of over 15% in$F_{m}$). Further experimental analyses also reveal that the proposed Sg-CoSOD paradigm has significant potential and promising prospects. Jie Wang 0095, Nana Yu, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Semantic Prompt Enhancement for Semi-Supervised Low-Light Salient Object DetectionabstractMost existing salient object detection (SOD) models are designed based on data collected in well-lit scenes, which is entirely inadequate for low-light conditions. Although recent models are designed for low-light conditions, they still have limitations. First, they simply integrate features without considering the impact of low-light scenes and fail to enhance the contextual information around salient objects. Second, in extremely dark scenes, it is difficult for the human eye to distinguish between the foreground and background, posing significant challenges for data labeling. To address these issues, we design a brightness Retinex enhancer (BRE) tailored for low-light SOD tasks and, for the first time, explore performing low-light SOD within a semi-supervised framework. By using sparse labeled semantic prompts to augment a large amount of unlabeled data, we mitigate the annotation burden while avoiding ineffective labeling in low-light conditions. More specifically, we first use Retinex decomposition to filter out the influence of illumination, while the semantic features extracted by a large model serve as semantic prompts to assist in enhancement. In addition, we introduce a context-guided encoder (CGE) to improve the model's understanding of salient objects. Finally, both labeled and unlabeled data undergo joint consistency training between the shared decoder (SD) and the perturbation decoder. The semi-supervised model enhances low-light SOD performance while also alleviating the burden of data annotation. Extensive experiments demonstrate that, compared with state-of-the-art fully supervised SOD models, the proposed semi-supervised model achieves highly competitive results across multiple test datasets. Nana Yu, Jie Wang 0095, Yahong Han, Weiping Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Degradation-removed multiscale fusion for low-light salient object detection
Nana Yu, Jie Wang 0095, Yahong Han |
Pattern Recognit. | 2 |
| 2024 | Depth-Assisted Semi-Supervised RGB-D Rail Surface Defect InspectionabstractVisual-based methods for rail surface defect inspection (RSDI) effectively improve the limitations of manual inspection, as they can intuitively display the locations and segmented areas of sensitive defects. The RGB-D RSDI task, which leverages the complementarity between RGB and depth (D) image information to enhance detection performance, has attracted widespread attention and achieved significant development. However, existing methods primarily depend on fully supervised training strategies that necessitate a substantial number of manually annotated pixel-level labels to supervise model training. Undoubtedly, extensive manual annotation is exceedingly time-consuming and labor-intensive, particularly considering the irregular shapes and textures of surface defects on rails, further compounding the burden of manual labeling. Therefore, in this paper, we aim to introduce the semi-supervised learning paradigm into this task. Towards the semi-supervised RGB-D RSDI task, a specific semi-supervised network for this task and an effective cross-modal fusion module are crucial to ensuring detection performance under the constraints of limited labeled samples. Thus, we propose a Depth-assisted Semi-Supervised RGB-D RSDI network (DSSNet) to simultaneously alleviate the annotation burden and achieve satisfactory detection performance. Specifically, adhering to the consistency training paradigm, we construct a semi-supervised RGB-D RSDI architecture for this task by optimizing structures, perturbation mechanisms, loss settings, etc. Furthermore, we propose a Depth-assisted Multi-scale Cross-modal Fusion Module (DMCFM) that conducts multi-scale exploration and cross-modal complementary fusion with the assistance of depth. Comprehensive experiments demonstrate that, compared to the latest 14 state-of-the-art fully supervised methods, the proposed DSSNet achieves highly competitive results while effectively alleviating an 80$\%$annotation burden. Jie Wang 0095, Guanwen Qiu, Jinwen Xi, Nana Yu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Weighted Guided Optional Fusion Network for RGB-T Salient Object DetectionabstractThere is no doubt that the rational and effective use of visible and thermal infrared image data information to achieve cross-modal complementary fusion is the key to improving the performance of RGB-T salient object detection (SOD). A meticulous analysis of the RGB-T SOD data reveals that it mainly consists of three scenarios in which both modalities (RGB and T) have a significant foreground and only a single modality (RGB or T) is disturbed. However, existing methods are obsessed with pursuing more effective cross-modal fusion based on treating both modalities equally. Obviously, the subjective use of equivalence has two significant limitations. Firstly, it does not allow for practical discrimination of which modality makes the dominant contribution to performance. While both modalities may have visually significant foregrounds, differences in their imaging properties will result in distinct performance contributions. Secondly, in a specific acquisition scenario, a pair of images with two modalities will contribute differently to the final detection performance due to their varying sensitivity to the same background interference. Intelligibly, for the RGB-T saliency detection task, it would be more reasonable to generate exclusive weights for the two modalities and select specific fusion mechanisms based on different weight configurations to perform cross-modal complementary integration. Consequently, we propose a weighted guided optional fusion network (WGOFNet) for RGB-T SOD. Specifically, a feature refinement module is first used to perform an initial refinement of the extracted multilevel features. Subsequently, a weight generation module (WGM) will generate exclusive network performance contribution weights for each of the two modalities, and an optional fusion module (OFM) will rely on this weight to perform particular integration of cross-modal information. Simple cross-level fusion is finally utilized to obtain the final saliency prediction map. Comprehensive experiments on three publicly available benchmark datasets demonstrate the proposed WGOFNet achieves superior performance compared with the state-of-the-art RGB-T SOD methods. The source code is available at: https://github.com/WJ-CV/WGOFNet . Jie Wang 0095, Jie Shi 0011, Jinwen Xi |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Saliency Prototype for RGB-D and RGB-T Salient Object DetectionabstractMost of the existing bi-modal (RGB-D or RGB-T) salient object detection methods attempt to integrate multimodality information through various fusion strategies. However, existing methods lack a clear definition of salient regions before feature fusion, which results in poor model robustness. To tackle this problem, we propose a novel prototype, the saliency prototype, which captures common characteristic information among salient objects. A prototype contains inherent characteristics information of multiple salient objects, which can be used for feature enhancement of various salient objects. By utilizing the saliency prototype, we provide a clearer definition of salient regions and enable the model to focus on these regions before feature fusion, avoiding the influence of complex backgrounds during the feature fusion stage. Additionally, we utilize the saliency prototypes to address the quality issue of auxiliary modality. Firstly, we apply the saliency prototypes obtained by the primary modality to perform semantic enhancement of the auxiliary modality. Secondly, we dynamically allocate weights for the auxiliary modality during the feature fusion stage in proportion to its quality. Thus, we develop a new bi-modal salient detection architecture Saliency Prototype Network (SPNet), which can be used for both RGB-D and RGB-T SOD. Extensive experimental results on RGB-D and RGB-T SOD datasets demonstrate the effectiveness of the proposed approach against the state-of-the-art. Our code is available at https://github.com/ZZ2490/SPNet. Jie Wang 0095, Yahong Han |
ACM Multimedia | 2 |
| 2022 | Unidirectional RGB-T salient object detection with intertwined driving of encoding and fusion
Jie Wang 0095, Kechen Song, Yanqi Bao, Yunhui Yan, Yahong Han |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Multi-Graph Fusion and Learning for RGBT Image Saliency DetectionabstractRGB and thermal infrared (RGBT) image saliency detection is a relatively new direction in the field of computer vision. Combining the advantages of RGB images and T images can significantly improve detection performance. Currently, there are only a few methods to work on RGBT saliency detection, and the number of image samples cannot meet the training requirements for deep learning, so it remains valuable to propose an effective unsupervised method. In this paper, we present an unsupervised RGBT saliency detection method based on multi-graph fusion and learning. Firstly, RGB images and T images are adaptively fused based on boundary information to produce more accurate superpixels. Next, a multi-graph fusion model is proposed to selectively learn useful information from multi-modal images. Finally, we implement the theory of finding good neighbors in the graph affinity and propose different algorithms for two stages of saliency ranking. Experimental results on three RGBT datasets show that the proposed method is effective compared with the state-of-the-art algorithms. Liming Huang, Kechen Song, Jie Wang 0095, Menghui Niu, Yunhui Yan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | CGFNet: Cross-Guided Fusion Network for RGB-T Salient Object DetectionabstractRGB salient object detection (SOD) has made great progress. However, the performance of this single-modal salient object detection will be significantly decreased when encountering some challenging scenes, such as low light or darkness. To deal with the above challenges, thermal infrared (T) image is introduced into the salient object detection. This fused modal is called RGB-T salient object detection. To achieve deep mining of the unique characteristics of single modal and the full integration of cross-modality information, a novel Cross-Guided Fusion Network (CGFNet) for RGB-T salient object detection is proposed. Specifically, a Cross-Scale Alternate Guiding Fusion (CSAGF) module is proposed to mine the high-level semantic information and provide global context support. Subsequently, we design a Guidance Fusion Module (GFM) to achieve sufficient cross-modality fusion by using single modal as the main guidance and the other modal as auxiliary. Finally, the Cross-Guided Fusion Module (CGFM) is presented and serves as the main decoding block. And each decoding block is consists of two parts with two modalities information of each being the main guidance, i.e., cross-shared Cross-Level Enhancement (CLE) and Global Auxiliary Enhancement (GAE). The main difference between the two parts is that the GFM using different modalities as the main guide. The comprehensive experimental results prove that our method achieves better performance than the state-of-the-art salient detection methods. The source code has released at:https://github.com/wangjie0825/CGFNet.git. Jie Wang 0095, Kechen Song, Yanqi Bao, Liming Huang, Yunhui Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Visible and thermal images fusion architecture for few-shot semantic segmentation
Yanqi Bao, Kechen Song, Jie Wang 0095, Liming Huang, Hongwen Dong, Yunhui Yan |
J. Vis. Commun. Image Represent. | 3 |