EDBT 2026 Demo / reviewers in the wild / expert
Yongri Piao
dblp:152/4090
· DBLP profile ↗
46ranked-venue papers
10as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 23 since 2021Artificial intelligence and machine learning · 21 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAM3-I: Segment Anything with InstructionsabstractJingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jincai Huang 0003, Wei Ji 0011, Qi Bi, Yongri Piao, Miao Zhang 0004, Xiaoqi Zhao 0003, Qiang Chen 0007, Shihao Zou, Huchuan Lu, Li Cheng 0001 |
ACL (1) | 7 |
| 2025 | Feature Selection and Dual Perturbation: Synergistic Approaches for Semi-Supervised Polyp SegmentationabstractSemi-supervised polyp segmentation (SSPS) is crucial for the computer-aided diagnosis of colorectal cancer, as it reduces the reliance on extensive labeled data. Although previous SSPS methods have achieved notable success, further investigation into critical issues remains necessary. In this paper, we focus on two key challenges in polyp segmentation: First, how can models extract discriminative features effectively? Second, how can we design a mechanism to suppress cognitive bias in SSPS? To address these, we propose a Feature Selection and Dual Perturbation Network integrating a feature selection and interaction module (FSIM) and a dual perturbation mechanism. The FSIM selects texture-rich information and aggregates discriminative features. Inspired by the immune response, the dual perturbation mechanism employs non-shared parameters to independently handle input and network perturbations. Moreover, a constrained loss function encourages effective collaboration among network components, enhancing robustness and reducing cognitive bias. Extensive experiments on multiple polyp datasets demonstrate our method consistently outperforms state-of-the-art SSPS approaches. Remarkably, even when using only 50 % of the labeled data, our approach surpasses several advanced fully supervised models. Miao Zhang 0004, Yidi Tang, Zhuangze Hou, Weibing Sun, Yongri Piao, Huchuan Lu |
BIBM | 6 |
| 2025 | DefMamba: Deformable Visual State Space ModelabstractRecently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and performance. However, most existing visual Mamba methods flatten images into 1D sequences using predefined scan orders, which results the model being less capable of utilizing the spatial structural information of the image during the feature extraction process. To address this issue, we proposed a novel visual foundation model called Def-Mamba. This model includes a multi-scale backbone structure and deformable mamba (DM) blocks, which dynamically adjust the scanning path to prioritize important information, thus enhancing the capture and processing of relevant input features. By combining a deformable scanning (DS) strategy, this model significantly improves its ability to learn image structures and detects changes in object details. Numerous experiments have shown that Def-Mamba achieves state-of-the-art performance in various visual tasks, including image classification, object detection, instance segmentation, and semantic segmentation. The code is open source on DefMamba . Leiye Liu, Miao Zhang 0004, Jihao Yin, Tingwei Liu, Wei Ji 0011, Yongri Piao, Huchuan Lu |
CVPR | 6 |
| 2025 | Hierarchical Scalable Receptive Fields for Efficient Neural Architecture Search in Salient Object Detection
Tingwei Liu, Miao Zhang 0004, Yongri Piao |
PRCV (2) | 4 |
| 2025 | Scene-Aware Background Decoupling via Collaborative Fusion for Video Salient Object DetectionabstractVideo Salient Object Detection (VSOD) faces significant challenges due to complex background disturbances in video sequences. Although notable progress has been made in this field, further efforts are needed to better address background interference. To address these challenges, our scene-aware background decoupling network (SBDNet) equipped with a scene-aware background decoupling strategy (SBDS) as well as collaborative fusion decoders (CFD). The SBDS includes a Dynamic Background Disentanglement Module (DBD) designed to effectively eliminate background distractions in real-world scenes. The DBD achieves this by utilizing a background mask to refine foreground elements and extracting semantic-enhanced context weights. The CFD enhances the fusion between CNN and Transformer decoders, ensuring high accuracy in saliency detection for video sequences. Extensive results demonstrate that our SBDNet significantly outperforms 14 state-of-the-art methods on four widely used benchmark datasets. Shuyao Wang, Tingwei Liu, Yongri Piao, Miao Zhang 0004 |
IEEE Signal Process. Lett. | 3 |
| 2025 | ProSegDiff: Prostate Segmentation Diffusion Network Based on Adaptive Adjustment of Injection FeaturesabstractRecently, methods based on Diffusion Probability Models (DPM) have achieved notable success in the field of medical image segmentation. However, most of these methods do not perform well in segmenting ambiguous areas when dealing with prostate segmentation tasks due to the low distinguishability of prostate images and the high overlap of its boundary with adjacent organs. To address this issue, this paper introduces a diffusion-based framework named ProSegDiff, ProSegDiff employs an Adapter to dynamically adjust features from the conditional network to align with the denoising process of the denoising network. Furthermore, the denoising process is conducted in the latent space to minimize the consumption of computational resources, and a proposed selection strategy is employed to identify the better results from multiple inferences. Extensive comparative experiments on four benchmark datasets demonstrate the effectiveness of this method, which achieves superior performance across four evaluation metrics. Jialong Zhong, Tingwei Liu, Yongri Piao, Weibing Sun, Huchuan Lu |
IEEE Signal Process. Lett. | 3 |
| 2025 | CNN-Transformer Rectified Collaborative Learning for Medical Image SegmentationabstractAutomatic and precise medical image segmentation (MIS) is of vital importance for clinical diagnosis and analysis. Current MIS methods mainly rely on the convolutional neural network (CNN) or self-attention mechanism (Transformer) for feature modeling. However, CNN-based methods suffer from the inaccurate localization owing to the limited global dependency while Transformer-based methods always present the coarse boundary for the lack of local emphasis. Although some CNN-Transformer hybrid methods are designed to synthesize the complementary local and global information for better performance, the combination of CNN and Transformer introduces numerous parameters and increases the computation cost. To this end, this paper proposes a CNN-Transformer rectified collaborative learning (CTRCL) framework to learn stronger CNN-based and Transformer-based models for MIS tasks via the bi-directional knowledge transfer between them. Specifically, we propose a rectified logit-wise collaborative learning (RLCL) strategy which introduces the ground truth to adaptively select and rectify the wrong regions in student soft labels for accurate knowledge transfer in the logit space. We also propose a class-aware feature-wise collaborative learning (CFCL) strategy to achieve effective knowledge transfer between CNN-based and Transformer-based models in the feature space by granting their intermediate features the similar capability of category perception. Extensive experiments on three popular MIS benchmarks demonstrate that our CTRCL outperforms most state-of-the-art collaborative learning methods under different evaluation metrics. The source code will be publicly available athttps://github.com/LanhooNg/CTRCL. Lanhu Wu, Miao Zhang 0004, Yongri Piao, Zhenyan Yao, Weibing Sun, Feng Tian 0001, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | PMNet: Predator-Mimicking Network for Video Camouflaged Object DetectionabstractThe predator has the ability to quickly respond to the misjudged decision and hunt the camouflaged target by analyzing its movement. Those decision compensation and movement analysis for hunting are closely tied to temporal and spatial information. This can be mirrored in the video camouflaged object detection (VCOD) task where the captured temporal information may be misjudged as well as the spatial information tends to be inaccurate in complex scenes. Thus, two key factors should be considered in the VCOD task: How can a model cope with the misjudged temporal information; How can spatial features interact with the temporal information to understand dynamic scenes? To this end, we propose a predator-mimicking network (PMNet) equipped with a selective temporal alignment module (STAM) and a temporal-spatial feedback module (T-SFM). The STAM is designed to alleviate the influence of the misjudged motion trajectory by adopting our adaptive selection mechanism from a novel perspective. In T-SFM, the temporal information works as the self-knowledge to provide assistance and interact with spatial features, enabling the model to effectively detect the camouflaged object. Experimental results demonstrate that our method achieves state-of-the-art performance on VCOD benchmarks. Furthermore, our model can be generalized in the video salient object detection (VSOD) task and also outperforms existing state-of-the-art methods. The source code will be publicly available athttps://github.com/LiuTingWed/CriDiff. Miao Zhang 0004, Beiqi Hu, Shunyu Yao 0004, Yongri Piao, Huchuan Lu |
IEEE Trans. Multim. | 4 |
| 2025 | Learning Discriminative Representation for Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) is the task of identifying and emphasizing the common salient objects in a collection of images. The current co-salient object detection frameworks often extract features and model interimage relations separately. Although these methods achieve promising performance in many scenes, separating the feature extraction and relation modeling falls short of obtaining discriminative features for co-salient objects, resulting in subperformance, especially in some complex and cluttered real-world scenes. In this article, we introduce a novel CoSOD framework to unify feature extraction and interimage relation modeling. We design an early token interaction module (ETIM) that bridges information flow between branches to simultaneously realize feature extraction and interimage information interaction. To further enhance our network's capability to distinguish co-salient objects from other irrelevant foreground objects, we introduce a pixel-to-group contrastive (PGC) learning method. This approach aids in eliminating the need for additional interaction modules while preserving features' discriminative power for co-salient objects. Our proposed CoSOD framework only includes a backbone embedded with ETIM, a decoder without interaction modules and a project head only used during the training phase. Extensive experiments on three challenging benchmarks, that is, CoCA, CoSOD3k, and Cosal2015, demonstrate that our proposed method can outperform current leading-edge models and achieve the new state-of-the-art. The source code is available at https://github.com/zhiwang98/LDRNet. Yongri Piao, Tingwei Liu, Jihao Yin, Miao Zhang 0004, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | BGDiff: Boundary-Guided Injection Diffusion Framework for Prostate SegmentationabstractRecently, the Diffusion Probabilistic Model (DPM)-based methods have achieved substantial success in the field of medical image segmentation. However, most of these methods are not effective in addressing the issue of blurred edges in prostate segmentation tasks. To address this issue, This paper proposes a framework based on the diffusion model, named BGDiff, which is based on Boundary Guided Injection Module(BGIM) and Adaptive Boundary Loss for prostate segmentation. The BGIM can establish connections between the denoising processes of adjacent steps, thereby providing stable guidance for boundary areas as the denoising progresses step by step, while the Adaptive Boundary Loss adjusts the loss weights for more challenging boundaries based on the model’s active feedback. Extensive experiments on four benchmark datasets demonstrate the effectiveness of the proposed method and achieve state-of-the-art performance on four evaluation metrics. The source code will be publicly available at https://github.com/zjlGO/BGDiff Jialong Zhong, Tingwei Liu, Miao Zhang 0004, Yongri Piao, Weibing Sun, Huchuan Lu |
BIBM | 4 |
| 2024 | CriDiff: Criss-Cross Injection Diffusion Framework via Generative Pre-train for Prostate Segmentation
Tingwei Liu, Miao Zhang 0004, Leiye Liu, Jialong Zhong, Shuyao Wang, Yongri Piao, Huchuan Lu |
MICCAI (8) | 6 |
| 2024 | Auto-USOD: Searching Topology for Underwater Salient Object Detection
Tingwei Liu, Runyu Wang, Miao Zhang 0004, Yongri Piao, Huchuan Lu |
PRCV (2) | 4 |
| 2024 | Edge-Guided Bidirectional-Attention Residual Network for Polyp Segmentation
Lanhu Wu, Miao Zhang 0004, Yongri Piao, Huchuan Lu |
PRCV (14) | 3 |
| 2024 | Focal Perception Transformer for Light Field Salient Object Detection
Miao Zhang 0004, Yongri Piao, Jihao Yin, Huchuan Lu |
PRCV (8) | 3 |
| 2023 | To Be Critical: Self-calibrated Weakly Supervised Learning for Salient Object Detection
Tingwei Liu, Miao Zhang 0004, Yongri Piao |
PRCV (11) | 4 |
| 2023 | Delving into Calibrated Depth for Accurate RGB-D Salient Object Detection
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | PANet: Patch-Aware Network for Light Field Salient Object DetectionabstractMost existing light field saliency detection methods have achieved great success by exploiting unique light field data-focus information in focal slices. However, they process light field data in a slicewise way, leading to suboptimal results because the relative contribution of different regions in focal slices is ignored. How we can comprehensively explore and integrate focused saliency regions that would positively contribute to accurate saliency detection. Answering this question inspires us to develop a new insight. In this article, we propose a patch-aware network to explore light field data in a regionwise way. First, we excavate focused salient regions with a proposed multisource learning module (MSLM), which generates a filtering strategy for integration followed by three guidances based on saliency, boundary, and position. Second, we design a sharpness recognition module (SRM) to refine and update this strategy and perform feature integration. With our proposed MSLM and SRM, we can obtain more accurate and complete saliency maps. Comprehensive experiments on three benchmark datasets prove that our proposed method achieves competitive performance over 2-D, 3-D, and 4-D salient object detection methods. The code and results of our method are available at https://github.com/OIPLab-DUT/IEEE-TCYB-PANet. Yongri Piao, Yongyao Jiang, Miao Zhang 0004, Huchuan Lu |
IEEE Trans. Cybern. | 1 |
| 2023 | Depth Injection Framework for RGBD Salient Object DetectionabstractDepth data with a predominance of discriminative power in location is advantageous for accurate salient object detection (SOD). Existing RGBD SOD methods have focused on how to properly use depth information for complementary fusion with RGB data, having achieved great success. In this work, we attempt a far more ambitious use of the depth information by injecting the depth maps into the encoder in a single-stream model. Specifically, we propose a depth injection framework (DIF) equipped with an Injection Scheme (IS) and a Depth Injection Module (DIM). The proposed IS enhances the semantic representation of the RGB features in the encoder by directly injecting depth maps into the high-level encoder blocks, while helping our model maintain computational convenience. Our proposed DIM acts as a bridge between the depth maps and the hierarchical RGB features of the encoder and helps the information of two modalities complement and guide each other, contributing to a great fusion effect. Experimental results demonstrate that our proposed method can achieve state-of-the-art performance on six RGBD datasets. Moreover, our method can achieve excellent performance on RGBT SOD and our DIM can be easily applied to single-stream SOD models and the transformer architecture, proving a powerful generalization ability. Shunyu Yao 0004, Miao Zhang 0004, Yongri Piao, Chaoyi Qiu, Huchuan Lu |
IEEE Trans. Image Process. | 3 |
| 2023 | Noise-Sensitive Adversarial Learning for Weakly Supervised Salient Object DetectionabstractWeakly supervised salient object detection (WSOD) aims at training saliency detection models with weak supervision. Normally, the WSOD methods use pseudo labels converted from image-level classification labels to train the saliency network. However, the converted pseudo labels always contain noise information compared to ground truth. Previous methods are directly affected by pseudo label noise to generate error-prone predictions. To mitigate this problem, we design a noise-robust adversarial learning framework and propose a noise-sensitive training strategy for the framework. The framework consists of a saliency network and a noise-robust discriminator network. With the guidance of noise-robust discriminator network, our saliency network is robust to noise information in pseudo labels. The proposed noise-sensitive training strategy can make good use of both superior and inferior samples in the pseudo label dataset. With the noise-sensitive training strategy, our framework can further balance the learning of saliency information and the robustness of noise information. Comprehensive experiments on five public datasets demonstrate that our method outperforms the existing image-level classification label based WSOD methods. Yongri Piao, Miao Zhang 0004, Yongyao Jiang, Huchuan Lu |
IEEE Trans. Multim. | 1 |
| 2023 | C$^{2}$DFNet: Criss-Cross Dynamic Filter Network for RGB-D Salient Object DetectionabstractThe ability to deal with intra and inter-modality features has been critical to the development of RGB-D salient object detection. While many works have advanced in leaps and bounds in this field, most existing methods have not taken their way down into the inherent differences between the RGB and depth data due to widely adopted conventional convolution in which fixed parameter kernels are applied during inference. To promote intra and inter-modality interaction conditioned on various scenarios, as RGB and depth data are processed independently and later fused interactively, we develop a new insight and a better model. In this paper, we introduce a criss-cross dynamic filter network by decoupling dynamic convolution. First, we propose a Model-specific Dynamic Enhanced Module (MDEM) that dynamically enhances the intra-modality features with global context guidance. Second, we propose a Scene-aware Dynamic Fusion Module (SDFM) to realize dynamic feature selection between two modalities. As a result, our model achieves accurate predictions of salient objects. Extensive experiments demonstrate that our method achieves competitive performance over 28 state-of-the-art RGB-D methods on 7 public datasets. Miao Zhang 0004, Shunyu Yao 0004, Beiqi Hu, Yongri Piao, Wei Ji 0011 |
IEEE Trans. Multim. | 4 |
| 2022 | Adaptive Co-teaching for Unsupervised Monocular Depth Estimation
Weisong Ren, Lijun Wang 0001, Yongri Piao, Miao Zhang 0004, Huchuan Lu, Ting Liu 0018 |
ECCV (1) | 3 |
| 2022 | PreyNet: Preying on Camouflaged ObjectsabstractSpecies often adopt various camouflage strategies to be seamlessly blended into the surroundings for self-protection. To figure out the concealment, predators have evolved excellent hunting skills. Exploring the intrinsic mechanisms of the predation behavior can offer more insightful glimpse into the task of camouflaged object detection (COD). In this work, we strive to seek answers for accurate COD and propose a PreyNet, which mimics the two processes of predation, namely, initial detection (sensory mechanism) and predator learning (cognitive mechanism). To exploit the sensory process, a bidirectional bridging interaction module (BBIM) is designed for selecting and aggregating initial features in an attentive manner. The predator learning process is formulated as a policy-and-calibration paradigm, with the goal of deciding on uncertain regions and encouraging targeted feature calibration. Besides, we obtain adaptive weight for multi-layer supervision during training via computing on the uncertainty estimation. Extensive experiments demonstrate that our model produces state-of-the-art results on several benchmarks. We further verify the scalability of the predator learning paradigm through applications on top-ranking salient object detection models. Our code is publicly available at \urlhttps://github.com/OIPLab-DUT/PreyNet. Miao Zhang 0004, Yongri Piao, Dongxiang Shi, Shusen Lin, Huchuan Lu |
ACM Multimedia | 3 |
| 2022 | Semi-Supervised Video Salient Object Detection Based on Uncertainty-Guided Pseudo LabelsabstractSemi-Supervised Video Salient Object Detection (SS-VSOD) is challenging because of the lack of temporal information in video sequences caused by sparse annotations. Most works address this problem by generating pseudo labels for unlabeled data. However, error-prone pseudo labels negatively affect the VOSD model. Therefore, a deeper insight into pseudo labels should be developed. In this work, we aim to explore 1) how to utilize the incorrect predictions in pseudo labels to guide the network to generate more robust pseudo labels and 2) how to further screen out the noise that still exists in the improved pseudo labels. To this end, we propose an Uncertainty-Guided Pseudo Label Generator (UGPLG), which makes full use of inter-frame information to ensure the temporal consistency of the pseudo labels and improves the robustness of the pseudo labels by strengthening the learning of difficult scenarios. Furthermore, we also introduce the adversarial learning to address the noise problems in pseudo labels, guaranteeing the positive guidance of pseudo labels during model training. Experimental results demonstrate that our methods outperform existing semi-supervised method and partial fully-supervised methods across five public benchmarks of DAVIS, FBMS, MCL, ViSal and SegTrack-V2. Yongri Piao, Chenyang Lu 0007, Miao Zhang 0004, Huchuan Lu |
NeurIPS | 1 |
| 2022 | DMRA: Depth-Induced Multi-Scale Recurrent Attention Network for RGB-D Saliency DetectionabstractIn this work, we propose a novel depth-induced multi-scale recurrent attention network for RGB-D saliency detection, named as DMRA. It achieves dramatic performance especially in complex scenarios. There are four main contributions of our network that are experimentally demonstrated to have significant practical merits. First, we design an effective depth refinement block using residual connections to fully extract and fuse cross-modal complementary cues from RGB and depth streams. Second, depth cues with abundant spatial information are innovatively combined with multi-scale contextual features for accurately locating salient objects. Third, a novel recurrent attention module inspired by Internal Generative Mechanism of human brain is designed to generate more accurate saliency results via comprehensively learning the internal semantic relation of the fused feature and progressively optimizing local details with memory-oriented scene understanding. Finally, a cascaded hierarchical feature fusion strategy is designed to promote efficient information interaction of multi-level contextual features and further improve the contextual representability of model. In addition, we introduce a new real-life RGB-D saliency dataset containing a variety of complex scenarios that has been widely used as a benchmark dataset in recent RGB-D saliency detection research. Extensive empirical experiments demonstrate that our method can accurately identify salient objects and achieve appealing performance against 18 state-of-the-art RGB-D saliency models on nine benchmark datasets. Wei Ji 0011, Ge Yan 0006, Yongri Piao, Shunyu Yao 0004, Miao Zhang 0004, Li Cheng 0001, Huchuan Lu |
IEEE Trans. Image Process. | 4 |
| 2022 | Exploring Spatial Correlation for Light Field Saliency Detection: Expansion From a Single ViewabstractPrevious 2D saliency detection methods extract salient cues from a single view and directly predict the expected results. Both traditional and deep-learning-based 2D methods do not consider geometric information of 3D scenes. Therefore the relationship between scene understanding and salient objects cannot be effectively established. This limits the performance of 2D saliency detection in challenging scenes. In this paper, we show for the first time that saliency detection problem can be reformulated as two sub-problems: light field synthesis from a single view and light-field-driven saliency detection. This paper first introduces a high-quality light field synthesis network to produce reliable 4D light field information. Then a novel light-field-driven saliency detection network is proposed, in which a Direction-specific Screening Unit (DSU) is tailored to exploit the spatial correlation among multiple viewpoints. The whole pipeline can be trained in an end-to-end fashion. Experimental results demonstrate that the proposed method outperforms the state-of-the-art 2D, 3D and 4D saliency detection methods. Our code is publicly available at https://github.com/OIPLab-DUT/ESCNet. Miao Zhang 0004, Yongri Piao, Huchuan Lu |
IEEE Trans. Image Process. | 3 |
| 2021 | Calibrated RGB-D Salient Object DetectionabstractComplex backgrounds and similar appearances between objects and their surroundings are generally recognized as challenging scenarios in Salient Object Detection (SOD). This naturally leads to the incorporation of depth information in addition to the conventional RGB image as input, known as RGB-D SOD or depth-aware SOD. Meanwhile, this emerging line of research has been considerably hindered by the noise and ambiguity that prevail in raw depth images. To address the aforementioned issues, we propose a Depth Calibration and Fusion (DCF) framework that contains two novel components: 1) a learning strategy to calibrate the latent bias in the original depth maps towards boosting the SOD performance; 2) a simple yet effective cross reference module to fuse features from both RGB and depth modalities. Extensive empirical experiments demonstrate that the proposed approach achieves superior performance against 27 state-of-the-art methods. Moreover, our depth calibration strategy alone can work as a preprocessing step; empirically it results in noticeable improvements when being applied to existing cutting-edge RGB-D SOD models. Source code is available at https://github.com/jiwei0921/DCF. Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Shunyu Yao 0004, Qi Bi, Kai Ma 0002, Yefeng Zheng 0001, Huchuan Lu, Li Cheng 0001 |
CVPR | 5 |
| 2021 | MFNet: Multi-filter Directive Network for Weakly Supervised Salient Object DetectionabstractWeakly supervised salient object detection (WSOD) targets to train a CNNs-based saliency network using only low-cost annotations. Existing WSOD methods take various techniques to pursue single "high-quality" pseudo label from low-cost annotations and then develop their saliency networks. Though these methods have achieved good performance, the generated single label is inevitably affected by adopted refinement algorithms and shows prejudiced characteristics which further influence the saliency networks. In this work, we introduce a new multiple-pseudo-label framework to integrate more comprehensive and accurate saliency cues from multiple labels, avoiding the aforementioned problem. Specifically, we propose a multi-filter directive network (MFNet) including a saliency network as well as multiple directive filters. The directive filter (DF) is designed to extract and filter more accurate saliency cues from the noisy pseudo labels. The multiple accurate cues from multiple DFs are then simultaneously propagated to the saliency network with a multi-guidance loss. Extensive experiments on five datasets over four metrics demonstrate that our method outperforms all the existing con-generic methods. Moreover, it is also worth noting that our framework is flexible enough to apply to existing methods and improve their performance. The code and results of our method are available at https://github.com/OIPLab-DUT/MFNet. Yongri Piao, Miao Zhang 0004, Huchuan Lu |
ICCV | 1 |
| 2021 | Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionabstractThe ability to capture inter-frame dynamics has been critical to the development of video salient object detection (VSOD). While many works have achieved great success in this field, a deeper insight into its dynamic nature should be developed. In this work, we aim to answer the following questions: How can a model adjust itself to dynamic variations as well as perceive fine differences in the real-world environment; How are the temporal dynamics well introduced into spatial information over time? To this end, we propose a dynamic context-sensitive filtering network (DCFNet) equipped with a dynamic context-sensitive filtering module (DCFM) and an effective bidirectional dynamic fusion strategy. The proposed DCFM sheds new light on dynamic filter generation by extracting location-related affinities between consecutive frames. Our bidirectional dynamic fusion strategy encourages the interaction of spatial and temporal information in a dynamic manner. Experimental results demonstrate that our proposed method can achieve state-of-the-art performance on most VSOD datasets while ensuring a real-time speed of 28 fps. The source code is publicly available at https://github.com/OIPLab-DUT/DCFNet. Miao Zhang 0004, Jie Liu 0044, Yongri Piao, Shunyu Yao 0004, Wei Ji 0011, Huchuan Lu, Zhongxuan Luo |
ICCV | 4 |
| 2021 | Auto-MSFNet: Search Multi-scale Fusion Network for Salient Object DetectionabstractMulti-scale features fusion plays a critical role in salient object detection. Most of existing methods have achieved remarkable performance by exploiting various multi-scale features fusion strategies. However, an elegant fusion framework requires expert knowledge and experience, heavily relying on laborious trial and error. In this paper, we propose a multi-scale features fusion framework based on Neural Architecture Search (NAS), named Auto-MSFNet. First, we design a novel search cell, named FusionCell to automatically decide multi-scale features aggregation. Rather than searching one repeatable cell stacked, we allow different FusionCells to flexibly integrate multi-level features. Simultaneously, considering features generated from CNNs are naturally spatial and channel-wise, we propose a new search space for efficiently focusing on the most relevant information. The search space mitigates incomplete object structures or over-predicted foreground regions caused by progressive fusion. Second, we propose a progressive polishing loss to further obtain exquisite boundaries by penalizing misalignment of salient object boundaries. Extensive experiments on five benchmark datasets demonstrate the effectiveness of the proposed method and achieve state-of-the-art performance on four evaluation metrics. The code and results of our method are available at https://github.com/OIPLab-DUT/Auto-MSFNet. Miao Zhang 0004, Tingwei Liu, Yongri Piao, Shunyu Yao 0004, Huchuan Lu |
ACM Multimedia | 3 |
| 2021 | Joint Semantic Mining for Weakly Supervised RGB-D Salient Object DetectionabstractTraining saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining per-pixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions. Wei Ji 0011, Qi Bi, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001 |
NeurIPS | 6 |
| 2021 | Dynamic Fusion Network for Light Field Depth Estimation
Yongri Piao, Xinxin Ji, Miao Zhang 0004 |
PRCV (2) | 2 |
| 2020 | Exploit and Replace: An Asymmetrical Two-Stream Architecture for Versatile Light Field Saliency DetectionabstractLight field saliency detection is becoming of increasing interest in recent years due to the significant improvements in challenging scenes by using abundant light field cues. However, high dimension of light field data poses computation-intensive and memory-intensive challenges, and light field data access is far less ubiquitous as RGB data. These may severely impede practical applications of light field saliency detection. In this paper, we introduce an asymmetrical two-stream architecture inspired by knowledge distillation to confront these challenges. First, we design a teacher network to learn to exploit focal slices for higher requirements on desktop computers and meanwhile transfer comprehensive focusness knowledge to the student network. Our teacher network is achieved relying on two tailor-made modules, namely multi-focusness recruiting module (MFRM) and multi-focusness screening module (MFSM), respectively. Second, we propose two distillation schemes to train a student network towards memory and computation efficiency while ensuring the performance. The proposed distillation schemes ensure better absorption of focusness knowledge and enable the student to replace the focal slices with a single RGB image in an user-friendly way. We conduct the experiments on three benchmark datasets and demonstrate that our teacher network achieves state-of-the-arts performance and student network (ResNet18) achieves Top-1 accuracies on HFUT-LFSD dataset and Top-4 on DUT-LFSD, which tremendously minimizes the model size by 56% and boosts the Frame Per Second (FPS) by 159%, compared with the best performing method. Yongri Piao, Zhengkun Rong, Miao Zhang 0004, Huchuan Lu |
AAAI | 1 |
| 2020 | A2dele: Adaptive and Attentive Depth Distiller for Efficient RGB-D Salient Object DetectionabstractExisting state-of-the-art RGB-D salient object detection methods explore RGB-D data relying on a two-stream architecture, in which an independent subnetwork is required to process depth data. This inevitably incurs extra computational costs and memory consumption, and using depth data during testing may hinder the practical applications of RGB-D saliency detection. To tackle these two dilemmas, we propose a depth distiller (A2dele) to explore the way of using network prediction and attention as two bridges to transfer the depth knowledge from the depth stream to the RGB stream. First, by adaptively minimizing the differences between predictions generated from the depth stream and RGB stream, we realize the desired control of pixel-wise depth knowledge transferred to the RGB stream. Second, to transfer the localization knowledge to RGB features, we encourage consistencies between the dilated prediction of the depth stream and the attention map from the RGB stream. As a result, we achieve a lightweight architecture without use of depth data at test time by embedding our A2dele. Our extensive experimental evaluation on five benchmarks demonstrate that our RGB stream achieves state-of-the-art performance, which tremendously minimizes the model size by 76% and runs 12 times faster, compared with the best performing method. Furthermore, our A2dele can be applied to existing RGB-D networks to significantly improve their efficiency while maintaining performance (boosts FPS by nearly twice for DMRA and 3 times for CPFP). Yongri Piao, Zhengkun Rong, Miao Zhang 0004, Weisong Ren, Huchuan Lu |
CVPR | 1 |
| 2020 | Select, Supplement and Focus for RGB-D Saliency DetectionabstractDepth data containing a preponderance of discriminative power in location have been proven beneficial for accurate saliency prediction. However, RGB-D saliency detection methods are also negatively influenced by randomly distributed erroneous or missing regions on the depth map or along the object boundaries. This offers the possibility of achieving more effective inference by well designed models. In this paper, we propose a new framework for accurate RGB-D saliency detection taking account of global location and local detail complementarities from two modalities. This is achieved by designing a complimentary interaction module (CIM) to discriminatively select useful representation from the RGB and depth data, and effectively integrate cross-modal features. Benefiting from the proposed CIM, the fused features can accurately locate salient objects with fine edge details. Moreover, we propose a compensation-aware loss to improve the network's confidence in detecting hard samples. Comprehensive experiments on six public datasets demonstrate that our method outperforms 18 state-of-the-art methods. Miao Zhang 0004, Weisong Ren, Yongri Piao, Zhengkun Rong, Huchuan Lu |
CVPR | 3 |
| 2020 | Accurate RGB-D Salient Object Detection via Collaborative Learning
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Huchuan Lu |
ECCV (18) | 4 |
| 2020 | RGB-D Salient Object Detection with Cross-Modality Modulation and Selection
Chongyi Li, Runmin Cong, Yongri Piao, Qianqian Xu 0001, Chen Change Loy |
ECCV (8) | 3 |
| 2020 | Asymmetric Two-Stream Architecture for Accurate RGB-D Saliency Detection
Miao Zhang 0004, Sun Xiao Fei, Jie Liu 0044, Yongri Piao, Huchuan Lu |
ECCV (28) | 5 |
| 2020 | Feature Reintegration over Differential Treatment: A Top-down and Adaptive Fusion Network for RGB-D Salient Object DetectionabstractMost methods for RGB-D salient object detection (SOD) utilize the same fusion strategy to explore the cross-modal complementary information at each level. However, this may ignore different feature contributions from two modalities on different levels towards prediction. In this paper, we propose a novel top-down multi-level fusion structure where different fusion strategies are utilized to effectively explore the low-level and high-level features. This is achieved by designing the interweave fusion module (IFM) to effectively integrate the global information and designing the gated select fusion module (GSFM) to discriminatively select useful local information by filtering out the unnecessary one from RGB and depth data. Moreover, we propose an adaptive fusion module (AFM) to reintegrate the fused cross-modal features of each level to predict a more accurate result. Comprehensive experiments on 7 challenging benchmark datasets demonstrate that our method achieves the competitive performance over 14 state-of-the-art RGB-D alternative methods. Miao Zhang 0004, Yu Zhang 0165, Yongri Piao, Beiqi Hu, Huchuan Lu |
ACM Multimedia | 3 |
| 2020 | Saliency Detection via Depth-Induced Cellular Automata on Light FieldabstractIncorrect saliency detection such as false alarms and missed alarms may lead to potentially severe consequences in various application areas. Effective separation of salient objects in complex scenes is a major challenge in saliency detection. In this paper, we propose a new method for saliency detection on light field to improve the saliency detection in challenging scenes. We construct an object-guided depth map, which acts as an inducer to efficiently incorporate the relations among light field cues, by using abundant light field cues. Furthermore, we enforce spatial consistency by constructing an optimization model, named Depth-induced Cellular Automata (DCA), in which the saliency value of each superpixel is updated by exploiting the intrinsic relevance of its similar regions. Additionally, the proposed DCA model enables inaccurate saliency maps to achieve a high level of accuracy. We analyze our approach on one publicly available dataset. Experiments show the proposed method is robust to a wide range of challenging scenes and outperforms the state-of-the-art 2D/3D/4D (light-field) saliency detection approaches. Yongri Piao, Miao Zhang 0004, Jingyi Yu 0001, Huchuan Lu |
IEEE Trans. Image Process. | 1 |
| 2020 | LFNet: Light Field Fusion Network for Salient Object DetectionabstractIn this work, we propose a novel light field fusion network-LFNet, a CNNs-based light field saliency model using 4D light field data containing abundant spatial and contextual information. The proposed method can reliably locate and identify salient objects even in a complex scene. Our LFNet contains a light field refinement module (LFRM) and a light field integration module (LFIM) which can fully refine and integrate focusness, depths and objectness cues from light field image. The LFRM learns the light field residual between light field and RGB images for refining features with useful light field cues, and then the LFIM weights each refined light field feature and learns spatial correlation between them to predict saliency maps. Our method can take full advantage of light field information and achieve excellent performance especially in complex scenes, e.g., similar foreground and background, multiple or transparent objects and low-contrast environment. Experiments show our method outperforms the state-of-the-art 2D, 3D and 4D methods across three light field datasets. Miao Zhang 0004, Wei Ji 0011, Yongri Piao, Yu Zhang 0165, Huchuan Lu |
IEEE Trans. Image Process. | 3 |
| 2019 | Deep Learning for Light Field Saliency DetectionabstractRecent research in 4D saliency detection is limited by the deficiency of a large-scale 4D light field dataset. To address this, we introduce a new dataset to assist the subsequent research in 4D light field saliency detection. To the best of our knowledge, this is to date the largest light field dataset in which the dataset provides 1465 all-focus images with human-labeled ground truth masks and the corresponding focal stacks for every light field image. To verify the effectiveness of the light field data, we first introduce a fusion framework which includes two CNN streams where the focal stacks and all-focus images serve as the input. The focal stack stream utilizes a recurrent attention mechanism to adaptively learn to integrate every slice in the focal stack, which benefits from the extracted features of the good slices. Then it is incorporated with the output map generated by the all-focus stream to make the saliency prediction. In addition, we introduce adversarial examples by adding noise intentionally into images to help train the deep network, which can improve the robustness of the proposed network. The noise is designed by users, which is imperceptible but can fool the CNNs to make the wrong prediction. Extensive experiments show the effectiveness and superiority of the proposed model on the popular evaluation metrics. The proposed method performs favorably compared with the existing 2D, 3D and 4D saliency detection methods on the proposed dataset and existing LFSD light field dataset. The code and results can be found at https://github.com/OIPLab-DUT/ ICCV2019_Deeplightfield_Saliency. Moreover, to facilitate research in this field, all images we collected are shared in a ready-to-use manner. Tiantian Wang 0002, Yongri Piao, Huchuan Lu, Lihe Zhang |
ICCV | 2 |
| 2019 | Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionabstractIn this work, we propose a novel depth-induced multi-scale recurrent attention network for saliency detection. It achieves dramatic performance especially in complex scenarios. There are three main contributions of our network that are experimentally demonstrated to have significant practical merits. First, we design an effective depth refinement block using residual connections to fully extract and fuse multi-level paired complementary cues from RGB and depth streams. Second, depth cues with abundant spatial information are innovatively combined with multi-scale context features for accurately locating salient objects. Third, we boost our model's performance by a novel recurrent attention module inspired by Internal Generative Mechanism of human brain. This module can generate more accurate saliency results via comprehensively learning the internal semantic relation of the fused feature and progressively optimizing local details with memory-oriented scene understanding. In addition, we create a large scale RGB-D dataset containing more complex scenarios, which can contribute to comprehensively evaluating saliency models. Extensive experiments on six public datasets and ours demonstrate that our method can accurately identify salient objects and achieve consistently superior performance over 16 state-of-the-art RGB and RGB-D approaches. Yongri Piao, Wei Ji 0011, Miao Zhang 0004, Huchuan Lu |
ICCV | 1 |
| 2019 | Deep Light-field-driven Saliency Detection from a Single ViewabstractPrevious 2D saliency detection methods extract salient cues from a single view and directly predict the expected results. Both traditional and deep-learning-based 2D methods do not consider geometric information of 3D scenes. Therefore the relationship between scene understanding and salient objects cannot be effectively established. This limits the performance of 2D saliency detection in challenging scenes. In this paper, we show for the first time that saliency detection problem can be reformulated as two sub-problems: light field synthesis from a single view and light-field-driven saliency detection. We propose a high-quality light field synthesis network to produce reliable 4D light field information. Then we propose a novel light-field-driven saliency detection network with two purposes, that is, i) richer saliency features can be produced for effective saliency detection; ii) geometric information can be considered for integration of multi-view saliency maps in a view-wise attention fashion. The whole pipeline can be trained in an end-to-end fashion. For training our network, we introduce the largest light field dataset for saliency detection, containing 1580 light fields that cover a wide variety of challenging scenes. With this new formulation, our method is able to achieve state-of-the-art performance. Yongri Piao, Zhengkun Rong, Miao Zhang 0004, Huchuan Lu |
IJCAI | 1 |
| 2019 | Memory-oriented Decoder for Light Field Salient Object DetectionabstractLight field data have been demonstrated in favor of many tasks in computer vision, but existing works about light field saliency detection still rely on hand-crafted features. In this paper, we present a deep-learning-based method where a novel memory-oriented decoder is tailored for light field saliency detection. Our goal is to deeply explore and comprehensively exploit internal correlation of focal slices for accurate prediction by designing feature fusion and integration mechanisms. The success of our method is demonstrated by achieving the state of the art on three datasets. We present this problem in a way that is accessible to members of the community and provide a large-scale light field dataset that facilitates comparisons across algorithms. The code and dataset will be made publicly available. Miao Zhang 0004, Ji Wei, Yongri Piao, Huchuan Lu |
NeurIPS | 4 |
| 2018 | An Information Geometry-Based Distance Between High-Dimensional Covariances for Scalable ClassificationabstractModeling images/videos with covariance matrices has attracted increasing attentions in various vision tasks, especially in visual classification. For covariances-based visual classification, measuring the distances between covariances is one of the key issues and has been studied for decades. Since the space of covariances is a Riemannian manifold, the geometrical structure of covariances should be favorably considered when designing distance metrics. Although this problem has been widely studied, designing an effective and efficient metric between high-dimensional covariances (HDCOV) for scalable classification is still an open problem. In this paper, we present an information geometry-based distance (IGBD) to tackle this challenge from the perspective of information geometry. Our idea is based on the fact that each covariance can be viewed as a zero-mean Gaussian distribution, and thus the distances between covariances are measured by those between the corresponding Gaussian distributions. The core of our method is to project each distribution, in the form of a set of random samples, to a vector on the tangent space of a common, known distribution on the statistical manifold, based on Fisher information metric and maximum likelihood method. On the tangent space, the Euclidean norm can be used to measure the distances between those sets of projection vectors (or equivalently distributions). The proposed IGBD for HDCOV is computationally efficient and easily combined with a linear support vector machine, suitable for scalable visual classification. The experiments are conducted on various kinds and sizes of benchmarks, and results show the proposed method is efficient and the combination of HDCOV can achieve very competitive performance. Qilong Wang 0001, Xiaoxiao Lu, Peihua Li, Zhenguo Gao, Yongri Piao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Computational 3D reconstruction of FAR and big size objects using synthetic apeture integral imagingabstractIn this paper, we present a three-dimensional reconstruction of far and large size objects in synthetic aperture integral imaging system. In the proposed method, the far and large size objects are recorded by the synthetic aperture integral imaging system with an additional Plano concave lens as a set of elemental images. In order to reconstruct the far and big size object, the reconstruction depth of the 3D images are dramatically reduced, because of the Plano concave lens can form a size and depth reduced virtual images. The effect of the proposed method is analyzed in detail and good experimental results confirmed the feasibility of the proposed method. Luyan Xing, Yongri Piao, Hongjia Qu, Miao Zhang 0004 |
ICIP | 2 |