EDBT 2026 Demo / reviewers in the wild / expert
Yi Liu 0038
dblp:97/4626-38
· DBLP profile ↗
41ranked-venue papers
13as first author
36since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Mamba-driven sifter for salient object detection
Yi Liu 0038, Dingwen Zhang, Shoukun Xu |
Expert Syst. Appl. | 3 |
| 2026 | AMK-CDiffNet: Adaptive-multiscale K-space cold diffusion network for fast MRI reconstruction
Bingchen Dong, Gengshen Wu, Xia Feng, Yi Liu 0038, Jungong Han |
Expert Syst. Appl. | 4 |
| 2026 | Integrating foundation models with capsule networks for enhanced weakly-supervised semantic segmentation
Zhuang Yao, Gengshen Wu, Yi Liu 0038, Shoukun Xu |
Expert Syst. Appl. | 5 |
| 2026 | Hybrid anchor graph learning and tensorized spectral embedding fusion for multi-view clustering
Guangqi Jiang, Wangjie Chen, Yi Liu 0038, Lin Shi 0007, Jinjia Peng, Huibing Wang |
Neurocomputing | 3 |
| 2026 | Mul-VMamba: Multimodal semantic segmentation using selection-fusion-based vision-Mamba
Yuanhui Guo, Yi Liu 0038, Hai Wang 0003, Chuan Hu 0003 |
Knowl. Based Syst. | 4 |
| 2026 | Dynamic routing towards few-shot point cloud semantic segmentation
Guangqi Jiang, Zhengyao Li, Gengshen Wu, Yi Liu 0038, Shoukun Xu |
Neural Networks | 4 |
| 2026 | Fixing Background Misclassification in Few-Shot Object Detection via Product of ExpertsabstractFew-shot object detection (FSOD) poses a significant challenge due to the difficulty of learning robust and discriminative object representations under limited supervision. A widely adopted solution is the two-stage fine-tuning framework, wherein knowledge acquired from a large-scale base dataset is transferred to a novel dataset containing only a small number of labeled instances. However, this framework is prone to systematically misclassifying novel objects as background, primarily due to incorrect background label caused by the domain gap between base and novel datasets-an issue exacerbated by the sparse representation of novel categories. In this work, we show that this inherent weakness can be exploited by explicitly redefining the category structure and transferring the representations learned during the base training stage. Building on this insight, we propose a simple yet effective framework grounded in the Product of Experts (PoE) formulation, which estimates the joint distribution over background and novel categories by combining the unnormalized logits from independently trained classifiers. Notably, it does not require modifications of the base model or repetition of the base training phase. Furthermore, we introduce a strategy for identifying additional novel-category instances within the base dataset, which effectively augmenting the training set for fine-tuning. The resulting method is architecture-agnostic, imposes negligible overhead, and integrates seamlessly with existing two-stage fine-tuning pipelines. Extensive experiments on PASCAL VOC and COCO demonstrate that the proposed method yields consistent improvements across different baselines, achieving significant gains over state-of-the-art FSOD approaches. Ding Sheng Ong, Yi Liu 0038, Changjing Shang, Guiguang Ding, Qiang Shen 0001, Jungong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Text-Centric multimodal sentiment analysis with asymmetric fine-tuning
Hanzhao Pan, Gengshen Wu, Yi Liu 0038, Jungong Han |
Pattern Recognit. | 3 |
| 2026 | Visual In-Context Learning for Underwater Image RestorationabstractUnderwater images often exhibit common visual degradations, such as color distortion, loss of details, and reduced sharpness, which inevitably compromise the effectiveness of underwater vision tasks. However, most underwater image restoration methods solely focus on learning degradation features from raw images, neglecting the incorporation of additional contextual information to guide restoration, which limits the capability of deep models to restore image quality. In this paper, we propose Visual In-Context Learning (VICL) for underwater image restoration, which leverages degradation information from context to improve image quality. In VICL, Degraded Context Extraction Block (DCEB) employs a self-attention mechanism to extract degradation information from context. In addition, Context Spatial Feature Fusion Block (CSFFB) consists of a Degraded Context Guidance Block (DCGB) and a Multi-Feature Fusion Block (MFFB). DCGB employs a cross-attention mechanism to fuse degraded context with spatial features for guiding underwater image restoration. MFFB replaces traditional encoder-decoder skip connections to better coordinate feature fusion. Extensive experiments on multiple underwater image benchmarks demonstrate that VICL outperforms state-of-the-art methods both quantitatively and visually. The code is available at:https://github.com/zhangao668/VICL. Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu |
IEEE Signal Process. Lett. | 3 |
| 2026 | Reliable Feature Imputation With Cross-View Relation Transfer for Deep Incomplete Multi-View ClassificationabstractIncomplete Multi-view Classification has sparked widespread interest in recent years, since multi-view data suffering from missing values are ubiquitous in real-world scenarios. While many imputation-based methods recover missing data by exploiting inter-sample structural information within individual views, they are inherently susceptible to unreliable or noisy samples, which can lead to low-quality imputation and degrade classification accuracy. Therefore, it is a challenge to effectively mine the multi-stage complex correlations for incomplete multi-view data to achieve reliable imputation and obtain discriminative representation. To address these issues, we present a novel imputation-based approach called Reliable Feature Imputation with Cross-view Relation Transfer for Deep Incomplete Multiview Classification (RFI-IMvC). Our framework fully exploits inter-view and intra-view structural information in multi-stage manner. Specifically, we propose a novel cross-view relation transfer strategy to recover reliable neighbor relationships and achieve high-quality imputation for missing data. Besides, to fully exploit the structural information in reconstructed multi-view data, we develop a dual graph learning module to mine high-order semantic correlation and facilitate interactions of complementary information from instances linked by hyperedge. Finally, inspired by prototype learning, we incorporate a class--level representation loss to further promote intra-class compactness. Extensive experiments on 7 real-world datasets demonstrate that our method outperforms state-of-the-art methods. Guangqi Jiang, Haodong Hou, Yi Liu 0038, Jinjia Peng, Huibing Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Gradient Balanced Part-Whole Relational Weakly Supervised Semantic Segmentation
Zhuang Yao, Guangqi Jiang, Lin Shi 0007, Gengshen Wu, Shoukun Xu, Yi Liu 0038 |
KSEM (1) | 6 |
| 2025 | Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
Yi Liu 0038, Dingwen Zhang, Shoukun Xu, Jungong Han |
Expert Syst. Appl. | 1 |
| 2025 | Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding
Yi Liu 0038, Shoukun Xu, Jungong Han |
Int. J. Comput. Vis. | 1 |
| 2025 | Mamba Capsule Routing Towards Part-Whole Relational Camouflaged Object Detection
Dingwen Zhang, Liangbo Cheng, Yi Liu 0038, Xinggang Wang, Junwei Han 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | MVG-FD: Multi-Modal Visual Guidance and Feature Decomposition for Underwater Image RestorationabstractUnderwater images are frequently affected by light absorption and scattering, which lead to color distortion, reduced contrast, and blurred details, significantly degrading overall image quality. Most underwater image restoration methods are confined to the pixel space of the raw modality, overlooking the important role of other modalities and different frequency-domain features. As a result, the representational capacity of deep learning models is not fully realized, affecting the generation of high-quality images. To address the above issues, we propose Multi-modal Visual Guidance and Feature Decomposition (MVG-FD) method for underwater image restoration. Specifically, we introduce Modality Visual Guidance (MVG) module, which integrates the complementary information provided by depth modality features into the raw features to guide the model in restoring the color of underwater images. Meanwhile, we design Feature Decomposition (FD) module, which utilizes Learnable Wavelet Decomposition (LWD) to decompose and extract the high-frequency bands of the raw features to help restore the texture details of the image. MVG-FD significantly improves PSNR and SSIM on existing datasets. The code is available at:https://github.com/zhangao668/MVG-FD. Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu |
IEEE Signal Process. Lett. | 3 |
| 2025 | SCMVC: Semantic Constraint-Based Spatial-Spectral Multiview Clustering for Hyperspectral ImagesabstractCross-view consensus representation plays a crucial role in hyperspectral image (HSI) clustering. Recently, multi-view contrastive cluster (MVCC) methods have leveraged contrastive loss to extract contextual consensus representations. However, these methods suffer from a critical limitation: MVCC frameworks often regard similar heterogeneous views as positive sample pairs while treating dissimilar homogeneous views as negative sample pairs. This misalignment leads to intra-class inconsistency and inter-class confusion. To address this problem, we propose a novel multi-view clustering method, termed Semantic Constraint-based Spatial-Spectral Multi-view Clustering (SCMVC). First, spatial views are designed to capture diverse features for contrastive clustering. Meanwhile, globally relevant information from the spectral view is extracted using a Transformer, which serves to enhance the representation of similar samples in the spatial multi-view. Then, SCMVC employs a semantic constraint-based joint loss function, comprising a semantic contrast loss and a semantic similarity consistency loss. The semantic contrast loss captures high-level, domain-invariant features from hyperspectral images, while the semantic similarity consistency loss enforces stricter constraints on the similarity of semantically related samples in feature space. Finally, SCMVC utilizes anchor points to guide similarity clustering, reducing randomness by predefining these points. This approach captures the directional characteristics of data, leading to a more stable clustering process. Abundant experiment studies on numerous benchmarks verify the superiority of SCMVC in comparison to some state-of-the-art clustering methods. The codes are available at SCMVC. Fulin Luo, Yi Liu 0038, Tan Guo, Chuan Fu, Qian Shi 0001, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Towards High-Quality MRI Reconstruction With Anisotropic Diffusion-Assisted Generative Adversarial Networks and Its Multi-Modal Images ExtensionabstractRecently, fast Magnetic Resonance Imaging reconstruction technology has emerged as a promising way to improve the clinical diagnostic experience by significantly reducing scan times. While existing studies have used Generative Adversarial Networks to achieve impressive results in reconstructing MR images, they still suffer from challenges such as blurred zones/boundaries and abnormal spots caused by inevitable noise in the reconstruction process. To this end, we propose a novel deep framework termed Anisotropic Diffusion-Assisted Generative Adversarial Networks, which aims to maximally preserve valid high-frequency information and structural details while minimizing noises in reconstructed images by optimizing a joint loss function in a unified framework. In doing so, it enables more authentic and accurate MR image generation. To specifically handle unforeseeable noises, an Anisotropic Diffused Reconstruction Module is developed and added aside the backbone network as a denoise assistant, which improves the final image quality by minimizing reconstruction losses between targets and iteratively denoised generative outputs with no extra computational complexity during the testing phase. To make the most of valuable MRI data, we extend its application to support multi-modal learning to boost reconstructed image quality by aggregating more valid information from images of diverse modalities. Extensive experiments on public datasets show that the proposed framework can achieve superior performance in polishing up the quality of reconstructed MR images. For example, the proposed method obtains average PSNR and mSSIM values of 35.785 dB and 0.9765 on the MRNet dataset, which are at least about 2.9 dB and 0.07 higher than those from the baselines. Yuyang Luo, Gengshen Wu, Yi Liu 0038, Jungong Han |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Capsule Networks With Residual Pose RoutingabstractCapsule networks (CapsNets) have been known difficult to develop a deeper architecture, which is desirable for high performance in the deep learning era, due to the complex capsule routing algorithms. In this article, we present a simple yet effective capsule routing algorithm, which is presented by a residual pose routing. Specifically, the higher-layer capsule pose is achieved by an identity mapping on the adjacently lower-layer capsule pose. Such simple residual pose routing has two advantages: 1) reducing the routing computation complexity and 2) avoiding gradient vanishing due to its residual learning framework. On top of that, we explicitly reformulate the capsule layers by building a residual pose block. Stacking multiple such blocks results in a deep residual CapsNets (ResCaps) with a ResNet-like architecture. Results on MNIST, AffNIST, SmallNORB, and CIFAR-10/100 show the effectiveness of ResCaps for image classification. Furthermore, we successfully extend our residual pose routing to large-scale real-world applications, including 3-D object reconstruction and classification, and 2-D saliency dense prediction. The source code has been released on https://github.com/liuyi1989/ResCaps. Yi Liu 0038, De Cheng, Dingwen Zhang, Shoukun Xu, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Prototype Enhancement for Few-Shot Point Cloud Semantic Segmentation
Zhengyao Li, Gengshen Wu, Yi Liu 0038 |
ICA3PP (2) | 3 |
| 2024 | Spatial-Frequency Integration Network with Dual Prompt Learning for Few-shot Image ClassificationabstractFew-shot image classification is a challenging task that aims to recognize image classes based on only a few training images. However, existing methods face the following two main challenges: (1) Ignoring the frequency domain information during image feature extraction. (2) It does not take the semantic gap between multiple modalities into consideration, which limits the classification performance. To overcome these limitations, we propose a novel method named Spatial-Frequency Integration Network with Dual Prompt Learning for few-shot image classification. Firstly, we introduce a spatial-frequency integration module that combines spatial domain and low-frequency information to extract discriminative image features from the image modality. Secondly, we design a dual prompting module, which integrates learnable prompts and hand-crafted prompts to improve the generalization of applications to new classes. Thirdly, we propose an image-text interaction module to enhance inter-modal complementary and consistency. Both theoretical and experimental validations confirm the effectiveness of the proposed method in few-shot image classification. Yi Liu 0038, Shoukun Xu, Guangqi Jiang |
ISPA | 1 |
| 2024 | Pose Convolutional Routing Towards Lightweight Capsule NetworksabstractAn important branch of artificial intelligence systems and architectures is Capsule Networks (CapsNets) have been known extremely large amount of parameters and computation because of the complex capsule routing algorithm, making it difficult deep architectures in the era of deep learning. To address this challenge, in this paper, we propose a simple yet effective capsule routing algorithm. Specifically, we activate the pose of the entity using its activation probability. On top of that, a convolution on the activated pose matrix to learn the high-level capsules’ pose matrices. Activations of the high-level capsules can be digged from their pose matrices via convolution and activation. Such mechanism generates fewer network parameters and lightweight computation, which make it practitable a deep CapsNets architecture. Experiments on CIFAR-10/100, Small-NORB, MINIST and even large-scale benchmark PASCAL VOC 2007, demonstrate the effectiveness of the proposed method. Chengxin Lv, Guangqi Jiang, Shoukun Xu, Yi Liu 0038 |
ISPA | 4 |
| 2024 | EMVCC: Enhanced Multi-View Contrastive Clustering for Hyperspectral ImagesabstractCross-view consensus representation plays a critical role in hyperspectral images (HSIs) clustering. Recent multi-view contrastive cluster methods utilize contrastive loss to extract contextual consensus representation. However, these methods have a fatal flaw: contrastive learning may treat similar heterogeneous views as positive sample pairs and dissimilar homogeneous views as negative sample pairs. At the same time, the data representation via self-supervised contrastive loss is not specifically designed for clustering. Thus, to tackle this challenge, we propose a novel multi-view clustering method, i.e., Enhanced Multi-View Contrastive Clustering (EMVCC). First, the spatial multi-view is designed to learn the diverse features for contrastive clustering, and the globally relevant information of spectrum-view is extracted by Transformer, enhancing the spatial multi-view differences between neighboring samples. Then, a joint self-supervised loss is designed to constrain the consensus representation from different perspectives to efficiently avoid false negative pairs. Specifically, to preserve the diversity of multi-view information, the features are enhanced by using probabilistic contrastive loss, and the data is projected into a semantic representation space, ensuring that the similar samples in this space are closer in distance. Finally, we design a novel clustering loss that aligns the view feature representation with high confidence pseudo-labels for promoting the network to learn cluster-friendly features. In the training process, the joint self-supervised loss is used to optimize the cross-view features.Abundant experiment studies on numerous benchmarks verify the superiority of EMVCC in comparison to some state-of-the-art clustering methods. The codes are available at https://github.com/YiLiu1999/EMVCC. Fulin Luo, Yi Liu 0038, Xiuwen Gong, Zhixiong Nan, Tan Guo |
ACM Multimedia | 2 |
| 2024 | Adaptive Unified Framework with Global Anchor Graph for Large-Scale Multi-view Clustering
Lin Shi 0007, Wangjie Chen, Yi Liu 0038, Lihua Zhuang, Guangqi Jiang |
PRCV (1) | 3 |
| 2024 | Deep unsupervised part-whole relational visual saliency
Yi Liu 0038, Dingwen Zhang, Shoukun Xu |
Neurocomputing | 1 |
| 2024 | Attention-Based Multi-Kernelized and Boundary-Aware Network for image semantic segmentation
Xuanchen Zhou, Gengshen Wu, Xin Sun 0003, Pengpeng Hu, Yi Liu 0038 |
Neurocomputing | 5 |
| 2024 | Dual enhanced semantic hashing for fast image retrieval
Sizhi Fang, Gengshen Wu, Yi Liu 0038, Xia Feng, Yinghui Kong |
Multim. Tools Appl. | 3 |
| 2024 | SDST: Self-Supervised Double-Structure Transformer for Hyperspectral Images ClusteringabstractDue to the lack of labeled information and the high spectral variability in high-dimensional hyperspectral images (HSI), HSI clustering has emerged as an effective unsupervised approach for HSI information extraction and classification. Deep clustering methods have achieved significant success in unsupervised HSI classification (HSIC) and have gained increasing attention. However, these methods have limitations in terms of robustness, adaptability, and feature representation when dealing with complex large-scale HSI datasets. Therefore, this paper introduced a novel Self-supervised Double-Structure Transformer (SDST) approach for hyperspectral image clustering. Specifically, in our approach, we designed a shared Autoformer structure based on autoencoder to learn the global properties of HSI data by fusing the multi-level features from autoencoder with Transformer. Furthermore, we proposed a siamese Dual-Former Graph Module with superpixel-level features for fewer nodes, which reveals long-dependency graph convolutional features, resulting in more precise graph structure features. By constructing graph with long-dependencies, this module significantly preserves the properties of global dependencies, while focusing on the local features of each superpixel to better represent the fine-grained local details. Finally, we designed a Joint Optimization Module to jointly optimize the double-structure model composed of the shared Autoformer Module and the siamese Dual-Former Graph Module. To validate the effectiveness of the proposed SDST method, we conducted a series of experiments on the Salinas, Botswana, Indian Pines, and Houston2013 datasets. The proposed SDST achieves competitive clustering accuracies compared with the state-of-the-art clustering methods. Codes: https://github.com/YiLiu1999/SDST. Fulin Luo, Yi Liu 0038, Tan Guo, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | TCGNet: Type-Correlation Guidance for Salient Object DetectionabstractContrast and part-whole relations induced by deep neural networks like Convolutional Neural Networks (CNNs) and Capsule Networks (CapsNets) have been known as two types of semantic cues for deep salient object detection. However, few works pay attention to their complementary properties in the context of saliency prediction. In this paper, we probe into this issue and propose a Type-Correlation Guidance Network (TCGNet) for salient object detection. Specifically, a Multi-Type Cue Correlation (MTCC) covering CNNs and CapsNets is designed to extract the contrast and part-whole relational semantics, respectively. Using MTCC, two correlation matrices containing complementary information are computed with these two types of semantics. In return, these correlation matrices are used to guide the learning of the above semantics to generate better saliency cues. Besides, a Type Interaction Attention (TIA) is developed to interact semantics from CNNs and CapsNets for the aim of saliency prediction. Experiments and analysis on five benchmarks show the superiority of the proposed approach. Codes has been released on https://github.com/liuyi1989/TCGNet. Yi Liu 0038, Ling Zhou 0002, Gengshen Wu, Shoukun Xu, Jungong Han |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Facial Expression Recognition on the High Aggregation SubgraphsabstractWith the development of deep learning technology, the performance of facial expression recognition (FER) has been significantly improved. The current main challenge comes from the confusion of facial expressions caused by the highly nonlinear changes of facial expressions. However, the existing FER methods based on Convolutional Neural Networks (CNN) often ignore the underlying relationship between expressions which is crucial to meliorate the performance of recognition for confusable expressions. And the methods based on Graph Convolutional Networks (GCN) can capture the relationship between vertices, but the aggregation degree of subgraphs generated by these methods is low. They are easy to include unconfident neighbors, which increases the learning difficulty of the network. To solve the above problems, this paper proposes a method to recognize facial expressions on the high aggregation subgraphs (HASs) by combing the advantages of CNN extracting features and GCN modeling complex graph patterns. Specifically, we formulate FER as a vertex prediction problem. Considering the importance of high-order neighbors and higher efficiency, we utilize vertex confidence to find high-order neighbors. Then we construct the HASs based on the top embedding features of these high-order neighbors. And we utilize the GCN to perform reasoning and infer the class of vertices for HASs without a large number of overlapping subgraphs. Our method captures the underlying relationship between expressions on the HASs and improves the accuracy and efficiency of FER. Experimental results on both the in-the-lab datasets and the in-the-wild datasets show that our method achieves higher recognition accuracy than several state-of-the-art methods. This highlights the benefit of the underlying relationship between expressions for FER. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Yi Liu 0038 |
IEEE Trans. Image Process. | 6 |
| 2022 | Part-Object Relational Visual SaliencyabstractRecent years have witnessed a big leap in automatic visual saliency detection attributed to advances in deep learning, especially Convolutional Neural Networks (CNNs). However, inferring the saliency of each image part separately, as was adopted by most CNNs methods, inevitably leads to an incomplete segmentation of the salient object. In this paper, we describe how to use the property of part-object relations endowed by the Capsule Network (CapsNet) to solve the problems that fundamentally hinge on relational inference for visual saliency detection. Concretely, we put in place a two-stream strategy, termed Two-Stream Part-Object RelaTional Network (TSPORTNet), to implement CapsNet, aiming to reduce both the network complexity and the possible redundancy during capsule routing. Additionally, taking into account the correlations of capsule types from the preceding training images, a correlation-aware capsule routing algorithm is developed for more accurate capsule assignments at the training stage, which also speeds up the training dramatically. By exploring part-object relationships, TSPORTNet produces a capsule wholeness map, which in turn aids multi-level features in generating the final saliency map. Experimental results on five widely-used benchmarks show that our framework consistently achieves state-of-the-art performance. The code can be found on https://github.com/liuyi1989/TSPORTNet. Yi Liu 0038, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Bi-Directional Progressive Guidance Network for RGB-D Salient Object DetectionabstractMost existing RGB-D salient detection models pay more attention to the quality of the depth images, while in some special cases, the quality of RGB images may even have greater impacts on saliency detection, which has long been ignored and underestimated. To address this problem, in this paper, we present a Bi-directional Progressive Guidance Network (BPGNet) for RGB-D salient object detection, where the qualities of both RGB and depth images are involved. Since it is usually difficult to determine which modality data have low quality in advance, a bi-directional framework based on progressive guidance (PG) strategy is employed to extract and enhance the unimodal features with the aid of another modality data via the alternative interactions between the saliency prediction results and the extracted features from the multi-modality input data. Specifically, the proposed PG strategy is achieved by using the proposed Global Context Awareness (GCA), Auxiliary Feature Extraction (AFE) and Cross-modality Feature Enhancement (CFE) modules. Benefiting from the proposed PG strategy, the disturbing information within the input RGB and depth images can be well suppressed, while the discriminative information within the input images gets enhanced. On top of that, a Fusion Prediction Module (FPM) is further designed to adaptively select those features with higher discriminability as well as enhancing the common information for the final saliency prediction. Experimental results demonstrate that our proposed model is comparable to those of state-of-the-art RGB-D SOD models. Yang Yang 0132, Yongjiang Luo, Yi Liu 0038, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Engaging Part-Whole Hierarchies and Contrast Cues for Salient Object DetectionabstractReal-world scenes always exhibit objects with clutter backgrounds, posing great challenges for deep salient object detection models. In this paper, we propose salient object detection by engaging two saliency cues,i.e., the part-whole hierarchies and contrast cues, resulting in a PWHCNet. Specifically, two branches, which consists of a Dynamic Grouping Capsules (DGC) branch and a DenseHRNet branch, are put in place to learn the part-whole hierarchies and contrast cues, respectively. Moreover, to help highlight the whole salient object in complex scenes, a Background Suppression (BS) module is proposed to guide the shallow features of DenseHRNet with the aid of the part-whole relational cues captured by DGC. Subsequently, these two saliency cues are integrated via a Self-Channel and Mutual-Spatial (SCMS) attention mechanism. Experimental results on five benchmarks demonstrate that the proposed PWHCNet achieves state-of-the-art performance while obtaining the whole salient objects with fine details. Qiang Zhang 0020, Mingxing Duanmu, Yongjiang Luo, Yi Liu 0038, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Disentangled Capsule Routing for Fast Part-Object Relational SaliencyabstractRecently, the Part-Object Relational (POR) saliency underpinned by the Capsule Network (CapsNet) has been demonstrated to be an effective modeling mechanism to improve the saliency detection accuracy. However, it is widely known that the current capsule routing operations have huge computational complexity, which seriously limited the usability of the POR saliency models in real-time applications. To this end, this paper takes an early step towards a fast POR saliency inference by proposing a novel disentangled part-object relational network. Concretely, we disentangle horizontal routing and vertical routing from the original omnidirectional capsule routing, thus generating Disentangled Capsule Routing (DCR). This mechanism enjoys two advantages. On one hand, DCR that disentangles orthogonal 1D (i.e., vertical and horizontal) routing greatly reduces parameters and routing complexity, resulting in much faster inference than omnidirectional 2D routing adopted by existing CapsNets. On the other hand, thanks to the light POR cues explored by DCR, we could conveniently integrate the part-object routing process to different feature layers in CNN, rather than just applying it to the small-scaled one as in previous works. This helps to increase saliency inference accuracy. Compared to previous POR saliency detectors, DPORTNet infers visual saliency (5 ∼ 9 ) × faster, and is more accurate. DPORTNet is available under the open-source license at https://github.com/liuyi1989/DCR. Yi Liu 0038, Dingwen Zhang, Nian Liu 0002, Shoukun Xu, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2021 | Exploring multi-scale deformable context and channel-wise attention for salient object detection
Yi Liu 0038, Mingxing Duanmu, Zhen Huo, Zuntian Chen, Qiang Zhang 0020 |
Neurocomputing | 1 |
| 2021 | Integrating Part-Object Relationship and Contrast for Camouflaged Object DetectionabstractObject detectors that solely rely on image contrast are struggling to detect camouflaged objects in images because of the high similarity between camouflaged objects and their surroundings. To address this issue, in this paper, we investigate the role of the part-object relationship for camouflaged object detection. Specifically, we propose a Part-Object relationship and Contrast Integrated Network (POCINet) covering both search and identification stages, where each stage adopts an appropriate scheme to engage the contrast information and part-object relational knowledge for camouflaged pattern decoding. Besides, we bridge these two stages via a Search-to-Identification Guidance (SIG) module, in which the search result, as well as decoded semantic knowledge, jointly enhances the features encoding ability of the identification stage. Experimental results demonstrate the superiority of our algorithm on three datasets. Notably, our algorithm raises Fβ of the best existing method by approximately 17 points on the CPD1K dataset. The source code will be released soon. Yi Liu 0038, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | Joint Cross-Modal and Unimodal Features for RGB-D Salient Object DetectionabstractRGB-D salient object detection is one of the basic tasks in computer vision. Most existing models focus on investigating efficient ways of fusing the complementary information from RGB and depth images for better saliency detection. However, for many real-life cases, where one of the input images has poor visual quality or contains affluent saliency cues, fusing cross-modal features does not help to improve the detection accuracy, when compared to using unimodal features only. In view of this, a novel RGB-D salient object detection model is proposed by simultaneously exploiting the cross-modal features from the RGB-D images and the unimodal features from the input RGB and depth images for saliency detection. To this end, a Multi-branch Feature Fusion Module is presented to effectively capture the cross-level and cross-modal complementary information between RGB-D images, as well as the cross-level unimodal features from the RGB images and the depth images separately. On top of that, a Feature Selection Module is designed to adaptively select those highly discriminative features for the final saliency prediction from the fused cross-modal features and the unimodal features. Extensive evaluations on four benchmark datasets demonstrate that the proposed model outperforms the state-of-the-art approaches by a large margin. Nianchang Huang, Yi Liu 0038, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Multim. | 2 |
| 2020 | Deep Salient Object Detection With Contextual Information GuidanceabstractIntegration of multi-level contextual information, such as feature maps and side outputs, is crucial for Convolutional Neural Networks (CNNs) based salient object detection. However, most existing methods either simply concatenate multi-level feature maps or calculate element-wise addition of multi-level side outputs, thus failing to take full advantages of them. In this work, we propose a new strategy for guiding multi-level contextual information integration, where feature maps and side outputs across layers are fully engaged. Specifically, shallower-level feature maps are guided by the deeper-level side outputs to learn more accurate properties of the salient object. In turn, the deeper-level side outputs can be propagated to high-resolution versions with spatial details complemented by means of shallower-level feature maps. Moreover, a group convolution module is proposed with the aim to achieve high-discriminative feature maps, in which the backbone feature maps are divided into a number of groups and then the convolution is applied to the channels of backbone feature maps within each group. Eventually, the group convolution module is incorporated in the guidance module to further promote the guidance role. Experiments on three public benchmark datasets verify the effectiveness and superiority of the proposed method over the state-of-the-art methods. Yi Liu 0038, Jungong Han, Qiang Zhang 0020, Caifeng Shan |
IEEE Trans. Image Process. | 1 |
| 2019 | Employing Deep Part-Object Relationships for Salient Object DetectionabstractDespite Convolutional Neural Networks (CNNs) based methods have been successful in detecting salient objects, their underlying mechanism that decides the salient intensity of each image part separately cannot avoid inconsistency of parts within the same salient object. This would ultimately result in an incomplete shape of the detected salient object. To solve this problem, we dig into part-object relationships and take the unprecedented attempt to employ these relationships endowed by the Capsule Network (CapsNet) for salient object detection. The entire salient object detection system is built directly on a Two-Stream Part-Object Assignment Network (TSPOANet) consisting of three algorithmic steps. In the first step, the learned deep feature maps of the input image are transformed to a group of primary capsules. In the second step, we feed the primary capsules into two identical streams, within each of which low-level capsules (parts) will be assigned to their familiar high-level capsules (object) via a locally connected routing. In the final step, the two streams are integrated in the form of a fully connected layer, where the relevant parts can be clustered together to form a complete salient object. Experimental results demonstrate the superiority of the proposed salient object detection network over the state-of-the-art methods. Yi Liu 0038, Qiang Zhang 0020, Dingwen Zhang, Jungong Han |
ICCV | 1 |
| 2019 | Salient object detection employing a local tree-structured low-rank representation and foreground consistency
Qiang Zhang 0020, Zhen Huo, Yi Liu 0038, Yunhui Pan, Caifeng Shan, Jungong Han |
Pattern Recognit. | 3 |
| 2019 | Salient Object Detection via Two-Stage GraphsabstractDespite recent advances made in salient object detection using graph theory, the approach still suffers from accuracy problems when the image is characterized by a complex structure, either in the foreground or background, causing erroneous saliency segmentation. This fundamental challenge is mainly attributed to the fact that most existing graph-based methods take only the adjacently spatial consistency among graph nodes into consideration. In this paper, we tackle this issue from a coarse-to-fine perspective and propose a two-stage-graphs approach for salient object detection, in which two graphs having the same nodes but different edges are employed. Specifically, a weighted joint robust sparse representation model, rather than the commonly used manifold ranking model, helps to compute the saliency value of each node in the first-stage graph, thereby providing a saliency map at the coarse level. In the second-stage graph, along with the adjacently spatial consistency, a new regionally spatial consistency among graph nodes is considered in order to refine the coarse saliency map, assuring uniform saliency assignment even in complex scenes. Particularly, the second stage is generic enough to be integrated in existing salient object detectors, enabling improved performance. Experimental results on benchmark data sets validate the effectiveness and superiority of the proposed scheme over related state-of-the-art methods. Yi Liu 0038, Jungong Han, Qiang Zhang 0020, Long Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Salient object detection based on super-pixel clustering and unified low-rank representation
Qiang Zhang 0020, Yi Liu 0038, Siyang Zhu, Jungong Han |
Comput. Vis. Image Underst. | 2 |