VLDB 2026 Research / reviewers in the wild / expert
Zhi Liu 0003
dblp:40/6686-3
· DBLP profile ↗
151ranked-venue papers
14as first author
50since 2021 · last 2026
0000-0002-8428-1131ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 114 · 12 first-author · 30 since 2021Artificial intelligence and machine learning · 34 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object DetectionabstractRecently segment anything model (SAM) has attracted widespread concerns, and it is often treated as a vision foundation model for universal segmentation. Some researchers have attempted to directly apply the foundation model to the RGB-D video salient object detection (RGB-D VSOD) task, which often encounters three challenges, including the dependence on manual prompts, the high memory consumption of sequential adapters, and the computational burden of memory attention. To address the limitations, we propose a novel method, namely Segment Anything Model with Depth-guided Adaptive Queries (SAM-DAQ), which adapts SAM2 to pop-out salient objects from videos by seamlessly integrating depth and temporal cues within a unified framework. Firstly, we deploy a parallel adapter-based multi-modal image encoder (PAMIE), which incorporates several depth-guided parallel adapters (DPAs) in a skip-connection way. Remarkably, we fine-tune the frozen SAM encoder under prompt-free conditions, where the DPA utilizes depth cues to facilitate the fusion of multi-modal features. Secondly, we deploy a query-driven temporal memory (QTM) module, which unifies the memory bank and prompt embeddings into a learnable pipeline. Concretely, by leveraging both frame-level queries and video-level queries simultaneously, the QTM module can not only selectively extract temporal consistency features but also iteratively update the temporal representations of the queries. Extensive experiments are conducted on three RGB-D VSOD datasets, and the results show that the proposed SAM-DAQ consistently outperforms state-of-the-art methods in terms of all evaluation metrics. Xiaofei Zhou 0003, Runmin Cong, Guodao Zhang, Zhi Liu 0003, Jiyong Zhang 0001 |
AAAI | 6 |
| 2026 | Multiscale Guided Cross-Stage Refinement SAM2 for Seamless Steel Pipe Internal Surface Defect DetectionabstractIn recent years, vision foundation models (e.g., SAM2) have achieved remarkable success across various downstream tasks and have been increasingly adopted in Internet of Things (IoT)-enabled visual perception systems. However, the decoders of existing vision foundation models fail to effectively handle multi-scale contextual cues, leading to unsatisfactory performance in detecting defects of varying sizes in IoT-based inspection scenarios. Besides, the insufficient preservation of spatial information degrades detection accuracy, especially when handling small defect regions, which is critical for reliable perception in IoT devices. To address these issues, we propose a Multi-scale Guided Cross-stage Refinement SAM2 (MGCR-SAM2) for seamless steel pipe (SSP) internal surface defect detection. Specifically, a Multi-scale Information Compensation (MIC) module is designed for multi-scale feature aggregation in the decoder, which integrates contextual information from different receptive fields to enhance the representation of defects with diverse scales. Furthermore, a cross-stage refinement strategy is adopted to progressively enhance detection precision. In the first stage, we construct an encoder–decoder architecture based on SAM2 and the proposed MIC module to adequately capture multi-scale features from SSP images for coarse localization. In the second stage, a Spatial Information Compensation (SIC) module and CNN-based refinement network are employed to enhance local details and refine coarse predictions. Extensive experiments on SSP2000 and SD-Saliency-900 datasets demonstrate the effectiveness of our MGCR-SAM2. Code is available at https://github.com/Kunye-Shen/MGCR-SAM2. Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
IEEE Internet Things J. | 3 |
| 2026 | Double assistant network with semi-supervised learning for schizophrenic atypical visual saliency prediction
Zhengye Wei, Zhi Liu 0003, Lihua Xu, Tianhong Zhang, Jijun Wang 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | SurfSyn: Domain foundation model for pixel-level surface defect detection driven by synthetic data
Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
Pattern Recognit. | 3 |
| 2026 | Condition-guided unified diffusion for unsupervised multi-class anomaly detection
Haocheng Gu, Zhi Liu 0003 |
Signal Process. Image Commun. | 3 |
| 2026 | Class Knowledge-Guided Lightweight Network for Salient Object Detection of Strip Steel Surface DefectabstractRecent lightweight salient object detection (SOD) models for strip steel surface defect typically adopt classification model as encoder to extract semantic features, followed by task-specific decoder for defect detection. Despite achieving promising results, these models often overlook the potential of leveraging class-level knowledge to enhance detection performance. In this paper, we propose a Class Knowledge-Guided Lightweight Network (CKLNet) for salient object detection of strip steel surface defect. Particularly, CKLNet employs a three-stage training strategy. In the first stage, a general SOD model is trained to perform initial detection of surface defect. Then, the second stage constructs a classification branch built upon the general SOD model to extract defect class knowledge. Embarking on this class knowledge, multiple class-specific decoders are trained in the third stage, where we can generate refined and class-aware defect predictions. Extensive experiments on two datasets demonstrate that CKLNet achieves superior detection performance and generalization capability when compared with the cutting-edge lightweight models, with real-time inference speed of 783 FPS on an RTX 2080Ti GPU. Moreover, unlike existing models that only enable defect localization, our CKLNet further provides accurate defect class identification. The source code is publicly available at https://github.com/Kunye-Shen/CKLNet. Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Uncertainty-Aware Audio-Visual Segmentation With Dynamic Fusion for Multimodal AlignmentabstractAudio-visual segmentation (AVS) aims to achieve precise object segmentation by leveraging multimodal cues. However, effective alignment and fusion of audio and visual features are often hindered by inherent uncertainty within multimodal data, such as data quality inconsistencies, semantic mismatches, and temporal or spatial misalignments. To address these challenges, we propose an Uncertainty-aware Audio-Visual Segmentation (UAVS) that dynamically handles uncertainty to improve segmentation accuracy and robustness. Our method employs CLIP-generated text embeddings to provide semantic cues of categories for audio features, reducing ambiguity in multimodal alignment. We then introduce a Mixture of Experts (MoE) model, mapping multimodal embedding samples to multi-dimensional Gaussian distributions to quantify uncertainty through variance and modeling feature confidence using the Gaussian probability density function, effectively capturing noise and semantic discrepancies across modalities. In addition, we design a dynamic path algorithm based on uncertainty, enabling the model to adaptively route samples to experts with high confidence. This algorithm enhances performance in complex, noisy, and ambiguous scenes. Extensive experiments conducted on three subsets of the AVSBench benchmark dataset demonstrate that our proposed method achieves competitive performance. Zhi Liu 0003, Xiaojun Chang |
IEEE Trans. Multim. | 2 |
| 2025 | PLGMNet: Parallel Local-Global Mamba Network for Real-Time Steel Surface Defect Detection
Chenlei Li, Xiaofei Zhou 0003, Yong Wu 0007, Deyang Liu, Jiyong Zhang 0001, Zhi Liu 0003 |
PRCV (17) | 7 |
| 2025 | Few-shot fine-tuning with auxiliary tasks for video anomaly detection
Jing Lv, Zhi Liu 0003, Gongyang Li |
Multim. Syst. | 2 |
| 2025 | Consistency-Queried Transformer for Audio-Visual SegmentationabstractAudio-visual segmentation (AVS) aims to segment objects in audio-visual content. The effective interaction between audio and visual features has garnered significant attention from the multimodal domain. Despite significant advancements, most existing AVS methods are hampered by multimodal inconsistencies. These inconsistencies primarily manifest as a mismatch between audio and visual information guided by audio cues, wherein visual features often dominate audio modality. To address this issue, we propose the Consistency-Queried Transformer (CQFormer), a novel framework for AVS tasks that leverages the transformer architecture. This framework features a Consistency Query Generator (CQG) and a Query-Aligned Matching (QAM) module. The Noise Contrastive Estimation (NCE) loss function enhances modality matching and consistency by minimizing the distributional differences between audio and visual features, facilitating effective fusion and interaction between these features. Additionally, introducing the consistency query during the decoding stage enhances consistency constraints and object-level semantic information, further improving the accuracy and stability of audio-visual segmentation. Extensive experiments on the popular benchmark of the audio-visual segmentation dataset demonstrate that the proposed CQFormer achieves state-of-the-art performance. Zhi Liu 0003, Xiaojun Chang |
IEEE Trans. Image Process. | 2 |
| 2025 | EMS: A Large-Scale Eye Movement Dataset, Benchmark, and New Model for Schizophrenia RecognitionabstractSchizophrenia (SZ) is a common and disabling mental illness, and most patients encounter cognitive deficits. The eye-tracking technology has been increasingly used to characterize cognitive deficits for its reasonable time and economic costs. However, there is no large-scale and publicly available eye movement dataset and benchmark for SZ recognition. To address these issues, we release a large-scale Eye Movement dataset for SZ recognition (EMS), which consists of eye movement data from 104 schizophrenics and 104 healthy controls (HCs) based on the free-viewing paradigm with 100 stimuli. We also conduct the first comprehensive benchmark, which has been absent for a long time in this field, to compare the related 13 psychosis recognition methods using six metrics. Besides, we propose a novel mean-shift-based network (MSNet) for eye movement-based SZ recognition, which elaborately combines the mean shift algorithm with convolution to extract the cluster center as the subject feature. In MSNet, first, a stimulus feature branch (SFB) is adopted to enhance each stimulus feature with similar information from all stimulus features, and then, the cluster center branch (CCB) is utilized to generate the cluster center as subject feature and update it by the mean shift vector. The performance of our MSNet is superior to prior contenders, thus, it can act as a powerful baseline to advance subsequent study. To pave the road in this research field, the EMS dataset, the benchmark results, and the code of MSNet are publicly available at https://github.com/YingjieSong1/EMS. Zhi Liu 0003, Gongyang Li, Qiang Wu 0001, Dan Zeng 0001, Lihua Xu, Tianhong Zhang, Jijun Wang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Distribution Alignment for Fully Test-Time Adaptation with Dynamic Online Data Streams
Ziqiang Wang 0003, Zhixiang Chi, Li Gu, Zhi Liu 0003, Konstantinos N. Plataniotis, Yang Wang 0003 |
ECCV (24) | 5 |
| 2024 | Audio-visual saliency prediction with multisensory perception and integration
Zhi Liu 0003, Gongyang Li |
Image Vis. Comput. | 2 |
| 2024 | Global semantic-guided network for saliency prediction
Zhi Liu 0003, Gongyang Li |
Knowl. Based Syst. | 2 |
| 2024 | Masked feature regeneration based asymmetric student-teacher network for anomaly detection
Haocheng Gu, Gongyang Li, Zhi Liu 0003 |
Multim. Tools Appl. | 3 |
| 2024 | Light Field Salient Object Detection With Sparse Views via Complementary and Discriminative Interaction Networkabstract4D light field data record the scene from multiple views, thus implicitly providing beneficial depth cue for salient object detection in challenging scenes. Existing light field salient object detection (LF SOD) methods usually use a large number of views to improve the detection accuracy. However, using so many views for LF SOD brings difficulties to its practical applications. Considering that adjacent views in a light field are actually with very similar contents, in this work, we propose defining a more efficient pattern of input views, i. e., key sparse views, and design a network to effectively explore the depth cue from sparse views for LF SOD. Specifically, we firstly introduce a low rank-based statistical analysis to the existing LF SOD datasets, which allows us to conclude a fixed yet universal pattern for our key sparse views, including the number and positions of views. These views maintain the sufficient depth cue, but greatly lower the number of views to be captured and processed, facilitating practical applications. Then, we propose an effective solution with a key Complementary and Discriminative Interaction Module (CDIM) for LF SOD from key sparse views, named CDINet. The CDINet follows a two-stream structure to extract the depth cue from the light field stream (i. e., sparse views) and the appearance cue from the RGB stream (i. e., center view), generating features and initial saliency maps for each stream. The CDIM is tailored for inter-stream interaction of both these features and saliency maps, using the depth cue to complement the missing salient regions in RGB stream and discriminate the background distraction, to enhance the final saliency map further. Extensive experiments on three LF multi-view datasets demonstrate that our CDINet not only outperforms the state-of-the-art 2D methods, but also achieves competitive performance as compared with the state-of-the-art 3D and 4D methods. The code and results of our method are available athttps://github.com/GilbertRC/LFSOD-CDINet. Gongyang Li, Ping An 0001, Zhi Liu 0003, Xinpeng Huang, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | TTAGaze: Self-Supervised Test-Time Adaptation for Personalized Gaze EstimationabstractIn this paper, we address the problem of personalized gaze estimation. Due to the anatomical differences between individuals, current personalized gaze models often rely on fine-tuning or fully-supervised methods with labeled calibration samples, which may not be practical in real-world applications. To tackle this limitation, we propose an approach called Self-Supervised Test-Time Adaptation for Personalized Gaze Estimation (TTAGaze), which enables adaptation with small unlabeled data at test time. Our goal is to develop a gaze estimation model specifically adapted to a target person using only a few unlabeled images. We call this setting as unsupervised few-shot personalized adaptation in gaze estimation, which is more aligned with real-world scenarios compared to existing approaches. Additionally, Our approach leverages self-supervised learning and meta-learning. The model consists of the main task (gaze estimation) and a self-supervised auxiliary task. During training, the two task are trained using a coupled method. At test time, adaptation is achieved by optimizing the self-supervised loss adapted to an unseen person with a few unlabeled data. The model parameters are learned via model-agnostic meta-learning (MAML) to facilitate effective unsupervised few-shot personalized adaptation in gaze estimation. Experimental results demonstrate that the proposed method outperforms alternative approaches on several widely-used benchmark datasets. Yong Wu 0007, Guang Chen 0001, Linwei Ye, Yuanning Jia, Zhi Liu 0003, Yang Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | MINet: Multiscale Interactive Network for Real-Time Salient Object Detection of Strip Steel Surface DefectsabstractThe automated surface defect detection is a fundamental task in industrial production, and the existing saliency-based works overcome the challenging scenes and give promising detection results. However, the cutting-edge efforts often suffer from large parameter size, heavy computational cost, and slow inference speed, which heavily limits the practical applications. To this end, we devise a multiscale interactive (MI) module, which employs depthwise convolution (DWConv) and pointwise convolution (PWConv) to independently extract and interactively fuse features of different scales, respectively. Particularly, the MI module can provide satisfactory characterization for defect regions with fewer parameters. Embarking on this module, we propose a lightweight multiscale interactive network (MINet) to conduct real-time salient object detection of strip steel surface defects. Comprehensive experimental results on SD-Saliency-900 dataset, which contains three kinds of strip steel surface defect detection images (i.e., inclusion, patches, and scratches), demonstrate that the proposed MINet presents comparable detection accuracy with the state-of-the-art methods while running at a GPU speed of 721 FPS and a CPU speed of 6.3 FPS for 368×368 images with only 0.28 M parameters. Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Context-Aware Interaction Network for RGB-T Semantic SegmentationabstractRGB-T semantic segmentation is a key technique for autonomous driving scenes understanding. For the existing RGB-T semantic segmentation methods, however, the effective exploration of the complementary relationship between different modalities is not implemented in the information interaction between multiple levels. To address such an issue, the Context-Aware Interaction Network (CAINet) is proposed for RGB-T semantic segmentation, which constructs interaction space to exploit auxiliary tasks and global context for explicitly guided learning. Specifically, we propose a Context-Aware Complementary Reasoning (CACR) module aimed at establishing the complementary relationship between multimodal features with the long-term context in both spatial and channel dimensions. Further, considering the importance of global contextual and detailed information, we propose the Global Context Modeling (GCM) module and Detail Aggregation (DA) module, and we introduce specific auxiliary supervision to explicitly guide the context interaction and refine the segmentation map. Extensive experiments on two benchmark datasets of MFNet and PST900 demonstrate that the proposed CAINet achieves state-of-the-art performance. The code is available athttps://github.com/YingLv1106/CAINet. Zhi Liu 0003, Gongyang Li |
IEEE Trans. Multim. | 2 |
| 2024 | ADMNet: Attention-Guided Densely Multi-Scale Network for Lightweight Salient Object DetectionabstractRecently, benefitting from the rapid development of deep learning technology, the research of salient object detection has achieved great progress. However, the performance of existing cutting-edge saliency models relies on large network size and high computational overhead. This is unamiable to real-world applications, especially the practical platforms with low cost and limited computing resources. In this paper, we propose a novel lightweight saliency model, namely Attention-guided Densely Multi-scale Network (ADMNet), to tackle this issue. Firstly, we design the multi-scale perception (MP) module to acquire different contextual features by using different receptive fields. Embarking on MP module, we build the encoder of our model, where each convolutional block adopts a dense structure to connect MP modules. Following this way, our model can provide powerful encoder features for the characterization of salient objects. Secondly, we employ dual attention (DA) module to equip the decoder blocks. Particularly, in DA module, the binarized coarse saliency inference of the decoder block (i.e., a hard spatial attention map) is first employed to filter out interference cues from the decoder feature, and then by introducing large receptive fields, the enhanced decoder feature is used to generate a soft spatial attention map, which further purifies the fused features. Following this way, the deep features are steered to give more concerns to salient regions. Extensive experiments on five public challenging datasets including ECSSD, DUT-OMRON, DUTS-TE, HKU-IS, and PASCAL-S clearly show that our model achieves comparable performance with the state-of-the-art saliency models while running at a 219.4fps GPU speed and a 1.76fps CPU speed for a 368×368 image with only 0.84 M parameters. Xiaofei Zhou 0003, Kunye Shen, Zhi Liu 0003 |
IEEE Trans. Multim. | 3 |
| 2023 | Few-Shot Learning of Compact Models via Task-Specific Meta DistillationabstractWe consider a new problem of few-shot learning of com-pact models. Meta-learning is a popular approach for few-shot learning. Previous work in meta-learning typically assumes that the model architecture during meta-training is the same as the model architecture used for final deployment. In this paper, we challenge this basic assumption. For final deployment, we often need the model to be small. But small models usually do not have enough capacity to effectively adapt to new tasks. In the mean time, we often have access to the large dataset and extensive computing power during meta-training since meta-training is typically per-formed on a server. In this paper, we propose task-specific meta distillation that simultaneously learns two models in meta-learning: a large teacher model and a small student model. These two models are jointly learned during meta-training. Given a new task during meta-testing, the teacher model is first adapted to this task, then the adapted teacher model is used to guide the adaptation of the student model. The adapted student model is used for final deployment. We demonstrate the effectiveness of our approach in few-shot image classification using model-agnostic meta-learning (MAML). Our proposed method outperforms other alternatives on several benchmark datasets. Yong Wu 0007, Shekhor Chanda, Mehrdad Hosseinzadeh, Zhi Liu 0003, Yang Wang 0003 |
WACV | 4 |
| 2023 | Exploring viewport features for semi-supervised saliency prediction in omnidirectional images
Mengke Huang, Gongyang Li, Zhi Liu 0003, Yong Wu 0007, Chen Gong 0002, Linchao Zhu, Yi Yang 0001 |
Image Vis. Comput. | 3 |
| 2023 | Atypical Salient Regions Enhancement Network for visual saliency prediction of individuals with Autism Spectrum Disorder
Huizhan Duan, Zhi Liu 0003, Weijie Wei 0001, Tianhong Zhang, Jijun Wang 0003, Lihua Xu, Haichun Liu |
Signal Process. Image Commun. | 2 |
| 2023 | Lightweight Distortion-Aware Network for Salient Object Detection in Omnidirectional ImagesabstractCompared with 2D image salient object detection (SOD), SOD in omnidirectional images (or 360° images) usually suffers from geometric distortion. Although existing omnidirectional image SOD (ODI-SOD) methods have improved the detection accuracy obviously, their application may be cumbersome in real scenes due to their high computational cost. To avoid distortion and reduce the computational cost simultaneously in ODI-SOD, we propose a novel lightweight distortion-aware network, named LDNet, in this letter. First, to extract features with less distortion from ODIs, we integrate the distortion-aware convolution and depth-wise separable convolution (DSConv) into distortion-aware DSConv (DDSConv) and replace the regular convolutions in the last two blocks of the ResNet-18 with DDSConvs to obtain our lightweight backbone network (LD-ResNet-18). To enhance spatial information in each channel of the extracted features at each level comprehensively, then, we propose a lightweight distortion-aware channel-wise enhancement (DCE) module (only 0.05M parameters) including DDSConvs with various dilation rates, channel shuffle operation and attention mechanism, and employ a high-to-low dense connection structure to modulate the enhanced multi-level features. Besides, we design a distortion-aware self-correlation (DSC) module (only 0.02M parameters) for mining the contextual dependency of the features via a coarse-fine strategy, and the correlated features are refined by DCE modules and integrated by another dense connection structure. The final saliency map is predicted from the densely integrated features. Compared with 12 state-of-the-art methods on two public datasets, our lightweight LDNet achieves competitive or even better performance with only 2.9M parameters and 3.4G FLOPs, which balances the efficiency and performance. Mengke Huang, Gongyang Li, Zhi Liu 0003, Linchao Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | RGB-T Semantic Segmentation With Location, Activation, and SharpeningabstractSemantic segmentation is important for scene understanding. To address the scenes of adverse illumination conditions of natural images, thermal infrared (TIR) images are introduced. Most existing RGB-T semantic segmentation methods follow three cross-modal fusion paradigms, i. e., encoder fusion, decoder fusion, and feature fusion. Some methods, unfortunately, ignore the properties of RGB and TIR features or the properties of features at different levels. In this paper, we propose a novel feature fusion-based network for RGB-T semantic segmentation, named LASNet, which follows three steps of location, activation, and sharpening. The highlight of LASNet is that we fully consider the characteristics of cross-modal features at different levels, and accordingly propose three specific modules for better segmentation. Concretely, we propose a Collaborative Location Module (CLM) for high-level semantic features, aiming to locate all potential objects. We propose a Complementary Activation Module for middle-level features, aiming to activate exact regions of different objects. We propose an Edge Sharpening Module (ESM) for low-level texture features, aiming to sharpen the edges of objects. Furthermore, in the training phase, we attach a location supervision and an edge supervision after CLM and ESM, respectively, and impose two semantic supervisions in the decoder part to facilitate network convergence. Experimental results on two public datasets demonstrate that the superiority of our LASNet over relevant state-of-the-art methods. The code and results of our method are available athttps://github.com/MathLee/LASNet. Gongyang Li, Yike Wang 0003, Zhi Liu 0003, Xinpeng Zhang 0001, Dan Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | SGFNet: Semantic-Guided Fusion Network for RGB-Thermal Semantic SegmentationabstractRecently, semantic segmentation based on RGB and thermal infrared (TIR) images has become a research hotspot because of its stability in the weak light environment. However, most of the current methods ignore the differences between the two modalities of data and do not use semantic information in multi-modal fusion. In this paper, we propose a novel Semantic-Guided Fusion Network (SGFNet) for RGB-Thermal semantic segmentation, which makes full use of semantic information in the multi-modal fusion. Our SGFNet consists of an asymmetric encoder with TIR branch and RGB branch and a decoder. We concentrate on enhancing the multi-modal feature representation in the encoder with a pattern of fusion and enhancement. Specifically, considering that TIR images are stable under weak light conditions, we first propose a Semantic Guidance Head to extract semantic information in the TIR branch. In the RGB branch, we propose a Multi-modal Coordination and Distillation Unit to fuse multi-modal features first. Then, we propose a Cross-level and Semantic-guided Enhancement Unit to enhance the fused features with cross-level information and semantic information. We arrange these two units at all stages of the RGB branch to generate features with strong representation abilities at different levels. For the decoder, to obtain large receptive fields and fine edges, we improve the Lawin ASPP decoder by introducing edge information extracted from the low-level features, proposing the edge-aware Lawin ASPP decoder. With our encoder and decoder working together, our SGFNet can identify objects accurately and segment objects finely. Extensive experiments on the MFNet dataset demonstrate the superior performance of the proposed SGFNet compared with state-of-the-art methods. The code and results of our method are available athttps://github.com/kw717/SGFNet. Yike Wang 0003, Gongyang Li, Zhi Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Adjacent Context Coordination Network for Salient Object Detection in Optical Remote Sensing ImagesabstractSalient object detection (SOD) in optical remote sensing images (RSIs), or RSI-SOD, is an emerging topic in understanding optical RSIs. However, due to the difference between optical RSIs and natural scene images (NSIs), directly applying NSI-SOD methods to optical RSIs fails to achieve satisfactory results. In this article, we propose a novel adjacent context coordination network (ACCoNet) to explore the coordination of adjacent features in an encoder-decoder architecture for RSI-SOD. Specifically, ACCoNet consists of three parts: 1) an encoder; 2) adjacent context coordination modules (ACCoMs); and 3) a decoder. As the key component of ACCoNet, ACCoM activates the salient regions of output features of the encoder and transmits them to the decoder. ACCoM contains a local branch and two adjacent branches to coordinate the multilevel features simultaneously. The local branch highlights the salient regions in an adaptive way, while the adjacent branches introduce global information of adjacent levels to enhance salient regions. In addition, to extend the capabilities of the classic decoder block (i.e., several cascaded convolutional layers), we extend it with two bifurcations and propose a bifurcation-aggregation block (BAB) to capture the contextual information in the decoder. Extensive experiments on two benchmark datasets demonstrate that the proposed ACCoNet outperforms 22 state-of-the-art methods under nine evaluation metrics, and runs up to 81 fps on a single NVIDIA Titan X GPU. The code and results of our method are available at https://github.com/MathLee/ACCoNet. Gongyang Li, Zhi Liu 0003, Dan Zeng 0001, Weisi Lin, Haibin Ling |
IEEE Trans. Cybern. | 2 |
| 2023 | Lightweight Salient Object Detection in Optical Remote-Sensing Images via Semantic Matching and Edge AlignmentabstractRecently, relying on convolutional neural networks (CNNs), many methods for salient object detection in optical remote-sensing images (ORSI-SOD) are proposed. However, most methods ignore the number of parameters and computational cost brought by CNNs, and only a few pay attention to portability and mobility. To facilitate practical applications, in this article, we propose a novel lightweight network for ORSI-SOD based on semantic matching and edge alignment, termed SeaNet. Specifically, SeaNet includes a lightweight MobileNet-V2 for feature extraction, a dynamic semantic matching module (DSMM) for high-level features, an edge self-alignment module (ESAM) for low-level features, and a portable decoder for inference. First, the high-level features are compressed into semantic kernels. Then, semantic kernels are used to activate salient object locations in two groups of high-level features through dynamic convolution operations in DSMM. Meanwhile, in ESAM, cross-scale edge information extracted from two groups of low-level features is self-aligned through$L_{2}$loss and used for detail enhancement. Finally, starting from the highest level features, the decoder infers salient objects based on the accurate locations and fine details contained in the outputs of the two modules. Extensive experiments on two public datasets demonstrate that our lightweight SeaNet not only outperforms most state-of-the-art lightweight methods, but also yields comparable accuracy with state-of-the-art conventional methods, while having only 2.76 M parameters and running with 1.7 G floating point operations (FLOPs) for$288 \times 288$inputs. Our code and results are available athttps://github.com/MathLee/SeaNet. Gongyang Li, Zhi Liu 0003, Xinpeng Zhang 0001, Weisi Lin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Salient Object Detection in Optical Remote Sensing Images Driven by TransformerabstractExisting methods for Salient Object Detection in Optical Remote Sensing Images (ORSI-SOD) mainly adopt Convolutional Neural Networks (CNNs) as the backbone, such as VGG and ResNet. Since CNNs can only extract features within certain receptive fields, most ORSI-SOD methods generally follow the local-to-contextual paradigm. In this paper, we propose a novel Global Extraction Local Exploration Network (GeleNet) for ORSI-SOD following the global-to-local paradigm. Specifically, GeleNet first adopts a transformer backbone to generate four-level feature embeddings with global long-range dependencies. Then, GeleNet employs a Direction-aware Shuffle Weighted Spatial Attention Module (D-SWSAM) and its simplified version (SWSAM) to enhance local interactions, and a Knowledge Transfer Module (KTM) to further enhance cross-level contextual interactions. D-SWSAM comprehensively perceives the orientation information in the lowest-level features through directional convolutions to adapt to various orientations of salient objects in ORSIs, and effectively enhances the details of salient objects with an improved attention mechanism. SWSAM discards the direction-aware part of D-SWSAM to focus on localizing salient objects in the highest-level features. KTM models the contextual correlation knowledge of two middle-level features of different scales based on the self-attention mechanism, and transfers the knowledge to the raw features to generate more discriminative features. Finally, a saliency predictor is used to generate the saliency map based on the outputs of the above three modules. Extensive experiments on three public datasets demonstrate that the proposed GeleNet outperforms relevant state-of-the-art methods. The code and results of our method are available at https://github.com/MathLee/GeleNet. Gongyang Li, Zhen Bai 0001, Zhi Liu 0003, Xinpeng Zhang 0001, Haibin Ling |
IEEE Trans. Image Process. | 3 |
| 2023 | Adaptive Group-Wise Consistency Network for Co-Saliency DetectionabstractCo-saliency detection focuses on detecting common and salient objects among a group of images. With the application of deep learning in co-saliency detection, more accurate and more effective models are proposed in an end-to-end manner. However, two major drawbacks in these models hinder the further performance improvement of co-saliency detection: 1) the static manner-based inference, and 2) the constant quantity of input images. To address these limitations, we present a novel Adaptive Group-wise Consistency Network (AGCNet) with the ability of content-adaptive adjustment for a given image group with random quantity of images. In AGCNet, we first introduce intra-saliency priors generated from any off-the-shelf salient object detection model. Then, an Adaptive Group-wise Consistency (AGC) module is proposed to capture group consistency for each individual image, and is applied on three-scale features to capture the group consistency from different perspectives. This module is composed of two key components, where the content-adaptive group consistency block breaks the above limitations to adaptively capture the global group consistency with the assistance of intra-saliency priors and the ranking-based fusion block combines the consistency with individual attributes of each image feature to generate discriminative group consistency feature for each image. Following AGC modules, a specially designed Aggregated Decoder aggregates the three-scale group consistency features to adapt to co-salient objects with diverse scales for preliminary detection. Finally, we incorporate two normal decoders to progressively refine the preliminary detection and generate the final co-saliency maps. Extensive experiments on four benchmark datasets demonstrate that our AGCNet achieves competitive performance as compared with 19 state-of-the-art models, and the proposed modules experimentally show substantial practical merits. Zhen Bai 0001, Zhi Liu 0003, Gongyang Li, Yang Wang 0003 |
IEEE Trans. Multim. | 2 |
| 2023 | RINet: Relative Importance-Aware Network for Fixation PredictionabstractFixation prediction aims to simulate human visual selection mechanism and estimate the visual saliency degree of regions in a scene. In semantically rich scenes, there are generally multiple salient regions. This condition requires a fixation prediction model to understand the relative importance relationship of multiple salient regions, that is, to identify which region is more important. In practice, existing fixation prediction models implicitly explore the relative importance relationship in the end-to-end training process while they do not work well. In this article, we propose a novel Relative Importance-aware Network (RINet) to explicitly explore the modeling of relative importance in fixation prediction. RINet perceives multi-scale local and global relative importance through the Hierarchical Relative Importance Enhancement (HRIE) module. Within a single scale subspace, on the one hand, HRIE module regards the similarity matrix as the local relative importance map to weight the input feature. On the other hand, HRIE module integrates a set of local relative importance maps into one map, defined as the global relative importance map, to grasp global relative importance. Moreover, we propose a Complexity-Relevant Focal (CRF) loss for network training. As such, we can progressively emphasize learning difficult samples for better handling the complicated scenarios, further improving the performance. The ablation studies confirm the contributions of key components of our RINet, and extensive experiments on five datasets demonstrate our RINet is superior to 28 relevant state-of-the-art models. Zhi Liu 0003, Gongyang Li, Dan Zeng 0001, Tianhong Zhang, Lihua Xu, Jijun Wang 0003 |
IEEE Trans. Multim. | 2 |
| 2023 | Spatio-Temporal Self-Attention Network for Video Saliency Predictionabstract3D convolutional neural networks have achieved promising results for video tasks in computer vision, including video saliency prediction that is explored in this paper. However, 3D convolution encodes visual representation merely on fixed local spacetime according to its kernel size, while human attention is always attracted by relational visual features at different time. To overcome this limitation, we propose a novel Spatio-Temporal Self-Attention 3D Network (STSANet) for video saliency prediction, in which multiple Spatio-Temporal Self-Attention (STSA) modules are employed at different levels of 3D convolutional backbone to directly capture long-range relations between spatio-temporal features of different time steps. Besides, we propose an Attentional Multi-Scale Fusion (AMSF) module to integrate multi-level features with the perception of context in semantic and spatio-temporal subspaces. Extensive experiments demonstrate the contributions of key components of our method, and the results on DHF1K, Hollywood-2, UCF, and DIEM benchmark datasets clearly prove the superiority of the proposed model compared with all state-of-the-art models. Ziqiang Wang 0003, Zhi Liu 0003, Gongyang Li, Yang Wang 0003, Tianhong Zhang, Lihua Xu, Jijun Wang 0003 |
IEEE Trans. Multim. | 2 |
| 2022 | Singing Voice Synthesis with Vibrato Modeling and Latent Energy RepresentationabstractThis paper proposes an expressive singing voice synthesis system by introducing explicit vibrato modeling and latent energy representation. Vibrato is essential to the naturalness of synthesized sound, due to the inherent characteristics of human singing. Hence, a deep learning-based vibrato model is introduced in this paper to control the vibrato's likeliness, rate, depth and phase in singing, where the vibrato likeliness represents the existence probability of vibrato and it would help improve the singing voice's naturalness. Actually, there is no annotated label about vibrato likeliness in existing singing corpus. We adopt a novel vibrato likeliness labeling method to label the vibrato likeliness automatically. Meanwhile, the power spectrogram of audio contains rich information that can improve the expressiveness of singing. An autoencoder-based latent energy bottleneck feature is proposed for expressive singing voice synthesis. Experimental results on the open dataset NUS48E show that both the vibrato modeling and the latent energy representation could significantly improve the expressiveness of singing voice. The audio samples are shown in the demo website11https://mango321321.github.io/ExpressiveSing/. Wei Zhang 0031, Zhengchen Zhang, Dan Zeng 0001, Zhi Liu 0003 |
MMSP | 6 |
| 2022 | Few-shot personalized saliency prediction using meta-learning
Xinhui Luo, Zhi Liu 0003, Weijie Wei 0001, Linwei Ye, Tianhong Zhang, Lihua Xu, Jijun Wang 0003 |
Image Vis. Comput. | 2 |
| 2022 | Referring Segmentation in Images and Videos With Cross-Modal Self-Attention NetworkabstractWe consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In this paper, we propose a cross-modal self-attention (CMSA) module to utilize fine details of individual words and the input image or video, which effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the visual input. We further propose a gated multi-level fusion (GMLF) module to selectively integrate self-attentive cross-modal features corresponding to different levels of visual features. This module controls the feature fusion of information flow of features at different levels with high-level and low-level semantic information related to different attentive words. Besides, we introduce cross-frame self-attention (CFSA) module to effectively integrate temporal information in consecutive frames which extends our method in the case of referring segmentation in videos. Experiments on benchmark datasets of four referring image datasets and two actor and action video segmentation datasets consistently demonstrate that our proposed approach outperforms existing state-of-the-art methods. Linwei Ye, Mrigank Rochan, Zhi Liu 0003, Xiaoqin Zhang 0002, Yang Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Gaze Estimation via Modulation-Based Adaptive Network With Auxiliary Self-LearningabstractGiven a face image, most of previous works in gaze estimation infer the gaze via a well-trained model with supervised training. However, the distribution of test data may be very different compared to that of training data since samples might be corrupted in real-world scenarios (e.g., taking a photo in strong light). This will lead to a gap between source domain (i.e., training data) and target domain (i.e., test data). In this paper, we first introduce self-supervised learning into our method for addressing challenging situations in gaze estimation. Moreover, existing appearance-based gaze estimation methods focus on directing towards the development of powerful regressors, which mainly utilize face and eye images simultaneously or face (eye) images only. However, the problem of inter cues between face and eye features has been largely overlooked. To this end, we propose a novel Modulation-based Adaptive Network (MANet) for gaze estimation, which uses high-level knowledge to filter the distractive information and bridges the intrinsic relationship between face and eye features. Further, we combine self-supervised learning and MANet to learn to adapt to challenging cases, such as abnormal lighting conditions and poor-quality images, by minimizing a self-supervised loss and a supervised loss jointly. The experimental results on several datasets demonstrate the effectiveness of our proposed approach with a real-time speed of 900fpson a PC with an NVIDIA Titan RTX GPU. Yong Wu 0007, Gongyang Li, Zhi Liu 0003, Mengke Huang, Yang Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Lightweight Salient Object Detection in Optical Remote Sensing Images via Feature CorrelationabstractSalient object detection in optical remote sensing images (ORSI-SOD) has been widely explored for understanding ORSIs. However, previous methods focus mainly on improving the detection accuracy while neglecting the cost in memory and computation, which may hinder their real-world applications. In this article, we propose a novel lightweight ORSI-SOD solution, named CorrNet, to address these issues. In CorrNet, we first lighten the backbone (VGG-16) and build a lightweight subnet for feature extraction. Then, following the coarse-to-fine strategy, we generate an initial coarse saliency map from high-level semantic features in a correlation module (CorrM). The coarse saliency map serves as the location guidance for low-level features. In CorrM, we mine the object location information between high-level semantic features through the cross-layer correlation operation. Finally, based on low-level detailed features, we refine the coarse saliency map in the refinement subnet equipped with dense lightweight refinement blocks (DLRBs) and produce the final fine saliency map. By reducing the parameters and computations of each component, CorrNet ends up having only 4.09M parameters and running with 21.09G FLOPs. Experimental results on two public datasets demonstrate that our lightweight CorrNet achieves competitive or even better performance compared with 26 state-of-the-art methods (including 16 large CNN-based methods and two lightweight methods), and meanwhile enjoys the clear memory and run-time efficiency. The code and results of our method are available athttps://github.com/MathLee/CorrNet. Gongyang Li, Zhi Liu 0003, Zhen Bai 0001, Weisi Lin, Haibin Ling |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multi-Content Complementation Network for Salient Object Detection in Optical Remote Sensing ImagesabstractIn the computer vision community, great progresses have been achieved in salient object detection from natural scene images (NSI-SOD); by contrast, salient object detection in optical remote sensing images (RSI-SOD) remains to be a challenging emerging topic. The unique characteristics of optical RSIs, such as scales, illuminations, and imaging orientations, bring significant differences between NSI-SOD and RSI-SOD. In this article, we propose a novel multi-content complementation network (MCCNet) to explore the complementarity of multiple content for RSI-SOD. Specifically, MCCNet is based on the general encoder–decoder architecture, and contains a novel key component named multi-content complementation module (MCCM), which bridges the encoder and the decoder. In MCCM, we consider multiple types of features that are critical to RSI-SOD, including foreground features, edge features, background features, and global image-level features, and exploit the content complementarity between them to highlight salient regions over various scales in RSI features through the attention mechanism. Besides, we comprehensively introduce pixel-level, map-level, and metric-aware losses in the training phase. Extensive experiments on two popular datasets demonstrate that the proposed MCCNet outperforms 23 state-of-the-art methods, including both NSI-SOD and RSI-SOD methods. The code and results of our method are available athttps://github.com/MathLee/MCCNet. Gongyang Li, Zhi Liu 0003, Weisi Lin, Haibin Ling |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Edge-Aware Multiscale Feature Integration Network for Salient Object Detection in Optical Remote Sensing ImagesabstractThe optical remote sensing images (RSIs) show various spatial resolutions and cluttered background, where salient objects with different scales, types, and orientations are presented in diverse RSI scenes. Therefore, it is inappropriate to directly extend cutting-edge saliency detection methods for conventional RGB images to optical RSIs. Besides, the existing saliency models targeting RSIs often render imperfect saliency maps, where some of them are with coarse boundary details. To solve this problem, this article attempts to introduce the edge information to precisely detect salient objects in RSIs. Accordingly, we propose an edge-aware multiscale feature integration network (EMFI-Net) for salient object detection by conducting multiscale feature integration under the explicit and implicit assistance of salient edge cues. Specifically, our network contains two parts including the encoder and decoder. First, the encoder extracts multiscale deep features from three RSIs with different resolutions, where the high-level deep semantic features from three RSIs are integrated using a cascaded feature fusion module. Second, the encoder explicitly enriches the multiscale deep features by integrating the salient edge cues extracted by a salient edge extraction module. Meanwhile, we also implicitly deploy an edge-aware constraint to the supervision of the saliency map prediction by introducing a hybrid loss function. Finally, the decoder integrates the enriched multiscale deep features in a coarse-to-fine way, yielding a high-quality saliency map. The experiments conducted on two public optical RSI datasets clearly prove the effectiveness and superiority of the proposed EMFI-Net against the state-of-the-art saliency models. Xiaofei Zhou 0003, Kunye Shen, Zhi Liu 0003, Chen Gong 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Circular Complement Network for RGB-D Salient Object Detection
Zhen Bai 0001, Zhi Liu 0003, Gongyang Li, Linwei Ye, Yang Wang 0003 |
Neurocomputing | 2 |
| 2021 | Deep saliency models : The quest for the loss function
Alexandre Bruckert, Hamed Rezazadegan Tavakoli, Zhi Liu 0003, Marc Christie, Olivier Le Meur |
Neurocomputing | 3 |
| 2021 | Personalized image observation behavior learning in fixation based personalized salient object segmentation
Gongyang Li, Weijie Wei 0001, Xiaofei Zhou 0003, Zhi Liu 0003 |
Neurocomputing | 5 |
| 2021 | Predicting atypical visual saliency for autism spectrum disorder via scale-adaptive inception module and discriminative region enhancement loss
Weijie Wei 0001, Zhi Liu 0003, Lijin Huang, Alexis Nebout, Olivier Le Meur, Tianhong Zhang, Jijun Wang 0003, Lihua Xu |
Neurocomputing | 2 |
| 2021 | SalED: Saliency prediction with a pithy encoder-decoder architecture sensing local and global information
Ziqiang Wang 0003, Zhi Liu 0003, Weijie Wei 0001, Huizhan Duan |
Image Vis. Comput. | 2 |
| 2021 | ATCC: Accurate tracking by criss-cross location attention
Yong Wu 0007, Zhi Liu 0003, Xiaofei Zhou 0003, Linwei Ye, Yang Wang 0003 |
Image Vis. Comput. | 2 |
| 2021 | Identify autism spectrum disorder via dynamic filter and deep spatiotemporal feature extraction
Weijie Wei 0001, Zhi Liu 0003, Lijin Huang, Ziqiang Wang 0003, Tianhong Zhang, Jijun Wang 0003, Lihua Xu |
Signal Process. Image Commun. | 2 |
| 2021 | EDBGAN: Image Inpainting via an Edge-Aware Dual Branch Generative Adversarial NetworkabstractAs deep learning technology develops rapidly, image inpainting methods have made significant progress in generating reasonable contents for images with large and irregular holes. Nevertheless, existing methods either use one encoder-decoder to generate results in one step without making good use of structure features which are helpful, or adopt two encoder-decoders to recover structures and textures subsequently where the second encoder-decoder used for recovering textures relies heavily on the first encoder-decoder. Thus, an one-stage method which simultaneously utilizes structures and textures is promising. In this letter, we propose a dual branch encoder-decoder, whose texture branch and edge branch can simultaneously extract features from the masked RGB image and its corresponding edge map. Moreover, we propose a lightweight mutually guided attention block (MGAB), which makes features of two branches guide each other from low level to high level. At the end of two branches, we apply dual attention block (DAB) to perform feature fusion, which uses self-attention mechanism and lightweight attention mechanism on features of texture branch and edge branch, respectively. Extensive experiments demonstrate that our proposed method is effective in recovering edges and textures and achieves the state-of-the-art performance. Minyu Chen 0001, Zhi Liu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Hierarchical Alternate Interaction Network for RGB-D Salient Object DetectionabstractExisting RGB-D Salient Object Detection (SOD) methods take advantage of depth cues to improve the detection accuracy, while pay insufficient attention to the quality of depth information. In practice, a depth map is often with uneven quality and sometimes suffers from distractors, due to various factors in the acquisition procedure. In this article, to mitigate distractors in depth maps and highlight salient objects in RGB images, we propose a Hierarchical Alternate Interactions Network (HAINet) for RGB-D SOD. Specifically, HAINet consists of three key stages: feature encoding, cross-modal alternate interaction, and saliency reasoning. The main innovation in HAINet is the Hierarchical Alternate Interaction Module (HAIM), which plays a key role in the second stage for cross-modal feature interaction. HAIM first uses RGB features to filter distractors in depth features, and then the purified depth features are exploited to enhance RGB features in turn. The alternate RGB-depth-RGB interaction proceeds in a hierarchical manner, which progressively integrates local and global contexts within a single feature scale. In addition, we adopt a hybrid loss function to facilitate the training of HAINet. Extensive experiments on seven datasets demonstrate that our HAINet not only achieves competitive performance as compared with 19 relevant state-of-the-art methods, but also reaches a real-time processing speed of 43 fps on a single NVIDIA Titan X GPU. The code and results of our method are available at https://github.com/MathLee/HAINet. Gongyang Li, Zhi Liu 0003, Minyu Chen 0001, Zhen Bai 0001, Weisi Lin, Haibin Ling |
IEEE Trans. Image Process. | 2 |
| 2021 | Personal Fixations-Based Object Segmentation With Object Localization and Boundary PreservationabstractAs a natural way for human-computer interaction, fixation provides a promising solution for interactive image segmentation. In this paper, we focus on Personal Fixations-based Object Segmentation (PFOS) to address issues in previous studies, such as the lack of appropriate dataset and the ambiguity in fixations-based interaction. In particular, we first construct a new PFOS dataset by carefully collecting pixel-level binary annotation data over an existing fixation prediction dataset, such dataset is expected to greatly facilitate the study along the line. Then, considering characteristics of personal fixations, we propose a novel network based on Object Localization and Boundary Preservation (OLBP) to segment the gazed objects. Specifically, the OLBP network utilizes an Object Localization Module (OLM) to analyze personal fixations and locates the gazed objects based on the interpretation. Then, a Boundary Preservation Module (BPM) is designed to introduce additional boundary information to guard the completeness of the gazed objects. Moreover, OLBP is organized in the mixed bottom-up and top-down manner with multiple types of deep supervision. Extensive experiments on the constructed PFOS dataset show the superiority of the proposed OLBP network over 17 state-of-the-art methods, and demonstrate the effectiveness of the proposed OLM and BPM components. The constructed PFOS dataset and the proposed OLBP network are available at https://github.com/MathLee/OLBPNet4PFOS. Gongyang Li, Zhi Liu 0003, Weijie Wei 0001, Yong Wu 0007, Mengke Huang, Haibin Ling |
IEEE Trans. Image Process. | 2 |
| 2021 | SAL: Selection and Attention Losses for Weakly Supervised Semantic SegmentationabstractTraining a fully supervised semantic segmentation network requires a large amount of expensive pixel-level annotations in manual labor. In this work, we focus on studying the semantic segmentation problem using only image-level supervision. An effective scheme for weakly supervised segmentation is employed to produce the proxy annotations via image tags firstly. Then the segmentation network is retrained on the generated noisy proxy annotations. However, learning from noisy annotations is risky, as proxy annotations of poor quality may deteriorate the performance of the baseline segmentation and classification networks. In order to train the segmentation network using noisy annotations more effectively, two novel loss functions are proposed in this paper, namely, the selection loss and attention loss. Firstly, a selection loss is designed by weighting the proxy annotations based on a coarse-to-fine strategy for evaluating the quality of segmentation masks. Secondly, an attention loss taking the clean image tags as supervision is utilized to correct the classification errors caused by ambiguous pixel-level labels. Finally, we propose an end-to-end semantic segmentation network SAL-Net guided by the above two losses. From the extensive experiments conducted on PASCAL VOC 2012 dataset, SAL-Net reaches state-of-the-art performance with mean IoU (mIoU) as 62.5% and 66.6% on the test set by taking VGG16 network and ResNet101 network as the baselines respectively, which demonstrates the superiority of the proposed algorithm over eight representative weakly supervised segmentation methods. The code and models are available at https://github.com/zmbhou/SALTMM. Lei Zhou 0003, Chen Gong 0002, Zhi Liu 0003, Keren Fu |
IEEE Trans. Multim. | 3 |
| 2020 | Cross-Modal Weighting Network for RGB-D Salient Object Detection
Gongyang Li, Zhi Liu 0003, Linwei Ye, Yang Wang 0003, Haibin Ling |
ECCV (17) | 2 |
| 2020 | Co-Saliency Detection Using Collaborative Feature Extraction And High-To-Low Feature IntegrationabstractCo-saliency detection, as a developing research branch of saliency detection, devotes to identify the common salient objects in a group of related images. The major challenge of co-saliency detection is how to effectively represent features considering both intra-image and inter-image information. In this paper, we propose a co-saliency detection model using collaborative feature extraction and high-to-low feature integration. We first feed the target image and its co-images into the Individual Feature Extraction Module (IFEM) to produce multi-level individual features. Then, to capture the collaborative inter-image information, the Collaborative Feature Extraction Module (CFEM) is applied to all highest-level individual features, generating the collaborative feature. Finally, we build a High-to-low Feature Integration Module (HFIM), which integrates the collaborative feature and multi-level individual features of the target image, to enrich the collaborative feature with individual intra-image information. Extensive experiments on two public datasets demonstrate that the proposed model achieves the state-of-the-art performance. Jingru Ren, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Cong Bai, Guangling Sun |
ICME | 2 |
| 2020 | Fixations based personal target objects segmentationabstractWith the development of the eye-tracking technique, the fixation becomes an emergent interactive mode in many human-computer interaction study field. For a personal target objects segmentation task, although the fixation can be taken as a novel and more convenient interactive input, it induces a heavy ambiguity problem of the input's indication so that the segmentation quality is severely degraded. In this paper, to address this challenge, we develop an "extraction-to-fusion" strategy based iterative lightweight neural network, whose input is composed by an original image, a fixation map and a position map. Our neural network consists of two main parts: The first extraction part is a concise interlaced structure of standard convolution layers and progressively higher dilated convolution layers to better extract and integrate local and global features of target objects. The second fusion part is a convolutional long short-term memory component to refine the extracted features and store them. Depending on the iteration framework, current extracted features are refined by fusing them with stored features extracted in the previous iterations, which is a feature transmission mechanism in our neural network. Then, current improved segmentation result is generated to further adjust the fixation map and the position map in the next iteration. Thus, the ambiguity problem induced by the fixations can be alleviated. Experiments demonstrate better segmentation performance of our method and effectiveness of each part in our model. Gongyang Li, Weijie Wei 0001, Zhi Liu 0003 |
MMAsia | 4 |
| 2020 | Attentional coarse-and-fine generative adversarial networks for image inpainting
Minyu Chen 0001, Zhi Liu 0003, Linwei Ye, Yang Wang 0003 |
Neurocomputing | 2 |
| 2020 | Co-saliency detection via integration of multi-layer convolutional features and inter-image propagation
Jingru Ren, Zhi Liu 0003, Xiaofei Zhou 0003, Cong Bai, Guangling Sun |
Neurocomputing | 2 |
| 2020 | Attention-guided RGBD saliency detection using appearance information
Xiaofei Zhou 0003, Gongyang Li, Chen Gong 0002, Zhi Liu 0003, Jiyong Zhang 0001 |
Image Vis. Comput. | 4 |
| 2020 | Weakly supervised instance segmentation using multi-stage erasing refinement and saliency-guided proposals ordering
Zhi Liu 0003, Gongyang Li, Linwei Ye, Lei Zhou 0003, Yang Wang 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Saliency detection using adversarial learning networks
Yong Wu 0007, Zhi Liu 0003, Xiaofei Zhou 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Effective schizophrenia recognition using discriminative eye movement features and model-metric based features
Lijin Huang, Weijie Wei 0001, Zhi Liu 0003, Tianhong Zhang, Jijun Wang 0003, Lihua Xu, Olivier Le Meur |
Pattern Recognit. Lett. | 3 |
| 2020 | FANet: Features Adaptation Network for 360$^{\circ }$ Omnidirectional Salient Object DetectionabstractSalient object detection (SOD) in 360° omnidirectional images has become an eye-catching problem because of the popularity of affordable 360° cameras. In this paper, we propose a Features Adaptation Network (FANet) to highlight salient objects in 360° omnidirectional images reliably. To utilize the feature extraction capability of convolutional neural networks and capture global object information, we input the equirectangular 360° images and corresponding cube-map 360° images to the feature extraction network (FENet) simultaneously to obtain multi-level equirectangular and cube-map features. Furthermore, we fuse these two kinds of features at each level of the FENetby a projection features adaptation (PFA) module, for selecting these two kinds of features adaptively. Finally, we combine the preliminary adaptation features at different levels by a multi-level features adaptation (MLFA) module, which weights these different-level features adaptively and produces the final saliency maps. Experiments show our FANet outperforms the state-of-the-art methods on the 360° omnidirectional SOD datasets. Mengke Huang, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Olivier Le Meur |
IEEE Signal Process. Lett. | 2 |
| 2020 | ICNet: Information Conversion Network for RGB-D Based Salient Object DetectionabstractRGB-D based salient object detection (SOD) methods leverage the depth map as a valuable complementary information for better SOD performance. Previous methods mainly resort to exploit the correlation between RGB image and depth map in three fusion domains: input images, extracted features, and output results. However, these fusion strategies cannot fully capture the complex correlation between the RGB image and depth map. Besides, these methods do not fully explore the cross-modal complementarity and the cross-level continuity of information, and treat information from different sources without discrimination. In this paper, to address these problems, we propose a novel Information Conversion Network (ICNet) for RGB-D based SOD by employing the siamese structure with encoder-decoder architecture. To fuse high-level RGB and depth features in an interactive and adaptive way, we propose a novel Information Conversion Module (ICM), which contains concatenation operations and correlation layers. Furthermore, we design a Cross-modal Depth-weighted Combination (CDC) block to discriminate the cross-modal features from different sources and to enhance RGB features with depth features at each level. Extensive experiments on five commonly tested datasets demonstrate the superiority of our ICNet over 15 state-of-theart RGB-D based SOD methods, and validate the effectiveness of the proposed ICM and CDC block. Gongyang Li, Zhi Liu 0003, Haibin Ling |
IEEE Trans. Image Process. | 2 |
| 2020 | Dual Convolutional LSTM Network for Referring Image SegmentationabstractWe consider referring image segmentation. It is a problem at the intersection of computer vision and natural language understanding. Given an input image and a referring expression in the form of a natural language sentence, the goal is to segment the object of interest in the image referred by the linguistic query. To this end, we propose a dual convolutional LSTM (ConvLSTM) network to tackle this problem. Our model consists of an encoder network and a decoder network, where ConvLSTM is used in both encoder and decoder networks to capture spatial and sequential information. The encoder network extracts visual and linguistic features for each word in the expression sentence, and adopts an attention mechanism to focus on words that are more informative in the multimodal interaction. The decoder network integrates the features generated by the encoder network at multiple levels as its input and produces the final precise segmentation mask. Experimental results on four challenging datasets demonstrate that the proposed network achieves superior segmentation performance compared with other state-of-the-art methods. Linwei Ye, Zhi Liu 0003, Yang Wang 0003 |
IEEE Trans. Multim. | 2 |
| 2019 | Cross-Modal Self-Attention Network for Referring Image SegmentationabstractWe consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expression in the image. Existing works in this area treat the language expression and the input image separately in their representations. They do not sufficiently capture long-range correlations between these two modalities. In this paper, we propose a cross-modal self-attention (CMSA) module that effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the input image. In addition, we propose a gated multi-level fusion module to selectively integrate self-attentive cross-modal features corresponding to different levels in the image. This module controls the information flow of features at different levels. We validate the proposed approach on four evaluation datasets. Our proposed approach consistently outperforms existing state-of-the-art methods. Linwei Ye, Mrigank Rochan, Zhi Liu 0003, Yang Wang 0003 |
CVPR | 3 |
| 2019 | Cleaning Adversarial Perturbations via Residual Generative Network for Face VerificationabstractDeep neural networks (DNNs) have recently achieved impressive performances on various applications. However, recent researches show that DNNs are vulnerable to adversarial perturbations injected into input samples. In this paper, we investigate a defense method for face verification: a deep residual generative network (ResGN) is learned to clean adversarial perturbations. We propose a novel training framework composed of ResGN, pre-trained VGG-Face network and FaceNet network. The parameters of ResGN are optimized by minimizing a joint loss consisting of a pixel loss, a texture loss and a verification loss, in which they measure content errors, subjective visual perception errors and verification task errors between cleaned image and legitimate image respectively. Specially, the latter two are provided by VGG-Face and FaceNet respectively and have essential contributions for improving verification performance of cleaned image. Empirical experiment results validate the effectiveness of the proposed defense method on the Labeled Faces in the Wild (LFW) benchmark dataset. Yuying Su, Guangling Sun, Weiqi Fan, Zhi Liu 0003 |
ICASSP | 5 |
| 2019 | Learn Image Object Co-segmentation with Multi-scale Feature FusionabstractImage object co-segmentation aims to segment common objects in a group of images. This paper proposes a novel neural network, which extracts multi-scale convolutional features at multiple layers via a modified VGG network and fuses them both within and across images as the intra-image and the inter-image features. Then these two kinds of features are further fused at each scale as the multi-scale co-features of common objects, and finally the multi-scale co-features are summed up and upsampled to obtain the co-segmentation results. To simplify the network and reduce the rapidly rising resource cost along with the inputs, the reduced input size, less downsampling and dilation convolution are adopted in the proposed model. Experimental results on the public dataset demonstrate that the proposed model achieves a comparable performance to the state-of-the-art co-segmentation methods while the computation cost has been effectively reduced. Zhi Liu 0003, Jian Zhang 0002, Xiaofei Zhou 0003 |
VCIP | 2 |
| 2019 | Saliency detection via multi-level integration and multi-scale fusion neural networks
Mengke Huang, Zhi Liu 0003, Linwei Ye, Xiaofei Zhou 0003, Yang Wang 0003 |
Neurocomputing | 2 |
| 2019 | Constrained fixation point based segmentation via deep neural network
Gongyang Li, Zhi Liu 0003, Weijie Wei 0001 |
Neurocomputing | 2 |
| 2019 | Superpixel based continuous conditional random field neural network for semantic segmentation
Lei Zhou 0003, Keren Fu, Zhi Liu 0003, Zhimin Yin, Jianli Zheng |
Neurocomputing | 3 |
| 2019 | Depth-aware saliency detection using convolutional neural networks
Zhi Liu 0003, Mengke Huang, Xiangyang Wang 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Weakly labeled fine-grained classification with hierarchy relationship of fine and coarse labels
Qihan Jiao, Zhi Liu 0003, Linwei Ye, Yang Wang 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Salient object segmentation based on depth-aware image layering
Huan Du, Zhi Liu 0003 |
Multim. Tools Appl. | 2 |
| 2019 | Integration of statistical detector and Gaussian noise injection detector for adversarial example detection in deep neural networks
Weiqi Fan, Guangling Sun, Yuying Su, Zhi Liu 0003 |
Multim. Tools Appl. | 4 |
| 2019 | Effective online refinement for video object segmentation
Gongyang Li, Zhi Liu 0003, Xiaofei Zhou 0003 |
Multim. Tools Appl. | 2 |
| 2019 | Saliency-aware inter-image color transfer for image manipulation
Zhi Liu 0003, Qihan Jiao, Olivier Le Meur, Wanlei Zhao |
Multim. Tools Appl. | 2 |
| 2019 | Video co-segmentation based on directed graph
Zhi Liu 0003, Xiaofei Zhou 0003, Wei Liu 0044, Xuemei Zou |
Multim. Tools Appl. | 2 |
| 2019 | Supervised learning based discrete hashing for image retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Shengyong Chen |
Pattern Recognit. | 4 |
| 2018 | Object Trajectory Proposal via Hierarchical Volume GroupingabstractObject trajectory proposal aims to locate category-independent object candidates in videos with a limited number of trajectories,i.e.,bounding box sequences. Most existing methods, which derive from combining object proposal with tracking, cannot handle object trajectory proposal effectively due to the lack of comprehensive objectness measurement through analyzing spatio-temporal characteristics over a whole video. In this paper, we propose a novel object trajectory proposal method using hierarchical volume grouping. Specifically, we first represent a given video with hierarchical volumes by mapping hierarchical regions with optical flow. Then, we filter the short volumes and background volumes, and combinatorially group the retained volumes into object candidates. Finally, we rank the object candidates using a multi-modal fusion scoring mechanism, which incorporates both appearance objectness and motion objectness, and generate the bounding boxes of the object candidates with the highest scores as the trajectory proposals. We validated the proposed method on a dataset consisting of 200 videos from ILSVRC2016-VID. The experimental results show that our method is superior to the state-of-the-art object trajectory proposal methods. Xu Sun 0009, Yuantian Wang, Tongwei Ren, Zhi Liu 0003, Zhengjun Zha, Gangshan Wu |
ICMR | 4 |
| 2018 | Video Saliency Detection Using Deep Convolutional Neural Networks
Xiaofei Zhou 0003, Zhi Liu 0003, Chen Gong 0002, Gongyang Li, Mengke Huang |
PRCV (2) | 2 |
| 2018 | Learning Semantic Segmentation with Diverse SupervisionabstractModels based on deep convolutional neural networks (CNN) have significantly improved the performance of semantic segmentation. However, learning these models requires a large amount of training images with pixel-level labels, which are very costly and time-consuming to collect. In this paper, we propose a method for learning CNNbased semantic segmentation models from images with several types of annotations that are available for various computer vision tasks, including image-level labels for classification, box-level labels for object detection and pixel-level labels for semantic segmentation. The proposed method is flexible and can be used together with any existing CNNbased semantic segmentation networks. Experimental evaluation on the challenging PASCAL VOC 2012 and SIFTflow benchmarks demonstrate that the proposed method can effectively make use of diverse training data to improve the performance of the learned models. Linwei Ye, Zhi Liu 0003, Yang Wang 0003 |
WACV | 2 |
| 2018 | Unsupervised image co-segmentation via guidance of simple images
Zhi Liu 0003, Jian Zhang 0002 |
Neurocomputing | 2 |
| 2018 | Saliency integration driven by similar images
Jingru Ren, Zhi Liu 0003, Xiaofei Zhou 0003, Guangling Sun, Cong Bai |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Video saliency detection via bagging-based prediction and spatiotemporal propagation
Xiaofei Zhou 0003, Zhi Liu 0003, Kai Li 0016, Guangling Sun |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | RGBD co-saliency detection via multiple kernel boosting and fusion
Lishan Wu, Zhi Liu 0003, Hangke Song, Olivier Le Meur |
Multim. Tools Appl. | 2 |
| 2018 | Spatiotemporal salient object detection by integrating with objectness
Tongbao Wu, Zhi Liu 0003, Xiaofei Zhou 0003, Kai Li 0016 |
Multim. Tools Appl. | 2 |
| 2018 | Efficient Intra Mode Selection for Depth-Map Coding Utilizing Spatiotemporal, Inter-Component and Inter-View Correlations in 3D-HEVCabstract3D-high efficiency video coding (HEVC) is developed for the compression of the multi-view video plus depth format, which is based on the latest generation of video coding standard, HEVC. It further adopts several new intra prediction modes, depth-modeling modes (DMMs) in intra candidate modes for a better representation of edges in depth maps, which introduces a drastic increase in the computational complexity. The procedure of depth intra mode decision together with DMMs and existing intra modes is a very time consuming part due to huge complexity of full rate distortion (RD) cost calculation. In this paper, a low complexity intra mode selection algorithm is proposed to reduce complexity of depth intra prediction in both intra-frames and inter-frames. An experimental analysis is first performed to study the inter-view correlation and the inter-component (texture video and its associated depth) correlation in intra coding information such as the intra mode and RD cost. All intra modes available in 3D-HEVC are classified into three activity classes assigned with different mode-weight factors, and the coding mode complexity of a coding unit (CU) is defined according to the intra mode information from available spatiotemporal, inter-view, and inter-component neighboring coded CUs. The coding mode complexity analysis is utilized to assign different candidate intra modes for different types of CUs. The optimal intra prediction mode and the RD cost value in current CU depth level are further used to skip unnecessary intra prediction sizes. Experimental results show that the proposed fast depth intra coding algorithm achieves 61% complexity reduction on intra prediction, while incurring a 0.2% Bjontegaard metric increase for coded and synthesized views compared to the test model of 3D-HEVC. Liquan Shen, Kai Li 0016, Guorui Feng, Ping An 0001, Zhi Liu 0003 |
IEEE Trans. Image Process. | 5 |
| 2018 | Improving Video Saliency Detection via Localized Estimation and Spatiotemporal RefinementabstractVideo saliency detection aims to pop out the most salient regions in every frame of a video. Up to now, many efforts have been made from various aspects for video saliency detection. Unfortunately, the existing video saliency models are very likely to fail in challenging videos with complicated motions and complex scenes. Therefore, in this paper, we propose a novel framework to improve the saliency detection results generated by existing video saliency models. The proposed framework consists of three key steps including localized estimation, spatiotemporal refinement, and saliency update. Specifically, the initial saliency map of each frame in a video is first generated by using an existing saliency model. Then, by considering the temporal consistency and strong correlation among adjacent frames, the localized estimation models, which are generated by training the random forest regressor within a local temporal window, are employed to generate the temporary saliency map. Finally, by taking the appearance and motion information of salient objects into consideration, the spatiotemporal refinement step is deployed to further improve the temporary saliency map and generate the final saliency map. Furthermore, such an improved saliency map is then utilized to update the initial saliency map and provide reliable cues for saliency detection in the next frame. The experimental results on four challenging datasets demonstrate that the proposed framework is able to consistently and significantly improve the saliency detection performance of various video saliency models, thereby achieving the state-of-the-art performance. Xiaofei Zhou 0003, Zhi Liu 0003, Chen Gong 0002, Wei Liu 0044 |
IEEE Trans. Multim. | 2 |
| 2017 | Semi-Global Weighted Least Squares in Image FilteringabstractSolving the global method of Weighted Least Squares (WLS) model in image filtering is both time- and memory-consuming. In this paper, we present an alternative approximation in a time- and memory- efficient manner which is denoted as Semi-Global Weighed Least Squares (SG-WLS). Instead of solving a large linear system, we propose to iteratively solve a sequence of subsystems which are one-dimensional WLS models. Although each subsystem is one-dimensional, it can take two-dimensional neighborhood information into account due to the proposed special neighborhood construction. We show such a desirable property makes our SG-WLS achieve close performance to the original two-dimensional WLS model but with much less time and memory cost. While previous related methods mainly focus on the 4-connected/8-connected neighborhood system, our SG-WLS can handle a more general and larger neighborhood system thanks to the proposed fast solution. We show such a generalization can achieve better performance than the 4-connected/8-connected neighborhood system in some applications. Our SG-WLS is ~ 20 times faster than the WLS model. For an image of M × N, the memory cost of SG-WLS is at most at the magnitude of max{1/M,1/N} of that of the WLS model. We show the effectiveness and efficiency of our SG-WLS in a range of applications. Wei Liu 0044, Chunhua Shen, Zhi Liu 0003, Jie Yang 0002 |
ICCV | 4 |
| 2017 | Two-Stage Saliency Fusion for Object Segmentation
Guangling Sun, Jingru Ren, Zhi Liu 0003 |
ICIG (1) | 3 |
| 2017 | Age-dependent saccadic models for predicting eye movementsabstractHow people look at visual information reveals fundamental information about themselves, their interests and their state of mind. While previous visual attention models output static 2-dimensional saliency maps, saccadic models predict not only what observers look at but also how they move their eyes to explore the scene. Here we demonstrate that saccadic models are a flexible framework that can be tailored to emulate the gaze patterns from childhood to adulthood. The proposed age-dependent saccadic model not only outputs human-like, i.e. age-specific visual scanpath, but also significantly outperforms other state-of-the-art saliency models. Olivier Le Meur, Antoine Coutrot, Adrien Le Roch, Andrea Helo, Pia Rämä, Zhi Liu 0003 |
ICIP | 6 |
| 2017 | Depth-aware object instance segmentationabstractWe consider the problem of object instance segmentation. The goal is to label each pixel in an image according to its object class as well as its object instance. The proposed approach consists of three steps including object instance detection, category-specific instance segmentation and depth-aware ordering. The novelty of the proposed approach is that it uses the depth information to resolve the ambiguity of pixel labels when two object instances are overlapping. Experimental results on the PASCAL VOC 2012 benchmark demonstrate the competitive performance of the proposed approach compared with other state-of-the-art methods. Linwei Ye, Zhi Liu 0003, Yang Wang 0003 |
ICIP | 2 |
| 2017 | Saliency-based navigation in omnidirectional imageabstractOmnidirectional images describe the color information at a given position from all directions. Affordable 360° cameras have recently been developed leading to an explosion of the 360° data shared on social networks. However, an omnidirectional image does not contain interesting content everywhere. Some part of the images are indeed more likely to be looked at by some users than others. Knowing these regions of interest might be useful for 360° image compression, streaming, retargeting or even editing. In this paper, we aim at modelling the user navigation within a 360° image, and detecting which parts of an omnidirectional content might draw users' attention. In particular, the paper proposes to aggregate and analyze 2D saliency detectors in different map projections, and also proposes a smooth navigation through the image to maximize saliency. Thomas Maugey, Olivier Le Meur, Zhi Liu 0003 |
MMSP | 3 |
| 2017 | A spiking neural network model for obstacle avoidance in simulated prosthetic vision
Chenjie Ge, Nikola K. Kasabov, Zhi Liu 0003, Jie Yang 0002 |
Inf. Sci. | 3 |
| 2017 | Adaptive saliency fusion based on quality assessment
Xiaofei Zhou 0003, Zhi Liu 0003, Guangling Sun, Xiangyang Wang 0003 |
Multim. Tools Appl. | 2 |
| 2017 | Saliency Detection for Unconstrained Videos Using Superpixel-Level Graph and Spatiotemporal PropagationabstractThis paper proposes an effective spatiotemporal saliency model for unconstrained videos with complicated motion and complex scenes. First, superpixel-level motion and color histograms as well as global motion histogram are extracted as the features for saliency measurement. Then a superpixel-level graph with the addition of a virtual background node representing the global motion is constructed, and an iterative motion saliency (MS) measurement method that utilizes the shortest path algorithm on the graph is exploited to reasonably generate MS maps. Temporal propagation of saliency in both forward and backward directions is performed using efficient operations on inter-frame similarity matrices to obtain the integrated temporal saliency maps with the better coherence. Finally, spatial propagation of saliency both locally and globally is performed via the use of intra-frame similarity matrices to obtain the spatiotemporal saliency maps with the even better quality. The experimental results on two video data sets with various unconstrained videos demonstrate that the proposed model consistently outperforms the state-of-the-art spatiotemporal saliency models on saliency detection performance. Zhi Liu 0003, Linwei Ye, Guangling Sun, Liquan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Visual Attention Saccadic Models Learn to Emulate Gaze Patterns From Childhood to AdulthoodabstractHow people look at visual information reveals fundamental information about themselves, their interests and their state of mind. While previous visual attention models output static 2D saliency maps, saccadic models aim to predict not only where observers look at but also how they move their eyes to explore the scene. In this paper, we demonstrate that saccadic models are a flexible framework that can be tailored to emulate observer's viewing tendencies. More specifically, we use fixation data from 101 observers split into five age groups (adults, 8-10 y.o., 6-8 y.o., 4-6 y.o., and 2 y.o.) to train our saccadic model for different stages of the development of human visual system. We show that the joint distribution of saccade amplitude and orientation is a visual signature specific to each age group, and can be used to generate age-dependent scan paths. Our age-dependent saccadic model does not only output human-like, age-specific visual scan paths, but also significantly outperforms other state-of-the-art saliency models. We demonstrate that the computational modeling of visual attention, through the use of saccadic model, can be efficiently adapted to emulate the gaze behavior of a specific group of observers. Olivier Le Meur, Antoine Coutrot, Zhi Liu 0003, Pia Rämä, Adrien Le Roch, Andrea Helo |
IEEE Trans. Image Process. | 3 |
| 2017 | Depth-Aware Salient Object Detection and Segmentation via Multiscale Discriminative Saliency Fusion and Bootstrap LearningabstractThis paper proposes a novel depth-aware salient object detection and segmentation framework via multiscale discriminative saliency fusion (MDSF) and bootstrap learning for RGBD images (RGB color images with corresponding Depth maps) and stereoscopic images. By exploiting low-level feature contrasts, mid-level feature weighted factors and high-level location priors, various saliency measures on four classes of features are calculated based on multiscale region segmentation. A random forest regressor is learned to perform the discriminative saliency fusion (DSF) and generate the DSF saliency map at each scale, and DSF saliency maps across multiple scales are combined to produce the MDSF saliency map. Furthermore, we propose an effective bootstrap learning-based salient object segmentation method, which is bootstrapped with samples based on the MDSF saliency map and learns multiple kernel support vector machines. Experimental results on two large datasets show how various categories of features contribute to the saliency detection performance and demonstrate that the proposed framework achieves the better performance on both saliency detection and salient object segmentation. Hangke Song, Zhi Liu 0003, Huan Du, Guangling Sun, Olivier Le Meur, Tongwei Ren |
IEEE Trans. Image Process. | 2 |
| 2017 | Salient Object Segmentation via Effective Integration of Saliency and ObjectnessabstractThis paper proposes an effective salient object segmentation method via the graph-based integration of saliency and objectness. Based on the superpixel segmentation result of the input image, a graph is built to represent superpixels using regular vertex, background seed vertex with the addition of a terminal vertex. The edge weights on the graph are defined by integrating the difference of appearance, saliency, and objectness between superpixels. Then, the object probability of each superpixel is measured by finding the shortest path from the corresponding vertex to the terminal vertex on the graph, and the resultant object probability map can generally better highlight salient objects and suppress background regions compared to both saliency map and objectness map. Finally, the object probability map is used to initialize salient object and background, and effectively incorporated into the framework of graph cut to obtain the final salient object segmentation result. Extensive experimental results on three public benchmark datasets show that the proposed method consistently improves the salient object segmentation performance and outperforms the state-of-the-art salient object segmentation methods. Furthermore, experimental results also demonstrate that the proposed graph-based integration method is more effective than other fusion schemes and robust to saliency maps generated using various saliency models. Linwei Ye, Zhi Liu 0003, Liquan Shen, Cong Bai, Yang Wang 0003 |
IEEE Trans. Multim. | 2 |
| 2017 | An Imbalance Compensation Framework for Background SubtractionabstractClass imbalance refers to the instance where the number of training samples for the majority classes is far more than that of the minority classes (relative imbalance), and the quality of training samples for the minority classes is inferior to that of the majority classes (absolute imbalance), which are further complicated by other imbalance factors, e.g., data overlapping. Video background subtraction aims to classify each pixel into two classes: foreground and background. This paper first reveals that background subtraction is a class imbalance problem, where the foreground and background are the minority and majority classes, respectively. By exploring spatial and temporal correlation inherent in video data, we present an imbalance compensation framework for background subtraction, which consists of two sequential modules, imbalance-compensated bilayer modeling, and imbalance-compensated Bayesian classification. In the first module, spatio-temporal oversampling (SOS) and selective downsampling (SDS) are proposed to compensate the imbalance at data level. SOS attempts to synthesize representative samples appended to the minority sample set, while SDS selectively deletes a number of majority samples in data overlapping areas. The rebalanced samples are then used to learn a bilayer model. In the second module, novel cost functions are proposed to compensate the effect of class imbalance at algorithm level. The cost functions are based on imbalance measurement, and used to construct the prior term in the Bayesian classification scheme. Experiments are conducted on public databases to demonstrate the effectiveness of the proposed method. Xiang Zhang 0006, Ce Zhu, Honggang Wu, Zhi Liu 0003, Yuanyuan Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Depth-aware saliency detection using discriminative saliency fusionabstractIn this paper, we propose a multi-stage depth-aware saliency model for salient region detection. We evaluate saliency on different features at low, mid and high levels, by taking account of primary depth and appearance contrasts, different feature weighted factors and location priors, respectively. Unlike most existing depth-aware saliency models that use a linear or experiential fusion formula to combine saliency maps from different features, we calculate saliency of each feature individually at each level and learn a discriminative saliency fusion (DSF) regressor based on random forest to estimate the saliency measures of regions. Both subjective and objective evaluations on two public datasets designed for depth-aware saliency detection demonstrate that the proposed saliency model consistently outperforms the state-of-the-art saliency models on saliency detection performance. Hangke Song, Zhi Liu 0003, Huan Du, Guangling Sun |
ICASSP | 2 |
| 2016 | Using independent component analysis and binocular combination for stereoscopic image quality assessmentabstractIn this paper, a full reference stereoscopic image quality assessment (FR-SIQA) method is proposed based on independent component analysis (ICA) and binocular combination. Image features that reflect the responds of simple cells in the cortex are extracted by ICA-based algorithm. Both image feature similarity (IFS) and local luminance consistency (LLC) are calculated to measure the structure and brightness distortions, respectively. To simulate the binocular fusion properties, the energy of image features and the global relative luminance information are selected as the basic of binocular combination to fuse the right-left IFS and LLC into a final index. Experimental results demonstrate that the proposed algorithm achieves high consistency with subjective assessment on two public available 3D image quality assessment databases. Xianqiu Geng, Liquan Shen, Ping An 0001, Zhi Liu 0003 |
VCIP | 4 |
| 2016 | Facial descriptor for Kinect depth using inner-inter-normal components local binary patterns and tensor histograms
Guangling Sun, Yong Dong, Xiaofei Zhou 0003, Zhi Liu 0003 |
Mach. Vis. Appl. | 4 |
| 2016 | RGBD Co-saliency Detection via Bagging-Based ClusteringabstractWith the additional depth information, RGBD co-saliency detection, which is an emerging and interesting issue in saliency detection, aims to discover the common salient objects in a set of RGBD images. This letter proposes a novel RGBD co-saliency model using bagging-based clustering. First, candidate object regions are generated based on RGBD single saliency maps and region pre-segmentation. Then, in order to make regional clustering more robust to different image sets, the feature bagging method is introduced to randomly generate multiple clustering results and the cluster-level weak co-saliency maps. Finally, a clustering quality (CQ) criterion is devised to adaptively integrate the weak co-saliency maps into the final co-saliency map for each image. Experimental results on a public RGBD co-saliency dataset show that the proposed co-saliency model significantly outperforms the state-of-the-art co-saliency models. Hangke Song, Zhi Liu 0003, Lishan Wu, Mengke Huang |
IEEE Signal Process. Lett. | 2 |
| 2016 | Saliency Detection Via Similar Image RetrievalabstractThis letter proposes a novel saliency detection framework by propagating saliency of similar images retrieved from large and diverse Internet image collections to boost saliency detection performance effectively. For the input image, a group of similar images is retrieved based on the saliency weighted color histograms and the Gist descriptor from Internet image collections. Then, a pixel-level correspondence process between images is performed to guide the saliency propagation from the retrieved images. Both initial saliency map and correspondence saliency map are exploited to select the training samples by using the graph cut-based segmentation. Finally, the training samples are input into a set of weak classifiers to learn the boosted classifier for generating the boosted saliency map, which is integrated with the initial saliency map to generate the final saliency map. Experimental results on two public image datasets demonstrate that the proposed model can achieve the better saliency detection performance than the state-of-the-art single-image saliency models and co-saliency models. Linwei Ye, Zhi Liu 0003, Xiaofei Zhou 0003, Liquan Shen, Jian Zhang 0002 |
IEEE Signal Process. Lett. | 2 |
| 2016 | Improving Saliency Detection Via Multiple Kernel Boosting and Adaptive FusionabstractThis letter proposes a novel framework to improve the saliency detection performance of an existing saliency model, which is used to generate the initial saliency map. First, a novel regional descriptor consisting of regional self-information, regional variance, and regional contrast on a number of features with local, global, and border context is proposed to describe the segmented regions at multiple scales. Then, regarding saliency computation as a regression problem, a multiple kernel boosting method based on support vector regression (MKB-SVR) is proposed to generate the complementary saliency map. Finally, an adaptive fusion method via learning a quality prediction model for saliency maps is proposed to effectively fuse the initial saliency map with the complementary saliency map and obtain the final saliency map with improvement on saliency detection performance. Experimental results on two public datasets with the state-of-the-art saliency models validate that the proposed method consistently improves the saliency detection performance of various saliency models. Xiaofei Zhou 0003, Zhi Liu 0003, Guangling Sun, Linwei Ye, Xiangyang Wang 0003 |
IEEE Signal Process. Lett. | 2 |
| 2015 | K-means based histogram using multiresolution feature vectors for color texture database retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Wanlei Zhao |
Multim. Tools Appl. | 3 |
| 2015 | Spatiotemporal saliency detection based on superpixel-level trajectory
Zhi Liu 0003, Xiang Zhang 0006, Olivier Le Meur, Liquan Shen |
Signal Process. Image Commun. | 2 |
| 2015 | Special issue on recent advances in saliency models, applications and evaluations
Zhi Liu 0003, Olivier Le Meur, Ali Borji, Hongliang Li 0001 |
Signal Process. Image Commun. | 1 |
| 2015 | Fast TU size decision algorithm for HEVC encoders using Bayesian theorem detection
Liquan Shen, Zhaoyang Zhang 0002, Xinpeng Zhang 0001, Ping An 0001, Zhi Liu 0003 |
Signal Process. Image Commun. | 5 |
| 2015 | Efficient Saliency-Model-Guided Visual Co-Saliency DetectionabstractThis letter proposes a novel framework to detect common salient objects in a group of images automatically and efficiently. Different from most existing co-saliency models which directly redesign algorithms for multiple images, the saliency model for a single image is fully exploited under the proposed framework to guide the co-saliency detection. Given single image saliency maps, a two-stage guided detection pipeline led by queries is proposed to obtain the guided saliency maps of the image set through a ranking scheme. Then the guided saliency maps generated by different queries are fused in a way that takes advantages of both averaging and multiplication. The proposed model makes existing saliency models work well in co-saliency scenarios. Experimental results on two benchmark databases demonstrate that the proposed framework outperforms the state-of-the-art models in terms of both accuracy and efficiency. Yijun Li 0003, Keren Fu, Zhi Liu 0003, Jie Yang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2015 | Co-Saliency Detection via Co-Salient Object Discovery and RecoveryabstractThis letter proposes a novel co-saliency model to effectively discover and highlight co-salient objects in a set of images. Based on the gross similarity which combines color features and SIFT descriptors, some co-salient object regions are first discovered in each image as exemplars, which are exploited to generate the exemplar saliency maps with the use of single-image saliency model. Then both local recovery and global recovery of co-salient object regions are performed by propagating the exemplar saliency to the matched regions, and border connectivity is further exploited to generate the region-level co-saliency maps. Finally, the foci of attention area based pixel-level saliency derivation is used to generate the pixel-level co-saliency maps with even better quality. Experimental results on two benchmark datasets demonstrate that the proposed co-saliency model outperforms the state-of-the-art co-saliency models. Linwei Ye, Zhi Liu 0003, Wanlei Zhao, Liquan Shen |
IEEE Signal Process. Lett. | 2 |
| 2015 | Unsupervised Joint Salient Region Detection and Object SegmentationabstractThis paper presents a novel unsupervised algorithm to detect salient regions and to segment out foreground objects from background. In contrast to previous unidirectional saliency-based object segmentation methods, in which only the detected saliency map is used to guide the object segmentation, our algorithm mutually exploits detection/segmentation cues from each other. To achieve this goal, an initial saliency map is generated by the proposed segmentation driven low-rank matrix recovery model. Such a saliency map is exploited to initialize object segmentation model, which is formulated as energy minimization of Markov random field. Mutually, the quality of saliency map is further improved by the segmentation result, and serves as a new guidance for the object segmentation. The optimal saliency map and the final segmentation are achieved by jointly optimizing the defined objective functions. Extensive evaluations on MSRA-B and PASCAL-1500 datasets demonstrate that the proposed algorithm achieves the state-of-the-art performance for both the salient region detection and the object segmentation. Wenbin Zou, Zhi Liu 0003, Kidiyo Kpalma, Joseph Ronsin, Yong Zhao 0010, Nikos Komodakis |
IEEE Trans. Image Process. | 2 |
| 2014 | Saliency Aggregation: Does Unity Make Strength?
Olivier Le Meur, Zhi Liu 0003 |
ACCV (4) | 2 |
| 2014 | Robust blurred face recognition using sample-wise kernel estimation and random compressed multi-scale local binary pattern histogramsabstractRobust face recognition from blurred images is an important and challenging problem. We propose sample-wise kernel estimation (SWKE). For each sharp gallery sample, an optimal blur kernel making the distance between probe and corresponding blurred sample minimum is estimated. Then, the kernel is applied to blur the sharp gallery sample. All such blurred samples compose the adaptive blurred sample set acting as references during recognition. Next, we present random compressed multi-scale local binary pattern histograms (RCMSLBPH). Due to the sparseness of the histogram of each sub-region and each feature image, it is possible to accomplish large-margin dimensionality reduced without impairing great information loss by using random Gaussian projection matrix. Finally, we design two stage framework involving coarse recognition and fine recognition to obtain an optimized effectiveness-efficiency trade-off. The experimental results confirm the improvement of the proposed SWKE, RCMSLBPH and two stage framework on public face database recognition. Guangling Sun, Zhi Liu 0003 |
ICIP | 2 |
| 2014 | Co-saliency detection based on region-level fusion and pixel-level refinementabstractThis paper addresses the problem of co-saliency detection, which aims to identify the common salient objects in a set of images and is important for many applications such as object co-segmentation and co-recognition. First, the segmentation driven low-rank matrix recovery model is used for intra saliency detection in each individual image of the image set, to highlight the regions whose features are sparse in each image. Then, a region-level fusion method, which exploits inter-region dissimilarities on color histograms and global consistency of regions over the image set, adjusts the intra saliency maps to obtain the region-level co-saliency maps, which can highlight co-salient object regions and suppress irrelevant regions. Finally, a pixel-level refinement method, which integrates color-spatial similarity between pixel and region with image border connectivity based object prior, generates the pixel-level co-saliency maps with better quality. Extensive experiments on two benchmark datasets demonstrate that the proposed co-saliency model consistently outperforms the state-of-the-art co-saliency models in both subjective and objective evaluation. Zhi Liu 0003, Wenbin Zou, Xiang Zhang 0006, Olivier Le Meur |
ICME | 2 |
| 2014 | Superpixel-Based Spatiotemporal Saliency DetectionabstractThis paper proposes a superpixel-based spatiotemporal saliency model for saliency detection in videos. Based on the superpixel representation of video frames, motion histograms and color histograms are extracted at the superpixel level as local features and frame level as global features. Then, superpixel-level temporal saliency is measured by integrating motion distinctiveness of superpixels with a scheme of temporal saliency prediction and adjustment, and superpixel-level spatial saliency is measured by evaluating global contrast and spatial sparsity of superpixels. Finally, a pixel-level saliency derivation method is used to generate pixel-level temporal and spatial saliency maps, and an adaptive fusion method is exploited to integrate them into the spatiotemporal saliency map. Experimental results on two public datasets demonstrate that the proposed model outperforms six state-of-the-art spatiotemporal saliency models in terms of both saliency detection and human fixation prediction. Zhi Liu 0003, Xiang Zhang 0006, Shuhua Luo, Olivier Le Meur |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Adaptive Inter-Mode Decision for HEVC Jointly Utilizing Inter-Level and Spatiotemporal CorrelationsabstractHigh Efficiency Video Coding (HEVC) adopts the quadtree structured coding unit (CU), which allows recursive splitting into four equally sized blocks. At each depth level, it enables SKIP mode, merge mode, inter 2N × 2N, inter 2N × N, inter N × 2N, inter 2N × nU, inter 2N × nD, inter nL x 2N, inter nR × 2N, inter N × N (only available for the smallest CU), intra 2N × 2N, and intra N × N (only available for the smallest CU) in inter-frames. Similar to H.264/AVC, the mode decision process in HEVC is performed using all the possible depth levels (or CU sizes) and prediction modes to find the one with the least rate distortion (RD) cost using Lagrange multiplier. This achieves the highest coding efficiency, but leads to a very high computational complexity. Since the optimal prediction mode is highly content dependent, it is not efficient to use all the modes. In this paper, we propose a fast inter-mode decision algorithm for HEVC by jointly using the inter-level correlation of quadtree structure and the spatiotemporal correlation. There exist strong correlations of the prediction mode, the motion vector and RD cost between different depth levels and between spatially temporally adjacent CUs. We statistically analyze the prediction mode distribution at each depth level and the coding information correlation among the adjacent CUs. Based on the analysis results, three adaptive inter-mode decision strategies are proposed including early SKIP mode decision, prediction size correlation-based mode decision and RD cost correlation-based mode decision. Experimental results show that the proposed overall algorithm can save 49%-52% computational complexity on average with negligible loss of coding efficiency, exhibiting applicability to various types of video sequences. Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Saliency Tree: A Novel Saliency Detection FrameworkabstractThis paper proposes a novel saliency detection framework termed as saliency tree. For effective saliency measurement, the original image is first simplified using adaptive color quantization and region segmentation to partition the image into a set of primitive regions. Then, three measures, i.e., global contrast, spatial sparsity, and object prior are integrated with regional similarities to generate the initial regional saliency for each primitive region. Next, a saliency-directed region merging approach with dynamic scale control scheme is proposed to generate the saliency tree, in which each leaf node represents a primitive region and each non-leaf node represents a non-primitive region generated during the region merging process. Finally, by exploiting a regional center-surround scheme based node selection criterion, a systematic saliency tree analysis including salient node selection, regional saliency adjustment and selection is performed to obtain final regional saliency measures and to derive the high-quality pixel-wise saliency map. Extensive experimental results on five datasets with pixel-wise ground truths demonstrate that the proposed saliency tree model consistently outperforms the state-of-the-art saliency models. Zhi Liu 0003, Wenbin Zou, Olivier Le Meur |
IEEE Trans. Image Process. | 1 |
| 2014 | Effective CU Size Decision for HEVC IntracodingabstractIn high efficiency video coding (HEVC), the tree structured coding unit (CU) is adopted to allow recursive splitting into four equally sized blocks. At each depth level (or CU size), it enables up to 35 intraprediction modes, including a planar mode, a dc mode, and 33 directional modes. The intraprediction via exhaustive mode search exploited in the test model of HEVC (HM) effectively improves coding efficiency, but results in a very high computational complexity. In this paper, a fast CU size decision algorithm for HEVC intracoding is proposed to speed up the process by reducing the number of candidate CU sizes required to be checked for each treeblock. The novelty of the proposed algorithm lies in the following two aspects: 1) an early determination of CU size decision with adaptive thresholds is developed based on the texture homogeneity and 2) a novel bypass strategy for intraprediction on large CU size is proposed based on the combination of texture property and coding information from neighboring coded CUs. Experimental results show that the proposed effective CU size decision algorithm achieves a computational complexity reduction up to 67%, while incurring only 0.06-dB loss on peak signal-to-noise ratio or 1.08% increase on bit rate compared with that of the original coding in HM. Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003 |
IEEE Trans. Image Process. | 3 |
| 2013 | Segmentation Driven Low-rank Matrix Recovery for Saliency DetectionabstractLow-rank matrix recovery (LRMR) model, aiming at decomposing a matrix into a low-rank matrix and a sparse one, has shown the potential to address the problem of saliency detection, where the decomposed low-rank matrix naturally corresponds to the background, and the sparse one captures salient objects. This is under the assumption that the background is consistent and objects are obviously distinctive. Unfortunately, in real images, the background may be cluttered and may have low contrast with objects. Thus directly applying the LRMR model to the saliency detection has limited robustness. This paper proposes a novel approach that exploits bottom-up segmentation as a guidance cue of the matrix recovery. This method is fully unsupervised, yet obtains higher performance than the supervised LRMR model. A new challenging dataset PASCAL-1500 is also introduced to validate the saliency detection performance. Extensive evaluations on the widely used MSRA-1000 dataset and also on the new PASCAL-1500 dataset demonstrate that the proposed saliency model outperforms the state-of-the-art models. Wenbin Zou, Kidiyo Kpalma, Zhi Liu 0003, Joseph Ronsin |
BMVC | 3 |
| 2013 | Cost-sensitive background subtractionabstractForeground and background are treated without distinction at classification stage in most background subtraction algorithms. However, correct classification of foreground is the primary requirement, and thus misclassification costs of the two classes should be different. Based on this fact, we present a new method to introduce cost sensitivity into background subtraction, where a cost matrix is created to represent the costs of misclassification. Some items in the cost matrix are not constants, but functions of foreground occurence at each pixel location. By the use of such non-constant costs, detection rate of foreground is improved while increase of false alarms is prevented at the same time. Experiments demonstrate the effectiveness of the proposed algorithm. Xiang Zhang 0006, Jian Cheng 0001, Zhi Liu 0003, Jie Yang 0002 |
ICIP | 3 |
| 2013 | A novel region merging based image segmentation approach for automatic object extractionabstractThis paper presents a novel region merging based automatic image segmentation approach, which is applicable for object extraction. From an initial over-segmentation result, we exploit the regional histogram based similarity measure as merging criterion and merging order determination scheme with three priorities, to efficiently perform region merging, which is recorded using a binary partition tree (BPT). Based on the analysis of BPT, an appropriate subset of BPT nodes is selected to represent a meaningful image segmentation result and object extraction result. Experimental results demonstrate the better segmentation performance of our approach. Lin Zha, Zhi Liu 0003, Shuhua Luo, Liquan Shen |
ISCAS | 2 |
| 2013 | Stretchability-aware block scaling for image retargeting
Huan Du, Zhi Liu 0003, Jianliang Jiang, Liquan Shen |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | A novel H.264 rate control algorithm with consideration of visual attention
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002 |
Multim. Tools Appl. | 2 |
| 2013 | An Effective CU Size Decision Method for HEVC EncodersabstractThe emerging high efficiency video coding standard (HEVC) adopts the quadtree-structured coding unit (CU). Each CU allows recursive splitting into four equal sub-CUs. At each depth level (CU size), the test model of HEVC (HM) performs motion estimation (ME) with different sizes including 2N × 2N, 2N × N, N × 2N and N × N. ME process in HM is performed using all the possible depth levels and prediction modes to find the one with the least rate distortion (RD) cost using Lagrange multiplier. This achieves the highest coding efficiency but requires a very high computational complexity. In this paper, we propose a fast CU size decision algorithm for HM. Since the optimal depth level is highly content-dependent, it is not efficient to use all levels. We can determine CU depth range (including the minimum depth level and the maximum depth level) and skip some specific depth levels rarely used in the previous frame and neighboring CUs. Besides, the proposed algorithm also introduces early termination methods based on motion homogeneity checking, RD cost checking and SKIP mode checking to skip ME on unnecessary CU sizes. Experimental results demonstrate that the proposed algorithm can significantly reduce computational complexity while maintaining almost the same RD performance as the original HEVC encoder. Liquan Shen, Zhi Liu 0003, Xinpeng Zhang 0001, Zhaoyang Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2012 | Region Diversity Maximization for Salient Object DetectionabstractSalient object detection is an important technique for many content-based applications, but it becomes a challenging work when handling the cluttered saliency maps, which cannot completely highlight salient object regions and cannot suppress background regions. In this letter, we propose a novel approach to detect salient object from saliency map without manually setting any parameters. Region diversity maximization is used as the objective function to direct the object detection, and the optimal window for locating the salient object is obtained using an efficient iterative search scheme. Experimental results on different saliency maps demonstrate the overall better detection performance and computational efficiency of our approach. Zhi Liu 0003, Huan Du, Xiang Zhang 0006, Liquan Shen |
IEEE Signal Process. Lett. | 2 |
| 2012 | Unsupervised Salient Object Segmentation Based on Kernel Density Estimation and Two-Phase Graph CutabstractIn this paper, we propose an unsupervised salient object segmentation approach based on kernel density estimation (KDE) and two-phase graph cut. A set of KDE models are first constructed based on the pre-segmentation result of the input image, and then for each pixel, a set of likelihoods to fit all KDE models are calculated accordingly. The color saliency and spatial saliency of each KDE model are then evaluated based on its color distinctiveness and spatial distribution, and the pixel-wise saliency map is generated by integrating likelihood measures of pixels and saliency measures of KDE models. In the first phase of salient object segmentation, the saliency map based graph cut is exploited to obtain an initial segmentation result. In the second phase, the segmentation is further refined based on an iterative seed adjustment method, which efficiently utilizes the information of minimum cut generated using the KDE model based graph cut, and exploits a balancing weight update scheme for convergence of segmentation refinement. Experimental results on a dataset containing 1000 test images with ground truths demonstrate the better segmentation performance of our approach. Zhi Liu 0003, Liquan Shen, Yinzhu Xue, King Ngi Ngan, Zhaoyang Zhang 0002 |
IEEE Trans. Multim. | 1 |
| 2011 | Interactive object segmentation using iterative adjustable graph cutabstractInteractive object segmentation is widely used for extracting any user-interested objects from natural images. A common problem with many interactive segmentation approaches is that the object segmentation quality is degraded due to inaccurate object/background seeds provided by the user. This paper proposes an iterative adjustable graph cut to efficiently solve this problem. First, object/background seeds are initialized based on the object segmentation result obtained with the user-specified scribbles as the interactive input. Then, an iterative seed adjustment scheme is exploited to correct inaccurate seeds and extract new suitable seeds via graph cut, in which the balancing weight between energy terms are adaptively updated to protect stable seeds and speedup the iteration process. Finally, suitable seeds are obtained and graph cut is used to segment the objects. Experimental results demonstrate the better segmentation performance of our approach even if user provides rather rough seeds. Zhi Liu 0003, Yinzhu Xue, Xiang Zhang 0006 |
VCIP | 2 |
| 2011 | Unsupervised image segmentation based on analysis of binary partition tree for salient object extraction
Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002 |
Signal Process. | 1 |
| 2011 | Low-Complexity Mode Decision for MVCabstractThe finalized international standard for multiview video coding (MVC) is an extension of H.264. In the joint model of MVC, variable size motion estimation (ME) and disparity estimation (DE) are introduced to achieve the highest coding efficiency with the cost of very high computational complexity. A low complexity mode decision algorithm is proposed to reduce complexity of ME and DE. An experimental analysis is performed to study inter-view correlation in the coding information such as the prediction mode and rate-distortion (RD) cost. Based on the correlation, we propose four efficient mode decision techniques, including early SKIP mode decision, adaptive early termination, fast mode size decision, and selective intra coding in inter frame. Experimental results show that the proposed algorithm can significantly reduce computational complexity of MVC while maintaining almost the same RD performance. Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Nonparametric saliency detection using kernel density estimationabstractThis paper proposes a nonparametric saliency model based on kernel density estimation (KDE) mainly aiming at content-based applications such as salient object segmentation. A set of KDE models are constructed on the basis of regions segmented using the mean shift algorithm. For each pixel, a set of color likelihood measures to all KDE models are calculated, and then the color saliency and spatial saliency of each KDE model are evaluated based on its color distinctiveness and spatial distribution. The final saliency map is generated by combining saliency measures of KDE models and color likelihood measures of pixels. Experimental results demonstrate the better saliency detection performance of our saliency model. Zhi Liu 0003, Yinzhu Xue, Liquan Shen, Zhaoyang Zhang 0002 |
ICIP | 1 |
| 2010 | An adaptive early termination of mode decision using inter-layer correlation in scalable video codingabstractThe scalable video coding (SVC) standard adopts the variable size motion estimation (ME) to select the best coding mode for each macroblock (MB). Although this technique achieves the highest possible coding efficiency, it results in extremely large computation complexity which obstructs SVC from the practical application. In this paper, we propose an adaptive early termination of fast mode decision algorithm in SVC. It makes use of the coding information of spatial neighbor MBs and the corresponding MBs in base layer to early terminate the mode decision procedure. Experimental results show that the proposed fast mode decision algorithm can achieve computational saving up to 67% with no significant loss of rate distortion (RD) performance. Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002 |
ICIP | 2 |
| 2010 | Unsupervised salient object segmentation from color imagesabstractThis paper proposes an efficient approach for unsupervised segmentation of salient objects from color images. A set of Gaussian models are first estimated based on a pre-segmentation result of the input image, and then for each pixel, a set of normalized color likelihood measures to each Gaussian model are calculated. The color saliency and spatial saliency of Gaussian models are exploited to generate the pixel-wise saliency map. By thresholding the saliency map, the pixels are classified into object seed pixels, background seed pixels and uncertain pixels to obtain the trimap. For each pixel, the probability belonging to salient object/background is evaluated using kernel density estimation, and the geodesic distances to salient object and background are calculated based on the object likelihood map. By comparing the two geodesic distances, uncertain pixels are finally classified into salient object or background. Experimental results demonstrate the better segmentation performance of the proposed approach. Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002 |
VCIP | 1 |
| 2010 | Automatic segmentation of focused objects from images with low depth of field
Zhi Liu 0003, Liquan Shen, Zhongmin Han, Zhaoyang Zhang 0002 |
Pattern Recognit. Lett. | 1 |
| 2010 | Early SKIP mode decision for MVC using inter-view correlation
Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001 |
Signal Process. Image Commun. | 2 |
| 2010 | Efficient SKIP Mode Detection for Coarse Grain Quality Scalable Video CodingabstractScalable video coding (SVC) was recently standardized by the Joint Video Team as an extension of H.264. In SVC, a computationally expensive exhaustive mode decision is employed to select the best coding mode for each macroblock (MB), which achieves a high coding efficiency. In order to reduce computational complexity, we propose an efficient SKIP mode detection approach for coarse grain quality SVC. It makes use of the coding information of spatial neighboring MBs and the co-located MB in base layer to predict the SKIP mode MB and early terminate its mode decision procedure. Experimental results show that the proposed early SKIP mode decision approach can achieve the average computational saving about 54% with almost no loss of rate distortion (RD) performance in the enhancement layer. Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2010 | View-Adaptive Motion Estimation and Disparity Estimation for Low Complexity Multiview Video CodingabstractThe emerging international standard for multiview video coding (MVC) is an extension of H.264/advanced video coding. In the joint mode of MVC, both motion estimation (ME) and disparity estimation (DE) are included in the encoding process. This achieves the highest coding efficiency but requires a very high computational complexity. In this letter, we propose a fast ME and DE algorithm that adaptively utilizes the inter-view correlation. The coding mode complexity and the motion homogeneity of a macroblock (MB) are first analyzed according to the coding modes and motion vectors from the corresponding MBs in the neighbor views, which are located by means of global disparity vector. According to the coding mode complexity and the motion homogeneity, the proposed algorithm adjusts the search strategies for different types of MBs in order to perform a precise search according to video content. Experimental results demonstrate that the proposed algorithm can save 85% computational complexity on average, with negligible loss of coding efficiency. Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Fast mode decision for multiview video codingabstractIn the draft of multi-view coding (MVC), variable size motion estimation and disparity estimation are employed to select the best coding mode for each macroblock. These techniques achieve the highest possible coding efficiency, but they result in extremely large computation complexity which obstructs MVC from practical application. This paper proposes a fast mode size decision algorithm for MVC in inter-frame coding. It makes use of the mode distribution correlation between neighbor views to deduct the executions of unnecessary modes. Experimental results show that the proposed fast mode decision algorithm reduces the computational complexity significantly with negligible coding efficiency. Liquan Shen, Tao Yan 0003, Zhi Liu 0003, Zhaoyang Zhang 0002, Ping An 0001 |
ICIP | 3 |
| 2009 | Frame-level bit allocation based on incremental PID algorithm and frame complexity estimation
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi |
J. Vis. Commun. Image Represent. | 2 |
| 2009 | Selective VS-MRF-ME and intra coding in H.264 based on spatiotemporal continuity of motion field
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi |
Signal Process. Image Commun. | 2 |
| 2009 | An Efficient Intermode Decision Algorithm Based on Motion Homogeneity for H.264/AVCabstractThe latest video coding standard H.264/AVC significantly outperforms previous standards in terms of coding efficiency. H.264/AVC adopts variable block sizes ranging from 4 times 4 to 16 times 16 in inter frame coding, and achieves significant gain in coding efficiency compared to coding a macroblock (MB) using regular block size. However, this new feature causes extremely high computation complexity when rate-distortion optimization (RDO) is performed using the scheme of full mode decision. This paper presents an efficient intermode decision algorithm based on motion homogeneity evaluated on a normalized motion vector (MV) field, which is generated using MVs from motion estimation on the block size of 4 times 4. Three directional motion homogeneity measures derived from the normalized MV field are exploited to determine a subset of candidate intermodes for each MB, and unnecessary RDO calculations on other intermodes can be skipped. Experimental results demonstrate that our algorithm can reduce the entire encoding time about 40% on average, without any noticeable loss of coding efficiency. Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Fast Inter Mode Decision Using Spatial Property of Motion FieldabstractVariable size motion estimation with multiple reference frames has been adopted by the new video coding standard H.264. It can achieve significant coding efficiency compared to coding a macroblock (MB) in regular size with single reference frame. On the other hand, it causes high computational complexity of motion estimation at the encoder. Rate distortion optimized (RDO) decision is one powerful method to choose the best coding mode among all combinations of block sizes and reference frames, but it requires extremely high computation. In this paper, a fast inter mode decision is proposed to decide best prediction mode utilizing the spatial continuity of motion field, which is generated by motion vectors from 4times4 motion estimation. Motion continuity of each MB is decided based on the motion edge map detected by the Sobel operator. Based on the motion continuity of a MB, only a small number of block sizes are selected in motion estimation and RDO computation process. Simulation results show that our algorithm can save more than 50% computational complexity, with negligible loss of coding efficiency. Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi |
IEEE Trans. Multim. | 2 |
| 2007 | A novel algorithm to fast mode decision with consideration about video texture in H.264abstractH.264 employs 7 different size block types for motion estimation that can significantly improve the coding performance compared with the previous video coding standards. However, H.264 requires extremely high computation with the R-D optimized decision since so many prediction modes are used. In this paper, a novel inter mode decision algorithm (NIMDA) is proposed that utilizes SADs of each 4X4 block and texture characteristic to reduce the candidate mode set after the 16 X16 prediction mode is tested. The simulation results show that the proposed algorithm reduces the entire encoding time by 64.52% with only negligible coding loss. Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002 |
AICCSA | 2 |
| 2007 | Video nature considerations for multi-frame selection algorithm in H.264abstractH.264 allows motion estimation performing on multiple reference frames. This new feature improves the prediction accuracy of inter-coding blocks significantly. However, the coding gain comes at the cost of a much higher computational complexity. The reference software JM adopts full search scheme, and the computational complexity of motion estimation increases linearly with the number of allowed reference frames. In fact, the reduction of prediction residues is highly dependent on the nature of sequences, not on the number of searched frames. In this paper, with consideration of video nature and the available information from previous searched reference frames, an adaptive multi-frame selection algorithm (AMFSA) is proposed to speed up the matching process for multiple reference frames in the H.264 video coding system. The proposed algorithm can effectively reduce 63.6% on average. Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002 |
AICCSA | 2 |
| 2007 | A Novel Video Object Tracking Approach Based on Kernel Density Estimation and Markov Random FieldabstractIn this paper, we propose a novel video object tracking approach based on kernel density estimation and Markov random field (MRF). The interested video objects are first segmented by the user, and a nonparametric model based on kernel density estimation is initialized for each video object and the remaining background, respectively. A temporal saliency map is also initialized for each object to memorize the temporal trajectory. Based on the probabilities evaluated on the non-parametric models, each pixel in the current frame is first classified into the corresponding video object or background using the maximum likelihood criterion. Starting from the initial classification result, a MRF model that combines spatial smoothness and temporal coherency is selectively exploited to generate more reliable video objects. The nonparametric model and the temporal saliency map for each video object are updated and propagated for the future tracking. Experimental results on several MPEG-4 test sequences demonstrate the good segmentation performance of our approach. Zhi Liu 0003, Liquan Shen, Zhongmin Han, Zhaoyang Zhang 0002 |
ICIP (3) | 1 |
| 2007 | An Adaptive and Fast H.264 Multi-Frame Selection Algorithm Based on Information from Previous SearchesabstractThe H.264 video coding standard adopts multiple reference frames for motion estimation. This new feature improves the prediction accuracy of inter-coding blocks significantly, but it results in a considerable increase in encoder complexity, mainly regarding to multi-frame selection and motion estimation. The reference software JM adopts the full search scheme, and the increased computation is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the nature of sequences, not on the number of searched frames. In this paper, we propose an adaptive and fast multi-frame selection algorithm (AFMFS) based on motion vectors and SAD information coming from previous searches to adaptively terminate the procedure of multiple reference frames selection. Compared with the full search algorithm and the flexible multi-reference frame search criterion (FMRFSC), simulation results show that the proposed algorithm can save 55.28% and 40.23% computation cost on average, respectively, while it still maintains similar coding efficiency. Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi |
ICME | 2 |
| 2007 | Real-time spatiotemporal segmentation of video objects in the H.264 compressed domain
Zhi Liu 0003, Zhaoyang Zhang 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2007 | An adaptive and fast fractional pixel search algorithm in H.264
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003, Wenjun Zhang 0001 |
Signal Process. | 3 |
| 2007 | An Adaptive and Fast Multiframe Selection Algorithm for H.264 Video CodingabstractH.264 allows motion estimation performing on multiple reference frames. This new feature improves the prediction accuracy of inter-coding blocks significantly, but it is extremely computational intensive. The reference software JM adopts full search scheme, and the increased computation load is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the content of sequences, not on the number of searched frames. In this letter, we propose an adaptive and fast multiframe selection algorithm (AFMFSA) to speed up the searching procedure for multiple reference frames. Simulation results show that the proposed algorithm can deduct 56.0%–74.2% computation load of motion estimation on average. Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002 |
IEEE Signal Process. Lett. | 2 |
| 2005 | Semi-automatic video object segmentation using seeded region merging and bidirectional projection
Zhi Liu 0003, Jie Yang 0002, Ningsong Peng |
Pattern Recognit. Lett. | 1 |
| 2005 | Mean shift blob tracking with kernel histogram filtering and hypothesis testing
Ningsong Peng, Jie Yang 0002, Zhi Liu 0003 |
Pattern Recognit. Lett. | 3 |
| 2005 | An efficient face segmentation algorithm based on binary partition tree
Zhi Liu 0003, Jie Yang 0002, Ningsong Peng |
Signal Process. Image Commun. | 1 |