Wenbin Zou

dblp:126/0718 · DBLP profile ↗
← Back
92ranked-venue papers
13as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 12 first-author · 32 since 2021Artificial intelligence and machine learning · 32 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic-guided policy network for zero-shot object goal visual navigation
Guoguang Hua, Yaqiong Ding, Yuhuan Chen, Dan Xiang, Wenbin Zou
Knowl. Based Syst.5
2026 MLWAC: A Modular, Low-coupling Waypoint-Angular Coordinated Network for visual navigation in unstructured environments
Yongdong Guo, Muxin Liao, Shishun Tian, Wenbin Zou, Chen Xu 0004
Knowl. Based Syst.7
2026 Training-Free VFM-Guided Dynamic Refinement for Domain Generalized Semantic Segmentation
abstract
Domain Generalized Semantic Segmentation (DGSS) has recently attracted lots of research attention aiming to achieve robust segmentation performance on unseen domains, which aligns with diverse real-world applications. Visual Foundation Model (VFM), depending on large-scale pre-trained data and exquisite training strategies with good generalizability, has been explored and mined in some downstream tasks like DGSS. However, existing VFM-based DGSS methods predominantly focus on fine-tuning the VFM to adapt them to the semantic segmentation task. While this improves task alignment, it inevitably introduces additional training overhead and still results in static prediction behavior at inference time, making the model unable to adapt to varying target-domain distributions during deployment. To address these issues, we propose aTraining-free VFM-guided Dynamic Refinement (TVDR) frameworkfor the DGSS task, operating at the inference stage. First, a maximum category voting module is proposed to smooth the segmentation result of large masks generated by Segment Anything Model (SAM). Second, a small-object under-segmentation optimization strategy is proposed to improve the generalization of the small objects. Finally, a fusion refinement module is designed to refine the segmentation results of the above strategies. Compared with existing methods, our strategy does not require any parameter updates for VFM and has the advantages of plug-and-play and flexible deployment. It provides an efficient and practical new paradigm for cross-domain semantic segmentation tasks. Extensive experiments on widely-used benchmarks verified the effectiveness of the proposed approaches. Code is available at: https://github.com/Hectoor/TVDR.
Yuhang Zhang 0011, Binbin Wei, Tiantian Zeng, Wenbin Zou
IEEE Trans. Circuits Syst. Video Technol.6
2026 Pancreas Segmentation With Multi-Phase Feature Aggregation and Modality Adaptive Transformer
abstract
Automatic pancreas segmentation can facilitate diagnosis and treatment of pancreatic diseases. The combination of non-contrast, arterial, and venous phases of CT imaging can enhance differentiation of the pancreas from its surrounding structures. However, existing multimodal methods, which try to integrate the multimodal information in computer-aided pancreas segmentation, often overlook the inter-modal relationships and have a limited capability for information fusion. In this paper, we propose a multi-phase pancreas segmentation method for incorporating Feature Aggregation Module (FAM) and Modality Adaptive Transformer (MAT). Specifically, we use the venous phase as the primary modality, while the non-contrast and arterial phases serve as supplementary modalities, based on clinical prior knowledge. Our FAM integrates spatial information from the primary and supplementary modalities, while our MAT adaptively enhances feature representation and establishes long-range dependencies among modalities. Our method outperforms state-of-the-art techniques on a large scale dataset. Based on the segmented pancreas region, We further perform a downstream task focused on pancreatic volume calculation. The prediction accuracy is on par with manual segmentation, demonstrating effectiveness and potential application of our proposed method.
Lulu Tan, Wenda Sheng, Wenbin Zou, Dengqiang Jia, Qianqian Chen 0002, Jing Sheng, Yangyang Qian, Qianjin Feng 0001, Zhuan Liao, Dinggang Shen
IEEE J. Biomed. Health Informatics4
2025 Prior-Guided Test Time Adaptation for Blind Image Quality Assessment
abstract
Current blind image quality assessment (BIQA) models usually lack adaptability to the test data with distribution shifts to the training data. This inspires an investigation into test time adaptation (TTA) methods to address distribution shifts between training and test data. However, existing methods mainly focus on simple feature alignment strategies, which may lead to incorrect knowledge generalization. To this issue, we propose a prior-guided test time adaptation (PGTA-IQA) for blind image quality assessment. Concretely, we extract the quality prior knowledge from the pre-trained BIQA model through clustering. The extracted quality prior knowledge forms the foundation for subsequent optimizations. These optimizations are carried out from two complementary perspectives: inter-cluster and intra-cluster. From the inter-cluster perspective, we propose a confident rank learning approach which consists of a relative quality matrix (RQM) and a confidence filtering strategy (CFS) to generate the high-confident quality rankings. From the intra-cluster perspective, we propose a selective feature alignment approach by only aligning the closest neighboring samples within the same cluster to reduce the impact of noisy labels. The experimental results demonstrate the effectiveness of the proposed approaches.
Shishun Tian, Fangjie Hou, Guanghui Yue 0001, Yuanhao Gong, Wenbin Zou, Ting Su 0004
ICME5
2025 Probabilistic Mixture of Hyperbolic Mamba for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) grapples with the dual challenge of learning new classes from minimal labeled training data while alleviating catastrophic forgetting of previous learned classes. Compared with previous methods employing static adaptation on specific parameters, current works verify that dynamic weights and sequence modeling in Selective State Space Models (SSMs) can capture distinctive feature drifts in FSCIL. However, the flattening operation in SSMs fragments the latent semantic relationship, where the resulting task isolation and representation degeneration are detrimental to FSCIL. Toward this issue, this paper presents a novel framework named Probabilistic Mixture of Hyperbolic State Space Experts (PmH-SSE) for FSCIL. First, since SSMs rely on scanning as an alternative to self-attention, the Hyperbolic state space model with multi-scale hybrid scan is built to facilitate few-shot learning by providing an extra Hyperbolic geometry that encodes hierarchical relationships. Moreover, we propose the probabilistic mixture of Mamba to increase the model's flexibility in handling non-stationary data streams in FSCIL and enhance the stability of high-parameter models in few-shot conditions. Finally, under the same experimental conditions, the proposed PmH-SSE demonstrates superior performance in comprehensive experiments. The codes are available at https://github.com/yawencui/PmH-SSE.
Yawen Cui, Wenbin Zou, Huiping Zhuang, Yi Wang 0068, Lap-Pui Chau
ACM Multimedia2
2025 Learning Content-enhanced Tokens for Domain Generalized Semantic Segmentation
abstract
Visual foundation models (VFMs) have demonstrated impressive generalization capabilities in computer vision tasks. Previous studies show that fine-tuning VFMs with learnable tokens can achieve better generalization performance than full-parameter fine-tuning. The problem we need to address is how to learn the tokens that focus on the content information while ignoring the influence of style. For this purpose, we propose a novel Dual-Branch Content-enhanced Token (DBCT) learning framework. Specifically, we construct a style-suppressing branch, which contains a Style-sensitive Channel Suppression (SCS) module to transform the frozen VFM features into style-suppressed features, enabling the learning of style-invariant tokens. In addition, to compensate for the content degradation caused by the style-suppressing branch, we introduce a content-preserving branch that directly takes the frozen VFM features as input to learn content-focused tokens. Meanwhile, we propose a Token-query Linking (TLink) strategy to connect the two sets of tokens with the queries in the decoder. Through extensive experiments, our method achieves advanced results on various benchmarks.
Shishun Tian, Wenbin Zou, Yuanhao Gong, Guanghui Yue 0001, Ting Su 0004
MMAsia3
2025 A progressive segmentation network for navigable areas with semantic-spatial information flow
Muxin Liao, Wenbin Zou
Expert Syst. Appl.3
2025 A terrain segmentation network for navigable areas with global strip reliability evaluation and dynamic fusion
Muxin Liao, Wenbin Zou
Expert Syst. Appl.3
2025 Contextual-aware terrain segmentation network for navigable areas with triple aggregation
Muxin Liao, Wenbin Zou
Expert Syst. Appl.3
2025 Task-Oriented Semantic Communication With Adaptive Semantic Reconstruction Network
abstract
In recent years, semantic communication has garnered significant attention for its potential to address challenges in traditional communication systems. However, in complex communication environments, semantic communication still faces challenges such as semantic information loss, low transmission efficiency, and poor adaptability. This paper proposes a novel Semantic Communication with Adaptive Semantic Reconstruction (SCASR) scheme to enhance transmission efficiency and adaptability in complex communication environments. First, a compression mechanism based on semantic importance is designed to achieve flexible and efficient semantic compression. Then, we develop an adaptive semantic reconstruction network to predict and reconstruct lost semantic information. Finally, we integrate an attention mechanism into the reconstruction network, dynamically adjusting parameter weights based on Signal-to-Noise Ratio (SNR), Semantic Compression Rate (SCR), and Packet Loss Rate (PLR) to improve reconstruction quality and adaptability. To evaluate the efficiency of SCASR, we conduct extensive simulation experiments on semantic segmentation tasks using the Cityscapes dataset. Results demonstrate that SCASR outperforms existing semantic communication and traditional schemes, offering higher Mean Intersection over Union (mIoU), and enhanced Semantic Transmission Benefit (STB).
Zhu Jin, Tiecheng Song, Wen-Kang Jia 0001, Wenbin Zou, Xiaoqin Song
IEEE Internet Things J.4
2025 Class-discriminative domain generalization for semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Rong You, Wenbin Zou, Xia Li 0006
Image Vis. Comput.6
2025 A global reweighting approach for cross-domain semantic segmentation
Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Guoguang Hua, Wenbin Zou, Chen Xu 0004
Signal Process. Image Commun.5
2025 Contextual Guidance Network for Real-Time Semantic Segmentation of Autonomous Driving
abstract
With the rise of mobile computing and the increasing demand for real-time applications, the need for efficient and accurate semantic segmentation models has become paramount. However, existing state-of-the-art models are often hindered by heavy computational requirements, rendering them impractical for real-time applications. To tackle this challenge, we introduce the Contextual Guidance Network (CGNet), an efficient and lightweight network designed specifically for real-time semantic segmentation in autonomous driving. CGNet primarily consists of two key components: the Contextual Guidance Module (CGM) and the Triple-Branch Residual Fusion Module (TRFM). The CGM is comprised of the Downsampling Refine Unit (DRU) and the Contextual Guidance Bottleneck (CGB), which are utilized to gather dense contextual information. The DRU functions as a downsampling tool to generate low-resolution images, while the CGB extracts rich contextual information from both spatial and channel dimensions. Additionally, the TRFM utilizes the Residual Fusion Module (RFM) to achieve feature fusion and enhance pixel prediction accuracy. Without bells and whistles, CGNet achieves impressive mean intersection over union (mIoU) scores of 77.11% with 1.00 million parameters at 86.71 frames per second (fps) on the Cityscapes dataset, 72.26% mIoU at 88.62 fps on the CamVid dataset, and 63.32% mIoU on the BDD100K dataset. Extensive experiments demonstrate that CGNet achieves a favorable tradeoff between segmentation accuracy, inference speed and computational cost, making it suitable for autonomous driving systems with limited hardware resources. The source code will be available on GitHub: https://github.com/lv881314/CGNet
Muxin Liao, Guoguang Hua, Yuhang Zhang 0011, Wenbin Zou
IEEE Trans. Intell. Transp. Syst.5
2025 Progressive Terrain Segmentation Network for Navigable Areas With Global Sparsity-Entropy and Fusion-Awareness
abstract
Precise segmentation of safe navigable areas is crucial for wild scene parsing in self-driving systems. Previous research has demonstrated that effective feature representation enhances model performance, yet few methods thoroughly explore the complementary relationships between features at different scales in complex wild environments. In this paper, we propose a Progressive Terrain Segmentation Network (PTSNet) for the segmentation of navigable areas, which introduces global contextual information as prior knowledge for fusion and delivers robust feature representation by progressively exploring the complementary relationships between multi-scale features from both spatial and channel perspectives. PTSNet consists of two main components: the Global Sparsity-Entropy Module (GSEM) and the Fusion-Awareness Module (FAM). The GSEM, based on self-attention, employs Top-K sparsity and entropy refinement to effectively capture global semantic information with long-range dependencies, and the information derived from GSEM act as prior knowledge to guide feature fusion. Additionally, we propose the FAM, which consists of the Attention Aggregation Unit (AAU) and the Contribution-aware Unit (CAU), to explore complementary relationships in multi-scale feature interactions and obtain sufficient scene information. Extensive experiments conducted on various wild datasets demonstrate that PTSNet outperforms state-of-the-art methods in accurately segmenting navigable areas, offering a new solution for the safe operation of self-driving systems in wild environments. The code will be available athttps://github.com/lv881314/PTSNet
Muxin Liao, Wenbin Zou
IEEE Trans. Intell. Transp. Syst.3
2025 Class-Balanced Sampling and Discriminative Stylization for Domain Generalization Semantic Segmentation
abstract
Existing domain generalization semantic segmentation (DGSS) methods have achieved remarkable performance on unseen domains by generating stylized images to increase the diversity of training data. However, since the training data is usually class-imbalanced, uniform style randomization is unable to generate diverse minority classes. This means that models may overfit to the minority classes, resulting in suboptimal performance on the minority classes. In addition, the image-level style randomization may also corrupt the class-discriminative regions of objects, leading to a loss of the class-discriminative representation. To address these issues, a novel class-balanced sampling and discriminative stylization (CSDS) approach is proposed for DGSS. Specifically, first, a pixel-level class-balanced sampling (PCS) strategy is proposed to adaptively sample patches of the minority classes from the source domain images and paste the sampled patches on the input images. Unlike existing class sampling strategies that fix the minority classes, the PCS strategy dynamically determines the minority classes by estimating the class distribution after each sampling. Then, a class-discriminative style randomization (CSR) strategy is proposed to increase the style diversity of the sampled patches while preserving the class-discriminative regions. Finally, since the pasting positions of the sampled patches are uncertain, which may confuse the semantic relations between the classes, a semantic consistency constraint is proposed to ensure the learning of reliable semantic relations. Extensive experiments demonstrate that the proposed approach achieves superior performance compared to existing DGSS methods on multiple benchmarks. The source code has been released onhttps://github.com/seabearlmx/CSDS.
Muxin Liao, Shishun Tian, Binbin Wei, Yuhang Zhang 0011, Wenbin Zou, Xia Li 0006
IEEE Trans. Intell. Transp. Syst.5
2025 Multi-Level Guided Discrepancy Learning for Source-Free Object Detection in Hazy Conditions
abstract
Haze deteriorates the quality of captured images, which severely limits the accuracy of clean image-trained object detectors in hazy conditions. Source-free domain adaptation (SFDA) aims to adapt a clean source image-trained detector to the unlabeled hazy target domain without access to the clean source domain data. However, existing source-free object detection (SFOD) methods encounter two issues when leveraging pseudo labeling paradigm: 1) the large domain shift between clear and hazy images may introduce noises in pseudo labels, 2) the single confidence threshold-based methods may ignore valuable information of the low-confidence samples. To address these issues, we propose a multi-level guided discrepancy learning-based approach for SFOD in hazy conditions, named MGDL. Specifically, we first propose a differentiated enhancement module (DEM) to intentionally augment the diversity of data styles. It uses dehazing and random perturbation to generate reliable high-quality labels and strengthen the tolerance to haze-related factors. To enhance the consistency constraints for discrepancy learning, we propose a multi-level guiding strategy (MLGS) which consists of a dual label-level and an instance-level guidance. Considering that false negatives may dominate in noisy labels, we propose a dual label-level guiding (DLG) strategy to excavate comprehensive useful information from high-and low-confidence samples. Besides, an instance-level contrastive learning (ICL) approach is proposed to guide the model to focus on objects and make the model insensitive to style at the same time. Extensive experiments conducted on multiple datasets have demonstrated that our method achieves superior performance over the state-of-the-art SFOD methods.
Shishun Tian, Tiantian Zeng, Wenbin Zou
IEEE Trans. Intell. Transp. Syst.4
2025 Prototypical Progressive Alignment and Reweighting for Generalizable Semantic Segmentation
abstract
Generalizable semantic segmentation, aims to excel on unseen target domains, as a critical focus due to the widespread practical applications requiring high generalizability. Class-wise prototypes, which depict class-wise centroids, as a type of domain-invariant information are key to improving the model generalizability due to its stability and representativeness. However, this manner faces some challenges. First, the existing methods adopt a coarse prototypical alignment form, potentially compromising performance. Second, the naive prototype generally serves as the class centroid generated by an average operation from source data batches, risks source domain overfitting, and may be detrimentally impacted by unrelated source data. Third, from a broader perspective, rather than just from a prototypical alignment perspective, the existing methods treat all samples equally, which is against the conclusion that different source features have different adaptation difficulties. To tackle these issues, we propose a novel method for generalizable semantic segmentation called Prototypical Progressive Alignment and Reweighting (PPAR) depending on the strong generalized representation of the Contrastive Language-Image Pretraining (CLIP) model. In particular, we first define the Original Text Prototype (OTP) and Visual Text Prototype (VTP) generated by the CLIP model, laying the foundation for the subsequent effective alignment strategy. Then, we propose a prototypical progressive alignment strategy by an easy-to-difficult alignment form to reduce domain-variant information progressively instead of directly. Finally, we propose a prototypical reweighting learning strategy that estimates the importance of the source data and corrects its learning weight to alleviate the influence of unrelated source features, i.e. alleviate negative transfer. Moreover, we also offer a theoretical insight into our method and it shows that our method compiles well on the domain generalization theory. Extensive experiments on several popular datasets demonstrate that our PPAR method achieves superior performance, proving the effectiveness of our method. The source code will available at: https://github.com/Hectoor/PPAR
Yuhang Zhang 0011, Muxin Liao, Shishun Tian, Wenbin Zou, Lu Zhang 0037, Chen Xu 0004
IEEE Trans. Intell. Transp. Syst.5
2025 Dual Residual-Guided Interactive Learning for the Quality Assessment of Enhanced Images
abstract
Image enhancement algorithms can facilitate computer vision tasks in real applications. However, various distortions may also be introduced by image enhancement algorithms. Therefore, the image quality assessment (IQA) plays a crucial role in accurately evaluating enhanced images to provide dependable feedback. Current enhanced IQA methods are mainly designed for single specific scenarios, resulting in limited performance in other scenarios. Besides, no-reference methods predict quality utilizing enhanced images alone, which ignores the existing degraded images that contain valuable information, are not reliable enough. In this work, we propose a degraded-reference image quality assessment method based on dual residual-guided interactive learning (DRGQA) for the enhanced images in multiple scenarios. Specifically, a global and local feature collaboration module (GLCM) is proposed to imitate the perception of observers to capture comprehensive quality-aware features by using convolutional neural networks (CNN) and Transformers in an interactive manner. Then, we investigate the structure damage and color shift distortions that commonly occur in the enhanced images and propose a dual residual-guided module (DRGM) to make the model concentrate on the distorted regions that are sensitive to human visual system (HVS). Furthermore, a distortion-aware feature enhancement module (DEM) is proposed to improve the representation abilities of features in deeper networks. Extensive experimental results demonstrate that our proposed DRGQA achieves superior performance with lower computational complexity compared to the state-of-the-art IQA methods.
Shishun Tian, Tiantian Zeng, Wenbin Zou, Xia Li 0006
IEEE Trans. Multim.4
2024 VQCNIR: Clearer Night Image Restoration with Vector-Quantized Codebook
abstract
Night photography often struggles with challenges like low light and blurring, stemming from dark environments and prolonged exposures. Current methods either disregard priors and directly fitting end-to-end networks, leading to inconsistent illumination, or rely on unreliable handcrafted priors to constrain the network, thereby bringing the greater error to the final result. We believe in the strength of data-driven high-quality priors and strive to offer a reliable and consistent prior, circumventing the restrictions of manual priors. In this paper, we propose Clearer Night Image Restoration with Vector-Quantized Codebook (VQCNIR) to achieve remarkable and consistent restoration outcomes on real-world and synthetic benchmarks. To ensure the faithful restoration of details and illumination, we propose the incorporation of two essential modules: the Adaptive Illumination Enhancement Module (AIEM) and the Deformable Bi-directional Cross-Attention (DBCA) module. The AIEM leverages the inter-channel correlation of features to dynamically maintain illumination consistency between degraded features and high-quality codebook features. Meanwhile, the DBCA module effectively integrates texture and structural information through bi-directional cross-attention and deformable convolution, resulting in enhanced fine-grained detail and structural fidelity across parallel decoders. Extensive experiments validate the remarkable benefits of VQCNIR in enhancing image quality under low-light conditions, showcasing its state-of-the-art performance on both synthetic and real-world datasets. The code is available at https://github.com/AlexZou14/VQCNIR.
Wenbin Zou, Hongxia Gao, Tian Ye 0001, Liang Chen 0026, Weipeng Yang 0002, Shasha Huang, Sixiang Chen
AAAI1
2024 Low-Light Image Enhancement via Weighted Low-Rank Tensor Regularized Retinex Model
abstract
Images captured under low light conditions are often affected by intense noise, which may become more pronounced during image enhancement, resulting in poor visual quality. The aim of this paper is to establish an effective low-light image enhancement model that can suppress noise and artifacts while preserving image details. To deal with intense noise, we propose a Weighted Low-Rank Tensor regularization Retinex (WLRT-Retinex) model, which introduces weighted low-rank tensor priors in the Retinex decomposition process to suppress noise and artifacts in the reflectance. Furthermore, since noise in dark areas is typically more severe, we introduce an illumination-aware weighting scheme in the total variation regularization term of the reflectance, which helps achieve adaptive denoising and preserve details in bright areas. Experiments on seven challenging datasets demonstrate the effectiveness of the proposed method, achieving better or comparable performance compared with state-of-the-art methods. Our code is available at https://github.com/YangWeipengscut/WLRT-Retinex.
Weipeng Yang 0002, Hongxia Gao, Wenbin Zou, Tongtong Liu 0003, Shasha Huang, Jianliang Ma
ICMR3
2024 Wave-Mamba: Wavelet State Space Model for Ultra-High-Definition Low-Light Image Enhancement
abstract
Ultra-high-definition (UHD) technology has attracted widespread attention due to its exceptional visual quality, but it also poses new challenges for low-light image enhancement (LLIE) techniques. UHD images inherently possess high computational complexity, leading existing UHD LLIE methods to employ high-magnification downsampling to reduce computational costs, which in turn results in information loss. The wavelet transform not only allows downsampling without loss of information, but also separates the image content from the noise. It enables state space models (SSMs) to avoid being affected by noise when modeling long sequences, thus making full use of the long-sequence modeling capability of SSMs. On this basis, we propose Wave-Mamba, a novel approach based on two pivotal insights derived from the wavelet domain: 1) most of the content information of an image exists in the low-frequency component, less in the high-frequency component. 2) The high-frequency component exerts a minimal influence on the outcomes of low-light enhancement. Specifically, to efficiently model global content information on UHD images, we proposed a low-frequency state space block (LFSSBlock) by improving SSMs to focus on restoring the information of low-frequency sub-bands. Moreover, we propose a high-frequency enhance block (HFEBlock) for high-frequency sub-band information, which uses the enhanced low-frequency information to correct the high-frequency information and effectively restore the correct high-frequency details. Through comprehensive evaluation, our method has demonstrated superior performance, significantly outshining current leading techniques while maintaining a more streamlined architecture. The code is available at https://github.com/AlexZou14/Wave-Mamba.
Wenbin Zou, Hongxia Gao, Weipeng Yang 0002, Tongtong Liu 0003
ACM Multimedia1
2024 Layout Relationship Decoupling Framework for Multi-target Domain Adaptative Semantic Segmentation
Yuhang Zhang 0011, Cuixin Yang, Muxin Liao, Shishun Tian, Wenbin Zou, Chen Xu 0004
MMAsia5
2024 Employing Multiple Priors in Retinex-Based Low-Light Image Enhancement
Weipeng Yang 0002, Hongxia Gao, Tongtong Liu 0003, Jianliang Ma, Wenbin Zou, Shasha Huang
EGSR (ST)5
2024 Strip and asymmetric aggregation network for unstructured terrain segmentation in wild environments
Shishun Tian, Yuhang Zhang 0011, Muxin Liao, Guoguang Hua, Wenbin Zou
Eng. Appl. Artif. Intell.6
2024 PDA: Progressive Domain Adaptation for Semantic Segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
Knowl. Based Syst.5
2024 Considering representation diversity and prediction consistency for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
Knowl. Based Syst.5
2024 Video Generalized Semantic Segmentation via Non-Salient Feature Reasoning and Consistency
Yuhang Zhang 0011, Muxin Liao, Shishun Tian, Rong You, Wenbin Zou, Chen Xu 0004
Knowl. Based Syst.6
2024 MI-RPN: Integrating multi-modalities and multi-scales information for region proposal
Shishun Tian, Wenbin Zou, Xia Li 0006
Multim. Tools Appl.3
2024 Fine-Grained Self-Supervision for Generalizable Semantic Segmentation
abstract
Unsupervised domain adaptative semantic segmentation is a powerful solution for the distribution shift problem between the source and target domains. However, such methods need specified target domain data that may be unavailable in actual applications due to excess expensive collection. Generalizable semantic segmentation as a new paradigm appears in recent research, which aims to generalize well on distinct unseen domains only using source domain data. The existing methods focus on learning domain-invariant features by using global distribution alignment strategies, which may lead to a decreased discriminability of the model. To cope with this challenge, we propose a fine-grained self-supervision (FGSS) framework for generalizable semantic segmentation that takes into account both discriminability and generalizability from the perspective of the intra-class relationship. The FGSS framework contains single-view and multi-view versions. In the single-view version, we propose a fine-grained self-supervision strategy to distinguish the sub-parts of the semantic class for better class discriminability. In the multi-view version, we propose a class prototype feature enhancement strategy to generate another view (i.e. another representation of the original representation). Then, we propose a multi-view mutual supervision loss to enforce consistency between different views and further enhance the generalizability of the model. Experimental results on five widely-used datasets, i.e., GTAV, SYNTHIA, BDD100K, Cityscapes, and Mapillary, demonstrate that our FGSS framework achieves superior performance compared to state-of-the-art methods.
Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Wenbin Zou, Chen Xu 0004
IEEE Trans. Circuits Syst. Video Technol.5
2024 Preserving Label-Related Domain-Specific Information for Cross-Domain Semantic Segmentation
abstract
Unsupervised domain adaptation semantic segmentation (UDASS) methods aim to learn domain-invariant information for alleviating the distribution shift problem between the source and target domains. However, ignoring the learning of domain-specific information that is label-related may limit the class discriminability on the target domain. We argue that a good representation for the UDASS task not only contains domain-invariant information but also preserves label-related domain-specific information. In this paper, a novel frequency spectrum domain adaptation approach via meta-learning (ML-FSDA) is proposed to achieve this goal for improving the class discriminability and generalization ability. ML-FSDA contains a frequency-spectrum meta-learning framework (FMF) and a class-aware domain-specific memory bank (CDMB). Specifically, first, inspired by the observation that the high-frequency component is consistent across different domains while the low-frequency component is much more domain-specific, the FMF aims to respectively learn label-related domain-specific and domain-invariant information from low-frequency and high-frequency images in a unified framework via the meta-learning strategy. Second, the CDMB is designed to preserve the label-related domain-specific information of each class in an external memory bank while the CDMB is updated in every iteration of the meta-training stage. Finally, the CDMB is utilized to embed the label-related domain-specific information into domain-invariant information at the class level during the meta-testing stage to enhance the class discriminability on the target domain. Extensive experiments demonstrate the effectiveness of ML-FSDA on two challenging cross-domain semantic segmentation benchmarks. Notably, for the GTA5 to Cityscapes task and the SYNTHIA to Cityscapes task, the proposed ML-FSDA achieves superior performance with 77.3% mIoU and 68.8% mIoU, respectively. The source code is released at https://github.com/seabearlmx/FSL.
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
IEEE Trans. Intell. Transp. Syst.5
2024 Calibration-Based Multi-Prototype Contrastive Learning for Domain Generalization Semantic Segmentation in Traffic Scenes
abstract
Prototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features for domain generalization semantic segmentation. These methods assume that the prototypes in different domains are invariant. However, the prototypes in different domains have discrepancies as well. First, the prototypes of the same class in different domains may be different. Second, the prototypes of different classes may be similar. To address these issues, a calibration-based multi-prototype contrastive learning (CMPCL) approach is proposed, which contains an uncertainty-guided multi-prototype contrastive learning (UMPCL) and a hard-weighted multi-prototype contrastive learning (HMPCL). Specifically, the UMPCL uses an uncertainty probability matrix, derived from element-wise discrepancies between the prototypes of the same class, to calibrate the weights of prototypes for alleviating the discrepancy between the prototypes of the same class in different domains. The HMPCL uses a hard-weighted matrix that is generated by the similarity between the prototypes of different classes, to calibrate the weights of the hard-aligned prototypes for alleviating the issue of similar prototypes between different classes, with hard-aligned prototypes referring to those exhibiting such similarity. Furthermore, since the learned class-wise domain-invariant features may overfit the prototype in the source domain, multi-prototype contrastive learning is used in the UMPCL and HMPCL to avoid this risk. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on multiple benchmarks of domain generalization semantic segmentation. The source code has been released onhttps://github.com/seabearlmx/CMPCL.
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
IEEE Trans. Intell. Transp. Syst.5
2024 "Where Does the Devil Lie?": Multimodal Multitask Collaborative Revision Network for Trusted Road Segmentation
abstract
Road segmentation is an essential component of navigation systems. Although recent advancements in road segmentation, the occurrence of failure segmentations remains inevitable. For safety-critical tasks, e.g., navigation, knowing when and where road segmentation fails is crucial. In this paper, we propose a novel trusted road segmentation architecture, namely Multimodal Multitask Collaborative Revision Network (M2CRN), to improve the trust of road segmentation. Our approach incorporates two strategies to predict and rectify segmentation errors. Firstly, a joint learning framework is devised to generate road segmentation results while estimating failure segmentation masks. Secondly, the road segmentation branch is equipped with an Uncertainty-Aware Revision Module (UARM), which eliminates the error in road segmentation. Additionally, we suppress the response of error regions in the road segmentation branch with an innovative design, called Adaptive Soft Error Suppression (ASES). To validate our methods, extensive experiments are conducted on three benchmark road segmentation datasets. The results demonstrate significant performance improvements with a real-time inference speed of 33.3 FPS, reaffirming the soundness of our revision model.
Guoguang Hua, Dalian Zheng, Shishun Tian, Wenbin Zou, Shenglan Liu 0001, Xia Li 0006
IEEE Trans. Multim.4
2023 Joint Edge-Guided and Spectral Transformation Network for Self-supervised X-Ray Image Restoration
Shasha Huang, Wenbin Zou, Hongxia Gao, Weipeng Yang 0002, Shicheng Niu, Tian Qi, Jianliang Ma
ICANN (2)2
2023 TRG-DQA: Texture Residual-Guided Dehazed Image Quality Assessment
abstract
Image dehazing algorithms have emerged to solve the visual impairment caused by haze. It is important to establish dehazed image quality assessment (DQA) methods that can accurately evaluate the dehazed image quality and the performance of dehazing algorithms. However, classical image quality assessment (IQA) and most hand-crafted feature based DQA methods may not be able to adequately measure complex distortions of dehazed images. To address this issue, this paper proposes a Texture Residual-Guided Dehazed image Quality Assessment (TRG-DQA) method. Specifically, we first introduce a global and local feature extraction module employing a combination of the Transformer and convolutional neural networks (CNN) for extracting the comprehensive features. Considering that texture residual maps represent haze density and artifact distortion information, we propose a residual-guided module to guide the model for efficient learning. Additionally, to mitigate the information loss issue that occurs in deeper networks, a distortion-aware feature enhancement module is proposed. Extensive experiments on six DQA databases demonstrate the proposed TRG-DQA achieves superior performance among all the state-of-the-art methods.
Tiantian Zeng, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Shishun Tian
ICIP3
2023 Blind Quality Assessment of Light Field Image Based on Spatio-Angular Textural Variation
abstract
Light Field Image Quality Assessment (LF-IQA) is vitally important to facilitate the development of immersive technologies. However, current state-of-the-art LF-IQA metrics still struggle to handle Light Field Image (LFI) with massive data in an efficient manner. To cope with this challenge, we propose a simple yet effective Blind LF-IQA metric based on Spatio-Angular Textural Variation, named SATV-BLiF. Given a distorted LFI, we first apply Local Binary Pattern (LBP) operator to measure the textural variation in the spatial and angular domains respectively. Then the generated spatial and angular textural matrices are merged and further transformed into statistical textural histogram features. Finally, Support Vector Regression (SVR) is employed to construct a nonlinear mapping function between the statistical textural histogram features and the perceptual quality score of the distorted LFI. Experimental results on three representative light field databases show that the proposed metric achieves state-of-the-art quality evaluation performance, while having much lower complexity than the existing No-Reference (NR) LF-IQA metrics. The code of the proposed SATV-BLiF metric is available at https://github.com/ZhengyuZhang96/SATV-BLiF.
Shishun Tian, Wenbin Zou, Yuhang Zhang 0011, Luce Morin, Lu Zhang 0037
ICIP3
2023 Need a dog for seeing eye? A Walk Viewpoint Dataset for Freespace Detection in Unstructured Environments
abstract
Freespace Detection (FD) is crucial for robust and safe autonomous navigation. However, existing datasets usually concentrate on structure road environments. The FD in unstructured environments, e.g., walk assistance for the visually-impaired, has been rarely investigated. In this paper, We propose a novel dataset called the Walk Viewpoint Dataset (WVD). Different from the previous datasets, we focus on the walk viewpoint, where FD can provide the potential for improving the walking of visually impaired people. The target regions of WVD are annotated with 20 categories by fine-grained labels, which consist of 3,737 images and depth images. Moreover, we propose a new annotation hierarchy, which allows different degrees of complexity and creates opportunities for new training methods. Finally, our study provides the statistical analysis of label characteristics and baseline analysis, which demonstrates its distinction compared to previous datasets. The dataset can be accessed through the project pages: http://www.sensingAI.com.cn.
Wenbin Zou, Guoguang Hua, Guangxu Chen, Zaiyue He, Guangli Liu, Huakun Li, Shishun Tian
ICME1
2023 Calibration-based Dual Prototypical Contrastive Learning Approach for Domain Generalization Semantic Segmentation
abstract
Prototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features recently. These methods are based on the assumption that the prototypes, which are represented as the central value of the same class in a certain domain, are domain-invariant. Since the prototypes of different domains have discrepancies as well, the class-wise domain-invariant features learned from the source domain by PCL need to be aligned with the prototypes of other domains simultaneously. However, the prototypes of the same class in different domains may be different while the prototypes of different classes may be similar, which may affect the learning of class-wise domain-invariant features. Based on these observations, a calibration-based dual prototypical contrastive learning (CDPCL) approach is proposed to reduce the domain discrepancy between the learned class-wise features and the prototypes of different domains for domain generalization semantic segmentation. It contains an uncertainty-guided PCL (UPCL) and a hard-weighted PCL (HPCL). Since the domain discrepancies of the prototypes of different classes may be different, we propose an uncertainty probability matrix to represent the domain discrepancies of the prototypes of all the classes. The UPCL estimates the uncertainty probability matrix to calibrate the weights of the prototypes during the PCL. Moreover, considering that the prototypes of different classes may be similar in some circumstances, which means these prototypes are hard-aligned, the HPCL is proposed to generate a hard-weighted matrix to calibrate the weights of the hard-aligned prototypes during the PCL. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on domain generalization segmentation tasks. The source code will be released at https://github.com/seabearlmx/CDPCL.
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
ACM Multimedia5
2023 Sequential Affinity Learning for Video Restoration
abstract
Video restoration networks aim to restore high-quality frame sequences from degraded ones. However, traditional video restoration methods heavily rely on temporal modeling operators or optical flow estimation, which limits their versatility. The aim of this work is to present a novel approach for video restoration that eliminates inefficient temporal modeling operators and pixel-level feature alignment in the network architecture. The proposed method, Sequential Affinity Learning Network (SALN), is designed based on an affinity mechanism that establishes direct correspondences between the Query frame, degraded sequence, and restored frames in latent space. This unique perspective allows for more accurate and effective restoration of video content without relying on temporal modeling operators or optical flow estimation techniques. Moreover, we enhanced the design of the channel-wise self-attention block to improve the decoder's performance for video restoration. Our method outperformed previous state-of-the-art methods by a significant margin in several classic video tasks, including video deraining, video dehazing, and video waterdrop removal, demonstrating excellent efficiency. As a novel network that differs significantly from previous video restoration methods, SALN aims to provide innovative ideas and directions for video restoration. Our contributions include proposing a novel affinity-based approach for video restoration, enhancing the design of the channel-wise self-attention block, and achieving state-of-the-art performance on several classic video tasks.
Tian Ye 0001, Sixiang Chen, Yun Liu 0002, Wenhao Chai, Jinbin Bai, Wenbin Zou, Yunchen Zhang, Mingchao Jiang, Erkang Chen, Chenghao Xue
ACM Multimedia6
2023 Channel Affinity Knowledge Distillation for Semantic Segmentation
abstract
In recent years, convolutional neural networks have achieved significant success in computer vision tasks. However, the deployment of these algorithms remains challenging. Knowledge distillation (KD) as a type of important method enables a tiny model to extract helpful information from a large model. Most existing KD methods based semantic segmentation aim to align predicted maps in the spatial domain, but channel distillation also may help to improve segmentation performance. Additionally, pairwise pixel affinity provides efficiently structured reasoning for semantic segmentation. Motivated by these considerations, we propose a novel Channel Affinity KD (CAKD) framework for semantic segmentation that focuses on channel and cross-channel affinity relationship distillation to better align the distribution of the student and teacher models. Extensive experiments demonstrate that our proposed approach outperforms state-of-the-art KD methods on Cityscapes, Pascal VOC, and ADE20k datasets.
Huakun Li, Yuhang Zhang 0011, Shishun Tian, Rong You, Wenbin Zou
MMSP6
2023 SPNet: An RGB-D Sequence Progressive Network for Road Semantic Segmentation
abstract
Road semantic segmentation is an essential component of autonomous driving and blind navigation. Although many excellent RGB-based road semantic segmentation algorithms have been proposed, these methods may not detect correctly due to the lack of geometric information. Recently, RGB-D road semantic segmentation methods attract more research attention. However, the existing RGB-D methods ignore the impact of unknown noise in sensors. To solve this problem, we propose an RGB-D Sequence Progressive Network (SPNet) for road semantic segmentation. Specifically, we first propose a sequence-based RGB-D feature extractor to alleviate the effect of noise. Then, We propose a multi-modal feature fusion (MMFF) module to enhance the feature representation of multi-modal data by further alleviating the effect of noise. Finally, we propose a semantic flow prediction (SFP) module that aims to align the multi-modal features in the decoder. Extensive experiments are conducted on several challenging datasets, including KITTI and GMRP. Our method achieves an F-score of 97.21% on the KITTI official leaderboard and ranked third in the official leaderboard.
Yuhang Zhang 0011, Guoguang Hua, Ruijing Long, Shishun Tian, Wenbin Zou
MMSP6
2023 Joint Priors-Based Restoration Method for Degraded Images Under Medium Propagation
Wenbin Zou, Hongxia Gao, Weipeng Yang 0002, Shasha Huang, Jianliang Ma
PRCV (11)2
2023 Enhancing Low-Light Images: A Variation-based Retinex with Modified Bilateral Total Variation and Tensor Sparse Coding
abstract
Abstract Low‐light conditions often result in the presence of significant noise and artifacts in captured images, which can be further exacerbated during the image enhancement process, leading to a decrease in visual quality. This paper aims to present an effective low‐light image enhancement model based on the variation Retinex model that successfully suppresses noise and artifacts while preserving image details. To achieve this, we propose a modified Bilateral Total Variation to better smooth out fine textures in the illuminance component while maintaining weak structures. Additionally, tensor sparse coding is employed as a regularization term to remove noise and artifacts from the reflectance component. Experimental results on extensive and challenging datasets demonstrate the effectiveness of the proposed method, exhibiting superior or comparable performance compared to state‐of‐the‐art approaches. Code, dataset and experimental results are available at https://github.com/YangWeipengscut/BTRetinex .
Weipeng Yang 0002, Hongxia Gao, Wenbin Zou, Shasha Huang, Jianliang Ma
Comput. Graph. Forum3
2023 Domain-invariant information aggregation for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006
Neurocomputing5
2023 A hybrid domain learning framework for unsupervised semantic segmentation
Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Wenbin Zou, Chen Xu 0004
Neurocomputing4
2023 Learning Shape-Invariant Representation for Generalizable Semantic Segmentation
abstract
Semantic segmentation assigns a category for each pixel and has achieved great success in a supervised manner. However, it fails to generalize well in new domains due to the domain gap. Domain adaptation is a popular way to solve this issue, but it needs target data and cannot handle unavailable domains. In domain generalization (DG), the model is trained without the target data and DG aims to generalize well in new unavailable domains. Recent works reveal that shape recognition is beneficial for generalization but still lack exploration in semantic segmentation. Meanwhile, the object shapes also exist a discrepancy in different domains, which is often ignored by the existing works. Thus, we propose a Shape-Invariant Learning (SIL) framework to focus on learning shape-invariant representation for better generalization. Specifically, we first define the structural edge, which considers both the object boundary and the inner structure of the object to provide more discrimination cues. Then, a shape perception learning strategy including a texture feature discrepancy reduction loss and a structural feature discrepancy enlargement loss is proposed to enhance the shape perception ability of the model by embedding the structural edge as a shape prior. Finally, we use shape deformation augmentation to generate samples with the same content and different shapes. Essentially, our SIL framework performs implicit shape distribution alignment at the domain-level to learn shape-invariant representation. Extensive experiments show that our SIL framework achieves state-of-the-art performance.
Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Guoguang Hua, Wenbin Zou, Chen Xu 0004
IEEE Trans. Image Process.5
2023 EDDMF: An Efficient Deep Discrepancy Measuring Framework for Full-Reference Light Field Image Quality Assessment
abstract
The increasing demand for immersive experience has greatly promoted the quality assessment research of Light Field Image (LFI). In this paper, we propose an efficient deep discrepancy measuring framework for full-reference light field image quality assessment. The main idea of the proposed framework is to efficiently evaluate the quality degradation of distorted LFIs by measuring the discrepancy between reference and distorted LFI patches. Firstly, a patch generation module is proposed to extract spatio-angular patches and sub-aperture patches from LFIs, which greatly reduces the computational cost. Then, we design a hierarchical discrepancy network based on convolutional neural networks to extract the hierarchical discrepancy features between reference and distorted spatio-angular patches. Besides, the local discrepancy features between reference and distorted sub-aperture patches are extracted as complementary features. After that, the angular-dominant hierarchical discrepancy features and the spatial-dominant local discrepancy features are combined to evaluate the patch quality. Finally, the quality of all patches is pooled to obtain the overall quality of distorted LFIs. To the best of our knowledge, the proposed framework is the first patch-based full-reference light field image quality assessment metric based on deep-learning technology. Experimental results on four representative LFI datasets show that our proposed framework achieves superior performance as well as lower computational complexity compared to other state-of-the-art metrics.
Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037
IEEE Trans. Image Process.3
2023 Multiple Relational Learning Network for Joint Referring Expression Comprehension and Segmentation
abstract
Multi-task learning is a successful learning framework which improves the performance of prediction models by leveraging knowledge among related tasks. Referring expression comprehension (REC) and segmentation (RES) are highly relevant tasks, which both are language-guided visual recognition tasks. However, their relations have not yet been fully exploited in previous works. In this paper, a Multiple Relational Learning Network (MRLN) is proposed for multi-task learning of REC and RES. First, a feature-feature interaction learning module is introduced to handle the complicated interactions among features. Moreover, we propose a feature-task dependence learning module, which associates the related features with target tasks. Furthermore, a task-task relationship learning module is designed, which captures the relationships among tasks automatically and guides the REC and RES fine-tuning adaptively. To verify our proposed approach, experiments are conducted on three benchmark datasets, i.e., RefCOCO, RefCOCO+, and RefCOCOg. Extensive experiments demonstrate that the multiple relationships are more appealing since it alleviates the prediction inconsistency issue in multi-task setup. In addition, the experimental results report the significant performance gains of MRLN over most existing methods, i.e., up to 83.46 % for REC and 63.62 % for RES over state-of-the-art methods, which demonstrate the validity and superiority of MRLN.
Guoguang Hua, Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Wenbin Zou
IEEE Trans. Multim.5
2023 Joint Wavelet Sub-Bands Guided Network for Single Image Super-Resolution
abstract
Since deep convolutional neural network (CNN) has achieved excellent results in single image super-resolution (SISR), an increasing number of methods based on CNN have been proposed. Most CNN-based methods are devoted to finding mapping based on pixel intensity while ignoring the importance of frequency information, which can reflect semantic information of images on different bands. This leads to less effectiveness in the reconstruction of high-frequency details. To address this problem, we propose a novel CNN-based super-resolution method named joint wavelet sub-bands guided network (JWSGN). We separate the different frequency information of the image by the WT and then recover this information by a multi-branch network. To recover finer edge details, we propose an edge extraction module, which estimates an edge feature map by using the similarity of all high-frequency sub-bands and then corrects the high-frequency features recovered from each branch by exploiting the edge feature map. Furthermore, we use the complementary relationship between different frequencies to calibrate the high-frequency sub-bands. Finally, the high-resolution image is obtained by inverse wavelet transform. Both qualitative and quantitative experiments show that our method performs excellent performance with the guidance of the edge extraction module.
Wenbin Zou, Liang Chen 0026, Yi Wu 0010, Yunchen Zhang, Yuxiang Xu
IEEE Trans. Multim.1
2022 Deeblif: Deep Blind Light Field Image Quality Assessment by Extracting Angular and Spatial Information
abstract
In the era of immersive media, the high-dimensional Light Field Image (LFI) puts forward higher requirements for Light Field Image Quality Assessment (LF-IQA). However, currently most existing LF-IQA metrics still rely on sophisticated hand-crafted feature extraction, which fail to predict the quality of LFI accurately. In this paper, we propose a patch-based Deep Blind Light Field image quality assessment metric (abbreviated as DeeBLiF), by employing a two-stream Convolutional Neural Network (CNN) model specifically designed for extracting the angular and spatial information of LFI. Firstly, the spatio-angular patches are generated as input data, which effectively reflect the spatio-angular information of LFI. After that, a two-stream CNN model is exploited to extract the patch features and further predict the patch scores. Finally, all the patch scores are pooled into an overall quality score of LFI. Experimental results on the LFI dataset demonstrate that the proposed DeeBLiF outperforms the state-of-the-art LF-IQA metrics. The code will be publicly available at https://github.com/ZhengyuZhang96/DeeBLiF.
Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037
ICIP3
2022 Exploring more concentrated and consistent activation regions for cross-domain semantic segmentation
Muxin Liao, Guoguang Hua, Shishun Tian, Yuhang Zhang 0011, Wenbin Zou, Xia Li 0006
Neurocomputing5
2022 RGB-D Gate-guided edge distillation for indoor semantic segmentation
Wenbin Zou, Yingqing Peng, Shishun Tian, Xia Li 0006
Multim. Tools Appl.1
2021 Temporal Pyramid Network for Pedestrian Trajectory Prediction with Multi-Supervision
abstract
Predicting human motion behavior in a crowd is important for many applications, ranging from the natural navigation of autonomous vehicles to intelligent security systems of video surveillance. All the previous works model and predict the trajectory with a single resolution, which is relatively ineffective and difficult to simultaneously exploit the long-range information (e.g., the destination of the trajectory), and the short-range information (e.g., the walking direction and speed at a certain time) of the motion behavior. In this paper, we propose a temporal pyramid network for pedestrian trajectory prediction through a squeeze modulation and a dilation modulation. Our hierarchical framework builds a feature pyramid with increasingly richer temporal information from top to bottom, which can better capture the motion behavior at various tempos. Furthermore, we propose a coarse-to-fine fusion strategy with multi-supervision. By progressively merging the top coarse features of global context to the bottom fine features of rich local context, our method can fully exploit both the long-range and short-range information of the trajectory. Experimental results on two benchmarks demonstrate the superiority of our method. Our code and models will be available upon acceptance.
Rongqin Liang, Yuanman Li, Xia Li 0006, Yi Tang 0008, Jiantao Zhou 0001, Wenbin Zou
AAAI6
2021 Quality assessment of DIBR-synthesized views: An overview
Shishun Tian, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Ting Su 0004, Luce Morin, Olivier Déforges
Neurocomputing3
2021 CIMask: Segmenting instances by class-specific semantic feature extraction and instance-specific attribute discrimination
Canqun Xiang, Wenbin Zou, Chen Xu 0004
Neurocomputing2
2021 STA3D: Spatiotemporally attentive 3D network for video saliency prediction
Wenbin Zou, Shengkai Zhuo, Yi Tang 0008, Shishun Tian, Xia Li 0006, Chen Xu 0004
Pattern Recognit. Lett.1
2021 Dual-Stream Multi-Path Recursive Residual Network for JPEG Image Compression Artifacts Reduction
abstract
JPEG is the most widely used lossy image compression standard. When using JPEG with high compression ratios, visual artifacts cannot be avoided. These artifacts not only degrade the user experience but also negatively affect many low-level image processing tasks. Recently, convolutional neural network (CNN)-based compression artifact removal approaches have achieved significant success, however, at the cost of high computational complexity due to an enormous number of parameters. To address this issue, we propose a dual-stream recursive residual network (STRRN) which consists of structure and texture streams for separately reducing the specific artifacts related to high-frequency or low-frequency image components. The outputs of these streams are combined and fed into an aggregation network to further enhance the restored images. By using parameter sharing, the proposed network reduces the total number of training parameters significantly. Moreover, experiments conducted on five commonly used datasets confirm that the proposed STRRN can efficiently reduce the compression artifacts, while using up to 4.6 times less training parameters and 5 times less running time compared to the state-of-the-art approaches.
Zhi Jin 0002, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.3
2021 SC-RPN: A Strong Correlation Learning Framework for Region Proposal
abstract
Current state-of-the-art two-stage detectors heavily rely on region proposals to guide the accurate detection for objects. In previous region proposal approaches, the interaction between different functional modules is correlated weakly, which limits or decreases the performance of region proposal approaches. In this paper, we propose a novel two-stage strong correlation learning framework, abbreviated as SC-RPN, which aims to set up stronger relationship among different modules in the region proposal task. Firstly, we propose a Light-weight IoU-Mask branch to predict intersection-over-union (IoU) mask and refine region classification scores as well, it is used to prevent high-quality region proposals from being filtered. Furthermore, a sampling strategy named Size-Aware Dynamic Sampling (SADS) is proposed to ensure sampling consistency between different stages. In addition, point-based representation is exploited to generate region proposals with stronger fitting ability. Without bells and whistles, SC-RPN achieves AR100014.5% higher than that of Region Proposal Network (RPN), surpassing all the existing region proposal approaches. We also integrate SC-RPN into Fast R-CNN and Faster R-CNN to test its effectiveness on object detection task, the experimental results achieve a gain of 3.2% and 3.8% in terms of mAP compared to the original ones.
Wenbin Zou, Yingqing Peng, Canqun Xiang, Shishun Tian, Lu Zhang 0037
IEEE Trans. Image Process.1
2020 MMA Regularization: Decorrelating Weights of Neural Networks by Maximizing the Minimal Angles
abstract
The strong correlation between neurons or filters can significantly weaken the generalization ability of neural networks. Inspired by the well-known Tammes problem, we propose a novel diversity regularization method to address this issue, which makes the normalized weight vectors of neurons or filters distributed on a hypersphere as uniformly as possible, through maximizing the minimal pairwise angles (MMA). This method can easily exert its effect by plugging the MMA regularization term into the loss function with negligible computational overhead. The MMA regularization is simple, efficient, and effective. Therefore, it can be used as a basic regularization method in neural network training. Extensive experiments demonstrate that MMA regularization is able to enhance the generalization ability of various modern models and achieves considerable performance improvements on CIFAR100 and TinyImageNet datasets. In addition, experiments on face verification show that MMA regularization is also effective for feature learning. Code is available at: https://github.com/wznpub/MMA_Regularization.
Zhennan Wang 0001, Canqun Xiang, Wenbin Zou, Chen Xu 0004
NeurIPS3
2020 Video salient object detection via spatiotemporal attention neural networks
Yi Tang 0008, Wenbin Zou, Yang Hua 0001, Zhi Jin 0002, Xia Li 0006
Neurocomputing2
2020 Deep and joint learning of longitudinal data for Alzheimer's disease prediction
Bai Ying Lei, Mengya Yang, Peng Yang 0011, Feng Zhou 0003, Wen Hou, Wenbin Zou, Xia Li 0006, Tianfu Wang 0001, Xiaohua Xiao, Shuqiang Wang
Pattern Recognit.6
2020 Salient video object detection using a virtual border and guided filter
Lu Zhang 0037, Wenbin Zou, Kidiyo Kpalma
Pattern Recognit.3
2020 DMA Regularization: Enhancing Discriminability of Neural Networks by Decreasing the Minimal Angle
abstract
Most of the discriminative feature learning methods are specifically developed for metric learning, however, the effectiveness may be not obvious for other tasks. In this letter, we propose a novel discrimination regularization method for image classification, which enhances the intra-class compactness and inter-class discrepancy simultaneously, through decreasing the minimalangle (DMA) between the feature vector and any one of the weight vectors in classification layer. This method can robustly improve the discriminability and generalizability of neural networks and easily exert its effect by plugging the DMA regularization term into the loss function with negligible computational overhead. The DMA regularization is simple, efficient, and effective. Therefore, it can be used as a basic regularization method for models based on neural networks. We evaluate DMA by applying it to various modern models on CIFAR10, CIFAR100, and TinyImageNet datasets, decreasing the test error rate by 0.2-0.4%, 0.2-1.5%, and 0.3-0.4% respectively. Code is available at: https://github.com/wznpub/DMA_Regularization.
Zhennan Wang 0001, Canqun Xiang, Wenbin Zou, Chen Xu 0004
IEEE Signal Process. Lett.3
2020 Matrix Capsule Convolutional Projection for Deep Feature Learning
abstract
Capsule projection network (CapProNet) has shown its ability to obtain semantic information, and spatial structural information from the raw images. However, the vector capsule of CapProNet has limitations in representing semantic information due to ignoring local information. Besides, the number of trainable parameters also increases greatly with the dimension of the feature vector. To that end, we propose a matrix capsule convolution projection (MCCP) module by replacing the feature vector with a feature matrix, of which each column represents a local feature. The feature matrix is then convoluted by columns into capsule subspaces to decrease the number of trainable parameters effectively. Furthermore, the CapDetNet is designed to explore the structural information encoding of the MCCP module based on object detection task. Experimental results demonstrate that the proposed MCCP outperforms the baselines in image classification, and CapDetNet achieves the 2.3% performance gain in object detection.
Canqun Xiang, Zhennan Wang 0001, Shishun Tian, Jianxin Liao, Wenbin Zou, Chen Xu 0004
IEEE Signal Process. Lett.5
2020 A Flexible Deep CNN Framework for Image Restoration
abstract
Image restoration is a long-standing problem in image processing and low-level computer vision. Recently, discriminative convolutional neural network (CNN)-based approaches have attracted considerable attention due to their superior performance. However, most of these frameworks are designed for one specific image restoration task; hence, they seldom show high performance on other image restoration tasks. To address this issue, we propose a flexible deep CNN framework that exploits the frequency characteristics of different types of artifacts. Hence, the same approach can be employed for a variety of image restoration tasks by adjusting the architecture. For reducing the artifacts with similar frequency characteristics, a quality enhancement network that adopts residual and recursive learning is proposed. Residual learning is utilized to speed up the training process and boost the performance; recursive learning is adopted to significantly reduce the number of training parameters as well as boost the performance. Moreover, lateral connections transmit the extracted features between different frequency streams via multiple paths. One aggregation network combines the outputs of these streams to further enhance the restored images. We demonstrate the capabilities of the proposed framework with three representative applications: image compression artifacts reduction (CAR), image denoising, and single image super-resolution (SISR). Extensive experiments confirm that the proposed framework outperforms the state-of-the-art approaches on benchmark datasets for these applications.
Zhi Jin 0002, Dmytro Bobkov, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Multim.4
2019 PR Product: A Substitute for Inner Product in Neural Networks
abstract
In this paper, we analyze the inner product of weight vector w and data vector x in neural networks from the perspective of vector orthogonal decomposition and prove that the direction gradient of w decreases with the angle between them close to 0 or π. We propose the Projection and Rejection Product (PR Product) to make the direction gradient of w independent of the angle and consistently larger than the one in standard inner product while keeping the forward propagation identical. As a reliable substitute for standard inner product, the PR Product can be applied into many existing deep learning modules, so we develop the PR Product version of fully connected layer, convolutional layer and LSTM layer. In static image classification, the experiments on CIFAR10 and CIFAR100 datasets demonstrate that the PR Product can robustly enhance the ability of various state-of-the-art classification networks. On the task of image captioning, even without any bells and whistles, our PR Product version of captioning model can compete or outperform the state-of-the-art models on MS COCO dataset. Code has been made available at: https://github.com/wzn0828/PR_Product.
Zhennan Wang 0001, Wenbin Zou, Chen Xu 0004
ICCV2
2019 An Efficient Quality Enhancement Solution for Stereo Images
Yingqing Peng, Zhi Jin 0002, Wenbin Zou, Yi Tang 0008, Xia Li 0006
ICIG (3)3
2019 Robust Plane Detection Using Depth Information From a Consumer Depth Camera
abstract
The emerging of depth-camera technology is paving the way for a variety of new applications and it is believed that plane detection is one of them. In fact, planes are common in man-made living structures, thus their accurate detection can benefit many visual-based applications. The use of depth information allows detecting planes characterized by complex pattern and texture, where the texture-based plane detection algorithms usually fail. In this paper, we propose a robust depth-driven plane detection (DPD) algorithm which consists of two parts: the growing-based plane detection and a two-stage refinement. The proposed approach starts from the seed patch with the highest planarity and uses the estimated equation of the growing plane and a dynamic threshold function to steer the growing process. Aided with this mechanism, each seed patch can grow to its maximum extent, and then the next seed patch starts to grow. This process is iteratively repeated so as to detect all the planes. Moreover, the refinement is proposed to tackle two common problems suffered by growing-based approaches, the over-growing problem, and the under-growing problem. Validated by extensive experiments, the proposed DPD algorithm is able to accurately detect planes and robust to various testing conditions. In terms of applications, it can be used as the pre-processing step for a variety of applications, such as, planar object recognition, super-resolution of the time-of-flight depth images with intrinsically low resolution.
Zhi Jin 0002, Tammam Tillo, Wenbin Zou, Yao Zhao 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.3
2019 Weakly Supervised Salient Object Detection With Spatiotemporal Cascade Neural Networks
abstract
Recently, deep learning techniques have substantially boosted the performance of salient object detection in still images. However, the salient object detection in videos by using traditional handcrafted features or deep learning features is not fully investigated, probably due to the lack of sufficient manually labeled video data for saliency modeling, especially for the data-driven deep learning. This paper proposes a novel weakly supervised approach to the salient object detection in a video, which can learn a robust saliency prediction model by using very limited manually labeled data and a large amount of weakly labeled data that could be easily generated in a supervised approach. Furthermore, we propose a spatiotemporal cascade neural network architecture for saliency modeling, in which two fully convolutional networks are cascaded to evaluate the visual saliency from both spatial and temporal cues to lead the optimal video saliency prediction. The proposed approach is extensively evaluated on the widely used challenging data sets, and the experiments demonstrate that our proposed approach substantially outperforms the state-of-the-art salient object detection models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Yuhuan Chen, Yang Hua 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.2
2019 An End-to-End Deep Learning Histochemical Scoring System for Breast Cancer TMA
abstract
One of the methods for stratifying different molecular classes of breast cancer is the Nottingham prognostic index plus, which uses breast cancer relevant biomarkers to stain tumor tissues prepared on tissue microarray (TMA). To determine the molecular class of the tumor, pathologists will have to manually mark the nuclei activity biomarkers through a microscope and use a semi-quantitative assessment method to assign a histochemical score (H-Score) to each TMA core. Manually marking positively stained nuclei is a time-consuming, imprecise, and subjective process, which will lead to inter-observer and intra-observer discrepancies. In this paper, we present an end-to-end deep learning system, which directly predicts the H-Score automatically. Our system imitates the pathologists' decision process and uses one fully convolutional network (FCN) to extract all nuclei region (tumor and non-tumor), a second FCN to extract tumor nuclei region, and a multi-column convolutional neural network, which takes the outputs of the first two FCNs and the stain intensity description image as an input and acts as the high-level decision making mechanism to directly output the H-Score of the input TMA image. To the best of our knowledge, this is the first end-to-end system that takes a TMA image as the input and directly outputs a clinical score. We will present experimental results, which demonstrate that the H-Scores predicted by our model have very high and statistically significant correlation with experienced pathologists' scores and that the H-Score discrepancy between our algorithm and the pathologists is on par with the inter-subject discrepancy between the pathologists.
Jingxin Liu 0005, Bolei Xu, Chi Zheng, Yuanhao Gong, Jonathan M. Garibaldi, Daniele Soria, Andrew R. Green, Ian O. Ellis, Wenbin Zou, Guoping Qiu
IEEE Trans. Medical Imaging9
2018 Video Salient Object Detection via Multiple Time-scale Analysis
abstract
This paper focuses on salient object detection in video by multiple time-scale analysis, which exploits the temporally consistent information under three different scales. In the first time-scale, we define an effective measure called motion contrast from both low-level cues and the optical flow fields. In the second time-scale, we propose a novel approach to repair the inaccurate motion contrast due to the mistake of optical flow. In the third time-scale, considering the low-contrast objects that stop moving for a certain amount of time and cannot remain prominent, we present a robust motion detection method based on point-tracking and trajectories clustering. Finally, the outcomes from the three time-scales jointly formulate the saliency detection by Bayesian inference. The proposed model is evaluated on the widely-used DAVIS and FBMS benchmark. Experiments demonstrate that our proposed model substantially outperforms the state-of-the-art saliency detection models.
Yuhuan Chen, Limin Huang, Wenbin Zou, Xia Li 0006, Guoping Qiu
ICPR3
2018 Multi-Scale Spatiotemporal Conv-LSTM Network for Video Saliency Detection
abstract
Recently, deep neural networks have been crucial techniques for image salient detection. However, two difficulties prevent the development of deep learning in video saliency detection. The first one is that the traditional static network cannot conduct a robust motion estimation in videos. The other is that the data-driven deep learning is in lack of sufficient manually annotated pixel-wise ground truths for video saliency network training. In this paper, we propose a multi-scale spatiotemporal convolutional LSTM network (MSST-ConvLSTM) to incorporate spatial and temporal cues for video salient objects detection. Furthermore, as manually pixel-wised labeling is very time-consuming, we sign lots of coarse labels, which are mixed with fine labels to train a robust saliency prediction model. Experiments on the widely used challenging benchmark datasets (e.g., FBMS and DAVIS) demonstrate that the proposed approach has competitive performance of video saliency detection compared with the state-of-the-art saliency models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Xia Li 0006
ICMR2
2018 Singular value decomposition based virtual representation for face recognition
Guiying Zhang, Wenbin Zou, Xianjie Zhang, Yong Zhao 0010
Multim. Tools Appl.2
2018 MS-CapsNet: A Novel Multi-Scale Capsule Network
abstract
Capsule network is a novel architecture to encode the properties and spatial relationships of the feature in an image, which shows encouraging results on image classification. However, the original capsule network is not suitable for some classification tasks, where the target objects are complex internal representations. Hence, we propose a multi-scale capsule network that is more robust and efficient for feature representation in image classification. The proposed multi-scale capsule network consists of two stages. In the first stage, structural and semantic information are obtained by multi-scale feature extraction. In the second stage, the hierarchy of features is encoded to multi-dimensional primary capsules. Moreover, we propose an improved dropout to enhance the robustness of the capsule network. Experimental results show that our method has a competitive performance on FashionMNIST and CIFAR10 datasets.
Canqun Xiang, Lu Zhang 0037, Yi Tang 0008, Wenbin Zou, Chen Xu 0004
IEEE Signal Process. Lett.4
2018 SCOM: Spatiotemporal Constrained Optimization for Salient Object Detection
abstract
This paper presents a novel model for video salient object detection called spatiotemporal constrained optimization model (SCOM), which exploits spatial and temporal cues, as well as a local constraint, to achieve a global saliency optimization. For a robust motion estimation of salient objects, we propose a novel approach to modeling the motion cues from optical flow field, the saliency map of the prior video frame and the motion history of change detection, which is able to distinguish the moving salient objects from diverse changing background regions. Furthermore, an effective objectness measure is proposed with intuitive geometrical interpretation to extract some reliable object and background regions, which provided as the basis to define the foreground potential, background potential, and the constraint to support saliency propagation. These potentials and the constraint are formulated into the proposed SCOM framework to generate an optimal saliency map for each frame in a video. The proposed model is extensively evaluated on the widely used challenging benchmark data sets. Experiments demonstrate that our proposed SCOM substantially outperforms the state-of-the-art saliency models.
Yuhuan Chen, Wenbin Zou, Yi Tang 0008, Xia Li 0006, Chen Xu 0004, Nikos Komodakis
IEEE Trans. Image Process.2
2017 Focus prior estimation for salient object detection
abstract
In the past five years, salient object detection has become one of the hot topics in the field of computer vision. Focus is a naturally strong indicator for the salient object detection task, but is not well studied. In this paper, a novel method is proposed to estimate the focus prior map for an arbitrary image. Different from the current edge density estimation based methods, the proposed method is based on the sparse defocus dictionary learning on a newly designed dataset. The focus strength is measured by the number of non-zero coefficients of the dictionary atoms. Objectness proposal method is introduced to improve the performance. Comparison with the other focusness estimation methods, the proposed focus prior map is more accurate and easier to be integrated by the other salient object detection methods. Experiments have confirmed the effectiveness and importance of the proposed focus prior.
Wenbin Zou, Chen Xu 0004
ICIP3
2017 Multi-modal metric learning for vehicle re-identification in traffic surveillance environment
abstract
Vehicle re-identification (Re-Id) aims to retrieve the same vehicle captured by disjoint cameras at different time instants from different locations, and is a challenging task mainly due to the high similarity among the captured vehicle images in surveillance environment. With the rapid development of Convolutional Neural Network (CNN), learning-based deep features have been adopted to combine with hand-crafted features to re-identify vehicles in traffic surveillance environment. However, the two kinds of features are in different feature space, and if they are fused directly together, their complementary correlation is not able to be fully explored. To address such an issue, this paper proposes a multi-modal metric learning architecture to fuse deep features and hand-crafted ones in an end-to-end optimization network, which achieves a more robust and discriminative feature representation for vehicle re-identification. The extensive experiments on a large-scale traffic surveillance vehicle dataset demonstrate that our proposed approach substantially outperforms the state-of-the-art methods on vehicle Re-Id.
Yi Tang 0008, Di Wu 0009, Zhi Jin 0002, Wenbin Zou, Xia Li 0006
ICIP4
2017 A CNN cascade for quality enhancement of compressed depth images
abstract
Transmitting depth images along with the corresponding textures enables a wide range of receiver-side 3D applications. Since each pixel on the depth images represents a corresponding 3D scene geometric information, when compressed during transmission the compression artifacts will lead to severe geometry distortions and visual perceptual degradation. To solve this problem, in this paper we proposed a convolutional neural network (CNN) cascade for suppressing the compression artifacts on depth images. According to the feature of depth images, we furthermore, adopt a weighted loss function for network training which can adaptively improve the learning efficiency and accuracy. Meanwhile, in order to overcome the limited training data problem, we audaciously trained our network on textures first and then finetune on the target depth images. To our best knowledge, few works have applied CNN on depth images targeting for compression artifacts reduction (CAR). Through extensive experiments, our proposed solution achieves higher quality for both reconstructed depth images and synthesized virtual views than the state-of-the-art methods.
Zhi Jin 0002, Lei Luo 0003, Yi Tang 0008, Wenbin Zou, Xia Li 0006
VCIP4
2017 Hierarchical Saliency Detection via Probabilistic Object Boundaries
abstract
Though there are many computational models proposed for saliency detection, few of them take object boundary information into account. This paper presents a hierarchical saliency detection model incorporating probabilistic object boundaries, which is based on the observation that salient objects are generally surrounded by explicit boundaries and show contrast with their surroundings. We perform adaptive thresholding operation on ultrametric contour map, which leads to hierarchical image segmentations, and compute the saliency map for each layer based on the proposed robust center bias, border bias, color dissimilarity and spatial coherence measures. After a linear weighted combination of multi-layer saliency maps, and Bayesian enhancement procedure, the final saliency map is obtained. Extensive experimental results on three challenging benchmark datasets demonstrate that the proposed model outperforms eight state-of-the-art saliency detection models.
Haijun Lei, Hai Xie, Wenbin Zou, Kidiyo Kpalma, Nikos Komodakis
Int. J. Pattern Recognit. Artif. Intell.3
2017 Diversity induced matrix decomposition model for salient object detection
Zhixiang He, Chen Xu 0004, Wenbin Zou, George Baciu
Pattern Recognit.5
2017 Enhanced Autofocusing in Optical Scanning Holography Based on Hologram Decomposition
abstract
Optical scanning holography is a compact and powerful method for capturing hologram of a wide three-dimensional (3-D) view scene. After a hologram has been taken, it is often necessary to determine the locations of the focal plane on which the objects are residing, so that the 3-D scene can be numerically reconstructed for further analysis or processing. Recent research has shown that automatic detection of the depth (focal plane) of objects represented in a hologram can be conducted with entropy minimization method. Despite the success, the method could fail if the entropy information of objects in a hologram are interfering with each other. In this paper, we propose a method based on the hologram decomposition to overcome this problem. Briefly, the hologram is decomposed into subholograms and the focal plane distance is determined separately for each subobject. Simulation results reveal that our proposed method has good accuracy and reliability.
Shuming Jiao, Peter Wai-Ming Tsang, Ting-Chung Poon, Jung-Ping Liu, Wenbin Zou, Xia Li 0006
IEEE Trans. Ind. Informatics5
2017 A Robust and Efficient Approach to License Plate Detection
abstract
This paper presents a robust and efficient method for license plate detection with the purpose of accurately localizing vehicle license plates from complex scenes in real time. A simple yet effective image downscaling method is first proposed to substantially accelerate license plate localization without sacrificing detection performance compared with that achieved using the original image. Furthermore, a novel line density filter approach is proposed to extract candidate regions, thereby significantly reducing the area to be analyzed for license plate localization. Moreover, a cascaded license plate classifier based on linear support vector machines using color saliency features is introduced to identify the true license plate from among the candidate regions. For performance evaluation, a data set consisting of 3977 images captured from diverse scenes under different conditions is also presented. Extensive experiments on the widely used Caltech license plate data set and our newly introduced data set demonstrate that the proposed approach substantially outperforms state-of-the-art methods in terms of both detection accuracy and run-time efficiency, increasing the detection ratio from 91.09% to 96.62% while decreasing the run time from 672 to 42 ms for processing an image with a resolution of 1082×728 . The executable code and our collected data set are publicly available.
Yule Yuan, Wenbin Zou, Yong Zhao 0010, Xin'an Wang, Xuefeng Hu, Nikos Komodakis
IEEE Trans. Image Process.2
2016 Saliency Detection via Diversity-Induced Multi-view Matrix Decomposition
Zhixiang He, Wenbin Zou, George Baciu
ACCV (1)4
2016 Background subtraction using dual-class backgrounds
abstract
This paper presents a novel approach to background subtraction which aims to extract moving objects in video stream. To this end, a novel background model is proposed by using both working backgrounds and candidate backgrounds, which can be transferred to each other according to an adaptive mechanism. The input image (video frame) is compared and evaluated with these dual-class backgrounds (DCB) to detect foreground objects. Furthermore, for robust background modeling a novel background updating scheme is proposed based on the life-value which represents the existing time of a background sample, and the access-time which represents the number of valid visits of a background sample. Experiments on a standard dataset demonstrated the effectiveness and robustness of the proposed approach by comparing it with the previous typical background subtraction techniques.
Bingshu Wang, Wenqian Zhu, Yong Zhao 0010, Wenbin Zou
ICARCV5
2015 HARF: Hierarchy-Associated Rich Features for Salient Object Detection
abstract
The state-of-the-art salient object detection models are able to perform well for relatively simple scenes, yet for more complex ones, they still have difficulties in highlighting salient objects completely from background, largely due to the lack of sufficiently robust features for saliency prediction. To address such an issue, this paper proposes a novel hierarchy-associated feature construction framework for salient object detection, which is based on integrating elementary features from multi-level regions in a hierarchy. Furthermore, multi-layered deep learning features are introduced and incorporated as elementary features into this framework through a compact integration scheme. This leads to a rich feature representation, which is able to represent the context of the whole object/background and is much more discriminative as well as robust for salient object detection. Extensive experiments on the most widely used and challenging benchmark datasets demonstrate that the proposed approach substantially outperforms the state-of-the-art on salient object detection.
Wenbin Zou, Nikos Komodakis
ICCV1
2015 Unsupervised Joint Salient Region Detection and Object Segmentation
abstract
This paper presents a novel unsupervised algorithm to detect salient regions and to segment out foreground objects from background. In contrast to previous unidirectional saliency-based object segmentation methods, in which only the detected saliency map is used to guide the object segmentation, our algorithm mutually exploits detection/segmentation cues from each other. To achieve this goal, an initial saliency map is generated by the proposed segmentation driven low-rank matrix recovery model. Such a saliency map is exploited to initialize object segmentation model, which is formulated as energy minimization of Markov random field. Mutually, the quality of saliency map is further improved by the segmentation result, and serves as a new guidance for the object segmentation. The optimal saliency map and the final segmentation are achieved by jointly optimizing the defined objective functions. Extensive evaluations on MSRA-B and PASCAL-1500 datasets demonstrate that the proposed algorithm achieves the state-of-the-art performance for both the salient region detection and the object segmentation.
Wenbin Zou, Zhi Liu 0003, Kidiyo Kpalma, Joseph Ronsin, Yong Zhao 0010, Nikos Komodakis
IEEE Trans. Image Process.1
2014 Co-saliency detection based on region-level fusion and pixel-level refinement
abstract
This paper addresses the problem of co-saliency detection, which aims to identify the common salient objects in a set of images and is important for many applications such as object co-segmentation and co-recognition. First, the segmentation driven low-rank matrix recovery model is used for intra saliency detection in each individual image of the image set, to highlight the regions whose features are sparse in each image. Then, a region-level fusion method, which exploits inter-region dissimilarities on color histograms and global consistency of regions over the image set, adjusts the intra saliency maps to obtain the region-level co-saliency maps, which can highlight co-salient object regions and suppress irrelevant regions. Finally, a pixel-level refinement method, which integrates color-spatial similarity between pixel and region with image border connectivity based object prior, generates the pixel-level co-saliency maps with better quality. Extensive experiments on two benchmark datasets demonstrate that the proposed co-saliency model consistently outperforms the state-of-the-art co-saliency models in both subjective and objective evaluation.
Zhi Liu 0003, Wenbin Zou, Xiang Zhang 0006, Olivier Le Meur
ICME3
2014 Saliency Tree: A Novel Saliency Detection Framework
abstract
This paper proposes a novel saliency detection framework termed as saliency tree. For effective saliency measurement, the original image is first simplified using adaptive color quantization and region segmentation to partition the image into a set of primitive regions. Then, three measures, i.e., global contrast, spatial sparsity, and object prior are integrated with regional similarities to generate the initial regional saliency for each primitive region. Next, a saliency-directed region merging approach with dynamic scale control scheme is proposed to generate the saliency tree, in which each leaf node represents a primitive region and each non-leaf node represents a non-primitive region generated during the region merging process. Finally, by exploiting a regional center-surround scheme based node selection criterion, a systematic saliency tree analysis including salient node selection, regional saliency adjustment and selection is performed to obtain final regional saliency measures and to derive the high-quality pixel-wise saliency map. Extensive experimental results on five datasets with pixel-wise ground truths demonstrate that the proposed saliency tree model consistently outperforms the state-of-the-art saliency models.
Zhi Liu 0003, Wenbin Zou, Olivier Le Meur
IEEE Trans. Image Process.2
2014 Online Glocal Transfer for Automatic Figure-Ground Segmentation
abstract
This paper addresses the problem of automatic figure-ground segmentation, which aims at automatically segmenting out all foreground objects from background. The underlying idea of this approach is to transfer segmentation masks of globally and locally (glocally) similar exemplars into the query image. For this purpose, we propose a novel high-level image representation method named as object-oriented descriptor. Using this descriptor, a set of exemplar images glocally similar to the query image is retrieved. Then, using over-segmented regions of these retrieved exemplars, a discriminative classifier is learned on-the-fly and subsequently used to predict foreground probability for the query image. Finally, the optimal segmentation is obtained by combining the online prediction with typical energy optimization of Markov random field. The proposed approach has been extensively evaluated on three datasets, including Pascal VOC 2010, VOC 2011 segmentation challenges, and iCoseg dataset. Experiments show that the proposed approach outperforms state-of-the-art methods and has the potential to segment large-scale images containing unknown objects, which never appear in the exemplar images.
Wenbin Zou, Cong Bai, Kidiyo Kpalma, Joseph Ronsin
IEEE Trans. Image Process.1
2013 Segmentation Driven Low-rank Matrix Recovery for Saliency Detection
abstract
Low-rank matrix recovery (LRMR) model, aiming at decomposing a matrix into a low-rank matrix and a sparse one, has shown the potential to address the problem of saliency detection, where the decomposed low-rank matrix naturally corresponds to the background, and the sparse one captures salient objects. This is under the assumption that the background is consistent and objects are obviously distinctive. Unfortunately, in real images, the background may be cluttered and may have low contrast with objects. Thus directly applying the LRMR model to the saliency detection has limited robustness. This paper proposes a novel approach that exploits bottom-up segmentation as a guidance cue of the matrix recovery. This method is fully unsupervised, yet obtains higher performance than the supervised LRMR model. A new challenging dataset PASCAL-1500 is also introduced to validate the saliency detection performance. Extensive evaluations on the widely used MSRA-1000 dataset and also on the new PASCAL-1500 dataset demonstrate that the proposed saliency model outperforms the state-of-the-art models.
Wenbin Zou, Kidiyo Kpalma, Zhi Liu 0003, Joseph Ronsin
BMVC1
2012 Semantic segmentation via sparse coding over hierarchical regions
abstract
The purpose of this paper is segmenting objects in an image and assigning a predefined semantic label to each object. There are two contributions in this paper. On one hand, semantic segmentation is guided by hierarchical regions instead of by single-level regions or multi-scale regions generated by multiple segmentations. On the other hand, sparse coding is introduced as high level description of the regions, which contributes to reduction of quantization error compared to traditional bag-of-visual-words method. Experiments on the challenging Microsoft Research Cambridge dataset (MSRC 21) show that our algorithm achieves state-of-the-art performance.
Wenbin Zou, Kidiyo Kpalma, Joseph Ronsin
ICIP1
2012 Semantic image segmentation using region bank
Wenbin Zou, Kidiyo Kpalma, Joseph Ronsin
ICPR1