Hongliang Li 0001

dblp:91/1905-1 · DBLP profile ↗
← Back
228ranked-venue papers
25as first author
93since 2021 · last 2026
0000-0002-7481-095XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 172 · 21 first-author · 69 since 2021Artificial intelligence and machine learning · 45 · 28 since 2021Systems, architecture and hardware · 13 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 4 since 2021Computer networks · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Parameter Merging with Gradient-Guided Supermasks in Online Continual Learning
abstract
Online continual learning (OCL) aims at learning a non-stationary data stream in a way of reading each data sample only once, and hence suffers from the trade-off of catastrophic forgetting and insufficient learning. In this work, we firstly analytically establish relationship between loss functions and model parameters from the Bayesian perspective. Based on our analysis, we subsequently propose a parameter merging method with gradient-guided supermasks. Our method leverages 1-order and 2-order gradient information to construct supermasks that determine the merging weights between the old and new models. Our method performs direct arithmetic operations on parameters to update models, beyond traditional gradient descent. We further discover that a widely-used premise that 1-order gradients can be negligible is invalid in OCL, due to slow convergence incurred by insufficient learning. Additionally, we utilize a dual-model dual-view distillation strategy that can align output distributions of the new and merged models for each sample, further enhancing model performance. Extensive experiments are conducted on four benchmarks in OCL settings, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-100. Experimental results demonstrate that our method is effective, and achieves a substantial boost over previous methods.
Benliu Qiu, Heqian Qiu, Lanxiao Wang, Taijin Zhao, Lili Pan 0001, Hongliang Li 0001
AAAI7
2026 Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze Prediction
abstract
Egocentric gaze prediction serves as a critical indicator for decoding human visual attention and cognitive processes, but its inherently limited field of view creates prediction challenges. Although exo-view data provides supplementary contextual information, it exhibits significant spatial and semantic gaps. Existing methods focus solely on isolated feature encoding in single-view paradigms, neglecting cross-view gaze correlations. To make up for this gap, we make the first exploration of cross-view gaze relationship for egocentric gaze prediction, and propose Ego-PMOVE, a novel Prompt-aware Mixture of View Experts network. Unlike prior cross-view studies that forcibly align cross-view features thereby introducing inference noise, we leverage the popular Mixture-of-Experts (MoE) and a set of flexible prompts to disentangle features from different views into three parallel experts: a view-shared expert directly modeling common semantic relationships, a view-discrepancy expert adaptively adjusting the spatial position, scale and shifts based on different view-specific features, and an egocentric expert extracting independent features to compensate for the case of missing exocentric data. To balance these experts, we further design a soft router to dynamically weight them for mining useful information while suppressing noise. A view-query gaze decoder then generates view-specific gaze attention maps, jointly optimized by gaze-heamap and cross-view contrastive loss that regularize both shared and divergent features for accurate gaze prediction. Extensive experiments across the multi-view EgoMe dataset and single-view Ego4D and EGTEA Gaze++ datasets demonstrate the effectiveness and generalizability of our approach.
Heqian Qiu, Lanxiao Wang, Taijin Zhao, Zhaofeng Shi, Linfeng Xu 0001, Hongliang Li 0001
AAAI7
2026 CMaP-SAM: Contraction mapping prior for SAM-driven few-shot segmentation
Fanman Meng, Liming Lei, Qingbo Wu 0001, Linfeng Xu 0001, Hongliang Li 0001
Neurocomputing8
2026 Zero-shot egocentric action recognition via chain-of-imagination prompts and inertial strengthening adaptor
Mingzhou He, Ruiqian Li, Qingbo Wu 0001, King Ngi Ngan, Fanman Meng, Hongliang Li 0001
Pattern Recognit.7
2026 LoRA-based continual learning with constraints on critical parameter changes
Shimou Ling, Liang Zhang 0054, Jiangwei Zhao, Lili Pan 0001, Hongliang Li 0001
Pattern Recognit.5
2026 Bridging the Gap between Vision and Text for unsupervised text-only captioning
Lanxiao Wang, Heqian Qiu, Haitao Wen, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
Pattern Recognit.6
2026 Zippo: RGB-Alpha Joint Modeling With a Unified Diffusion Model
abstract
Recent advances in generative models have sparked growing interest in moving beyond pure image generation toward transparent image generation, i.e., joint generation of image and its alpha mask. However, most existing approaches adopt a two-stage pipeline, where a diffusion-based model first generates an RGB image and a subsequent matting head predicts the alpha mask. This separation not only leads to error accumulation and inaccurate predictions but also overlooks the intrinsic correlation between the cross-modal data. In this work, we introduce Zippo, a unified diffusion framework, zipping color and transparency distributions into a single diffusion model, by learning joint distribution of RGB image and alpha mask. Zippo not only generates high-fidelity images but also produces plausible and sharp alpha masks. In practice, Zippo inflates the latent space into a unified representation that encodes cross-modal data, and builds upon it with a modality-aware diffusion process that flexibly switches between RGB and alpha domains. In this process, conditioning on one modality while denoising the other allows the model to generate RGB images from alpha masks and predict transparency from input images. In addition to single-modality prediction, we further design a modality-aware noise reassignment strategy to empower Zippo with the joint generation capability of RGB images and their corresponding alpha masks under text guidance. With these techniques, Zippo supports a wide range of transparent image generation tasks, including image-alpha joint generation, image matting, and alpha mask conditioned image generation. Extensive experiments demonstrate that Zippo not only delivers superior visual fidelity but also achieves competitive performance in visual downstream prediction, highlighting joint image-alpha modeling as a powerful alternative to traditional paradigms.
Kangyang Xie, Chenchen Jing, Cheng Peng 0011, Ming Yang 0007, Heqian Qiu, Hongliang Li 0001, Hao Chen 0041
IEEE Trans. Circuits Syst. Video Technol.9
2026 DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
abstract
Continual learning of vision-language models (VLMs) focuses on leveraging cross-modal pretrained knowledge to incrementally adapt to expanding downstream tasks and datasets, while tackling the challenge of knowledge forgetting. Existing research often focuses on connecting visual features with specific class text in downstream tasks, overlooking the latent relationships between general and specialized knowledge. Our findings reveal that forcing models to optimize inappropriate visual-text matches exacerbates forgetting of VLM's recognition ability. To tackle this issue, we propose DesCLIP, which leverages general attribute (GA) descriptions to guide the understanding of specific class objects, enabling VLMs to establish robust <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-GA-class</i> trilateral associations rather than relying solely on <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-class</i> connections. Specifically, we introduce a language assistant to generate concrete GA description candidates via proper request prompts. Then, an anchor-based embedding filter is designed to obtain highly relevant GA description embeddings, which are leveraged as the paired text embeddings for visual-textual instance matching, thereby tuning the visual encoder. Correspondingly, the class text embeddings are gradually calibrated to align with these shared GA description embeddings. Extensive experiments demonstrate the advancements and efficacy of our proposed method, with comprehensive empirical evaluations highlighting its superior performance in VLM-based recognition compared to existing continual learning methods.
Chiyuan He, Zihuan Qiu, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001
IEEE Trans. Multim.6
2026 On the Adversarial Robustness of Learning-Based Image Compression Against Rate-Distortion Attacks
abstract
Despite demonstrating superior Rate-Distortion (RD) performance, Learning-based Image Compression (LIC) algorithms have been found to be vulnerable to malicious perturbations in recent studies. However, the adversarial attacks considered in existing literature remain divergent from real-world scenarios, both in terms of the attack direction and bitrate. Additionally, existing methods focus solely on empirical observations of the model vulnerability, neglecting to identify the origin of it. These limitations hinder the comprehensive investigation and in-depth understanding of the adversarial robustness of LIC algorithms. To address the aforementioned issues, this paper considers the arbitrary nature of the attack direction and the uncontrollable compression ratio faced by adversaries, and presents two practical rate-distortion attack paradigms,i.e., Specific-ratio Rate-Distortion Attack (SRDA) and Agnostic-ratio Rate-Distortion Attack (ARDA). To the best of our knowledge, we are the first to conduct joint rate-distortion attacks on LIC algorithms. Using the performance variations as indicators, we evaluate the adversarial robustness of eight predominant LIC algorithms against diverse attacks. Furthermore, we propose two novel analytical tools for in-depth analysis,i.e., Entropy Causal Intervention and Layer-wise Distance Magnify Ratio, and reveal thathyperpriorsignificantly increases the bitrate andInverse Generalized Divisive Normalization (IGDN)significantly amplifies input perturbations when under attack. Lastly, we examine the efficacy of adversarial training and introduce the use of online updating for defense. By comparing their advantages and disadvantages, we provide a reference for constructing more robust LIC algorithms against the rate-distortion attacks.
Qingbo Wu 0001, Lei Wang 0186, Fanman Meng, King Ngi Ngan, Li Zhuo 0001, Hongliang Li 0001
IEEE Trans. Multim.8
2025 Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
abstract
Unlike traditional Multimodal Class-Incremental Learning (MCIL) methods that focus only on vision and text, this paper explores MCIL across vision, audio and text modalities, addressing challenges in integrating complementary information and mitigating catastrophic forgetting. To tackle these issues, we propose an MCIL method based on multimodal pre-trained models. Firstly, a Multimodal Incremental Feature Extractor (MIFE) based on Mixture-of-Experts (MoE) structure is introduced to achieve effective incremental fine-tuning for AudioCLIP. Secondly, to enhance feature discriminability and generalization, we propose an Adaptive Audio-Visual Fusion Module (AAVFM) that includes a masking threshold mechanism and a dynamic feature fusion mechanism, along with a strategy to enhance text diversity. Thirdly, a novel multimodal class-incremental contrastive training loss is proposed to optimize cross-modal alignment in MCIL. Finally, two MCIL-specific evaluation metrics are introduced for comprehensive assessment. Extensive experiments on three multimodal datasets validate the effectiveness of our method.
Zihuan Qiu, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001
ICASSP4
2025 Cmp: Composable Meta Prompt for Sam-Based Cross-Domain Few-Shot Segmentation
abstract
Cross-Domain Few-Shot Segmentation (CD-FSS) remains challenging due to limited data and domain shifts. Recent foundation models like the Segment Anything Model (SAM) have shown remarkable zero-shot generalization capability in general segmentation tasks, making it a promising solution for few-shot scenarios. However, adapting SAM to CD-FSS faces two critical challenges: reliance on manual prompt and limited cross-domain ability. Therefore, we propose the Composable Meta-Prompt (CMP) framework that introduces three key modules: (i) the Reference Complement and Transformation (RCT) module for semantic expansion, (ii) the Composable Meta-Prompt Generation (CMPG) module for automated meta-prompt synthesis, and (iii) the Frequency-Aware Interaction (FAI) module for domain discrepancy mitigation. Evaluations across four cross-domain datasets demonstrate CMP’s state-of-the-art performance, achieving 71.8% and 74.5% mIoU in 1-shot and 5-shot scenarios respectively.
Fanman Meng, Chunjin Yang, Qingbo Wu 0001, Hongliang Li 0001
ICIP7
2025 DPM-CLIP: Zero-Shot Multimodal Egocentric Activity Recognition based on Dual-Prediction Mechanism
abstract
Advancements in Zero-shot Multimodal Egocentric Activity Recognition (ZS-MM-EAR) largely rely on Vision-Language Model (VLM). However, existing methods struggle with VLM’s inadequate representation of egocentric activities, including challenges in capturing egocentric-specific features, adapting to domain shifts between egocentric video and pre-training data, and effectively leveraging complementary data such as Inertial Measurement Unit (IMU). To address these issues, we propose DPM-CLIP, a ZS-MM-EAR method tailored for vision, text and IMU modalities. Firstly, we design an attribute-driven text augmentation module that leverages a Large Language Model (LLM) to generate fine-grained textual descriptions of activities. Secondly, we construct an Instance-feature Repository (IFR) to store base class features and generate pseudo-features for novel classes through feature center migration. Finally, we introduce a dual-prediction mechanism with a prediction correction module to enhance generalization and recognition accuracy. Extensive experiments on the UESTC-MMEA-CL dataset validate the effectiveness of the proposed method.
Zihuan Qiu, Mingzhou He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001
ICIP8
2025 Your Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low Bitrate
abstract
With the help of powerful generative models, Semantic Image Compression (SIC) has achieved impressive performance at ultra-low bitrate. However, due to coarse-grained visual-semantic alignment and inherent randomness, the reliability of SIC is seriously concerned for reconstructing completely different object instances, even they are semantically consistent with original images. To tackle this issue, we propose a novel Referring Semantic Image Compression (RSIC) framework to improve the fidelity of user-specified content while retaining extreme compression ratios. Specifically, RSIC consists of three modules: Global Description Encoding (GDE), Referring Guidance Encoding (RGE), and Guided Generative Decoding (GGD). GDE and RGE encode global semantic information and local features, respectively, while GGD handles the non-uniformly guided generative process based on the encoded information. In this way, our RSIC achieves flexible customized compression according to user demands, which better balance the local fidelity, global realism, semantic alignment, and bit overhead. Extensive experiments on three datasets verify the compression efficiency and flexibility of the proposed method.
Qingbo Wu 0001, Mingzhou He, King Ngi Ngan, Fanman Meng, Hongliang Li 0001
ISCAS8
2025 Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
abstract
Even from an early age, humans naturally adapt between exocentric (Exo) and egocentric (Ego) perspectives to understand daily procedural activities. Inspired by this cognitive ability, we propose a novel Unsupervised Ego-Exo Dense Procedural Activity Captioning (UE^2 DPAC) task, which aims to transfer knowledge from the labeled source view to predict the time segments and descriptions of action sequences for the target view without annotations. Despite previous works endeavoring to address the fully-supervised single-view or cross-view dense video captioning, they lapse in the proposed task due to the significant inter-view gap caused by temporal misalignment and irrelevant object interference. Hence, we propose a Gaze Consensus-guided Ego-Exo Adaptation Network (GCEAN) that injects the gaze information into the learned representations for the fine-grained Ego-Exo alignment. Specifically, we propose a Score-based Adversarial Learning Module (SALM) that incorporates a discriminative scoring network and compares the scores of distinct views to learn unified view-invariant representations from a global level. Then, the Gaze Consensus Construction Module (GCCM) utilizes the gaze to progressively calibrate the learned representations to highlight the regions of interest and extract the corresponding temporal contexts. Moreover, we adopt hierarchical gaze-guided consistency losses to construct gaze consensus for the explicit temporal and spatial adaptation between the source and target views. To support our research, we propose a new EgoMe-UE^2 DPAC benchmark, and extensive experiments demonstrate the effectiveness of our method, which outperforms many related methods by a large margin. Code is available at https://github.com/ZhaofengSHI/GCEAN.
Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001
ACM Multimedia6
2025 DFR: A Decompose-Fuse-Reconstruct Framework for Multi-Modal Few-Shot Segmentation
abstract
This paper presents DFR (Decompose, Fuse and Reconstruct), a novel framework that addresses the fundamental challenge of effectively utilizing multi-modal guidance in few-shot segmentation (FSS). While existing approaches primarily rely on visual support samples or textual descriptions, their single or dual-modal paradigms limit exploitation of rich perceptual information available in real-world scenarios. To overcome this limitation, the proposed approach leverages the Segment Anything Model (SAM) to systematically integrate visual, textual, and audio modalities for enhanced semantic understanding. The DFR framework introduces three key innovations: 1) Multi-modal Decompose: a hierarchical decomposition scheme that extracts visual region proposals via SAM, expands textual semantics into fine-grained descriptors, and processes audio features for contextual enrichment; 2) Multi-modal Contrastive Fuse: a fusion strategy employing contrastive learning to maintain consistency across visual, textual, and audio modalities while enabling dynamic semantic interactions between foreground and background features; 3) Dual-path Reconstruct: an adaptive integration mechanism combining semantic guidance from tri-modal fused tokens with geometric cues from multi-modal location priors. Extensive experiments across visual, textual, and audio modalities under both synthetic and real settings demonstrate DFR's substantial performance improvements over state-of-the-art methods.
Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
MMSP7
2025 DBAB: A Dual-Branch Adaptive Balance Framework with Optimized Plasticity Branch for Class-Incremental Learning
abstract
In the context of class-incremental learning, the primary challenge for models is to overcome catastrophic forgetting. Leveraging the strong generalization ability of frozen pre-trained models can significantly enhance model performance and alleviate catastrophic forgetting during training. To enable models to better adapt to downstream tasks, fine-tuning pre-trained models for new tasks is a common approach. However, current works struggle to balance plasticity and generalization performance after fine-tuning, as fine-tuning causes model parameters to overwrite knowledge of old tasks. This paper proposes a Dual-Branch Adaptive Balance (DBAB) framework, which consists of a plasticity branch fine-tuned from a pre-trained model and optimized for downstream tasks, and a generalization branch with frozen pre-trained parameters. The framework designs an adaptive balance mechanism for the dual branches, introduces learnable balance coefficients to dynamically fuse class prototype distances from both branches, and devises a loss function for training and regularizing the balance coefficients. This ensures a better balance between plasticity and generalization during the incremental learning process. To optimize the plasticity branch in the DBAB framework, an Adaptive Plasticity Module (APM) is proposed. Considering the heterogeneity of embedding distributions in continuous learning of downstream tasks, APM employs Mahalanobis distance for anisotropic feature alignment, uses a covariance matrix to dynamically adapt to the heterogeneous distributions of new tasks, and improves and stabilizes the Mahalanobis distance-based classification method.Experimental results show that DBAB outperforms multiple state-of-the-art (SOTA) methods on benchmark datasets such as CIFAR100 and CUB200, demonstrating significant performance improvements.
Heqian Qiu, Chenghao Qi, Ruisong Dai, Hongliang Li 0001
MMSP5
2025 OrthCal: Synergizing Orthogonal Contrastive Learning and Prototype Calibration for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) requires models to progressively learn novel classes with limited samples while mitigating catastrophic forgetting of base classes. Existing methods face dual challenges: novel classes are prone to misclassification into base classes because the strong discriminability of base classes distracts the classification of novel classes, and the feature space lacks sufficient generalization capacity. This paper proposes the OrthCal framework built on the deep integration of orthogonal contrastive learning and a prototype calibration strategy to improve the performance during incremental sessions. Our three-stage optimization includes: 1) pretraining with hybrid supervised and self-supervised contrastive learning to construct geometrically constrained orthogonal pseudo-targets. 2) Dynamic prototype calibration, using semantic similarity among base classes to adjust novel class prototypes without additional training. 3) Hybrid loss design optimizing orthogonality constraints, perturbation-sensitive contrastive loss, and calibrated prototypes jointly to address challenges arising from data limitations during incremental sessions. Experiments on miniImageNet and CIFAR100 demonstrate that OrthCal achieves state-of-the-art performance. The framework provides a unified solution for feature space optimization and prototype calibration in FSCIL.
Ruisong Dai, Chenghao Qi, Heqian Qiu, Hongliang Li 0001
MMSP6
2025 D3Net: Dual-Path Decoupling-Distillation for Adaptive Fusion in Continual Egocentric Learning
abstract
Egocentric continual action recognition faces severe challenges such as sudden viewpoint changes, occlusions, and complex backgrounds. In such scenarios, relying solely on visual modalities is susceptible to interference and lacks sufficient recognition robustness. To overcome the limitations of unimodal approaches, multimodal fusion methods are widely adopted, significantly enhancing recognition performance. However, existing multimodal schemes generally suffer from insufficient exploration of cross-modal complementarity and the vulnerability of modal independence. To address this, this paper proposes a Dual-path Decoupling-Distillation NetWork (D3Net), aiming to achieve more effective dynamic fusion of modal information and knowledge transfer.D3Net first explicitly separates the shared and private features of modalities through a dual-path decoupling module, combined with a dynamic gating mechanism to adaptively adjust the modal fusion weights. Secondly, it designs a complementary distillation module, leveraging cross-modal contrastive learning to effectively mitigate the issues of poor unimodal robustness and vulnerability to interference. Finally, through a cross-task distillation mechanism, it efficiently extracts knowledge from old tasks, alleviating the catastrophic forgetting problem during learning. Experimental results demonstrate that D3Net achieves an average accuracy of 83.97% under the 8×4 task configuration on the UESTC MMEA CL dataset, surpassing baseline method by 5.17%.
Chenghao Qi, Heqian Qiu, Zhaofeng Shi, Lanxiao Wang, Hongliang Li 0001
MMSP7
2025 Efficient Polyp Detection via Wavelet-Driven Boundary Enhancement and Temporal Consistency
abstract
Accurate early detection of polyps plays a critical role in preventing, diagnosing, and treating colorectal cancer. Although significant progress has been made, the accurate and efficient detection of polyps remains a challenging task. Existing image-based methods, while computationally efficient, typically rely on single-frame inputs and struggle to handle polyps with ambiguous boundaries or varying sizes. On the other hand, video-based approaches leverage temporal information to improve detection robustness, but often incur high computational costs and compromise real-time performance due to the need to process multiple frames simultaneously. Moreover, both types of methods are susceptible to dynamic artifacts caused by endoscopic camera movement, which can lead to polyp-like false positives. To address these issues, we propose BEC-Net, a novel Boundary-Enhanced network with adjacent-Frame Contrastive Learning for accurate and efficient polyp detection. Specifically, we design a Wavelet-Based Boundary-Aware feature Fusion (WBAF) module to enhance the representation of polyp boundaries and improve generalization across diverse appearances. To accommodate scale variation, we introduce a Context-Aware Gated Aggregation (CAGA) module that adaptively integrates multi-scale contextual information. Furthermore, we propose an Adjacent-Frame Contrastive Learning (AFCL) strategy that utilizes temporal consistency between adjacent frames to suppress polyp-like artifacts without increasing inference cost. Extensive experiments on large-scale colonoscopy benchmarks demonstrate that our method outperforms state-of-the-art approaches in both accuracy and real-time performance.
Heqian Qiu, Lanxiao Wang, Chenghao Qi, Ruisong Dai, Hongliang Li 0001
MMSP6
2025 MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
abstract
Continual model merging integrates independently fine-tuned models sequentially without access to the original training data, offering a scalable and efficient solution for continual learning. However, existing methods face two critical challenges: parameter interference among tasks, which leads to catastrophic forgetting, and limited adaptability to evolving test distributions. To address these issues, we introduce the task of Test-Time Continual Model Merging (TTCMM), which leverages a small set of unlabeled test samples during inference to alleviate parameter conflicts and handle distribution shifts. We propose MINGLE, a novel framework for TTCMM. MINGLE employs a mixture-of-experts architecture with parameter-efficient, low-rank experts, which enhances adaptability to evolving test distributions while dynamically merging models to mitigate conflicts. To further reduce forgetting, we propose Null-Space Constrained Gating, which restricts gating updates to subspaces orthogonal to prior task representations, thereby suppressing activations on old tasks and preserving past knowledge. We further introduce an Adaptive Relaxation Strategy that adjusts constraint strength dynamically based on interference signals observed during test-time adaptation, striking a balance between stability and adaptability. Extensive experiments on standard continual merging benchmarks demonstrate that MINGLE achieves robust generalization, significantly reduces forgetting, and consistently surpasses previous state-of-the-art methods by 7–9% on average across diverse task orders. Our code is available at: https://github.com/zihuanqiu/MINGLE
Zihuan Qiu, Yi Xu 0008, Chiyuan He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001
NeurIPS7
2025 GRSDet: Learning to Generate Local Reverse Samples for Few-shot Object Detection
Hefei Mei, Taijin Zhao, Shiyuan Tang, Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Fanman Meng, Hongliang Li 0001
Neurocomputing8
2025 Adaptively forget with crossmodal and textual distillation for class-incremental video captioning
Huiyu Xiong, Lanxiao Wang, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001
Neurocomputing6
2025 High efficiency deep image compression via channel-wise scale adaptive latent representation learning
Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
Signal Process. Image Commun.4
2025 MCCE-REC: MLLM-Driven Cross-Modal Contrastive Entropy Model for Zero-Shot Referring Expression Comprehension
abstract
Zero-shot referring expression comprehension (zero-shot REC) is a crucial yet challenging task in the field of multi-modal understanding, which aims to locate an object described by a referring expression without training on task-specific datasets. Existing methods take advantage of a pre-trained CLIP model to align cropped proposal regions with referring expressions. However, our analysis reveals that this aligning way heavily biases toward certain salient visual regions due to CLIP focusing on global-level image-text matching. To mitigate this bias, we propose MCCE-REC, an MLLM-driven cross-modal contrastive entropy model for training-free zero-shot REC. Benefiting from the remarkable in-context comprehension ability of the multi-modal large language model (MLLM), we design a set of referring prompts for MLLM to generate diverse detailed informative, and contrastive cues related to referring objects. Based on these cues, on the one hand, we propose a multi-cues cross-modal interaction network, which associates the visual features and referring object textual features from multiple perspectives and perceives surrounding context object information in a parameter-free manner, avoiding bias towards salient features. On the other hand, we introduce a contrastive similarity entropy selection mechanism that compares the positive and negative cues to suppress biased regions with high similarity scores and emphasizes accurate regions correlating with referring descriptions. Extensive experiments demonstrate our MCCE-REC outperforms existing zero-shot methods by a significant margin on various REC datasets.
Heqian Qiu, Lanxiao Wang, Taijin Zhao, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Cognition Transferring and Decoupling for Text-Supervised Egocentric Semantic Segmentation
abstract
In this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective potential, the egocentric scenes contain dense wearer-object relations and inter-object interference. However, most recent third-view methods leverage the frozen Contrastive Language-Image Pre-training (CLIP) model, which is pre-trained on the semantic-oriented third-view data and lapses in the egocentric view due to the “relation insensitive” problem. Hence, we propose a Cognition Transferring and Decoupling Network (CTDN) that first learns the egocentric wearer-object relations via correlating the image and text. Besides, a Cognition Transferring Module (CTM) is developed to distill the cognitive knowledge from the large-scale pre-trained model to our model for recognizing egocentric objects with various semantics. Based on the transferred cognition, the Foreground-background Decoupling Module (FDM) disentangles the visual representations to explicitly discriminate the foreground and background regions to mitigate false activation areas caused by foreground-background interferential objects during egocentric relation learning. Extensive experiments on four TESS benchmarks demonstrate the effectiveness of our approach, which outperforms many recent related methods by a large margin. Code will be available athttps://github.com/ZhaofengSHI/CTDN.
Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Class Incremental Learning With Less Forgetting Direction and Equilibrium Point
abstract
Catastrophic forgetting is the core problem of class incremental learning (CIL). Existing work mainly adopts memory replay, knowledge distillation, and dynamic architecture to alleviate this problem, but seldom from the aspect of parameter regularization. However, existing parameter regularization methods struggle to achieve an appropriate balance between old and new tasks. To bring it back to CIL, we first propose constrained incremental learning with less forgetting direction (LFD) to leave more plasticity for the new task under a strong stability constraint for old tasks. Specifically, the new parameters are constrained to be close to the LFD of old tasks instead of a single group of old parameters. To validate the effectiveness of this regularization, we investigate the connectivity between the old parameters and the new parameters, and additionally find that a higher accuracy interval exists along the linear connection. Therefore, we further propose a post-processing procedure to find an equilibrium point in this interval for better balance between old and new tasks. Extensive classification experiments on CIFAR-100, ImageNet-100, and ImageNet-1K show our method can significantly improve performance compared with existing CIL methods and the object detection experiments on PASCAL-VOC show its broad generality on other tasks.
Haitao Wen, Heqian Qiu, Lanxiao Wang, Haoyang Cheng, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Geodesic-Aligned Gradient Projection for Continual Task Learning
abstract
Deep networks notoriously suffer from performance deterioration on previous tasks when learning from sequential tasks, i.e., catastrophic forgetting. Recent methods of gradient projection show that the forgetting is resulted from the gradient interference on old tasks and accordingly propose to update the network in an orthogonal direction to the task space. However, these methods assume the task space is invariant and neglect the gradual change between tasks, resulting in sub-optimal gradient projection and a compromise of the continual learning capacity. To tackle this problem, we propose to embed each task subspace into a non-Euclidean manifold, which can naturally capture the change of tasks since the manifold is intrinsically non-static compared to the Euclidean space. Subsequently, we analytically derive the accumulated projection between any two subspaces on the manifold along the geodesic path by integrating an infinite number of intermediate subspaces. Building upon this derivation, we propose a novel geodesic-aligned gradient projection (GAGP) method that harnesses the accumulated projection to mitigate catastrophic forgetting. The proposed method utilizes the geometric structure information on the task manifold by capturing the gradual change between the new and the old tasks. Empirical studies on image classification demonstrate that the proposed method alleviates catastrophic forgetting and achieves on-par or better performance compared to the state-of-the-art approaches.
Benliu Qiu, Heqian Qiu, Haitao Wen, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
IEEE Trans. Image Process.8
2025 GCSTG: Generating Class-Confusion-Aware Samples With a Tree-Structure Graph for Few-Shot Object Detection
abstract
Few-Shot Object Detection (FSOD) aims to detect the objects of novel classes using only a few manually annotated samples. With the few novel class samples, learning the inter-class relationships among foreground and constructing the corresponding class hierarchy in FSOD is a challenging task. The poor construction of the class hierarchy will result in the inter-class confusion problem, which has been identified as a primary cause of inferior performance in novel classes by recent FSOD methods. In this work, we further find that the intra-super-class confusion, where samples are misclassified as classes within their associated super-classes, is the main challenge in solving the confusion problem. To solve this issue, this work generates class-confusion-aware samples with a pre-defined tree-structure graph, for helping models to construct a precise class hierarchy. In precise, for generating class-confusion-aware samples, we add the noise into available samples and update the noise to maximize confidence scores on associated confusion categories of samples. Then, a confusion-aware curriculum learning strategy is proposed to make generated samples gradually participate in the training, which benefits the model convergence while learning the generated samples. Experimental results show that our method can be used as a plug-in in recent FSOD methods and consistently improve the model performance.
Longrong Yang, Hanbin Zhao, Hongliang Li 0001, Liang Qiao 0001, Xi Li 0001
IEEE Trans. Image Process.3
2025 Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding Confusion
abstract
Continual learning (CL) aims to enable deep neural networks (DNNs) to learn new data without forgetting previously learned knowledge. The key to achieving this goal is to avoid confusion at the feature level, i.e., to avoid confusion within old tasks and between new and old tasks. Existing prototype-based CL methods generate pseudo features for old knowledge replay by adding Gaussian noise to the centroids of old classes. However, the distribution in the feature space exhibits anisotropy during the incremental process, which prevents the pseudo features from faithfully reproducing the distribution of old knowledge in the feature space, leading to confusion at the classification boundaries within old tasks. To address this issue, we propose the distribution-level memory recall (DMR) method, which uses a Gaussian mixture model to precisely fit the feature distribution of old knowledge at the distribution level and generate pseudo features in the next stage. Furthermore, resistance to confusion at the distribution level is crucial for multimodal learning. Multimodal imbalance, which refers to uneven optimization processes among encoders of different modalities, results in significant differences in feature responses between modalities; this exacerbates confusion within old tasks in prototype-based CL methods. Therefore, we mitigate the multimodal imbalance problem by using the intermodal guidance and intramodal mining (IGIM) method to guide weaker modalities with prior information from dominant modalities and further explore useful information within modalities. To avoid confusion between new and old tasks, we propose using the confusion index to quantitatively describe a model's ability to distinguish between new and old tasks, and we use the incremental mixup feature enhancement (IMFE) method to enhance pseudo features with new sample features, alleviating classification confusion between new and old knowledge. We conduct extensive experiments on the CIFAR100, ImageNet100, TinyImageNet, ImageNet-1K and UESTC-MMEA-CL datasets and achieve state-of-the-art results.
Shaoxu Cheng, Kanglei Geng, Chiyuan He, Zihuan Qiu, Linfeng Xu 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001
IEEE Trans. Multim.10
2025 Cross-Modal Cognitive Consensus Guided Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented reality, and intelligent robot systems. The pioneering work conducts this task through dense feature-level audio-visual interaction, which ignores the dimension gap between different modalities. More specifically, the audio clip could only provide aGlobalsemantic label in each sequence, but the video frame covers multiple semantic objects across differentLocalregions, which leads to mislocalization of the representationally similar but semantically different object. In this paper, we propose a Cross-modal Cognitive Consensus guided Network (C3N) to align the audio-visual semantics from the global dimension and progressively inject them into the local regions via an attention mechanism. Firstly, a Cross-modal Cognitive Consensus Inference Module (C3IM) is developed to extract a unified-modal label by integrating audio/visual classification confidence and similarities of modality-agnostic label embeddings. Then, we feed the unified-modal label back to the visual backbone as the explicit semantic-level guidance via a Cognitive Consensus guided Attention Module (CCAM), which highlights the local features corresponding to the interested object. Extensive experiments on the Single Sound Source Segmentation (S4) setting and Multiple Sound Source Segmentation (MS3) setting of the AVSBench dataset demonstrate the effectiveness of the proposed method, which achieves state-of-the-art performance.
Zhaofeng Shi, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001
IEEE Trans. Multim.5
2025 Learning With Noisy Low-Cost MOS for Image Quality Assessment via Dual-Bias Calibration
abstract
Learning-based Image Quality Assessment (IQA) models have obtained impressive performance with the help of reliable subjective quality labels, where Mean Opinion Score (MOS) is the most popular choice. However, in view of the subjective bias of individual annotators, the Labor-Abundant MOS (LA-MOS) typically requires large collections of opinion scores from multiple annotators for each image, which significantly increases the learning cost. In this paper, we aim to learn robust IQA models from Low-Cost MOS (LC-MOS), which only requires very few opinion scores or even a single opinion score for each image. More specifically, we consider the LC-MOS as the noisy observation of LA-MOS and enforce the IQA model learned from LC-MOS to approach the unbiased estimation of LA-MOS. Thus, we represent the subjective bias between LC-MOS and LA-MOS, and the model bias between IQA predictions learned from LC-MOS and LA-MOS (i.e., dual-bias) as two latent variables with unknown parameters. By means of the expectation-maximization-based alternating optimization, we can jointly estimate the parameters of the dual-bias, which suppresses the misleading of LC-MOS via a gated dual-bias calibration (GDBC) module. To the best of our knowledge, this is the first exploration of robust IQA model learning from noisy low-cost labels. Theoretical analysis and extensive experiments on four popular IQA datasets show that the proposed method is robust toward different bias rates and annotation numbers and significantly outperforms the other Learning-based IQA models when only LC-MOS is available. Furthermore, we also achieve comparable performance with respect to the other models learned with LA-MOS.
Lei Wang 0029, Qingbo Wu 0001, Desen Yuan, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
IEEE Trans. Multim.5
2024 Dual-Consistency Model Inversion for Non-Exemplar Class Incremental Learning
abstract
Non-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge without forgetting previously acquired ones when historical data are un-available. One of the generative NECIL methods is to in-vert the images of old classes for joint training. However, these synthetic images suffer significant domain shifts compared with real data, hampering the recognition of old classes. In this paper, we present a novel method termed Dual-Consistency Model Inversion (DCMI) to generate better synthetic samples of old classes through two pivotal consistency alignments: (1) the semantic consistency between the synthetic images and the corresponding prototypes, and (2) domain consistency between synthetic and real images of new classes. Besides, we introduce Prototypical Routing (PR) to provide task-prior information and generate unbi-ased and accurate predictions. Our comprehensive experiments across diverse datasets consistently showcase the superiority of our method over previous state-of-the-art approaches.
Zihuan Qiu, Yi Xu 0008, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001
CVPR4
2024 Prompt-Driven Referring Image Segmentation with Instance Contrasting
abstract
Referring image segmentation (RIS) aims to segment the target referent described by natural language. Recently, large-scale pre-trained models, e.g., CLIP and SAM, have been successfully applied in many downstream tasks, but they are not well adapted to RIS task due to inter-task differences. In this paper, we propose a new prompt-driven framework named Prompt-RIS, which bridges CLIP and SAM end-to-end and transfers their rich knowledge and powerful capabilities to RIS task through prompt learning. To adapt CLIP to pixel-level task, we first propose a Cross-Modal Prompting method, which acquires more comprehensive vision-language interaction and fine-grained text-to-pixel alignment by performing bidirectional prompting. Then, the prompt-tuned CLIP generates masks, points, and text prompts for SAM to generate more accurate mask predictions. Moreover, we further propose Instance Contrastive Learning to improve the model's discriminability to different instances and robustness to diverse languages describing the same instance. Extensive experiments demonstrate that the performance of our method outperforms the state-of-the-art methods consistently in both general and open-vocabulary settings.
Chao Shang 0001, Zichen Song 0002, Heqian Qiu, Lanxiao Wang, Fanman Meng, Hongliang Li 0001
CVPR6
2024 Class Incremental Learning with Multi-Teacher Distillation
abstract
Distillation strategies are currently the primary approaches for mitigating forgetting in class incremental learning (CIL). Existing methods generally inherit previous knowledge from a single teacher. However, teachers with different mechanisms are talented at different tasks, and inheriting diverse knowledge from them can enhance compatibility with new knowledge. In this paper, we propose the MTD method to find multiple diverse teachers for CIL. Specifically, we adopt weight permutation, feature perturbation, and diversity regularization techniques to ensure diverse mechanisms in teachers. To reduce time and memory consumption, each teacher is represented as a small branch in the model. We adapt existing CIL distillation strategies with MTD and extensive experiments on CIFAR-100, ImageNet-100, and ImageNet-1000 show significant performance improvement. Our code is available at https://github.com/HaitaoWen/CLearning.
Haitao Wen, Lili Pan 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Hongliang Li 0001
CVPR7
2024 Vision-Sensor Attention Based Continual Multimodal Egocentric Activity Recognition
abstract
Continual learning aims to equip deep neural networks (DNNs) with the capability to continuously learn new knowledge without catastrophic forgetting. Currently, there is significant attention on multimodal continual activity recognition from a egocentric perspective. However, the issue of modality imbalance can lead to exacerbated forgetting in multimodal continual learning. To address this, we propose an exemplar-free vision-sensor Attention-based Incremental Discriminability enhancement (AID) method. Firstly, we employ a Vision-Sensor attention module to enhance the time-frequency information of sensor modality and fuse them with vision modality. This alleviates the modality imbalance problem, yielding more discriminative and generalizable representations. Simultaneously, to prevent the classifier from overfitting to old class prototypes, we enhance old prototypes with features from new classes, thereby enhancing classifier discriminability. We validate the effectiveness of this method through numerous experiments with various task settings on the UESTC-MMEA-CL dataset.
Shaoxu Cheng, Chiyuan He, Kailong Chen, Linfeng Xu 0001, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001
ICASSP5
2024 A Text Detector Based on the Specific Text Prompt
abstract
Nowadays, the prompt tuning has emerged as a novel new paradigm for adapting the original large-scale Contrastive Language-Image Pre-trained (CLIP) model into the downstream task as text detection. However, the learnable prompt adopted by the existing methods of prompt-tuning represents blurry and abstract meanings instead of fine-grained text feature. In this paper, we propose a powerful and robust text detector, called STP-TD, utilizing the specific text prompt and a learnable visual mask to fully apply the prior knowledge of CLIP model into the downstream task of text detection. STP-TD aims to make the split prompt character focusing on an ordered image token by Transformer mechanism. It is anticipated that a prompt character could stand for the fine-grained text feature of an image token, thus a distance loss is added to rectify prompt through optimization. Additionally, STP-TD firstly proposes a learnable visual mask to refine the text region in advance. Meanwhile, a synergetic framework is introduced as a bridge between visual branch and text branch. We also adopt the pixel-text matching process to align every pixel of image feature with the text feature. The experiments are conducted on the datasets ICDAR2015 and TotalText and outperform the state of the art.
Xingtao Lin, Chuanyang Gong, Lanxiao Wang, Heqian Qiu, Shengyu Tong, Hongliang Li 0001
ICIP6
2024 Video Class-Incremental Learning With Clip Based Transformer
abstract
Vision Language Pre-training Models have shown significant potential in various domains, but there are few attempts to introduce it in the field of continual learning for video action recognition. We propose Video Class-Incremental Learner with CLIP based Transformer (VCIL-CT), which uses CLIP based vision transformer to train action recognition task by class-incremental learning pipeline. To specifically address the issue of catastrophic forgetting in transformer, we introduce Attention Distillation which distilling the attention feature from each transformer decoder. In the process of incremental learning of classes, there may be a problem of high bias towards new classes, we incorporate Class Balance Module to prevent bias on new task. Furthermore, we adopt Exemplar Augment strategy to improve exemplar quality on data replay step. We evaluate our proposed method based on the incremental action recognition benchmark presented by TCD, using UCF101, HMDB51, and UESTC-MMEA-CL datasets, and demonstrate the effectiveness of our algorithm compared to existing state-of-the-art continuous learning methods for action recognition.
Shuyun Lu, Lanxiao Wang, Heqian Qiu, Xingtao Lin, Hefei Mei, Hongliang Li 0001
ICIP7
2024 Attribute-Prompting Multi-Modal Object Reasoning Transformer for Remote Sensing Visual Grounding
abstract
Remote sensing visual grounding (RSVG) task aims to locate the particular object in a remote sensing image referred to a natural language expression, which requires to precisely fuse and align features from different modalities. However, existing methods usually use object-based multi-modal fusion, which is limited to capturing the detailed object characteristics in remote sensing images, resulting in object confusion with similar objects. To address this problem, we propose an attribute-prompting multi-modal object reasoning network for RSVG. Specifically, we first develop a learnable attribute prompter to adaptively explore diverse and rich attribute information according to common object characteristics in RS. With the help of attribute prompts, we design an attribute-prompting multi-modal fusion encoder to build fine-grained interactive and alignment between the visual and language features to avoid object confusion. Furthermore, we design a multi-modal progressive object reasoning decoder to gradually query more comprehensive object features for accurate object localization. Experimental results demonstrate that the proposed method achieves significant improvements.
Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Taijin Zhao, Hongliang Li 0001
IGARSS5
2024 DP-RSCAP: Dual Prompt-Based Scene and Entity Network for Remote Sensing Image Captioning
abstract
As a challenging task towards remote sensing image analysis, the core problem of remote sensing image captioning is how to accurately transform the vision information into text information. Existing methods usually achieve it based on the simple multi-task learning strategy or visual attention mechanism, which ignores the importance of intermediate connection information for cross-modal transformation. To solve above problem, we propose a novel dual prompt-based scene and entity network (DP-RSCap) which aims to fully utilize the ability of cross-modal alignment in vision-language model build text prior information as intermediate connection to narrow the gap between different modalities and improve the quality of caption. Specifically, we first introduce an entity-concept prompt exporter to obtain explicit entity concepts in images. Then, we design a scene class prompt generator which can predict scene class and obtain fine-grained visual semantic features. Finally, we further design a dual prompt-based caption decoder to align and merge the visual semantic feature and dual prompts information as explicit intermediate connections, which can assist in generating precise caption. Extensive experiments on the challenging RSICD demonstrate the superior ability of our model.
Lanxiao Wang, Heqian Qiu, Minjian Zhang 0003, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
IGARSS6
2024 Robust Real-World Image Dehazing via Knowledge Guided Conditional Diffusion Model Finetuning
abstract
Due to the domain gap, the dehazing models trained from the synthetic images suffer poor generalization performance on real-world images. To address this issue, we pro-pose a Knowledge guided Conditional Diffusion (KCDiff) model finetuning method, which enables both the domain knowledge adaptation from the synthetic images and general knowledge guidance from the real-world images. More specifically, our KCDiff comprises two modules, i.e., the Conditional Image Generation (CIG) and Dehazing Instruction Generation (DIG). For CIG, we freeze a pre-trained latent diffusion model, add learnable conditioning control layers with Low-Rank Adaptation (LoRA) blocks, and include skip connections with zero-initialized convolutional layers, all of which play a fundamental role in image dehazing. Meanwhile, the DIG utilizes a large vision-language model LLaVA to extract the semantic content of the input hazy image and redescribe it in clear weather, which serves as the control instruction of CIG. To mitigate potential artifacts in CIG caused by misinterpretation of DIG's instructions, we further enforce depth and physical model-based reconstruction consistency constraints on both dehazing and hazy images. In the training phase, CIG is trained with the paired synthetic images to adapt the diffusion prior to the domain knowledge of image dehazing and finetuned with unpaired real-world images to suppress the domain gap with general knowledge guidance from the atmospheric scattering and depth perception. Experiments on real-world databases demonstrate the superiority of the proposed method over many state-of-the-art image dehazing models.
Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Fanman Meng, Hongliang Li 0001
MMSP8
2024 IoU-CLIP: IoU-Aware Language-Image Model Tuning for Open Vocabulary Object Detection
abstract
Open vocabulary object detection (OVD), which detects novel categories through detectors trained on base categories, has achieved remarkable advancement attributable to large-scale vision-language models, such as CLIP. The prior OVD works mainly focused on improving the classification accuracy of proposals, ignoring the ability of localization for novel categories. In this work, we propose IoU-aware language-image model tuning (IoU-CLIP) for open vocabulary object detection. Specifically, we construct a region image dataset with different IoU and adopt IoU values as labels to fine-tune the CLIP model to learn IoU-aware and class-agnostic semantic prompts and visual embeddings. The fine-tuned IoU-CLIP can predict IoU scores for proposals, which interact with classification scores. Meanwhile, IoU-aware and class-agnostic visual embeddings are utilized for box regression to enhance the generalization of the localization capability. We evaluate our method on the COCO and LVIS OVD benchmarks, outperforming the baseline (RegionCLIP) by 5.5% AP50and 5.8% AP on novel categories, respectively, achieving state-of-the-art performance.
Mingzhou He, Qingbo Wu 0001, King Ngi Ngan, Fanman Meng, Heqian Qiu, Hongliang Li 0001
VCIP7
2024 Proposal-level Correction Guided by CLIP for Few-shot Object Detection
abstract
Few-shot object detection aims at detecting previously unseen objects given only a few annotated samples. Most existing approaches treat the model obtained from the base training stage with abundant data as a container of prior knowledge that can be transferred to novel objects. Knowledge with similar properties is also contained in Contrastive Language-Image Pretraining (CLIP). In this paper, we utilize this external prior knowledge to generate proposal-level classification scores to improve the detection results. We notice that these scores can hardly reflect the quality of proposal localization, so we combine them with the ones from a conventional detector to obtain the ability to distinguish the background. Moreover, we propose a new score fusion module with regularization to alleviate the ambiguity of detection results generated by a trivial element-wise multiplication fusion method. To further improve the quality of classification scores in our proposed branch, we add learnable prompts to mitigate the inaccurate classification problem we observe. We conduct extensive experiments on the PASCAL VOC dataset and demonstrate the effectiveness of our approach.
Ruihang Wang, Taijin Zhao, Hefei Mei, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001
VCIP6
2024 VLM-guided Explicit-Implicit Complementary novel class semantic learning for few-shot object detection
Taijin Zhao, Heqian Qiu, Lanxiao Wang, Hefei Mei, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
Expert Syst. Appl.8
2024 Exploiting Inter-Sample Affinity for Knowability-Aware Universal Domain Adaptation
Yifan Wang 0020, Lin Zhang 0041, Ran Song 0001, Hongliang Li 0001, Paul L. Rosin, Wei Zhang 0021
Int. J. Comput. Vis.4
2024 Advancing zero-shot semantic segmentation through attribute correlations
Runtong Zhang, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, Hongliang Li 0001
Neurocomputing6
2024 Closed-Loop Training for Projected GAN
abstract
Projected GAN, a pre-trained GAN, has been found to perform well in generating images with only a few training samples. However, it struggles with extended training, which may lead to decreased performance over time. This is because the pre-trained discriminator consistently surpasses the generator, creating an unstable training environment. In this work, we propose a solution to this issue by introducing closed-loop control (CLC) into the dynamics of Projected GAN, stabilizing training, and improving generation performance. Our proposed method consistently reduces the Fréchet Inception Distance (FID) of the previous methods; for example, it reduces the FID of Projected GAN by 4.31 on the Obama dataset. Our finding is fundamental and can be used in other pre-trained GANs.
Jiangwei Zhao, Liang Zhang 0054, Lili Pan 0001, Hongliang Li 0001
IEEE Signal Process. Lett.4
2024 TridentCap: Image-Fact-Style Trident Semantic Framework for Stylized Image Captioning
abstract
Stylized image captioning (SIC) aims to generate captions with target style for images. The biggest challenge is that the collection and annotation of stylized data are pretty difficult and time-consuming. Most existing methods learn massive factual captions or additional stylized bookcorpus independently to assist in generating stylized caption, which ignore core relationships between existing image-fact-style trident data. In this paper, we propose a novel image-fact-style trident semantic framework TridentCap for stylized image captioning, which includes an image-fact semantic fusion encoder (SFE) and a trident stylization decoder (TSD). Unlike existing methods, we directly mine the core relationship in image-fact-style trident data and use factual semantic and image to build cross-modal semantic feature space, achieving the coherence between image and text. Specifically, SFE aims to learn the image-related prior language knowledge information from factual text and leverage fine-grained region-level semantic correlations of image and factual text to achieve cross-modal semantic information alignment and integration. TSD is designed to decouple the dual-source fused semantic feature based on the target style to achieve stylized caption generation. In addition, we design a pseudo labels filter (PLF) to obtain and expand massive image-fact-style trident data by building pseudo stylized annotations for all image-fact data in traditional caption datasets, which can further strengthen stylized caption learning. It is a generic algorithm to solve the problem of insufficient data and can be used into any existing stylized caption models. We conduct extensive experiments on SentiCap and FlickrStyle datasets, which achieve consistently improvement on almost all metrics. Our code will be released at: https://github.com/WangLanxiao/TridentCap_Code.
Lanxiao Wang, Heqian Qiu, Benliu Qiu, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Robust Unpaired Image Dehazing via Adversarial Deformation Constraint
abstract
Due to the flexible training requirement and the appealing generalization ability, unpaired image dehazing has received increasing attention in coping with real-world hazy images. However, most of the existing methods rely on the loose dehazing-hazing cycle constraint, which makes it hard to eliminate poor-quality dehazing results when using a powerful hazing network in the training process. To address this issue, this paper proposes a simple yet efficient Adversarial Deformation Constraint (ADC). More specifically, we sequentially perform two operations, i.e., dehazing and deformation, on a hazy image. In the training process, the dehazing branch is desired to be deformation-unaware, which requires that the output of these two operations remains constant regardless of their performing order. Adversarially, the deformation branch tends to maximize the difference in the outputs of these two operations when their performing orders are different. Through an additive image decomposition model, we verify that the ADC could regularize the solution space to push the dehazing error towards zero. Finally, by incorporating ADC into the common dehazing-hazing cycle constraint, we significantly improve the robustness of unpaired image dehazing. Experiments on multiple benchmark hazy image databases demonstrate the superiority of ADC over many state-of-the-art image dehazing methods. The source code of the proposed ADC-Net will be released on https://github.com/whrws/ADC-Net.
Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu
IEEE Trans. Circuits Syst. Video Technol.5
2024 Continual Cross-Domain Image Compression via Entropy Prior Guided Knowledge Distillation and Scalable Decoding
abstract
Learning based image compression has achieved impressive rate-distortion performance in recent years. However, due to the disposable learning strategy and rigid network architecture, existing methods perform poorly for compressing the images of different domains when they emerge with the expanding real-world applications, such as, natural, oil painting, medical images and so on. To cope with this open-world challenge, this paper proposes a continual cross-domain image compression method based on entropy prior guided knowledge distillation and scalable decoding network, which perform well in balancing the plasticity, stability and compatibility. Firstly, we generate pseudo-samples of old domains by reusing their entropy priors. These pseudo-samples serve as guides for knowledge distillation in the old domains, ensuring that the bit rate and reconstruction of the new model align with those of the old model. This approach assists the updated model in retaining its capability to compress and reconstruct old images. Secondly, we develop a scalable decoding network via dynamic pruning and masked recovery, which could effectively infer an old entropy decoder from the latestly updated model. It ensures that the updated model could decode image features from binary strings encoded by old entropy encoders. Experiments on five image datasets with different domains demonstrate the effectiveness of the proposed method and its superiority over representative continual learning methods. Code of the proposed method is available athttps://github.com/wuchenhaoo/Continual_Cross-domain_Image_Compression/.
Qingbo Wu 0001, Rui Ma 0030, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu
IEEE Trans. Circuits Syst. Video Technol.5
2024 Oriented-DINO: Angle Decoupling Prediction and Consistency Optimizing for Oriented Detection Transformer
abstract
Considering the arbitrary orientation of remote sensing objects, accurate angle prediction plays a crucial role in achieving precise oriented object detection (OOD) of aerial scenes. Existing transformer-based methods typically adopt an iterative refinement mechanism to update angle prediction and perform bipartite graph matching based on the combined matching costs. However, these methods may suffer from angle error accumulation across decoder layers and inconsistency between the L1 cost and the rotated intersection-of-union (IoU) cost, thus resulting in inaccurate angle prediction. To address these problems, this article proposes a novel transformer-based OOD method named Oriented-DINO (ODINO), which comprises three important components: error-mitigating angle decoupling prediction (EADP) module, nonlinear angle-conversion consistency optimizer (NACO), and query-driven diversity (QD) loss. To mitigate the angle error, the EADP module decouples angle prediction from the iterative box refinement process and uses independent branches to directly predict the angle. To address the issue of inconsistent matching, the NACO module uses a nonlinear function for angle conversion in matching cost calculation. This approach effectively alleviates the matching cost discrepancy in angle boundary case, while preserving the consistency in other instances. To avoid highly overlapped predictions triggered by similar queries, we introduce the QD loss to encourage the generation of diverse object queries, thus avoiding redundant predictions and enhancing prediction accuracy. Extensive experimental results demonstrate that our method achieves superior performance on OOD task.
Minjian Zhang 0003, Heqian Qiu, Lanxiao Wang, Haoyang Cheng, Taijin Zhao, Hongliang Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and Beyond
abstract
Few-shot segmentation (FSS) aims to segment the novel class with a few annotated images. Due to CLIP's advantages of aligning visual and textual information, the integration of CLIP can enhance the generalization ability of FSS model. However, even with the CLIP model, the existing CLIP-based FSS methods are still subject to the biased prediction towards base class, which is caused by the class-specific feature level interactions. To solve this issue, we propose a visual and textual Prior Guided Mask Assemble Network (PGMA-Net). It employs a class-agnostic mask assembly process to alleviate the bias, and formulates diverse tasks into a unified manner by assembling the prior through affinity. Specifically, the class-relevant textual and visual features are first transformed to class-agnostic prior in the form of probability map. Then, a Prior-Guided Mask Assemble Module (PGMAM) including multiple General Assemble Units (GAUs) is introduced. It considers diverse and plug-and-play interactions, such as visual-textual, inter- and intra-image, training-free, and high-order ones. Lastly, to ensure the class-agnostic ability, a Hierarchical Decoder with Channel-Drop Mechanism (HDCDM) is proposed to flexibly exploit the assembled masks and low-level features, without relying on any class-specific information. It achieves new state-of-the-art results in the FSS task, with mIoU of 77.6 on$\rm{PASCAL-}5^{i}$and 59.4 on$\rm{COCO-}20^{i}$in 1-shot scenario. Beyond this, we show that without extra re-training, the proposed PGMA-Net can solve bbox-level and cross-domain FSS, co-segmentation, zero-shot segmentation (ZSS) tasks, leading an any-shot segmentation framework capable of accommodating diverse weak or pixel annotations.
Fanman Meng, Runtong Zhang, Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Linfeng Xu 0001
IEEE Trans. Multim.5
2024 CrowdCaption++: Collective-Guided Crowd Scenes Captioning
abstract
Crowd scenes analysis plays an important role in various fields, including public security, smart cities, and intelligent transportation systems. However, traditional crowd scenes captioning methods mainly focus on a single and prominent crowd collective, which limits their ability to describe the different crowd collectives in complex crowd scenes. To address this issue, we propose a collective-guided crowd scenes captioning model (CrowdCaption++) to explore a more comprehensive and detailed description. We design a crowd features encoder (CFE) including double-query features encoder and foreground crowd features encoder, which uses double-query attention module (DQ-ATT) to capture more representative visual features and extracts foreground crowd features to avoid interference from background for collectives prediction. Moreover, we build a collective-guided captioning decoder (CCD) to generate captions of different crowd collectives without requiring extra alignment between crowd collectives and captions. To achieve this, we first design a crowd collectives predictor to identify multiple potential crowd collectives and create crowd collectives guidance information. Finally, we use the crowd collectives guidance information to merge useful visual features and further generate corresponding caption. We evaluate our approach on the latest crowd scenes dataset CrowdCaption and demonstrate that our model can achieve a comprehensive understanding and describe the different crowd collectives in complex crowd scenes.
Lanxiao Wang, Hongliang Li 0001, Minjian Zhang 0003, Heqian Qiu, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001
IEEE Trans. Multim.2
2024 Towards Continual Egocentric Activity Recognition: A Multi-Modal Egocentric Activity Dataset for Continual Learning
abstract
With the rapid development of wearable cameras, it is now feasible to considerably increase the collection of egocentric video for first-person visual perception. However, the development is hindered by a shortage of multi-modal egocentric activity datasets. Furthermore, the catastrophic forgetting problem of multimodal continual activity learning, as a branch of continual learning, has not been thoroughly explored, which makes accumulating a larger collection of multi-modal activity data more urgent. To address this shortage, we propose a multi-modal egocentric activity dataset for continual activity learning named UESTC-MMEA-CL in this paper. The dataset is collected using our self-developed glasses with a first-person camera and wearable sensors, and it contains synchronized data of video, accelerometers, and gyroscopes for 32 types of daily activities performed by 10 participants who wore our glasses. Statistical analysis of the sensor data is given to show the auxiliary effects of activity recognition. We report the results of egocentric activity recognition of three modalities (RGB, acceleration, and gyroscope) separately and jointly on a base network architecture. We thoroughly evaluated four baseline methods with different multimodal combinations to explore the catastrophic forgetting in continual learning on UESTC-MMEA-CL. We hope that the UESTC-MMEA-CL dataset can act as a facilitator for future studies on continual learning for first-person activity recognition in wearable applications. You can download preliminary data fromhttps://ivipclab.github.io/publication_uestc-mmea-cl/mmea-cl. The data is currently used to solve the problems of multimodal continual learning of activities.
Linfeng Xu 0001, Qingbo Wu 0001, Lili Pan 0001, Fanman Meng, Hongliang Li 0001, Chiyuan He, Hanxin Wang, Shaoxu Cheng
IEEE Trans. Multim.5
2024 InfoUCL: Learning Informative Representations for Unsupervised Continual Learning
abstract
Unsupervised continual learning (UCL) has made remarkable progress over the past two years, significantly expanding the application of continual learning (CL). However, existing UCL approaches have only focused on transferring continual strategies from supervised to unsupervised. They have overlooked the relationship issue between visual features and representational continuity. This work draws attention to the texture bias problem in existing UCL methods. To address this problem, we propose a new UCL framework called InfoUCL, in which we develop InfoDrop contrastive loss to guide continual learners to extract more informative shape features of objects and discard useless texture features simultaneously. The proposed InfoDrop contrastive loss is general and can be combined with various UCL methods. Extensive experiments on various benchmarks have demonstrated that our InfoUCL framework can lead to higher classification accuracy and superior robustness to catastrophic forgetting.
Liang Zhang 0054, Jiangwei Zhao, Qingbo Wu 0001, Lili Pan 0001, Hongliang Li 0001
IEEE Trans. Multim.5
2024 Learning Offset Probability Distribution for Accurate Object Detection
abstract
Object detection combines object classification and object localization problems. Current object detection methods heavily depend on regression networks to locate objects, which are optimized with various regression loss functions to predict offsets between candidate boxes and objects. However, these regression losses are difficult to assign the appropriate penalties for samples with large offset errors, resulting in suboptimal regression networks and inaccurate object offsets. In this article, we consider object location as offset bin classification problem, and propose a distance-aware offset bin classification network optimized with multiple binary cross entropy losses to learn various offset probability distribution, including single label distribution and distance-aware label distribution. On one hand, it provides gradient contributions for different samples based on the bounded probability instead of previous incalculable offset error. On the other hand, it explores the distance correlations between discrete offset bins to facilitate network learning. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin, in which the probability should be higher for the offset bin closer to the target offsets, and vice versa. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the precision of prediction. We conduct extensive experiments to evaluate the effectiveness of our method. In addition, our method can be conveniently and flexibly inserted into existing object detection methods, which consistently achieves a large gain based on popular anchor-based and anchor-free methods on the PASCAL VOC, MS-COCO, KITTI, and CrowdHuman datasets. Code will be released at: https://github.com/QiuHeqian/DBC .
Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 CafeBoost: Causal Feature Boost to Eliminate Task-Induced Bias for Class Incremental Learning
abstract
Continual learning requires a model to incrementally learn a sequence of tasks and aims to predict well on all the learned tasks so far, which notoriously suffers from the catastrophic forgetting problem. In this paper, we find a new type of bias appearing in continual learning, coined as task-induced bias. We place continual learning into a causal framework, based on which we find the task-induced bias is reduced naturally by two underlying mechanisms in task and domain incremental learning. However, these mechanisms do not exist in class incremental learning (CIL), in which each task contains a unique subset of classes. To eliminate the task-induced bias in CIL, we devise a causal intervention operation so as to cut off the causal path that causes the task-induced bias, and then implement it as a causal debias module that transforms biased features into unbiased ones. In addition, we propose a training pipeline to incorporate the novel module into existing methods and jointly optimize the entire architecture. Our overall approach does not rely on data replay, and is simple and convenient to plug into existing methods. Extensive empirical study on CIFAR-100 and ImageNet shows that our approach can improve accuracy and reduce forgetting of well-established methods by a large margin.
Benliu Qiu, Hongliang Li 0001, Haitao Wen, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Lili Pan 0001
CVPR2
2023 Incrementer: Transformer for Class-Incremental Semantic Segmentation with Knowledge Distillation Focusing on Old Class
abstract
Class-incremental semantic segmentation aims to incrementally learn new classes while maintaining the capability to segment old ones, and suffers catastrophic forgetting since the old-class labels are unavailable. Most existing methods are based on convolutional networks and prevent forgetting through knowledge distillation, which (1) need to add additional convolutional layers to predict new classes, and (2) ignore to distinguish different regions corresponding to old and new classes during knowledge distillation and roughly distill all the features, thus limiting the learning of new classes. Based on the above observations, we propose a new transformer framework for class-incremental semantic segmentation, dubbed Incrementer, which only needs to add new class tokens to the transformer decoder for new-class learning. Based on the Incrementer, we propose a new knowledge distillation scheme that focuses on the distillation in the old-class regions, which reduces the constraints of the old model on the new-class learning, thus improving the plasticity. Moreover, we propose a class deconfusion strategy to alleviate the overfitting to new classes and the confusion of similar classes. Our method is simple and effective, and extensive experiments show that our method outperforms the SOTAs by a large margin (5~15 absolute points boosts on both Pascal VOC and ADE20k). We hope that our Incrementer can serve as a new strong pipeline for class-incremental semantic segmentation.
Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Heqian Qiu, Lanxiao Wang
CVPR2
2023 Contrastive Continuity on Augmentation Stability Rehearsal for Continual Self-Supervised Learning
abstract
Self-supervised learning has attracted a lot of attention recently, which is able to learn powerful representations without any manual annotations. However, self-supervised learning needs to develop the ability to continuously learn to cope with a variety of real-world challenges, i.e., Continual Self-Supervised Learning (CSSL). Catastrophic forgetting is a notorious problem in CSSL, where the model tends to forget the learned knowledge. In practice, simple rehearsal or regularization will bring extra negative effects while alleviating catastrophic forgetting in CSSL, e.g., overfitting on the rehearsal samples or hindering the model from encoding fresh information. In order to address catastrophic forgetting without overfitting on the rehearsal samples, we propose Augmentation Stability Rehearsal (ASR) in this paper, which selects the most representative and discriminative samples by estimating the augmentation stability for rehearsal. Meanwhile, we design a matching strategy for ASR to dynamically update the rehearsal buffer. In addition, we further propose Contrastive Continuity on Augmentation Stability Rehearsal (C2ASR) based on ASR. We show that C2ASR is an upper bound of the Information Bottleneck (IB) principle, which suggests that C2ASR essentially preserves as much information shared among seen task streams as possible to prevent catastrophic forgetting and dismisses the redundant information between previous task streams and current task stream to free up the ability to encode fresh information. Our method obtains a great achievement compared with state-of-the-art CSSL methods on a variety of CSSL benchmarks.
Haoyang Cheng, Haitao Wen, Xiaoliang Zhang 0002, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001
ICCV6
2023 Optimizing Mode Connectivity for Class Incremental Learning
abstract
Class incremental learning (CIL) is one of the most challenging scenarios in continual learning. Existing work mainly focuses on strategies like memory replay, regularization, or dynamic architecture but ignores a crucial aspect: mode connectivity. Recent studies have shown that different minima can be connected by a low-loss valley, and ensembling over the valley shows improved performance and robustness. Motivated by this, we try to investigate the connectivity in CIL and find that the high-loss ridge exists along the linear connection between two adjacent continual minima. To dodge the ridge, we propose parameter-saving OPtimizing Connectivity (OPC) based on Fourier series and gradient projection for finding the low-loss path between minima. The optimized path provides infinite low-loss solutions. We further propose EOPC to ensemble points within a local bent cylinder to improve performance on learned tasks. Our scheme can serve as a plug-in unit, extensive experiments on CIFAR-100, ImageNet-100, and ImageNet-1K show consistent improvements when adapting EOPC to existing representative CIL methods. Our code is available at https://github.com/HaitaoWen/EOPC.
Haitao Wen, Haoyang Cheng, Heqian Qiu, Lanxiao Wang, Lili Pan 0001, Hongliang Li 0001
ICML6
2023 PTCP: Alleviate Layer Collapse in Pruning at Initialization via Parameter Threshold Compensation and Preservation
Xinpeng Hao, Shiyuan Tang, Heqian Qiu, Hefei Mei, Benliu Qiu, Chuanyang Gong, Hongliang Li 0001
ICONIP (11)8
2023 Novel-Registrable Weights and Region-Level Contrastive Learning for Incremental Few-shot Object Detection
Shiyuan Tang, Hefei Mei, Heqian Qiu, Xinpeng Hao, Taijin Zhao, Benliu Qiu, Haoyang Cheng, Chuanyang Gong, Hongliang Li 0001
ICONIP (11)10
2023 CFS: Character Feature Summarization Model for Real-time End-to-end Text Spotting
abstract
Most real-time end-to-end text spotting methods employ sequence models as their recognition heads. However, these models generate characters one by one, which is inefficient when there are many characters. To solve this problem, we propose a Character Feature Summarization (CFS) Model, which can predict fixed-length characters in parallel, regardless of length. Specifically, we propose a Character Feature Summarization Module (CFSM) consisting of a Global Feature Capture and a Historical Feature Summarizer to extract and summarize global character features, enabling getting characters by simple linear prediction. We use Multi-stage Testing, cascading multiple CFSMs to obtain multi-stage summarized global character features to obtain several predictions for better convergence. The Result Selector is used to select the most likely result. Experiments on the Total-Text dataset show that CFS achieves a 3.53% improvement on the "Full" while being 3.6 times faster than ABCNet v2’s head.
Chuanyang Gong, Heifei Mei, Heqian Qiu, Xinpeng Hao, Shiyuan Tang, Hongliang Li 0001
VCIP7
2023 ISM-Net: Mining incremental semantics for class incremental learning
Zihuan Qiu, Linfeng Xu 0001, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001
Neurocomputing6
2023 GFR: Generic feature representations for class incremental learning
abstract
Class incremental learning (CIL) aims to continuously learn new classes while maintaining discrimination for old classes with sequentially coming data. Due to the lack of old-class samples, existing CIL methods fail to learn discriminative representations for both old and new classes simultaneously, resulting in a severe performance drop in old classes, which is the well-known catastrophic forgetting phenomenon. Different from most existing works, we facilitate CIL by learning generic feature representations that perform well in seen and unseen classes. Specifically, we prove that representations with a substantial number of significant singular values benefit CIL via better old knowledge reservation. However, the overly uniform singular value spectrum will hurt the discrimination of current tasks. Furthermore, we propose that increasing the embedding dimension can enhance the number of significant singular values and validate this assumption from two perspectives: adopting different pooling techniques and devising a wider network. Meanwhile, we also prove that satisfactory current task accuracy and old knowledge reservation can be achieved simultaneously. Finally, the simple yet effective generic feature representation regulation (GFR) is devised and incorporated into two baselines. Extensive experiments are conducted on CIFAR100, ImageNet-Subset, and ImageNet. The results show that the proposed method boosts the performance of both baselines with a large margin (2.00%-9.58% on CIFAR100, 0.68%-7.10% on ImageNet-Subset and 1.18%-5.04% on ImageNet) which outperforms existing SOTAs.
Linfeng Xu 0001, Zihuan Qiu, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001
Neurocomputing6
2023 Disturbed Augmentation Invariance for Unsupervised Visual Representation Learning
abstract
Contrastive learning has gained great prominence recently, which achieves excellent performance by simple augmentation invariance. However, the simple contrastive pairs suffer from lacking of diversity due to the mechanical augmentation strategies. In this paper, we propose Disturbed Augmentation Invariance (DAI for abbreviation), which constructs disturbed contrastive pairs by generating appropriate disturbed views for each augmented view in the feature space to increase the diversity. In practice, we establish a multivariate normal distribution for each augmented view, whose mean is corresponding augmented view and covariance matrix is estimated from its nearest neighbors in the dataset. Then we sample random vectors from this distribution as the disturbed views to construct disturbed contrastive pairs. In order to avoid extra computational cost with the increase of disturbed contrastive pairs, we utilize an upper bound of the trivial disturbed augmentation invariance loss to construct the DAI loss. In addition, we propose Bottleneck version of Disturbed Augmentation Invariance (BDAI for abbreviation) inspired by the Information Bottleneck principle, which further refines the extracted information and learns a compact representation by additionally increasing the variance of the original contrastive pair. In order to make BDAI work effectively, we design a statistical strategy to control the balance between the amount of the information shared by all disturbed contrastive pairs and the compactness of the representation. Our approach gets a consistent improvement over the popular contrastive learning methods on a variety of downstream tasks, e.g. image classification, object detection and instance segmentation.
Haoyang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Heqian Qiu, Xiaoliang Zhang 0002, Fanman Meng, Taijin Zhao
IEEE Trans. Circuits Syst. Video Technol.2
2023 CrossDet++: Growing Crossline Representation for Object Detection
abstract
In object detection, precise object representation is a key factor to successfully classify and locate objects of an image. Existing methods usually use rectangular anchor boxes or a set of points to represent objects. However, these methods either introduce background noise or miss the continuous appearance information inside the object, and thus cause incorrect detection results. In this paper, we propose a novel anchor-free object detection network, called CrossDet++, which uses a set of growing crosslines along horizontal and vertical axes as object representations. An object can be flexibly represented as crosslines in different combinations, which inspires us to select the expressive crossline to effectively reduce the interference of noise. Meanwhile, the crossline representation takes into account the continuous adjacent object information, which is useful to enhance the discriminability of object features and find the object boundaries. Based on the learned crosslines, we propose an axis-query crossline growing module to adaptively capture features of crosslines and query surrounding pixels related to the line features for subsequent growing of crosslines. Their growing offsets and scales can be supervised by a decoupled regression mechanism, which limits the regression target to a specific direction for decreasing the optimization difficulty. During the training, we design a semantic-guided label assignment to emphasize the importance of crossline targets with higher semantic richness, further improving the detection performance. The experiment results demonstrate the effectiveness of our proposed method. Code can be available at:https://github.com/QiuHeqian/CrossDet.
Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Jianhua Cui, Zichen Song 0002, Lanxiao Wang, Minjian Zhang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2023 Cross-Modal Recurrent Semantic Comprehension for Referring Image Segmentation
abstract
Referring image segmentation aims to segment the target object from the image according to the description of language expression. Due to the diversity of language expressions, word sequences in different orders often express different semantic information. The previous methods focus more on matching different words to different visual regions in the image separately, ignoring the global semantic understanding of language expression based on the sequence structure. To address this problem, we redesign a new recurrent network structure for referring image segmentation, called Cross-Modal Recurrent Semantic Comprehension Network (CRSCNet), to obtain a more comprehensive global semantic understanding through iterative cross-modal semantic reasoning. Specifically, in each iteration, we first propose a Dynamic SepConv to extract relevant visual features guided by language and further propose Language Attentional Feature Modulation to improve the feature discriminability, then propose a Cross-Modal Semantic Reasoning module to perform global semantic reasoning by capturing both linguistic and visual information, and finally updates and corrects the visual features of the predicted object based on semantic information. Moreover, we further propose a Cross-Modal ASPP to capture richer visual information referred to in the global semantics of the language expression from larger receptive fields. Extensive experiments demonstrate that our proposed network significantly outperforms previous state-of-the-art methods on multiple datasets.
Chao Shang 0001, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, Taijin Zhao, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.2
2023 Task-Specific Loss for Robust Instance Segmentation With Noisy Class Labels
abstract
Deep learning methods have achieved significant progress in the presence of correctly annotated datasets in instance segmentation. However, object classes in large-scale datasets are sometimes ambiguous, which easily causes confusion. Besides, limited experience and knowledge of annotators can lead to mislabeled object semantic classes. To solve this issue, a novel method is proposed in this paper, which considers different roles of noisy class labels in different sub-tasks. Our method is based on two basic observations: firstly, the foreground-background annotation of a sample is correct even though its class label is noisy. Secondly, symmetric loss benefits the model robustness to noisy labels but harms the learning of hard samples, while cross entropy loss is the opposite. Based on the two basic observations, in the foreground-background sub-task, cross entropy loss is used to fully exploit correct gradient guidance. In the foreground-instance sub-task, symmetric loss is used to prevent incorrect gradient guidance provided by noisy class labels. Furthermore, we apply contrastive self-supervised loss to update features of all foreground, to compensate for insufficient guidance provided by partially correct labels especially in the highly noisy setting. Extensive experiments conducted with three popular datasets (i.e., Pascal VOC, Cityscapes and COCO) have demonstrated the effectiveness of our method in a wide range of noisy class label scenarios.
Longrong Yang, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.2
2023 DRDet: Dual-Angle Rotated Line Representation for Oriented Object Detection
abstract
In aerial scenes, oriented object detection is sensitive to the orientation of objects, which makes the formulation of orientation-aware object representation become a critical problem. Existing methods mostly adopt rectangle anchor or discrete points as object representation, which may lead to the feature aliasing between overlapping objects and ignore the orientation information of objects. To solve these issues, we propose a novel anchor-free oriented object detection network named DRDet, which adopts Dual-angle Rotated Lines (DRL) as object representation. Different from other object representations, DRL can adaptively rotate and extend to the boundary of the object according to its orientation and shape, which explicitly introduces the orientation information into the formulation of object representation. And it can adaptively cope with the geometric deformation of objects. Based on the dual-angle rotated lines, we design an Orientation-guided Feature Encoder (OFE) to encode discriminant object feature along each rotated line, respectively. Instead of encoding rectangle feature, the OFE module adopts line features for orientation-guided feature encoding, which can alleviate the feature aliasing between neighboring objects or background. To further enhance the flexibility of dual-angle rotated lines, we design a Dual-angle Decoder (DD) that predicts two angle offsets according to the orientation-guided feature and converts the angle offsets and regression offsets into dual-angle rotated line representation, which can help to guide the adaptive rotation of each rotated line, respectively. Our proposed method achieves consistent improvement on both DOTA and HRSC2016 datasets. Extensive experimental results verify the effectiveness of our method in oriented object detection.
Minjian Zhang 0003, Heqian Qiu, Hefei Mei, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Unsupervised Visual Representation Learning via Multi-Dimensional Relationship Alignment
abstract
Recently, contrastive learning based on augmentation invariance and instance discrimination has made great achievements, owing to its excellent ability to learn beneficial representations without any manual annotations. However, the natural similarity among instances conflicts with instance discrimination which treats each instance as a unique individual. In order to explore the natural relationship among instances and integrate it into contrastive learning, we propose a novel approach in this paper, Relationship Alignment (RA for abbreviation), which forces different augmented views of current batch instances to main a consistent relationship with other instances. In order to perform RA effectively in existing contrastive learning framework, we design an alternating optimization algorithm where the relationship exploration step and alignment step are optimized respectively. In addition, we add an equilibrium constraint for RA to avoid the degenerate solution, and introduce the expansion handler to make it approximately satisfied in practice. In order to better capture the complex relationship among instances, we additionally propose Multi-Dimensional Relationship Alignment (MDRA for abbreviation), which aims to explore the relationship from multiple dimensions. In practice, we decompose the final high-dimensional feature space into a cartesian product of several low-dimensional subspaces and perform RA in each subspace respectively. We validate the effectiveness of our approach on multiple self-supervised learning benchmarks and get consistent improvements compared with current popular contrastive learning methods. On the most commonly used ImageNet linear evaluation protocol, our RA obtains significant improvements over other methods, our MDRA gets further improvements based on RA to achieve the best performance. The source code of our approach will be released soon.
Haoyang Cheng, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Xiaoliang Zhang 0002, Fanman Meng, King Ngi Ngan
IEEE Trans. Image Process.2
2023 Forgetting to Remember: A Scalable Incremental Learning Framework for Cross-Task Blind Image Quality Assessment
abstract
Recent years have witnessed the great success of blind image quality assessment (BIQA) in various task-specific scenarios, which present invariable distortion types and evaluation criteria. However, due to the rigid structure and learning framework, they cannot apply to the cross-task BIQA scenario, where the distortion types and evaluation criteria keep changing in practical applications. This paper proposes a scalable incremental learning framework (SILF) that could sequentially conduct BIQA across multiple evaluation tasks with limited memory capacity. More specifically, we develop a dynamic parameter isolation strategy to sequentially update the task-specific parameter subsets, which are non-overlapped with each other. Each parameter subset is temporarily settled toRememberone evaluation preference toward its corresponding task, and the previously settled parameter subsets can be adaptively reused in the following BIQA to achieve better performance based on the task relevance. To suppress the unrestrained expansion of memory capacity in sequential tasks learning, we develop a scalable memory unit by gradually and selectively pruning unimportant neurons from previously settled parameter subsets, which enable us toForgetpart of previous experiences and free the limited memory capacity for adapting to the emerging new tasks. Extensive experiments on eleven IQA datasets demonstrate that our proposed method significantly outperforms the other state-of-the-art methods in cross-task BIQA. The source code of the proposed method is available atgithub.com/maruiperfect/SILF.
Rui Ma 0030, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
IEEE Trans. Multim.4
2023 What Happens in Crowd Scenes: A New Dataset About Crowd Scenes for Image Captioning
abstract
Making machines endowed with eyes and brains to effectively understand and analyze crowd scenes is of paramount importance for building a smart city to serve people. This is of far-reaching significance for the guidance of dense crowds and accident prevention, such as crowding and stampedes. As a typical multimodal scene understanding task, image captioning has always attracted widespread attention. However, crowd scene understanding captioning is rarely studied due to the unobtainability of related datasets. Therefore, it is difficult to know what happens in crowd scenes. In order to fill this research gap, we propose a crowd scenes caption dataset named CrowdCaption which has the advantages of crowd-topic scenes, comprehensive and complex caption descriptions, typical relationships and detailed grounding annotations. The complexity and diversity of the descriptions and the specificity of the crowd scenes make this dataset extremely challenging to most current methods. Thus, we propose a Multi-hierarchical Attribute Guided Crowd Caption Network (MAGC) based on crowd objects, actions, and status (such as position, dress, posture, etc.) aiming to generate crowd-specific detailed descriptions. We conduct extensive experiments on our CrowdCaption dataset, and our proposed method reaches the state-of-the-art (SoTA) performance. We hope the CrowdCaption dataset can assist future studies related to crowd scenes in the multimodal domain.
Lanxiao Wang, Hongliang Li 0001, Wenzhe Hu, Xiaoliang Zhang 0002, Heqian Qiu, Fanman Meng, Qingbo Wu 0001
IEEE Trans. Multim.2
2023 Efficient Geometry Surface Coding in V-PCC
abstract
In recent video-based point cloud compression (V-PCC), 3D point clouds are projected onto 2D images and compressed by High-Efficiency Video Coding (HEVC). However, HEVC was originally designed for natural visual signals, which is a suboptimal framework for point clouds. Therefore, there are still problems in geometry information compression in V-PCC: (1) The distortion based on the sum of squared error (SSE) in the existing rate-distortion optimization (RDO) is inconsistent with the geometric quality measurement; (2) The existing prediction cannot explore the fixed relationship between the corresponding far layer and near layer depth, which means that the far layer depth can be always not less than the corresponding near layer depth. In this paper, we present an efficient geometry surface coding (EGSC) method for V-PCC to address the problems. Firstly, an error projection (EP) model is designed to establish the relationship between the SSE-based distortion and the geometry quality metric. Secondly, an EP-based RDO is employed to improve the geometry information compression by estimating the point normals with gradients. Finally, an occupancy-map driven scheme is proposed to improve the prediction accuracy of merge modes. Experimental results show that the proposed method achieves an average of over 10% bit-rate saving compared with the V-PCC reference software.
Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, King Ngi Ngan, Weisi Lin
IEEE Trans. Multim.4
2023 Bias-Correction Feature Learner for Semi-Supervised Instance Segmentation
abstract
Instance segmentation is heavily reliant on large-scale annotated datasets to yield an ideal accuracy. However, annotated data are difficult to collect. To expand the annotated data, a straightforward idea is to introduce semi-supervised learning, which uses a trained model to obtain initial proposals on unlabeled images and then use initial proposals to generate pseudo labels. However, existing methods inevitably introduce the bias for the model learning, i.e., the foreground in initial low-confident proposals (low-confident foreground) is arbitrarily assigned as background. This bias makes the foreground and background closer in the feature space, which degenerates the model accuracy. To address this issue, this paper discards incorrect supervision and designs a bias-correction feature learner. Specifically, on the one hand, low-confident foreground does not participate in supervised learning. On the other hand, we extract possible foreground regions from all initial proposals to construct high-quality positive pairs which depict objects of the same category in contrastive learning. Then, positive pairs are pulled closer in the feature space. This helps models extract closely clustered foreground features. Experimental results demonstrate the effectiveness of our method on the public datasets (i.e., COCO, Cityscapes and Pascal VOC).
Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu, Linfeng Xu 0001
IEEE Trans. Multim.2
2022 RefCrowd: Grounding the Target in Crowd with Referring Expressions
abstract
Crowd understanding has aroused the widespread interest in vision domain due to its important practical significance. Unfortunately, there is no effort to explore crowd understanding in multi-modal domain that bridges natural language and computer vision. Referring expression comprehension (REF) is such a representative multi-modal task. Current REF studies focus more on grounding the target object from multiple distinctive categories in general scenarios. It is difficult to applied to complex real-world crowd understanding. To fill this gap, we propose a new challenging dataset, called RefCrowd, which towards looking for the target person in crowd with referring expressions. It not only requires to sufficiently mine natural language information, but also requires to carefully focus on subtle differences between the target and a crowd of persons with similar appearance, so as to realize fine-grained mapping from language to vision. Furthermore, we propose a Fine-grained Multi-modal Attribute Contrastive Network (FMAC) to deal with REF in crowd understanding. It first decomposes the intricate visual and language features into attribute-aware multi-modal features, and then captures discriminative but robustness fine-grained attribute features to effectively distinguish these subtle differences between similar persons. The proposed method outperforms existing state-of-the-art (SoTA) methods on our RefCrowd dataset and existing REF datasets. In addition, we implement an end-to-end REF toolbox for the deeper research in multi-modal domain. Our dataset and code can be available at: https://qiuheqian.github.io/datasets/refcrowd/.
Heqian Qiu, Hongliang Li 0001, Taijin Zhao, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng
ACM Multimedia2
2022 DE-CrossDet: Divisible and Extensible Crossline Representation for Object Detection
abstract
Object detection aims to localize and classify objects. Suitable object representation plays an important role in accurate detection. Because a complete crossline inevitably passes through the noise of backgrounds or other objects, object features directly extracted by the whole crossline are often confused. In this paper, we present a new feature extraction method, DE-Crossline, which can enhance the original crossline representation to capture more accurate object information. Specifically, we divide the crossline into several segments, each of which extracts the maximum activation key point respectively to reduce the impact of noise mentioned above. Furthermore, considering various shapes and sizes of objects, we design a Deformable Width Extension Module to learn a suitable width of each crossline, so as to capture richer object information. Extensive experiments prove the effectiveness of our proposed method. The total performance of our proposed detector can reach 49.0% AP, using ResNet-101 as backbone on the MS-COCO dataset.
Hefei Mei, Hongliang Li 0001, Heqian Qiu, Jianhua Cui, Longrong Yang
VCIP2
2022 Mining Regional Relation from Pixel-wise Annotation for Scene Parsing
abstract
Scene parsing is an important and challenging task in computer vision, which assigns semantic labels to each pixel in the entire scene. Existing scene parsing methods only utilize pixel-wise annotation as the supervision of neural network, thus, some similar categories are easy to be misclassified in the complex scenes without the utilization of regional relation. To tackle these above challenging problems, a Regional Relation Network (RRNet) is proposed in this paper, which aims to boost the scene parsing performance by mining regional relation from pixel-wise annotation. Specifically, the pixel-wise annotation is divided into a lot of fixed regions, so that intra- and inter-regional relation are able to be extracted as the supervision of network. We firstly design an intra-regional relation module to predict category distribution in each fixed region, which is helpful for reducing the misclassification phenomenon in regions. Secondly, an inter-regional relation module is proposed to learn the relationships among each region in scene images. With the guideline of relation information extracted from the ground truth, the network is able to learn more discriminative relation representations. To validate our proposed model, we conduct experiments on three typical datasets, including NYU-depth-v2, PASCAL-Context and ADE20k. The achieved competitive results on all three datasets demonstrate the effectiveness of our method.
Zichen Song 0002, Hongliang Li 0001, Heqian Qiu, Xiaoliang Zhang 0002
VCIP2
2022 Instance-level Context Attention Network for instance segmentation
Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Heqian Qiu, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan
Neurocomputing2
2022 Real-time panoptic segmentation with relationship between adjacent pixels and boundary prediction
Xiaoliang Zhang 0002, Hongliang Li 0001, Lanxiao Wang, Haoyang Cheng, Heqian Qiu, Wenzhe Hu, Fanman Meng, Qingbo Wu 0001
Neurocomputing2
2022 Category boundary re-decision by component labels to improve generation of class activation map
Runtong Zhang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan
Neurocomputing3
2022 Blind Image Deblurring via Superpixel Segmentation Prior
abstract
We present an effective blind image deblurring algorithm based on superpixel segmentation prior (SSP). The motivation of this work is an interesting observation that the more rough the segmentation boundaries are, the clearer the image will be. Intuitively, the blurry images have less image details, which results in more smooth segmentation boundaries. The clear images have more vivid textural details and obtain more rough boundaries. The segmentation roughness could be defined as the length of image segmentation boundaries. However, the segmentation boundary length is not differentiable, making it difficult to integrate into existing joint optimization framework. Therefore, we transform the segmentation boundary length into the segmentation entropy to guide the process of image deblurring. With the image becomes clearer, its boundary becomes more rough, while the segmentation entropy is much smaller. The analysis of relationship between segmentation entropy and segmentation boundaries is detailed. Benefiting from the convexity of segmentation entropy, we propose a novel algorithm by integrating half-quadratic split and gradient descent to alternately minimize energy function. Extensive experiments show that the proposed method achieves best performance with the state-of-the-art blind deblurring methods on natural and face image deblurring.
Bing Luo 0003, Zhongzhe Cheng, Guangrong Zhang, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 POS-Trends Dynamic-Aware Model for Video Caption
abstract
Video caption aims to generate descriptive sentences about the video, and the most critical problem is how to achieve accurate word prediction with standardized and coherent syntax structure, which requires the model to thoroughly understand video content and precisely map them into corresponding sentence components. Many existing methods usually fuse different video features into a single visual feature for generating sentences. However, they ignore the word dataset prior information in the annotations (such as Part-Of-Speech) and they also ignore the association between sentence components and types of visual features. To solve these problems, we propose a POS-trends dynamic-aware model (PDA) to fully exploit the word dataset prior information in the captions to predict POS tag, so as to assist generating captions. We propose a POS feature extraction (PFE) module to use different filters to extract different POS-trends features, predict POS tags and fuse visual features. Furthermore, we propose a visual-dynamic-aware (VDA) module to dynamically adjust the mapping way of words and supplement the visual information into the local features. The fusion features provide directional visual information to generate correct words, and the predicted POS tags to guide the decoding process to generate a more standardized and coherent syntax structure. A large number of experiments based on MSVD, MSR-VTT and VATEX demonstrated that our method outperforms the state-of-the-art methods in BLEU-4, ROUGE-L, METEOR, CIDEr. Code can be available at:https://github.com/WangLanxiao/PDA-for-video-caption.
Lanxiao Wang, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.2
2022 Occupancy Map Guided Fast Video-Based Dynamic Point Cloud Coding
abstract
In video-based dynamic point cloud compression (V-PCC), 3D point clouds are projected into patches, and then the patches are padded into 2D images suitable for the video compression framework. However, the patch projection-based method produces a large number of empty pixels; the far and near components are projected to generate different 2D images (video frames), respectively. As a result, the generated video is with high resolutions and double frame rates, so the V-PCC has huge computational complexity. This paper proposes an occupancy map guided fast V-PCC method. Firstly, the relationship between the prediction coding and block complexity is studied based on a local linear image gradient model. Secondly, according to the V-PCC strategies of patch projection and block generation, we investigate the differences of rate-distortion characteristics between different types of blocks, and the temporal correlations between the far and near layers. Finally, by taking advantage of the fact that occupancy maps can explicitly indicate the block types, we propose an occupancy map guided fast coding method, in which coding is performed on the different types of blocks. Experiments have tested typical dynamic point clouds, and shown that the proposed method achieves an average 43.66% time-saving at the cost of only 0.27% and 0.16% Bjontegaard Delta (BD) rate increment under the geometry Point-to-Point (D1) error and attribute Luma Peak-Signal-Noise-Ratio (PSNR), respectively.
Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.4
2022 Segmenting Beyond the Bounding Box for Instance Segmentation
abstract
Instance segmentation needs to locate all instances in an image correctly and segment each instance precisely. Currently, the most dominant methods for instance segmentation take object detection as a pre-task. However, they rely on the accuracy of object detection incredibly. If the pre-task cannot predict an accurate bounding box, the performance of instance segmentation will degenerate. In this paper, we present a novel method for instance segmentation to solve this problem, which is calledSegmentingBeyond theBoundingBox (S3B-Net). Our S3B-Net designs a sub-network to help instance segmentation methods based on object detection to segment the part of an instance beyond the bounding box. Specifically, the sub-network first predicts a two-dimensional pixel embedding for each pixel. Then, the Gaussian function is employed to calculate a pixel’s probability belongs to a corresponding instance according to the two-dimensional pixel embedding. Finally, the output of the sub-network combines with the output of instance segmentation based on object detection to generate a more precise instance mask. Our sub-network can easily extend on the existing instance segmentation method based on object detection to segment instance beyond the bounding box. We do our experiments on dominant instance segmentation datasets, such as the COCO dataset and Cityscapes dataset. The results show that our method can achieve 6.8 points gain compared with the baseline Mask R-CNN with ResNet-50-FPN in Cityscapes datasets, and 1.7 points gain with ResNet-101-FPN-DCN in COCO datasets. Our S3B-Net outperforms the previous state-of-the-art instance segmentation method, which proves our method is competitive. The source code of our method will be made available.
Xiaoliang Zhang 0002, Hongliang Li 0001, Fanman Meng, Zichen Song 0002, Linfeng Xu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Bal-R$^2$CNN: High Quality Recurrent Object Detection With Balance Optimization
abstract
It is a common practice to refine object detection results using recurrent detection paradigm. We evaluate the recurrent detection on Faster R-CNN, but the improvement is far away from expected. We consider that the performance bottleneck is fromimbalance optimizationcaused by the biased distribution of training data. Low-IoU-skewed RPN proposals could suppress the contribution of High-IoU examples at the training stage. Besides, data imbalance and statistical discrepancy on regression targets between low-IoU and high-IoU examples are not considered in the regression task; this design could impede localization quality. In this work, we propose Bal-R$^2$CNN for high-quality recurrent object detection. There are two new components in Bal-R$^2$CNN.Self-iteration box samplingcollects object boxes from recurrent steps and increases the number of high-IoU training examples.IoU-sensitive bounding-box regressionsends proposal boxes with different IoUs to specified regression branches for more accurate bounding-box prediction. Both two new components could inducebalanced optimizationand be helpful. With the resulting Bal-R$^2$CNN detector, evaluation on PASCAL VOC and MSCOCO reveal that our method has a significant improvement on the existing solution and could reach a better performance than several state-of-the-art methods.
Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu
IEEE Trans. Multim.2
2021 CrossDet: Crossline Representation for Object Detection
abstract
Object detection aims to accurately locate and classify objects in an image, which requires precise object representations. Existing methods usually use rectangular anchor boxes or a set of points to represent objects. However, these methods either introduce background noise or miss the continuous appearance information inside the object, and thus cause incorrect detection results. In this paper, we propose a novel anchor-free object detection network, called Cross-Det, which uses a set of growing cross lines along horizontal and vertical axes as object representations. An object can be flexibly represented as cross lines in different combinations. It not only can effectively reduce the interference of noise, but also take into account the continuous object information, which is useful to enhance the discriminability of object features and find the object boundaries. Based on the learned cross lines, we propose a crossline extraction module to adaptively capture features of cross lines. Furthermore, we design a decoupled regression mechanism to regress the localization along the horizontal and vertical directions respectively, which helps to decrease the optimization difficulty because the optimization space is limited to a specific direction. Our method achieves consistently improvement on the PASCAL VOC and MS-COCO datasets. The experiment results demonstrate the effectiveness of our proposed method. Code can be available at: https://github.com/QiuHeqian/CrossDet.
Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Jianhua Cui, Zichen Song 0002, Lanxiao Wang, Minjian Zhang 0003
ICCV2
2021 Remember and Reuse: Cross-Task Blind Image Quality Assessment via Relevance-aware Incremental Learning
abstract
Existing blind image quality assessment (BIQA) methods have made great progress in various task-specific applications, including the synthetic, authentic, or over-enhanced distortion evaluations. However, limited by the static model and once-for-all learning strategy, they failed to perform the cross-task evaluations in many practical applications, where diverse evaluation criteria and distortion types are constantly emerging. To address this issue, in this paper, we propose a dynamic Remember and Reuse (R&R) network, which efficiently performs the cross-task BIQA based on a novel relevance-aware incremental learning strategy. Given multiple evaluation tasks across different distortion types or databases, our R&R network sequentially updates the parameters for every task one by one. After each update step, part of task-specific parameters is settled, which ensures R&R Remembers their dedicated evaluation preferences. The remaining parameters are pruned for the dynamic usage of the subsequent tasks. To further exploit the correlation between different tasks, we feed the training data of a new task to previously settled parameters. Better prediction accuracy is considered as higher task relevance and vice versa. Then, we selectively Reuse parts of previously settled parameters, whose proportion is adaptively determined by the task relevance. Extensive experiments show that the proposed method efficiently achieves the cross-task BIQA without catastrophic forgetting, and significantly outperforms many state-of-the-art methods. Code is available at https://github.com/maruiperfect/R-R-Net.
Rui Ma 0030, Hanxiao Luo, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
ACM Multimedia5
2021 Few-Shot Segmentation via Complementary Prototype Learning and Cascaded Refinement
Hanxiao Luo, Hui Li 0080, Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Fanman Meng, Linfeng Xu 0001
PRCV (4)4
2021 Hierarchical class grouping with orthogonal constraint for class activation map generation
Fanman Meng, Kaixu Huang, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan
Neural Comput. Appl.3
2021 High-Quality R-CNN Object Detection Using Multi-Path Detection Calibration Network
abstract
Object proposals are used in two-stage detectors, such as R-CNN, to generate detection results, including category predictions and refined bounding-boxes. As a result, classification scores are assigned to refined bounding-boxes rather than object proposals. However, this procedure ignores the discrepancy of data distribution between object proposals and refined bounding-boxes. We consider this discrepancy could limit the detection accuracy. Specifically, the foreground/background imbalance on object proposals and inaccurate information from low-IoU proposals could hinder the category prediction. In this paper, we propose a detector called the Multi-Path Detection Calibration Network (PDC-Net) to address this problem. The key idea behind PDC-Net is calibrating detection results from R-CNN by considering the statistical discrepancy between object proposals and refined bounding-boxes. PDC-Net is built on Faster R-CNN. The core component in PDC-Net is the multi-path detection head, in which the base detector (from Faster R-CNN) generates detection results from object proposals and multiple calibration detectors fix incorrect outputs from the base detector using refined bounding-boxes. Experiments reveal that PDC-Net can boost detection results. Our method could reach 83.1% and 43.3% mAP respectively on PASCAL VOC and MSCOCO benchmarks, which is comparable to several state-of-the-art methods.
Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan, Linfeng Xu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Robust Texture Description Using Local Grouped Order Pattern and Non-Local Binary Pattern
abstract
Local binary pattern (LBP) and its many variants have shown effectiveness for texture classification. However, most of these LBP methods focus on encoding local intensity differences between a central pixel and its neighboring sampling points and consequently have two major problems: 1) they are unable to describe the intensity order relationships among neighboring sampling points, and 2) they fail to capture long-range pixel interactions that take place outside a compact neighborhood. In view of these problems, in this paper we propose two novel operators, called local grouped order pattern (LGOP) and non-local binary pattern (NLBP), for texture description. For the first problem, LGOP groups the neighboring sampling points by referring to a dominant direction and encodes the groupwise intensity order relationships. For the second problem, NLBP computes several anchors based on global image statistics and progressively encodes non-local intensity differences between the neighboring sampling points and anchors. Finally, we combine LGOP and NLBP via central pixel encoding to construct discriminative histogram features as texture descriptor LGONBP. Experiments on four texture benchmark databases (i.e., Outex, CUReT, UMD and KTH-TIPS) demonstrate the superiority of LGONBP over state-of-the-art LBP variants for texture classification under both noise-free and noisy conditions. The code is available athttps://github.com/stc-cqupt/LGONBP.
Tiecheng Song, Jie Feng 0007, Lin Luo 0007, Chenqiang Gao, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Non-Homogeneous Haze Removal via Artificial Scene Prior and Bidimensional Graph Reasoning
abstract
Due to the lack of natural scene and haze prior information, it is greatly challenging to completely remove the haze from a single image without distorting its visual content. Fortunately, the real-world haze usually presents non-homogeneous distribution, which provides us with many valuable clues in partial well-preserved regions. In this paper, we propose a Non-Homogeneous Haze Removal Network (NHRN) via artificial scene prior and bidimensional graph reasoning. Firstly, we employ the gamma correction iteratively to simulate artificial multiple shots under different exposure conditions, whose haze degrees are different and enrich the underlying scene prior. Secondly, beyond utilizing the local neighboring relationship, we build a bidimensional graph reasoning module to conduct non-local filtering in the spatial and channel dimensions of feature maps, which models their long-range dependency and propagates the natural scene prior between the well-preserved nodes and the nodes contaminated by haze. To the best of our knowledge, this is the first exploration to remove non-homogeneous haze via the graph reasoning based framework. We evaluate our method on different benchmark datasets. The results demonstrate that our method achieves superior performance over many state-of-the-art algorithms for both the single image dehazing and hazy image understanding tasks. The source code of the proposed NHRN is available on https://github.com/whrws/NHRNet.
Qingbo Wu 0001, Hui Li 0080, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
IEEE Trans. Image Process.5
2021 Query Reconstruction Network for Referring Expression Image Segmentation
abstract
Referring expression image segmentation aims at segmenting out the object described by a natural language query. Due to the diversity of visual content and language descriptions, it is very challenging to accurately model the correspondence between the vision and language, which inevitably produces some undesired segmentation objects from the queries. In this paper, we propose a query reconstruction network (QRN) to build more consistent corresponding relations between the language queries and object segmentation results. QRN not only generates segmentations from the queries and images but also reversely reconstructs the queries from the segmentations and the images. Through query reconstruction, QRN can confirm the vision-language consistency between the segmentations and queries. In the inference stage, for inconsistent segmentations and queries, we propose an iterative segmentation correction (ISC) method to correct them. ISC takes the difference between the reconstructed and input queries as a loss to optimize the proposed QRN. Then, the proposed QRN can generate new segmentations and queries. By iterative optimization, the segmentations can be gradually corrected. Extensive experiments on four referring expression image segmentation databases demonstrate the effectiveness of the proposed method.
Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan
IEEE Trans. Multim.2
2020 Offset Bin Classification Network for Accurate Object Detection
abstract
Object detection combines object classification and object localization problems. Most existing object detection methods usually locate objects by leveraging regression networks trained with Smooth L1loss function to predict offsets between candidate boxes and objects. However, this loss function applies the same penalties on different samples with large errors, which results in suboptimal regression networks and inaccurate offsets. In this paper, we propose an offset bin classification network optimized with cross entropy loss to predict more accurate offsets. It not only provides different penalties for different samples but also avoids the gradient explosion problem caused by the samples with large errors. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the prediction precision. Extensive experiments on the PASCAL VOC and MS-COCO datasets demonstrate the effectiveness of our proposed method. Our method outperforms the baseline methods by a large margin.
Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi
CVPR2
2020 Learning with Noisy Class Labels for Instance Segmentation
Longrong Yang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Qishang Cheng
ECCV (14)3
2020 Single Image Dehazing Via Artificial Multiple Shots And Multidimensional Context
abstract
The main challenge for single image dehazing is the lack of effective prior information for restoration. To address this issue, in this paper, we propose to generate artificial multiple shots for simulating the images captured under different haze degrees, and two context reasoning modules are developed to describe the relationship across different spatial regions and artificial shots. It brings two benefits in the inhomogeneous haze distribution. First, within one shot, the regions occluded in one location could be recovered with the help of other clear regions, which share the similar structures. Second, for the same spatial location, the regions distorted in one shot could be restored by means of other shots with clear content. We evaluate our method on different benchmark datasets. The results demonstrate that our method achieves superior performance over many state-of-the-art dehazing algorithms.
Qingbo Wu 0001, Hui Li 0080, King Ngi Ngan, Hongliang Li 0001, Fanman Meng
ICIP5
2020 Region Adaptive Two-Shot Network For Single Image Dehazing
abstract
Existing single image dehazing methods typically adopt a one-shot strategy by indiscriminately applying the same filters to all local regions, which easily cause under-/over-dehazing across different regions by ignoring the inhomogeneity and asymmetry of illumination and detail distortions. In this paper, we propose a region adaptive two-shot network (RATNet) to address this issue. In the first shot, a lightweight subnetwork is utilized to conduct the regular global filtering, which could remove parts of haze but also distort some image details. In the second shot, a two-branch subnetwork is developed to restore the illumination and details of the initially renovated image respectively. The final dehazed image is obtained by fusing the outputs of the previous two branches, whose region-variant weights are adaptively learned by minimizing the difference between the haze-free image and our fused result. Experiments on four dehazing benchmark datasets show that our RATNet significantly outperforms many state-of-the-art dehazing approaches.
Hui Li 0080, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng
ICME4
2020 CODAN: Counting-driven Attention Network for Vehicle Detection in Congested Scenes
abstract
Although recent object detectors have shown excellent performance for vehicle detection, they are incompetent for scenarios with a relatively large number of vehicles. In this paper, we explore the dense vehicle detection given the number of vehicles. Existing crowd counting methods cannot directly applied for dense vehicle detection due to insufficient description of density map, and the lack of effective constraint for mining the spatial awareness of dense vehicles. Inspired by these observations, a conceptually simple yet efficient framework, called CODAN, is proposed for dense vehicle detection. The proposed approach is composed of three major components: (i) an efficient strategy for generating multi-scale density maps (MDM) is designed to represent the vehicle counting, which can capture the global semantics and spatial information of dense vehicles, (ii) a multi-branch attention module (MAM) is proposed to bridging the gap between object counting and vehicle detection framework, (iii) with the well-designed density maps as explicit supervision, an effective counting-awareness loss (C-Loss) is employed to guide the attention learning by building the pixel-level constrain. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods. The impressive results indicate that vehicle detection and counting can be mutually supportive, which is an important and meaningful finding.
Wei Li 0110, Zhenting Wang, Xiao Wu 0001, Ji Zhang 0027, Qiang Peng, Hongliang Li 0001
ACM Multimedia6
2020 Language-Aware Fine-Grained Object Representation for Referring Expression Comprehension
abstract
Referring expression comprehension expects to accurately locate an object described by a language expression, which requires precise language-aware visual object representations. However, existing methods usually use rectangular object representations, such as object proposal regions and grid regions. They ignore some fine-grained object information like shapes and poses, which are often described in language expressions and important to localize objects. Additionally, rectangular object regions usually contain background contents and irrelevant foreground features, which also decrease the localization performance. To address these problems, we propose a language-aware deformable convolution model (LDC) to learn language-aware fine-grained object representations. Rather than extracting rectangular object representations, LDC adaptively samples a set of key points based on the image and language to represent objects. This type of object representations can capture more fine-grained object information (e.g., shapes and poses) and suppress noises in accordance with language and thus, boosts the object localization performance. Based on the language-aware fine-grained object representation, we next design a bidirectional interaction model (BIM) that leverages a modified co-attention mechanism to build cross-modal bidirectional interactions to further improve the language and object representations. Furthermore, we propose a hierarchical fine-grained representation network (HFRN) to learn language-aware fine-grained object representations and cross-modal bidirectional interactions at local word level and global sentence level, respectively. Our proposed method outperforms the state-of-the-art methods on the RefCOCO, RefCOCO+ and RefCOCOg datasets.
Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Hengcan Shi, Taijin Zhao, King Ngi Ngan
ACM Multimedia2
2020 Multi-stage Tag Guidance Network in Video Caption
abstract
Recently, video caption plays an important role in computer vision tasks. We participate in Pre-training for Video Captioning Challenge which aims to produce at least one sentence for each challenge video based on the pretraining models. In this work, we propose a tag guidance module to learn a representation which can better build the interaction in cross-modal between visual content and textual sentences. First, we utilize three types of features extraction networks to fully capture the information of 2D, 3D and object information. Second, to prevent overfitting and time issues, the entire process of training is divided into two stages. The first stage trains all data, and the second stage introduces a random dropout. Furthermore, we train a CNN-based network to pick out the best candidate results. In summary, we were ranked third place in Pre-training for Video Captioning Challenge which proved the effectiveness of our model.
Lanxiao Wang, Chao Shang 0001, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001
ACM Multimedia6
2020 A multi-scale language embedding network for proposal-free referring expression comprehension
abstract
Referring expression comprehension (REC) is a task that aims to find the location of an object specified by a language expression. Current solutions for REC can be classified into proposal-based methods and proposal-free methods. Proposal-free methods are popular recently because of its flexibility and lightness. Nevertheless, existing proposal-free works give little consideration to visual context. As REC is a context sensitive task, it is hard for current proposal-free methods to comprehend expressions that describe objects by the relative position with surrounding things. In this paper, we propose a multi-scale language embedding network for REC. Our method adopts the proposal-free structure, which directly feeds fused visual-language features into a detection head to predict the bounding box of the target. In the fusion process, we propose a grid fusion module and a grid-context fusion module to compute the similarity between language features and visual features in different size regions. Meanwhile, we extra add fully interacted vision-language information and position information to strength the feature fusion. This novel fusion strategy can help to utilize context flexibly therefore the network can deal with varied expressions, especially expressions that describe objects by things around. Our proposed method outperforms the state-of-the-art methods on Refcoco, Refcoco+ and Refcocog datasets.
Taijin Zhao, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, King Ngi Ngan
MMAsia2
2020 A New Local Transformation Module for Few-Shot Segmentation
Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Xiaolong Xu 0004
MMM (2)3
2020 Haze-robust image understanding via context-aware deep feature refinement
abstract
Image understanding under the foggy scene is greatly challenging due to inhomogeneous visibility deterioration. Although various image dehazing methods have been proposed, they usually aim to improve image visibility (such as, PSNR/SSIM) in the pixel space rather than the feature space, which is critical for the perception of computer vision. Due to this mismatch, existing dehazing methods are limited or even adverse in facilitating the foggy scene understanding. In this paper, we propose a generalized deep feature refinement module to minimize the difference between clear images and hazy images in the feature space. It is consistent with the computer perception and can be embedded into existing detection or segmentation backbones for joint optimization. Our feature refinement module is built upon the graph convolutional network, which is favorable in capturing the contextual information and beneficial for distinguishing different semantic objects. We validate our method on the detection and segmentation tasks under foggy scenes. Extensive experimental results show that our method outperforms the state-of-the-art dehazing based pretreatments and the fine-tuning results on hazy images.
Hui Li 0080, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
MMSP5
2020 A Unified Single Image De-raining Model via Region Adaptive Coupled Network
abstract
Single image de-raining is quite challenging due to the diversity of rain types and inhomogeneous distributions of rainwater. By means of dedicated models and constraints, existing methods perform well for specific rain type. However, their generalization capability is highly limited as well. In this paper, we propose a unified de-raining model by selectively fusing the clean background of the input rain image and the well restored regions occluded by various rains. This is achieved by our region adaptive coupled network (RACN), whose two branches integrate the features of each other in different layers to jointly generate the spatial-variant weight and restored image respectively. On the one hand, the weight branch could lead the restoration branch to focus on the regions with higher contributions for de-raining. On the other hand, the restoration branch could guide the weight branch to keep off the regions with over-/under-filtering risks. Extensive experiments show that our method outperforms many state-of-the-art de-raining algorithms on diverse rain types including the rain streak, raindrop and rain-mist.
Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
VCIP4
2020 A New Bounding Box based Pseudo Annotation Generation Method for Semantic Segmentation
abstract
This paper proposes a fusion-based method to generate pseudo-annotations from bounding boxes for semantic segmentation. The idea is to first generate diverse foreground masks by multiple bounding box segmentation methods, and then combine these masks to generate pseudo-annotations. Existing methods generate foreground masks from bounding boxes by classical segmentation methods driving by low-level features and own local information, which is hard to generate accurate and diverse results for the fusion. Different from the traditional methods, multiple class-agnostic models are modeled to learn the objectiveness cues by using existing labeled pixel-level annotations and then to fuse. Firstly, the classical Fully Convolutional Network (FCN) that densely predicts the pixels' labels is used. Then, two new sparse prediction based class-agnostic models are proposed, which simplify the segmentation task as sparsely predicting the boundary points through predicting the distance from the bounding box border to the object boundary in Cartesian Coordinate System and the Polar Coordinate System, respectively. Finally, a voting-based strategy is proposed to combine these segmentation results to form better pseudo-annotations. We conduct experiments on PASCAL VOC 2012 dataset. The mIoU of the proposed method is 68.7%, which outperforms the state-of-the-art method by 1.9%.
Xiaolong Xu 0004, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan
VCIP3
2020 Mono is Enough: Instance Segmentation from Single Annotated Sample
abstract
With the help of various Deep Neural Networks, instance segmentation has achieved significant progress. How-ever, these successes are heavily reliant on large-scale manually annotated samples, which are extremely time-consuming and expensive. To address this issue, we propose a highly efficient anisotropic data augmentation method, which generates high quality training data from a single manually annotated sample. Instead of equivalently modifying foreground and background like traditional data augmentation methods, we focus on enriching the diversities of foreground appearance and positional relation between foreground and background, which are beneficial for the classification and localization sub-tasks respectively. All foreground instances of the source annotated sample undergo various rotation, brightness change, rescale, distortion and frequency-component mixup (FCM). Then, these modified instances are randomly embedded into background, which serve as new training samples. Experiments on Cityscapes dataset show that our method significantly outperforms traditional data augmentation methods.
Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan
VCIP2
2020 Mining Larger Class Activation Map with Common Attribute Labels
abstract
Class Activation Map (CAM) is the visualization of target regions generated from classification networks. However, classification network trained by class-level labels only has high responses to a few features of objects and thus the network cannot discriminate the whole target. We think that original labels used in classification tasks are not enough to describe all features of the objects. If we annotate more detailed labels like class-agnostic attribute labels for each image, the network may be able to mine larger CAM. Motivated by this idea, we propose and design common attribute labels, which are lower-level labels summarized from original image-level categories to describe more details of the target. Moreover, it should be emphasized that our proposed labels have good generalization on unknown categories since attributes (such as head, body, etc.) in some categories (such as dog, cat, etc.) are common and class-agnostic. That is why we call our proposed labels as common attribute labels, which are lower-level and more general compared with traditional labels. We finish the annotation work based on the PASCAL VOC2012 dataset and design a new architecture to successfully classify these common attribute labels. Then after fusing features of attribute labels into original categories, our network can mine larger CAMs of objects. Our method achieves better CAM results in visual and higher evaluation scores compared with traditional methods.
Runtong Zhang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan
VCIP3
2020 Hybrid-loss supervision for deep neural network
Qishang Cheng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan
Neurocomputing2
2020 Discriminative deep metric learning for asymmetric discrete hashing
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan
Neurocomputing2
2020 Parametric Deformable Exponential Linear Units for deep neural networks
Qishang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Lei Ma 0004, King Ngi Ngan
Neural Networks2
2020 Group Maximum Differentiation Competition: Model Comparison with Few Samples
abstract
In many science and engineering fields that require computational models to predict certain physical quantities, we are often faced with the selection of the best model under the constraint that only a small sample set can be physically measured. One such example is the prediction of human perception of visual quality, where sample images live in a high dimensional space with enormous content variations. We propose a new methodology for model comparison named group maximum differentiation (gMAD) competition. Given multiple computational models, gMAD maximizes the chances of falsifying a "defender" model using the rest models as "attackers". It exploits the sample space to find sample pairs that maximally differentiate the attackers while holding the defender fixed. Based on the results of the attacking-defending game, we introduce two measures, aggressiveness and resistance, to summarize the performance of each model at attacking other models and defending attacks from other models, respectively. We demonstrate the gMAD competition using three examples-image quality, image aesthetics, and streaming video quality-of-experience. Although these examples focus on visually discriminable quantities, the gMAD methodology can be extended to many other fields, and is especially useful when the sample space is large, the physical measurement is expensive and the cost of computational prediction is low.
Kede Ma, Zhengfang Duanmu, Zhou Wang 0001, Qingbo Wu 0001, Wentao Liu 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006
IEEE Trans. Pattern Anal. Mach. Intell.7
2020 Guest Editorial Introduction to the Special Section on Intelligent Visual Content Analysis and Understanding
abstract
Visual content analysis and understanding attract tremendous attention because of its potentially wide range of applications including human activity analysis, automated photo face tagging, multicamera tracking, crowded counting, and biometric security. With recent progress in end-to-end differentiable learning, the accuracy of algorithms has been significantly improved and even outperforms humans in some tasks. In addition, multimodality methods, targeting on making full use of various visual data sources, are further investigated. These developments contribute to the innovations of two core modules for a typical intelligent vision system, i.e., image and video description and recognition, which are critical for the success of the visual content analysis and understanding in more complex and challenging open world.
Hongliang Li 0001, Lu Fang 0001, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 HeadNet: An End-to-End Adaptive Relational Network for Head Detection
abstract
Head detection plays an important role in localizing and identifying persons from visual data. Most existing methods treat head detection as a specific form of object detection. Head detection is nontrivial due to the considerable difficulty in building the local and global information under conditions of unconstrained pose and orientation. To address these issues, this paper presents an effective adaptive relational network to capture context information, which is greatly helpful to suppress missed detection. We show that the fundamental contextual properties, such as the global shape priors from different heads and the local adjacent relationship between the head and shoulders, can be systematically quantified by visual operators. Specifically, we propose a two-step search algorithm to quantify the global intergroup conflict with adaptive scale, pose and viewpoint. Meanwhile, a structured feature module is introduced to capture the local relation of intraindividual stability. Finally, the global priors and local relation are integrated seamlessly into a single-stage head detector that is end-to-end trainable. An extensive ablation analysis demonstrates the effectiveness of our approach. We achieve state-of-the-art results on two challenging datasets, i.e., HollywoodHeads and Brainwash.
Wei Li 0110, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.2
2020 Weakly Supervised Semantic Segmentation by a Class-Level Multiple Group Cosegmentation and Foreground Fusion Strategy
abstract
Weakly supervised semantic segmentation uses image-level labels to extract object regions. The existing methods focus on efficiently training CNN-based segmentation networks using the image-level labels. In contrast to the existing methods, this paper proposes a new fusion-based method, which first segments the foregrounds of each image by multiple group cosegmentation and then generates the semantic segmentation by combining the foregrounds. Specifically, a new CNN-based multiple group cosegmentation network is first proposed to segment foregrounds employing two cues, the discriminative cue and the local-to-global cue. Then, the fusion method is proposed to simply perform semantic segmentation based on the multiple group cosegmentation results. Experiments on the PASCAL VOC 2012 and MS COCO 2017 datasets demonstrate the effectiveness of the proposed method with mIoU values that are obviously larger than those of the existing methods.
Fanman Meng, Kunming Luo, Hongliang Li 0001, Qingbo Wu 0001, Xiaolong Xu 0004
IEEE Trans. Circuits Syst. Video Technol.3
2020 Subjective and Objective De-Raining Quality Assessment Towards Authentic Rain Image
abstract
Images acquired by outdoor vision systems easily suffer poor visibility and annoying interference due to the rainy weather, which brings great challenge for accurately understanding and describing the visual contents. Recent researches have devoted great efforts on the task of rain removal for improving the image visibility. However, there is very few exploration about the quality assessment of de-rained image, even it is crucial for accurately measuring the performance of various de-raining algorithms. In this paper, we first create a de-raining quality assessment (DQA) database that collects 206 authentic rain images and their de-rained versions produced by 6 representative single image rain removal algorithms. Then, a subjective study is conducted on our DQA database, which collects the subject-rated scores of all de-rained images. To quantitatively measure the quality of de-rained image with non-uniform artifacts, we propose a bi-directional feature embedding network (B-FEN) which integrates the features of global perception and local difference together. Experiments confirm that the proposed method significantly outperforms many existing universal blind image quality assessment models. To help the research towards perceptually preferred de-raining algorithm, we will publicly release our DQA database and B-FEN source code on https://github.com/wqb-uestc.
Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Hierarchical Context Features Embedding for Object Detection
abstract
Pixel-level segmentation has been widely used to improve object detection. Most of the existing methods refine detection features by adding the constraint of the segmentation branch or by simply embedding high-level segmentation features into detection features within the local receptive field. However, noisy segmentation features are unavoidable in real-word applications and can easily cause false positives. To address this problem, we propose a novel hierarchical context embedding module to effectively embed segmentation features into detection features. The idea of this module is to capture hierarchical context information that includes local objects or parts and nonlocal context features by learning multiple attention maps, and subsequently utilize interdependencies between features to recalibrate noisy segmentation features. Furthermore, we use this module in the proposed gated encoder-decoder network that adaptively aggregates feature maps of different resolutions based on the gate mechanism so that we can embed multiscale segmentation feature maps into detection features for more accurate detection of objects of all sizes. Experimental results demonstrate the effectiveness of the proposed method on the Pascal VOC 2012Seg dataset, the Pascal VOC dataset and the MS COCO dataset.
Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan, Hengcan Shi
IEEE Trans. Multim.2
2019 Scene Parsing via Integrated Classification Model and Variance-Based Regularization
abstract
Scene Parsing is a challenging task in computer vision, which can be formulated as a pixel-wise classification problem. Existing deep-learning-based methods usually use one general classifier to recognize all object categories. However, the general classifier easily makes some mistakes in dealing with some confusing categories that share similar appearances or semantics. In this paper, we propose an integrated classification model and a variance-based regularization to achieve more accurate classifications. On the one hand, the integrated classification model contains multiple classifiers, not only the general classifier but also a refinement classifier to distinguish the confusing categories. On the other hand, the variance-based regularization differentiates the scores of all categories as large as possible to reduce misclassifications. Specifically, the integrated classification model includes three steps. The first is to extract the features of each pixel. Based on the features, the second step is to classify each pixel across all categories to generate a preliminary classification result. In the third step, we leverage a refinement classifier to refine the classification result, focusing on differentiating the high-preliminary-score categories. An integrated loss with the variance-based regularization is used to train the model. Extensive experiments on three common scene parsing datasets demonstrate the effectiveness of the proposed method.
Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, Zichen Song 0002
CVPR2
2019 Beyond Synthetic Data: A Blind Deraining Quality Assessment Metric Towards Authentic Rain Image
abstract
Deraining quality assessment (DQA) plays an important role in evaluating and guiding the design of the image deraining algorithm. Due to the absence of rain-free image in the real rainy weather, the existing deraining algorithms are typically tested on several synthetic data by simulating very limited types of rain streaks, which are far from sufficient to measure the practicability of a deraining algorithm. In this paper, we first build a subjective DQA database that collects diverse authentic rain images and their derained versions. Then, a blind quality metric is developed to predict the deraining quality. Since the deraining artifacts are anisotropic and variable, we propose to describe the image via a bi-directional gated fusion network (B-GFN), which adaptively integrates the multi-scale cues of deraining artifact. Experiments confirm the effectiveness of the proposed method and its superiority with respect to many state-of-the-art blind image quality metrics.
Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Hongliang Li 0001, Fanman Meng
ICIP4
2019 Blind Image Sharpness Assessment And Enhancement via Deep Auxiliary Learning
abstract
In this paper, we propose an unified deep auxiliary learning network to train the blind image sharpness assessment (BISA) metric and enhancer simultaneously. Instead of using the BISA as a parameter tuner like existing works, the proposed method aims to exploit the complementary information between two tasks and boost both of their performance. On the one hand, the enhancement subnetwork tries to separate a blurry image into the clear version and disparity map, which provide additional mask effect and blurry degree information for accurate BISA. On the other hand, the BISA subnetwork help determine the enhancement degree by feeding sharpness-aware features to the enhancer, which is helpful for avoiding under-/over-enhancing. Experimental results on three publicly available databases show that the proposed method outperforms many state-of-the-art algorithms in both the BISA and sharpness enhancement tasks.
Qingbo Wu 0001, Rui Ma 0030, King Ngi Ngan, Hongliang Li 0001, Fanman Meng
ICME4
2019 Incorporating Non-local and Task-specific Features for Instance Segmentation
abstract
This paper proposes a novel instance segmentation model, which improves the instance segmentation by considering two aspects. One is a new non-local features module to recover detailed information that is lost in the deep convolutional operations. The other is to introduce attention mechanism to generate specific features adaptive to each task. The proposed method is verified on three well-known datasets, namely Pascal VOC, Cityscapes and COCO. The experiments show that the method using the proposed modules outperforms baseline Mask R-CNN on all of the datasets without bells and whistles.
Longrong Yang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
MMSP4
2019 A New Few-shot Segmentation Network Based on Class Representation
abstract
This paper studies few-shot segmentation, which is a task of predicting foreground mask of unseen classes by a few of annotations only, aided by a set of rich annotations already existed. The existing methods mainly focus the task on "how to transfer segmentation cues from support images (labeled images) to query images (unlabeled images)", and try to learn efficient and general transfer module that can be easily extended to unseen classes. However, it is proved to be a challenging task to learn the transfer module that is general to various classes. This paper solves few-shot segmentation in a new perspective of "how to represent unseen classes by existing classes", and formulates few-shot segmentation as the representation process that represents unseen classes (in terms of forming the foreground prior) by existing classes precisely. Based on such idea, we propose a new class representation based few-shot segmentation framework, which firstly generates class activation map of unseen class based on the knowledge of existing classes, and then uses the map as foreground probability map to extract the foregrounds from query image. A new two-branch based few-shot segmentation network is proposed. Moreover, a new CAM generation module that extracts the CAM of unseen classes rather than the classical training classes is raised. We validate the effectiveness of our method on Pascal VOC 2012 dataset, the value FB-IoU of one-shot and five-shot arrives at 69.2% and 70.1% respectively, which outperforms the state-of-the-art method.
Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Qingbo Wu 0001
VCIP3
2018 Key-Word-Aware Network for Referring Expression Image Segmentation
Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001
ECCV (6)2
2018 Boosting Scene Parsing Performance via Reliable Scale Prediction
abstract
Segmenting objects on suitable scales is a key factor to improve the scene parsing performance. Existing methods either simply average multi-scale results or predict scales by weakly-supervised models, due to the lack of scale labels. In this paper, we propose a novel fully-supervised Scale Prediction Model. On one hand, the proposed Scale Prediction Model learns parsing scales by the strong scale supervision, which is automatically generated from the scene parsing ground truth without any extra manually annotation. On the other hand, we explore the relationship between scale and object class, and propose to use the object class information to further improve the reliability of the scale prediction. The proposed Scale Prediction Model improves 23.1%, 20.1% and 29.3% scale prediction accuracies on the NYU Depth v2, PASCAL-Context and SIFT Flow datasets, respectively. Based on the Scale Prediction Model, we design a Scale Parsing Net (SPNet) for scene parsing, which segments each object on the scale predicted by the Scale Prediction Model. Moreover, SPNet leverages the intermediate result (i.e., the object class) to refine the parsing results. The experiment results show that SPNet outperforms many state-of-the-art methods on multiple scene parsing datasets.
Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan
ACM Multimedia2
2018 Multi-task Learning for Deep Semantic Hashing
abstract
Deep learning to hash has emerged as a popular technique for large-scale image retrieval. Existing deep learning to hash methods seek to solve the single retrieval task within one stream framework or jointly solve the retrieval task and the classification task within two stream framework. Consequently, the semantic information is not fully exploited to generate compact and discriminative hash codes. In this paper, we propose a multi-task learning architecture for deep semantic hashing (MLDH), which incorporates the retrieval task and the classification task within one-stream framework. Specifically, we introduce a COCO loss to learn compact binary codes for the classification task. For the retrieval task, we introduce a pairwise loss to learn discriminative binary codes. Finally, these two tasks are investigated into one-stream deep learning framework. Extensive experiments show that MLDH can outperform state-of-the-art methods on benchmark datasets.
Lei Ma 0004, Hongliang Li 0001, Qingbo Wu 0001, Chao Shang 0001, King Ngi Ngan
VCIP2
2018 Weakly Supervised Semantic Segmentation by Multiple Group Cosegmentation
abstract
Weakly supervised semantic segmentation aims at segmenting images by image-level labels. The existing methods try to train an end-to-end CNN network, which needs to handle multiple classes that is difficult. In addition, the existing methods are sensitive to the image-level cues such as discriminative regions and the pseudo-annotations. To avoid these drawbacks, this paper proposes a new strategy, which first obtains the foregrounds of each class by multiple group cosegmentation, and then combines the results to form the semantic segmentation. In our method, three new aspects are considered. (1) we solve semantic segmentation by each class that is easy to handle. (2) we extract discriminative regions more globally by context analysis. (3) we learn local-to-global segmentation network to segment the object from local discriminative priors. A new CNN network for multiple group cosegmentation is proposed. Two subnetworks such as global context based discriminative region extraction network and local-to-global segmentation network are designed. A simple combination method based on the discriminative map is proposed to finally obtain the semantic segmentation results. We verify the proposed method on Pascal VOC dataset. The experimental results show that the proposed method can obtain mIOU value 0.563 and 0.603 (without CRF post-processing) on the validation and test dataset that outperforms many existing weakly supervised semantic segmentation methods.
Kunming Luo, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
VCIP4
2018 Global and local semantics-preserving based deep hashing for cross-modal retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan
Neurocomputing2
2018 Interactive object segmentation in two phases
King Ngi Ngan, Songnan Li, Hongliang Li 0001
Signal Process. Image Commun.4
2018 Boundary-Guided Optimization Framework for Saliency Refinement
abstract
Salient object detection has made a rapid progress in recent years. To improve the quality of initial saliency maps, existing algorithms typically refine them via a neighbor-constrained smoothing model, which assigns similar saliency values to neighboring regions. Since the adjacent regions could also cross the boundary between the salient object and background, these spatial distance-based methods easily cause false detection by involving the background regions that are close to the salient objects. To address this problem, we propose a boundary-guided optimization framework to jointly improve the region smoothness and correct the false detect regions. Specifically, we introduce a latent segmentation variable to regularize the consistency between the refined saliency map and the latent segmentation mask, which penalizes high (low) saliency values of the regions lying outside (inside) the estimated object boundary. To optimize the proposed objective function, we decompose the primary problem into two subproblems: submodular optimization problem and convex optimization problem. The submodular optimization problem can be quickly optimized using the off-the-shelf technique while the convex optimization can be solved with a closed-form solution. The experimental results show that the proposed method consistently improve the performances of eight state-of-the-art salient object detection algorithms on three datasets, including latest deep convolutional neural network-based algorithms. Meanwhile, our method outperforms state-of-the-art saliency refinement algorithms.
Liangzhi Tang, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan
IEEE Signal Process. Lett.2
2018 An Unsupervised Method to Extract Video Object via Complexity Awareness and Object Local Parts
abstract
Existing unsupervised video object segmentation generates object information from the whole video, which ignores analysis of the local clips. However, we observe that local clips and their relationships are also useful for the video object segmentation. For example, the simple background clips can be used to improve the segmentation of complex background clips. In this paper, we propose a novel unsupervised segmentation framework to segment the primary object based on two aspects, i.e., the complexity awareness of video clips and their segmentation propagation. The first one is used to select the simple clips with smooth backgrounds and the second one generates an object prior from the simple clips and propagates the object prior to help and improve the segmentation of the complex clips. A complexity awareness method using the static cues and the dynamic cues are proposed to evaluate the complexity of the video frames. A new object prior learning model based on the local part structure is designed and a local part-based prior propagation is proposed for the complex clip segmentation. To verify our method, we collect a new challenging video segmentation data set, in which each video contains diverse backgrounds. Experimental results demonstrate that our method outperforms several state-of-the-art methods both on a classical data set and our new data set.
Bing Luo 0003, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.2
2018 Globally Measuring the Similarity of Superpixels by Binary Edge Maps for Superpixel Clustering
abstract
This paper proposes an edge-based superpixel similarity measurement, which globally evaluates the similarity between superpixels by binary edge maps. The basic idea is to assess whether the superpixels are surrounded by the same edges. To this end, we first describe the edge spatial distributions by directional regions and then use the directional regions to represent the surrounding relationships of superpixels and edges by their traverse relationships, which form the histogram feature. Finally, the similarity is simply calculated by the distances between the features. To verify the proposed similarity measurement, we use our global similarity measurement to perform superpixel clustering. Two clustering methods, the directed graph clustering (DGC) and spectral clustering (ultrametric contour map) are combined to achieve the clustering process. The combination of our global similarity measurement and DGC to form a new three-layer-based superpixel generation method, which can quickly generate the superpixel from edge maps, is highlighted. We verify the global similarity measurement by the BSDS500 dataset. The experimental results demonstrate that the proposed global similarity measurement can improve the clustering accuracy in terms of larger intersection-over-union-criterion-based values. The code can be downloaded from https://github.com/FanmanMeng/Superpixel-Similarity-Measurement.
Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, Chao Huang 0003, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.2
2018 LETRIST: Locally Encoded Transform Feature Histogram for Rotation-Invariant Texture Classification
abstract
Classifying texture images, especially those with significant rotation, illumination, scale, and viewpoint changes, is a fundamental and challenging problem in computer vision. This paper proposes a simple yet effective image descriptor, called Locally Encoded TRansform feature hISTogram (LETRIST), for texture classification. LETRIST is a histogram representation that explicitly encodes the joint information within an image across feature and scale spaces. The proposed representation is training-free, low-dimensional, yet discriminative and robust for texture description. It consists of the following major steps. First, a set of transform features is constructed to characterize local texture structures and their correlation by applying linear and non-linear operators on the extremum responses of directional Gaussian derivative filters in scale space. Established on the basis of steerable filters, the constructed transform features are exactly rotationally invariant as well as computationally efficient. Second, the scalar quantization via binary or multi-level thresholding is adopted to quantize these transform features into texture codes. Two quantization schemes are designed, both of which are robust to image rotation and illumination changes. Third, the cross-scale joint coding is explored to aggregate the discrete texture codes into a compact histogram representation, i.e., LETRIST. Experimental results on the Outex, CUReT, KTH-TIPS, and UIUC texture data sets show that LETRIST consistently produces better or comparable classification results than the state-of-the-art approaches. Impressively, recognition rates of 100.00% and 99.00% have been achieved on the Outex and KTH-TIPS data sets, respectively. In addition, the noise robustness is evaluated on the Outex and CUReT data sets. The source code is publicly available athttps://github.com/stc-cqupt/letrist.
Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Jianfei Cai 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Blind Image Quality Assessment Using Local Consistency Aware Retriever and Uncertainty Aware Evaluator
abstract
Blind image quality assessment (BIQA) aims to automatically predict the perceptual quality of a digital image without accessing its pristine reference. Previous studies mainly focus on extracting various quality-relevant image features. By contrast, the explorations on highly efficient learning model are still very limited. Motivated by the fact that it is difficult to approximate a complex and large data set via a global parametric model, we propose a novel local learning method for BIQA to improve quality prediction performance. More specifically, we search for the perceptually similar neighbors of a test image to serve as its unique training set. Unlike the widely used k nearest neighbors principle, which only measures the similarity between the testing and training samples, the local consistency of the selected training data is also considered to generate smoother sample space. The image quality is estimated via a sparse Gaussian process. As an additional benefit, the uncertainty of the predicted score is jointly inferred, which can subsequently drive more robust perceptual image processing applications, such as deblocking investigated in this paper. Extensive experiments demonstrate that the proposed learning model leads to consistent quality prediction improvements over many state-of-the-art BIQA algorithms.
Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Kede Ma
IEEE Trans. Circuits Syst. Video Technol.2
2018 A Perceptually Weighted Rank Correlation Indicator for Objective Image Quality Assessment
abstract
In the field of objective image quality assessment (IQA), Spearman's ρ and Kendall's τ, which straightforwardly assign uniform weights to all quality levels and assume that each pair of images is sortable, are the two most popular rank correlation indicators. These indicators can successfully measure the average accuracy of an IQA metric for ranking multiple processed images. However, two important perceptual properties are ignored. First, the sorting accuracy (SA) of high-quality images is usually more important than that of poor-quality images in many real-world applications, where only top-ranked images are pushed to the users. Second, due to the subjective uncertainty in making judgments, two perceptually similar images are usually barely sortable, and their ranks do not contribute to the evaluation of an IQA metric. To more accurately compare different IQA algorithms, in this paper, we explore a perceptually weighted rank correlation indicator, which rewards the capability of correctly ranking high-quality images and suppresses the attention towards insensitive rank mistakes. Specifically, we focus on activating a 'valid' pairwise comparison of images whose quality difference exceeds a given sensory threshold (ST). Meanwhile, each image pair is assigned a unique weight that is determined by both the quality level and rank deviation. By modifying the perception threshold, we can illustrate the sorting accuracy with a sophisticated SA-ST curve rather than a single rank correlation coefficient. The proposed indicator offers new insight into interpreting visual perception behavior. Furthermore, the applicability of our indicator is validated for recommending robust IQA metrics for both degraded and enhanced image data.
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan
IEEE Trans. Image Process.2
2018 Generic Proposal Evaluator: A Lazy Learning Strategy Toward Blind Proposal Quality Assessment
abstract
Existing detection or recognition systems typically select one state-of-the-art proposal algorithm to produce massive object-covered candidate windows, and a quality metric specifically designed for this algorithm is utilized to single out small amounts of proposals. However, in practice, the accuracies of different proposal algorithms significantly change from one image content to another one. To obtain more robust proposal results, a generic proposal evaluator (GPE) is highly desired, which could choose optimal candidate windows across multiple proposal algorithms. In this paper, we propose a lazy learning strategy to train the GPE, which aims to blindly estimate the quality of each proposal without accessing to its manual annotation. Unlike the traditional end-to-end framework that learns a universal model from all training samples, we try to build query-specific training subset for each given proposal, where only its k-nearest-neighborhoods are collected from all labeled candidate windows. Benefits from the capability of updating the regression parameters for different visual contents, the proposed method delivers a higher quality prediction accuracy even with respect to the deep neural network learned by end-to-end method. Experimental results confirm that the proposed algorithm significantly outperforms many state-of-the-art proposal quality metrics.
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan
IEEE Trans. Intell. Transp. Syst.2
2018 Seeds-Based Part Segmentation by Seeds Propagation and Region Convexity Decomposition
abstract
Object part segmentation is an important and challenging task in computer vision. The existing supervised part segmentation methods need pixel level training data which leads to a huge workload for the user. In this paper a weakly supervised part segmentation method is proposed which segments part regions from multiple images by only several seeds on an image. Two aspects such as seed propagation among multiple images and part generation from seeds are considered. The first aspect is to generate part seeds in each image in terms of seed propagation which is accomplished by part matching combined with latent object regions. We fuse the local part matching and global shape cosegmentation to avoid the noise propagation. The second aspect is to segment part regions from object regions and part seeds which is formulated as the object shape decomposition model. The shape convexity analysis and seed location are fused to accomplish the decomposition and the final part segmentation. The proposed method is verified on the PASCAL 2010 dataset Bird dataset Cat-Dog dataset and UCF Sports Actions dataset. Experimental results demonstrate the effectiveness of the proposed method with larger intersection over union (IOU) values compared with existing weakly supervised part generation methods.
Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan, Jianfei Cai 0001
IEEE Trans. Multim.2
2018 Hierarchical Parsing Net: Semantic Scene Parsing From Global Scene to Objects
abstract
This paper proposes a novel Hierarchical Parsing Net (HPN) for semantic scene parsing. Unlike previous methods, which separately classify each object, HPN leverages global scene semantic information and the context among multiple objects to enhance scene parsing. On the one hand, HPN uses the global scene category to constrain the semantic consistency between the scene and each object. On the other hand, the context among all objects is also modeled to avoid incompatible object predictions. Specifically, HPN consists of four steps. In the first step, we extract scene and local appearance features. Based on these appearance features, the second step is to encode a contextual feature for each object, which models both the scene-object context (the context between the scene and each object) and the interobject context (the context among different objects). In the third step, we classify the global scene and then use the scene classification loss and a backpropagation algorithm to constrain the scene feature encoding. In the fourth step, a label map for scene parsing is generated from the local appearance and contextual features. Our model outperforms many state-of-the-art deep scene parsing networks on five scene parsing databases.
Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan
IEEE Trans. Multim.2
2017 Blind proposal quality assessment via deep objectness representation and local linear regression
abstract
The quality of object proposal plays an important role in boosting the performance of many computer vision tasks, such as, object detection and recognition. Due to the absence of manually annotated bounding-box in practice, the quality metric towards blind assessment of object proposal is highly desirable for singling out the optimal proposals. In this paper, we propose a blind proposal quality assessment algorithm based on the Deep Objectness Representation and Local Linear Regression (DORLLR). Inspired by the hierarchy model of the human vision system, a deep convolutional neural network is developed to extract the objectness-aware image feature. Then, the local linear regression method is utilized to map the image feature to a quality score, which tries to evaluate each individual test window based on its k-nearest-neighbors. Experimental results on a large-scale IoU labeled dataset verify that the proposed method significantly outperforms the state-of-the-art blind proposal evaluation metrics.
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Linfeng Xu 0001
ICME2
2017 Store classification using Text-Exemplar-Similarity and Hypotheses-Weighted-CNN
Chao Huang 0003, Hongliang Li 0001, Wei Li 0110, Qingbo Wu 0001, Linfeng Xu 0001
J. Vis. Commun. Image Represent.2
2017 Manifold-ranking embedded order preserving hashing for image semantic retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001
J. Vis. Commun. Image Represent.2
2017 Improving object proposals with top-down cues
Wei Li 0110, Hongliang Li 0001, Bing Luo 0003, Hengcan Shi, Qingbo Wu 0001, King Ngi Ngan
Signal Process. Image Commun.2
2017 Gaze-Based Object Segmentation
abstract
This letter addresses the problem of object segmentation with a gaze map and a group of candidate regions. First, we analyze distribution characteristics of gazes when people look at an image. Then, we summarize different cases of the group of candidate regions. Based on our analysis and summary, we develop three measures to evaluate the likelihood of a candidate region belonging to the target object and a pooling method to create a likelihood map of this object. Finally, the measures and pooling method are integrated with a proposed iterative strategy for generating the segmentation result. Experimental results demonstrate that our method can handle different types of gaze maps and different groups of candidate regions, and the overall performance of our method is better than that of the state-of-the-art method.
King Ngi Ngan, Hongliang Li 0001
IEEE Signal Process. Lett.3
2017 Waterloo Exploration Database: New Challenges for Image Quality Assessment Models
abstract
The great content diversity of real-world digital images poses a grand challenge to image quality assessment (IQA) models, which are traditionally designed and validated on a handful of commonly used IQA databases with very limited content variation. To test the generalization capability and to facilitate the wide usage of IQA techniques in real-world applications, we establish a large-scale database named the Waterloo Exploration Database, which in its current state contains 4744 pristine natural images and 94 880 distorted images created from them. Instead of collecting the mean opinion score for each image via subjective testing, which is extremely difficult if not impossible, we present three alternative test criteria to evaluate the performance of IQA models, namely, the pristine/distorted image discriminability test, the listwise ranking consistency test, and the pairwise preference consistency test (P-test). We compare 20 well-known IQA models using the proposed criteria, which not only provide a stronger test in a more challenging testing environment for existing models, but also demonstrate the additional benefits of using the proposed database. For example, in the P-test, even for the best performing no-reference IQA model, more than 6 million failure cases against the model are "discovered" automatically out of over 1 billion test pairs. Furthermore, we discuss how the new database may be exploited using innovative approaches in the future, to reveal the weaknesses of existing IQA models, to provide insights on how to improve the models, and to shed light on how the next-generation IQA models may be developed. The database and codes are made publicly available at: https://ece.uwaterloo.ca/~k29ma/exploration/.
Kede Ma, Zhengfang Duanmu, Qingbo Wu 0001, Zhou Wang 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006
IEEE Trans. Image Process.6
2017 Weakly Supervised Part Proposal Segmentation From Multiple Images
abstract
Weakly supervised local part segmentation is challenging, due to the difficulty of modeling multiple local parts from image level prior. In this paper, we propose a new weakly supervised local part proposal segmentation method based on the observation that local parts will keep fixed along the object pose variations. Hence, the local part can be segmented by capturing object pose variations. Based on such observation, a new local part proposal segmentation model is proposed. Three aspects, such as shape similarity-based cosegmentation, shape matching-based part detection and segmentation, and graph matching-based part assignment are considered. A part segmentation energy function is first proposed. Four terms, such as MRF-based single image segmentation term, shape feature-based foreground consistency term, NCuts-based part segmentation term, and two-order graphs matching based part consistency term, are contained. Then, a three sub-minimization-based energy minimization method is proposed to accomplish approximation solution. Finally, we verify our method based on three image data sets (PASCAL VOC 2008 Part data set, UCB Bird data set, and Cat-Dog data set), and one video data set (UCF Sports) data set. The experimental results demonstrate a better segmentation performance compared with the existing object cosegmentation and part proposal generation methods.
Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, King Ngi Ngan
IEEE Trans. Image Process.2
2017 Objective Quality Assessment of Image Retargeting by Incorporating Fidelity Measures and Inconsistency Detection
abstract
The tremendous growth in mobile devices has resulted in huge generation and usage of digital images. Image quality assessment is thus an important issue for mobile media applications. In this paper, we focus on the quality evaluation of images generated by content-aware image retargeting, in which the reference and the distorted images are of different sizes. Through retargeting, many types of deformation inconsistency lead to shape distortion, deformation artifacts, and content information loss, worsening its perceptual quality. The deformation inconsistency occurs on different levels of the retargeted images. Limited by the accuracy of the alignment between the original and retargeted images, previous methods only focus on pixel-level and patch-level fidelity analyses and fail to detect deformation inconsistency. In this paper, we improve the alignment algorithm and propose a three-level representation of the retargeting process. Based on the analysis of this three-level representation, both fidelity measures and inconsistency detection are combined to determine the final retargeting quality. The proposed algorithm is validated on the public data sets RetargetMe and CUHK. Experimental results demonstrate that inconsistency detection contributes to accurately assessing the image retargeting perceptual quality. This inspires us to investigate more about deformation inconsistency to formulate the objective quality of image retargeting.
Yichi Zhang 0014, King Ngi Ngan, Lin Ma 0002, Hongliang Li 0001
IEEE Trans. Image Process.4
2017 PBC: Polygon-Based Classifier for Fine-Grained Categorization
abstract
Fine-grained categorization is a challenging task mainly due to two factors: first, objects share similar appearances between different categories; second, objects present significant pose variation within the same category. To address these challenges, we propose a method to automatically detect discriminative and pose-invariant regions, which is referred to as a polygon-based classifier (PBC). In the first stage, we generate a set of polygons that are composed of multiple parts. For each polygon, a classifier is trained based on deep features of a convolutional network. Then, a greedy algorithm is employed to select the discriminative and complementary polygon-based classifiers that deliver highest classification accuracy for fine-grained object categories. In the second stage, the confusing classes of the first stage are selected and employed to train the polygon-based classifiers. Then, a greedy algorithm is employed to select discriminative classifiers. For the test images, we use the classifiers trained in the first stage to obtain a coarse result. Then, the classifiers of the second stage are adopted to distinguish the confusing classes of the coarse result. In our experiments, the proposed approach is evaluated on three well-known fine-grained datasets. The experiments show that our approach outperforms the state-of-the-art methods.
Chao Huang 0003, Hongliang Li 0001, Yurui Xie, Qingbo Wu 0001, Bing Luo 0003
IEEE Trans. Multim.2
2017 Video Object Segmentation via Global Consistency Aware Query Strategy
abstract
In this paper, we propose a video object segmentation method via global consistency aware query strategy. The aim is to obtain higher segmentation accuracy with less user annotation. Intuitively, we hope to annotate some frames to obtain better segmentation performance than to annotate other frames, which can be modeled by active learning framework. Specifically, we first generate a sample space of potential annotation regions via an object proposals method for each frame. Then, the annotation likelihood for the region is calculated in terms of annotation history and global consistency for the object in the video. Third, the segmentation result of the annotation region can be obtained by minimizing an MRF energy function. Fourth, the algorithm will provide the user with the most valuable frame to annotate, which has high annotation likelihood and large segmentation result change. Finally, the annotation is added to the framework to begin the next iteration. Experiments on a number of video sequences demonstrate that the proposed method can reduce the user effort and obtain the higher segmentation accuracy compared with the state-of-the-art methods.
Bing Luo 0003, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Chao Huang 0003
IEEE Trans. Multim.2
2017 Learning Efficient Binary Codes From High-Level Feature Representations for Multilabel Image Retrieval
abstract
Due to the efficiency and effectiveness of hashing technologies, they have become increasingly popular in large-scale image semantic retrieval. However, existing hash methods suppose that the data distributions satisfy the manifold assumption that semantic similar samples tend to lie on a low-dimensional manifold, which will be weakened due to the large intraclass variation. Moreover, these methods learn hash functions by relaxing the discrete constraints on binary codes to real value, which will introduce large quantization loss. To tackle the above problems, this paper proposes a novel unsupervised hashing algorithm to learn efficient binary codes from high-level feature representations. More specifically, we explore nonnegative matrix factorization for learning high-level visual features. Ultimately, binary codes are generated by performing binary quantization in the high-level feature representations space, which will map images with similar (visually or semantically) high-level feature representations to similar binary codes. To solve the corresponding optimization problem involving nonnegative and discrete variables, we develop an efficient optimization algorithm to reduce quantization loss with guaranteed convergence in theory. Extensive experiments show that our proposed method outperforms the state-of-the-art hashing methods on several multilabel real-world image datasets.
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan
IEEE Trans. Multim.2
2017 Blind Image Quality Assessment Based on Rank-Order Regularized Regression
abstract
Blind image quality assessment (BIQA) aims to estimate the subjective quality of a query image without access to the reference image. Existing learning-based methods typically train a regression function by minimizing the average error between subjective opinion scores and model predictions. However, minimizing average error does not necessarily lead to correct quality rank-orders between the test images, which is a highly desirable property of image quality models. In this paper, we propose a novel rank-order regularized regression model to address this problem. The key idea is to introduce a pairwise rank-order constraint into the maximum margin regression framework, aiming to better preserve the correct perceptual preference. To the best of our knowledge, this is the first attempt to incorporate rank-order constraints into margin-based quality regression model. By combing with a new local spatial structure feature, we achieve highly consistent quality prediction with human perception. Experimental results show that the proposed method outperforms many state-of-the-art BIQA metrics on popular publicly available IQA databases (i.e., LIVE-II, TID2013, VCL@FER, LIVEMD, and ChallengeDB).
Qingbo Wu 0001, Hongliang Li 0001, Zhou Wang 0001, Fanman Meng, Bing Luo 0003, Wei Li 0110, King Ngi Ngan
IEEE Trans. Multim.2
2016 Group MAD Competition? A New Methodology to Compare Objective Image Quality Models
abstract
Objective image quality assessment (IQA) models aim to automatically predict human visual perception of image quality and are of fundamental importance in the field of image processing and computer vision. With an increasing number of IQA models proposed, how to fairly compare their performance becomes a major challenge due to the enormous size of image space and the limited resource for subjective testing. The standard approach in literature is to compute several correlation metrics between subjective mean opinion scores (MOSs) and objective model predictions on several well-known subject-rated databases that contain distorted images generated from a few dozens of source images, which however provide an extremely limited representation of real-world images. Moreover, most IQA models developed on these databases often involve machine learning and/or manual parameter tuning steps to boost their performance, and thus their generalization capabilities are questionable. Here we propose a novel methodology to compare IQA models. We first build a database that contains 4,744 source natural images, together with 94,880 distorted images created from them. We then propose a new mechanism, namely group MAximum Differentiation (gMAD) competition, which automatically selects subsets of image pairs from the database that provide the strongest test to let the IQA models compete with each other. Subjective testing on the selected subsets reveals the relative performance of the IQA models and provides useful insights on potential ways to improve them. We report the gMAD competition results between 16 well-known IQA models, but the framework is extendable, allowing future IQA models to be added into the competition.
Kede Ma, Qingbo Wu 0001, Zhou Wang 0001, Zhengfang Duanmu, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006
CVPR6
2016 Part propagation for local part segmentation
abstract
Segment propagation transfers object priors among images, which is an important prior generation manner in image segmentation. The existing propagation methods focus on object foreground propagation, while the detailed part propagation is deficiency, which is caused by the challenges that not only the multiple part regions, but also their relationships need to be transferred. In this paper, a part propagation method is proposed. Two level propagations such as object level propagation, and part level propagation are successively used for the part propagation. The object level propagation is to transfer global shape information among images, which is formulated as graph matching based edge fragments matching problem, with dynamic programming solution. The part level propagation is to transfer the more detailed part labels, which is formulated as pixel level structure matching problem, and is efficiently solved by traditional dense pixel matching methods. The proposed method is verified on 15 challenging classes selected from PASCAL 2010 dataset, Bird dataset and Cat-Dog dataset. The experimental results demonstrate the effectiveness of the proposed method.
Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, Jianfei Cai 0001, Chao Huang 0003
VCIP2
2016 Q-DNN: A quality-aware deep neural network for blind assessment of enhanced images
abstract
Image enhancement is widely popular due to its capability of producing "better" visual quality for specific applications. Although many enhancement algorithms have been developed in recent years, the studies towards blind assessment of enhanced images are still very lacking. In this paper, we propose a data-driven blind image quality assessment (BIQA) method based on the quality-aware deep neural network (Q-DNN). Unlike the conventional hand-crafted features designed for measuring the degradation level of specific distortion types, a supervised learning model is utilized in our Q-DNN, which is capable of adaptively updating the feature extractor and quality regressor for describing the visual artifacts caused by different image enhancement tasks. Experimental results on two challenging enhanced image databases show that the proposed method is significantly superior to the state-of-the-art BIQA metrics.
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan
VCIP2
2016 Cosegmentation of multiple image groups
Fanman Meng, Jianfei Cai 0001, Hongliang Li 0001
Comput. Vis. Image Underst.3
2016 Hybrid human detection and recognition in surveillance
Qiang Liu 0015, Wei Zhang 0021, Hongliang Li 0001, King Ngi Ngan
Neurocomputing3
2016 Feature discovering for image classification via wavelet-like pattern decomposition
Yurui Xie, Hongliang Li 0001, Chao Huang 0003, Bo Wu 0013, Linfeng Xu 0001
J. Vis. Commun. Image Represent.2
2016 Blind Image Quality Assessment Based on Multichannel Feature Fusion and Label Transfer
abstract
In this paper, we propose an efficient blind image quality assessment (BIQA) algorithm, which is characterized by a new feature fusion scheme and a k-nearest-neighbor (KNN)-based quality prediction model. Our goal is to predict the perceptual quality of an image without any prior information of its reference image and distortion type. Since the reference image is inaccessible in many applications, the BIQA is quite desirable in this context. In our method, a new feature fusion scheme is first introduced by combining an image's statistical information from multiple domains (i.e., discrete cosine transform, wavelet, and spatial domains) and multiple color channels (i.e., Y, Cb, and Cr). Then, the predicted image quality is generated from a nonparametric model, which is referred to as the label transfer (LT). Based on the assumption that similar images share similar perceptual qualities, we implement the LT with an image retrieval procedure, where a query image's KNNs are searched for from some annotated images. The weighted average of the KNN labels (e.g., difference mean opinion score or mean opinion score) is used as the predicted quality score. The proposed method is straightforward and computationally appealing. Experimental results on three publicly available databases (i.e., LIVE II, TID2008, and CSIQ) show that the proposed method is highly consistent with human perception and outperforms many representative BIQA metrics.
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 Low-Delay Rate Control for Consistent Quality Using Distortion-Based Lagrange Multiplier
abstract
Video quality fluctuation plays a significant role in human visual perception, and hence, many rate control approaches have been widely developed to maintain consistent quality for video communication. This paper presents a novel rate control framework based on the Lagrange multiplier in high-efficiency video coding. With the assumption of constant quality control, a new relationship between the distortion and the Lagrange multiplier is established. Based on the proposed distortion model and buffer status, we obtain a computationally feasible solution to the problem of minimizing the distortion variation across video frames at the coding tree unit level. Extensive simulation results show that our method outperforms the rate control used in HEVC Test Model (HM) by providing a more accurate rate regulation, lower video quality fluctuation, and stabler buffer fullness. The average peak signal-to-noise ratio (PSNR) and PSNR deviation improvements are about 0.37 dB and 57.14% in the low-delay (P and B) video communication, where the complexity overhead is ∼ 4.44% .
Miaohui Wang, King Ngi Ngan, Hongliang Li 0001
IEEE Trans. Image Process.3
2015 A highly efficient method for blind image quality assessment
abstract
Blind image quality assessment (BIQA) has attracted a great deal of attention due to the increasing demand in industry and the promising recent progress in academia. To bridge the gap between academic research accomplishment and industrial needs, high efficiency BIQA approaches that allow for real-time computation are highly desirable. In this paper, we propose a novel BIQA method by selecting statistical features extracted from binary patterns of local image structures. This allows us to largely reduce the feature space to eventually one dimension. Somewhat surprisingly, such a single feature, faster-than-real-time approach named local pattern statistics index (LPSI) exhibits impressive generalization ability across different distortion types and achieves competitive quality prediction performance in comparison with state-of-the-art approaches on public databases such as LIVE II and TID2008.
Qingbo Wu 0001, Zhou Wang 0001, Hongliang Li 0001
ICIP3
2015 Improved block level adaptive quantization for high efficiency video coding
abstract
As the concept of block level adaptivity becomes an important feature in recent video CODECs, block level adaptive quantization (BLAQ) is being considered in the High Efficiency Video Coding (HEVC) standard. The BLAQ is based on the assumption that each block should have its own quantization parameter (QP), which can adapt to the local content of video sequences much better, and hence the video encoder with adaptive QP can perform a better perceptual quality. However, in the HEVC reference software, the BLAQ is required to obtain a proper QP for each block by the rate distortion optimization (RDO) scheme and so the computational complexity of the encoder increases significantly. In this paper, an improved BLAQ algorithm is proposed to obtain the adaptive QP for each block. The simulation results show that the proposed method can save more bits as well as require lower computational complexity, compared to the traditional method.
Miaohui Wang, King Ngi Ngan, Hongliang Li 0001, Huanqiang Zeng
ISCAS3
2015 Object Segmentation from Long Video Sequences
abstract
Most existing video segmentation methods are focused on extracting the primary objects in test video sequences. They assumed that only one object appeared through the whole video sequences, which is impractical in many applications. In this paper, we focus on the object segmentation from the long video sequences which consist of many different scenes, shot cuts and various motion patterns, etc. In order to solve this problem, we propose a framework to segment the objects in relative video shots, while discarding the irrelative video shots. A graph is constructed to model the video object detection and final segmentation is obtained by getting the superpixels in the detection boxes. We also introduce a new long video segmentation dataset which corresponds to the pixel-wise ground truth. The experiments demonstrate that our proposed method can deal with the object segmentation in long video sequence.
Bing Luo 0003, Hongliang Li 0001, Tiecheng Song, Chao Huang 0003
ACM Multimedia2
2015 No reference image quality assessment metric via multi-domain structural information and piecewise regression
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Shuyuan Zhu
J. Vis. Commun. Image Represent.2
2015 Exploring space-frequency co-occurrences via local quantized patterns for texture representation
Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003
Pattern Recognit.2
2015 Special issue on recent advances in saliency models, applications and evaluations
Zhi Liu 0003, Olivier Le Meur, Ali Borji, Hongliang Li 0001
Signal Process. Image Commun.4
2015 An Efficient Frame-Content Based Intra Frame Rate Control for High Efficiency Video Coding
abstract
Rate control plays an important role in the rapid development of high-fidelity video services. As the High Efficiency Video Coding (HEVC) standard has been finalized, many rate control algorithms are being developed to promote its commercial use. The HEVC encoder adopts a new R-lambda based rate control model to reduce the bit estimation error. However, the R-lambda model fails to consider the frame-content complexity that ultimately degrades the performance of the bit rate control. In this letter, a gradient based R-lambda (GRL) model is proposed for the intra frame rate control, where the gradient can effectively measure the frame-content complexity and enhance the performance of the traditional R-lambda method. In addition, a new coding tree unit (CTU) level bit allocation method is developed. The simulation results show that the proposed GRL method can reduce the bit estimation error and improve the video quality in HEVC all intra frame coding.
Miaohui Wang, King Ngi Ngan, Hongliang Li 0001
IEEE Signal Process. Lett.3
2015 Constrained Directed Graph Clustering and Segmentation Propagation for Multiple Foregrounds Cosegmentation
abstract
This paper proposes a new constrained directed graph clustering (DGC) method and segmentation propagation method for the multiple foreground cosegmentation. We solve the multiple object cosegmentation with the perspective of classification and propagation, where the classification is used to obtain the object prior of each class and the propagation is used to propagate the prior to all images. In our method, the DGC method is designed for the classification step, which adds clustering constraints in cosegmentation to prevent the clustering of the noise data. A new clustering criterion such as the strongly connected component search on the graph is introduced. Moreover, a linear time strongly connected component search algorithm is proposed for the fast clustering performance. Then, we extract the object priors from the clusters, and propagate these priors to all the images to obtain the foreground maps, which are used to achieve the final multiple objects extraction. We verify our method on both the cosegmentation and clustering tasks. The experimental results show that the proposed method can achieve larger accuracy compared with both the existing cosegmentation methods and clustering methods.
Fanman Meng, Hongliang Li 0001, Shuyuan Zhu, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.2
2015 Visual Quality Evaluation of Image Object Segmentation: Subjective Assessment and Objective Measure
abstract
A visual quality evaluation of image object segmentation as one member of the visual quality evaluation family has been studied over the years. Researchers aim at developing the objective measures that can evaluate the visual quality of object segmentation results in agreement with human quality judgments. It is also significant to construct a platform for evaluating the performance of the objective measures in order to analyze their pros and cons. In this paper, first, we present a novel subjective object segmentation visual quality database, in which a total of 255 segmentation results were evaluated by more than thirty human subjects. Then, we propose a novel full-reference objective measure for an object segmentation visual quality evaluation, which involves four human visual properties. Finally, our measure is compared with some state-of-the-art objective measures on our database. The experiment demonstrates that the proposed measure performs better in matching subjective judgments. Moreover, the database is available publicly for other researchers in the field to evaluate their measures.
King Ngi Ngan, Songnan Li, Raveendran Paramesran, Hongliang Li 0001
IEEE Trans. Image Process.5
2015 Fast HEVC Inter CU Decision Based on Latent SAD Estimation
abstract
The emerging high efficiency video coding (HEVC) standard has improved compression performance significantly in comparison with H.264/AVC. However, more intensive computational complexity has been introduced by adopting a number of new coding tools. In this paper, a fast inter CU decision is proposed based on the latent sum of absolute differences (SAD) estimation. Firstly, a two-layer motion estimation (ME) method is designed to take advantage of the latent SAD cost. The new ME method can obtain the SAD costs for both the upper CU and its sub-CUs. Secondly, a concept of motion compensation rate- distortion (R-D) cost is defined, and an exponential model is proposed to express the relationship between the motion compensation R-D cost and the SAD cost. Then, a fast CU decision approach is designed based on the exponential model. The fast CU decision is implemented by comparing a derived threshold with the SAD cost difference between the upper and sub SAD costs. Experimental results show that the proposed algorithm achieves an average of 52% and 58.4% reductions of the coding time at the cost of 1.61% and 2% bit-rate increases under the low delay and random access conditions, respectively.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan
IEEE Trans. Multim.2
2014 On Multiple Image Group Cosegmentation
Fanman Meng, Jianfei Cai 0001, Hongliang Li 0001
ACCV (4)3
2014 Fast and efficient inter CU decision for high efficiency video coding
abstract
In this paper, a graph cut based fast Coding Unit (CU) decision algorithm is proposed for HEVC inter frames. Firstly, a feature called pyramid variance of the absolute difference (PVAD) is designed for the CU selection. Secondly, the CU decision is modeled as a Markov Random Field (MRF) inference problem, which can be optimized by the graph cut algorithm. Thirdly, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show the effectiveness of the proposed method.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Bing Zeng 0001, Shuyuan Zhu, Qingbo Wu 0001
ICIP2
2014 Using mid-high level cues to detect salient object
abstract
This paper proposes a novel saliency object detection method by using the mid-level and high-level visual cues. In the mid-level objectness evaluation, we generate three complementary saliency maps, such as the multi-scale segmentation cue, the background cue and the spatial color distribution cue. The first cue is used to highlight the objects via the local region segment. The second cue uses the background priors to detect the saliency information. The third cue is to capture the spatial color distribution. For the high-level visual cue, we propose an objectness evaluation model to distinguish the object and the background. All the saliency cues are finally combined to achieve the saliency detection. The experimental results show that the proposed method outperforms the state-of-the-art saliency object detection methods.
Hongliang Li 0001, Yurui Xie, Bing Luo 0003, Liangzhi Tang, Bing Zeng 0001, King Ngi Ngan, Fanman Meng
ICME1
2014 Cosegmentation from similar backgrounds
abstract
Recently, the common objects are often required to be extracted from a group of images in many applications, such as video coding and model training. Co-segmentation is a new and efficient method for this requirement. In realistic applications, we observe that the images usually contain similar backgrounds (namely similar scene co-segmentation), such as the city landmark images collected from the web or the key frames sampled from a video. Meanwhile, the existing co-segmentation has not paid so much attention on the similar scene co-segmentation, and the insufficiently accurate segments may be provided by the existing methods. In this paper, we propose an active contours based co-segmentation model to provide foregrounds from the similar backgrounds. We combine the background consistency constraint with the foreground consistency constraint to form the energy function, and use the method of level-set and the calculus of variations to minimize the model. We also speed up the model by the hierarchical structure and the superpixel technique. We test the method on both the image and video dataset. The results show that the proposed model can obtain larger IOU values than the state-of-the-art co-segmentation methods.
Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Nini Rao
ISCAS2
2014 Texture classification using joint statistical representation in space-frequency domain with local quantized patterns
abstract
Despite its success in texture analysis, Local Binary Pattern (LBP) is operated in the original image space, and it fails to capture deeper pixel interactions to provide a more discriminative description. In this paper, we propose to explore the joint statistical representation in the space-frequency domain with local quantized patterns for texture classification. The proposed method consists of two channels. In each channel, the multi-resolution spatial filters are employed to generate multi-scale spatial maps and the local Fourier transform is subsequently applied to extract local frequency features (spectral maps). The global thresholding is adopted to quantize the spatial and spectral maps into different levels, which are then jointly encoded to built a space-frequency co-occurrence histogram. Finally, the two-channel feature histograms are combined to represent the texture. Experiments on the Outex texture database demonstrate the robustness of our method to image rotation and illumination changes, and our method outperforms the state of the art in terms of the classification accuracy.
Tiecheng Song, Hongliang Li 0001, Bing Zeng 0001, Moncef Gabbouj
ISCAS2
2014 No reference image quality metric via distortion identification and multi-channel label transfer
abstract
In this paper, we propose a no reference image quality assessment (NR-IQA) algorithm based on distortion identification (DI) and multi-channel label transfer (LT). First, the distortion type classification is used to obtain the query image's probabilities of belonging to each distortion type. Then, the distortion specific label transfer is implemented in multiple distortion category channels. Based on the hypothesis that the similar images share the similar subjective qualities, the label transfer predicts the subjective quality of the query image by pooling the labels of its k-nearest neighbors (KNN) retrieved from the annotated samples. A weighting average of the multi-channel label transfer's outputs is computed to obtain the final perceptual quality score. The weight is the query image's probability that belongs to the corresponding distortion type. The experimental results show that the proposed method outperforms representative NR-IQA approaches and some full-reference metrics.
Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Moncef Gabbouj
ISCAS2
2014 Noise-Robust Texture Description Using Local Contrast Patterns via Global Measures
abstract
This letter presents a noise-robust descriptor by exploring a set of local contrast patterns (LCPs) via global measures for texture classification. To handle image noise, the directed and undirected difference masks are designed to calculate three types of local intensity contrasts: directed, undirected, and maximum difference responses. To describe pixel-wise features, these responses are separately quantized and encoded into specific patterns based on different global measures. These resulting patterns (i.e., LCPs) are jointly encoded to form our final texture representation. Experiments are conducted on the well-known Outex and CUReT databases in the presence of high levels of noise. Compared to many state-of-the-art methods, the proposed descriptor achieves superior texture classification performance while enjoying a compact feature representation.
Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003, Bing Zeng 0001, Moncef Gabbouj
IEEE Signal Process. Lett.2
2014 Unsupervised Multiclass Region Cosegmentation via Ensemble Clustering and Energy Minimization
abstract
The problem of unsupervised segmentation of multi-class regions can be significantly boosted when they irregularly recur in multiple images. The existing segmentation methods are either weakly supervised, such as tagging images with object classes, or are limited by the assumption that each image contains all the object instances. In this paper, we propose a new method to cosegment multiclass regions from a group of images without the assumption about object configurations. The key idea is to discover the unknown object-like proposals via a robust ensemble clustering scheme. The proposals are then used to derive unary and pairwise energy potentials across all the images, which can be minimized with the α-expansion. Experimental evaluation on a number of image groups demonstrates the good performance of the proposed method on the multiclass region cosegmentation.
Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003
IEEE Trans. Circuits Syst. Video Technol.1
2014 Semantic Annotation of Satellite Images Using Author-Genre-Topic Model
abstract
In this paper, we propose a novel hierarchical generative model, named author-genre-topic model (AGTM), to perform satellite image annotation. Different from the existing author-topic model in which each author and topic are associated with the multinomial distributions over topics and words, in AGTM, each genre, author, and topic are associated with the multinomial distributions over authors, topics, and words, respectively. The bias of the distribution of the authors with respect to the topics can be rectified by incorporating the distribution of the genres with respect to the authors. Therefore, the classification accuracy of documents is improved when the information of genre is introduced. By representing the images with several visual words, the AGTM can be used for satellite image annotation. The labels of classes and scenes of the images correspond to the authors and the genres of the documents, respectively. The labels of classes and scenes of test images can be estimated, and the accuracy of satellite image annotation is improved when the information of scenes is introduced in the training images. Experimental results demonstrate the good performance of the proposed method.
Hongliang Li 0001, Guanghui Liu 0001, Liaoyuan Zeng
IEEE Trans. Geosci. Remote. Sens.2
2014 Repairing Bad Co-Segmentation Using Its Quality Evaluation and Segment Propagation
abstract
In this paper, we improve co-segmentation performance by repairing bad segments based on their quality evaluation and segment propagation. Starting from co-segmentation results of the existing co-segmentation method, we first perform co-segmentation quality evaluation to score each segment. Good segments can be filter out based on the scores. Then, a propagation method is designed to transfer good segments to the rest bad ones so as to repair the bad segmentation. In our method, the quality evaluation is implemented by the measurements of foreground consistency and segment completeness. Two propagation methods such as global propagation and local region propagation are then defined to achieve the more accurate propagation. We verify the proposed method using four state-of-the-arts co-segmentation methods and two public datasets such as ICoseg dataset and MSRC dataset. The experimental results demonstrate the effectiveness of the proposed quality evaluation method. Furthermore, the proposed method can significantly improve the performance of existing methods with larger intersection-over-union score values.
Hongliang Li 0001, Fanman Meng, Bing Luo 0003, Shuyuan Zhu
IEEE Trans. Image Process.1
2014 MRF-Based Fast HEVC Inter CU Decision With the Variance of Absolute Differences
abstract
The newly developed High Efficiency Video Coding (HEVC) Standard has improved video coding performance significantly in comparison to its predecessors. However, more intensive computation complexity is introduced by implementing a number of new coding tools. In this paper, a fast coding unit (CU) decision based on Markov random field (MRF) is proposed for HEVC inter frames. First, it is observed that the variance of the absolute difference (VAD) is proportional with the rate-distortion (R-D) cost. The VAD based feature is designed for the CU selection. Second, the decision of CU splittings is modeled as an MRF inference problem, which can be optimized by the Graphcut algorithm. Third, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show that the proposed algorithm can achieve about 53% reduction of the coding time with negligible coding performance degradation, which outperforms the state-of-the-art algorithms significantly.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Shuyuan Zhu, Qingbo Wu 0001, Bing Zeng 0001
IEEE Trans. Multim.2
2014 A Fast HEVC Inter CU Selection Method Based on Pyramid Motion Divergence
abstract
The newly developed HEVC video coding standard can achieve higher compression performance than the previous video coding standards, such as MPEG-4, H.263 and H.264/AVC. However, HEVC's high computational complexity raises concerns about the computational burden on real-time application. In this paper, a fast pyramid motion divergence (PMD) based CU selection algorithm is presented for HEVC inter prediction. The PMD features are calculated with estimated optical flow of the downsampled frames. Theoretical analysis shows that PMD can be used to help selecting CU size. A k nearest neighboring like method is used to determine the CU splittings. Experimental results show that the fast inter prediction method speeds up the inter coding significantly with negligible loss of the peak signal-to-noise ratio.
Jian Xiong 0005, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng
IEEE Trans. Multim.2
2013 Complexity awareness based feature adaptive co-segmentation
abstract
In this paper, we achieve co-segmentation by learning adaptive feature model for each image group. A novel feature adaptive co-segmentation method and an image complexity awareness method are proposed. We also propose a linear feature model and an expectation-minimization (EM) based algorithm for adaptive feature learning. In the EM based algorithm, two aspects such as the accuracy confidence of the simple image segmentation and the fitness of the learned model to the simple image segmentation are considered. L1-regularized least squares optimization is also combined for the minimization. By testing on several well-known datasets, the error rates of the final co-segmentation are verified to be lower than the existing state-of-the-art co-segmentation methods.
Fanman Meng, Hongliang Li 0001
ICIP2
2013 Segmenting specific object based on logo detection
abstract
This paper proposes a method to segment object with logos. In the method, we firstly locate the logos by SIFT matching. Then, the object boundary is extracted based on the logo location. Finally, we model the object prior based on the boundary, and introduce the prior into Markov random field segmentation method to segment the object. To verify the proposed method, we collect a logo dataset from the web such as Flickr and Google. The experimental results demonstrate the effectiveness of the proposed method.
Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001
ISCAS2
2013 Saliency detection using a central stimuli sensitivity based model
abstract
In this paper, a novel method is proposed to predict attention in image scenes by using a central stimuli sensitivity based saliency model. The proposed method is based on the general “center-surround” visual attention mechanism and the spatial frequency response of the human visual system (HVS). Following three biologically inspired principles, the saliency value is computed by two “scatter matrices” which are used to measure the similarity and distinctness within and between two classes, i.e., the center and surrounding regions, respectively. In order to detect salient objects with different size, the saliency of a pixel is estimated via the saliency support region of the pixel, which is the most salient region centered at the pixel with respect to the surrounding region. The proposed method which is compliant with human perceptual characteristics enables the prediction of human fixations. Experimental results on three eye tracking datasets verify the effectiveness of the method and show that the proposed method outperforms the state-of-the-art methods on the visual saliency detection task.
Linfeng Xu 0001, Hongliang Li 0001, Liaoyuan Zeng, Zhengning Wang, Guanghui Liu 0001
ISCAS2
2013 Saliency detection using joint spatial-color constraint and multi-scale segmentation
Linfeng Xu 0001, Hongliang Li 0001, Liaoyuan Zeng, King Ngi Ngan
J. Vis. Commun. Image Represent.2
2013 WaveLBP based hierarchical features for image classification
Tiecheng Song, Hongliang Li 0001
Pattern Recognit. Lett.2
2013 Two-layer average-to-peak ratio based saliency detection
Hongliang Li 0001, Linfeng Xu 0001, Guanghui Liu 0001
Signal Process. Image Commun.1
2013 Mode dependent down-sampling and interpolation scheme for high efficiency video coding
Qingbo Wu 0001, Hongliang Li 0001
Signal Process. Image Commun.2
2013 Face Hallucination via Similarity Constraints
abstract
In this letter, we present a new face hallucination method based on similarity constraints to produce a high-resolution (HR) face image from an input low-resolution (LR) face image. This method is modeled as a local linear filtering process by incorporating four constraint functions at patch level. The first two constraints focus on checking if the training images are similar to the input face image. The third is defined in the HR face image, which is to impose the smoothness constraint between neighboring hallucinated patches. The final constraint computes the spatial distance to reduce the effect of patches that are far from the hallucinating patch. Experimental evaluation on a number of face images demonstrates the good performance of the proposed method on the face hallucination task.
Hongliang Li 0001, Linfeng Xu 0001, Guanghui Liu 0001
IEEE Signal Process. Lett.1
2013 Local Polar DCT Features for Image Description
abstract
We present a novel feature descriptor, Local Polar DCT Features (LPDF), which is robust to a variety of image transformations. Specifically, the local patch is quantized in the designed polar geometric structure and the 2-D DCT features are then extracted and rearranged. A subset of the resulting DCT coefficients is selected as our compact LPDF descriptor. We perform a comprehensive performance evaluation with state-of-the-art methods, i.e., SIFT, DAISY, LIOP, and GLOH on the standard Oxford dataset and two additional test image pairs. Experimental results demonstrate the superiority of proposed descriptor under various image transformations, even with very low dimensions.
Tiecheng Song, Hongliang Li 0001
IEEE Signal Process. Lett.2
2013 Image Cosegmentation by Incorporating Color Reward Strategy and Active Contour Model
abstract
The design of robust and efficient cosegmentation algorithms is challenging because of the variety and complexity of the objects and images. In this paper, we propose a new cosegmentation model by incorporating a color reward strategy and an active contour model. A new energy function corresponding to the curve is first generated with two considerations: the foreground similarity between the image pairs and the background consistency in each of the image pair. Furthermore, a new foreground similarity measurement based on the rewarding strategy is proposed. Then, we minimize the energy function value via a mutual procedure which uses dynamic priors to mutually evolve the curves. The proposed method is evaluated on many images from commonly used databases. The experimental results demonstrate that the proposed model can efficiently segment the common objects from the image pairs with generally lower error rate than many existing and conventional cosegmentation methods.
Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan
IEEE Trans. Cybern.2
2013 Feature Adaptive Co-Segmentation by Complexity Awareness
abstract
In this paper, we propose a novel feature adaptive co-segmentation method that can learn adaptive features of different image groups for accurate common objects segmentation. We also propose image complexity awareness for adaptive feature learning. In the proposed method, the original images are first ranked according to the image complexities that are measured by superpixel changing cue and object detection cue. Then, the unsupervised segments of the simple images are used to learn the adaptive features, which are achieved using an expectation-minimization algorithm combining l 1-regularized least squares optimization with the consideration of the confidence of the simple image segmentation accuracies and the fitness of the learned model. The error rate of the final co-segmentation is tested by the experiments on different image groups and verified to be lower than the existing state-of-the-art co-segmentation methods.
Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Liaoyuan Zeng, Qingbo Wu 0001
IEEE Trans. Image Process.2
2013 Co-Salient Object Detection From Multiple Images
abstract
In this paper, we propose a novel method to discover co-salient objects from a group of images, which is modeled as a linear fusion of an intra-image saliency (IaIS) map and an inter-image saliency (IrIS) map. The first term is to measure the salient objects from each image using multiscale segmentation voting. The second term is designed to detect the co-salient objects from a group of images. To compute the IrIS map, we perform the pairwise similarity ranking based on an image pyramid representation. A minimum spanning tree is then constructed to determine the image matching order. For each region in an image, we design three types of visual descriptors, which are extracted from the local appearance, e.g., color, color co-occurrence and shape properties. The final region matching problem between the images is formulated as an assignment problem that can be optimized by linear programming. Experimental evaluation on a number of images demonstrates the good performance of the proposed method on co-salient object detection.
Hongliang Li 0001, Fanman Meng, King Ngi Ngan
IEEE Trans. Multim.1
2013 From Logo to Object Segmentation
abstract
This paper proposes a method to segment object from the web images using logo detection. The method consists of three steps. In the first step, the logos are located from the original images by SIFT matching. Based on the logo location and the object shape model, the second step extracts the object boundary from the image. In the third step, we use the object boundary to model the object appearance, which is then used in the MRF based segmentation method to finally achieve the object segmentation. The key of our method is the object boundary extraction, which is achieved by searching a variation of the shape model that best fits the local edge of the image. Affine transform is used to consider the variations among the objects. Meanwhile, the Nelder-Mead simplex method with a simple initial rough search is used to run the boundary search. To verify the proposed method, we collect a LogoSeg dataset from the web such as Flickr and Google. The MOMI dataset is also used for the verification. The experimental results demonstrate that the proposed logo detection based segmentation method can improve the performance of the object segmentation.
Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan
IEEE Trans. Multim.2
2012 Image co-segmentation via active contours
abstract
In this paper, a new co-segmentation model by incorporating active contours based method and rewarding strategy is represented. We first generate co-segmentation energy function from two aspects. One is foreground similarity between image pairs. The other is background consistency in each single image. Then, we optimize the energy function through a mutual optimization approach. We verify the proposed method on the images commonly used in co-segmentation research. Experimental results demonstrate the effectiveness of our method.
Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001
ISCAS2
2012 Mode dependent deblocking filter for video coding
abstract
In this paper, a mode dependent deblocking filter is proposed for suppressing the effects of quantization noise. The proposed filter employs Wiener Filter as an in-loop filter which can minimize the mean square error between the original image and the decoded image. In addition, to adapt to different local features of the decoded image efficiently, we design our filter based on the intra mode combination of each pair of 4×4 blocks. Experiment results show that the proposed filter achieves superior coding gains relative to H.264/AVC high profile with negligible complexity increase.
Qingbo Wu 0001, Hongliang Li 0001
ISCAS2
2012 Saliency detection from joint embedding of spatial and color cues
abstract
Visual saliency detection provides an important methodology for many computer vision applications. In this paper, we propose a novel method to detect salient regions from an image. To detect pixel-level saliency, this method uses joint embedding of spatial and color cues, i.e., spatial constraint based saliency, color double-opponent saliency, and similarity distribution based saliency. Finally, a multi-layer structure is adopted to merge the three terms into a saliency map. In order to make the saliency map consistent, we perform a region-based saliency detection by incorporating a multi-scale segmentation technique. The proposed method was evaluated on the MSRA benchmark images. Experimental results show that our method outperforms the state-of-the-art methods on visual saliency detection by achieving both higher precision and better recall.
Linfeng Xu 0001, Hongliang Li 0001, Zhengning Wang
ISCAS2
2012 Automatic Annotation of Multispectral Satellite Images Using Author-Topic Model
abstract
In this letter, we propose a new method for the annotation of multispectral satellite images. This method performs the multispectral image annotation by incorporating a graphical model. To obtain the annotated image, first, we use a set of images with defined semantic concepts to represent the training set. Second, the images are represented by several visual words based on the color and texture features. Finally, an author-topic model is exploited to estimate probabilities of semantic classes for the regions in the test images and categorize them into the semantic concepts. Experimental evaluation on the multispectral images demonstrates the good performance of the proposed method on the multispectral image annotation.
Hongliang Li 0001, Guanghui Liu 0001
IEEE Geosci. Remote. Sens. Lett.2
2012 Segmenting focused objects based on the Amplitude Decomposition Model
Tiantang Chen, Hongliang Li 0001
Pattern Recognit. Lett.2
2012 Global salient information maximization for saliency detection
Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan
Signal Process. Image Commun.2
2012 Object Co-Segmentation Based on Shortest Path Algorithm and Saliency Model
abstract
Segmenting common objects that have variations in color, texture and shape is a challenging problem. In this paper, we propose a new model that efficiently segments common objects from multiple images. We first segment each original image into a number of local regions. Then, we construct a digraph based on local region similarities and saliency maps. Finally, we formulate the co-segmentation problem as the shortest path problem, and we use the dynamic programming method to solve the problem. The experimental results demonstrate that the proposed model can efficiently segment the common objects from a group of images with generally lower error rate than many existing and conventional co-segmentation methods.
Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan
IEEE Trans. Multim.2
2011 Directional samples reordering for intra residual transform
abstract
In this paper, a directional samples reordering (DSR) based algorithm was proposed for intra residual data transform. To make the residual more suitable for discrete cosine transform, the diagonal edge in arbitrary size of intra prediction block can be rotated to a regular horizontal edge by reordering the samples in the block. The intra residual data will be used to implement 2D discrete cosine transform after DSR procedure. Experimental results show that up to 0.5967 dB and on average 0.4372 dB gain can be achieved for CIF sequence in high bitrate with the proposed algorithm at high complexity mode than H.264 intra coding.
Qingbo Wu 0001, Hongliang Li 0001, Tiantang Chen
MMSP2
2011 Automatic body segmentation with graph cut and self-adaptive initialization level set (SAILS)
Qiang Liu 0015, Hongliang Li 0001, King Ngi Ngan
J. Vis. Commun. Image Represent.2
2011 Soft-Change Detection in Optical Satellite Images
abstract
In this letter, we propose a novel approach for unsupervised change detection in multitemporal optical satellite images. Unlike the traditional methods, the proposed method, called the soft-change detection, models the change detection as a transparency computation problem and assigns to each pixel a set of soft labels. In order to extract the pixel opacity, we optimize an objective function by exploiting the Bayesian matting method. Comparisons between the proposed method and the state-of-the-art methods are reported. Experimental results demonstrate the effectiveness of the proposed method.
Hongliang Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2011 Learning to Extract Focused Objects From Low DOF Images
abstract
This paper proposes an approach to extract focused objects (i.e., attention objects) from low depth-of-field images. To recognize the focused object, we decompose the image into multiple regions, which are described by using three types of visual descriptors. Each descriptor is extracted from a representation of some aspects of local appearance, e.g., a spatially localized texture, color, or geometrical property. Therefore, the focus detection of a region can be achieved by the classification of extracted visual descriptors based on a binary classifier. We employ a boosting algorithm to learn the classifier with a cascade of decision structure. Given a test image, initial segmentation can be achieved using obtained classification results. Finally, we apply a post-processing technique to improve the results by incorporating region grouping and pixel-level segmentation. Experimental evaluation on a number of images demonstrates the performance advantages of the proposed method, when compared with state-of-the-art methods.
Hongliang Li 0001, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.1
2011 A Co-Saliency Model of Image Pairs
abstract
In this paper, we introduce a method to detect co-saliency from an image pair that may have some objects in common. The co-saliency is modeled as a linear combination of the single-image saliency map (SISM) and the multi-image saliency map (MISM). The first term is designed to describe the local attention, which is computed by using three saliency detection techniques available in literature. To compute the MISM, a co-multilayer graph is constructed by dividing the image pair into a spatial pyramid representation. Each node in the graph is described by two types of visual descriptors, which are extracted from a representation of some aspects of local appearance, e.g., color and texture properties. In order to evaluate the similarity between two nodes, we employ a normalized single-pair SimRank algorithm to compute the similarity score. Experimental evaluation on a number of image pairs demonstrates the good performance of the proposed method on the co-saliency detection task.
Hongliang Li 0001, King Ngi Ngan
IEEE Trans. Image Process.1
2011 Guided Face Cartoon Synthesis
abstract
In this paper, we propose a new method, called guided synthesis, to synthesize a face cartoon from a face photo. The guided synthesis is defined as a local linear model, which generates a cartoon image by incorporating the content of guidance images taken from the training set. Our synthesis operation is achieved based on four weight functions. The first is a photo-photo weight that aims to measure the similarity between an input photo patch and a training photo patch. The second is defined as a photo-cartoon weight, which is used to compute the likelihood by computing the similarity between a cartoon patch and an input photo patch. The third weight is defined in the synthesized photos, which is to set a smoothness constraint between neighboring synthesized patches. The final weight is designed to evaluate the similarity of a synthesized patch to an input patch based on the spatial distance. Experimental evaluation on a number of face photos demonstrates the good performance of the proposed method on the face cartoon synthesis.
Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan
IEEE Trans. Multim.1
2010 Parallel-Filtering Based Equalization of OFDM over Doubly Selective Channels
abstract
Time selectivity of multipath channels severely degrades the performance of OFDM systems. In this paper, a piece-wise channel approximation is presented for improving the structure of the frequency-domain channel gain matrix. The proposed model is exploited to simplify the zero-forcing (ZF) equalizer as a parallel-filtering implementation with very low complexity. The linearized case of the parallel-filtering based equalizer is studied and applied to the DVB-H receiver design. The simulation indicates that the developed equalizer enforces considerably the immunity of the DVB-H receiver to the time selectivity with very little complexity increased.
Guanghui Liu 0001, Hongliang Li 0001, Mingzhen Wang
GLOBECOM2
2010 Learn to segment attention object from low DoF image
abstract
In this paper, a novel segmentation algorithm is proposed to extract attention object (i.e., focus object) from Low depth of field image. In order to recognize the focus object, we first decompose the image into multiple segments that are described by visual words. Each visual word is computed from a filter bank to represent the high frequency components. The boosting method is then used to generate a strong classifier for each training image. Given a test image, we employ the voting algorithm to achieve the attention decision according to obtained strong classifiers. To extract focus objects from the test image, two-level segmentation method is proposed, which includes region and pixel levels segmentation. Experimental evaluation on test images shows that the proposed method is capable of segmenting the attention object quite effectively.
Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan
ISCAS1
2009 FaceSeg: Automatic Face Segmentation for Real-Time Video
abstract
Segmenting human faces automatically is very important for face recognition and verification, security system, and computer vision. In this paper, we present an accurate segmentation system for cutting human faces out from video sequences in real-time. First, a learning based face detector is developed to rapidly find human faces. To speed up the detection process, a face rejection cascade is constructed to remove most of negative samples while retaining all the face samples. Then, we develop a coarse-to-fine segmentation approach to extract the faces based on a min-cut optimization. Finally, a new matting algorithm is proposed to estimate the alpha-matte based on an adaptive trimap generation method. Experimental results demonstrate the effectiveness and robustness of our proposed method that can compete with the well-known interactive methods in real-time.
Hongliang Li 0001, King Ngi Ngan, Qiang Liu 0015
IEEE Trans. Multim.1
2008 Saliency model-based face segmentation and tracking in head-and-shoulder video sequences
Hongliang Li 0001, King Ngi Ngan
J. Vis. Commun. Image Represent.1
2008 An efficient intra-mode selection algorithm for H.264 based on edge classification and rate-distortion estimation
King Ngi Ngan, Hongliang Li 0001
Signal Process. Image Commun.3
2008 Fast and Efficient Method for Block Edge Classification and Its Application in H.264/AVC Video Coding
abstract
Edge is an important feature in video classification which finds applications in video representation and coding. In H.264/AVC, intra-prediction mode decision (a computationally intensive process) is based on the orientation of the edges in the macroblock. In this paper, we first investigate the difference properties derived from three coefficients in the non-normalized Haar transform (NHT) domain and present a fast and efficient method to classify block edge using these properties. The proposed method significantly reduces the number of computational operations in the edge models determination with no multiplications and less addition operations. The use of edge classification for fast intra-prediction mode decision in H.264/AVC video coding is then presented. The experimental results show the effectiveness of the proposed method.
Hongliang Li 0001, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.1
2007 An Efficient Intra Mode Selection Algorithm For H.264 Based On Fast Edge Classification
abstract
The H.264/AVC is the newest video coding standard recommended by ITU-T and MPEG. Compared with all existing video coding standards, H.264 can achieve superior performance by using many advanced techniques. Intra mode selection is an important feature in H.264 standard and can reduce the spatial redundancy in intra frame significantly. An efficient rate distortion optimization (RDO) technique is employed in H.264 to choose the best mode for each MB, but the computational cost increases drastically. In this paper, a fast intra mode selection algorithm is introduced. By using a fast edge detection method which is based on non-normalized Haar transform (NHT), edge for each sub-block can be extracted. Based on the local edge information, only few intra modes are chosen as mode candidates. A fast RDO algorithm is also proposed in this paper. By combing these two methods, computational load is reduced remarkably. Experimental results show that this fast intra mode selection scheme can lessen about 80% encoding time with little loss of bit-rate and visual quality.
Hongliang Li 0001, King Ngi Ngan
ISCAS2
2007 Unsupervized Video Segmentation With Low Depth of Field
abstract
In this paper, a novel segmentation algorithm based on matting model is proposed to extract the focused objects in low depth-of-field (DoF) video images. The proposed algorithm is fully automatic and can be used to partition the video image into focused objects and defocused background. This method consists of three stages. The first stage is to generate a saliency map of the input image by the reblurring model. In the second stage, bilateral and morphological filtering are employed to smooth and accentuate the salient regions. Then a trimap with three regions is calculated by an adaptive thresholding method. The third stage involves the proposed adaptive error control matting scheme to extract the boundaries of the focused objects accurately. Experimental evaluation on test sequences shows that the proposed method is capable of segmenting the focused region effectively and accurately.
Hongliang Li 0001, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.1
2007 A Multiple Visual Models Based Perceptive Analysis Framework for Multilevel Video Summarization
abstract
In this paper, we propose a generic framework to human perception analysis in video understanding based on multiple visual cues. Video features that prominently influence human perception, such as motion, contrast, special scenes, and statistical rhythm, are first extracted and modeled. A perception curve that corresponds to human perception change is then constructed from these individual models using linear or priority based fusion approach. As an important application of the perceptive analysis framework, a feasible scheme for video summarization is implemented in order to demonstrate the validity, robustness, and generality of the proposed framework. The frames that correspond to the peak points in these individual models and the fusion curve are extracted as multilevel summarizations that include video keywords, keyframes, and dynamic segments. The subjective evaluations from a supplementary volunteer study on video summarizations indicate that the analysis framework is effective and offer a promising approach to semantic video management, access, and understanding
Junyong You, Guizhong Liu, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2006 Unsupervised Segmentation of Defocused Video Based on Matting Model
abstract
In this paper, an unsupervised segmentation algorithm based on matting model is proposed to extract the focused objects in the low depth of field (DOF) video images. The proposed algorithm is fully automatic and can be used to partition the video image into focused objects and defocused background. This method consists of three stages. The first stage is to generate the saliency map from the input image. In the second stage, bilateral and morphological filtering are employed to smooth and lift the saliency regions. Then a trimap with three regions is calculated by an adaptive thresholding method. The third stage involves the Poisson matting scheme to extract the boundaries of the focused objects accurately. Experimental evaluation on test sequences shows that the proposed method is capable of segmenting the focused region quite effectively and accurately.
Hongliang Li 0001, King Ngi Ngan
ICIP1
2006 Face segmentation in head-and-shoulder video sequences based on facial saliency map
abstract
In this paper, a novel face segmentation algorithm is proposed based on facial saliency map (FSM) for head-and-shoulder type video application. This method consists of three stages. The first stage is to generate the saliency map of input video image by our proposed facial attention model. In the second stage, a geometric model and a eye-map built from chrominance components are employed to localize face region according to the saliency map. The third stage involves the adaptive boundary correction and the final face contour extraction. Experimental evaluation on test sequences shows that the proposed method is capable of segmenting and tracking the face area quite effectively.
Hongliang Li 0001, King Ngi Ngan
ISCAS1
2006 Fast and efficient method for block edge classification
abstract
Advanced multimedia applications will have to provide the user with the flexibility to rapidly access and manipulate the multimedia data. In order to achieve this goal, especially in limited computing environments, such as mobile-phone, the computational cost of extracting the visual features for image segmentation must be greatly reduced. In this paper, we investigate difference properties from three coefficients in the non-normalised Haar transform (NHT) domain and present a fast and efficient method to classify block edge using these properties. The proposed method significantly reduces the number of computational operations in the edge models determination with no multiplications and less addition operations. The experimental results are presented to show the effectiveness of the proposed method.
Hongliang Li 0001, King Ngi Ngan
IWCMC1
2006 A new texture generation method based on pseudo-DCT coefficients
abstract
In this paper, a new method for generating different texture images is presented. This method involves a simple transform from a certain one-dimensional (1-D) signal to an expected two-dimensional (2-D) image. Unlike traditional methods, the input signal is generated by a simple 1-D function in our work instead of a sample texture. We first transform the 1-D input signal into frequency domain using fast Fourier transform. Based on the sufficient analysis in 2-D discrete cosine transform (DCT) domain, where each of the coefficients expresses a texture feature in a certain direction, the 2-D pseudo-DCT coefficients are then constructed by appropriately rearranging the Fourier coefficients in terms of their frequency components. Finally, the corresponding texture image can be produced by 2-D inverse DCT algorithm. We applied the proposed method to generate several stochastic textures (i.e., cloud, illumination, and sand), and several structural texture images. Experimental results indicate the good performance of the proposed method.
Hongliang Li 0001, Guizhong Liu
IEEE Trans. Image Process.1
2005 A novel PDE-based rate-distortion model for rate control
abstract
This paper presents a novel rate-distortion (R-D) model for rate control of video coding. First, we investigate the well-known heat conduction equation (HCE) in the heat conduction process for a thin bar. HCE describes the dynamic distribution of temperature in a thin bar using a partial differential equation (PDE). Motivated by the heat conduction process in a thin bar, a new rate transmission equation (RTE) is proposed to describe the dynamic behaviors of the bit-rate fluctuation during video coding. Second, a particular solution of RTE is employed as our proposed R-D model, which consists of two variables, i.e., the distortion D and the source statistical character M. Based on our model, a corresponding rate-control scheme is proposed for MPEG coding. Finally, extensive experimental results are reported to show that, compared with the well-known MPEG-4 Q2 rate-control scheme, our proposed work achieves better buffer status, higher control accuracy, and stabler and more consistent picture quality.
Guizhong Liu, Hongliang Li 0001, Yongli Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2005 Optimization of integer wavelet transforms based on difference correlation structures
abstract
In this paper, a novel lifting integer wavelet transform based on difference correlation structure (DCCS-LIWT) is proposed. First, we establish a relationship between the performance of a linear predictor and the difference correlations of an image. The obtained results provide a theoretical foundation for the following construction of the optimal lifting filters. Then, the optimal prediction lifting coefficients in the sense of least-square prediction error are derived. DCCS-LIWT puts heavy emphasis on image inherent dependence. A distinct feature of this method is the use of the variance-normalized autocorrelation function of the difference image to construct a linear predictor and adapt the predictor to varying image sources. The proposed scheme also allows respective calculations of the lifting filters for the horizontal and vertical orientations. Experimental evaluation shows that the proposed method produces better results than the other well-known integer transforms for the lossless image compression.
Hongliang Li 0001, Guizhong Liu
IEEE Trans. Image Process.1
2004 MRF based construction of statistical operator and its application
Hongliang Li 0001, Guizhong Liu, Yongli Li 0001, Xingsong Hou
Sci. China Ser. F Inf. Sci.1
2004 Adaptive scene-detection algorithm for VBR video stream
abstract
Several scene-detection algorithms, which are only based on bit rate fluctuations, have been proposed. All of them are presented on the fixed thresholds, which are obtained by the empirical records of the video characteristics. Due to the sensitivity of these methods to the accuracy of the records, which are generally obtained by testing several values repeatedly, bad performance evaluation might be observed for the actual scene detection, especially for real-time video traffic. In this paper, we review the previous works in this area, and study the correlation between the scene duration and the scene change at the frame level, and simultaneously investigate the local statistical characteristics of scenes such as variance and peak bit rate etc. Based on this analysis, an effective decision function is first constructed for the scene segmentation. Then, we propose a scene-detection algorithm using the defined dynamic threshold model, which can capture the statistical properties of the scene changes. Experimental results using 15 variable bit rate MPEG video traces indicate good performances of the proposed algorithm with significantly improved scene-detection accuracy.
Hongliang Li 0001, Guizhong Liu, Yongli Li 0001
IEEE Trans. Multim.1
2003 An effective burstiness estimation model for VBR video stream
abstract
Burstiness property plays an important role in the variable bit rate (VBR) video traffic. One of the most significant issues in providing quality-of-service (QoS) guarantees for the real-time multimedia communications in a high-speed network is to estimate the burstiness accurately. In this paper, an efficient burstiness estimation model (fractal exponent model) for VBR video stream is proposed. In order to improve the accuracy of the burstiness measure, a new operator is firstly introduced, which is similar to the convolution operator. Then, using burstiness index derived from the fractal exponent function, we can estimate the burstiness characteristics of VBR video correctly. Simulation results using 20 VBR MPEG video traces indicate good performance of the proposed algorithm with significantly improved burstiness estimation accuracy.
Hongliang Li 0001, Guizhong Liu, Yongli Li 0001
ICME1
2002 An embedded wavelet packet image coding algorithm
abstract
In this paper, we presented a novel wavelet packet image coding approach which provides the functionality of fine granular bitstream scalability. The proposed progressive wavelet packet image coding scheme consists of three parts, wavelet packet decomposition, a quadtree sorting procedure for classifying wavelet coefficients and universal trellis-coded quantization for quantizing the sorted coefficients. The image coding results, calculated in PSNR and images reconstructed by the decoding algorithm, are either comparable to or surpass previous results, due to the flexible representation ability of wavelet packet, the effective quadtree classifier and the improved granular fidelity of the UTCQ over scalar quantization.
Xingsong Hou, Guizhong Liu, Hongliang Li 0001, Yongli Li 0001
ICASSP3
2002 Wavelet-based analysis of hurst parameter estimation for self-similar traffic
abstract
In order to guarantee quality of service (QoS) over Internet, traffic analysis, traffic management have been active research areas. A lot of facts show the Internet traffic and variable bit rate videos streaming all are characterized by self-similar property. Hurst parameter as an important factor that reflects the self-similar property is a key to traffic management and QoS. In this paper existing wavelet methods for the estimation of the Hurst parameter of self-similar traffic is systematically analyzed and examined. The effects of wavelet functions, vanishing moments and wavelet decomposition levels to the results of wavelet methods for acquiring the Hurst parameter are investigated via numerical experiments. Some useful conclusions are drawn on the relationship between the accuracy of the methods and the selection of the order of vanishing moments and the selection of wavelet functions.
Yongli Li 0001, Guizhong Liu, Hongliang Li 0001, Xingsong Hou
ICASSP3
2002 An effective approach to edge classification from DCT domain
abstract
In the field of content-based visual information analysis, the detection of visual features is a significant topic. In order to process video data efficiently, visual features extraction is required. Many advanced video applications require direct manipulation of compressed video data. An effective approach that detects edges in MPEG compressed images is proposed. First, DCT coefficients of 8/spl times/8 subblocks are analyzed on their meaning in determining boundaries. Then, we consider two types of ideal linear edges cutting through a block of size 8/spl times/8. Based on these edge models, we derive an edge detection approach from ten normalized DCT coefficients, obtaining the general rules of edge classification. Finally, we test the proposed algorithm on different images, and compare our method with other edge classification approaches. Simulations show that our approach can be used to estimate the edge information of images from their DCT coefficients more effectively than those proposed previously.
Hongliang Li 0001, Guizhong Liu, Yongli Li 0001
ICIP (1)1
2002 The construction of a statistical prediction lifting operator and its application
abstract
A new method of nonseparable nonlinear wavelet decomposition is proposed, which is suited for the task of image compression, especially for lossless coding applications. It is based on a certain statistical operator that is defined here according to the Markov random field theory. In contrast to the previous nonlinear predictors such as the median or morphological operators, this statistical operator can sufficiently take advantage of the statistical correlation between neighboring pixels. It can be used to realize integer-valued wavelet transforms, which can avoid quantization with the image detail signals being zero (or almost zero) in the smooth gray-level variation areas at a big probability. Numerical results show that the entropy of the coefficients in the transform domain obtained with this new method is smaller than that obtained with the other nonlinear transform methods.
Hongliang Li 0001, Guizhong Liu, Yongli Li 0001, Xingsong Hou
ICIP (1)1
2002 Structure Based Adaptive Prediction of VBR MPEG Traffic for RCBR Network
abstract
With the development of the Internet, variable bit rate (VBR) video will be the major component of future multimedia services. In order to guarantee quality of service (QoS) in real-time transmission, on-line prediction of VBR video traffic integrated with a mechanism for dynamic resource allocation in RCBR (renegotiate constant bit rate) networks has been an active research area. We exploit the short-range dependence (SRD) and long-range dependence (LRD) characteristics of MPEG video and develop a novel structure-based LMS algorithm to predict the bandwidth required by the future frame and group of pictures (GOP). Compared to the LMS algorithm, the modified LMS algorithm can predict the bit rate of frames quickly and accurately without any delay. Through analysis of the queuing of video traffic in the network buffer, dynamic bandwidth allocation using prediction in RCBR networks exhibits better performance than fixed bandwidth allocation.
Yongli Li 0001, Guizhong Liu, Hongliang Li 0001
LCN3
2002 Scene Based M/G/1 Queuing Model of VBR Video Traffic
abstract
To guarantee quality of service (QoS) in future integrated service networks, traffic sources must be characterized to capture the traffic characteristics relevant to network performance. Recent studies reveal that multimedia traffic shows burstiness over multiple time scales and long-range dependence (LRD). Its burstiness and self-similarity make it hard to control and guarantee QoS, so analysis of its queuing behavior attracts much interest and attention. A new scene based M/G/1 queuing model is put forward to analyze the queuing behavior of video streams in the network buffer. Integrated with the scene length distribution, the variance of scene size and the average bit rate of GOPs, the proposed model can accurately describe the queuing behavior of video traffic overall.
Yongli Li 0001, Guizhong Liu, Hongliang Li 0001
LCN3