VLDB 2026 Research / reviewers in the wild / expert
Fanman Meng
dblp:46/7666
· DBLP profile ↗
154ranked-venue papers
20as first author
79since 2021 · last 2026
0000-0002-3016-2567ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 115 · 12 first-author · 53 since 2021Artificial intelligence and machine learning · 36 · 4 first-author · 25 since 2021Systems, architecture and hardware · 5 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Class-agnostic and semantic-aware fusing network with optimal transport for weakly supervised object localization
Lei Ma 0004, Hongbo Wen, Hanyu Hong, Fanman Meng, Qingbo Wu 0001 |
Expert Syst. Appl. | 4 |
| 2026 | CMaP-SAM: Contraction mapping prior for SAM-driven few-shot segmentation
Fanman Meng, Liming Lei, Qingbo Wu 0001, Linfeng Xu 0001, Hongliang Li 0001 |
Neurocomputing | 2 |
| 2026 | Unsupervised deep hashing based on multi-scale aggregation and optimal transport matching for image retrieval
Lei Ma 0004, Hao Pei, Lei Wang 0068, Ying Zhu 0002, Yu Shi 0004, Hanyu Hong, Xinyu Dai, Fanman Meng, Qingbo Wu 0001 |
Neurocomputing | 8 |
| 2026 | Zero-shot egocentric action recognition via chain-of-imagination prompts and inertial strengthening adaptor
Mingzhou He, Ruiqian Li, Qingbo Wu 0001, King Ngi Ngan, Fanman Meng, Hongliang Li 0001 |
Pattern Recognit. | 6 |
| 2026 | Bridging the Gap between Vision and Text for unsupervised text-only captioning
Lanxiao Wang, Heqian Qiu, Haitao Wen, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
Pattern Recognit. | 4 |
| 2026 | DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual RecognitionabstractContinual learning of vision-language models (VLMs) focuses on leveraging cross-modal pretrained knowledge to incrementally adapt to expanding downstream tasks and datasets, while tackling the challenge of knowledge forgetting. Existing research often focuses on connecting visual features with specific class text in downstream tasks, overlooking the latent relationships between general and specialized knowledge. Our findings reveal that forcing models to optimize inappropriate visual-text matches exacerbates forgetting of VLM's recognition ability. To tackle this issue, we propose DesCLIP, which leverages general attribute (GA) descriptions to guide the understanding of specific class objects, enabling VLMs to establish robust <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-GA-class</i> trilateral associations rather than relying solely on <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-class</i> connections. Specifically, we introduce a language assistant to generate concrete GA description candidates via proper request prompts. Then, an anchor-based embedding filter is designed to obtain highly relevant GA description embeddings, which are leveraged as the paired text embeddings for visual-textual instance matching, thereby tuning the visual encoder. Correspondingly, the class text embeddings are gradually calibrated to align with these shared GA description embeddings. Extensive experiments demonstrate the advancements and efficacy of our proposed method, with comprehensive empirical evaluations highlighting its superior performance in VLM-based recognition compared to existing continual learning methods. Chiyuan He, Zihuan Qiu, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | On the Adversarial Robustness of Learning-Based Image Compression Against Rate-Distortion AttacksabstractDespite demonstrating superior Rate-Distortion (RD) performance, Learning-based Image Compression (LIC) algorithms have been found to be vulnerable to malicious perturbations in recent studies. However, the adversarial attacks considered in existing literature remain divergent from real-world scenarios, both in terms of the attack direction and bitrate. Additionally, existing methods focus solely on empirical observations of the model vulnerability, neglecting to identify the origin of it. These limitations hinder the comprehensive investigation and in-depth understanding of the adversarial robustness of LIC algorithms. To address the aforementioned issues, this paper considers the arbitrary nature of the attack direction and the uncontrollable compression ratio faced by adversaries, and presents two practical rate-distortion attack paradigms,i.e., Specific-ratio Rate-Distortion Attack (SRDA) and Agnostic-ratio Rate-Distortion Attack (ARDA). To the best of our knowledge, we are the first to conduct joint rate-distortion attacks on LIC algorithms. Using the performance variations as indicators, we evaluate the adversarial robustness of eight predominant LIC algorithms against diverse attacks. Furthermore, we propose two novel analytical tools for in-depth analysis,i.e., Entropy Causal Intervention and Layer-wise Distance Magnify Ratio, and reveal thathyperpriorsignificantly increases the bitrate andInverse Generalized Divisive Normalization (IGDN)significantly amplifies input perturbations when under attack. Lastly, we examine the efficacy of adversarial training and introduce the use of online updating for defense. By comparing their advantages and disadvantages, we provide a reference for constructing more robust LIC algorithms against the rate-distortion attacks. Qingbo Wu 0001, Lei Wang 0186, Fanman Meng, King Ngi Ngan, Li Zhuo 0001, Hongliang Li 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | SSWMNet: Solving the Speech Separation Problem While the Target is Wearing a MaskabstractSingle-channel speech separation remains one of the most challenging tasks in the field of speech signal processing. In many situations, such as during epidemics that involve respiratory diseases (e.g., COVID-19 or influenza A), individuals are required to wear masks while communicating. Is it possible to address the challenge of speech separation when the target speaker is wearing a mask? Can audio–visual approaches achieve better speech separation performance than that of audio-only approaches in scenarios where speakers are wearing masks? To address the aforementioned questions, we first construct a large-scale multimodal dataset, termed Speech Separation while Wearing a Mask (SSWM), which includes both the audio modality and the visual modality with masked faces. We explore two strategies for addressing the problem of facial occlusion. One strategy involves utilizing occluded faces—which lack critical visual cues such as mouth movements—directly as supervisory information for self-supervised speech separation; the other strategy involves the use of Wav2Lip to first generate visual information, which is then used as supervisory guidance for self-supervised speech separation. Building upon these two strategies, we propose the SSWM network (SSWMNet), which can flexibly choose to either utilize occluded facial images directly or employ Wav2Lip to generate visual information. The experimental results demonstrate that the proposed speech separation method in which Wav2Lip is used for visual information generation outperforms the approach of utilizing occluded faces directly for self-supervised speech separation. Both proposed audio–visual methods outperform the audio-only speech separation approach, which operates without the aid of visual information. Availability—SSWMNet is available at https://github.com/fanmanqian/SSWMNetwork . Fanman Meng, Kang Qin, Huazhong Shu, Lotfi Senhadji, Jiasong Wu |
ACM Trans. Internet Techn. | 1 |
| 2025 | Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive FusionabstractUnlike traditional Multimodal Class-Incremental Learning (MCIL) methods that focus only on vision and text, this paper explores MCIL across vision, audio and text modalities, addressing challenges in integrating complementary information and mitigating catastrophic forgetting. To tackle these issues, we propose an MCIL method based on multimodal pre-trained models. Firstly, a Multimodal Incremental Feature Extractor (MIFE) based on Mixture-of-Experts (MoE) structure is introduced to achieve effective incremental fine-tuning for AudioCLIP. Secondly, to enhance feature discriminability and generalization, we propose an Adaptive Audio-Visual Fusion Module (AAVFM) that includes a masking threshold mechanism and a dynamic feature fusion mechanism, along with a strategy to enhance text diversity. Thirdly, a novel multimodal class-incremental contrastive training loss is proposed to optimize cross-modal alignment in MCIL. Finally, two MCIL-specific evaluation metrics are introduced for comprehensive assessment. Extensive experiments on three multimodal datasets validate the effectiveness of our method. Zihuan Qiu, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001 |
ICASSP | 3 |
| 2025 | Cmp: Composable Meta Prompt for Sam-Based Cross-Domain Few-Shot SegmentationabstractCross-Domain Few-Shot Segmentation (CD-FSS) remains challenging due to limited data and domain shifts. Recent foundation models like the Segment Anything Model (SAM) have shown remarkable zero-shot generalization capability in general segmentation tasks, making it a promising solution for few-shot scenarios. However, adapting SAM to CD-FSS faces two critical challenges: reliance on manual prompt and limited cross-domain ability. Therefore, we propose the Composable Meta-Prompt (CMP) framework that introduces three key modules: (i) the Reference Complement and Transformation (RCT) module for semantic expansion, (ii) the Composable Meta-Prompt Generation (CMPG) module for automated meta-prompt synthesis, and (iii) the Frequency-Aware Interaction (FAI) module for domain discrepancy mitigation. Evaluations across four cross-domain datasets demonstrate CMP’s state-of-the-art performance, achieving 71.8% and 74.5% mIoU in 1-shot and 5-shot scenarios respectively. Fanman Meng, Chunjin Yang, Qingbo Wu 0001, Hongliang Li 0001 |
ICIP | 2 |
| 2025 | DPM-CLIP: Zero-Shot Multimodal Egocentric Activity Recognition based on Dual-Prediction MechanismabstractAdvancements in Zero-shot Multimodal Egocentric Activity Recognition (ZS-MM-EAR) largely rely on Vision-Language Model (VLM). However, existing methods struggle with VLM’s inadequate representation of egocentric activities, including challenges in capturing egocentric-specific features, adapting to domain shifts between egocentric video and pre-training data, and effectively leveraging complementary data such as Inertial Measurement Unit (IMU). To address these issues, we propose DPM-CLIP, a ZS-MM-EAR method tailored for vision, text and IMU modalities. Firstly, we design an attribute-driven text augmentation module that leverages a Large Language Model (LLM) to generate fine-grained textual descriptions of activities. Secondly, we construct an Instance-feature Repository (IFR) to store base class features and generate pseudo-features for novel classes through feature center migration. Finally, we introduce a dual-prediction mechanism with a prediction correction module to enhance generalization and recognition accuracy. Extensive experiments on the UESTC-MMEA-CL dataset validate the effectiveness of the proposed method. Zihuan Qiu, Mingzhou He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001 |
ICIP | 5 |
| 2025 | Your Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low BitrateabstractWith the help of powerful generative models, Semantic Image Compression (SIC) has achieved impressive performance at ultra-low bitrate. However, due to coarse-grained visual-semantic alignment and inherent randomness, the reliability of SIC is seriously concerned for reconstructing completely different object instances, even they are semantically consistent with original images. To tackle this issue, we propose a novel Referring Semantic Image Compression (RSIC) framework to improve the fidelity of user-specified content while retaining extreme compression ratios. Specifically, RSIC consists of three modules: Global Description Encoding (GDE), Referring Guidance Encoding (RGE), and Guided Generative Decoding (GGD). GDE and RGE encode global semantic information and local features, respectively, while GGD handles the non-uniformly guided generative process based on the encoded information. In this way, our RSIC achieves flexible customized compression according to user demands, which better balance the local fidelity, global realism, semantic alignment, and bit overhead. Extensive experiments on three datasets verify the compression efficiency and flexibility of the proposed method. Qingbo Wu 0001, Mingzhou He, King Ngi Ngan, Fanman Meng, Hongliang Li 0001 |
ISCAS | 7 |
| 2025 | Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationabstractEven from an early age, humans naturally adapt between exocentric (Exo) and egocentric (Ego) perspectives to understand daily procedural activities. Inspired by this cognitive ability, we propose a novel Unsupervised Ego-Exo Dense Procedural Activity Captioning (UE^2 DPAC) task, which aims to transfer knowledge from the labeled source view to predict the time segments and descriptions of action sequences for the target view without annotations. Despite previous works endeavoring to address the fully-supervised single-view or cross-view dense video captioning, they lapse in the proposed task due to the significant inter-view gap caused by temporal misalignment and irrelevant object interference. Hence, we propose a Gaze Consensus-guided Ego-Exo Adaptation Network (GCEAN) that injects the gaze information into the learned representations for the fine-grained Ego-Exo alignment. Specifically, we propose a Score-based Adversarial Learning Module (SALM) that incorporates a discriminative scoring network and compares the scores of distinct views to learn unified view-invariant representations from a global level. Then, the Gaze Consensus Construction Module (GCCM) utilizes the gaze to progressively calibrate the learned representations to highlight the regions of interest and extract the corresponding temporal contexts. Moreover, we adopt hierarchical gaze-guided consistency losses to construct gaze consensus for the explicit temporal and spatial adaptation between the source and target views. To support our research, we propose a new EgoMe-UE^2 DPAC benchmark, and extensive experiments demonstrate the effectiveness of our method, which outperforms many related methods by a large margin. Code is available at https://github.com/ZhaofengSHI/GCEAN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
ACM Multimedia | 5 |
| 2025 | DFR: A Decompose-Fuse-Reconstruct Framework for Multi-Modal Few-Shot SegmentationabstractThis paper presents DFR (Decompose, Fuse and Reconstruct), a novel framework that addresses the fundamental challenge of effectively utilizing multi-modal guidance in few-shot segmentation (FSS). While existing approaches primarily rely on visual support samples or textual descriptions, their single or dual-modal paradigms limit exploitation of rich perceptual information available in real-world scenarios. To overcome this limitation, the proposed approach leverages the Segment Anything Model (SAM) to systematically integrate visual, textual, and audio modalities for enhanced semantic understanding. The DFR framework introduces three key innovations: 1) Multi-modal Decompose: a hierarchical decomposition scheme that extracts visual region proposals via SAM, expands textual semantics into fine-grained descriptors, and processes audio features for contextual enrichment; 2) Multi-modal Contrastive Fuse: a fusion strategy employing contrastive learning to maintain consistency across visual, textual, and audio modalities while enabling dynamic semantic interactions between foreground and background features; 3) Dual-path Reconstruct: an adaptive integration mechanism combining semantic guidance from tri-modal fused tokens with geometric cues from multi-modal location priors. Extensive experiments across visual, textual, and audio modalities under both synthetic and real settings demonstrate DFR's substantial performance improvements over state-of-the-art methods. Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
MMSP | 2 |
| 2025 | MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model MergingabstractContinual model merging integrates independently fine-tuned models sequentially without access to the original training data, offering a scalable and efficient solution for continual learning. However, existing methods face two critical challenges: parameter interference among tasks, which leads to catastrophic forgetting, and limited adaptability to evolving test distributions. To address these issues, we introduce the task of Test-Time Continual Model Merging (TTCMM), which leverages a small set of unlabeled test samples during inference to alleviate parameter conflicts and handle distribution shifts. We propose MINGLE, a novel framework for TTCMM. MINGLE employs a mixture-of-experts architecture with parameter-efficient, low-rank experts, which enhances adaptability to evolving test distributions while dynamically merging models to mitigate conflicts. To further reduce forgetting, we propose Null-Space Constrained Gating, which restricts gating updates to subspaces orthogonal to prior task representations, thereby suppressing activations on old tasks and preserving past knowledge. We further introduce an Adaptive Relaxation Strategy that adjusts constraint strength dynamically based on interference signals observed during test-time adaptation, striking a balance between stability and adaptability. Extensive experiments on standard continual merging benchmarks demonstrate that MINGLE achieves robust generalization, significantly reduces forgetting, and consistently surpasses previous state-of-the-art methods by 7–9% on average across diverse task orders. Our code is available at: https://github.com/zihuanqiu/MINGLE Zihuan Qiu, Yi Xu 0008, Chiyuan He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001 |
NeurIPS | 4 |
| 2025 | Broad feature extraction and multi-directional imbalanced weighted broad learning system for the unsupervised stereo matching method
Fanman Meng, Tiejun Yang, Huifang Hou, Quan Pan 0001 |
Expert Syst. Appl. | 3 |
| 2025 | GRSDet: Learning to Generate Local Reverse Samples for Few-shot Object Detection
Hefei Mei, Taijin Zhao, Shiyuan Tang, Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 7 |
| 2025 | High efficiency deep image compression via channel-wise scale adaptive latent representation learning
Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
Signal Process. Image Commun. | 5 |
| 2025 | Optimal Transport Quantization Based on Cross-X Semantic Hypergraph Learning for Fine-Grained Image RetrievalabstractLarge-scale fine-grained image retrieval aims to learn compact discriminative feature representations based on mining the subtle distinctions between visually similar objects. However, existing fine-grained image retrieval methods focus on enhancing the attention to the discriminative regions within single images, which barely exploit the high-order relational information between the global features and local region features across different images. Thus, the over-fitting problem of complex personalized differences cannot be effectively solved. In addition, existing unconstrained vector quantization methods tend to assign unquantized feature vectors to a few major codewords, which are unable to effectively distinguish the quantized features and reduce the redundant information. To address these issues, we propose a novel optimal transport quantization method based on cross-X semantic hypergraph learning for large-scale fine-grained image retrieval. Specifically, we first introduce a cross-layer multi-scale aggregation module to extract the global features and local region features. Subsequently, we build a semantic hypergraph to model the high-order correlations between the global features and local region features extracted from different layers, different scales and different images, which can alleviate the over-fitting problem of complex personalized differences by suppressing sample-level and background noise. Moreover, we introduce an error regularization term into the progressive asymmetric quantization loss to reduce the quantization errors and preserve the semantic similarity. Finally, we attempt to introduce the code balance and uncorrelated constraints into the multi-codebook quantization framework to improve the utilization efficiency of codewords and reduce the redundant information, which can be approximated by solving the optimal transport problem. Experimental results on several fine-grained image datasets demonstrate that the proposed method outperforms the state-of-the-art fine-grained image retrieval methods. Lei Ma 0004, Yu Shi 0004, Fanman Meng, Qingbo Wu 0001, Hanyu Hong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | MCCE-REC: MLLM-Driven Cross-Modal Contrastive Entropy Model for Zero-Shot Referring Expression ComprehensionabstractZero-shot referring expression comprehension (zero-shot REC) is a crucial yet challenging task in the field of multi-modal understanding, which aims to locate an object described by a referring expression without training on task-specific datasets. Existing methods take advantage of a pre-trained CLIP model to align cropped proposal regions with referring expressions. However, our analysis reveals that this aligning way heavily biases toward certain salient visual regions due to CLIP focusing on global-level image-text matching. To mitigate this bias, we propose MCCE-REC, an MLLM-driven cross-modal contrastive entropy model for training-free zero-shot REC. Benefiting from the remarkable in-context comprehension ability of the multi-modal large language model (MLLM), we design a set of referring prompts for MLLM to generate diverse detailed informative, and contrastive cues related to referring objects. Based on these cues, on the one hand, we propose a multi-cues cross-modal interaction network, which associates the visual features and referring object textual features from multiple perspectives and perceives surrounding context object information in a parameter-free manner, avoiding bias towards salient features. On the other hand, we introduce a contrastive similarity entropy selection mechanism that compares the positive and negative cues to suppress biased regions with high similarity scores and emphasizes accurate regions correlating with referring descriptions. Extensive experiments demonstrate our MCCE-REC outperforms existing zero-shot methods by a significant margin on various REC datasets. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Cognition Transferring and Decoupling for Text-Supervised Egocentric Semantic SegmentationabstractIn this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective potential, the egocentric scenes contain dense wearer-object relations and inter-object interference. However, most recent third-view methods leverage the frozen Contrastive Language-Image Pre-training (CLIP) model, which is pre-trained on the semantic-oriented third-view data and lapses in the egocentric view due to the “relation insensitive” problem. Hence, we propose a Cognition Transferring and Decoupling Network (CTDN) that first learns the egocentric wearer-object relations via correlating the image and text. Besides, a Cognition Transferring Module (CTM) is developed to distill the cognitive knowledge from the large-scale pre-trained model to our model for recognizing egocentric objects with various semantics. Based on the transferred cognition, the Foreground-background Decoupling Module (FDM) disentangles the visual representations to explicitly discriminate the foreground and background regions to mitigate false activation areas caused by foreground-background interferential objects during egocentric relation learning. Extensive experiments on four TESS benchmarks demonstrate the effectiveness of our approach, which outperforms many recent related methods by a large margin. Code will be available athttps://github.com/ZhaofengSHI/CTDN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Geodesic-Aligned Gradient Projection for Continual Task LearningabstractDeep networks notoriously suffer from performance deterioration on previous tasks when learning from sequential tasks, i.e., catastrophic forgetting. Recent methods of gradient projection show that the forgetting is resulted from the gradient interference on old tasks and accordingly propose to update the network in an orthogonal direction to the task space. However, these methods assume the task space is invariant and neglect the gradual change between tasks, resulting in sub-optimal gradient projection and a compromise of the continual learning capacity. To tackle this problem, we propose to embed each task subspace into a non-Euclidean manifold, which can naturally capture the change of tasks since the manifold is intrinsically non-static compared to the Euclidean space. Subsequently, we analytically derive the accumulated projection between any two subspaces on the manifold along the geodesic path by integrating an infinite number of intermediate subspaces. Building upon this derivation, we propose a novel geodesic-aligned gradient projection (GAGP) method that harnesses the accumulated projection to mitigate catastrophic forgetting. The proposed method utilizes the geometric structure information on the task manifold by capturing the gradual change between the new and the old tasks. Empirical studies on image classification demonstrate that the proposed method alleviates catastrophic forgetting and achieves on-par or better performance compared to the state-of-the-art approaches. Benliu Qiu, Heqian Qiu, Haitao Wen, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | MTDA-STGCN: Modern Temporal and Dual-Attention-Based Spatiotemporal Graph Convolutional Network for 4D Trajectory PredictionabstractFour-dimensional (4D) trajectory prediction plays a critical role in modern air traffic management, enabling applications such as conflict detection, anomaly monitoring, and congestion mitigation. However, existing methods have limited information sources when modeling potential spatial correlations between aircraft in complex airspace scenarios, and their final trajectory inference ability is weak, resulting in lower prediction accuracy. Faced with these challenges, we propose Modern Temporal and Dual Attention based Spatiotemporal Graph Convolutional Network (MTDA-STGCN), which employs a self-attention mechanism to reconstruct the adjacency matrix to enhance the ability of capturing global node correlations. This adjacency matrix reconstructed with the self-attention mechanism is dynamically optimized throughout the training process of network, offering a more nuanced reflection of the inter-node relationships compared to traditional algorithms. Subsequently, our model uses graph attention to extract additional global features for modeling accuracy interactions between aircraft. Finally, the output is input into the Modern Temporal Prediction Network (MTPN) to obtain the predicted trajectory probability distribution. The experiments on real-world ADS-B datasets demonstrate that MTDA-STGCN outperforms existing 4D trajectory prediction algorithms on all datasets. The proposed dual-attention framework significantly enhances the capture of node spatial correlations, while the MTPN module effectively improves the accuracy of the predicted results. Yuheng Kuang, Shuxuan Yuan, Yuding Zhang, Fanman Meng, Zhengning Wang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding ConfusionabstractContinual learning (CL) aims to enable deep neural networks (DNNs) to learn new data without forgetting previously learned knowledge. The key to achieving this goal is to avoid confusion at the feature level, i.e., to avoid confusion within old tasks and between new and old tasks. Existing prototype-based CL methods generate pseudo features for old knowledge replay by adding Gaussian noise to the centroids of old classes. However, the distribution in the feature space exhibits anisotropy during the incremental process, which prevents the pseudo features from faithfully reproducing the distribution of old knowledge in the feature space, leading to confusion at the classification boundaries within old tasks. To address this issue, we propose the distribution-level memory recall (DMR) method, which uses a Gaussian mixture model to precisely fit the feature distribution of old knowledge at the distribution level and generate pseudo features in the next stage. Furthermore, resistance to confusion at the distribution level is crucial for multimodal learning. Multimodal imbalance, which refers to uneven optimization processes among encoders of different modalities, results in significant differences in feature responses between modalities; this exacerbates confusion within old tasks in prototype-based CL methods. Therefore, we mitigate the multimodal imbalance problem by using the intermodal guidance and intramodal mining (IGIM) method to guide weaker modalities with prior information from dominant modalities and further explore useful information within modalities. To avoid confusion between new and old tasks, we propose using the confusion index to quantitatively describe a model's ability to distinguish between new and old tasks, and we use the incremental mixup feature enhancement (IMFE) method to enhance pseudo features with new sample features, alleviating classification confusion between new and old knowledge. We conduct extensive experiments on the CIFAR100, ImageNet100, TinyImageNet, ImageNet-1K and UESTC-MMEA-CL datasets and achieve state-of-the-art results. Shaoxu Cheng, Kanglei Geng, Chiyuan He, Zihuan Qiu, Linfeng Xu 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
IEEE Trans. Multim. | 9 |
| 2025 | Cross-Modal Cognitive Consensus Guided Audio-Visual SegmentationabstractAudio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented reality, and intelligent robot systems. The pioneering work conducts this task through dense feature-level audio-visual interaction, which ignores the dimension gap between different modalities. More specifically, the audio clip could only provide aGlobalsemantic label in each sequence, but the video frame covers multiple semantic objects across differentLocalregions, which leads to mislocalization of the representationally similar but semantically different object. In this paper, we propose a Cross-modal Cognitive Consensus guided Network (C3N) to align the audio-visual semantics from the global dimension and progressively inject them into the local regions via an attention mechanism. Firstly, a Cross-modal Cognitive Consensus Inference Module (C3IM) is developed to extract a unified-modal label by integrating audio/visual classification confidence and similarities of modality-agnostic label embeddings. Then, we feed the unified-modal label back to the visual backbone as the explicit semantic-level guidance via a Cognitive Consensus guided Attention Module (CCAM), which highlights the local features corresponding to the interested object. Extensive experiments on the Single Sound Source Segmentation (S4) setting and Multiple Sound Source Segmentation (MS3) setting of the AVSBench dataset demonstrate the effectiveness of the proposed method, which achieves state-of-the-art performance. Zhaofeng Shi, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Learning With Noisy Low-Cost MOS for Image Quality Assessment via Dual-Bias CalibrationabstractLearning-based Image Quality Assessment (IQA) models have obtained impressive performance with the help of reliable subjective quality labels, where Mean Opinion Score (MOS) is the most popular choice. However, in view of the subjective bias of individual annotators, the Labor-Abundant MOS (LA-MOS) typically requires large collections of opinion scores from multiple annotators for each image, which significantly increases the learning cost. In this paper, we aim to learn robust IQA models from Low-Cost MOS (LC-MOS), which only requires very few opinion scores or even a single opinion score for each image. More specifically, we consider the LC-MOS as the noisy observation of LA-MOS and enforce the IQA model learned from LC-MOS to approach the unbiased estimation of LA-MOS. Thus, we represent the subjective bias between LC-MOS and LA-MOS, and the model bias between IQA predictions learned from LC-MOS and LA-MOS (i.e., dual-bias) as two latent variables with unknown parameters. By means of the expectation-maximization-based alternating optimization, we can jointly estimate the parameters of the dual-bias, which suppresses the misleading of LC-MOS via a gated dual-bias calibration (GDBC) module. To the best of our knowledge, this is the first exploration of robust IQA model learning from noisy low-cost labels. Theoretical analysis and extensive experiments on four popular IQA datasets show that the proposed method is robust toward different bias rates and annotation numbers and significantly outperforms the other Learning-based IQA models when only LC-MOS is available. Furthermore, we also achieve comparable performance with respect to the other models learned with LA-MOS. Lei Wang 0029, Qingbo Wu 0001, Desen Yuan, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Dual-Consistency Model Inversion for Non-Exemplar Class Incremental LearningabstractNon-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge without forgetting previously acquired ones when historical data are un-available. One of the generative NECIL methods is to in-vert the images of old classes for joint training. However, these synthetic images suffer significant domain shifts compared with real data, hampering the recognition of old classes. In this paper, we present a novel method termed Dual-Consistency Model Inversion (DCMI) to generate better synthetic samples of old classes through two pivotal consistency alignments: (1) the semantic consistency between the synthetic images and the corresponding prototypes, and (2) domain consistency between synthetic and real images of new classes. Besides, we introduce Prototypical Routing (PR) to provide task-prior information and generate unbi-ased and accurate predictions. Our comprehensive experiments across diverse datasets consistently showcase the superiority of our method over previous state-of-the-art approaches. Zihuan Qiu, Yi Xu 0008, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001 |
CVPR | 3 |
| 2024 | Prompt-Driven Referring Image Segmentation with Instance ContrastingabstractReferring image segmentation (RIS) aims to segment the target referent described by natural language. Recently, large-scale pre-trained models, e.g., CLIP and SAM, have been successfully applied in many downstream tasks, but they are not well adapted to RIS task due to inter-task differences. In this paper, we propose a new prompt-driven framework named Prompt-RIS, which bridges CLIP and SAM end-to-end and transfers their rich knowledge and powerful capabilities to RIS task through prompt learning. To adapt CLIP to pixel-level task, we first propose a Cross-Modal Prompting method, which acquires more comprehensive vision-language interaction and fine-grained text-to-pixel alignment by performing bidirectional prompting. Then, the prompt-tuned CLIP generates masks, points, and text prompts for SAM to generate more accurate mask predictions. Moreover, we further propose Instance Contrastive Learning to improve the model's discriminability to different instances and robustness to diverse languages describing the same instance. Extensive experiments demonstrate that the performance of our method outperforms the state-of-the-art methods consistently in both general and open-vocabulary settings. Chao Shang 0001, Zichen Song 0002, Heqian Qiu, Lanxiao Wang, Fanman Meng, Hongliang Li 0001 |
CVPR | 5 |
| 2024 | Vision-Sensor Attention Based Continual Multimodal Egocentric Activity RecognitionabstractContinual learning aims to equip deep neural networks (DNNs) with the capability to continuously learn new knowledge without catastrophic forgetting. Currently, there is significant attention on multimodal continual activity recognition from a egocentric perspective. However, the issue of modality imbalance can lead to exacerbated forgetting in multimodal continual learning. To address this, we propose an exemplar-free vision-sensor Attention-based Incremental Discriminability enhancement (AID) method. Firstly, we employ a Vision-Sensor attention module to enhance the time-frequency information of sensor modality and fuse them with vision modality. This alleviates the modality imbalance problem, yielding more discriminative and generalizable representations. Simultaneously, to prevent the classifier from overfitting to old class prototypes, we enhance old prototypes with features from new classes, thereby enhancing classifier discriminability. We validate the effectiveness of this method through numerous experiments with various task settings on the UESTC-MMEA-CL dataset. Shaoxu Cheng, Chiyuan He, Kailong Chen, Linfeng Xu 0001, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001 |
ICASSP | 6 |
| 2024 | DP-RSCAP: Dual Prompt-Based Scene and Entity Network for Remote Sensing Image CaptioningabstractAs a challenging task towards remote sensing image analysis, the core problem of remote sensing image captioning is how to accurately transform the vision information into text information. Existing methods usually achieve it based on the simple multi-task learning strategy or visual attention mechanism, which ignores the importance of intermediate connection information for cross-modal transformation. To solve above problem, we propose a novel dual prompt-based scene and entity network (DP-RSCap) which aims to fully utilize the ability of cross-modal alignment in vision-language model build text prior information as intermediate connection to narrow the gap between different modalities and improve the quality of caption. Specifically, we first introduce an entity-concept prompt exporter to obtain explicit entity concepts in images. Then, we design a scene class prompt generator which can predict scene class and obtain fine-grained visual semantic features. Finally, we further design a dual prompt-based caption decoder to align and merge the visual semantic feature and dual prompts information as explicit intermediate connections, which can assist in generating precise caption. Extensive experiments on the challenging RSICD demonstrate the superior ability of our model. Lanxiao Wang, Heqian Qiu, Minjian Zhang 0003, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IGARSS | 4 |
| 2024 | Robust Real-World Image Dehazing via Knowledge Guided Conditional Diffusion Model FinetuningabstractDue to the domain gap, the dehazing models trained from the synthetic images suffer poor generalization performance on real-world images. To address this issue, we pro-pose a Knowledge guided Conditional Diffusion (KCDiff) model finetuning method, which enables both the domain knowledge adaptation from the synthetic images and general knowledge guidance from the real-world images. More specifically, our KCDiff comprises two modules, i.e., the Conditional Image Generation (CIG) and Dehazing Instruction Generation (DIG). For CIG, we freeze a pre-trained latent diffusion model, add learnable conditioning control layers with Low-Rank Adaptation (LoRA) blocks, and include skip connections with zero-initialized convolutional layers, all of which play a fundamental role in image dehazing. Meanwhile, the DIG utilizes a large vision-language model LLaVA to extract the semantic content of the input hazy image and redescribe it in clear weather, which serves as the control instruction of CIG. To mitigate potential artifacts in CIG caused by misinterpretation of DIG's instructions, we further enforce depth and physical model-based reconstruction consistency constraints on both dehazing and hazy images. In the training phase, CIG is trained with the paired synthetic images to adapt the diffusion prior to the domain knowledge of image dehazing and finetuned with unpaired real-world images to suppress the domain gap with general knowledge guidance from the atmospheric scattering and depth perception. Experiments on real-world databases demonstrate the superiority of the proposed method over many state-of-the-art image dehazing models. Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Fanman Meng, Hongliang Li 0001 |
MMSP | 7 |
| 2024 | IoU-CLIP: IoU-Aware Language-Image Model Tuning for Open Vocabulary Object DetectionabstractOpen vocabulary object detection (OVD), which detects novel categories through detectors trained on base categories, has achieved remarkable advancement attributable to large-scale vision-language models, such as CLIP. The prior OVD works mainly focused on improving the classification accuracy of proposals, ignoring the ability of localization for novel categories. In this work, we propose IoU-aware language-image model tuning (IoU-CLIP) for open vocabulary object detection. Specifically, we construct a region image dataset with different IoU and adopt IoU values as labels to fine-tune the CLIP model to learn IoU-aware and class-agnostic semantic prompts and visual embeddings. The fine-tuned IoU-CLIP can predict IoU scores for proposals, which interact with classification scores. Meanwhile, IoU-aware and class-agnostic visual embeddings are utilized for box regression to enhance the generalization of the localization capability. We evaluate our method on the COCO and LVIS OVD benchmarks, outperforming the baseline (RegionCLIP) by 5.5% AP50and 5.8% AP on novel categories, respectively, achieving state-of-the-art performance. Mingzhou He, Qingbo Wu 0001, King Ngi Ngan, Fanman Meng, Heqian Qiu, Hongliang Li 0001 |
VCIP | 5 |
| 2024 | VLM-guided Explicit-Implicit Complementary novel class semantic learning for few-shot object detection
Taijin Zhao, Heqian Qiu, Lanxiao Wang, Hefei Mei, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Class similarity weighted knowledge distillation for few shot incremental learning
Feidu Akmel, Fanman Meng, Qingbo Wu 0001, Runtong Zhang, Maregu Assefa |
Neurocomputing | 2 |
| 2024 | Advancing zero-shot semantic segmentation through attribute correlations
Runtong Zhang, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, Hongliang Li 0001 |
Neurocomputing | 2 |
| 2024 | Few-shot class incremental learning via prompt transfer and knowledge distillation
Feidu Akmel, Fanman Meng, Runtong Zhang, Asebe Teka, Elias Lemuye |
Image Vis. Comput. | 2 |
| 2024 | Blessing few-shot segmentation via semi-supervised learning with noisy support images
Runtong Zhang, Hongyuan Zhu 0002, Hanwang Zhang, Chen Gong 0002, Joey Tianyi Zhou, Fanman Meng |
Pattern Recognit. | 6 |
| 2024 | CSLNSpeech: Solving the extended speech separation problem with the help of Chinese sign language
Jiasong Wu, Taotao Li, Fanman Meng, Youyong Kong, Guanyu Yang 0001, Lotfi Senhadji, Huazhong Shu |
Speech Commun. | 4 |
| 2024 | TridentCap: Image-Fact-Style Trident Semantic Framework for Stylized Image CaptioningabstractStylized image captioning (SIC) aims to generate captions with target style for images. The biggest challenge is that the collection and annotation of stylized data are pretty difficult and time-consuming. Most existing methods learn massive factual captions or additional stylized bookcorpus independently to assist in generating stylized caption, which ignore core relationships between existing image-fact-style trident data. In this paper, we propose a novel image-fact-style trident semantic framework TridentCap for stylized image captioning, which includes an image-fact semantic fusion encoder (SFE) and a trident stylization decoder (TSD). Unlike existing methods, we directly mine the core relationship in image-fact-style trident data and use factual semantic and image to build cross-modal semantic feature space, achieving the coherence between image and text. Specifically, SFE aims to learn the image-related prior language knowledge information from factual text and leverage fine-grained region-level semantic correlations of image and factual text to achieve cross-modal semantic information alignment and integration. TSD is designed to decouple the dual-source fused semantic feature based on the target style to achieve stylized caption generation. In addition, we design a pseudo labels filter (PLF) to obtain and expand massive image-fact-style trident data by building pseudo stylized annotations for all image-fact data in traditional caption datasets, which can further strengthen stylized caption learning. It is a generic algorithm to solve the problem of insufficient data and can be used into any existing stylized caption models. We conduct extensive experiments on SentiCap and FlickrStyle datasets, which achieve consistently improvement on almost all metrics. Our code will be released at: https://github.com/WangLanxiao/TridentCap_Code. Lanxiao Wang, Heqian Qiu, Benliu Qiu, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Robust Unpaired Image Dehazing via Adversarial Deformation ConstraintabstractDue to the flexible training requirement and the appealing generalization ability, unpaired image dehazing has received increasing attention in coping with real-world hazy images. However, most of the existing methods rely on the loose dehazing-hazing cycle constraint, which makes it hard to eliminate poor-quality dehazing results when using a powerful hazing network in the training process. To address this issue, this paper proposes a simple yet efficient Adversarial Deformation Constraint (ADC). More specifically, we sequentially perform two operations, i.e., dehazing and deformation, on a hazy image. In the training process, the dehazing branch is desired to be deformation-unaware, which requires that the output of these two operations remains constant regardless of their performing order. Adversarially, the deformation branch tends to maximize the difference in the outputs of these two operations when their performing orders are different. Through an additive image decomposition model, we verify that the ADC could regularize the solution space to push the dehazing error towards zero. Finally, by incorporating ADC into the common dehazing-hazing cycle constraint, we significantly improve the robustness of unpaired image dehazing. Experiments on multiple benchmark hazy image databases demonstrate the superiority of ADC over many state-of-the-art image dehazing methods. The source code of the proposed ADC-Net will be released on https://github.com/whrws/ADC-Net. Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Continual Cross-Domain Image Compression via Entropy Prior Guided Knowledge Distillation and Scalable DecodingabstractLearning based image compression has achieved impressive rate-distortion performance in recent years. However, due to the disposable learning strategy and rigid network architecture, existing methods perform poorly for compressing the images of different domains when they emerge with the expanding real-world applications, such as, natural, oil painting, medical images and so on. To cope with this open-world challenge, this paper proposes a continual cross-domain image compression method based on entropy prior guided knowledge distillation and scalable decoding network, which perform well in balancing the plasticity, stability and compatibility. Firstly, we generate pseudo-samples of old domains by reusing their entropy priors. These pseudo-samples serve as guides for knowledge distillation in the old domains, ensuring that the bit rate and reconstruction of the new model align with those of the old model. This approach assists the updated model in retaining its capability to compress and reconstruct old images. Secondly, we develop a scalable decoding network via dynamic pruning and masked recovery, which could effectively infer an old entropy decoder from the latestly updated model. It ensures that the updated model could decode image features from binary strings encoded by old entropy encoders. Experiments on five image datasets with different domains demonstrate the effectiveness of the proposed method and its superiority over representative continual learning methods. Code of the proposed method is available athttps://github.com/wuchenhaoo/Continual_Cross-domain_Image_Compression/. Qingbo Wu 0001, Rui Ma 0030, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and BeyondabstractFew-shot segmentation (FSS) aims to segment the novel class with a few annotated images. Due to CLIP's advantages of aligning visual and textual information, the integration of CLIP can enhance the generalization ability of FSS model. However, even with the CLIP model, the existing CLIP-based FSS methods are still subject to the biased prediction towards base class, which is caused by the class-specific feature level interactions. To solve this issue, we propose a visual and textual Prior Guided Mask Assemble Network (PGMA-Net). It employs a class-agnostic mask assembly process to alleviate the bias, and formulates diverse tasks into a unified manner by assembling the prior through affinity. Specifically, the class-relevant textual and visual features are first transformed to class-agnostic prior in the form of probability map. Then, a Prior-Guided Mask Assemble Module (PGMAM) including multiple General Assemble Units (GAUs) is introduced. It considers diverse and plug-and-play interactions, such as visual-textual, inter- and intra-image, training-free, and high-order ones. Lastly, to ensure the class-agnostic ability, a Hierarchical Decoder with Channel-Drop Mechanism (HDCDM) is proposed to flexibly exploit the assembled masks and low-level features, without relying on any class-specific information. It achieves new state-of-the-art results in the FSS task, with mIoU of 77.6 on$\rm{PASCAL-}5^{i}$and 59.4 on$\rm{COCO-}20^{i}$in 1-shot scenario. Beyond this, we show that without extra re-training, the proposed PGMA-Net can solve bbox-level and cross-domain FSS, co-segmentation, zero-shot segmentation (ZSS) tasks, leading an any-shot segmentation framework capable of accommodating diverse weak or pixel annotations. Fanman Meng, Runtong Zhang, Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Linfeng Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Deep Progressive Asymmetric Quantization Based on Causal Intervention for Fine-Grained Image RetrievalabstractIn the field of computer vision, fine-grained image retrieval is an extremely challenging task due to the inherently subtle intra-class object variations. In addition, the high-dimensional real-valued features extracted from large-scale fine-grained image datasets slow the retrieval speed and increase the storage cost. To solve above issues, existing fine-grained image retrieval methods mainly focus on finding more discriminative local regions for generating discriminative and compact hash codes, which achieve limited fine-grained image retrieval performance due to the large quantization errors and the confounding granularities and context of discriminative parts, i.e., the correct recognition of fine-grained objects mainly attribute to the discriminative parts and their context. To learn robust causal features and reduce the quantization errors, we propose a deep progressive asymmetric quantization (DPAQ) method based on causal intervention to learn compact and robust descriptions for fine-grained image retrieval task. Specifically, we introduce a structural causal model to learn robust casual features via causal intervention for fine-grained visual recognition. Subsequently, we design a progressive asymmetric quantization layer in the feature embedding space, which can preserve the semantic information and reduce the quantization errors sufficiently. Finally, we incorporate both the fine-grained image classification and retrieval tasks into an end-to-end deep learning architecture for generating robust and compact descriptions. Experimental results on several fine-grained image retrieval datasets demonstrate that the proposed DPAQ method performs the best for fine-grained image retrieval task and surpasses the state-of-the art fine-grained hashing methods by a large margin. Lei Ma 0004, Hanyu Hong, Fanman Meng, Qingbo Wu 0001, Jinmeng Wu |
IEEE Trans. Multim. | 3 |
| 2024 | Logit Variated Product Quantization Based on Parts Interaction and Metric Learning With Knowledge Distillation for Fine-Grained Image RetrievalabstractImage retrieval with fine-grained categories is an extremely challenging task due to the high intraclass variance and low interclass variance. Most previous works have focused on localizing discriminative image regions in isolation, but have rarely exploited correlations across the different discriminative regions to alleviate intraclass differences. In addition, the intraclass compactness of embedding features is ensured by extra regularization terms that only exist during the training phase, which appear to generalize less well in the inference phase. Finally, the information granularity of the distance measure should distinguish subtle visual differences and the correlation between the embedding features and the quantized features should be maximized sufficiently. To address the above issues, we propose a logit variated product quantization method based on part interaction and metric learning with knowledge distillation for fine-grained image retrieval. Specifically, we introduce a causal context module into the deep navigator to generate discriminative regions and utilize a channelwise cross-part fusion transformer to model the part correlations while alleviating intraclass differences. Subsequently, we design a logit variation module based on a weighted sum scheme to further reduce the intraclass variance of the embedding features directly and enhance the learning power of the quantization model. Finally, we propose a novel product quantization loss based on metric learning and knowledge distillation to enhance the correlation between the embedding features and the quantized features and allow the quantization features to learn more knowledge from the embedding features. The experimental results on several fine-grained datasets demonstrate that the proposed method is superior to state-of-the-art fine-grained image retrieval methods. Lei Ma 0004, Hanyu Hong, Fanman Meng, Qingbo Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | CrowdCaption++: Collective-Guided Crowd Scenes CaptioningabstractCrowd scenes analysis plays an important role in various fields, including public security, smart cities, and intelligent transportation systems. However, traditional crowd scenes captioning methods mainly focus on a single and prominent crowd collective, which limits their ability to describe the different crowd collectives in complex crowd scenes. To address this issue, we propose a collective-guided crowd scenes captioning model (CrowdCaption++) to explore a more comprehensive and detailed description. We design a crowd features encoder (CFE) including double-query features encoder and foreground crowd features encoder, which uses double-query attention module (DQ-ATT) to capture more representative visual features and extracts foreground crowd features to avoid interference from background for collectives prediction. Moreover, we build a collective-guided captioning decoder (CCD) to generate captions of different crowd collectives without requiring extra alignment between crowd collectives and captions. To achieve this, we first design a crowd collectives predictor to identify multiple potential crowd collectives and create crowd collectives guidance information. Finally, we use the crowd collectives guidance information to merge useful visual features and further generate corresponding caption. We evaluate our approach on the latest crowd scenes dataset CrowdCaption and demonstrate that our model can achieve a comprehensive understanding and describe the different crowd collectives in complex crowd scenes. Lanxiao Wang, Hongliang Li 0001, Minjian Zhang 0003, Heqian Qiu, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Towards Continual Egocentric Activity Recognition: A Multi-Modal Egocentric Activity Dataset for Continual LearningabstractWith the rapid development of wearable cameras, it is now feasible to considerably increase the collection of egocentric video for first-person visual perception. However, the development is hindered by a shortage of multi-modal egocentric activity datasets. Furthermore, the catastrophic forgetting problem of multimodal continual activity learning, as a branch of continual learning, has not been thoroughly explored, which makes accumulating a larger collection of multi-modal activity data more urgent. To address this shortage, we propose a multi-modal egocentric activity dataset for continual activity learning named UESTC-MMEA-CL in this paper. The dataset is collected using our self-developed glasses with a first-person camera and wearable sensors, and it contains synchronized data of video, accelerometers, and gyroscopes for 32 types of daily activities performed by 10 participants who wore our glasses. Statistical analysis of the sensor data is given to show the auxiliary effects of activity recognition. We report the results of egocentric activity recognition of three modalities (RGB, acceleration, and gyroscope) separately and jointly on a base network architecture. We thoroughly evaluated four baseline methods with different multimodal combinations to explore the catastrophic forgetting in continual learning on UESTC-MMEA-CL. We hope that the UESTC-MMEA-CL dataset can act as a facilitator for future studies on continual learning for first-person activity recognition in wearable applications. You can download preliminary data fromhttps://ivipclab.github.io/publication_uestc-mmea-cl/mmea-cl. The data is currently used to solve the problems of multimodal continual learning of activities. Linfeng Xu 0001, Qingbo Wu 0001, Lili Pan 0001, Fanman Meng, Hongliang Li 0001, Chiyuan He, Hanxin Wang, Shaoxu Cheng |
IEEE Trans. Multim. | 4 |
| 2024 | Learning Offset Probability Distribution for Accurate Object DetectionabstractObject detection combines object classification and object localization problems. Current object detection methods heavily depend on regression networks to locate objects, which are optimized with various regression loss functions to predict offsets between candidate boxes and objects. However, these regression losses are difficult to assign the appropriate penalties for samples with large offset errors, resulting in suboptimal regression networks and inaccurate object offsets. In this article, we consider object location as offset bin classification problem, and propose a distance-aware offset bin classification network optimized with multiple binary cross entropy losses to learn various offset probability distribution, including single label distribution and distance-aware label distribution. On one hand, it provides gradient contributions for different samples based on the bounded probability instead of previous incalculable offset error. On the other hand, it explores the distance correlations between discrete offset bins to facilitate network learning. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin, in which the probability should be higher for the offset bin closer to the target offsets, and vice versa. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the precision of prediction. We conduct extensive experiments to evaluate the effectiveness of our method. In addition, our method can be conveniently and flexibly inserted into existing object detection methods, which consistently achieves a large gain based on popular anchor-based and anchor-free methods on the PASCAL VOC, MS-COCO, KITTI, and CrowdHuman datasets. Code will be released at: https://github.com/QiuHeqian/DBC . Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | CafeBoost: Causal Feature Boost to Eliminate Task-Induced Bias for Class Incremental LearningabstractContinual learning requires a model to incrementally learn a sequence of tasks and aims to predict well on all the learned tasks so far, which notoriously suffers from the catastrophic forgetting problem. In this paper, we find a new type of bias appearing in continual learning, coined as task-induced bias. We place continual learning into a causal framework, based on which we find the task-induced bias is reduced naturally by two underlying mechanisms in task and domain incremental learning. However, these mechanisms do not exist in class incremental learning (CIL), in which each task contains a unique subset of classes. To eliminate the task-induced bias in CIL, we devise a causal intervention operation so as to cut off the causal path that causes the task-induced bias, and then implement it as a causal debias module that transforms biased features into unbiased ones. In addition, we propose a training pipeline to incorporate the novel module into existing methods and jointly optimize the entire architecture. Our overall approach does not rely on data replay, and is simple and convenient to plug into existing methods. Extensive empirical study on CIFAR-100 and ImageNet shows that our approach can improve accuracy and reduce forgetting of well-established methods by a large margin. Benliu Qiu, Hongliang Li 0001, Haitao Wen, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Lili Pan 0001 |
CVPR | 6 |
| 2023 | Incrementer: Transformer for Class-Incremental Semantic Segmentation with Knowledge Distillation Focusing on Old ClassabstractClass-incremental semantic segmentation aims to incrementally learn new classes while maintaining the capability to segment old ones, and suffers catastrophic forgetting since the old-class labels are unavailable. Most existing methods are based on convolutional networks and prevent forgetting through knowledge distillation, which (1) need to add additional convolutional layers to predict new classes, and (2) ignore to distinguish different regions corresponding to old and new classes during knowledge distillation and roughly distill all the features, thus limiting the learning of new classes. Based on the above observations, we propose a new transformer framework for class-incremental semantic segmentation, dubbed Incrementer, which only needs to add new class tokens to the transformer decoder for new-class learning. Based on the Incrementer, we propose a new knowledge distillation scheme that focuses on the distillation in the old-class regions, which reduces the constraints of the old model on the new-class learning, thus improving the plasticity. Moreover, we propose a class deconfusion strategy to alleviate the overfitting to new classes and the confusion of similar classes. Our method is simple and effective, and extensive experiments show that our method outperforms the SOTAs by a large margin (5~15 absolute points boosts on both Pascal VOC and ADE20k). We hope that our Incrementer can serve as a new strong pipeline for class-incremental semantic segmentation. Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Heqian Qiu, Lanxiao Wang |
CVPR | 3 |
| 2023 | Instance-Wise Adaptive Tuning and Caching for Vision-Language ModelsabstractLarge-scale vision-language models (LVLMs) pre-trained on massive image-text pairs have achieved remarkable success in visual representations. However, existing paradigms to transfer LVLMs to downstream tasks encounter two primary challenges. Firstly, the text features remain fixed after being calculated and cannot be adjusted according to image features, which decreases the model’s adaptability. Secondly, the model’s output solely depends on the similarity between the text and image features, leading to excessive reliance on LVLMs. To address these two challenges, we introduce a novel two-branch model named the Instance-Wise Adaptive Tuning and Caching (ATC). Specifically, one branch implements our proposed ConditionNet, which guides image features to form an adaptive textual cache that adjusts based on image features, achieving instance-wise inference and improving the model’s adaptability. The other branch introduces the similarities between images and incorporates a learnable visual cache, designed to decouple new and previous knowledge, allowing the model to acquire new knowledge while preserving prior knowledge. The model’s output is jointly determined by the two branches, thus overcoming the limitations of existing methods that rely solely on LVLMs. Additionally, our method requires limited computing resources to tune parameters, yet outperforms existing methods on 11 benchmark datasets. Chunjin Yang, Fanman Meng, Runtong Zhang |
ECAI | 2 |
| 2023 | MFAT: A Multi-Level Feature Aggregated Transformer for Person Re-IdentificationabstractRecently, with the development of the Transformer, re-identification (ReID) has great success in various applications. Existing works prefer to utilize the Transformer’s highest-level information as its discriminative feature, which focuses on a few concentrated parts or areas. However, in ReID filed, under such various scenes and camera views, only using a few concentrated parts to distinguish the query person is insufficient. Meanwhile, we find that Transformer’s lower-level information is also helpful for the recognition accuracy of the query person, especially, when the scene changes greatly. Therefore, we propose a Multi-level Feature Aggregated Transformer for person re-identification (MFAT) with high performance. To aggregate multi-level information, two novel modules are carefully designed. (i) The Global Content and Structure Aggregation (GCSA) module is proposed to aggregate multi-level information in a global manner. (ii) The Local Convolution Aggregation (LCA) module which consists of a series of convolutional blocks, is introduced to aggregate multi-level features with local operations. To the best of our knowledge, this is the first work to aggregate multi-level features with a Transformer backbone for person ReID task. Experiment results show that our method has achieved state-of-the-art on three person ReID benchmarks, with both Pyramid Vision Transformer (PVT) and Vision Transformer (ViT) backbones. Bowen Tan, Linfeng Xu 0001, Zihuan Qiu, Qingbo Wu 0001, Fanman Meng |
ICASSP | 5 |
| 2023 | The Elliptic Energy Loss for Rotated Object Detection in Aerial ImagesabstractRotated object detection is a promising yet challenging task in computer vision. Existing algorithms mainly train the rotated detector by the Ln-norm loss, which is inconsistent with the evaluation metric of Intersection over Union (IoU). However, the concise and efficient solution using the loss based on IoU between oriented boxes is hindered by its non-differentiability. In this paper, we propose Elliptic Energy Loss, a differentiable loss based on curve energy, to fit with the evaluation metric IoU. Specifically, given a pair of predicted and ground truth boxes, we first convert them to curve representations using the elliptic transformation. Then, the curve energy is calculated to measure the similarity between the predicted and ground truth curves. Finally, the curve energy is used as regression loss to optimize rotated detectors. We conduct experiments with different detectors on DOTA and HRSC2016 datasets, which demonstrate that the performance is significantly improved by our proposed loss. The code is available at https://github.com/zhangc-uestc/EEL. Kunming Luo, Fanman Meng, Qingbo Wu 0001 |
ICIP | 3 |
| 2023 | Semi-Supervised Few-Shot Segmentation with Noisy Support ImagesabstractMotivated by the semi-supervised learning that uses the unlabeled data and pseudo annotations to improve the image classification, this paper proposes a new semi-supervised few-shot segmentation (FSS) framework of which the training process uses not only the annotated images, but also the unlabeled images, e.g. images from other available datasets, to enhance the training of the FSS model. Furthermore, in the test phase, more support images and pseudo-annotations can also be generated by the proposed framework to enrich the support set of novel classes and therefore benefit the inference. However, unlabeled images are not a free lunch. The noisy intra-class samples and inter-class samples existed in the unlabeled images as well as the interferences of the bad quality of pseudo annotations make it difficult to utilize the correct images and pseudo annotations for a certain class. To this end, we further propose a ranking algorithm consisting of an inter-class confidence term and an intra-class confidence term to efficiently utilize the pseudo annotations of the class with high quality. Extensive experiments on COCO-20idataset demonstrate that the proposed semi-supervised FSS framework is superior to many state-of-the-art methods. Runtong Zhang, Hongyuan Zhu 0002, Hanwang Zhang, Chen Gong 0002, Joey Tianyi Zhou, Fanman Meng |
ICIP | 6 |
| 2023 | ISM-Net: Mining incremental semantics for class incremental learning
Zihuan Qiu, Linfeng Xu 0001, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 5 |
| 2023 | GFR: Generic feature representations for class incremental learningabstractClass incremental learning (CIL) aims to continuously learn new classes while maintaining discrimination for old classes with sequentially coming data. Due to the lack of old-class samples, existing CIL methods fail to learn discriminative representations for both old and new classes simultaneously, resulting in a severe performance drop in old classes, which is the well-known catastrophic forgetting phenomenon. Different from most existing works, we facilitate CIL by learning generic feature representations that perform well in seen and unseen classes. Specifically, we prove that representations with a substantial number of significant singular values benefit CIL via better old knowledge reservation. However, the overly uniform singular value spectrum will hurt the discrimination of current tasks. Furthermore, we propose that increasing the embedding dimension can enhance the number of significant singular values and validate this assumption from two perspectives: adopting different pooling techniques and devising a wider network. Meanwhile, we also prove that satisfactory current task accuracy and old knowledge reservation can be achieved simultaneously. Finally, the simple yet effective generic feature representation regulation (GFR) is devised and incorporated into two baselines. Extensive experiments are conducted on CIFAR100, ImageNet-Subset, and ImageNet. The results show that the proposed method boosts the performance of both baselines with a large margin (2.00%-9.58% on CIFAR100, 0.68%-7.10% on ImageNet-Subset and 1.18%-5.04% on ImageNet) which outperforms existing SOTAs. Linfeng Xu 0001, Zihuan Qiu, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 5 |
| 2023 | GAB-Net: A Robust Detector for Remote Sensing Object Detection Under Dramatic Scale Variation and Complex BackgroundsabstractDetecting objects in remote sensing images (RSIs), characterized by dramatic scale variation and complex backgrounds, has always been a challenging problem. These challenges can be further summarized into three aspects: 1) scale variation among objects; 2) feature fusion misalignment due to the semantic gap between adjacent feature layers and noise from backgrounds; and 3) boundary uncertainty under ambiguous and complex backgrounds. To alleviate these problems, we first utilize a global–local feature enhancement module (GLFEM) to capture local features with multiple receptive fields through cheap pooling operation and obtain global features through nonlocal block, thus alleviating the scale variation issues. Subsequently, attentional feature fusion alignment (AFFA) module is designed to align adjacent feature levels in the feature pyramid from pixel and channel levels. Finally, boundary-uncertainty aware head (BUAH) with distribution focal loss (DFL) is adopted to solve the boundary uncertainty problems. After fusing GLFEM, AFFA, and BUAH modules, we obtain GAB-Net. GAB-Net outperforms state-of-the-art methods on the Dior and NWPU VHR-10 datasets, achieving mAP scores of 73.8% and 89.8%, respectively, without adding high computational costs. The code is available at:https://github.com/Hong-yu-Zhang/GAB-Net. Yunbo Rao, Jie Shao 0001, Fanman Meng, Naveed Ahmad 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Multi-directional broad learning system for the unsupervised stereo matching method
Niu Ying, Fanman Meng, Tiejun Yang, Xiaozhen Ren, Cao Kun |
Pattern Recognit. | 3 |
| 2023 | Disturbed Augmentation Invariance for Unsupervised Visual Representation LearningabstractContrastive learning has gained great prominence recently, which achieves excellent performance by simple augmentation invariance. However, the simple contrastive pairs suffer from lacking of diversity due to the mechanical augmentation strategies. In this paper, we propose Disturbed Augmentation Invariance (DAI for abbreviation), which constructs disturbed contrastive pairs by generating appropriate disturbed views for each augmented view in the feature space to increase the diversity. In practice, we establish a multivariate normal distribution for each augmented view, whose mean is corresponding augmented view and covariance matrix is estimated from its nearest neighbors in the dataset. Then we sample random vectors from this distribution as the disturbed views to construct disturbed contrastive pairs. In order to avoid extra computational cost with the increase of disturbed contrastive pairs, we utilize an upper bound of the trivial disturbed augmentation invariance loss to construct the DAI loss. In addition, we propose Bottleneck version of Disturbed Augmentation Invariance (BDAI for abbreviation) inspired by the Information Bottleneck principle, which further refines the extracted information and learns a compact representation by additionally increasing the variance of the original contrastive pair. In order to make BDAI work effectively, we design a statistical strategy to control the balance between the amount of the information shared by all disturbed contrastive pairs and the compactness of the representation. Our approach gets a consistent improvement over the popular contrastive learning methods on a variety of downstream tasks, e.g. image classification, object detection and instance segmentation. Haoyang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Heqian Qiu, Xiaoliang Zhang 0002, Fanman Meng, Taijin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Cross-Modal Recurrent Semantic Comprehension for Referring Image SegmentationabstractReferring image segmentation aims to segment the target object from the image according to the description of language expression. Due to the diversity of language expressions, word sequences in different orders often express different semantic information. The previous methods focus more on matching different words to different visual regions in the image separately, ignoring the global semantic understanding of language expression based on the sequence structure. To address this problem, we redesign a new recurrent network structure for referring image segmentation, called Cross-Modal Recurrent Semantic Comprehension Network (CRSCNet), to obtain a more comprehensive global semantic understanding through iterative cross-modal semantic reasoning. Specifically, in each iteration, we first propose a Dynamic SepConv to extract relevant visual features guided by language and further propose Language Attentional Feature Modulation to improve the feature discriminability, then propose a Cross-Modal Semantic Reasoning module to perform global semantic reasoning by capturing both linguistic and visual information, and finally updates and corrects the visual features of the predicted object based on semantic information. Moreover, we further propose a Cross-Modal ASPP to capture richer visual information referred to in the global semantics of the language expression from larger receptive fields. Extensive experiments demonstrate that our proposed network significantly outperforms previous state-of-the-art methods on multiple datasets. Chao Shang 0001, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, Taijin Zhao, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Task-Specific Loss for Robust Instance Segmentation With Noisy Class LabelsabstractDeep learning methods have achieved significant progress in the presence of correctly annotated datasets in instance segmentation. However, object classes in large-scale datasets are sometimes ambiguous, which easily causes confusion. Besides, limited experience and knowledge of annotators can lead to mislabeled object semantic classes. To solve this issue, a novel method is proposed in this paper, which considers different roles of noisy class labels in different sub-tasks. Our method is based on two basic observations: firstly, the foreground-background annotation of a sample is correct even though its class label is noisy. Secondly, symmetric loss benefits the model robustness to noisy labels but harms the learning of hard samples, while cross entropy loss is the opposite. Based on the two basic observations, in the foreground-background sub-task, cross entropy loss is used to fully exploit correct gradient guidance. In the foreground-instance sub-task, symmetric loss is used to prevent incorrect gradient guidance provided by noisy class labels. Furthermore, we apply contrastive self-supervised loss to update features of all foreground, to compensate for insufficient guidance provided by partially correct labels especially in the highly noisy setting. Extensive experiments conducted with three popular datasets (i.e., Pascal VOC, Cityscapes and COCO) have demonstrated the effectiveness of our method in a wide range of noisy class label scenarios. Longrong Yang, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | DRDet: Dual-Angle Rotated Line Representation for Oriented Object DetectionabstractIn aerial scenes, oriented object detection is sensitive to the orientation of objects, which makes the formulation of orientation-aware object representation become a critical problem. Existing methods mostly adopt rectangle anchor or discrete points as object representation, which may lead to the feature aliasing between overlapping objects and ignore the orientation information of objects. To solve these issues, we propose a novel anchor-free oriented object detection network named DRDet, which adopts Dual-angle Rotated Lines (DRL) as object representation. Different from other object representations, DRL can adaptively rotate and extend to the boundary of the object according to its orientation and shape, which explicitly introduces the orientation information into the formulation of object representation. And it can adaptively cope with the geometric deformation of objects. Based on the dual-angle rotated lines, we design an Orientation-guided Feature Encoder (OFE) to encode discriminant object feature along each rotated line, respectively. Instead of encoding rectangle feature, the OFE module adopts line features for orientation-guided feature encoding, which can alleviate the feature aliasing between neighboring objects or background. To further enhance the flexibility of dual-angle rotated lines, we design a Dual-angle Decoder (DD) that predicts two angle offsets according to the orientation-guided feature and converts the angle offsets and regression offsets into dual-angle rotated line representation, which can help to guide the adaptive rotation of each rotated line, respectively. Our proposed method achieves consistent improvement on both DOTA and HRSC2016 datasets. Extensive experimental results verify the effectiveness of our method in oriented object detection. Minjian Zhang 0003, Heqian Qiu, Hefei Mei, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Unsupervised Visual Representation Learning via Multi-Dimensional Relationship AlignmentabstractRecently, contrastive learning based on augmentation invariance and instance discrimination has made great achievements, owing to its excellent ability to learn beneficial representations without any manual annotations. However, the natural similarity among instances conflicts with instance discrimination which treats each instance as a unique individual. In order to explore the natural relationship among instances and integrate it into contrastive learning, we propose a novel approach in this paper, Relationship Alignment (RA for abbreviation), which forces different augmented views of current batch instances to main a consistent relationship with other instances. In order to perform RA effectively in existing contrastive learning framework, we design an alternating optimization algorithm where the relationship exploration step and alignment step are optimized respectively. In addition, we add an equilibrium constraint for RA to avoid the degenerate solution, and introduce the expansion handler to make it approximately satisfied in practice. In order to better capture the complex relationship among instances, we additionally propose Multi-Dimensional Relationship Alignment (MDRA for abbreviation), which aims to explore the relationship from multiple dimensions. In practice, we decompose the final high-dimensional feature space into a cartesian product of several low-dimensional subspaces and perform RA in each subspace respectively. We validate the effectiveness of our approach on multiple self-supervised learning benchmarks and get consistent improvements compared with current popular contrastive learning methods. On the most commonly used ImageNet linear evaluation protocol, our RA obtains significant improvements over other methods, our MDRA gets further improvements based on RA to achieve the best performance. The source code of our approach will be released soon. Haoyang Cheng, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Xiaoliang Zhang 0002, Fanman Meng, King Ngi Ngan |
IEEE Trans. Image Process. | 6 |
| 2023 | Forgetting to Remember: A Scalable Incremental Learning Framework for Cross-Task Blind Image Quality AssessmentabstractRecent years have witnessed the great success of blind image quality assessment (BIQA) in various task-specific scenarios, which present invariable distortion types and evaluation criteria. However, due to the rigid structure and learning framework, they cannot apply to the cross-task BIQA scenario, where the distortion types and evaluation criteria keep changing in practical applications. This paper proposes a scalable incremental learning framework (SILF) that could sequentially conduct BIQA across multiple evaluation tasks with limited memory capacity. More specifically, we develop a dynamic parameter isolation strategy to sequentially update the task-specific parameter subsets, which are non-overlapped with each other. Each parameter subset is temporarily settled toRememberone evaluation preference toward its corresponding task, and the previously settled parameter subsets can be adaptively reused in the following BIQA to achieve better performance based on the task relevance. To suppress the unrestrained expansion of memory capacity in sequential tasks learning, we develop a scalable memory unit by gradually and selectively pruning unimportant neurons from previously settled parameter subsets, which enable us toForgetpart of previous experiences and free the limited memory capacity for adapting to the emerging new tasks. Extensive experiments on eleven IQA datasets demonstrate that our proposed method significantly outperforms the other state-of-the-art methods in cross-task BIQA. The source code of the proposed method is available atgithub.com/maruiperfect/SILF. Rui Ma 0030, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | What Happens in Crowd Scenes: A New Dataset About Crowd Scenes for Image CaptioningabstractMaking machines endowed with eyes and brains to effectively understand and analyze crowd scenes is of paramount importance for building a smart city to serve people. This is of far-reaching significance for the guidance of dense crowds and accident prevention, such as crowding and stampedes. As a typical multimodal scene understanding task, image captioning has always attracted widespread attention. However, crowd scene understanding captioning is rarely studied due to the unobtainability of related datasets. Therefore, it is difficult to know what happens in crowd scenes. In order to fill this research gap, we propose a crowd scenes caption dataset named CrowdCaption which has the advantages of crowd-topic scenes, comprehensive and complex caption descriptions, typical relationships and detailed grounding annotations. The complexity and diversity of the descriptions and the specificity of the crowd scenes make this dataset extremely challenging to most current methods. Thus, we propose a Multi-hierarchical Attribute Guided Crowd Caption Network (MAGC) based on crowd objects, actions, and status (such as position, dress, posture, etc.) aiming to generate crowd-specific detailed descriptions. We conduct extensive experiments on our CrowdCaption dataset, and our proposed method reaches the state-of-the-art (SoTA) performance. We hope the CrowdCaption dataset can assist future studies related to crowd scenes in the multimodal domain. Lanxiao Wang, Hongliang Li 0001, Wenzhe Hu, Xiaoliang Zhang 0002, Heqian Qiu, Fanman Meng, Qingbo Wu 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Bias-Correction Feature Learner for Semi-Supervised Instance SegmentationabstractInstance segmentation is heavily reliant on large-scale annotated datasets to yield an ideal accuracy. However, annotated data are difficult to collect. To expand the annotated data, a straightforward idea is to introduce semi-supervised learning, which uses a trained model to obtain initial proposals on unlabeled images and then use initial proposals to generate pseudo labels. However, existing methods inevitably introduce the bias for the model learning, i.e., the foreground in initial low-confident proposals (low-confident foreground) is arbitrarily assigned as background. This bias makes the foreground and background closer in the feature space, which degenerates the model accuracy. To address this issue, this paper discards incorrect supervision and designs a bias-correction feature learner. Specifically, on the one hand, low-confident foreground does not participate in supervised learning. On the other hand, we extract possible foreground regions from all initial proposals to construct high-quality positive pairs which depict objects of the same category in contrastive learning. Then, positive pairs are pulled closer in the feature space. This helps models extract closely clustered foreground features. Experimental results demonstrate the effectiveness of our method on the public datasets (i.e., COCO, Cityscapes and Pascal VOC). Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu, Linfeng Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | RefCrowd: Grounding the Target in Crowd with Referring ExpressionsabstractCrowd understanding has aroused the widespread interest in vision domain due to its important practical significance. Unfortunately, there is no effort to explore crowd understanding in multi-modal domain that bridges natural language and computer vision. Referring expression comprehension (REF) is such a representative multi-modal task. Current REF studies focus more on grounding the target object from multiple distinctive categories in general scenarios. It is difficult to applied to complex real-world crowd understanding. To fill this gap, we propose a new challenging dataset, called RefCrowd, which towards looking for the target person in crowd with referring expressions. It not only requires to sufficiently mine natural language information, but also requires to carefully focus on subtle differences between the target and a crowd of persons with similar appearance, so as to realize fine-grained mapping from language to vision. Furthermore, we propose a Fine-grained Multi-modal Attribute Contrastive Network (FMAC) to deal with REF in crowd understanding. It first decomposes the intricate visual and language features into attribute-aware multi-modal features, and then captures discriminative but robustness fine-grained attribute features to effectively distinguish these subtle differences between similar persons. The proposed method outperforms existing state-of-the-art (SoTA) methods on our RefCrowd dataset and existing REF datasets. In addition, we implement an end-to-end REF toolbox for the deeper research in multi-modal domain. Our dataset and code can be available at: https://qiuheqian.github.io/datasets/refcrowd/. Heqian Qiu, Hongliang Li 0001, Taijin Zhao, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng |
ACM Multimedia | 6 |
| 2022 | Dynamic Perceptual Quality Ranking based Autofocus Method for ProjectorabstractProjectors are widespread in families, schools and companies nowadays. Common projectors generally use laser ranging to realize autofocus, yet this kind of autofocus is limited by the highly-cost temperature-sensitive devices, so long-term operation in projectors “high-temperature environment may lead to autofocus failure. To find an more efficient and lower-cost solution, we made an autofocus dataset, which contains 8 category folders with a total of 80 videos, each containing 174 pictures with labels. For the autofocus task, the most intuitive idea is to use an image quality assessment algorithm to score each image with the best value corresponding to the best focal length, but the fact is that the image quality assessment algorithm will fail due to the large number of categories and the large span of video contents. In this paper, we proposed a light-weight dynamic perceptual quality ranking based autofocus method to achieve fascinate results on the dataset. Lanjiang Wang, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001 |
MMSP | 3 |
| 2022 | Instance-level Context Attention Network for instance segmentation
Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Heqian Qiu, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
Neurocomputing | 3 |
| 2022 | Real-time panoptic segmentation with relationship between adjacent pixels and boundary prediction
Xiaoliang Zhang 0002, Hongliang Li 0001, Lanxiao Wang, Haoyang Cheng, Heqian Qiu, Wenzhe Hu, Fanman Meng, Qingbo Wu 0001 |
Neurocomputing | 7 |
| 2022 | Category boundary re-decision by component labels to improve generation of class activation map
Runtong Zhang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 2 |
| 2022 | ASFlow: Unsupervised Optical Flow Learning With Adaptive Pyramid SamplingabstractWe present an unsupervised optical flow estimation method by proposing an adaptive pyramid sampling in the deep pyramid network. Specifically, in the pyramid downsampling, we propose a Content-Aware Pooling (CAP) module, which promotes local feature gathering by avoiding cross region pooling, so that the learned features become more representative. In the pyramid upsampling, we propose an Adaptive Flow Upsampling (AFU) module, where cross edge interpolation can be avoided, producing sharp motion boundaries. Equipped with these two modules, our method achieves the best performance for unsupervised optical flow estimation on multiple leading benchmarks, including MPI-Sintel, KITTI 2012 and KITTI 2015. Particularly, we achieve EPE=1.5 on KITTI 2012 and F1=9.67% KITTI 2015, which outperform the previous state-of-the-art methods by 16.7% and 13.1%, respectively. Shuaicheng Liu, Kunming Luo, Ao Luo, Chuan Wang 0001, Fanman Meng, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | POS-Trends Dynamic-Aware Model for Video CaptionabstractVideo caption aims to generate descriptive sentences about the video, and the most critical problem is how to achieve accurate word prediction with standardized and coherent syntax structure, which requires the model to thoroughly understand video content and precisely map them into corresponding sentence components. Many existing methods usually fuse different video features into a single visual feature for generating sentences. However, they ignore the word dataset prior information in the annotations (such as Part-Of-Speech) and they also ignore the association between sentence components and types of visual features. To solve these problems, we propose a POS-trends dynamic-aware model (PDA) to fully exploit the word dataset prior information in the captions to predict POS tag, so as to assist generating captions. We propose a POS feature extraction (PFE) module to use different filters to extract different POS-trends features, predict POS tags and fuse visual features. Furthermore, we propose a visual-dynamic-aware (VDA) module to dynamically adjust the mapping way of words and supplement the visual information into the local features. The fusion features provide directional visual information to generate correct words, and the predicted POS tags to guide the decoding process to generate a more standardized and coherent syntax structure. A large number of experiments based on MSVD, MSR-VTT and VATEX demonstrated that our method outperforms the state-of-the-art methods in BLEU-4, ROUGE-L, METEOR, CIDEr. Code can be available at:https://github.com/WangLanxiao/PDA-for-video-caption. Lanxiao Wang, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Segmenting Beyond the Bounding Box for Instance SegmentationabstractInstance segmentation needs to locate all instances in an image correctly and segment each instance precisely. Currently, the most dominant methods for instance segmentation take object detection as a pre-task. However, they rely on the accuracy of object detection incredibly. If the pre-task cannot predict an accurate bounding box, the performance of instance segmentation will degenerate. In this paper, we present a novel method for instance segmentation to solve this problem, which is calledSegmentingBeyond theBoundingBox (S3B-Net). Our S3B-Net designs a sub-network to help instance segmentation methods based on object detection to segment the part of an instance beyond the bounding box. Specifically, the sub-network first predicts a two-dimensional pixel embedding for each pixel. Then, the Gaussian function is employed to calculate a pixel’s probability belongs to a corresponding instance according to the two-dimensional pixel embedding. Finally, the output of the sub-network combines with the output of instance segmentation based on object detection to generate a more precise instance mask. Our sub-network can easily extend on the existing instance segmentation method based on object detection to segment instance beyond the bounding box. We do our experiments on dominant instance segmentation datasets, such as the COCO dataset and Cityscapes dataset. The results show that our method can achieve 6.8 points gain compared with the baseline Mask R-CNN with ResNet-50-FPN in Cityscapes datasets, and 1.7 points gain with ResNet-101-FPN-DCN in COCO datasets. Our S3B-Net outperforms the previous state-of-the-art instance segmentation method, which proves our method is competitive. The source code of our method will be made available. Xiaoliang Zhang 0002, Hongliang Li 0001, Fanman Meng, Zichen Song 0002, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Bal-R$^2$CNN: High Quality Recurrent Object Detection With Balance OptimizationabstractIt is a common practice to refine object detection results using recurrent detection paradigm. We evaluate the recurrent detection on Faster R-CNN, but the improvement is far away from expected. We consider that the performance bottleneck is fromimbalance optimizationcaused by the biased distribution of training data. Low-IoU-skewed RPN proposals could suppress the contribution of High-IoU examples at the training stage. Besides, data imbalance and statistical discrepancy on regression targets between low-IoU and high-IoU examples are not considered in the regression task; this design could impede localization quality. In this work, we propose Bal-R$^2$CNN for high-quality recurrent object detection. There are two new components in Bal-R$^2$CNN.Self-iteration box samplingcollects object boxes from recurrent steps and increases the number of high-IoU training examples.IoU-sensitive bounding-box regressionsends proposal boxes with different IoUs to specified regression branches for more accurate bounding-box prediction. Both two new components could inducebalanced optimizationand be helpful. With the resulting Bal-R$^2$CNN detector, evaluation on PASCAL VOC and MSCOCO reveal that our method has a significant improvement on the existing solution and could reach a better performance than several state-of-the-art methods. Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Multim. | 4 |
| 2021 | Remember and Reuse: Cross-Task Blind Image Quality Assessment via Relevance-aware Incremental LearningabstractExisting blind image quality assessment (BIQA) methods have made great progress in various task-specific applications, including the synthetic, authentic, or over-enhanced distortion evaluations. However, limited by the static model and once-for-all learning strategy, they failed to perform the cross-task evaluations in many practical applications, where diverse evaluation criteria and distortion types are constantly emerging. To address this issue, in this paper, we propose a dynamic Remember and Reuse (R&R) network, which efficiently performs the cross-task BIQA based on a novel relevance-aware incremental learning strategy. Given multiple evaluation tasks across different distortion types or databases, our R&R network sequentially updates the parameters for every task one by one. After each update step, part of task-specific parameters is settled, which ensures R&R Remembers their dedicated evaluation preferences. The remaining parameters are pruned for the dynamic usage of the subsequent tasks. To further exploit the correlation between different tasks, we feed the training data of a new task to previously settled parameters. Better prediction accuracy is considered as higher task relevance and vice versa. Then, we selectively Reuse parts of previously settled parameters, whose proportion is adaptively determined by the task relevance. Extensive experiments show that the proposed method efficiently achieves the cross-task BIQA without catastrophic forgetting, and significantly outperforms many state-of-the-art methods. Code is available at https://github.com/maruiperfect/R-R-Net. Rui Ma 0030, Hanxiao Luo, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
ACM Multimedia | 6 |
| 2021 | Few-Shot Segmentation via Complementary Prototype Learning and Cascaded Refinement
Hanxiao Luo, Hui Li 0080, Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Fanman Meng, Linfeng Xu 0001 |
PRCV (4) | 6 |
| 2021 | Behaviour detection in crowded classroom scenes via enhancing features robust to scale and perspective variationsabstractAbstract Detecting human behaviours in images of crowded classroom scenes is a challenging task, due to the large variations of humans in scale and pose perspective. In this paper, two modules are proposed to tackle these two variations. First, an attention‐based RoI (region‐of‐interest) extractor is designed to handle scale variation. Feature fusion and attention mechanism are used to improve the RoI feature with more local and global information. Second, a transformation‐based detection head is introduced to handle perspective variation. The spatial transformation is adopted to extract consistent representation under various perspectives. Moreover, since there is a lack of proper datasets for human behaviour detection in classroom scenes, a new dataset is created, namely CLBD. The experiments on the proposed dataset demonstrate that the modules obtain significant improvements of performance over the state‐of‐the‐art detectors. Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, Qianghua Liao |
IET Image Process. | 2 |
| 2021 | Hierarchical class grouping with orthogonal constraint for class activation map generation
Fanman Meng, Kaixu Huang, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
Neural Comput. Appl. | 1 |
| 2021 | Non-Homogeneous Haze Removal via Artificial Scene Prior and Bidimensional Graph ReasoningabstractDue to the lack of natural scene and haze prior information, it is greatly challenging to completely remove the haze from a single image without distorting its visual content. Fortunately, the real-world haze usually presents non-homogeneous distribution, which provides us with many valuable clues in partial well-preserved regions. In this paper, we propose a Non-Homogeneous Haze Removal Network (NHRN) via artificial scene prior and bidimensional graph reasoning. Firstly, we employ the gamma correction iteratively to simulate artificial multiple shots under different exposure conditions, whose haze degrees are different and enrich the underlying scene prior. Secondly, beyond utilizing the local neighboring relationship, we build a bidimensional graph reasoning module to conduct non-local filtering in the spatial and channel dimensions of feature maps, which models their long-range dependency and propagates the natural scene prior between the well-preserved nodes and the nodes contaminated by haze. To the best of our knowledge, this is the first exploration to remove non-homogeneous haze via the graph reasoning based framework. We evaluate our method on different benchmark datasets. The results demonstrate that our method achieves superior performance over many state-of-the-art algorithms for both the single image dehazing and hazy image understanding tasks. The source code of the proposed NHRN is available on https://github.com/whrws/NHRNet. Qingbo Wu 0001, Hui Li 0080, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Learning with Noisy Class Labels for Instance Segmentation
Longrong Yang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Qishang Cheng |
ECCV (14) | 2 |
| 2020 | Single Image Dehazing Via Artificial Multiple Shots And Multidimensional ContextabstractThe main challenge for single image dehazing is the lack of effective prior information for restoration. To address this issue, in this paper, we propose to generate artificial multiple shots for simulating the images captured under different haze degrees, and two context reasoning modules are developed to describe the relationship across different spatial regions and artificial shots. It brings two benefits in the inhomogeneous haze distribution. First, within one shot, the regions occluded in one location could be recovered with the help of other clear regions, which share the similar structures. Second, for the same spatial location, the regions distorted in one shot could be restored by means of other shots with clear content. We evaluate our method on different benchmark datasets. The results demonstrate that our method achieves superior performance over many state-of-the-art dehazing algorithms. Qingbo Wu 0001, Hui Li 0080, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICIP | 6 |
| 2020 | Region Adaptive Two-Shot Network For Single Image DehazingabstractExisting single image dehazing methods typically adopt a one-shot strategy by indiscriminately applying the same filters to all local regions, which easily cause under-/over-dehazing across different regions by ignoring the inhomogeneity and asymmetry of illumination and detail distortions. In this paper, we propose a region adaptive two-shot network (RATNet) to address this issue. In the first shot, a lightweight subnetwork is utilized to conduct the regular global filtering, which could remove parts of haze but also distort some image details. In the second shot, a two-branch subnetwork is developed to restore the illumination and details of the initially renovated image respectively. The final dehazed image is obtained by fusing the outputs of the previous two branches, whose region-variant weights are adaptively learned by minimizing the difference between the haze-free image and our fused result. Experiments on four dehazing benchmark datasets show that our RATNet significantly outperforms many state-of-the-art dehazing approaches. Hui Li 0080, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICME | 5 |
| 2020 | Language-Aware Fine-Grained Object Representation for Referring Expression ComprehensionabstractReferring expression comprehension expects to accurately locate an object described by a language expression, which requires precise language-aware visual object representations. However, existing methods usually use rectangular object representations, such as object proposal regions and grid regions. They ignore some fine-grained object information like shapes and poses, which are often described in language expressions and important to localize objects. Additionally, rectangular object regions usually contain background contents and irrelevant foreground features, which also decrease the localization performance. To address these problems, we propose a language-aware deformable convolution model (LDC) to learn language-aware fine-grained object representations. Rather than extracting rectangular object representations, LDC adaptively samples a set of key points based on the image and language to represent objects. This type of object representations can capture more fine-grained object information (e.g., shapes and poses) and suppress noises in accordance with language and thus, boosts the object localization performance. Based on the language-aware fine-grained object representation, we next design a bidirectional interaction model (BIM) that leverages a modified co-attention mechanism to build cross-modal bidirectional interactions to further improve the language and object representations. Furthermore, we propose a hierarchical fine-grained representation network (HFRN) to learn language-aware fine-grained object representations and cross-modal bidirectional interactions at local word level and global sentence level, respectively. Our proposed method outperforms the state-of-the-art methods on the RefCOCO, RefCOCO+ and RefCOCOg datasets. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Hengcan Shi, Taijin Zhao, King Ngi Ngan |
ACM Multimedia | 4 |
| 2020 | A New Local Transformation Module for Few-Shot Segmentation
Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Xiaolong Xu 0004 |
MMM (2) | 2 |
| 2020 | Haze-robust image understanding via context-aware deep feature refinementabstractImage understanding under the foggy scene is greatly challenging due to inhomogeneous visibility deterioration. Although various image dehazing methods have been proposed, they usually aim to improve image visibility (such as, PSNR/SSIM) in the pixel space rather than the feature space, which is critical for the perception of computer vision. Due to this mismatch, existing dehazing methods are limited or even adverse in facilitating the foggy scene understanding. In this paper, we propose a generalized deep feature refinement module to minimize the difference between clear images and hazy images in the feature space. It is consistent with the computer perception and can be embedded into existing detection or segmentation backbones for joint optimization. Our feature refinement module is built upon the graph convolutional network, which is favorable in capturing the contextual information and beneficial for distinguishing different semantic objects. We validate our method on the detection and segmentation tasks under foggy scenes. Extensive experimental results show that our method outperforms the state-of-the-art dehazing based pretreatments and the fine-tuning results on hazy images. Hui Li 0080, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
MMSP | 6 |
| 2020 | A Unified Single Image De-raining Model via Region Adaptive Coupled NetworkabstractSingle image de-raining is quite challenging due to the diversity of rain types and inhomogeneous distributions of rainwater. By means of dedicated models and constraints, existing methods perform well for specific rain type. However, their generalization capability is highly limited as well. In this paper, we propose a unified de-raining model by selectively fusing the clean background of the input rain image and the well restored regions occluded by various rains. This is achieved by our region adaptive coupled network (RACN), whose two branches integrate the features of each other in different layers to jointly generate the spatial-variant weight and restored image respectively. On the one hand, the weight branch could lead the restoration branch to focus on the regions with higher contributions for de-raining. On the other hand, the restoration branch could guide the weight branch to keep off the regions with over-/under-filtering risks. Extensive experiments show that our method outperforms many state-of-the-art de-raining algorithms on diverse rain types including the rain streak, raindrop and rain-mist. Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
VCIP | 5 |
| 2020 | A New Bounding Box based Pseudo Annotation Generation Method for Semantic SegmentationabstractThis paper proposes a fusion-based method to generate pseudo-annotations from bounding boxes for semantic segmentation. The idea is to first generate diverse foreground masks by multiple bounding box segmentation methods, and then combine these masks to generate pseudo-annotations. Existing methods generate foreground masks from bounding boxes by classical segmentation methods driving by low-level features and own local information, which is hard to generate accurate and diverse results for the fusion. Different from the traditional methods, multiple class-agnostic models are modeled to learn the objectiveness cues by using existing labeled pixel-level annotations and then to fuse. Firstly, the classical Fully Convolutional Network (FCN) that densely predicts the pixels' labels is used. Then, two new sparse prediction based class-agnostic models are proposed, which simplify the segmentation task as sparsely predicting the boundary points through predicting the distance from the bounding box border to the object boundary in Cartesian Coordinate System and the Polar Coordinate System, respectively. Finally, a voting-based strategy is proposed to combine these segmentation results to form better pseudo-annotations. We conduct experiments on PASCAL VOC 2012 dataset. The mIoU of the proposed method is 68.7%, which outperforms the state-of-the-art method by 1.9%. Xiaolong Xu 0004, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
VCIP | 2 |
| 2020 | Mono is Enough: Instance Segmentation from Single Annotated SampleabstractWith the help of various Deep Neural Networks, instance segmentation has achieved significant progress. How-ever, these successes are heavily reliant on large-scale manually annotated samples, which are extremely time-consuming and expensive. To address this issue, we propose a highly efficient anisotropic data augmentation method, which generates high quality training data from a single manually annotated sample. Instead of equivalently modifying foreground and background like traditional data augmentation methods, we focus on enriching the diversities of foreground appearance and positional relation between foreground and background, which are beneficial for the classification and localization sub-tasks respectively. All foreground instances of the source annotated sample undergo various rotation, brightness change, rescale, distortion and frequency-component mixup (FCM). Then, these modified instances are randomly embedded into background, which serve as new training samples. Experiments on Cityscapes dataset show that our method significantly outperforms traditional data augmentation methods. Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
VCIP | 4 |
| 2020 | Mining Larger Class Activation Map with Common Attribute LabelsabstractClass Activation Map (CAM) is the visualization of target regions generated from classification networks. However, classification network trained by class-level labels only has high responses to a few features of objects and thus the network cannot discriminate the whole target. We think that original labels used in classification tasks are not enough to describe all features of the objects. If we annotate more detailed labels like class-agnostic attribute labels for each image, the network may be able to mine larger CAM. Motivated by this idea, we propose and design common attribute labels, which are lower-level labels summarized from original image-level categories to describe more details of the target. Moreover, it should be emphasized that our proposed labels have good generalization on unknown categories since attributes (such as head, body, etc.) in some categories (such as dog, cat, etc.) are common and class-agnostic. That is why we call our proposed labels as common attribute labels, which are lower-level and more general compared with traditional labels. We finish the annotation work based on the PASCAL VOC2012 dataset and design a new architecture to successfully classify these common attribute labels. Then after fusing features of attribute labels into original categories, our network can mine larger CAMs of objects. Our method achieves better CAM results in visual and higher evaluation scores compared with traditional methods. Runtong Zhang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
VCIP | 2 |
| 2020 | Discriminative deep metric learning for asymmetric discrete hashing
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 3 |
| 2020 | An efficient and compact 3D local descriptor based on the weighted height image
Tiecheng Sun, Guanghui Liu 0001, Shuaicheng Liu, Fanman Meng, Liaoyuan Zeng, Ru Li 0002 |
Inf. Sci. | 4 |
| 2020 | HeadNet: An End-to-End Adaptive Relational Network for Head DetectionabstractHead detection plays an important role in localizing and identifying persons from visual data. Most existing methods treat head detection as a specific form of object detection. Head detection is nontrivial due to the considerable difficulty in building the local and global information under conditions of unconstrained pose and orientation. To address these issues, this paper presents an effective adaptive relational network to capture context information, which is greatly helpful to suppress missed detection. We show that the fundamental contextual properties, such as the global shape priors from different heads and the local adjacent relationship between the head and shoulders, can be systematically quantified by visual operators. Specifically, we propose a two-step search algorithm to quantify the global intergroup conflict with adaptive scale, pose and viewpoint. Meanwhile, a structured feature module is introduced to capture the local relation of intraindividual stability. Finally, the global priors and local relation are integrated seamlessly into a single-stage head detector that is end-to-end trainable. An extensive ablation analysis demonstrates the effectiveness of our approach. We achieve state-of-the-art results on two challenging datasets, i.e., HollywoodHeads and Brainwash. Wei Li 0110, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Weakly Supervised Semantic Segmentation by a Class-Level Multiple Group Cosegmentation and Foreground Fusion StrategyabstractWeakly supervised semantic segmentation uses image-level labels to extract object regions. The existing methods focus on efficiently training CNN-based segmentation networks using the image-level labels. In contrast to the existing methods, this paper proposes a new fusion-based method, which first segments the foregrounds of each image by multiple group cosegmentation and then generates the semantic segmentation by combining the foregrounds. Specifically, a new CNN-based multiple group cosegmentation network is first proposed to segment foregrounds employing two cues, the discriminative cue and the local-to-global cue. Then, the fusion method is proposed to simply perform semantic segmentation based on the multiple group cosegmentation results. Experiments on the PASCAL VOC 2012 and MS COCO 2017 datasets demonstrate the effectiveness of the proposed method with mIoU values that are obviously larger than those of the existing methods. Fanman Meng, Kunming Luo, Hongliang Li 0001, Qingbo Wu 0001, Xiaolong Xu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Subjective and Objective De-Raining Quality Assessment Towards Authentic Rain ImageabstractImages acquired by outdoor vision systems easily suffer poor visibility and annoying interference due to the rainy weather, which brings great challenge for accurately understanding and describing the visual contents. Recent researches have devoted great efforts on the task of rain removal for improving the image visibility. However, there is very few exploration about the quality assessment of de-rained image, even it is crucial for accurately measuring the performance of various de-raining algorithms. In this paper, we first create a de-raining quality assessment (DQA) database that collects 206 authentic rain images and their de-rained versions produced by 6 representative single image rain removal algorithms. Then, a subjective study is conducted on our DQA database, which collects the subject-rated scores of all de-rained images. To quantitatively measure the quality of de-rained image with non-uniform artifacts, we propose a bi-directional feature embedding network (B-FEN) which integrates the features of global perception and local difference together. Experiments confirm that the proposed method significantly outperforms many existing universal blind image quality assessment models. To help the research towards perceptually preferred de-raining algorithm, we will publicly release our DQA database and B-FEN source code on https://github.com/wqb-uestc. Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Hierarchical Context Features Embedding for Object DetectionabstractPixel-level segmentation has been widely used to improve object detection. Most of the existing methods refine detection features by adding the constraint of the segmentation branch or by simply embedding high-level segmentation features into detection features within the local receptive field. However, noisy segmentation features are unavoidable in real-word applications and can easily cause false positives. To address this problem, we propose a novel hierarchical context embedding module to effectively embed segmentation features into detection features. The idea of this module is to capture hierarchical context information that includes local objects or parts and nonlocal context features by learning multiple attention maps, and subsequently utilize interdependencies between features to recalibrate noisy segmentation features. Furthermore, we use this module in the proposed gated encoder-decoder network that adaptively aggregates feature maps of different resolutions based on the gate mechanism so that we can embed multiscale segmentation feature maps into detection features for more accurate detection of objects of all sizes. Experimental results demonstrate the effectiveness of the proposed method on the Pascal VOC 2012Seg dataset, the Pascal VOC dataset and the MS COCO dataset. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan, Hengcan Shi |
IEEE Trans. Multim. | 4 |
| 2019 | Instance Segmentation by Learning Deep Feature in Embedding SpaceabstractThe proposal-based framework is the mainstream architecture for instance segmentation. However, such architecture typically ignores the interference between objects, which fails to correctly segment overlapping objects with same category or appearance. In this paper, we propose a new instance segmentation network named Instance Discrimination Network (ID-Net) to consider the interference between objects by mapping pixels into an embedding space so that the pixels from different objects can be distinguished more accurately. To identify the foreground object in RoI, we learn a discriminative deep feature that can represent the embedding vectors corresponding to the foreground. Then, we get the foreground confidence map by calculating the similarities between the deep feature and embeddings. The experiments on PASCAL VOC and COCO datasets demonstrate the effectiveness of our method. Chao Shang 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001 |
ICIP | 3 |
| 2019 | Beyond Synthetic Data: A Blind Deraining Quality Assessment Metric Towards Authentic Rain ImageabstractDeraining quality assessment (DQA) plays an important role in evaluating and guiding the design of the image deraining algorithm. Due to the absence of rain-free image in the real rainy weather, the existing deraining algorithms are typically tested on several synthetic data by simulating very limited types of rain streaks, which are far from sufficient to measure the practicability of a deraining algorithm. In this paper, we first build a subjective DQA database that collects diverse authentic rain images and their derained versions. Then, a blind quality metric is developed to predict the deraining quality. Since the deraining artifacts are anisotropic and variable, we propose to describe the image via a bi-directional gated fusion network (B-GFN), which adaptively integrates the multi-scale cues of deraining artifact. Experiments confirm the effectiveness of the proposed method and its superiority with respect to many state-of-the-art blind image quality metrics. Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICIP | 5 |
| 2019 | Blind Image Sharpness Assessment And Enhancement via Deep Auxiliary LearningabstractIn this paper, we propose an unified deep auxiliary learning network to train the blind image sharpness assessment (BISA) metric and enhancer simultaneously. Instead of using the BISA as a parameter tuner like existing works, the proposed method aims to exploit the complementary information between two tasks and boost both of their performance. On the one hand, the enhancement subnetwork tries to separate a blurry image into the clear version and disparity map, which provide additional mask effect and blurry degree information for accurate BISA. On the other hand, the BISA subnetwork help determine the enhancement degree by feeding sharpness-aware features to the enhancer, which is helpful for avoiding under-/over-enhancing. Experimental results on three publicly available databases show that the proposed method outperforms many state-of-the-art algorithms in both the BISA and sharpness enhancement tasks. Qingbo Wu 0001, Rui Ma 0030, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICME | 5 |
| 2019 | Incorporating Non-local and Task-specific Features for Instance SegmentationabstractThis paper proposes a novel instance segmentation model, which improves the instance segmentation by considering two aspects. One is a new non-local features module to recover detailed information that is lost in the deep convolutional operations. The other is to introduce attention mechanism to generate specific features adaptive to each task. The proposed method is verified on three well-known datasets, namely Pascal VOC, Cityscapes and COCO. The experiments show that the method using the proposed modules outperforms baseline Mask R-CNN on all of the datasets without bells and whistles. Longrong Yang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
MMSP | 2 |
| 2019 | A New Few-shot Segmentation Network Based on Class RepresentationabstractThis paper studies few-shot segmentation, which is a task of predicting foreground mask of unseen classes by a few of annotations only, aided by a set of rich annotations already existed. The existing methods mainly focus the task on "how to transfer segmentation cues from support images (labeled images) to query images (unlabeled images)", and try to learn efficient and general transfer module that can be easily extended to unseen classes. However, it is proved to be a challenging task to learn the transfer module that is general to various classes. This paper solves few-shot segmentation in a new perspective of "how to represent unseen classes by existing classes", and formulates few-shot segmentation as the representation process that represents unseen classes (in terms of forming the foreground prior) by existing classes precisely. Based on such idea, we propose a new class representation based few-shot segmentation framework, which firstly generates class activation map of unseen class based on the knowledge of existing classes, and then uses the map as foreground probability map to extract the foregrounds from query image. A new two-branch based few-shot segmentation network is proposed. Moreover, a new CAM generation module that extracts the CAM of unseen classes rather than the classical training classes is raised. We validate the effectiveness of our method on Pascal VOC 2012 dataset, the value FB-IoU of one-shot and five-shot arrives at 69.2% and 70.1% respectively, which outperforms the state-of-the-art method. Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Qingbo Wu 0001 |
VCIP | 2 |
| 2018 | Adaptive Multi-Scale Information Flow for Object Detection
Wei Li 0110, Qingbo Wu 0001, Fanman Meng |
BMVC | 4 |
| 2018 | Key-Word-Aware Network for Referring Expression Image Segmentation
Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001 |
ECCV (6) | 3 |
| 2018 | Boosting Scene Parsing Performance via Reliable Scale PredictionabstractSegmenting objects on suitable scales is a key factor to improve the scene parsing performance. Existing methods either simply average multi-scale results or predict scales by weakly-supervised models, due to the lack of scale labels. In this paper, we propose a novel fully-supervised Scale Prediction Model. On one hand, the proposed Scale Prediction Model learns parsing scales by the strong scale supervision, which is automatically generated from the scene parsing ground truth without any extra manually annotation. On the other hand, we explore the relationship between scale and object class, and propose to use the object class information to further improve the reliability of the scale prediction. The proposed Scale Prediction Model improves 23.1%, 20.1% and 29.3% scale prediction accuracies on the NYU Depth v2, PASCAL-Context and SIFT Flow datasets, respectively. Based on the Scale Prediction Model, we design a Scale Parsing Net (SPNet) for scene parsing, which segments each object on the scale predicted by the Scale Prediction Model. Moreover, SPNet leverages the intermediate result (i.e., the object class) to refine the parsing results. The experiment results show that SPNet outperforms many state-of-the-art methods on multiple scene parsing datasets. Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
ACM Multimedia | 4 |
| 2018 | Weakly Supervised Semantic Segmentation by Multiple Group CosegmentationabstractWeakly supervised semantic segmentation aims at segmenting images by image-level labels. The existing methods try to train an end-to-end CNN network, which needs to handle multiple classes that is difficult. In addition, the existing methods are sensitive to the image-level cues such as discriminative regions and the pseudo-annotations. To avoid these drawbacks, this paper proposes a new strategy, which first obtains the foregrounds of each class by multiple group cosegmentation, and then combines the results to form the semantic segmentation. In our method, three new aspects are considered. (1) we solve semantic segmentation by each class that is easy to handle. (2) we extract discriminative regions more globally by context analysis. (3) we learn local-to-global segmentation network to segment the object from local discriminative priors. A new CNN network for multiple group cosegmentation is proposed. Two subnetworks such as global context based discriminative region extraction network and local-to-global segmentation network are designed. A simple combination method based on the discriminative map is proposed to finally obtain the semantic segmentation results. We verify the proposed method on Pascal VOC dataset. The experimental results show that the proposed method can obtain mIOU value 0.563 and 0.603 (without CRF post-processing) on the validation and test dataset that outperforms many existing weakly supervised semantic segmentation methods. Kunming Luo, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
VCIP | 2 |
| 2018 | Global and local semantics-preserving based deep hashing for cross-modal retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 3 |
| 2018 | An Unsupervised Method to Extract Video Object via Complexity Awareness and Object Local PartsabstractExisting unsupervised video object segmentation generates object information from the whole video, which ignores analysis of the local clips. However, we observe that local clips and their relationships are also useful for the video object segmentation. For example, the simple background clips can be used to improve the segmentation of complex background clips. In this paper, we propose a novel unsupervised segmentation framework to segment the primary object based on two aspects, i.e., the complexity awareness of video clips and their segmentation propagation. The first one is used to select the simple clips with smooth backgrounds and the second one generates an object prior from the simple clips and propagates the object prior to help and improve the segmentation of the complex clips. A complexity awareness method using the static cues and the dynamic cues are proposed to evaluate the complexity of the video frames. A new object prior learning model based on the local part structure is designed and a local part-based prior propagation is proposed for the complex clip segmentation. To verify our method, we collect a new challenging video segmentation data set, in which each video contains diverse backgrounds. Experimental results demonstrate that our method outperforms several state-of-the-art methods both on a classical data set and our new data set. Bing Luo 0003, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Globally Measuring the Similarity of Superpixels by Binary Edge Maps for Superpixel ClusteringabstractThis paper proposes an edge-based superpixel similarity measurement, which globally evaluates the similarity between superpixels by binary edge maps. The basic idea is to assess whether the superpixels are surrounded by the same edges. To this end, we first describe the edge spatial distributions by directional regions and then use the directional regions to represent the surrounding relationships of superpixels and edges by their traverse relationships, which form the histogram feature. Finally, the similarity is simply calculated by the distances between the features. To verify the proposed similarity measurement, we use our global similarity measurement to perform superpixel clustering. Two clustering methods, the directed graph clustering (DGC) and spectral clustering (ultrametric contour map) are combined to achieve the clustering process. The combination of our global similarity measurement and DGC to form a new three-layer-based superpixel generation method, which can quickly generate the superpixel from edge maps, is highlighted. We verify the global similarity measurement by the BSDS500 dataset. The experimental results demonstrate that the proposed global similarity measurement can improve the clustering accuracy in terms of larger intersection-over-union-criterion-based values. The code can be downloaded from https://github.com/FanmanMeng/Superpixel-Similarity-Measurement. Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, Chao Huang 0003, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | LETRIST: Locally Encoded Transform Feature Histogram for Rotation-Invariant Texture ClassificationabstractClassifying texture images, especially those with significant rotation, illumination, scale, and viewpoint changes, is a fundamental and challenging problem in computer vision. This paper proposes a simple yet effective image descriptor, called Locally Encoded TRansform feature hISTogram (LETRIST), for texture classification. LETRIST is a histogram representation that explicitly encodes the joint information within an image across feature and scale spaces. The proposed representation is training-free, low-dimensional, yet discriminative and robust for texture description. It consists of the following major steps. First, a set of transform features is constructed to characterize local texture structures and their correlation by applying linear and non-linear operators on the extremum responses of directional Gaussian derivative filters in scale space. Established on the basis of steerable filters, the constructed transform features are exactly rotationally invariant as well as computationally efficient. Second, the scalar quantization via binary or multi-level thresholding is adopted to quantize these transform features into texture codes. Two quantization schemes are designed, both of which are robust to image rotation and illumination changes. Third, the cross-scale joint coding is explored to aggregate the discrete texture codes into a compact histogram representation, i.e., LETRIST. Experimental results on the Outex, CUReT, KTH-TIPS, and UIUC texture data sets show that LETRIST consistently produces better or comparable classification results than the state-of-the-art approaches. Impressively, recognition rates of 100.00% and 99.00% have been achieved on the Outex and KTH-TIPS data sets, respectively. In addition, the noise robustness is evaluated on the Outex and CUReT data sets. The source code is publicly available athttps://github.com/stc-cqupt/letrist. Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Jianfei Cai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | A Perceptually Weighted Rank Correlation Indicator for Objective Image Quality AssessmentabstractIn the field of objective image quality assessment (IQA), Spearman's ρ and Kendall's τ, which straightforwardly assign uniform weights to all quality levels and assume that each pair of images is sortable, are the two most popular rank correlation indicators. These indicators can successfully measure the average accuracy of an IQA metric for ranking multiple processed images. However, two important perceptual properties are ignored. First, the sorting accuracy (SA) of high-quality images is usually more important than that of poor-quality images in many real-world applications, where only top-ranked images are pushed to the users. Second, due to the subjective uncertainty in making judgments, two perceptually similar images are usually barely sortable, and their ranks do not contribute to the evaluation of an IQA metric. To more accurately compare different IQA algorithms, in this paper, we explore a perceptually weighted rank correlation indicator, which rewards the capability of correctly ranking high-quality images and suppresses the attention towards insensitive rank mistakes. Specifically, we focus on activating a 'valid' pairwise comparison of images whose quality difference exceeds a given sensory threshold (ST). Meanwhile, each image pair is assigned a unique weight that is determined by both the quality level and rank deviation. By modifying the perception threshold, we can illustrate the sorting accuracy with a sophisticated SA-ST curve rather than a single rank correlation coefficient. The proposed indicator offers new insight into interpreting visual perception behavior. Furthermore, the applicability of our indicator is validated for recommending robust IQA metrics for both degraded and enhanced image data. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Image Process. | 3 |
| 2018 | Generic Proposal Evaluator: A Lazy Learning Strategy Toward Blind Proposal Quality AssessmentabstractExisting detection or recognition systems typically select one state-of-the-art proposal algorithm to produce massive object-covered candidate windows, and a quality metric specifically designed for this algorithm is utilized to single out small amounts of proposals. However, in practice, the accuracies of different proposal algorithms significantly change from one image content to another one. To obtain more robust proposal results, a generic proposal evaluator (GPE) is highly desired, which could choose optimal candidate windows across multiple proposal algorithms. In this paper, we propose a lazy learning strategy to train the GPE, which aims to blindly estimate the quality of each proposal without accessing to its manual annotation. Unlike the traditional end-to-end framework that learns a universal model from all training samples, we try to build query-specific training subset for each given proposal, where only its k-nearest-neighborhoods are collected from all labeled candidate windows. Benefits from the capability of updating the regression parameters for different visual contents, the proposed method delivers a higher quality prediction accuracy even with respect to the deep neural network learned by end-to-end method. Experimental results confirm that the proposed algorithm significantly outperforms many state-of-the-art proposal quality metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Seeds-Based Part Segmentation by Seeds Propagation and Region Convexity DecompositionabstractObject part segmentation is an important and challenging task in computer vision. The existing supervised part segmentation methods need pixel level training data which leads to a huge workload for the user. In this paper a weakly supervised part segmentation method is proposed which segments part regions from multiple images by only several seeds on an image. Two aspects such as seed propagation among multiple images and part generation from seeds are considered. The first aspect is to generate part seeds in each image in terms of seed propagation which is accomplished by part matching combined with latent object regions. We fuse the local part matching and global shape cosegmentation to avoid the noise propagation. The second aspect is to segment part regions from object regions and part seeds which is formulated as the object shape decomposition model. The shape convexity analysis and seed location are fused to accomplish the decomposition and the final part segmentation. The proposed method is verified on the PASCAL 2010 dataset Bird dataset Cat-Dog dataset and UCF Sports Actions dataset. Experimental results demonstrate the effectiveness of the proposed method with larger intersection over union (IOU) values compared with existing weakly supervised part generation methods. Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan, Jianfei Cai 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Hierarchical Parsing Net: Semantic Scene Parsing From Global Scene to ObjectsabstractThis paper proposes a novel Hierarchical Parsing Net (HPN) for semantic scene parsing. Unlike previous methods, which separately classify each object, HPN leverages global scene semantic information and the context among multiple objects to enhance scene parsing. On the one hand, HPN uses the global scene category to constrain the semantic consistency between the scene and each object. On the other hand, the context among all objects is also modeled to avoid incompatible object predictions. Specifically, HPN consists of four steps. In the first step, we extract scene and local appearance features. Based on these appearance features, the second step is to encode a contextual feature for each object, which models both the scene-object context (the context between the scene and each object) and the interobject context (the context among different objects). In the third step, we classify the global scene and then use the scene classification loss and a backpropagation algorithm to constrain the scene feature encoding. In the fourth step, a label map for scene parsing is generated from the local appearance and contextual features. Our model outperforms many state-of-the-art deep scene parsing networks on five scene parsing databases. Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2017 | Blind proposal quality assessment via deep objectness representation and local linear regressionabstractThe quality of object proposal plays an important role in boosting the performance of many computer vision tasks, such as, object detection and recognition. Due to the absence of manually annotated bounding-box in practice, the quality metric towards blind assessment of object proposal is highly desirable for singling out the optimal proposals. In this paper, we propose a blind proposal quality assessment algorithm based on the Deep Objectness Representation and Local Linear Regression (DORLLR). Inspired by the hierarchy model of the human vision system, a deep convolutional neural network is developed to extract the objectness-aware image feature. Then, the local linear regression method is utilized to map the image feature to a quality score, which tries to evaluate each individual test window based on its k-nearest-neighbors. Experimental results on a large-scale IoU labeled dataset verify that the proposed method significantly outperforms the state-of-the-art blind proposal evaluation metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Linfeng Xu 0001 |
ICME | 3 |
| 2017 | Segmentation quality evaluation based on multi-scale convolutional neural networksabstractSegmentation quality evaluation is an important task in image segmentation. The existing evaluation methods formulate segmentation quality as regression model, and recent Convolutional Neural Network (CNN) based evaluation methods show superior performance. However, designing efficient CNN-based segmentation evaluation model is still under exploited. In this paper, we propose two types of CNN structures such as double-net and multi-scale network for segmentation quality evaluation. We observe that learning the local and global information and considering multi-scale image are useful for segmentation quality evaluation. To train and verify the proposed networks, we construct a novel objective segmentation quality evaluation dataset with large amount of data by combining several proposal generation methods. The experimental results demonstrate that the proposed method obtains larger Linear Correlation Coefficient (LCC) value than several state-of-art segmentation quality evaluation methods. Fanman Meng, Qingbo Wu 0001 |
VCIP | 2 |
| 2017 | L2SSP: Robust keypoint description using local second-order statistics with soft-pooling
Tiecheng Song, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003, Yongjun Xu 0002 |
Neurocomputing | 2 |
| 2017 | Manifold-ranking embedded order preserving hashing for image semantic retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Semi-supervised manifold-embedded hashing with joint feature representation and classifier learning
Tiecheng Song, Jianfei Cai 0001, Chenqiang Gao, Fanman Meng, Qingbo Wu 0001 |
Pattern Recognit. | 5 |
| 2017 | Weakly Supervised Part Proposal Segmentation From Multiple ImagesabstractWeakly supervised local part segmentation is challenging, due to the difficulty of modeling multiple local parts from image level prior. In this paper, we propose a new weakly supervised local part proposal segmentation method based on the observation that local parts will keep fixed along the object pose variations. Hence, the local part can be segmented by capturing object pose variations. Based on such observation, a new local part proposal segmentation model is proposed. Three aspects, such as shape similarity-based cosegmentation, shape matching-based part detection and segmentation, and graph matching-based part assignment are considered. A part segmentation energy function is first proposed. Four terms, such as MRF-based single image segmentation term, shape feature-based foreground consistency term, NCuts-based part segmentation term, and two-order graphs matching based part consistency term, are contained. Then, a three sub-minimization-based energy minimization method is proposed to accomplish approximation solution. Finally, we verify our method based on three image data sets (PASCAL VOC 2008 Part data set, UCB Bird data set, and Cat-Dog data set), and one video data set (UCF Sports) data set. The experimental results demonstrate a better segmentation performance compared with the existing object cosegmentation and part proposal generation methods. Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, King Ngi Ngan |
IEEE Trans. Image Process. | 1 |
| 2017 | Video Object Segmentation via Global Consistency Aware Query StrategyabstractIn this paper, we propose a video object segmentation method via global consistency aware query strategy. The aim is to obtain higher segmentation accuracy with less user annotation. Intuitively, we hope to annotate some frames to obtain better segmentation performance than to annotate other frames, which can be modeled by active learning framework. Specifically, we first generate a sample space of potential annotation regions via an object proposals method for each frame. Then, the annotation likelihood for the region is calculated in terms of annotation history and global consistency for the object in the video. Third, the segmentation result of the annotation region can be obtained by minimizing an MRF energy function. Fourth, the algorithm will provide the user with the most valuable frame to annotate, which has high annotation likelihood and large segmentation result change. Finally, the annotation is added to the framework to begin the next iteration. Experiments on a number of video sequences demonstrate that the proposed method can reduce the user effort and obtain the higher segmentation accuracy compared with the state-of-the-art methods. Bing Luo 0003, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Chao Huang 0003 |
IEEE Trans. Multim. | 3 |
| 2017 | Learning Efficient Binary Codes From High-Level Feature Representations for Multilabel Image RetrievalabstractDue to the efficiency and effectiveness of hashing technologies, they have become increasingly popular in large-scale image semantic retrieval. However, existing hash methods suppose that the data distributions satisfy the manifold assumption that semantic similar samples tend to lie on a low-dimensional manifold, which will be weakened due to the large intraclass variation. Moreover, these methods learn hash functions by relaxing the discrete constraints on binary codes to real value, which will introduce large quantization loss. To tackle the above problems, this paper proposes a novel unsupervised hashing algorithm to learn efficient binary codes from high-level feature representations. More specifically, we explore nonnegative matrix factorization for learning high-level visual features. Ultimately, binary codes are generated by performing binary quantization in the high-level feature representations space, which will map images with similar (visually or semantically) high-level feature representations to similar binary codes. To solve the corresponding optimization problem involving nonnegative and discrete variables, we develop an efficient optimization algorithm to reduce quantization loss with guaranteed convergence in theory. Extensive experiments show that our proposed method outperforms the state-of-the-art hashing methods on several multilabel real-world image datasets. Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2017 | Blind Image Quality Assessment Based on Rank-Order Regularized RegressionabstractBlind image quality assessment (BIQA) aims to estimate the subjective quality of a query image without access to the reference image. Existing learning-based methods typically train a regression function by minimizing the average error between subjective opinion scores and model predictions. However, minimizing average error does not necessarily lead to correct quality rank-orders between the test images, which is a highly desirable property of image quality models. In this paper, we propose a novel rank-order regularized regression model to address this problem. The key idea is to introduce a pairwise rank-order constraint into the maximum margin regression framework, aiming to better preserve the correct perceptual preference. To the best of our knowledge, this is the first attempt to incorporate rank-order constraints into margin-based quality regression model. By combing with a new local spatial structure feature, we achieve highly consistent quality prediction with human perception. Experimental results show that the proposed method outperforms many state-of-the-art BIQA metrics on popular publicly available IQA databases (i.e., LIVE-II, TID2013, VCL@FER, LIVEMD, and ChallengeDB). Qingbo Wu 0001, Hongliang Li 0001, Zhou Wang 0001, Fanman Meng, Bing Luo 0003, Wei Li 0110, King Ngi Ngan |
IEEE Trans. Multim. | 4 |
| 2016 | QualityNet: Segmentation quality evaluation with deep convolutional networksabstractThis paper proposes a deep convolutional network for quality evaluation of object segmentation. In this work, we first present two large-scale datasets for object quality evaluation. The segmentation results of the proposed datasets are generated from simple and complex segmentation datasets respectively. Then, we propose three types of deep convolutional networks to learn object segmentation quality. The weighted mask layer is proposed to utilize both the information of foreground and background. One advantage of our network is that the segmentation quality is evaluated without using the groundtruth segmentation results during test period. This may help us obtain better segmentation result based on the predicted segmentation quality score. The promising results on the proposed datasets indicate the efficiency of our network. Chao Huang 0003, Qingbo Wu 0001, Fanman Meng |
VCIP | 3 |
| 2016 | Part propagation for local part segmentationabstractSegment propagation transfers object priors among images, which is an important prior generation manner in image segmentation. The existing propagation methods focus on object foreground propagation, while the detailed part propagation is deficiency, which is caused by the challenges that not only the multiple part regions, but also their relationships need to be transferred. In this paper, a part propagation method is proposed. Two level propagations such as object level propagation, and part level propagation are successively used for the part propagation. The object level propagation is to transfer global shape information among images, which is formulated as graph matching based edge fragments matching problem, with dynamic programming solution. The part level propagation is to transfer the more detailed part labels, which is formulated as pixel level structure matching problem, and is efficiently solved by traditional dense pixel matching methods. The proposed method is verified on 15 challenging classes selected from PASCAL 2010 dataset, Bird dataset and Cat-Dog dataset. The experimental results demonstrate the effectiveness of the proposed method. Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, Jianfei Cai 0001, Chao Huang 0003 |
VCIP | 1 |
| 2016 | Q-DNN: A quality-aware deep neural network for blind assessment of enhanced imagesabstractImage enhancement is widely popular due to its capability of producing "better" visual quality for specific applications. Although many enhancement algorithms have been developed in recent years, the studies towards blind assessment of enhanced images are still very lacking. In this paper, we propose a data-driven blind image quality assessment (BIQA) method based on the quality-aware deep neural network (Q-DNN). Unlike the conventional hand-crafted features designed for measuring the degradation level of specific distortion types, a supervised learning model is utilized in our Q-DNN, which is capable of adaptively updating the feature extractor and quality regressor for describing the visual artifacts caused by different image enhancement tasks. Experimental results on two challenging enhanced image databases show that the proposed method is significantly superior to the state-of-the-art BIQA metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
VCIP | 3 |
| 2016 | Cosegmentation of multiple image groups
Fanman Meng, Jianfei Cai 0001, Hongliang Li 0001 |
Comput. Vis. Image Underst. | 1 |
| 2016 | Person re-identification based on multi-region-set ensembles
Wei Li 0110, Chao Huang 0003, Bing Luo 0003, Fanman Meng, Tiecheng Song, Hengcan Shi |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Beyond pixels: A comprehensive survey from bottom-up to semantic image segmentation and cosegmentation
Hongyuan Zhu 0002, Fanman Meng, Jianfei Cai 0001, Shijian Lu |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Blind Image Quality Assessment Based on Multichannel Feature Fusion and Label TransferabstractIn this paper, we propose an efficient blind image quality assessment (BIQA) algorithm, which is characterized by a new feature fusion scheme and a k-nearest-neighbor (KNN)-based quality prediction model. Our goal is to predict the perceptual quality of an image without any prior information of its reference image and distortion type. Since the reference image is inaccessible in many applications, the BIQA is quite desirable in this context. In our method, a new feature fusion scheme is first introduced by combining an image's statistical information from multiple domains (i.e., discrete cosine transform, wavelet, and spatial domains) and multiple color channels (i.e., Y, Cb, and Cr). Then, the predicted image quality is generated from a nonparametric model, which is referred to as the label transfer (LT). Based on the assumption that similar images share similar perceptual qualities, we implement the LT with an image retrieval procedure, where a query image's KNNs are searched for from some annotated images. The weighted average of the KNN labels (e.g., difference mean opinion score or mean opinion score) is used as the predicted quality score. The proposed method is straightforward and computationally appealing. Experimental results on three publicly available databases (i.e., LIVE II, TID2008, and CSIQ) show that the proposed method is highly consistent with human perception and outperforms many representative BIQA metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | No reference image quality assessment metric via multi-domain structural information and piecewise regression
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Shuyuan Zhu |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Exploring space-frequency co-occurrences via local quantized patterns for texture representation
Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003 |
Pattern Recognit. | 3 |
| 2015 | Constrained Directed Graph Clustering and Segmentation Propagation for Multiple Foregrounds CosegmentationabstractThis paper proposes a new constrained directed graph clustering (DGC) method and segmentation propagation method for the multiple foreground cosegmentation. We solve the multiple object cosegmentation with the perspective of classification and propagation, where the classification is used to obtain the object prior of each class and the propagation is used to propagate the prior to all images. In our method, the DGC method is designed for the classification step, which adds clustering constraints in cosegmentation to prevent the clustering of the noise data. A new clustering criterion such as the strongly connected component search on the graph is introduced. Moreover, a linear time strongly connected component search algorithm is proposed for the fast clustering performance. Then, we extract the object priors from the clusters, and propagate these priors to all the images to obtain the foreground maps, which are used to achieve the final multiple objects extraction. We verify our method on both the cosegmentation and clustering tasks. The experimental results show that the proposed method can achieve larger accuracy compared with both the existing cosegmentation methods and clustering methods. Fanman Meng, Hongliang Li 0001, Shuyuan Zhu, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Fast HEVC Inter CU Decision Based on Latent SAD EstimationabstractThe emerging high efficiency video coding (HEVC) standard has improved compression performance significantly in comparison with H.264/AVC. However, more intensive computational complexity has been introduced by adopting a number of new coding tools. In this paper, a fast inter CU decision is proposed based on the latent sum of absolute differences (SAD) estimation. Firstly, a two-layer motion estimation (ME) method is designed to take advantage of the latent SAD cost. The new ME method can obtain the SAD costs for both the upper CU and its sub-CUs. Secondly, a concept of motion compensation rate- distortion (R-D) cost is defined, and an exponential model is proposed to express the relationship between the motion compensation R-D cost and the SAD cost. Then, a fast CU decision approach is designed based on the exponential model. The fast CU decision is implemented by comparing a derived threshold with the SAD cost difference between the upper and sub SAD costs. Experimental results show that the proposed algorithm achieves an average of 52% and 58.4% reductions of the coding time at the cost of 1.61% and 2% bit-rate increases under the low delay and random access conditions, respectively. Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2014 | On Multiple Image Group Cosegmentation
Fanman Meng, Jianfei Cai 0001, Hongliang Li 0001 |
ACCV (4) | 1 |
| 2014 | Automatic image co-segmentation using geometric mean saliencyabstractMost existing high-performance co-segmentation algorithms are usually complicated due to the way of co-labelling a set of images and the requirement to handle quite a few parameters for effective co-segmentation. In this paper, instead of relying on the complex process of co-labelling multiple images, we perform segmentation on individual images but based on a combined saliency map that is obtained by fusing singleimage saliency maps of a group of similar images. Particularly, a new multiple image based saliency map extraction, namely geometric mean saliency (GMS) method, is proposed to obtain the global saliency maps. In GMS, we transmit the saliency information among the images using the warping technique. Experiments show that our method is able to outperform state-of-the-art methods on three benchmark co-segmentation datasets. Koteswar Rao Jerripothula, Jianfei Cai 0001, Fanman Meng, Junsong Yuan 0001 |
ICIP | 3 |
| 2014 | Fast and efficient inter CU decision for high efficiency video codingabstractIn this paper, a graph cut based fast Coding Unit (CU) decision algorithm is proposed for HEVC inter frames. Firstly, a feature called pyramid variance of the absolute difference (PVAD) is designed for the CU selection. Secondly, the CU decision is modeled as a Markov Random Field (MRF) inference problem, which can be optimized by the graph cut algorithm. Thirdly, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show the effectiveness of the proposed method. Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Bing Zeng 0001, Shuyuan Zhu, Qingbo Wu 0001 |
ICIP | 3 |
| 2014 | Using mid-high level cues to detect salient objectabstractThis paper proposes a novel saliency object detection method by using the mid-level and high-level visual cues. In the mid-level objectness evaluation, we generate three complementary saliency maps, such as the multi-scale segmentation cue, the background cue and the spatial color distribution cue. The first cue is used to highlight the objects via the local region segment. The second cue uses the background priors to detect the saliency information. The third cue is to capture the spatial color distribution. For the high-level visual cue, we propose an objectness evaluation model to distinguish the object and the background. All the saliency cues are finally combined to achieve the saliency detection. The experimental results show that the proposed method outperforms the state-of-the-art saliency object detection methods. Hongliang Li 0001, Yurui Xie, Bing Luo 0003, Liangzhi Tang, Bing Zeng 0001, King Ngi Ngan, Fanman Meng |
ICME | 7 |
| 2014 | Favorite object extraction using web imagesabstractIn this paper, we propose a framework to discover and segment favorite object from the natural images. The main idea is to first generate the shape based common template of the favorite object using the images collected from the web. Then, the common template is used to extract the favorite object from the original images. In the common template generation, co-segmentation is used to provide the initial segments. The median graph theory is employed to construct the common template. We also propose a new shape descriptor namely directional shape representation to handle shape variations. We test our method on the images collected from image datasets and web. Experimental results demonstrate the effectiveness of the proposed method. Fanman Meng, Bing Luo 0003, Chao Huang 0003, Liangzhi Tang, Bing Zeng 0001, Nini Rao |
ISCAS | 1 |
| 2014 | Cosegmentation from similar backgroundsabstractRecently, the common objects are often required to be extracted from a group of images in many applications, such as video coding and model training. Co-segmentation is a new and efficient method for this requirement. In realistic applications, we observe that the images usually contain similar backgrounds (namely similar scene co-segmentation), such as the city landmark images collected from the web or the key frames sampled from a video. Meanwhile, the existing co-segmentation has not paid so much attention on the similar scene co-segmentation, and the insufficiently accurate segments may be provided by the existing methods. In this paper, we propose an active contours based co-segmentation model to provide foregrounds from the similar backgrounds. We combine the background consistency constraint with the foreground consistency constraint to form the energy function, and use the method of level-set and the calculus of variations to minimize the model. We also speed up the model by the hierarchical structure and the superpixel technique. We test the method on both the image and video dataset. The results show that the proposed model can obtain larger IOU values than the state-of-the-art co-segmentation methods. Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Nini Rao |
ISCAS | 1 |
| 2014 | Bird breed classification and annotation using saliency based graphical model
Chao Huang 0003, Fanman Meng, Shuyuan Zhu |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Noise-Robust Texture Description Using Local Contrast Patterns via Global MeasuresabstractThis letter presents a noise-robust descriptor by exploring a set of local contrast patterns (LCPs) via global measures for texture classification. To handle image noise, the directed and undirected difference masks are designed to calculate three types of local intensity contrasts: directed, undirected, and maximum difference responses. To describe pixel-wise features, these responses are separately quantized and encoded into specific patterns based on different global measures. These resulting patterns (i.e., LCPs) are jointly encoded to form our final texture representation. Experiments are conducted on the well-known Outex and CUReT databases in the presence of high levels of noise. Compared to many state-of-the-art methods, the proposed descriptor achieves superior texture classification performance while enjoying a compact feature representation. Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003, Bing Zeng 0001, Moncef Gabbouj |
IEEE Signal Process. Lett. | 3 |
| 2014 | Unsupervised Multiclass Region Cosegmentation via Ensemble Clustering and Energy MinimizationabstractThe problem of unsupervised segmentation of multi-class regions can be significantly boosted when they irregularly recur in multiple images. The existing segmentation methods are either weakly supervised, such as tagging images with object classes, or are limited by the assumption that each image contains all the object instances. In this paper, we propose a new method to cosegment multiclass regions from a group of images without the assumption about object configurations. The key idea is to discover the unknown object-like proposals via a robust ensemble clustering scheme. The proposals are then used to derive unary and pairwise energy potentials across all the images, which can be minimized with the α-expansion. Experimental evaluation on a number of image groups demonstrates the good performance of the proposed method on the multiclass region cosegmentation. Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Repairing Bad Co-Segmentation Using Its Quality Evaluation and Segment PropagationabstractIn this paper, we improve co-segmentation performance by repairing bad segments based on their quality evaluation and segment propagation. Starting from co-segmentation results of the existing co-segmentation method, we first perform co-segmentation quality evaluation to score each segment. Good segments can be filter out based on the scores. Then, a propagation method is designed to transfer good segments to the rest bad ones so as to repair the bad segmentation. In our method, the quality evaluation is implemented by the measurements of foreground consistency and segment completeness. Two propagation methods such as global propagation and local region propagation are then defined to achieve the more accurate propagation. We verify the proposed method using four state-of-the-arts co-segmentation methods and two public datasets such as ICoseg dataset and MSRC dataset. The experimental results demonstrate the effectiveness of the proposed quality evaluation method. Furthermore, the proposed method can significantly improve the performance of existing methods with larger intersection-over-union score values. Hongliang Li 0001, Fanman Meng, Bing Luo 0003, Shuyuan Zhu |
IEEE Trans. Image Process. | 2 |
| 2014 | MRF-Based Fast HEVC Inter CU Decision With the Variance of Absolute DifferencesabstractThe newly developed High Efficiency Video Coding (HEVC) Standard has improved video coding performance significantly in comparison to its predecessors. However, more intensive computation complexity is introduced by implementing a number of new coding tools. In this paper, a fast coding unit (CU) decision based on Markov random field (MRF) is proposed for HEVC inter frames. First, it is observed that the variance of the absolute difference (VAD) is proportional with the rate-distortion (R-D) cost. The VAD based feature is designed for the CU selection. Second, the decision of CU splittings is modeled as an MRF inference problem, which can be optimized by the Graphcut algorithm. Third, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show that the proposed algorithm can achieve about 53% reduction of the coding time with negligible coding performance degradation, which outperforms the state-of-the-art algorithms significantly. Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Shuyuan Zhu, Qingbo Wu 0001, Bing Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | A Fast HEVC Inter CU Selection Method Based on Pyramid Motion DivergenceabstractThe newly developed HEVC video coding standard can achieve higher compression performance than the previous video coding standards, such as MPEG-4, H.263 and H.264/AVC. However, HEVC's high computational complexity raises concerns about the computational burden on real-time application. In this paper, a fast pyramid motion divergence (PMD) based CU selection algorithm is presented for HEVC inter prediction. The PMD features are calculated with estimated optical flow of the downsampled frames. Theoretical analysis shows that PMD can be used to help selecting CU size. A k nearest neighboring like method is used to determine the CU splittings. Experimental results show that the fast inter prediction method speeds up the inter coding significantly with negligible loss of the peak signal-to-noise ratio. Jian Xiong 0005, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng |
IEEE Trans. Multim. | 4 |
| 2013 | Complexity awareness based feature adaptive co-segmentationabstractIn this paper, we achieve co-segmentation by learning adaptive feature model for each image group. A novel feature adaptive co-segmentation method and an image complexity awareness method are proposed. We also propose a linear feature model and an expectation-minimization (EM) based algorithm for adaptive feature learning. In the EM based algorithm, two aspects such as the accuracy confidence of the simple image segmentation and the fitness of the learned model to the simple image segmentation are considered. L1-regularized least squares optimization is also combined for the minimization. By testing on several well-known datasets, the error rates of the final co-segmentation are verified to be lower than the existing state-of-the-art co-segmentation methods. Fanman Meng, Hongliang Li 0001 |
ICIP | 1 |
| 2013 | Segmenting specific object based on logo detectionabstractThis paper proposes a method to segment object with logos. In the method, we firstly locate the logos by SIFT matching. Then, the object boundary is extracted based on the logo location. Finally, we model the object prior based on the boundary, and introduce the prior into Markov random field segmentation method to segment the object. To verify the proposed method, we collect a logo dataset from the web such as Flickr and Google. The experimental results demonstrate the effectiveness of the proposed method. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001 |
ISCAS | 1 |
| 2013 | Object co-segmentation based on directed graph clusteringabstractIn this paper, we develop a new algorithm to segment multiple common objects from a group of images. Our method consists of two aspects: directed graph clustering and prior propagation. The clustering is used to cluster the local regions of the original images and generate the foreground priors from these clusterings. The second step propagates the prior of each class and locates the common objects from the images in terms of foreground map. Finally, we use the foreground map as the unary term of Markov random field segmentation and segment the common objects by graph-cuts algorithm. We test our method on FlickrMFC and ICoseg datasets. The experimental results show that the proposed method can achieve larger accuracy compared with several state-of-arts co-segmentation methods. Fanman Meng, Bing Luo 0003, Chao Huang 0003 |
VCIP | 1 |
| 2013 | Robust texture representation by using binary code ensembleabstractIn this paper, we present a robust texture representation by exploring an ensemble of binary codes. The proposed method, called Locally Enhanced Binary Coding (LEBC), is training-free and needs no costly data-to-cluster assignments. Given an input image, a set of features that describe different pixel-wise properties, is first extracted so as to be robust to rotation and illumination changes. Then, these features are binarized and jointly encoded into specific pixel labels. Meanwhile, the Local Binary Pattern (LBP) operator is utilized to encode the neighboring relationship. Finally, based on the statistics of these pixel labels and LBP labels, a joint histogram is built and used for texture representation. Extensive experiments have been conducted on the Outex, CUReT and UIUC texture databases. Impressive classification results have been achieved compared with state-of-the-art LBP-based and even learning-based algorithms. Tiecheng Song, Fanman Meng, Bing Luo 0003, Chao Huang 0003 |
VCIP | 2 |
| 2013 | Image Cosegmentation by Incorporating Color Reward Strategy and Active Contour ModelabstractThe design of robust and efficient cosegmentation algorithms is challenging because of the variety and complexity of the objects and images. In this paper, we propose a new cosegmentation model by incorporating a color reward strategy and an active contour model. A new energy function corresponding to the curve is first generated with two considerations: the foreground similarity between the image pairs and the background consistency in each of the image pair. Furthermore, a new foreground similarity measurement based on the rewarding strategy is proposed. Then, we minimize the energy function value via a mutual procedure which uses dynamic priors to mutually evolve the curves. The proposed method is evaluated on many images from commonly used databases. The experimental results demonstrate that the proposed model can efficiently segment the common objects from the image pairs with generally lower error rate than many existing and conventional cosegmentation methods. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
IEEE Trans. Cybern. | 1 |
| 2013 | Feature Adaptive Co-Segmentation by Complexity AwarenessabstractIn this paper, we propose a novel feature adaptive co-segmentation method that can learn adaptive features of different image groups for accurate common objects segmentation. We also propose image complexity awareness for adaptive feature learning. In the proposed method, the original images are first ranked according to the image complexities that are measured by superpixel changing cue and object detection cue. Then, the unsupervised segments of the simple images are used to learn the adaptive features, which are achieved using an expectation-minimization algorithm combining l 1-regularized least squares optimization with the consideration of the confidence of the simple image segmentation accuracies and the fitness of the learned model. The error rate of the final co-segmentation is tested by the experiments on different image groups and verified to be lower than the existing state-of-the-art co-segmentation methods. Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Liaoyuan Zeng, Qingbo Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Co-Salient Object Detection From Multiple ImagesabstractIn this paper, we propose a novel method to discover co-salient objects from a group of images, which is modeled as a linear fusion of an intra-image saliency (IaIS) map and an inter-image saliency (IrIS) map. The first term is to measure the salient objects from each image using multiscale segmentation voting. The second term is designed to detect the co-salient objects from a group of images. To compute the IrIS map, we perform the pairwise similarity ranking based on an image pyramid representation. A minimum spanning tree is then constructed to determine the image matching order. For each region in an image, we design three types of visual descriptors, which are extracted from the local appearance, e.g., color, color co-occurrence and shape properties. The final region matching problem between the images is formulated as an assignment problem that can be optimized by linear programming. Experimental evaluation on a number of images demonstrates the good performance of the proposed method on co-salient object detection. Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Multim. | 2 |
| 2013 | From Logo to Object SegmentationabstractThis paper proposes a method to segment object from the web images using logo detection. The method consists of three steps. In the first step, the logos are located from the original images by SIFT matching. Based on the logo location and the object shape model, the second step extracts the object boundary from the image. In the third step, we use the object boundary to model the object appearance, which is then used in the MRF based segmentation method to finally achieve the object segmentation. The key of our method is the object boundary extraction, which is achieved by searching a variation of the shape model that best fits the local edge of the image. Affine transform is used to consider the variations among the objects. Meanwhile, the Nelder-Mead simplex method with a simple initial rough search is used to run the boundary search. To verify the proposed method, we collect a LogoSeg dataset from the web such as Flickr and Google. The MOMI dataset is also used for the verification. The experimental results demonstrate that the proposed logo detection based segmentation method can improve the performance of the object segmentation. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 1 |
| 2012 | Image co-segmentation via active contoursabstractIn this paper, a new co-segmentation model by incorporating active contours based method and rewarding strategy is represented. We first generate co-segmentation energy function from two aspects. One is foreground similarity between image pairs. The other is background consistency in each single image. Then, we optimize the energy function through a mutual optimization approach. We verify the proposed method on the images commonly used in co-segmentation research. Experimental results demonstrate the effectiveness of our method. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001 |
ISCAS | 1 |
| 2012 | Object Co-Segmentation Based on Shortest Path Algorithm and Saliency ModelabstractSegmenting common objects that have variations in color, texture and shape is a challenging problem. In this paper, we propose a new model that efficiently segments common objects from multiple images. We first segment each original image into a number of local regions. Then, we construct a digraph based on local region similarities and saliency maps. Finally, we formulate the co-segmentation problem as the shortest path problem, and we use the dynamic programming method to solve the problem. The experimental results demonstrate that the proposed model can efficiently segment the common objects from a group of images with generally lower error rate than many existing and conventional co-segmentation methods. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 1 |