EDBT 2026 Demo / reviewers in the wild / expert
Linfeng Xu 0001
dblp:119/4096-1
· DBLP profile ↗
57ranked-venue papers
5as first author
35since 2021 · last 2026
0000-0002-9934-0958ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 3 first-author · 27 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Systems, architecture and hardware · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze PredictionabstractEgocentric gaze prediction serves as a critical indicator for decoding human visual attention and cognitive processes, but its inherently limited field of view creates prediction challenges. Although exo-view data provides supplementary contextual information, it exhibits significant spatial and semantic gaps. Existing methods focus solely on isolated feature encoding in single-view paradigms, neglecting cross-view gaze correlations. To make up for this gap, we make the first exploration of cross-view gaze relationship for egocentric gaze prediction, and propose Ego-PMOVE, a novel Prompt-aware Mixture of View Experts network. Unlike prior cross-view studies that forcibly align cross-view features thereby introducing inference noise, we leverage the popular Mixture-of-Experts (MoE) and a set of flexible prompts to disentangle features from different views into three parallel experts: a view-shared expert directly modeling common semantic relationships, a view-discrepancy expert adaptively adjusting the spatial position, scale and shifts based on different view-specific features, and an egocentric expert extracting independent features to compensate for the case of missing exocentric data. To balance these experts, we further design a soft router to dynamically weight them for mining useful information while suppressing noise. A view-query gaze decoder then generates view-specific gaze attention maps, jointly optimized by gaze-heamap and cross-view contrastive loss that regularize both shared and divergent features for accurate gaze prediction. Extensive experiments across the multi-view EgoMe dataset and single-view Ego4D and EGTEA Gaze++ datasets demonstrate the effectiveness and generalizability of our approach. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Zhaofeng Shi, Linfeng Xu 0001, Hongliang Li 0001 |
AAAI | 6 |
| 2026 | CMaP-SAM: Contraction mapping prior for SAM-driven few-shot segmentation
Fanman Meng, Liming Lei, Qingbo Wu 0001, Linfeng Xu 0001, Hongliang Li 0001 |
Neurocomputing | 7 |
| 2026 | DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual RecognitionabstractContinual learning of vision-language models (VLMs) focuses on leveraging cross-modal pretrained knowledge to incrementally adapt to expanding downstream tasks and datasets, while tackling the challenge of knowledge forgetting. Existing research often focuses on connecting visual features with specific class text in downstream tasks, overlooking the latent relationships between general and specialized knowledge. Our findings reveal that forcing models to optimize inappropriate visual-text matches exacerbates forgetting of VLM's recognition ability. To tackle this issue, we propose DesCLIP, which leverages general attribute (GA) descriptions to guide the understanding of specific class objects, enabling VLMs to establish robust <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-GA-class</i> trilateral associations rather than relying solely on <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-class</i> connections. Specifically, we introduce a language assistant to generate concrete GA description candidates via proper request prompts. Then, an anchor-based embedding filter is designed to obtain highly relevant GA description embeddings, which are leveraged as the paired text embeddings for visual-textual instance matching, thereby tuning the visual encoder. Correspondingly, the class text embeddings are gradually calibrated to align with these shared GA description embeddings. Extensive experiments demonstrate the advancements and efficacy of our proposed method, with comprehensive empirical evaluations highlighting its superior performance in VLM-based recognition compared to existing continual learning methods. Chiyuan He, Zihuan Qiu, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive FusionabstractUnlike traditional Multimodal Class-Incremental Learning (MCIL) methods that focus only on vision and text, this paper explores MCIL across vision, audio and text modalities, addressing challenges in integrating complementary information and mitigating catastrophic forgetting. To tackle these issues, we propose an MCIL method based on multimodal pre-trained models. Firstly, a Multimodal Incremental Feature Extractor (MIFE) based on Mixture-of-Experts (MoE) structure is introduced to achieve effective incremental fine-tuning for AudioCLIP. Secondly, to enhance feature discriminability and generalization, we propose an Adaptive Audio-Visual Fusion Module (AAVFM) that includes a masking threshold mechanism and a dynamic feature fusion mechanism, along with a strategy to enhance text diversity. Thirdly, a novel multimodal class-incremental contrastive training loss is proposed to optimize cross-modal alignment in MCIL. Finally, two MCIL-specific evaluation metrics are introduced for comprehensive assessment. Extensive experiments on three multimodal datasets validate the effectiveness of our method. Zihuan Qiu, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001 |
ICASSP | 5 |
| 2025 | DPM-CLIP: Zero-Shot Multimodal Egocentric Activity Recognition based on Dual-Prediction MechanismabstractAdvancements in Zero-shot Multimodal Egocentric Activity Recognition (ZS-MM-EAR) largely rely on Vision-Language Model (VLM). However, existing methods struggle with VLM’s inadequate representation of egocentric activities, including challenges in capturing egocentric-specific features, adapting to domain shifts between egocentric video and pre-training data, and effectively leveraging complementary data such as Inertial Measurement Unit (IMU). To address these issues, we propose DPM-CLIP, a ZS-MM-EAR method tailored for vision, text and IMU modalities. Firstly, we design an attribute-driven text augmentation module that leverages a Large Language Model (LLM) to generate fine-grained textual descriptions of activities. Secondly, we construct an Instance-feature Repository (IFR) to store base class features and generate pseudo-features for novel classes through feature center migration. Finally, we introduce a dual-prediction mechanism with a prediction correction module to enhance generalization and recognition accuracy. Extensive experiments on the UESTC-MMEA-CL dataset validate the effectiveness of the proposed method. Zihuan Qiu, Mingzhou He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001 |
ICIP | 6 |
| 2025 | MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model MergingabstractContinual model merging integrates independently fine-tuned models sequentially without access to the original training data, offering a scalable and efficient solution for continual learning. However, existing methods face two critical challenges: parameter interference among tasks, which leads to catastrophic forgetting, and limited adaptability to evolving test distributions. To address these issues, we introduce the task of Test-Time Continual Model Merging (TTCMM), which leverages a small set of unlabeled test samples during inference to alleviate parameter conflicts and handle distribution shifts. We propose MINGLE, a novel framework for TTCMM. MINGLE employs a mixture-of-experts architecture with parameter-efficient, low-rank experts, which enhances adaptability to evolving test distributions while dynamically merging models to mitigate conflicts. To further reduce forgetting, we propose Null-Space Constrained Gating, which restricts gating updates to subspaces orthogonal to prior task representations, thereby suppressing activations on old tasks and preserving past knowledge. We further introduce an Adaptive Relaxation Strategy that adjusts constraint strength dynamically based on interference signals observed during test-time adaptation, striking a balance between stability and adaptability. Extensive experiments on standard continual merging benchmarks demonstrate that MINGLE achieves robust generalization, significantly reduces forgetting, and consistently surpasses previous state-of-the-art methods by 7–9% on average across diverse task orders. Our code is available at: https://github.com/zihuanqiu/MINGLE Zihuan Qiu, Yi Xu 0008, Chiyuan He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001 |
NeurIPS | 5 |
| 2025 | High efficiency deep image compression via channel-wise scale adaptive latent representation learning
Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
Signal Process. Image Commun. | 6 |
| 2025 | Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding ConfusionabstractContinual learning (CL) aims to enable deep neural networks (DNNs) to learn new data without forgetting previously learned knowledge. The key to achieving this goal is to avoid confusion at the feature level, i.e., to avoid confusion within old tasks and between new and old tasks. Existing prototype-based CL methods generate pseudo features for old knowledge replay by adding Gaussian noise to the centroids of old classes. However, the distribution in the feature space exhibits anisotropy during the incremental process, which prevents the pseudo features from faithfully reproducing the distribution of old knowledge in the feature space, leading to confusion at the classification boundaries within old tasks. To address this issue, we propose the distribution-level memory recall (DMR) method, which uses a Gaussian mixture model to precisely fit the feature distribution of old knowledge at the distribution level and generate pseudo features in the next stage. Furthermore, resistance to confusion at the distribution level is crucial for multimodal learning. Multimodal imbalance, which refers to uneven optimization processes among encoders of different modalities, results in significant differences in feature responses between modalities; this exacerbates confusion within old tasks in prototype-based CL methods. Therefore, we mitigate the multimodal imbalance problem by using the intermodal guidance and intramodal mining (IGIM) method to guide weaker modalities with prior information from dominant modalities and further explore useful information within modalities. To avoid confusion between new and old tasks, we propose using the confusion index to quantitatively describe a model's ability to distinguish between new and old tasks, and we use the incremental mixup feature enhancement (IMFE) method to enhance pseudo features with new sample features, alleviating classification confusion between new and old knowledge. We conduct extensive experiments on the CIFAR100, ImageNet100, TinyImageNet, ImageNet-1K and UESTC-MMEA-CL datasets and achieve state-of-the-art results. Shaoxu Cheng, Kanglei Geng, Chiyuan He, Zihuan Qiu, Linfeng Xu 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Cross-Modal Cognitive Consensus Guided Audio-Visual SegmentationabstractAudio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented reality, and intelligent robot systems. The pioneering work conducts this task through dense feature-level audio-visual interaction, which ignores the dimension gap between different modalities. More specifically, the audio clip could only provide aGlobalsemantic label in each sequence, but the video frame covers multiple semantic objects across differentLocalregions, which leads to mislocalization of the representationally similar but semantically different object. In this paper, we propose a Cross-modal Cognitive Consensus guided Network (C3N) to align the audio-visual semantics from the global dimension and progressively inject them into the local regions via an attention mechanism. Firstly, a Cross-modal Cognitive Consensus Inference Module (C3IM) is developed to extract a unified-modal label by integrating audio/visual classification confidence and similarities of modality-agnostic label embeddings. Then, we feed the unified-modal label back to the visual backbone as the explicit semantic-level guidance via a Cognitive Consensus guided Attention Module (CCAM), which highlights the local features corresponding to the interested object. Extensive experiments on the Single Sound Source Segmentation (S4) setting and Multiple Sound Source Segmentation (MS3) setting of the AVSBench dataset demonstrate the effectiveness of the proposed method, which achieves state-of-the-art performance. Zhaofeng Shi, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Learning With Noisy Low-Cost MOS for Image Quality Assessment via Dual-Bias CalibrationabstractLearning-based Image Quality Assessment (IQA) models have obtained impressive performance with the help of reliable subjective quality labels, where Mean Opinion Score (MOS) is the most popular choice. However, in view of the subjective bias of individual annotators, the Labor-Abundant MOS (LA-MOS) typically requires large collections of opinion scores from multiple annotators for each image, which significantly increases the learning cost. In this paper, we aim to learn robust IQA models from Low-Cost MOS (LC-MOS), which only requires very few opinion scores or even a single opinion score for each image. More specifically, we consider the LC-MOS as the noisy observation of LA-MOS and enforce the IQA model learned from LC-MOS to approach the unbiased estimation of LA-MOS. Thus, we represent the subjective bias between LC-MOS and LA-MOS, and the model bias between IQA predictions learned from LC-MOS and LA-MOS (i.e., dual-bias) as two latent variables with unknown parameters. By means of the expectation-maximization-based alternating optimization, we can jointly estimate the parameters of the dual-bias, which suppresses the misleading of LC-MOS via a gated dual-bias calibration (GDBC) module. To the best of our knowledge, this is the first exploration of robust IQA model learning from noisy low-cost labels. Theoretical analysis and extensive experiments on four popular IQA datasets show that the proposed method is robust toward different bias rates and annotation numbers and significantly outperforms the other Learning-based IQA models when only LC-MOS is available. Furthermore, we also achieve comparable performance with respect to the other models learned with LA-MOS. Lei Wang 0029, Qingbo Wu 0001, Desen Yuan, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Dual-Consistency Model Inversion for Non-Exemplar Class Incremental LearningabstractNon-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge without forgetting previously acquired ones when historical data are un-available. One of the generative NECIL methods is to in-vert the images of old classes for joint training. However, these synthetic images suffer significant domain shifts compared with real data, hampering the recognition of old classes. In this paper, we present a novel method termed Dual-Consistency Model Inversion (DCMI) to generate better synthetic samples of old classes through two pivotal consistency alignments: (1) the semantic consistency between the synthetic images and the corresponding prototypes, and (2) domain consistency between synthetic and real images of new classes. Besides, we introduce Prototypical Routing (PR) to provide task-prior information and generate unbi-ased and accurate predictions. Our comprehensive experiments across diverse datasets consistently showcase the superiority of our method over previous state-of-the-art approaches. Zihuan Qiu, Yi Xu 0008, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001 |
CVPR | 5 |
| 2024 | Vision-Sensor Attention Based Continual Multimodal Egocentric Activity RecognitionabstractContinual learning aims to equip deep neural networks (DNNs) with the capability to continuously learn new knowledge without catastrophic forgetting. Currently, there is significant attention on multimodal continual activity recognition from a egocentric perspective. However, the issue of modality imbalance can lead to exacerbated forgetting in multimodal continual learning. To address this, we propose an exemplar-free vision-sensor Attention-based Incremental Discriminability enhancement (AID) method. Firstly, we employ a Vision-Sensor attention module to enhance the time-frequency information of sensor modality and fuse them with vision modality. This alleviates the modality imbalance problem, yielding more discriminative and generalizable representations. Simultaneously, to prevent the classifier from overfitting to old class prototypes, we enhance old prototypes with features from new classes, thereby enhancing classifier discriminability. We validate the effectiveness of this method through numerous experiments with various task settings on the UESTC-MMEA-CL dataset. Shaoxu Cheng, Chiyuan He, Kailong Chen, Linfeng Xu 0001, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001 |
ICASSP | 4 |
| 2024 | Advancing zero-shot semantic segmentation through attribute correlations
Runtong Zhang, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, Hongliang Li 0001 |
Neurocomputing | 5 |
| 2024 | Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and BeyondabstractFew-shot segmentation (FSS) aims to segment the novel class with a few annotated images. Due to CLIP's advantages of aligning visual and textual information, the integration of CLIP can enhance the generalization ability of FSS model. However, even with the CLIP model, the existing CLIP-based FSS methods are still subject to the biased prediction towards base class, which is caused by the class-specific feature level interactions. To solve this issue, we propose a visual and textual Prior Guided Mask Assemble Network (PGMA-Net). It employs a class-agnostic mask assembly process to alleviate the bias, and formulates diverse tasks into a unified manner by assembling the prior through affinity. Specifically, the class-relevant textual and visual features are first transformed to class-agnostic prior in the form of probability map. Then, a Prior-Guided Mask Assemble Module (PGMAM) including multiple General Assemble Units (GAUs) is introduced. It considers diverse and plug-and-play interactions, such as visual-textual, inter- and intra-image, training-free, and high-order ones. Lastly, to ensure the class-agnostic ability, a Hierarchical Decoder with Channel-Drop Mechanism (HDCDM) is proposed to flexibly exploit the assembled masks and low-level features, without relying on any class-specific information. It achieves new state-of-the-art results in the FSS task, with mIoU of 77.6 on$\rm{PASCAL-}5^{i}$and 59.4 on$\rm{COCO-}20^{i}$in 1-shot scenario. Beyond this, we show that without extra re-training, the proposed PGMA-Net can solve bbox-level and cross-domain FSS, co-segmentation, zero-shot segmentation (ZSS) tasks, leading an any-shot segmentation framework capable of accommodating diverse weak or pixel annotations. Fanman Meng, Runtong Zhang, Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Linfeng Xu 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | CrowdCaption++: Collective-Guided Crowd Scenes CaptioningabstractCrowd scenes analysis plays an important role in various fields, including public security, smart cities, and intelligent transportation systems. However, traditional crowd scenes captioning methods mainly focus on a single and prominent crowd collective, which limits their ability to describe the different crowd collectives in complex crowd scenes. To address this issue, we propose a collective-guided crowd scenes captioning model (CrowdCaption++) to explore a more comprehensive and detailed description. We design a crowd features encoder (CFE) including double-query features encoder and foreground crowd features encoder, which uses double-query attention module (DQ-ATT) to capture more representative visual features and extracts foreground crowd features to avoid interference from background for collectives prediction. Moreover, we build a collective-guided captioning decoder (CCD) to generate captions of different crowd collectives without requiring extra alignment between crowd collectives and captions. To achieve this, we first design a crowd collectives predictor to identify multiple potential crowd collectives and create crowd collectives guidance information. Finally, we use the crowd collectives guidance information to merge useful visual features and further generate corresponding caption. We evaluate our approach on the latest crowd scenes dataset CrowdCaption and demonstrate that our model can achieve a comprehensive understanding and describe the different crowd collectives in complex crowd scenes. Lanxiao Wang, Hongliang Li 0001, Minjian Zhang 0003, Heqian Qiu, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Towards Continual Egocentric Activity Recognition: A Multi-Modal Egocentric Activity Dataset for Continual LearningabstractWith the rapid development of wearable cameras, it is now feasible to considerably increase the collection of egocentric video for first-person visual perception. However, the development is hindered by a shortage of multi-modal egocentric activity datasets. Furthermore, the catastrophic forgetting problem of multimodal continual activity learning, as a branch of continual learning, has not been thoroughly explored, which makes accumulating a larger collection of multi-modal activity data more urgent. To address this shortage, we propose a multi-modal egocentric activity dataset for continual activity learning named UESTC-MMEA-CL in this paper. The dataset is collected using our self-developed glasses with a first-person camera and wearable sensors, and it contains synchronized data of video, accelerometers, and gyroscopes for 32 types of daily activities performed by 10 participants who wore our glasses. Statistical analysis of the sensor data is given to show the auxiliary effects of activity recognition. We report the results of egocentric activity recognition of three modalities (RGB, acceleration, and gyroscope) separately and jointly on a base network architecture. We thoroughly evaluated four baseline methods with different multimodal combinations to explore the catastrophic forgetting in continual learning on UESTC-MMEA-CL. We hope that the UESTC-MMEA-CL dataset can act as a facilitator for future studies on continual learning for first-person activity recognition in wearable applications. You can download preliminary data fromhttps://ivipclab.github.io/publication_uestc-mmea-cl/mmea-cl. The data is currently used to solve the problems of multimodal continual learning of activities. Linfeng Xu 0001, Qingbo Wu 0001, Lili Pan 0001, Fanman Meng, Hongliang Li 0001, Chiyuan He, Hanxin Wang, Shaoxu Cheng |
IEEE Trans. Multim. | 1 |
| 2024 | Learning Offset Probability Distribution for Accurate Object DetectionabstractObject detection combines object classification and object localization problems. Current object detection methods heavily depend on regression networks to locate objects, which are optimized with various regression loss functions to predict offsets between candidate boxes and objects. However, these regression losses are difficult to assign the appropriate penalties for samples with large offset errors, resulting in suboptimal regression networks and inaccurate object offsets. In this article, we consider object location as offset bin classification problem, and propose a distance-aware offset bin classification network optimized with multiple binary cross entropy losses to learn various offset probability distribution, including single label distribution and distance-aware label distribution. On one hand, it provides gradient contributions for different samples based on the bounded probability instead of previous incalculable offset error. On the other hand, it explores the distance correlations between discrete offset bins to facilitate network learning. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin, in which the probability should be higher for the offset bin closer to the target offsets, and vice versa. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the precision of prediction. We conduct extensive experiments to evaluate the effectiveness of our method. In addition, our method can be conveniently and flexibly inserted into existing object detection methods, which consistently achieves a large gain based on popular anchor-based and anchor-free methods on the PASCAL VOC, MS-COCO, KITTI, and CrowdHuman datasets. Code will be released at: https://github.com/QiuHeqian/DBC . Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | MFAT: A Multi-Level Feature Aggregated Transformer for Person Re-IdentificationabstractRecently, with the development of the Transformer, re-identification (ReID) has great success in various applications. Existing works prefer to utilize the Transformer’s highest-level information as its discriminative feature, which focuses on a few concentrated parts or areas. However, in ReID filed, under such various scenes and camera views, only using a few concentrated parts to distinguish the query person is insufficient. Meanwhile, we find that Transformer’s lower-level information is also helpful for the recognition accuracy of the query person, especially, when the scene changes greatly. Therefore, we propose a Multi-level Feature Aggregated Transformer for person re-identification (MFAT) with high performance. To aggregate multi-level information, two novel modules are carefully designed. (i) The Global Content and Structure Aggregation (GCSA) module is proposed to aggregate multi-level information in a global manner. (ii) The Local Convolution Aggregation (LCA) module which consists of a series of convolutional blocks, is introduced to aggregate multi-level features with local operations. To the best of our knowledge, this is the first work to aggregate multi-level features with a Transformer backbone for person ReID task. Experiment results show that our method has achieved state-of-the-art on three person ReID benchmarks, with both Pyramid Vision Transformer (PVT) and Vision Transformer (ViT) backbones. Bowen Tan, Linfeng Xu 0001, Zihuan Qiu, Qingbo Wu 0001, Fanman Meng |
ICASSP | 2 |
| 2023 | ISM-Net: Mining incremental semantics for class incremental learning
Zihuan Qiu, Linfeng Xu 0001, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 2 |
| 2023 | GFR: Generic feature representations for class incremental learningabstractClass incremental learning (CIL) aims to continuously learn new classes while maintaining discrimination for old classes with sequentially coming data. Due to the lack of old-class samples, existing CIL methods fail to learn discriminative representations for both old and new classes simultaneously, resulting in a severe performance drop in old classes, which is the well-known catastrophic forgetting phenomenon. Different from most existing works, we facilitate CIL by learning generic feature representations that perform well in seen and unseen classes. Specifically, we prove that representations with a substantial number of significant singular values benefit CIL via better old knowledge reservation. However, the overly uniform singular value spectrum will hurt the discrimination of current tasks. Furthermore, we propose that increasing the embedding dimension can enhance the number of significant singular values and validate this assumption from two perspectives: adopting different pooling techniques and devising a wider network. Meanwhile, we also prove that satisfactory current task accuracy and old knowledge reservation can be achieved simultaneously. Finally, the simple yet effective generic feature representation regulation (GFR) is devised and incorporated into two baselines. Extensive experiments are conducted on CIFAR100, ImageNet-Subset, and ImageNet. The results show that the proposed method boosts the performance of both baselines with a large margin (2.00%-9.58% on CIFAR100, 0.68%-7.10% on ImageNet-Subset and 1.18%-5.04% on ImageNet) which outperforms existing SOTAs. Linfeng Xu 0001, Zihuan Qiu, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 2 |
| 2023 | DRDet: Dual-Angle Rotated Line Representation for Oriented Object DetectionabstractIn aerial scenes, oriented object detection is sensitive to the orientation of objects, which makes the formulation of orientation-aware object representation become a critical problem. Existing methods mostly adopt rectangle anchor or discrete points as object representation, which may lead to the feature aliasing between overlapping objects and ignore the orientation information of objects. To solve these issues, we propose a novel anchor-free oriented object detection network named DRDet, which adopts Dual-angle Rotated Lines (DRL) as object representation. Different from other object representations, DRL can adaptively rotate and extend to the boundary of the object according to its orientation and shape, which explicitly introduces the orientation information into the formulation of object representation. And it can adaptively cope with the geometric deformation of objects. Based on the dual-angle rotated lines, we design an Orientation-guided Feature Encoder (OFE) to encode discriminant object feature along each rotated line, respectively. Instead of encoding rectangle feature, the OFE module adopts line features for orientation-guided feature encoding, which can alleviate the feature aliasing between neighboring objects or background. To further enhance the flexibility of dual-angle rotated lines, we design a Dual-angle Decoder (DD) that predicts two angle offsets according to the orientation-guided feature and converts the angle offsets and regression offsets into dual-angle rotated line representation, which can help to guide the adaptive rotation of each rotated line, respectively. Our proposed method achieves consistent improvement on both DOTA and HRSC2016 datasets. Extensive experimental results verify the effectiveness of our method in oriented object detection. Minjian Zhang 0003, Heqian Qiu, Hefei Mei, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Forgetting to Remember: A Scalable Incremental Learning Framework for Cross-Task Blind Image Quality AssessmentabstractRecent years have witnessed the great success of blind image quality assessment (BIQA) in various task-specific scenarios, which present invariable distortion types and evaluation criteria. However, due to the rigid structure and learning framework, they cannot apply to the cross-task BIQA scenario, where the distortion types and evaluation criteria keep changing in practical applications. This paper proposes a scalable incremental learning framework (SILF) that could sequentially conduct BIQA across multiple evaluation tasks with limited memory capacity. More specifically, we develop a dynamic parameter isolation strategy to sequentially update the task-specific parameter subsets, which are non-overlapped with each other. Each parameter subset is temporarily settled toRememberone evaluation preference toward its corresponding task, and the previously settled parameter subsets can be adaptively reused in the following BIQA to achieve better performance based on the task relevance. To suppress the unrestrained expansion of memory capacity in sequential tasks learning, we develop a scalable memory unit by gradually and selectively pruning unimportant neurons from previously settled parameter subsets, which enable us toForgetpart of previous experiences and free the limited memory capacity for adapting to the emerging new tasks. Extensive experiments on eleven IQA datasets demonstrate that our proposed method significantly outperforms the other state-of-the-art methods in cross-task BIQA. The source code of the proposed method is available atgithub.com/maruiperfect/SILF. Rui Ma 0030, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Bias-Correction Feature Learner for Semi-Supervised Instance SegmentationabstractInstance segmentation is heavily reliant on large-scale annotated datasets to yield an ideal accuracy. However, annotated data are difficult to collect. To expand the annotated data, a straightforward idea is to introduce semi-supervised learning, which uses a trained model to obtain initial proposals on unlabeled images and then use initial proposals to generate pseudo labels. However, existing methods inevitably introduce the bias for the model learning, i.e., the foreground in initial low-confident proposals (low-confident foreground) is arbitrarily assigned as background. This bias makes the foreground and background closer in the feature space, which degenerates the model accuracy. To address this issue, this paper discards incorrect supervision and designs a bias-correction feature learner. Specifically, on the one hand, low-confident foreground does not participate in supervised learning. On the other hand, we extract possible foreground regions from all initial proposals to construct high-quality positive pairs which depict objects of the same category in contrastive learning. Then, positive pairs are pulled closer in the feature space. This helps models extract closely clustered foreground features. Experimental results demonstrate the effectiveness of our method on the public datasets (i.e., COCO, Cityscapes and Pascal VOC). Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu, Linfeng Xu 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Eldnet: Establishment and Refinement of Edge Likelihood Distributions for Camouflaged Object DetectionabstractCamouflaged object detection (COD) focuses on detecting objects assimilated into surroundings. Distractions from the background make it extremely challenging to find the correct semantics and border details. Some of the current methods aim to explore edge details directly which has the risk of leading to boundary errors under strong background interference. While other conservative methods find rough semantic locations, then implicitly infer the edge against the background which cannot fully pay attention to the spatial details manifesting as partial semantic loss and blurring. Different from the existence, we propose a framework that generates progressively refined edge likelihood maps to guide the feature fusion of camouflaged objects. To get robust and refined edge likelihood maps, a module is first trained to fit the reasonable edge likelihood distribution based on semantic position clues and then a progressive architecture is designed to refine the possible edge area which gradually approaches the real edge. Our method can avoid catastrophic errors during exploring the fine edge, but also improves border details, solving the border blurring. Experiments on four datasets demonstrate the significant improvements and generalization ability of ours compared with state-of-the-art methods. Chiyuan He, Linfeng Xu 0001, Zihuan Qiu |
ICIP | 2 |
| 2022 | Spatial-Semantic Attention for Grounded Image CaptioningabstractGrounded image captioning models usually process high-dimensional vectors from the feature extractor to generate descriptions. However, mere vectors do not provide adequate information. The model needs more explicit information for grounded image captioning. Besides high dimensional vectors, the feature extractor also predicts the locations and categories of the objects, which contains low-level spatial information and high-level semantic information. To this end, we propose a new attention module called Spatial-Semantic (SS) Attention, which utilizes the predictions from the backbone network to help the model attend to the correct objects. Specifically, the SS attention module collects the position of proposals and the class probabilities from the feature extractor as spatial and semantic information to assist attention weighting. In addition, we propose a grounding loss to supervise the SS attention. Our method achieves high performance on captioning and grounding metrics and outperforms some powerful previous models on the Flickr30k Entities dataset. Wenzhe Hu, Lanxiao Wang, Linfeng Xu 0001 |
ICIP | 3 |
| 2022 | Cross-Domain Object Detection with Missing Classes in Target DomainabstractMany existing methods focus on detecting either objects from different domains or those of rare classes, but it's difficult for them to tackle the two issues together. However, in the real world, due to the difficulty of collecting samples of special classes, deep learning practitioners have to use simulated images to substitute for them. To deal with this scenario, in this paper, we research a new task: cross-domain object detection with missing classes in target domain, where there are only partial classes have images and annotations in the target domain. We devise a simple but effective play-and-plug method to address this new task, named the three-stage learning approach with domain and class information preservation. In addition, extensive experiments demonstrate our method is effective and can boost the performance when added to existing unsupervised domain adaptation object detectors. Benliu Qiu, Heqian Qiu, Haitao Wen, Zichen Song 0002, Linfeng Xu 0001 |
MMSP | 5 |
| 2022 | Dynamic Perceptual Quality Ranking based Autofocus Method for ProjectorabstractProjectors are widespread in families, schools and companies nowadays. Common projectors generally use laser ranging to realize autofocus, yet this kind of autofocus is limited by the highly-cost temperature-sensitive devices, so long-term operation in projectors “high-temperature environment may lead to autofocus failure. To find an more efficient and lower-cost solution, we made an autofocus dataset, which contains 8 category folders with a total of 80 videos, each containing 174 pictures with labels. For the autofocus task, the most intuitive idea is to use an image quality assessment algorithm to score each image with the best value corresponding to the best focal length, but the fact is that the image quality assessment algorithm will fail due to the large number of categories and the large span of video contents. In this paper, we proposed a light-weight dynamic perceptual quality ranking based autofocus method to achieve fascinate results on the dataset. Lanjiang Wang, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001 |
MMSP | 4 |
| 2022 | Instance-level Context Attention Network for instance segmentation
Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Heqian Qiu, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
Neurocomputing | 6 |
| 2022 | Segmenting Beyond the Bounding Box for Instance SegmentationabstractInstance segmentation needs to locate all instances in an image correctly and segment each instance precisely. Currently, the most dominant methods for instance segmentation take object detection as a pre-task. However, they rely on the accuracy of object detection incredibly. If the pre-task cannot predict an accurate bounding box, the performance of instance segmentation will degenerate. In this paper, we present a novel method for instance segmentation to solve this problem, which is calledSegmentingBeyond theBoundingBox (S3B-Net). Our S3B-Net designs a sub-network to help instance segmentation methods based on object detection to segment the part of an instance beyond the bounding box. Specifically, the sub-network first predicts a two-dimensional pixel embedding for each pixel. Then, the Gaussian function is employed to calculate a pixel’s probability belongs to a corresponding instance according to the two-dimensional pixel embedding. Finally, the output of the sub-network combines with the output of instance segmentation based on object detection to generate a more precise instance mask. Our sub-network can easily extend on the existing instance segmentation method based on object detection to segment instance beyond the bounding box. We do our experiments on dominant instance segmentation datasets, such as the COCO dataset and Cityscapes dataset. The results show that our method can achieve 6.8 points gain compared with the baseline Mask R-CNN with ResNet-50-FPN in Cityscapes datasets, and 1.7 points gain with ResNet-101-FPN-DCN in COCO datasets. Our S3B-Net outperforms the previous state-of-the-art instance segmentation method, which proves our method is competitive. The source code of our method will be made available. Xiaoliang Zhang 0002, Hongliang Li 0001, Fanman Meng, Zichen Song 0002, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Remember and Reuse: Cross-Task Blind Image Quality Assessment via Relevance-aware Incremental LearningabstractExisting blind image quality assessment (BIQA) methods have made great progress in various task-specific applications, including the synthetic, authentic, or over-enhanced distortion evaluations. However, limited by the static model and once-for-all learning strategy, they failed to perform the cross-task evaluations in many practical applications, where diverse evaluation criteria and distortion types are constantly emerging. To address this issue, in this paper, we propose a dynamic Remember and Reuse (R&R) network, which efficiently performs the cross-task BIQA based on a novel relevance-aware incremental learning strategy. Given multiple evaluation tasks across different distortion types or databases, our R&R network sequentially updates the parameters for every task one by one. After each update step, part of task-specific parameters is settled, which ensures R&R Remembers their dedicated evaluation preferences. The remaining parameters are pruned for the dynamic usage of the subsequent tasks. To further exploit the correlation between different tasks, we feed the training data of a new task to previously settled parameters. Better prediction accuracy is considered as higher task relevance and vice versa. Then, we selectively Reuse parts of previously settled parameters, whose proportion is adaptively determined by the task relevance. Extensive experiments show that the proposed method efficiently achieves the cross-task BIQA without catastrophic forgetting, and significantly outperforms many state-of-the-art methods. Code is available at https://github.com/maruiperfect/R-R-Net. Rui Ma 0030, Hanxiao Luo, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
ACM Multimedia | 7 |
| 2021 | Few-Shot Segmentation via Complementary Prototype Learning and Cascaded Refinement
Hanxiao Luo, Hui Li 0080, Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Fanman Meng, Linfeng Xu 0001 |
PRCV (4) | 7 |
| 2021 | Behaviour detection in crowded classroom scenes via enhancing features robust to scale and perspective variationsabstractAbstract Detecting human behaviours in images of crowded classroom scenes is a challenging task, due to the large variations of humans in scale and pose perspective. In this paper, two modules are proposed to tackle these two variations. First, an attention‐based RoI (region‐of‐interest) extractor is designed to handle scale variation. Feature fusion and attention mechanism are used to improve the RoI feature with more local and global information. Second, a transformation‐based detection head is introduced to handle perspective variation. The spatial transformation is adopted to extract consistent representation under various perspectives. Moreover, since there is a lack of proper datasets for human behaviour detection in classroom scenes, a new dataset is created, namely CLBD. The experiments on the proposed dataset demonstrate that the modules obtain significant improvements of performance over the state‐of‐the‐art detectors. Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, Qianghua Liao |
IET Image Process. | 4 |
| 2021 | A Deep Supervised Edge Optimization Algorithm for Salt Body SegmentationabstractSeismic image analysis plays a vital role in a wide range of industrial applications and has received widespread attention. One of the main challenges of seismic image analysis is detecting underground salt structures, which is essential for identifying oil and gas reservoirs and planning drilling paths. Currently traditional seismic image analysis still requires professionals to analyze the salt body. Convolutional neural networks have been successfully applied in many fields, and several attempts have been made in the field of seismic imaging. In this letter, we propose a deep-supervised method which effectively segments the salt body. We design an edge prediction branch to predict the boundary of the salt body, which guides feature learning through the supervision of boundary loss, so that the network can make the features on both sides of the semantic boundary distinguishable. We show that our approach outperforms state-of-the-art methods on the TGS Salt Identification Challenge data set and experimental results demonstrate the effectiveness of the proposed method. The source code is available at GitHub.The source code is available at.11The source code is available athttps://github.com/Gjiangtao/A-Deep-Supervised-Edge-Optimization-Algorithm-for-Salt-Body-Segmentation. Jiangtao Guo 0003, Linfeng Xu 0001, Jisheng Ding, Shengxuan Dai |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | High-Quality R-CNN Object Detection Using Multi-Path Detection Calibration NetworkabstractObject proposals are used in two-stage detectors, such as R-CNN, to generate detection results, including category predictions and refined bounding-boxes. As a result, classification scores are assigned to refined bounding-boxes rather than object proposals. However, this procedure ignores the discrepancy of data distribution between object proposals and refined bounding-boxes. We consider this discrepancy could limit the detection accuracy. Specifically, the foreground/background imbalance on object proposals and inaccurate information from low-IoU proposals could hinder the category prediction. In this paper, we propose a detector called the Multi-Path Detection Calibration Network (PDC-Net) to address this problem. The key idea behind PDC-Net is calibrating detection results from R-CNN by considering the statistical discrepancy between object proposals and refined bounding-boxes. PDC-Net is built on Faster R-CNN. The core component in PDC-Net is the multi-path detection head, in which the base detector (from Faster R-CNN) generates detection results from object proposals and multiple calibration detectors fix incorrect outputs from the base detector using refined bounding-boxes. Experiments reveal that PDC-Net can boost detection results. Our method could reach 83.1% and 43.3% mAP respectively on PASCAL VOC and MSCOCO benchmarks, which is comparable to several state-of-the-art methods. Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Non-Homogeneous Haze Removal via Artificial Scene Prior and Bidimensional Graph ReasoningabstractDue to the lack of natural scene and haze prior information, it is greatly challenging to completely remove the haze from a single image without distorting its visual content. Fortunately, the real-world haze usually presents non-homogeneous distribution, which provides us with many valuable clues in partial well-preserved regions. In this paper, we propose a Non-Homogeneous Haze Removal Network (NHRN) via artificial scene prior and bidimensional graph reasoning. Firstly, we employ the gamma correction iteratively to simulate artificial multiple shots under different exposure conditions, whose haze degrees are different and enrich the underlying scene prior. Secondly, beyond utilizing the local neighboring relationship, we build a bidimensional graph reasoning module to conduct non-local filtering in the spatial and channel dimensions of feature maps, which models their long-range dependency and propagates the natural scene prior between the well-preserved nodes and the nodes contaminated by haze. To the best of our knowledge, this is the first exploration to remove non-homogeneous haze via the graph reasoning based framework. We evaluate our method on different benchmark datasets. The results demonstrate that our method achieves superior performance over many state-of-the-art algorithms for both the single image dehazing and hazy image understanding tasks. The source code of the proposed NHRN is available on https://github.com/whrws/NHRNet. Qingbo Wu 0001, Hui Li 0080, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Image Process. | 7 |
| 2020 | Human Detection In Dense Scene Of ClassroomsabstractThe intelligent classroom project has gained a lot of attention in recent years. One of the critical applications is to recognize the position and category of an object accurately in the classrooms. Object detector based on CNN is widely used to address this issue to realize more accurate performance. But in the scene of classrooms, dense distribution of people, like students and teachers, always leads to serious class-class or class-in occlusion and brings poor accuracy to current typical detectors. To solve this problem, we propose an original method named Dense Occlusion Object Detection network, which consists of Dense Anchor Generation Model (DAG) and Discriminative Part Selection Model (DPS). The DAG model could exploit crucial semantic information from feature map to generate accurate anchor boxes by locating center point and predicting box size of an object in the anchor-free manner. The DPS model aims to mitigate the occlusion problem through dividing the proposal into some smaller parts and processing them respectively so that we can select the most discriminative ones to update the confidence score of the input proposal. Experimental results show that the proposed method outperforms the state-of-the-art methods for human detection on our CScene dataset. Jisheng Ding, Linfeng Xu 0001, Jiangtao Guo 0003, Shengxuan Dai |
ICIP | 2 |
| 2020 | Haze-robust image understanding via context-aware deep feature refinementabstractImage understanding under the foggy scene is greatly challenging due to inhomogeneous visibility deterioration. Although various image dehazing methods have been proposed, they usually aim to improve image visibility (such as, PSNR/SSIM) in the pixel space rather than the feature space, which is critical for the perception of computer vision. Due to this mismatch, existing dehazing methods are limited or even adverse in facilitating the foggy scene understanding. In this paper, we propose a generalized deep feature refinement module to minimize the difference between clear images and hazy images in the feature space. It is consistent with the computer perception and can be embedded into existing detection or segmentation backbones for joint optimization. Our feature refinement module is built upon the graph convolutional network, which is favorable in capturing the contextual information and beneficial for distinguishing different semantic objects. We validate our method on the detection and segmentation tasks under foggy scenes. Extensive experimental results show that our method outperforms the state-of-the-art dehazing based pretreatments and the fine-tuning results on hazy images. Hui Li 0080, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
MMSP | 7 |
| 2020 | A Unified Single Image De-raining Model via Region Adaptive Coupled NetworkabstractSingle image de-raining is quite challenging due to the diversity of rain types and inhomogeneous distributions of rainwater. By means of dedicated models and constraints, existing methods perform well for specific rain type. However, their generalization capability is highly limited as well. In this paper, we propose a unified de-raining model by selectively fusing the clean background of the input rain image and the well restored regions occluded by various rains. This is achieved by our region adaptive coupled network (RACN), whose two branches integrate the features of each other in different layers to jointly generate the spatial-variant weight and restored image respectively. On the one hand, the weight branch could lead the restoration branch to focus on the regions with higher contributions for de-raining. On the other hand, the restoration branch could guide the weight branch to keep off the regions with over-/under-filtering risks. Extensive experiments show that our method outperforms many state-of-the-art de-raining algorithms on diverse rain types including the rain streak, raindrop and rain-mist. Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
VCIP | 6 |
| 2020 | HeadNet: An End-to-End Adaptive Relational Network for Head DetectionabstractHead detection plays an important role in localizing and identifying persons from visual data. Most existing methods treat head detection as a specific form of object detection. Head detection is nontrivial due to the considerable difficulty in building the local and global information under conditions of unconstrained pose and orientation. To address these issues, this paper presents an effective adaptive relational network to capture context information, which is greatly helpful to suppress missed detection. We show that the fundamental contextual properties, such as the global shape priors from different heads and the local adjacent relationship between the head and shoulders, can be systematically quantified by visual operators. Specifically, we propose a two-step search algorithm to quantify the global intergroup conflict with adaptive scale, pose and viewpoint. Meanwhile, a structured feature module is introduced to capture the local relation of intraindividual stability. Finally, the global priors and local relation are integrated seamlessly into a single-stage head detector that is end-to-end trainable. An extensive ablation analysis demonstrates the effectiveness of our approach. We achieve state-of-the-art results on two challenging datasets, i.e., HollywoodHeads and Brainwash. Wei Li 0110, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Subjective and Objective De-Raining Quality Assessment Towards Authentic Rain ImageabstractImages acquired by outdoor vision systems easily suffer poor visibility and annoying interference due to the rainy weather, which brings great challenge for accurately understanding and describing the visual contents. Recent researches have devoted great efforts on the task of rain removal for improving the image visibility. However, there is very few exploration about the quality assessment of de-rained image, even it is crucial for accurately measuring the performance of various de-raining algorithms. In this paper, we first create a de-raining quality assessment (DQA) database that collects 206 authentic rain images and their de-rained versions produced by 6 representative single image rain removal algorithms. Then, a subjective study is conducted on our DQA database, which collects the subject-rated scores of all de-rained images. To quantitatively measure the quality of de-rained image with non-uniform artifacts, we propose a bi-directional feature embedding network (B-FEN) which integrates the features of global perception and local difference together. Experiments confirm that the proposed method significantly outperforms many existing universal blind image quality assessment models. To help the research towards perceptually preferred de-raining algorithm, we will publicly release our DQA database and B-FEN source code on https://github.com/wqb-uestc. Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Hierarchical Context Features Embedding for Object DetectionabstractPixel-level segmentation has been widely used to improve object detection. Most of the existing methods refine detection features by adding the constraint of the segmentation branch or by simply embedding high-level segmentation features into detection features within the local receptive field. However, noisy segmentation features are unavoidable in real-word applications and can easily cause false positives. To address this problem, we propose a novel hierarchical context embedding module to effectively embed segmentation features into detection features. The idea of this module is to capture hierarchical context information that includes local objects or parts and nonlocal context features by learning multiple attention maps, and subsequently utilize interdependencies between features to recalibrate noisy segmentation features. Furthermore, we use this module in the proposed gated encoder-decoder network that adaptively aggregates feature maps of different resolutions based on the gate mechanism so that we can embed multiscale segmentation feature maps into detection features for more accurate detection of objects of all sizes. Experimental results demonstrate the effectiveness of the proposed method on the Pascal VOC 2012Seg dataset, the Pascal VOC dataset and the MS COCO dataset. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan, Hengcan Shi |
IEEE Trans. Multim. | 5 |
| 2019 | Instance Segmentation by Learning Deep Feature in Embedding SpaceabstractThe proposal-based framework is the mainstream architecture for instance segmentation. However, such architecture typically ignores the interference between objects, which fails to correctly segment overlapping objects with same category or appearance. In this paper, we propose a new instance segmentation network named Instance Discrimination Network (ID-Net) to consider the interference between objects by mapping pixels into an embedding space so that the pixels from different objects can be distinguished more accurately. To identify the foreground object in RoI, we learn a discriminative deep feature that can represent the embedding vectors corresponding to the foreground. Then, we get the foreground confidence map by calculating the similarities between the deep feature and embeddings. The experiments on PASCAL VOC and COCO datasets demonstrate the effectiveness of our method. Chao Shang 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001 |
ICIP | 4 |
| 2018 | Hierarchical Parsing Net: Semantic Scene Parsing From Global Scene to ObjectsabstractThis paper proposes a novel Hierarchical Parsing Net (HPN) for semantic scene parsing. Unlike previous methods, which separately classify each object, HPN leverages global scene semantic information and the context among multiple objects to enhance scene parsing. On the one hand, HPN uses the global scene category to constrain the semantic consistency between the scene and each object. On the other hand, the context among all objects is also modeled to avoid incompatible object predictions. Specifically, HPN consists of four steps. In the first step, we extract scene and local appearance features. Based on these appearance features, the second step is to encode a contextual feature for each object, which models both the scene-object context (the context between the scene and each object) and the interobject context (the context among different objects). In the third step, we classify the global scene and then use the scene classification loss and a backpropagation algorithm to constrain the scene feature encoding. In the fourth step, a label map for scene parsing is generated from the local appearance and contextual features. Our model outperforms many state-of-the-art deep scene parsing networks on five scene parsing databases. Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 5 |
| 2017 | Blind proposal quality assessment via deep objectness representation and local linear regressionabstractThe quality of object proposal plays an important role in boosting the performance of many computer vision tasks, such as, object detection and recognition. Due to the absence of manually annotated bounding-box in practice, the quality metric towards blind assessment of object proposal is highly desirable for singling out the optimal proposals. In this paper, we propose a blind proposal quality assessment algorithm based on the Deep Objectness Representation and Local Linear Regression (DORLLR). Inspired by the hierarchy model of the human vision system, a deep convolutional neural network is developed to extract the objectness-aware image feature. Then, the local linear regression method is utilized to map the image feature to a quality score, which tries to evaluate each individual test window based on its k-nearest-neighbors. Experimental results on a large-scale IoU labeled dataset verify that the proposed method significantly outperforms the state-of-the-art blind proposal evaluation metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Linfeng Xu 0001 |
ICME | 5 |
| 2017 | Store classification using Text-Exemplar-Similarity and Hypotheses-Weighted-CNN
Chao Huang 0003, Hongliang Li 0001, Wei Li 0110, Qingbo Wu 0001, Linfeng Xu 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | Manifold-ranking embedded order preserving hashing for image semantic retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | Feature discovering for image classification via wavelet-like pattern decomposition
Yurui Xie, Hongliang Li 0001, Chao Huang 0003, Bo Wu 0013, Linfeng Xu 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2015 | An effective vector model for global-contrast-based saliency detection
Linfeng Xu 0001, Liaoyuan Zeng, Huiping Duan |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Integrating bottom-up and top-down visual stimulus for saliency detection in news video
Bo Wu 0013, Linfeng Xu 0001 |
Multim. Tools Appl. | 2 |
| 2014 | Semantic superpixel extraction via a discriminative sparse representation
Yurui Xie, Chao Huang 0003, Linfeng Xu 0001 |
Multim. Tools Appl. | 3 |
| 2013 | A visual attention model for news videoabstractIn this paper, a novel method is proposed to perform saliency detection in news video. This method comprises bottom-up attention model which considers low level features to produce bottom-up saliency map and top-down attention model which utilizes high level factors to generate top-down saliency map. In bottom-up attention model, color image is represented as quaternion. Then the quaternion discrete cosine transform is used to detect static saliency in multi-scale and two color spaces. Meanwhile, the multi-scale local and global motion conspicuity maps are computed. To suppress the background motion noise, a novel histogram of average optical flow is proposed to calculate motion contrast. Then, the static saliency map and motion saliency map are fused after normalization. In top-down attention model, we explore high level factors of news video and generate the top-down saliency map based on these factors. Finally, the bottom-up and top-down saliency maps are integrated after normalization. Experiment results show that our method outperforms several state-of-the-art methods in saliency detection of news videos. Bo Wu 0013, Linfeng Xu 0001, Guanghui Liu 0001 |
ISCAS | 2 |
| 2013 | Saliency detection using a central stimuli sensitivity based modelabstractIn this paper, a novel method is proposed to predict attention in image scenes by using a central stimuli sensitivity based saliency model. The proposed method is based on the general “center-surround” visual attention mechanism and the spatial frequency response of the human visual system (HVS). Following three biologically inspired principles, the saliency value is computed by two “scatter matrices” which are used to measure the similarity and distinctness within and between two classes, i.e., the center and surrounding regions, respectively. In order to detect salient objects with different size, the saliency of a pixel is estimated via the saliency support region of the pixel, which is the most salient region centered at the pixel with respect to the surrounding region. The proposed method which is compliant with human perceptual characteristics enables the prediction of human fixations. Experimental results on three eye tracking datasets verify the effectiveness of the method and show that the proposed method outperforms the state-of-the-art methods on the visual saliency detection task. Linfeng Xu 0001, Hongliang Li 0001, Liaoyuan Zeng, Zhengning Wang, Guanghui Liu 0001 |
ISCAS | 1 |
| 2013 | Mode dependent loop filter for intra prediction coding in H.264/AVC
Qingbo Wu 0001, Linfeng Xu 0001, Liaoyuan Zeng, Jian Xiong 0005 |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Saliency detection using joint spatial-color constraint and multi-scale segmentation
Linfeng Xu 0001, Hongliang Li 0001, Liaoyuan Zeng, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | Two-layer average-to-peak ratio based saliency detection
Hongliang Li 0001, Linfeng Xu 0001, Guanghui Liu 0001 |
Signal Process. Image Commun. | 2 |
| 2013 | Face Hallucination via Similarity ConstraintsabstractIn this letter, we present a new face hallucination method based on similarity constraints to produce a high-resolution (HR) face image from an input low-resolution (LR) face image. This method is modeled as a local linear filtering process by incorporating four constraint functions at patch level. The first two constraints focus on checking if the training images are similar to the input face image. The third is defined in the HR face image, which is to impose the smoothness constraint between neighboring hallucinated patches. The final constraint computes the spatial distance to reduce the effect of patches that are far from the hallucinating patch. Experimental evaluation on a number of face images demonstrates the good performance of the proposed method on the face hallucination task. Hongliang Li 0001, Linfeng Xu 0001, Guanghui Liu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2012 | Saliency detection from joint embedding of spatial and color cuesabstractVisual saliency detection provides an important methodology for many computer vision applications. In this paper, we propose a novel method to detect salient regions from an image. To detect pixel-level saliency, this method uses joint embedding of spatial and color cues, i.e., spatial constraint based saliency, color double-opponent saliency, and similarity distribution based saliency. Finally, a multi-layer structure is adopted to merge the three terms into a saliency map. In order to make the saliency map consistent, we perform a region-based saliency detection by incorporating a multi-scale segmentation technique. The proposed method was evaluated on the MSRA benchmark images. Experimental results show that our method outperforms the state-of-the-art methods on visual saliency detection by achieving both higher precision and better recall. Linfeng Xu 0001, Hongliang Li 0001, Zhengning Wang |
ISCAS | 1 |