VLDB 2026 Research / reviewers in the wild / expert
Lihuo He
dblp:64/8741
· DBLP profile ↗
74ranked-venue papers
5as first author
46since 2021 · last 2026
0000-0002-0555-3574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 2 first-author · 29 since 2021Artificial intelligence and machine learning · 31 · 4 first-author · 18 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code GenerationabstractJiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang, Haibo Qiu, Lihuo He, Dengpan Ye, Xinbo Gao, Jing Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chi Zhang 0080, Qiming Zhang 0001, Haibo Qiu, Lihuo He, Dengpan Ye, Xinbo Gao 0001, Jing Zhang 0037 |
ACL (1) | 6 |
| 2026 | HVS-inspired blind image quality index with prominent perception learning and multi-level progressive integration
Taiyang Chen, Bo Hu 0008, Chunyi Li 0001, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
Neurocomputing | 5 |
| 2026 | DINO-PCB: Two-stage vision foundation model pretraining and distillation for real-time circuit-board defect detection
Junjie Ke, Lihuo He, Jing Zhang 0037, Yuqi Ji, Hui Chen 0013, Jie Li 0001, Sicheng Zhao, Guiguang Ding, Xinbo Gao 0001 |
Pattern Recognit. | 2 |
| 2026 | EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality AssessmentabstractModeling visual perception in a manner consistent with human subjective evaluation has become a central direction in both video quality assessment (VQA) and broader visual understanding tasks. While free-energy-guided self-repair mechanisms—reflecting human observational experience—have proven effective in image quality assessment, extending them to VQA remains non-trivial. In addition, biologically inspired paradigms such as holistic perception, local analysis, and gaze-driven scanning have achieved notable success in high-level vision tasks, yet their potential within the VQA context remains largely underexplored. To address these issues, we propose EyeSimVQA, a novel VQA framework that incorporates free-energy-based self-repair. It adopts a dual-branch architecture, with an aesthetic branch for global perceptual evaluation and a technical branch for fine-grained structural and semantic analysis. Each branch integrates specialized enhancement modules tailored to distinct visual inputs—resized full-frame images and patch-based fragments—to simulate adaptive repair behaviors. We also explore a principled strategy for incorporating high-level visual features without disrupting the original backbone. In addition, we design a biologically inspired prediction head that models sweeping gaze dynamics to better fuse global and local representations for quality prediction. Experiments on five public VQA benchmarks demonstrate that EyeSimVQA achieves competitive or superior performance compared to state-of-the-art methods, while offering improved interpretability through its biologically grounded design. Our code will be publicly available at https://github.com/handsomewzy/EyeSim-VQA. Zhaoyang Wang 0003, Wen Lu 0004, Jie Li 0001, Lihuo He, Maoguo Gong, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Efficient and Accurate Object Detection With Asymmetric Progressive Semi-Decoupled Head and Harmonic Focal LossabstractEfficiently and accurately recognizing interesting objects within the image and regressing bounding boxes to enclose them has been a persistent pursuit in object detection. However, existing detectors fail to achieve both aspects simultaneously due to insufficient task interaction and suboptimal classification behavior. To solve the problem, this paper proposes a novel detector with Efficient Asymmetric Progressive Semi-Decoupled Head (EAPSDH) and Harmonic Focal Loss (HFL). Specifically, we generalize the detection head into a progressive asymmetric paradigm that performs hierarchical and dynamically recalibrated interaction between classification and localization, enabling iterative mutual enhancement in an efficient manner beyond the prior designs. Meanwhile, HFL is proposed to improve classifier optimization by addressing the imbalance between positive and negative samples. HFL dynamically increases the loss weights of positive samples, amplifying their gradient contributions during classifier training, which significantly reduces classification error. By jointly improving task-specific feature representation and classification optimization, EAPSDH and HFL complement each other to alleviate the inconsistency between classification and localization performance, resulting in an efficient and accurate one-stage detector termed EADet. Experimental results on the MS COCO database demonstrate that EADet effectively mitigates the inconsistency between classification and localization performance. Furthermore, EADet achieves a strong trade-off between accuracy and speed, reaching 47.4 AP at 33.2 FPS on the MS COCO with ResNet-101 under the $2\times $ training schedule, demonstrating its effectiveness compared with recent state-of-the-art detectors. Code will be available at https://github.com/HB-X/EADet. Bo Han 0004, Lihuo He, Junjie Ke, Jiehao Tang, Di Wang 0011, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Toward Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge MiningabstractExisting object detection methods struggle to generalize across increasingly data domains while simultaneously adapting to the emergence of novel categories. To tackle this challenge, adaptive open-set object detection (AOOD) has been introduced, which employs supervised training on base categories within the source domain while enabling unsupervised adaptation to both base and novel categories in the target domain. However, existing AOOD approaches are still hindered by several limitations, including insufficient cross-domain feature representation, inter-category ambiguity in novel classes, and inherent feature bias toward the source domain. To overcome these issues, this paper proposes a category-level collaboration knowledge mining strategy designed to comprehensively exploit both inter-class and intra-class feature relationships across domains. Specifically, a clustering-based memory bank (CMB) is initially constructed to aggregate class prototype features, class auxiliary features, and intra-class disparity features, thereby embedding rich category-level knowledge into a unified memory structure. The CMB is iteratively updated through unsupervised clustering, which facilitates the modeling of intra-category relationships and enhances its capacity for cross-domain knowledge representation. Subsequently, a base-to-novel selection metric (BNSM) is designed to identify features corresponding to novel categories within the source domain by regulating the relationships between the novel categories and each base category. The selected features are then leveraged to initialize the object detector for the classification of novel categories. Finally, an adaptive feature assignment (AFA) strategy is introduced to transfer the learned category-level knowledge to the target domain, enabling the assignment of category labels to features. The memory bank is updated asynchronously with these assigned features to mitigate source domain bias. Extensive experiments conducted on diverse domain datasets demonstrate that the proposed method consistently outperforms state-of-the-art AOOD approaches, achieving performance gains of 1.1 to 5.5 mAP. Code is available at https://github.com/Jandsome/CCKM. Yuqi Ji, Junjie Ke, Lihuo He, Lizhi Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | IBCL-VQA: Video Quality Assessment with In-Batch Contrastive Learning and Two-Phase Feature FusionabstractVideo Quality Assessment (VQA) technology is of significant importance for improving video transmission, storage, and processing. Although Convolutional Neural Networks (CNNs)-based and Transformer-based methods have achieved significant progress, they still suffer from some drawbacks. The previous methods treated video data as independent samples, thereby neglecting the close or distant relationships between different quality levels and consequently constraining the model’s discriminative ability; meanwhile, most existing fusion strategies utilize fixed architectures that lack the ability to adapt to the data for optimal integration, resulting in insufficient utilization of spatio-temporal information and limited expression capabilities of the fused features. To address the above issues, this article proposes a VQA method with In-Batch Contrastive Learning and Two-Phase Feature Fusion (IBCL-VQA). Firstly, the spatial features are extracted through two branches, which not only preserve the global semantics but also focus on the local regions. The temporal characteristics are obtained through a pre-trained video recognition model. Secondly, we propose an in-batch contrastive learning mechanism which, through the principles of maximizing intra-class similarity and minimizing inter-class similarity, combined with a dynamically adjusted penalty strategy, models the correlation between video quality levels. Thirdly, a two-phase feature fusion strategy, consisting of the Gated Spatio-temporal Attention Unit (GSTU) and the Adaptive Fusion Cell (AFC), is further proposed. The former achieves spatio-temporal feature fusion through dynamic weight allocation, and the latter adaptively integrates the features from the two branches based on data characteristics. Finally, a regression module outputs the quality score. Experimental results on five real-world VQA datasets demonstrate the superior performance of the IBCL-VQA. Furthermore, the strong generalizability is verified through cross-database testing. The code and pre-trained weights will be publicly available at: https://github.com/BoHu90/IBCL-VQA . Bo Hu 0008, Leida Li, Lihuo He, Xinbo Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | CognitionCapturer: Decoding Visual Stimuli from Human EEG Signal with Multimodal InformationabstractElectroencephalogram (EEG) signals have attracted significant attention from researchers due to their non-invasive nature and high temporal sensitivity in decoding visual stimuli. However, most recent studies have focused solely on the relationship between EEG and image data pairs, neglecting the valuable "beyond-image-modality" information embedded in EEG signals. This results in the loss of critical multimodal information in EEG. To address the limitation, this paper proposes a unified framework that fully leverages multimodal data to represent EEG signals, named CognitionCapturer. Specifically, CognitionCapturer trains modality expert encoders for each modality to extract cross-modal information from the EEG modality. Then, it introduces a diffusion prior to map the EEG embedding space to the CLIP embedding space, followed by using a pretrained generative model, the proposed framework can reconstruct visual stimuli with high semantic and structural fidelity. Notably, the framework does not require any fine-tuning of the generative models and can be extended to incorporate more modalities. Through extensive experiments, we demonstrate that CognitionCapturer outperforms state-of-the-art methods both qualitatively and quantitatively. Kaifan Zhang, Lihuo He, Wen Lu 0004, Di Wang 0011, Xinbo Gao 0001 |
AAAI | 2 |
| 2025 | AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated ImagesabstractWith the advancement of AI-generated content technologies, AI-generated images (AGIs) have become increasingly influential in artistic creation and visual communication. However, the aesthetic quality of AGIs varies significantly due to technical limitations and the influence of user input, underscoring the urgent need for systematic aesthetic evaluation of AGIs. In addition, it is difficult to ensure the consistency of text-to-image, which compresses the application space of AGIs. To address these issues, a fine-grained dataset for Aesthetic and Alignment evaluation of AGIs (AGIAA-2K) is presented. This dataset contains 2,064 images generated using 172 well-designed prompts across six different AGI models, with each image annotated based on the subjective experiment. Then, the rationality of the dataset is verified by data analysis. Finally, the performances of the existing algorithms are evaluated in terms of image aesthetic assessment and text-to-image alignment of AGIs. The results demonstrate that these algorithms cannot effectively evaluate these two aspects. The AGIAA-2K is available at https://github.com/BoHu90/AGIAA-2K. Bo Hu 0008, Nanxiang Li, Lihuo He, Wen Lu 0004, Leida Li, Xinbo Gao 0001 |
ICASSP | 3 |
| 2025 | A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality AssessmentabstractWide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the perception and experience, which may seriously hinder its application in fields such as competitive sports. Up to now, few explorations focus on the quality assessment issue of wide-angle video. This deficiency primarily stems from the absence of a specialized dataset for wide-angle videos. To bridge this gap, we construct the first Multi-annotated and multi-modal Wide-angle Video quality assessment (MWV) dataset. Then, the performances of state-of-the-art video quality methods on the MWV dataset are investigated by inter-dataset testing and intra-dataset testing. Experimental results show that these methods impose significant limitations on their applicability. Bo Hu 0008, Chunyi Li 0001, Lihuo He, Leida Li, Xinbo Gao 0001 |
ICASSP | 4 |
| 2025 | A Two-Stage AIGC Image Quality Assessment with T2I Correspondence and Visual PerceptionabstractImage quality assessment (IQA) of artificial intelligence-generated content (AIGC) has recently attracted significant research attention. Unlike general-purpose IQA, which primarily focuses on evaluating image content, AIGCIQA often requires addressing both the Text-to-Image (T2I) correspondence and the perceptual quality of images. To address this requirement, this paper proposes a novel two-stage AIGCIQA method. The first stage evaluates the alignment of the AI-generated images (AIGIs) with their corresponding descriptions, serving as an indicator of overall image quality. Specifically, positive and negative prompts are constructed to describe the T2I correspondence degree, and then a CLIP model is employed to predict the degree based on these image-prompt pairs. The second stage refines the perceptual quality assessment by integrating both global and local degradation features of AIGIs. Importantly, the contribution of local features is measured according to their correlation with the overall image, ensuring key regions are adequately represented in the quality prediction. Experimental results on AGIQA-1K, AGIQA-3K, and AIGCIQA2023 demonstrate the superior performance of the proposed method. Jili Xia, Lihuo He, Bo Hu 0008, Bo Han 0004, Xinbo Gao 0001 |
ICASSP | 2 |
| 2025 | Mixture-of-Modality-Experts for Unified Image Aesthetic Assessment with Multi-Level AdaptationabstractMulti-modal image aesthetic assessment (MIAA) has gained significant progress, by predicting aesthetic based on both an image and its text comments. However, most MIAA methods are not applicable, when there are no text comments available. To combat this challenge, we propose a unified image aesthetic assessment (IAA) framework, termed AesFormer, by using mixtures of vision-language Transformers. Specially, AesFormer first learns aligned image-text representations through contrastive learning, and uses a vision-language head for MIAA prediction. Afterward, we propose a multi-level adaptation (MLA) method to adapt the learned MIAA model to the case without text comments, and use another vision head for vison-only IAA (VIAA) prediction. Extensive experimental results show that AesFormer significantly outperforms previous methods in both MIAA and VIAA tasks, on diverse benchmarking datasets. Our code has been released at: https://github.com/AiArt-Gao/AesFormer Fei Gao 0006, Xiaodan Zhang 0005, Lihuo He, Nannan Wang 0001 |
ICME | 5 |
| 2025 | MACA-VQA: Quality Assessment of UGC Videos via Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention FusionabstractUser-generated content (UGC) videos often exhibit complex distortions and diverse content, posing significant challenges for traditional video quality assessment (VQA) methods. Approaches that directly merge distortion and semantic information risk feature conflicts and the loss of details. In addition, simple concatenation of spatiotemporal features fails to capture vital interactions, limiting predictive accuracy. Motivated by these challenges, this paper proposes a Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention Fusion framework for VQA, named MACA-VQA. Specifically, a novel multi-level adaptive strategy progressively incorporates distortion information into each Transformer layer of the CLIP model, enabling layer-wise fusion of semantic and distortion features. Furthermore, a newly introduced cross-attention fusion mechanism dynamically integrates spatiotemporal features, capturing complex, multidimensional interactions. Extensive experiments demonstrate that MACA-VQA achieves state-of-the-art performance on multiple public datasets, validating its effectiveness and robustness in both intra-dataset and inter-dataset scenarios. The source code is available at https://github.com/BoHu90/MACA-VQA Bo Hu 0008, Yimeng Zhao, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
ICME | 4 |
| 2025 | Cross-Space Alignment-based Attribute Artifact Removal for V-PCCabstractIn this paper, we propose a cross-space alignment-based attribute artifact removal method for V-PCC. Firstly, we construct a cross-space frame alignment module to align the adjacent projected 2D frames based on the temporal continuity between the 3D point clouds. Then, we design a multi-frame fusion network that is developed based on U-Net to fuse the current frame with the aligned frames to remove the artifact that occurred in the attribute. Experimental results demonstrate that our proposed method achieves superior performance for the enhancement of point clouds. Yu Liu 0091, Mingfei Hu, Zeliang Li, Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Lihuo He, Fan Zhang 0017 |
ISCAS | 7 |
| 2025 | Low-light Image Enhancement Quality Assessment: A Real-World Dataset and An Objective MethodabstractLow-light Image Enhancement (LIE) technology adaptively improves brightness while preserving texture details and suppressing noise artifacts, thereby reducing visual degradation caused by insufficient illumination. While deep learning-based image enhancement algorithms have made significant progress, a key gap remains in establishing standardized methods for fairly evaluating and comparing their performance. To bridge this gap, this paper systematically investigates enhanced low-light image quality assessment from both subjective and objective dimensions. First, we introduce a Real-world Low-light Image Enhancement quality assessment dataset (RLIE), which contains 1540 images from 154 scenarios, each with a subjective score given by the subjects. Based on this, we propose a low light enhanced image quality assessment method based on Multi-level Illumination Injection and Hierarchical Discrepancy Perception (MIIHDP). The core idea of this method is to hierarchically inject separated illumination information into the feature extraction process, then tailor the processing of difference information at different scales to obtain a more comprehensive representation. Finally, extensive statistical analyses demonstrate the rationality of the proposed RLIE dataset, and experimental results show the superior performance of the proposed MIIHDP compared with state-of-the-arts. Our dataset and code are released at: https://github.com/BoHu90/RLIE. Chunyi Li 0001, Bo Hu 0008, Taiyang Chen, Leida Li, Lihuo He, Xinbo Gao 0001 |
ACM Multimedia | 5 |
| 2025 | Boosting Temporal Sentence Grounding via Causal InferenceabstractTemporal Sentence Grounding (TSG) aims to identify relevant moments in an untrimmed video that semantically correspond to a given textual query. Despite existing studies having made substantial progress, they often overlook the issue of spurious correlations between video and textual queries. These spurious correlations arise from two primary factors: (1) inherent biases in the textual data, such as frequent co-occurrences of specific verbs or phrases, and (2) the model's tendency to overfit to salient or repetitive patterns in video content. Such biases mislead the model into associating textual cues with incorrect visual moments, resulting in unreliable predictions and poor generalization to out-of-distribution examples. To overcome these limitations, we propose a novel TSG framework, causal intervention and counterfactual reasoning that utilizes causal inference to eliminate spurious correlations and enhance the model's robustness. Specifically, we first formulate the TSG task from a causal perspective with a structural causal model. Then, to address unobserved confounders reflecting textual biases toward specific verbs or phrases, a textual causal intervention is proposed, utilizing do-calculus to estimate the causal effects. Furthermore, visual counterfactual reasoning is performed by constructing a counterfactual scenario that focuses solely on video features, excluding the query and fused multi-modal features. This allows us to debias the model by isolating and removing the influence of the video from the overall effect. Experiments on public datasets demonstrate the superiority of the proposed method. The code is available at https://github.com/Tangkfan/CICR. Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao 0001 |
ACM Multimedia | 2 |
| 2025 | Blind image quality assessment for in-the-wild images by integrating distorted patch selection and multi-scale-and-granularity fusion
Jili Xia, Lihuo He, Xinbo Gao 0001, Bo Hu 0008 |
Knowl. Based Syst. | 2 |
| 2025 | Global-aware Fragment Representation Aggregation Network for image-text retrieval
Di Wang 0011, Jiabo Tian, Yumin Tian, Lihuo He |
Pattern Recognit. | 5 |
| 2025 | Blind Quality Assessment of Wide-Angle Videos Based on Deformation Representation Learning and Multi-Dimensional Feature FusionabstractWide-angle videos shot with short-focus lenses often exhibit deformation distortions, which poses significant challenges for video quality assessment (VQA). Although current VQA methods focus primarily on video content and distortion perception, there has been little explicit research on the impact of deformation characteristics on the perception of wide-angle video quality. To this end, this paper makes the first attempt to construct a novel wide-angle video quality assessment method based on deformation representation learning and multi-dimensional feature fusion, termed DRLMF. Specifically, we first analyze the deformation distribution characteristics of wide-angle videos based on the deformation camera model. Based on this, a three-stream video perception and assessment network is proposed. The first branch extracts global semantics using the image encoder of CLIP. The second branch introduces an effective deformation region selection strategy and proposes an interpretable deformation representation learning module. This module leverages the perception advantages of convolutional neural networks (CNNs) in local distortions and considers the correlation between patch size and distortion perception. The third branch extracts motion features using an action recognition network. Finally, an effective multi-dimensional feature fusion module is proposed to integrate more refined and richer semantic, deformation, and motion features. Extensive experiments on wide-angle VQA datasets and standard video datasets show that the DRLMF outperforms the state-of-the-arts in terms of prediction monotonicity and accuracy. The codes will be available at https://github.com/BoHu90/DRLMF. Bo Hu 0008, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Vision-Language Models Empowered Nighttime Object Detection With Consistency Sampler and Hallucination Feature GeneratorabstractCurrent object detectors often suffer performance degradation when applied to cross-domain scenarios, particularly under challenging visual conditions such as nighttime scenes. This is primarily due to the I3 problems: Inadequate sampling of instance-level features, Indistinguishable feature representation across domains and Inaccurate generation for identical category participation. To address these challenges, we propose a domain-adaptive detection framework that enables robust generalization across different visual domains without introducing any additional inference overhead. The framework comprises three key components. Specifically, the centerness-category consistency sampler alleviates inadequate sampling by selecting representative instance-level features, while the paired centerness consistency loss enforces alignment between classification and localization. Second, VLM-based orthogonality enhancement leverages frozen vision-language encoders with an orthogonal projection loss to improve cross-domain feature distinguishability. Third, hallucination feature generator synthesizes robust instance-level features for missing categories, ensuring balanced category participation across domains. Extensive experiments on multiple datasets covering various domain adaptation and generalization settings demonstrate that our method consistently outperforms state-of-the-art detectors, achieving up to 5.5 mAP improvement, with particularly strong gains in nighttime adaptation. Lihuo He, Junjie Ke, Jie Li 0001, Qi Wang 0009, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | AI-Generated Image Quality Assessment Based on Task-Specific Prompt and Multi-Granularity SimilarityabstractRecently, AI-generated images (AIGIs), synthesized based on initial textual prompts, have attracted widespread attention. However, due to limitations in current generation techniques, these images often exhibit degraded perceptual quality and semantic misalignment with the guiding prompts. Therefore, evaluating both perceptual quality and text-to-image alignment is essential for optimizing the performance of generative models. Existing methods design textual prompts solely based on the initial prompt for both perceptual and alignment quality tasks, and compute only coarse-grained similarity between the designed prompt and the generated image. However, such task-agnostic prompts overlook the distinctions between the perceptual and alignment quality tasks, and coarse-level similarity fails to capture semantic details, leading to suboptimal evaluation performance. To address these challenges, we propose a novel AIGI quality assessment framework, termed TPMS, which incorporates task-specific prompt and multi-granularity similarity computation. The task-specific prompt constructs dedicated prompts for perceptual and alignment quality respectively, allowing the model to capture distinct quality cues tailored to each evaluation task. Multi-granularity similarity measures the coarse-level similarity between the generated image and task-specific prompts to capture global quality characteristics, and the fine-level similarity between the generated image and the initial prompt to enhance semantic detail awareness. By integrating these two complementary similarities, TPMS enables precise and robust quality prediction. Extensive experiments on four widely-used AIGI quality benchmarks validate the effectiveness and superiority of the proposed framework. Jili Xia, Lihuo He, Cheng Deng 0002, Leida Li, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Progressive Semi-Decoupled Detector for Accurate Object DetectionabstractInconsistent accuracy between classification and localization tasks is a common challenge in modern object detection. Task decoupling, which employs distinct features or labeling strategies for each task, is a widely used approach to address this issue. Although it has led to noteworthy advancements, this approach is insufficient as it neglects task interdependence and lacks an explicit consistency constraint. To bridge this gap, this paper proposes the Progressive Semi-Decoupled Detector (ProSDD) to enhance both classification and localization accuracy. Specifically, a new detection head is designed that incorporates feature suppression and enhancement mechanism (FSEM) and bidirectional interaction module (BIM). Compared with the decoupled head, it not only filters out task-irrelevant information and enhances task-related information, but also avoids excessive decoupling at the feature level. Moreover, both FSEM and BIM are used multiple times, thus forming a progressive semi-decoupled head. Then, a novel consistency loss is proposed and integrated into the loss function of object detection, ensuring harmonic performance in classification and localization. Experimental results demonstrate that the proposed ProSDD effectively alleviates inconsistent accuracy and achieves high-quality object detection. Taking the pretrained ResNet-50 as the backbone, ProSDD achieves a remarkable 43.3 AP on the MS COCO dataset, surpassing contemporary state-of-the-art detectors by a substantial margin under the equivalent configurations. Code is available athttps://github.com/HB-X/ProSDD. Bo Han 0004, Lihuo He, Junjie Ke, Jinjian Wu, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Boosting Modal-Specific Representations for Sentiment Analysis With Incomplete ModalitiesabstractMultimodal sentiment analysis aims at exploiting complementary information from multiple modalities or data sources to enhance the understanding and interpretation of sentiment. While existing multi-modal fusion techniques offer significant improvements in sentiment analysis, real-world scenarios often involve missing modalities, introducing complexity due to uncertainty of which modalities may be absent. To tackle the challenge of incomplete modality-specific feature extraction caused by missing modalities, this paper proposes a Cosine Margin-Aware Network (CMANet) which centers on the Cosine Margin-Aware Distillation (CMAD) module. The core module measures distance between samples and the classification boundary, enabling CMANet to focus on samples near the boundary. So, it effectively captures the unique features of different modal combinations. To address the issue of modality imbalance during modality-specific feature extraction, this paper proposes a Weak Modality Regularization (WMR) strategy, which aligns the feature distributions between strong and weak modalities at the dataset-level, while also enhancing the prediction loss of samples at the sample-level. This dual mechanism improves the recognition robustness of weak modality combination. Extensive experiments demonstrate that the proposed method outperforms the previous best model, MMIN, with a 3.82% improvement in unweighted accuracy. These results underscore the robustness of the approach under conditions of uncertain and missing modalities. Lihuo He, Fei Gao 0006, Kaifan Zhang, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Dual Semantic Reconstruction Network for Weakly Supervised Temporal Sentence GroundingabstractWeakly supervised temporal sentence grounding aims to identify semantically relevant video moments in an untrimmed video corresponding to a given sentence query without exact timestamps. Neuropsychology research indicates that the way the human brain handles information varies based on the grammatical categories of words, highlighting the importance of separately considering nouns and verbs. However, current methodologies primarily utilize pre-extracted video features to reconstruct randomly masked queries, neglecting the distinction between grammatical classes. This oversight could hinder forming meaningful connections between linguistic elements and the corresponding components in the video. To address this limitation, this paper introduces the dual semantic reconstruction network (DSRN) model. DSRN processes video features by distinctly correlating object features with nouns and motion features with verbs, thereby mimicking the human brain's parsing mechanism. It begins with a feature disentanglement module that separately extracts object-aware and motion-aware features from video content. Then, in a dual-branch structure, these disentangled features are used to generate separate proposals for objects and motions through two dedicated proposal generation modules. A consistency constraint is proposed to ensure a high level of agreement between the boundaries of object-related and motion-related proposals. Subsequently, the DSRN independently reconstructs masked nouns and verbs from the sentence queries using the generated proposals. Finally, an integration block is applied to synthesize the two types of proposals, distinguishing between positive and negative instances through contrastive learning. Experiments on the Charades-STA and ActivityNet Captions datasets demonstrate that the proposed method achieves state-of-the-art performance. Kefan Tang, Lihuo He, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Visual-Language Multi-Task Blind Image Quality Assessment With Local Quality WeightingabstractThe objective of blind image quality assessment (BIQA) is to develop a model capable of automatically evaluating image quality without requiring any reference knowledge. While multi-task learning has been widely utilized in BIQA, it has predominantly remained unimodal. This paper delves into the Visual-Language multi-task BIQA model, where distortion knowledge can be captured through image-text contrastive learning. Specifically, Visual-Language auxiliary tasks targeting distortion type and quality level are introduced, respectively, where both positive and negative image-text pairs are constructed for the target distorted image. Subsequently, image-text correspondences are learned in the embedding space while simultaneously evaluating image quality. Notably, in the auxiliary task learning, the proposed method not only brings the image and its corresponding positive text prompt closer but also pushes away the image from its negative text prompts, thereby facilitating the extraction of pertinent distortion features. In the quality assessment task, a patch-wise strategy is employed during the training phase. Differing from conventional BIQA methods, a novel NSS-guided quality weighting is introduced to gauge the correlation between patch quality and global quality, thereby enabling precise quality prediction. Extensive experiments are conducted on six IQA datasets, and the experimental results verify the superiority of the proposed method. Jili Xia, Lihuo He, Bo Hu 0008, Leida Li, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Multi-object behavior recognition based on object detection for dense crowds
Min Dang, Gang Liu 0006, Qijie Xu, Ke Li 0024, Di Wang 0011, Lihuo He |
Expert Syst. Appl. | 6 |
| 2024 | Weighted parallel decoupled feature pyramid network for object detection
Bo Han 0004, Lihuo He, Junjie Ke, Chenwei Tang, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2024 | Blind image quality assessment based on hierarchical dependency learning and quality aggregationabstractImage quality assessment (IQA) aims to build a quality prediction model to assess image quality automatically rather than artificially. Due to a lack of reference images, blind image quality assessment (BIQA) has become an attractive yet challenging research topic. Inspired by the hierarchical perception mechanism in the human visual system , some existing BIQA methods aggregate multi-stage features of a convolutional neural network (CNN). However, they are regardless of the latent dependencies. To solve this problem, we propose a novel BIQA method based on hierarchical dependency learning and quality aggregation (HDLaQA). The proposed method includes multi-stage feature extraction, hierarchical dependency learning, and quality aggregation. In multi-stage feature extraction, a CNN is used as the feature extractor and multi-stage features are output for further learning. In hierarchical dependency learning, spatial and channel dependencies among the multi-stage features are modeled. To this end, a dual-head spatial dependency (DSD) module is designed to harvest the spatial dependencies between the adjacent-stage features and deliver these dependencies to the next stage. Moreover, exponential bilinear pooling (EBP) is presented to learn the channel dependencies, which is more stable than commonly used BP. In quality aggregation, multiple quality scores are predicted based on the learned dependencies, and multiple learnable weights are used to measure the importance of the predicted scores for final quality evaluation. Experimental results on seven IQA databases demonstrate the competitiveness of the proposed method on both synthetic and authentic distortions. Jili Xia, Lihuo He, Xinbo Gao 0001, Bo Hu 0008 |
Neurocomputing | 2 |
| 2024 | ProFPN: Progressive feature pyramid network with soft proposal assignment for object detection
Junjie Ke, Lihuo He, Bo Han 0004, Jie Li 0001, Xinbo Gao 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Improving generalized zero-shot learning via cluster-based semantic disentangling representation
Wentao Feng, Rong Xiao 0001, Lihuo He, Zhenan He 0001, Jiancheng Lv 0001, Chenwei Tang |
Pattern Recognit. | 4 |
| 2024 | General Deformable RoI Pooling and Semi-Decoupled Head for Object DetectionabstractObject detection aims to classify interest objects within an image and pinpoint their positions using predicted rectangular bounding boxes. However, classification and localization tasks are heterogeneous, not only spatially misaligned but also differing in properties and feature requirements. Modern detectors commonly share the spatial region and detection head for both tasks, making them challenging to achieve optimal performance altogether, resulting in inconsistent accuracy. Specifically, the predicted bounding box may have higher classification confidence but lower localization quality, or vice versa. To tackle this issue, the spatial decoupling mechanism via general deformable RoI pooling is first proposed. This mechanism separately pursues the favorable regions for classification and localization, and subsequently extracts the corresponding features. Then, the semi-decoupled head is designed. Compared to the decoupled head that utilizes independent classification and localization networks, potentially leading to excessive decoupling and compromised detection performance, the semi-decoupled head enables the networks to mutually enhance each other while concentrating on their respective tasks. In addition, the semi-decoupled head also introduces a redundancy suppression module to filter out redundant task-irrelevant information of features extracted by separate networks and reinforce task-related information. By combining the spatial decoupling mechanism with the semi-decoupled head, the proposed detector achieves an impressive 43.7 AP in Faster R-CNN framework with ResNet-101 as backbone network. Without bells and whistles, extensive experimental results on the popular MS COCO dataset demonstrate that the proposed detector suppresses the baseline by a significant margin and outperforms some state-of-the-art detectors. Code is available athttps://github.com/HB-X/gdpool_semi_dehead. Bo Han 0004, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | VLDadaptor: Domain Adaptive Object Detection With Vision-Language Model DistillationabstractDomain adaptive object detection (DAOD) aims to develop a detector trained on labeled source domains to identify objects in unlabeled target domains. A primary challenge in DAOD is the domain shift problem. Most existing methods learn domain-invariant features within single domain embedding space, often resulting in heavy model biases due to the intrinsic data properties of source domains. To mitigate the model biases, this paper proposes VLDadaptor, a domain adaptive object detector based on vision-language models (VLMs) distillation. Firstly, the proposed method integrates domain-mixed contrastive knowledge distillation between the visual encoder of CLIP and the detector by transferring category-level instance features, which guarantees the detector can extract domain-invariant visual instance features across domains. Then, VLDadaptor employs domain-mixed consistency distillation between the text encoder of CLIP and detector by aligning text prompt embeddings with visual instance features, which helps to maintain the category-level feature consistency among the detector, text encoder and the visual encoder of VLMs. Finally, the proposed method further promotes the adaptation ability by adopting a prompt-based memory bank to generate semantic-complete features for graph matching. These contributions enable VLDadaptor to extract visual features into the visual-language embedding space without any evident model bias towards specific domains. Extensive experimental results demonstrate that the proposed method achieves state-of-the-art performance on Pascal VOC to Clipart adaptation tasks and exhibits high accuracy on driving scenario tasks with significantly less training time. Junjie Ke, Lihuo He, Bo Han 0004, Jie Li 0001, Di Wang 0011, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Gist, Content, Target-Oriented: A 3-Level Human-Like Framework for Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to locate corresponding moments in an untrimmed video via a given natural language query. While most existing approaches treat this task as a cross-modal content matching or boundary prediction problem, recent studies have started to solve the VMR problem from a reading comprehension perspective. However, the cross-modal interaction processes of existing models are either insufficient or overly complex. Therefore, we reanalyze human behaviors in the document fragment location task of reading comprehension, and design a specific module for each behavior to propose a 3-level human-like moment retrieval framework (Tri-MRF). Specifically, we summarize human behaviors such as grasping the general structures of the document and the question separately, cross-scanning to mark the direct correspondences between keywords in the document and in the question, and summarizing to obtain the overall correspondences between document fragments and the question. Correspondingly, the proposed Tri-MRF model contains three modules: 1) a gist-oriented intra-modal comprehension module is used to establish contextual dependencies within each modality; 2) a content-oriented fine-grained comprehension module is used to explore direct correspondences between clips and words; and 3) a target-oriented integrated comprehension module is used to verify the overall correspondence between the candidate moments and the query. In addition, we introduce a biconnected GCN feature enhancement module to optimize query-guided moment representations. Extensive experiments conducted on three benchmarks, TACoS, ActivityNet Captions and Charades-STA demonstrate that the proposed framework outperforms State-of-the-Art methods. Di Wang 0011, Xiantao Lu, Quan Wang 0006, Yumin Tian, Bo Wan 0002, Lihuo He |
IEEE Trans. Multim. | 6 |
| 2024 | Dual-Perspective Fusion Network for Aspect-Based Multimodal Sentiment AnalysisabstractAspect-based multimodal sentiment analysis (ABMSA) is an important sentiment analysis task that analyses aspect-specific sentiment in data with different modalities (usually multimodal data with text and images). Previous works usually ignore the overall sentiment tendency when analyzing the sentiment of each aspect term. However, the overall sentiment tendency is highly correlated with aspect-specific sentiment. In addition, existing methods neglect to explore and make full use of the fine-grained multimodal information closely related to aspect terms. To address these limitations, we propose a dual-perspective fusion network (DPFN) that considers both global and local fine-grained sentiment information in multimodal data. From the global perspective, we use text-image caption pairs to obtain a global representation containing information about the overall sentiment tendencies. From the local fine-grained perspective, we construct two graph structures to explore the fine-grained information in texts and images. Finally, aspect-level sentiment polarities can be obtained by analyzing the combination of global and local fine-grained sentiment information. Experimental results on two multimodal Twitter datasets show that the proposed DPFN model outperforms state-of-the-art methods. Di Wang 0011, Changning Tian, Lin Zhao 0003, Lihuo He, Quan Wang 0006 |
IEEE Trans. Multim. | 5 |
| 2024 | Deep Hierarchical Multimodal Metric LearningabstractMultimodal metric learning aims to transform heterogeneous data into a common subspace where cross-modal similarity computing can be directly performed and has received much attention in recent years. Typically, the existing methods are designed for nonhierarchical labeled data. Such methods fail to exploit the intercategory correlations in the label hierarchy and, therefore, cannot achieve optimal performance on hierarchical labeled data. To address this problem, we propose a novel metric learning method for hierarchical labeled multimodal data, named deep hierarchical multimodal metric learning (DHMML). It learns the multilayer representations for each modality by establishing a layer-specific network corresponding to each layer in the label hierarchy. In particular, a multilayer classification mechanism is introduced to enable the layerwise representations to not only preserve the semantic similarities within each layer, but also retain the intercategory correlations across different layers. In addition, an adversarial learning mechanism is proposed to bridge the cross-modality gap by producing indistinguishable features for different modalities. Through integration of the multilayer classification and adversarial learning mechanisms, DHMML can obtain hierarchical discriminative modality-invariant representations for multimodal data. Experiments on two benchmark datasets are used to demonstrate the superiority of the proposed DHMML method over several state-of-the-art methods. Di Wang 0011, Aqiang Ding, Yumin Tian, Quan Wang 0006, Lihuo He, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Language-Guided Visual Aggregation Network for Video Question AnsweringabstractVideo Question Answering (VideoQA) aims to comprehend intricate relationships, actions, and events within video content, as well as the inherent links between objects and scenes, to answer text-based questions accurately. Transferring knowledge from the cross-modal pre-trained model CLIP is a natural approach, but its dual-tower structure hinders fine-grained modality interaction, posing challenges for direct application to VideoQA tasks. To address this issue, we introduce a Language-Guided Visual Aggregation (LGVA) network. It employs CLIP as an effective feature extractor to obtain language-aligned visual features with different granularities and avoids resource-intensive video pre-training. The LGVA network progressively aggregates visual information in a bottom-up manner, focusing on both regional and temporal levels, and ultimately facilitating accurate answer prediction. More specifically, it employs local cross-attention to combine pre-extracted question tokens and region embeddings, pinpointing the object of interest in the question. Then, graph attention is utilized to aggregate regions at the frame level and integrate additional captions for enhanced detail. Following this, global cross-attention is used to merge sentence and frame-level embeddings, identifying the video segment relevant to the question. Ultimately, contrastive learning is applied to optimize the similarities between aggregated visual and answer embeddings, unifying upstream and downstream tasks. Our method conserves resources by avoiding large-scale video pre-training and simultaneously demonstrates commendable performance on the NExT-QA, MSVD-QA, MSRVTT-QA, TGIF-QA, and ActivityNet-QA datasets, even outperforming some end-to-end trained models. Our code is available at https://github.com/ecoxial2007/LGVA_VideoQA. Di Wang 0011, Quan Wang 0006, Bo Wan 0002, Lingling An, Lihuo He |
ACM Multimedia | 6 |
| 2023 | TETFN: A text enhanced transformer fusion network for multimodal sentiment analysis
Di Wang 0011, Xutong Guo, Yumin Tian, Lihuo He, Xuemei Luo |
Pattern Recognit. | 5 |
| 2023 | Cross-Modal Enhancement Network for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) plays an important role in many applications, such as intelligent question-answering, computer-assisted psychotherapy and video understanding, and has attracted considerable attention in recent years. It leverages multimodal signals including verbal language, facial gestures, and acoustic behaviors to identify sentiments in videos. Language modality typically outperforms nonverbal modalities in MSA. Therefore, strengthening the significance of language in MSA will be a vital way to promote recognition accuracy. Considering that the meaning of a sentence often varies in different nonverbal contexts, combining nonverbal information with text representations is conducive to understanding the exact emotion conveyed by an utterance. In this paper, we propose a Cross-modal Enhancement Network (CENet) model to enhance text representations by integrating visual and acoustic information into a language model. Specifically, it embeds a Cross-modal Enhancement (CE) module, which enhances each word representation according to long-range emotional cues implied in unaligned nonverbal data, into a transformer-based pre-trained language model. Moreover, a feature transformation strategy is introduced for acoustic and visual modalities to reduce the distribution differences between the initial representations of verbal and nonverbal modalities, thereby facilitating the fusion of distinct modalities. Extensive experiments on benchmark datasets demonstrate the significant gains of CENet over state-of-the-art methods. Di Wang 0011, Shuai Liu 0009, Quan Wang 0006, Yumin Tian, Lihuo He, Xinbo Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Hierarchical Semantic Structure Preserving Hashing for Cross-Modal RetrievalabstractCross-modal hashing has become a vital technique in cross-modal retrieval due to its fast query speed and low storage cost in recent years. Generally, most of the priors supervised cross-modal hashing methods are flat methods which are designed for non-hierarchical labeled data. They treat different categories independently and ignore the inter-category correlations. In practical applications, many instances are labeled with hierarchical categories. The hierarchical label structure provides rich information among different categories. To rationally take use of category correlations, hierarchical cross-modal hashing is proposed. However, existing methods intend to preserve instance-pairwise or class-pairwise similarities, which cannot fully explore the semantic correlations among different categories and make the learned hash codes less discriminative. In this paper, we propose a deep cross-modal hashing method named hierarchical semantic structure preserving hashing (HSSPH), which directly exploits the label hierarchy information to learn discriminative hash codes. Specifically, HSSPH learns a set of class-wise hash codes for each layer. By augmenting class-wise codes with labels, it generates layer-wise prototype codes which reflect the semantic structure of each layer. In order to enhance the discriminative ability of hash codes, HSSPH supervises the hash codes learning with both labels and semantic structures to preserve the hierarchical semantics. Besides, efficient optimization algorithms are developed to directly learn the discrete hash codes for each instance and each class. Extensive experiments on two benchmark datasets show the superiority of HSSPH over several state-of-the-art methods. Di Wang 0011, Caiping Zhang, Quan Wang 0006, Yumin Tian, Lihuo He, Lin Zhao 0003 |
IEEE Trans. Multim. | 5 |
| 2023 | Deep Hybrid 2-D-3-D CNN Based on Dual Second-Order Attention With Camera Spectral Sensitivity Prior for Spectral Super-ResolutionabstractA largely ignored fact in spectral super-resolution (SSR) is that the subsistent mapping methods neglect the auxiliary prior of camera spectral sensitivity (CSS) and only pay attention to wider or deeper network framework design while ignoring to excavate the spatial and spectral dependencies among intermediate layers, hence constraining representational capability of convolutional neural networks (CNNs). To conquer these drawbacks, we propose a novel deep hybrid 2-D-3-D CNN based on dual second-order attention with CSS prior (HSACS), which can excavate sufficient spatial-spectral context information. Specifically, dual second-order attention embedded in the residual block for more powerful spatial-spectral feature representation and relation learning is composed of a brand new trainable 2-D second-order channel attention (SCA) or 3-D second-order band attention (SBA) and a structure tensor attention (STA). Concretely, the band and channel attention modules are developed to adaptively recalibrate the band-wise and interchannel features via employing second-order band or channel feature statistics for more discriminative representations. Besides, the STA is promoted to rebuild the significant high-frequency spatial details for enough spatial feature extraction. Moreover, the CSS is first employed as a superior prior to avoid its effect of SSR quality, on the strength of which the resolved RGB can be calculated naturally through the super-reconstructed hyperspectral image (HSI); then, the final loss consists of the discrepancies of RGB and the HSI as a finer constraint. Experimental results demonstrate the superiority and progressiveness of the presented approach in terms of quantitative metrics and visual effect over SOTA SSR methods. Jiaojiao Li 0001, Chaoxiong Wu, Rui Song 0003, Yunsong Li 0001, Weiying Xie, Lihuo He, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Feature Erasing and Diffusion Network for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) aims at matching occluded person images to holistic ones across different camera views. Target Pedestrians (TP) are often disturbed by Non-Pedestrian Occlusions (NPO) and Non-Target Pedestrians (NTP). Previous methods mainly focus on increasing the model's robustness against NPO while ignoring feature contamination from NTP. In this paper, we propose a novel Feature Erasing and Diffusion Network (FED) to simultaneously handle challenges from NPO and NTP. Specifically, aided by the NPO augmentation strategy that simulates NPO on holistic pedestrian images and gen-erates precise occlusion masks, NPO features are explicitly eliminated by our proposed Occlusion Erasing Module (OEM). Subsequently, we diffuse the pedestrian representations with other memorized features to synthesize the NTP characteristics in the feature space through the novel Feature Diffusion Module (FDM). With the guidance of the occlusion scores from OEM, the feature diffusion process is conducted on visible body parts, thereby improving the quality of the synthesized NTP characteristics. We can greatly improve the model's perception ability towards TP and alleviate the influence of NPO and NTP by jointly optimizing OEM and FDM. Furthermore, the proposed FDM works as an auxiliary module for training and will not be engaged in the inference phase, thus with high flexibility. Experiments on occluded and holistic person ReID benchmarks demonstrate the superiority of FED over state-of-the-art methods. Zhikang Wang, Feng Zhu 0006, Shixiang Tang, Rui Zhao 0001, Lihuo He, Jiangning Song |
CVPR | 5 |
| 2022 | Robust Video-Based Person Re-Identification by Hierarchical MiningabstractVideo-based person re-identification (Re-ID) aims at retrieving the person through the video sequences across non-overlapping cameras. Some characteristics of pedestrians are not consecutive across frames due to the variations of viewpoints, postures, and occlusions over time. However, existing methods ignore such data peculiarity and the networks tend to only learn those salient consecutive characteristics among frames in video sequences. As a result, the learned representations fail to cover all the characteristics of pedestrians, thus lacking integrity and discrimination. To tackle this problem, we present a novel deep architecture termed Hierarchical Mining Network (HMN), which mines as many pedestrians’ characteristics by referring to the temporal and intra-class knowledge. It consists of a novel Attentive Temporal Module (ATM) and a Dynamic Supervising Branch (DSB), with a Balancing Triplet Loss (BTL) assisting the training. The proposed ATM, with pedestrian perceiving capacity, is capable of evaluating each activation of features through temporal analysis, so that the temporally scattered characteristics of pedestrians can be better aggregated and the contaminated ones can be eliminated. Then, the DSB along with the BTL further enhances the integrity of representations by multiple supervision. Specifically, the DSB perceives the diversities of intra-class samples in each mini-batch and generates targeted supervising signals for them, in which process the BTL guarantees the signals with smaller intra-class variations and larger inter-class variations. Comprehensive experiments on two video-based datasets, i.e., MARS, and DukeMTMC-VideoReID, demonstrate the contribution of each component and the superiority of the proposed HMN over the state-of-the-arts. Benchmarking our model on three popular image-based datasets, i.e., Market1501, DukeMTMC-Reid, and MSMT17 additionally verifies the promising generalizability of the proposed DSB and BTL. Zhikang Wang, Lihuo He, Xiaoguang Tu, Jian Zhao 0006, Xinbo Gao 0001, Shengmei Shen, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Pseudo-Label Guided Collective Matrix Factorization for Multiview ClusteringabstractMultiview clustering has aroused increasing attention in recent years since real-world data are always comprised of multiple features or views. Despite the existing clustering methods having achieved promising performance, there still remain some challenges to be solved: 1) most existing methods are unscalable to large-scale datasets due to the high computational burden of eigendecomposition or graph construction and 2) most methods learn latent representations and cluster structures separately. Such a two-step learning scheme neglects the correlation between the two learning stages and may obtain a suboptimal clustering result. To address these challenges, a pseudo-label guided collective matrix factorization (PLCMF) method that jointly learns latent representations and cluster structures is proposed in this article. The proposed PLCMF first performs clustering on each view separately to obtain pseudo-labels that reflect the intraview similarities of each view. Then, it adds a pseudo-label constraint on collective matrix factorization to learn unified latent representations, which preserve the intraview and interview similarities simultaneously. Finally, it intuitively incorporates latent representation learning and cluster structure learning into a joint framework to directly obtain clustering results. Besides, the weight of each view is learned adaptively according to data distribution in the joint framework. In particular, the joint learning problem can be solved with an efficient iterative updating method with linear complexity. Extensive experiments on six benchmark datasets indicate the superiority of the proposed method over state-of-the-art multiview clustering methods in both clustering accuracy and computational efficiency. Di Wang 0011, Songwei Han, Quan Wang 0006, Lihuo He, Yumin Tian, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | MSCAN: Multimodal Self-and-Collaborative Attention Network for image aesthetic prediction tasks
Xiaodan Zhang 0005, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
Neurocomputing | 3 |
| 2021 | Video quality assessment with dense features and ranking pooling
Yu Zhang 0062, Lihuo He, Wen Lu 0004, Jie Li 0001, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2021 | Beyond Vision: A Multimodal Recurrent Attention Convolutional Neural Network for Unified Image Aesthetic Prediction TasksabstractOver the past few years, image aesthetic prediction has attracted increasing attention because of its wide applications, such as image retrieval, photo album management and aesthetic-driven image enhancement. However, previous studies in this area only achieve limited success because 1) they primarily depend on visual features and ignore textual information. 2) they tend to focus equally on to each part of images and ignore the selective attention mechanism. This paper overcomes these limitations by proposing a novel multimodal recurrent attention convolutional neural network (MRACNN). More specifically, the MRACNN consists of two streams: the vision stream and the language stream. The former employs the recurrent attention network to tune out irrelevant information and focuses on some key regions to extract visual features. The latter utilizes the Text-CNN to capture the high-level semantics of user comments. Finally, a multimodal factorized bilinear (MFB) pooling approach is used to achieve effective fusion of textual and visual features. Extensive experiments demonstrate that the proposed MRACNN significantly outperforms state-of-the-art methods for unified aesthetic prediction tasks: (i) aesthetic quality classification; (ii) aesthetic score regression; and (iii) aesthetic score distribution prediction. Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Jie Li 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Learning More Accurate Features for Semantic Segmentation in CycleNet
Linzi Qu, Lihuo He, Junji Ke, Xinbo Gao 0001, Wen Lu 0004 |
ACCV (1) | 2 |
| 2020 | Joint and individual matrix factorization hashing for large-scale cross-modal retrieval
Di Wang 0011, Quan Wang 0006, Lihuo He, Xinbo Gao 0001, Yumin Tian |
Pattern Recognit. | 3 |
| 2020 | Deep multi-label learning for image distortion identification
Xinbo Gao 0001, Wen Lu 0004, Lihuo He |
Signal Process. | 4 |
| 2020 | Objective Video Quality Assessment Combining Transfer Learning With CNNabstractNowadays, video quality assessment (VQA) is essential to video compression technology applied to video transmission and storage. However, small-scale video quality databases with imbalanced samples and low-level feature representations for distorted videos impede the development of VQA methods. In this paper, we propose a full-reference (FR) VQA metric integrating transfer learning with a convolutional neural network (CNN). First, we imitate the feature-based transfer learning framework to transfer the distorted images as the related domain, which enriches the distorted samples. Second, to extract high-level spatiotemporal features of the distorted videos, a six-layer CNN with the acknowledged learning ability is pretrained and finetuned by the common features of the distorted image blocks (IBs) and video blocks (VBs), respectively. Notably, the labels of the distorted IBs and VBs are predicted by the classic FR metrics. Finally, based on saliency maps and the entropy function, we conduct a pooling stage to obtain the quality scores of the distorted videos by weighting the block-level scores predicted by the trained CNN. In particular, we introduce a preprocessing and a postprocessing to reduce the impact of inaccurate labels predicted by the FR-VQA metric. Due to feature learning in the proposed framework, two kinds of experimental schemes including train-test iterative procedures on one database and tests on one database with training other databases are carried out. The experimental results demonstrate that the proposed method has high expansibility and is on a par with some state-of-the-art VQA metrics on two widely used VQA databases with various compression distortions. Yu Zhang 0062, Xinbo Gao 0001, Lihuo He, Wen Lu 0004, Ran He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Multi-scale Spatial-temporal Network for Person Re-identificationabstractVideo-based person re-identification (ReID) is an important task, which has received much attention in recent years due to its efficiency in the field of surveillance. Researchers have employed many effective approaches for video-based person ReID, but there are still two problems. Firstly, the same pedestrian in the video sequences differs in size. Secondly, traditional RNNs can only process one-dimension features, which are not suitable for dealing with video sequences. To solve above problems, we propose a new network called Multi-scale Spatial-Temporal Network (MSTN), which combines multi-scale feature extractor and CLSTM together to tackle the discrepant sizes of pedestrians and extract more representative temporal information for the video sequences. We conduct the experiments on the iLIDS-VID, PRID-2011 and MARS datasets, and our approach outperforms state-of-the-art methods by a large margin. Zhikang Wang, Lihuo He, Xinbo Gao 0001, Yuanfei Huang |
ICASSP | 2 |
| 2019 | Beauty Aware Network: An Unsupervised Method for Makeup Product RetrievalabstractMakeup product retrieval has gained more and more attention for its wide application prospects. However, the challenging problem is that the dataset crawled from Internet doesn't have annotated labels. Therefore, existing methods are unable to obtain well-trained networks. To solve this problem, this paper proposes a trainable network named Beauty Aware Network (BAN) for makeup product retrieval. The core of proposed method is using an unsupervised cluster method to train the beauty classification network. And then a covariance pooling layer is introduced to leverage the statistical information. Finally, a multi-layer fusion strategy is used to capture informative clues in images. The proposed method can get simpler but more efficient features for beauty product retrieval with less computation cost. The experiments conduct on Perfect-500k dataset which has more than half-million images. The results demonstrate the effectiveness of beauty aware network by competitive performance. Linzi Qu, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
ACM Multimedia | 3 |
| 2019 | Label Consistent Matrix Factorization Hashing for Large-Scale Cross-Modal Similarity SearchabstractMultimodal hashing has attracted much interest for cross-modal similarity search on large-scale multimedia data sets because of its efficiency and effectiveness. Recently, supervised multimodal hashing, which tries to preserve the semantic information obtained from the labels of training data, has received considerable attention for its higher search accuracy compared with unsupervised multimodal hashing. Although these algorithms are promising, they are mainly designed to preserve pairwise similarities. When semantic labels of training data are given, the algorithms often transform the labels into pairwise similarities, which gives rise to the following problems: (1) constructing pairwise similarity matrix requires enormous storage space and a large amount of calculation, making these methods unscalable to large-scale data sets; (2) transforming labels into pairwise similarities loses the category information of the training data. Therefore, these methods do not enable the hash codes to preserve the discriminative information reflected by labels and, hence, the retrieval accuracies of these methods are affected. To address these challenges, this paper introduces a simple yet effective supervised multimodal hashing method, called label consistent matrix factorization hashing (LCMFH), which focuses on directly utilizing semantic labels to guide the hashing learning procedure. Considering that relevant data from different modalities have semantic correlations, LCMFH transforms heterogeneous data into latent semantic spaces in which multimodal data from the same category share the same representation. Therefore, hash codes quantified by the obtained representations are consistent with the semantic labels of the original data and, thus, can have more discriminative power for cross-modal similarity search tasks. Thorough experiments on standard databases show that the proposed algorithm outperforms several state-of-the-art methods. Di Wang 0011, Xinbo Gao 0001, Xiumei Wang 0002, Lihuo He |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Fusion global and local deep representations with neural attention for aesthetic quality assessmentabstractIn recent years, deep-learning based aesthetics assessment methods have shown promising results. However, existing methods can only achieve limited success because 1) most of the methods take one fixed-size patch as the training example, which loses the fine grained details and the holistic layout information, and 2) most of the methods ignore ordinal issues in image aesthetic assessment, i.e. image scored 5.3 is more likely to be in the high quality class than image scored 4.5. To address these challenges, we presents a novel convolutional networks with two branches to encode global and local features . The first branch not only captures the spatial layout information but also feedbacks the top-down neural attention. The second branch selects the important attended region to extract the fine details features. A sobel-based attention layer is integrated with the second branch to enhance fine details encoding. Regarding the second problem, we combine the strength of classification approach and regression approach by a multi-task learning framework. Extensive experiments on challenging Aesthetic and Visual Analysis (AVA) dataset and Photo.net dataset indicate the effectiveness of the proposed method. Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He |
Signal Process. Image Commun. | 5 |
| 2019 | Blind Video Quality Assessment With Weakly Supervised Learning and Resampling StrategyabstractDue to the 3D spatiotemporal regularities of natural videos and small-scale video quality databases, effective objective video quality assessment (VQA) metrics are difficult to obtain but highly desirable. In this paper, we propose a general-purpose no-reference VQA framework that is based on weakly supervised learning with a convolutional neural network (CNN) and a resampling strategy. First, an eight-layer CNN is trained by weakly supervised learning to construct the relationship between the deformations of the 3D discrete cosine transform of video blocks and the corresponding weak labels judged by a full-reference (FR) VQA metric. Thus, the CNN obtains the quality assessment capacity converted from the FR-VQA metric, and the effective features of the distorted videos can be extracted through the trained network. Then, we map the frequency histogram calculated from the quality score vectors predicted by the trained network onto the perceptual quality. Especially, to improve the performance of the mapping function, we transfer the frequency histogram of the distorted images and videos to resample the training set. The experiments are carried out on several widely used VQA databases. The experimental results demonstrate that the proposed method is on a par with some state-of-the-art VQA metrics and has promising robustness. Yu Zhang 0062, Xinbo Gao 0001, Lihuo He, Wen Lu 0004, Ran He 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A Gated Peripheral-Foveal Convolutional Neural Network for Unified Image Aesthetic PredictionabstractLearning fine-grained details is a key issue in image aesthetic assessment. Most of the previous methods extract the fine-grained details via random cropping strategy, which may undermine the integrity of semantic information. Extensive studies show that humans perceive fine-grained details with a mixture of foveal vision and peripheral vision. Fovea has the highest possible visual acuity and is responsible for seeing the details. The peripheral vision is used for perceiving the broad spatial scene and selecting the attended regions for the fovea. Inspired by these observations, we propose a gated peripheral-foveal convolutional neural network. It is a dedicated double-subnet neural network (i.e., a peripheral subnet and a foveal subnet). The former aims to mimic the functions of peripheral vision to encode the holistic information and provide the attended regions. The latter aims to extract fine-grained features on these key regions. Considering that the peripheral vision and foveal vision play different roles in processing different visual stimuli, we further employ a gated information fusion network to weigh their contributions. The weights are determined through the fully connected layers followed by a sigmoid function. We conduct comprehensive experiments on the standard Aesthetic Visual Analysis (AVA) dataset and Photo.net dataset for unified aesthetic prediction tasks: 1) aesthetic quality classification; 2) aesthetic score regression; and 3) aesthetic score distribution prediction. The experimental results demonstrate the effectiveness of the proposed method. Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He |
IEEE Trans. Multim. | 4 |
| 2018 | Blind Image Quality Assessment Based on Visuo-Spatial Series StatisticsabstractExisting blind image quality assessment (BIQA) methods based on statistics attach limited attention to the relative position of pixels. Features in these BIQA methods are too flimsy to characterize quite a few distortions with strong locality or complexity. However, psychological studies have shown that according to the relative position within visual field, the cognitive system generates visuo-spatial serial memory used for cognitive tasks, e.g., subjective image quality assessment. Inspired by the visuo-spatial series generated by human visual system (HVS), we propose a BIQA method based on imitation Visuo-spatial Series Statistics (VSS). The proposed method simulates visual system to construct visuo-spatial series based on the relative position of pixels, and use statistical features of visuo-spatial series to predict image quality. Extensive experiments demonstrate the proposed method has a superior performance compared to the state-of-the-art BIQA methods. Ziheng Zhou 0004, Wen Lu 0004, Lihuo He, Xinbo Gao 0001 |
ICASSP | 3 |
| 2018 | Single Image Super Resolution Based on Deep Residual Network via Lateral ModulesabstractRecently, convolutional neural networks have demonstrated high-quality reconstruction for single image super resolution (SISR). In this paper, we propose a Deep Residual Network via lateral modules (DRNLM), DRNLM is the structure with lateral modules, progressive and symmetric residual blocks (convolutional residual blocks and deconvolutional residual blocks). First, DRNLM introduces lateral modules, which are used to transmit low-level features (coarse residue) into high-level features (fine residue) effectively, thus finer residue can be obtained for better image reconstruction. Second, considering more channels can stack more details, progressive channels that vary from 64 to 256 are utilized in DRNLM through residual blocks. Third, symmetric residual blocks have same dimensions of input and output, which can ensure the gradient ranging within certain limits when the network goes deeper. Extensive experiments demonstrate that the proposed method outperforms the existing methods in accuracy and visual impression. Rui Wang 0173, Wen Lu 0004, Yuanfei Huang, Xinbo Gao 0001, Lihuo He |
ICIP | 6 |
| 2018 | Spatiotemporal Masking for Objective Video Quality Assessment
Ran He 0002, Wen Lu 0004, Yu Zhang 0062, Xinbo Gao 0001, Lihuo He |
PRCV (1) | 5 |
| 2018 | Dominant vanishing point detection in the wild with application in composition analysis
Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Qi Liu 0054 |
Neurocomputing | 4 |
| 2018 | Single Image Super-Resolution via Multiple Mixture Prior ModelsabstractExample learning-based single image super-resolution (SR) is a promising method for reconstructing a high-resolution (HR) image from a single-input low-resolution (LR) image. Lots of popular SR approaches are more likely either time-or space-intensive, which limit their practical applications. Hence, some research has focused on a subspace view and delivered state-of-the-art results. In this paper, we utilize an effective way with mixture prior models to transform the large nonlinear feature space of LR images into a group of linear subspaces in the training phase. In particular, we first partition image patches into several groups by a novel selective patch processing method based on difference curvature of LR patches, and then learning the mixture prior models in each group. Moreover, different prior distributions have various effectiveness in SR, and in this case, we find that student-t prior shows stronger performance than the well-known Gaussian prior. In the testing phase, we adopt the learned multiple mixture prior models to map the input LR features into the appropriate subspace, and finally reconstruct the corresponding HR image in a novel mixed matching way. Experimental results indicate that the proposed approach is both quantitatively and qualitatively superior to some state-of-the-art SR methods. Yuanfei Huang, Jie Li 0001, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
IEEE Trans. Image Process. | 4 |
| 2018 | Single Image Dehazing With Depth-Aware Non-Local Total Variation RegularizationabstractSingle image dehazing can benefit many computer vision applications hence has attracted much more attention in recent years. However, it still remains a challenging task due to its double uncertainty of scene transmission and scene radiance. The existing image dehazing methods usually impair edges in the estimated transmission which leads to halo effects in the dehazing results. Besides, most existing methods suffer from noise and artifacts amplification in dense haze region after dehazing. To address these challenges, we propose a transmission adaptive regularized image recovery method for high quality single image dehazing. An initial transmission map is first obtained by a boundary constraint on the haze model. Then it is refined by applying a non-local total variation (NLTV) regularization to keep depth structures while smoothing excessive details. Noticing that the artifacts amplification effect depends on scene transmission, a transmission adaptive regularized recovery method based on NLTV is proposed to simultaneously suppress visual artifacts and preserve image details in the final dehazing result. An efficient alternating optimization algorithm is also proposed to solve the regularization model. Thorough experimental results demonstrate that the proposed method can effectively suppress visual artifacts for degraded hazy images, and yields high-quality results comparative to the state-of-the-art dehazing methods both quantitatively and qualitatively. Qi Liu 0054, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
IEEE Trans. Image Process. | 3 |
| 2017 | Video quality assessment by compact representation of energy in 3D-DCT domain
Lihuo He, Wen Lu 0004, Changcheng Jia |
Neurocomputing | 1 |
| 2017 | Single image super resolution based on sparse domain selection
Wen Lu 0004, Huxing Sun, Rui Wang 0173, Lihuo He, Ming-Jong Jou, Shensian Syu, JiShiang Li |
Neurocomputing | 4 |
| 2017 | Haze removal for a single visible remote sensing image
Qi Liu 0054, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
Signal Process. | 3 |
| 2016 | Fast image quality assessment via supervised iterative quantization method
Lihuo He, Di Wang 0011, Qi Liu 0054, Wen Lu 0004 |
Neurocomputing | 1 |
| 2016 | On combining visual perception and color structure based image quality assessment
Wen Lu 0004, Tianjiao Xu, Yuling Ren, Lihuo He |
Neurocomputing | 4 |
| 2016 | Statistical modeling in the shearlet domain for blind image quality assessment
Wen Lu 0004, Tianjiao Xu, Yuling Ren, Lihuo He |
Multim. Tools Appl. | 4 |
| 2016 | Multimodal Discriminative Binary Embedding for Large-Scale Cross-Modal RetrievalabstractMultimodal hashing, which conducts effective and efficient nearest neighbor search across heterogeneous data on large-scale multimedia databases, has been attracting increasing interest, given the explosive growth of multimedia content on the Internet. Recent multimodal hashing research mainly aims at learning the compact binary codes to preserve semantic information given by labels. The overwhelming majority of these methods are similarity preserving approaches which approximate pairwise similarity matrix with Hamming distances between the to-be-learnt binary hash codes. However, these methods ignore the discriminative property in hash learning process, which results in hash codes from different classes undistinguished, and therefore reduces the accuracy and robustness for the nearest neighbor search. To this end, we present a novel multimodal hashing method, named multimodal discriminative binary embedding (MDBE), which focuses on learning discriminative hash codes. First, the proposed method formulates the hash function learning in terms of classification, where the binary codes generated by the learned hash functions are expected to be discriminative. And then, it exploits the label information to discover the shared structures inside heterogeneous data. Finally, the learned structures are preserved for hash codes to produce similar binary codes in the same class. Hence, the proposed MDBE can preserve both discriminability and similarity for hash codes, and will enhance retrieval accuracy. Thorough experiments on benchmark data sets demonstrate that the proposed method achieves excellent accuracy and competitive computational efficiency compared with the state-of-the-art methods for large-scale cross-modal retrieval task. Di Wang 0011, Xinbo Gao 0001, Xiumei Wang 0002, Lihuo He |
IEEE Trans. Image Process. | 4 |
| 2015 | Semantic Topic Multimodal Hashing for Cross-Media Retrieval
Di Wang 0011, Xinbo Gao 0001, Xiumei Wang 0002, Lihuo He |
IJCAI | 4 |
| 2012 | Sparse representation for blind image quality assessmentabstractBlind image quality assessment (BIQA) is an important yet difficult task in image processing related applications. Existing algorithms for universal BIQA learn a mapping from features of an image to the corresponding subjective quality or divide the image into different distortions before mapping. Although these algorithms are promising, they face the following problems: (1) they require a large number of samples (pairs of distorted image and its subjective quality) to train a robust mapping; (2) they are sensitive to different datasets; and (3) they have to be retrained when new training samples are available. In this paper, we introduce a simple yet effective algorithm based upon the sparse representation of natural scene statistics (NSS) feature. It consists of three key steps: extracting NSS features in the wavelet domain, representing features via sparse coding, and weighting differential mean opinion scores by the sparse coding coefficients to obtain the final visual quality values. Thorough experiments on standard databases show that the proposed algorithm outperforms representative BIQA algorithms and some full-reference metrics. Lihuo He, Dacheng Tao, Xuelong Li 0001, Xinbo Gao 0001 |
CVPR | 1 |
| 2012 | Local Structure Divergence Index for Image Quality Assessment
Fei Gao 0006, Dacheng Tao, Xuelong Li 0001, Xinbo Gao 0001, Lihuo He |
ICONIP (5) | 5 |
| 2012 | Color Fractal Structure Model for Reduced-Reference Colorful Image Quality Assessment
Lihuo He, Dongxue Wang, Xuelong Li 0001, Dacheng Tao, Xinbo Gao 0001, Fei Gao 0006 |
ICONIP (2) | 1 |
| 2010 | A novel image quality metric based on morphological component analysisabstractDue to that human eye has different perceptual characteristics for different morphological components, so a novel image quality metric is proposed by incorporating morphological component analysis (MCA) and human visual system (HVS), which is capable of assessing the image with different types of distortion. Firstly, reference and distorted images are decomposed into texture and cartoon components by MCA respectively. Then these components are changed into perceptual features by just noticeable difference (JND) which integrates masking features, luminance adaptation and contrast sensitive function (CSF). Finally, the difference between reference and distorted images' perceptual features is quantified using a pooling strategy, and then the final result of the image quality is obtained. Experimental results demonstrate that the performance of the metric prevail over some existing methods on LIVE database II. Xuelong Li 0001, Lihuo He, Wen Lu 0004, Xinbo Gao 0001, Dacheng Tao |
SMC | 2 |