Chao Huang 0008

dblp:18/4087-8 · DBLP profile ↗
← Back
60ranked-venue papers
19as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 11 first-author · 36 since 2021Artificial intelligence and machine learning · 34 · 11 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly Detection
abstract
Industrial anomaly detection is a critical component of modern manufacturing, yet the scarcity of defective samples restricts traditional detection methods to scenario-specific applications. Although Vision-Language Models (VLMs) demonstrate significant advantages in generalization capabilities, their performance in industrial anomaly detection remains limited. To address this challenge, we propose IAD-R1, a universal post-training framework applicable to VLMs of different architectures and parameter scales, which substantially enhances their anomaly detection capabilities. IAD-R1 employs a two-stage training strategy: the Perception Activation Supervised Fine-Tuning (PA-SFT) stage utilizes a meticulously constructed high-quality Chain-of-Thought dataset (Expert-AD) for training, enhancing anomaly perception capabilities and establishing reasoning-to-answer correlations; the Structured Control Group Relative Policy Optimization (SC-GRPO) stage employs carefully designed reward functions to achieve a capability leap from "Anomaly Perception" to "Anomaly Interpretation". Experimental results demonstrate that IAD-R1 achieves significant improvements across 7 VLMs, the largest improvement was on the DAGM dataset, with average accuracy 43.3% higher than the 0.5B baseline. Notably, the 0.5B parameter model trained with IAD-R1 surpasses commercial models including GPT-4.1 and Claude-Sonnet-4 in zero-shot settings, demonstrating the effectiveness and superiority of IAD-R1.
Yunkang Cao, Chengliang Liu 0003, Yuan Xiong, Xinghui Dong, Chao Huang 0008
AAAI6
2026 Detecting Fake News in Short Videos Through Multi-View Aggregation
abstract
The increasing prominence of short video platforms has positioned them as a primary channel for public awareness of current events, while also facilitating the widespread dissemination of fake news, thus highlighting the critical need for automated detection technologies. In contrast to fake news confined to text and images, short video news encompasses multiple modalities and extensive information, presenting heightened challenges. Most existing research emphasizes the analysis of news content or user comments alone, while overlooking the crucial role of publishers, leading to poor model performance when handling fake news lacking obvious false signals. Therefore, we propose a Publisher Profiling Module to identify new false signals. To enable a more comprehensive detection of misinformation, we design a Multi-View Aggregation (MVA) model, simultaneously evaluating news from three distinct perspectives: sentiment analysis, content understanding, and publisher profiling. Late fusion is applied at the decision level to leverage the complementary strengths of these perspectives, addressing the limitations of single-view methods. Our experiments conducted on the FakeSV and FVC datasets demonstrate the superior performance of the proposed method.
Yuan Xiong, Chengliang Liu 0003, Jie Wen 0001, Chao Huang 0008
AAAI5
2026 Causal Interventional Prompt Tuning for Few-Shot Out-of-Distribution Generalization
abstract
Fine-tuning pre-trained vision-language models (VLMs) has shown substantial benefits in a wide range of downstream tasks, often achieving impressive performance with minimal labeled data. Parameter-efficient fine-tuning techniques, in particular, have demonstrated their effectiveness in enhancing downstream task performance. However, these methods frequently struggle to generalize to out-of-distribution (OOD) data due to their reliance on non-causal representations, which can introduce biases and spurious correlations that negatively impact decision-making. Such spurious factors hinder the model's generalization ability beyond the training distribution. To address these challenges, in this paper, we propose a novel causal intervention-based prompt tuning method to adapt VLMs to few-shot OOD generalization. Specifically, we leverage the front-door adjustment technique from causal inference to mitigate the effects of spurious correlations and enhance the model's focus on causal relationships. Built upon VLMs, our approach begins by decoupling causal and non-causal representations in the vision-language alignment process. The causal representation that captures only essential semantically relevant information can serve as a mediator variable between the input image and output label, mitigating the biases from the latent confounder. To further enrich this causal representation, we propose a novel text-based diversity augmentation technique that uses textual features to provide additional semantic context. This augmentation technique can enhance the diversity of the causal representation, making it more robust and generalizable to various OOD scenarios. Experimental results across multiple OOD datasets demonstrate that our method significantly outperforms existing approaches, achieving state-of-the-art generalization performance.
Jie Wen 0001, Chao Huang 0008, Chengliang Liu 0003, Yong Xu 0001, Xiaochun Cao
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Disentangling Consistent and Specific Information for Double Incomplete Multi-View Multi-Label Classification
abstract
As a prominent research topic, multi-view multi-label classification (MvMlC) aims to assign multiple labels to samples by integrating information from various perspectives. However, in real-world scenarios, MvMlC frequently faces the learning challenge of data with missing views and labels, typically resulting from sensor malfunctions, or the costly and time-consuming process of manual annotation. In addition, learning robust representations that are both consistent across views and specific to individual views remains a challenge. To address these issues, we propose a novel double incomplete multi-view multi-label classification framework based on Disentangling Consistent and Specific Information (DCSI). Specifically, we employ a dual-channel encoder with identical architecture but distinct objectives to extract cross-view consistent information and view-specific unique information from all views, respectively. Meanwhile, a view discriminator is constructed to decouple these two types of information, facilitating the extraction of pure consistent and specific information. Moreover, we meticulously design fusion strategies tailored to each representation type. Regarding consistent representations, we propose a dynamic-confidence-aware fusion mechanism that assesses the reliability of each view's representations in relation to the classification task, enabling the model to prioritize information from trustworthy representations. For specific representations, in light of their complementary rather than redundant property, we suggest treating such representations from each view equally to ensure fairness. Through experimental validation on five datasets, the results demonstrate that our method outperforms existing state-of-the-art methods.
Jie Wen 0001, Lian Zhao, Xiaohuan Lu, Chengliang Liu 0003, Li Shen 0008, Chao Huang 0008, Yong Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Visual anomaly detection under complex view-illumination interplay: A large-scale benchmark
Yunkang Cao, Xiaohao Xu, Yihan Sun 0007, Yuxiang Tan, Xiaonan Huang, Chao Huang 0008, Weiming Shen 0001
Pattern Recognit.9
2026 DyC-CLIP: Dynamic context-aware multi-modal prompt learning for zero-shot anomaly detection
Fangjun Huang, Chao Huang 0008
Pattern Recognit.3
2026 Noise-Induced Cross-Modal Information Interaction and Dual-Prompt Learning for Medical Image Segmentation
abstract
Accurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling is both labor-intensive and reliant on domain-specific expertise. To address this limitation without requiring additional annotations, we propose a novel multimodal segmentation framework that leverages medical text annotations as an auxiliary modality to complement visual information. In particular, our approach introduces a learnable encoding strategy for joint distribution modeling of image and text, which enables discriminative fusion and effectively suppresses cross-modal redundancy. Moreover, we innovatively design a frequency-domain prompt encoder based on the discrete wavelet transform (DWT) to capture multi-frequency features, thereby significantly enhancing the model's ability to delineate fine-grained boundaries. Overall, our framework integrates cross-attention for effective cross-modal interaction, employs joint distribution modeling to enable discriminative and redundancy-reduced multimodal fusion, and incorporates auxiliary supervision to strengthen the learning of task-relevant features. Extensive experiments on nine public datasets across three clinical tasks-including cell, lung infection, and polyp segmentation-demonstrate that our method achieves competitive segmentation performance while maintaining favorable computational efficiency. Comprehensive ablation studies and feature distribution visualizations further validate the effectiveness and robustness of our proposed components. The code will be made publicly available at https://github.com/chenpeng052/MDFP.
Chao Huang 0008, Jie Wen 0001, Wei Wang 0335, Li Shen 0008, Wenqi Ren, Xiaochun Cao, Chengliang Liu 0003
IEEE Trans. Image Process.1
2026 Latent Fingerprint Quality Assessment for Criminal Investigations: A Benchmark Dataset and Method
abstract
Fingerprint biometrics plays a crucial role in biometric identification, especially in applications such as criminal investigations. Although recent progress in recognition methodology has significantly enhanced automated fingerprint recognition, these systems still rely heavily on the quality of the input fingerprints. In criminal investigations, fingerprints are often of low quality due to their incidental deposition from natural oils and sweat, rather than being deliberately captured under controlled conditions. This degradation can significantly impact usability and identification accuracy, underscoring the need for effective Fingerprint Quality Assessment (FQA) methods. In this paper, we establish the Crime Scene Fingerprints quality assessment Dataset (CSFD-10k), the largest dataset of its kind, containing 11,500 fingerprint images from real criminal investigations. Of these, 10,000 samples are assigned Mean Opinion Scores (MOSs) for correlation testing, while the remaining 1,500 are labeled based on matching performance for generalizability testing. All labels are provided by frontline criminal police officers. Using this dataset, we propose a deep neural network-based Dual-Branch FQA (DB-FQA) framework that integrates image-level and edge-level features. The DB-FQA enhances ridge details by transforming raw grayscale fingerprints into edge maps using the Logical/Linear operator. A dual-branch network processes both the raw fingerprint and the edge map, and the Multi-scale Adaptive Cross feature Fusion (MACF) module fuses these features, guided by the edge map to highlight quality-related regions of interest. Extensive experiments demonstrate the robustness and superiority of our proposed method, offering substantial support for forensic fingerprint biometrics. The code and dataset are available at https://github.com/wzhsysu/FIQA.
Chao Huang 0008, Ye Zhang 0017, Peibei Cao, Zhihua Wang 0002, Yang Yu 0014, Xiaochun Cao
IEEE Trans. Image Process.1
2026 Early Pneumoconiosis Recognition From CT Images via Distance-Similarity Graph Encoding and Dynamic-Scored Adaptive Pooling
abstract
Accurate recognition of early-stage pneumoconiosis presents significant challenges due to the irregular morphology, diffuse distribution, and small size of pulmonary lesions. Existing 2D methods struggle to focus on lesion-level 3D characteristics and inter-slice correlations in localized weak lesion regions, resulting in incomplete feature extraction and inaccurate calculation of lesion volume. To obtain complete 3D fine-grained lesion features in the entire lung, this paper proposes an early pneumoconiosis recognition network (EPRNet) to enhance fine-grained feature acquisition abilities and discover inter-slice correlations, thereby improving early pneumoconiosis recognition accuracy in a more structured and flexible manner. Specifically, to obtain the fine-grained 3D features of early pneumoconiosis more comprehensively, a distance-similarity graph encoding module is proposed to construct and encode the relationships of the distributed tiny lesions within CT slices, integrating the spatial positions and the corresponding feature similarities of the lesions to improve the accuracy of pneumoconiosis feature representation. To adaptively preserve accurate graph representations of the correlations of the lesions across inter-CT slices, a hierarchical dynamic-scored adaptive pooling module is proposed to discover the potential long distance correlations between cross-slices, obtaining spatial semantic information of diffused lesions in the entire lung. Experimental results based on multiple datasets demonstrate that EPRNet achieves state-of-the-art performance while exhibiting better generalization. The ablation experiments also prove the effectiveness of each module in EPRNet.
Wenchao Jiang, Junhang Li, Quankeng Huang, Chao Huang 0008, Ji He 0001, Song Guo 0001
IEEE Trans. Image Process.4
2026 MCFL: Multimodal Collaborative Fusion Learning for Hashimoto's Thyroiditis Recognition
abstract
Ultrasound imaging and biochemical examinations are the primary methods for diagnosing Hashimoto's thyroiditis (HT). However, neither of them is sufficient to accurately diagnose HT alone. Most existing multimodal models for HT diagnosis focus primarily on extracting and concatenating features from different modalities, which are ineffective due to the dimensional imbalance of the features between the textual and image data. To address this issue, we propose a novel Multimodal Collaborative Fusion Learning (MCFL) approach, which can enhance and recalibrate the biochemical indicators using ultrasound images, effectively improving the significance and specificity of biochemical indicators for the diagnosis of HT. Specifically, MCFL first constructs a novel INNet to convert the image-level characteristics of the HT ultrasound image into two numerical indicators, i.e., the Local prominent inflammatory (Lpi) and the Global diffuse lesion (Gdl), unifying image data and textual data into a single representation space. Then, a decision tree-based optimization strategy is employed to supervise the training of INNet, interactively recalibrating biochemical indicators with the guidance of the two numerical indicators mentioned above and obtaining a more accurate feature representation of HT. Finally, based on the deep Q-learning framework, a reward mechanism is established to guide the HT diagnostic process, in which the experience replay mechanism and the $\epsilon $ -greedy strategy are utilized collaboratively to improve the accuracy and robustness of the model. Extensive experiments are conducted on a multimodal dataset from multiple medical centers, and the results demonstrate that MCFL achieves state-of-the-art performance, setting a new benchmark.
Wenchao Jiang, Guanjie Zhou, Honghua Bai, Ji He 0001, Chao Huang 0008, Song Guo 0001
IEEE Trans. Image Process.5
2026 Toward a Completely Blind Attacker for No-Reference Image Quality Assessment Models
abstract
No-reference image quality assessment (NR-IQA) models are critically vulnerable to adversarial attacks, posing significant risks to downstream vision systems. However, existing attack methods suffer from high computational costs, reliance on Mean Opinion Score (MOS) annotations, and poor cross-model transferability. To overcome these limitations, we propose Degrade-to-OverReconstruct (DOR), a novel prior knowledge-driven black-box attack framework operating in a "completely blind" manner, requiring neither MOS labels nor surrogate models, inducing significant prediction bias solely based on distortion statistics. Specifically, DOR generates universal adversarial examples by first applying mild degradation to preserve global structure and then employing aggressive over-reconstruction using a Residual Denoising Diffusion Model (RDDM) to adaptively disrupt intrinsic Natural Scene Statistics (NSS)-a shared foundation across NR-IQA models. Extensive experiments on synthetic (LIVE, TID2013) and authentic (CLIVE) datasets demonstrate DOR's strong attack performance and superior transferability against leading NR-IQA models that cover diverse deep neural network architectures. Our work pioneers a diffusion model-based "completely blind" attack paradigm, offering a practical, MOS-free solution for adversarial robustness assessment of NR-IQA models in real-world deployments.
Xinyu Ruan, Hangwei Chen, Chao Huang 0008, Wenqi Ren, Qiuping Jiang
IEEE Trans. Image Process.3
2026 Harnessing Multi-Modal Large Language Models for Measuring and Interpreting Color Differences
abstract
The accurate measurement of perceptual color differences (CDs) between two images plays an important role in modern smartphone photography. Although traditional CD metrics provide numerical scores to quantify color variations, they often lack the ability to offer intuitive insights or explanations that reflect the factors behind these differences in a way that aligns with human perception and reasoning. Here, we present CD-Reasoning, an innovative method designed not merely to compute numerical CD scores but also to provide a detailed rationale for the observed CDs between images. This method surpasses simple numerical quantification, delivering a more profound and explanatory analysis that bridges quantitative assessments with the qualitative reasoning characteristic of human perception. The development of the CD-Reasoning model begins with the compilation of a multi-modal CD dataset dubbed M-SPCD based on the existing SPCD, where we collect textual descriptions that detail the quantification of CDs across seven pivotal attributes: white balance, brightness contrast, color contrast, overall brightness, overall color, shadow detail, and highlight detail. Utilizing the newly curated M-SPCD dataset, we enhance the capabilities of cutting-edge Multimodal Large Language Models (MLLMs) to not only accurately assess numerical CD scores but also to provide in-depth reasoning that explains the CDs between two images. Extensive experiments demonstrate that the proposed CD-Reasoning not only achieves superior accuracy compared to state-of-the-art CD metrics but also significantly exceeds leading MLLMs in CD interpreting. Source codes will be available at https://github.com/LongYu-LY/CD-Reasoning.
Zhihua Wang 0002, Qiuping Jiang, Chao Huang 0008, Xiaochun Cao
IEEE Trans. Image Process.4
2025 Multi-view Evidential Learning-based Medical Image Segmentation
abstract
Medical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limited embedded knowledge scope. Vision foundation models have been demonstrated to be effective in extracting generalizable knowledge, but they cannot extract domain-specific knowledge without fine-tuning. In this work, we propose a novel multi-view evidential learning-based framework, which can extract both domain-specific and generalizable knowledge from multi-view features by combining the advantages of traditional and vision foundation models. Specifically, a novel multi-view state space model (MV-SSM) is designed to extract task-related knowledge while removing redundant information within multi-view features. The proposed MV-SSM utilizes Mamba, a state space model, to model cross-view contextual dependencies between domain-specific and generalizable features. Additionally, evidential learning is adopted to quantify the segmentation uncertainty of the model for boundary. In special, variational Dirichlet is introduced to characterize the distribution of the result probabilities, parameterized with collected evidence to quantify uncertainty. As a result, the model can reduce the segmentation uncertainties of boundaries by optimizing the parameters of the Dirichlet distribution. Experimental results on three datasets show that our method obtains superior segmentation performance.
Chao Huang 0008, Yushu Shi, Wai Keung Wong, Chengliang Liu 0003, Wei Wang 0169, Zhihua Wang 0002, Jie Wen 0001
AAAI1
2025 Federated Weakly Supervised Video Anomaly Detection with Multimodal Prompt
abstract
Video anomaly detection (VAD) aims at locating the abnormal events in videos. Recently, the Weakly Supervised VAD has made great progress, which only requires video-level annotations when training. In practical applications, different institutions may have different types of abnormal videos. However, the abnormal videos cannot be circulated on the internet due to privacy protection. To train a more generalized anomaly detector that can identify various anomalies, it is reasonable to introduce federated learning into WSVAD. In this paper, we propose Global and Local Context-driven Federated Learning, a new paradigm for privacy protected weakly supervised video anomaly detection. Specifically, we utilize the vision-language association of CLIP to detect whether the video frame is abnormal. Instead of leveraging handcrafted text prompts for CLIP, we propose a text prompt generator. The generated prompt is simultaneously influenced by text and visual. On the one hand, the text provides global context related to anomaly, which improves the model's ability of generalization. On the other hand, the visual provides personalized local context because different clients may have videos with different types of anomalies or scenes. The generated prompt ensures global generalization while processing personalized data from different clients. Extensive experiments show that the proposed method achieves remarkable performance.
Benfeng Wang, Chao Huang 0008, Jie Wen 0001, Wei Wang 0169, Yong Xu 0001
AAAI2
2025 Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language Models
abstract
With advancements in visual language models (VLMs) and large language models (LLMs), video anomaly detection (VAD) has progressed beyond binary classification to fine-grained categorization and multidimensional analysis. However, existing methods focus mainly on coarse-grained detection, lacking anomaly explanations. To address these challenges, we propose Ex-VAD, an Explainable Fine-grained Video Anomaly Detection approach that combines fine-grained classification with detailed explanations of anomalies. First, we use a VLM to extract frame-level captions, and an LLM converts them to video-level explanations, enhancing the model's explainability. Second, integrating textual explanations of anomalies with visual information greatly enhances the model's anomaly detection capability. Finally, we apply label-enhanced alignment to optimize feature fusion, enabling precise fine-grained detection. Extensive experimental results on the UCF-Crime and XD-Violence datasets demonstrate that Ex-VAD significantly outperforms existing State-of-The-Art methods.
Chao Huang 0008, Yushu Shi, Jie Wen 0001, Wei Wang 0169, Yong Xu 0001, Xiaochun Cao
ICML1
2025 Omni-Dimensional State Space Model-driven SAM for Pixel-level Anomaly Detection
abstract
Pixel-level anomaly detection is indispensable in industrial defect detection and medical diagnosis. Recently, Segment Anything Model (SAM) has achieved promising results in many vision tasks. However, direct application of the SAM to pixel-level anomaly detection tasks results in unsatisfactory performance, meanwhile SAM needs the manual prompt. Although some automatically prompt-based SAM has been proposed, these automated prompting approaches merely utilize partial image features as prompts and fail to incorporate crucial features such as multi-scale image features to generate more suitable prompts. In this paper, we propose a novel Omni Dimensional State Space Model-driven SAM (ODS-SAM) for pixel-level anomaly detection. Specifically, the proposed method adopts the SAM architecture, ensuring easy implementation and avoiding the need for fine-tuning. A State-Space Model-based residual Omni Dimensional module is designed to automatically generate suitable prompts. This module can effectively leverage multi-scale and global information, facilitating an iterative search for optimal prompts in the prompt space. The identified optimal prompts are then fed into SAM as high-dimensional tensors. Experimental results demonstrate that the proposed ODS-SAM outperforms state-of-the-art models on both industrial and medical image datasets.
Chao Huang 0008, Qianyi Li, Jie Wen 0001, Bob Zhang 0001
IJCAI1
2025 Towards VLM-based Hybrid Explainable Prompt Enhancement for Zero-Shot Industrial Anomaly Detection
abstract
Zero-Shot Industrial Anomaly Detection (ZSIAD) aims to identify and localize anomalies in industrial images from unseen categories. Owing to the powerful generalization capabilities, Vision-Language Models (VLMs) have achieved growing interest in ZSIAD. To guide the model toward understanding and localizing the semantically complex industrial anomalies, existing VLM-based methods have attempted to provide additional prompts to the model through learnable text prompt templates. However, these zero-shot methods lack detailed descriptions of specific anomalies, making it difficult to classify and segment the diverse range of industrial anomalies accurately. To address the aforementioned issue, we firstly propose the multi-stage prompt generation agent for ZSIAD. Specifically, we leverage the Multi-modal Language Large Model (MLLM) to articulate the detailed differential information between normal and test samples, which can provide detailed text prompts to the model through further refinement and anti-false alarm constraint. Moreover, we introduce the Visual Fundamental Model (VFM) to generate anomaly-related attention prompts for more accurate localization of anomalies with varying sizes and shapes. Extensive experiments on seven real-world industrial anomaly detection datasets have shown that the proposed method not only outperforms recent SOTA methods, but also its explainable prompts provide the model with a more intuitive basis for anomaly identification.
Weichao Cai, Weiliang Huang, Yunkang Cao, Chao Huang 0008, Bob Zhang 0001, Jie Wen 0001
IJCAI4
2025 Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple Annotators
abstract
Existing deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this challenge, we propose a Deep opinion-Unaware BIQA model by learning and adapting from Multiple Annotators, termed DUBMA, thereby eliminating the need for human annotations. Specifically, we first generate a large-scale set of distorted image pairs and then assign relative quality rankings using existing full-reference IQA models. The resulting dataset is subsequently employed for training our DUBMA. Due to the inherent discrepancies between synthetic and real-world distortions, a domain shift may occur. To address this, we propose an outlier-robust unsupervised domain adaptation approach leveraging optimal transport. This strategy effectively reduces the gap between synthetic and real-world distortion domains, thereby boosting the model’s adaptability and overall performance. Extensive experiments show that DUBMA outperforms existing opinion-unaware BIQA methods in terms of prediction accuracy across multiple datasets.
Zhihua Wang 0002, Xuelin Liu, Jiebin Yan, Jie Wen 0001, Wei Wang 0169, Chao Huang 0008
IJCAI6
2025 Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought
abstract
Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLMs to reason about anomalies step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks.
Chao Huang 0008, Benfeng Wang, Wei Wang 0169, Jie Wen 0001, Chengliang Liu 0003, Li Shen 0008, Xiaochun Cao
NeurIPS1
2025 Rethinking Joint Maximum Mean Discrepancy for Visual Domain Adaptation
abstract
In domain adaption (DA), joint maximum mean discrepancy (JMMD), as a famous distribution-distance metric, aims to measure joint probability distribution difference between the source domain and target domain, while it is still not fully explored and especially hard to be applied into a subspace-learning framework as its empirical estimation involves a tensor-product operator whose partial derivative is difficult to obtain. To solve this issue, we deduce a concise JMMD based on the Representer theorem that avoids the tensor-product operator and obtains two essential findings. First, we reveal the uniformity of JMMD by proving that previous marginal, class conditional, and weighted class conditional probability distribution distances are three special cases of JMMD with different label reproducing kernels. Second, inspired by graph embedding, we observe that the similarity weights, which strengthen the intra-class compactness in the graph of Hilbert Schmidt independence criterion (HSIC), take opposite signs in the graph of JMMD, revealing why JMMD degrades the feature discrimination. This motivates us to propose a novel loss JMMD-HSIC by jointly considering JMMD and HSIC to promote discrimination of JMMD. Extensive experiments on several cross-domain datasets could demonstrate the validity of our revealed theoretical results and the effectiveness of our proposed JMMD-HSIC.
Wei Wang 0335, Haifeng Xia, Chao Huang 0008, Zhengming Ding, Cong Wang 0018, Xiaochun Cao
NeurIPS3
2025 Prototype-guided and dynamic-aware video anomaly detection
Chao Huang 0008, Qianyi Li, Bob Zhang 0001
Neural Networks1
2025 Toward Efficient Test Time Adaptation With Hierarchical Distribution Alignment
abstract
A model trained in a source domain often experiences a decline in effectiveness when deployed in a different target domain, primarily due to the discrepancies between the source and target domain characteristics. Test time adaptation (TTA) provides a practical solution for addressing the domain gap by adapting the models during the test phase. Existing TTA approaches mainly focus on aligning image features into a unified feature space. However, they generally only manage to achieve broad, coarse-grained alignment across domains while overlooking the more detailed, fine-grained feature clusters within each category. Furthermore, these methods are susceptible to settling at local optima because significant details can be lost when image features are abstracted into distribution parameters. To surpass these challenges, we introduce a novel approach that ensures hierarchical cross-domain alignment at three distinct levels: category-level, subcategory-level, and sample-level. Simple category-level alignment is inadequate due to the presence of various subcategories within each category, which possess distinct semantic properties identified through unsupervised clustering in our approach. Advancing further, we enhance our method by creating synthesized features from the initially extracted category-specific features, aiming for precise sample-level alignment. During our optimization process, we redefine TTA as essentially a feature matching problem, concentrating on the calculation of feature matching probabilities. Through hierarchical distribution alignment across these levels, our method maintains the semantic consistency of cross-domain image features from a broad to a detailed scale. Unlike prior test-time adaptation methods such as Tent, our method leverages source data only once after pre-training to fit feature distributions. During the testing phase, source data is completely discarded, and the model relies solely on test sample features. This design ensures privacy preservation and makes the method well-suited for privacy-sensitive applications. Our experimental evaluations on recognized datasets demonstrate that our approach significantly surpasses other established TTA methods in performance. Our code is accessible at https://github.com/yaboliudotug/HDA-TTA.
Chao Huang 0008, Yong Xu 0001, Xiaochun Cao
IEEE Trans. Image Process.2
2025 Deep Label Propagation With Nuclear Norm Maximization for Visual Domain Adaptation
abstract
Domain adaptation aims to leverage abundant label information from a source domain to an unlabeled target domain with two different distributions. Existing methods usually rely on a classifier to generate high-quality pseudo-labels for the target domain, facilitating the learning of discriminative features. Label propagation (LP), as an effective classifier, propagates labels from the source domain to the target domain by designing a smooth function over a similarity graph, which represents structural relationships among data points in feature space. However, LP has not been thoroughly explored in deep neural network-based domain adaptation approaches. Additionally, the probability labels generated by LP are low-confident and LP is sensitive to class imbalance problem. To address these problems, we propose a novel approach for domain adaptation named deep label propagation with nuclear norm maximization (DLP-NNM). Specifically, we employ the constraint of nuclear norm maximization to enhance both label confidence and class diversity in LP and propose an efficient algorithm to solve the corresponding optimization problem. Subsequently, we utilize the proposed LP to guide the classifier layer in a deep discriminative adaptation network using the cross-entropy loss. As such, the network could produce more reliable predictions for the target domain, thereby facilitating more effective discriminative feature learning. Extensive experimental results on three cross-domain benchmark datasets demonstrate that the proposed DLP-NNM surpasses existing state-of-the-art domain adaptation approaches.
Wei Wang 0335, Cong Wang 0018, Chao Huang 0008, Zhengming Ding, Feiping Nie 0001, Xiaochun Cao
IEEE Trans. Image Process.4
2025 Optimal Graph Learning-Based Label Propagation for Cross-Domain Image Classification
abstract
Label propagation (LP) is a popular semi-supervised learning technique that propagates labels from a training dataset to a test one using a similarity graph, assuming that nearby samples should have similar labels. However, the recent cross-domain problem assumes that training (source domain) and test data sets (target domain) follow different distributions, which may unexpectedly degrade the performance of LP due to small similarity weights connecting the two domains. To address this problem, we propose optimal graph learning-based label propagation (OGL2P), which optimizes one cross-domain graph and two intra-domain graphs to connect the two domains and preserve domain-specific structures, respectively. During label propagation, the cross-domain graph draws two labels close if they are nearby in feature space and from different domains, while the intra-domain graph pulls two labels close if they are nearby in feature space and from the same domain. This makes label propagation more insensitive to cross-domain problems. During graph embedding, we optimize the three graphs using features and labels in the embedded subspace to extract locally discriminative and domain-invariant features and make the graph construction process robust to noise in the original feature space. Notably, as a more relaxed constraint, locally discriminative and domain-invariant can somewhat alleviate the contradiction between discriminability and domain-invariance. Finally, we conduct extensive experiments on five cross-domain image classification datasets to verify that OGL2P outperforms some state-of-the-art cross-domain approaches.
Wei Wang 0335, Mengzhu Wang, Chao Huang 0008, Cong Wang 0018, Jie Mu, Feiping Nie 0001, Xiaochun Cao
IEEE Trans. Image Process.3
2025 A Lesion-Fusion Neural Network for Multi-View Diabetic Retinopathy Grading
abstract
As the most common complication of diabetes, diabetic retinopathy (DR) is one of the main causes of irreversible blindness. Automatic DR grading plays a crucial role in early diagnosis and intervention, reducing the risk of vision loss in people with diabetes. In these years, various deep-learning approaches for DR grading have been proposed. Most previous DR grading models are trained using the dataset of single-field fundus images, but the entire retina cannot be fully visualized in a single field of view. There are also problems of scattered location and great differences in the appearance of lesions in fundus images. To address the limitations caused by incomplete fundus features, and the difficulty in obtaining lesion information. This work introduces a novel multi-view DR grading framework, which solves the problem of incomplete fundus features by jointly learning fundus images from multiple fields of view. Furthermore, the proposed model combines multi-view inputs such as fundus images and lesion snapshots. It utilizes heterogeneous convolution blocks (HCB) and scalable self-attention classes (SSAC), which enhance the ability of the model to obtain lesion information. The experimental results show that our proposed method performs better than the benchmark methods on the large-scale dataset.
Xiaoling Luo 0001, Qihao Xu, Zhihua Wang 0002, Chao Huang 0008, Chengliang Liu 0003, Xiaopeng Jin, Jianguo Zhang 0001
IEEE J. Biomed. Health Informatics4
2025 Multimodal Evidential Learning for Open-World Weakly-Supervised Video Anomaly Detection
abstract
Efforts in weakly-supervised video anomaly detection center on detecting abnormal events within videos by coarse-grained labels, which has been successfully applied to many real-world applications. However, a significant limitation of most existing methods is that they are only effective for specific objects in specific scenarios, which makes them prone to misclassification or omission when confronted with previously unseen anomalies. Relative to conventional anomaly detection tasks, Open-world Weakly-supervised Video Anomaly Detection (OWVAD) poses greater challenges due to the absence of labels and fine-grained annotations for unknown anomalies. To address the above problem, we propose a multi-scale evidential vision-language model to achieve open-world video anomaly detection. Specifically, we leverage generalized visual-language associations derived from CLIP to harness the full potential of large pre-trained models in addressing the OWVAD task. Subsequently, we integrate a multi-scale temporal modeling module with a multimodal evidence collector to achieve precise frame-level detection of both seen and unseen anomalies. Extensive experiments on two widely-utilized benchmarks have conclusively validated the effectiveness of our method. The code will be made publicly available.
Chao Huang 0008, Weiliang Huang, Qiuping Jiang, Wei Wang 0335, Jie Wen 0001, Bob Zhang 0001
IEEE Trans. Multim.1
2024 Attention-Induced Embedding Imputation for Incomplete Multi-View Partial Multi-Label Classification
abstract
As a combination of emerging multi-view learning methods and traditional multi-label classification tasks, multi-view multi-label classification has shown broad application prospects. The diverse semantic information contained in heterogeneous data effectively enables the further development of multi-label classification. However, the widespread incompleteness problem on multi-view features and labels greatly hinders the practical application of multi-view multi-label classification. Therefore, in this paper, we propose an attention-induced missing instances imputation technique to enhance the generalization ability of the model. Different from existing incomplete multi-view completion methods, we attempt to approximate the latent features of missing instances in embedding space according to cross-view joint attention, instead of recovering missing views in kernel space or original feature space. Accordingly, multi-view completed features are dynamically weighted by the confidence derived from joint attention in the late fusion phase. In addition, we propose a multi-view multi-label classification framework based on label-semantic feature learning, utilizing the statistical weak label correlation matrix and graph attention network to guide the learning process of label-specific features. Finally, our model is compatible with missing multi-view and partial multi-label data simultaneously and extensive experiments on five datasets confirm the advancement and effectiveness of our embedding imputation method and multi-view multi-label classification model.
Chengliang Liu 0003, Jinlong Jia, Jie Wen 0001, Xiaoling Luo 0001, Chao Huang 0008, Yong Xu 0001
AAAI6
2024 HACDR-Net: Heterogeneous-Aware Convolutional Network for Diabetic Retinopathy Multi-Lesion Segmentation
abstract
Diabetic Retinopathy (DR), the leading cause of blindness in diabetic patients, is diagnosed by the condition of retinal multiple lesions. As a difficult task in medical image segmentation, DR multi-lesion segmentation faces the main concerns as follows. On the one hand, retinal lesions vary in location, shape, and size. On the other hand, because some lesions occupy only a very small part of the entire fundus image, the high proportion of background leads to difficulties in lesion segmentation. To solve the above problems, we propose a heterogeneous-aware convolutional network (HACDR-Net) that composes heterogeneous cross-convolution, heterogeneous modulated deformable convolution, and optional near-far-aware convolution. Our network introduces an adaptive aggregation module to summarize the heterogeneous feature maps and get diverse lesion areas in the heterogeneous receptive field along the channels and space. In addition, to solve the problem of the highly imbalanced proportion of focal areas, we design a new medical image segmentation loss function, Noise Adjusted Loss (NALoss). NALoss balances the predictive feature distribution of background and lesion by jointing Gaussian noise and hard example mining, thus enhancing awareness of lesions. We conduct the experiments on the public datasets IDRiD and DDR, and the experimental results show that the proposed method achieves better performance than other state-of-the-art methods. The code is open-sourced on github.com/xqh180110910537/HACDR-Net.
Qihao Xu, Xiaoling Luo 0001, Chao Huang 0008, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001
AAAI3
2024 Diffusion-based Missing-view Generation With the Application on Incomplete Multi-view Clustering
abstract
As a branch of clustering, multi-view clustering has received much attention in recent years. In practical applications, a common phenomenon is that partial views of some samples may be missing in the collected multi-view data, which poses a severe challenge to design the multi-view learning model and explore complementary and consistent information. Currently, most of the incomplete multi-view clustering methods only focus on exploring the information of available views while few works study the missing view recovery for incomplete multi-view learning. To this end, we propose an innovative diffusion-based missing view generation (DMVG) network. Moreover, for the scenarios with high missing rates, we further propose an incomplete multi-view data augmentation strategy to enhance the recovery quality for the missing views. Extensive experimental results show that the proposed DMVG can not only accurately predict missing views, but also further enhance the subsequent clustering performance in comparison with several state-of-the-art incomplete multi-view clustering methods.
Jie Wen 0001, Wai Keung Wong, Guoqing Chao, Chao Huang 0008, Lunke Fei, Yong Xu 0001
ICML5
2024 Partial Multi-View Multi-Label Classification via Semantic Invariance Learning and Prototype Modeling
abstract
The difficulty of partial multi-view multi-label learning lies in coupling the consensus of multi-view data with the task relevance of multi-label classification, under the condition where partial views and labels are unavailable. In this paper, we seek to compress cross-view representation to maximize the proportion of shared information to better predict semantic tags. To achieve this, we establish a model consistent with the information bottleneck theory for learning cross-view shared representation, minimizing non-shared information while maintaining feature validity to help increase the purity of task-relevant information. Furthermore, we model multi-label prototype instances in the latent space and learn label correlations in a data-driven manner. Our method outperforms existing state-of-the-art methods on multiple public datasets while exhibiting good compatibility with both partial and complete data. Finally, we experimentally reveal the importance of condensing shared information under the premise of information balancing, in the process of multi-view information encoding and compression.
Chengliang Liu 0003, Gehui Xu, Jie Wen 0001, Chao Huang 0008, Yong Xu 0001
ICML5
2024 Batch Singular Value Polarization and Weighted Semantic Augmentation for Universal Domain Adaptation
abstract
As a more challenging domain adaptation setting, universal domain adaptation (UniDA) introduces category shift on top of domain shift, which needs to identify unknown category in the target domain and avoid misclassifying target samples into source private categories. To this end, we propose a novel UniDA approach named Batch Singular value Polarization and Weighted Semantic Augmentation (BSP-WSA). Specifically, we adopt an adversarial classifier to identify the target unknown category and align feature distributions between the two domains. Then, we propose to perform SVD on the classifier's outputs to maximize larger singular values while minimizing those smaller ones, which could prevent target samples from being wrongly assigned to source private classes. To better bridge the domain gap, we propose a weighted semantic augmentation approach for UniDA to generate data on common categories between the two domains. Extensive experiments on three benchmarks demonstrate that BSP-WSA could outperform existing state-of-the-art UniDA approaches.
Wangzi Qi, Wei Wang 0169, Chao Huang 0008, Jie Wen 0001, Cong Wang 0018
ICML3
2024 Long Short-Term Dynamic Prototype Alignment Learning for Video Anomaly Detection
Chao Huang 0008, Jie Wen 0001, Chengliang Liu 0003
IJCAI1
2024 Multimodal Representation Distribution Learning for Medical Image Segmentation
Chao Huang 0008, Weichao Cai, Qiuping Jiang, Zhihua Wang 0002
IJCAI1
2024 Optimal Graph Learning and Nuclear Norm Maximization for Deep Cross-Domain Robust Label Propagation
Wei Wang 0335, Chao Huang 0008, Yang Cao 0011, Cong Wang 0018, Xiaochun Cao
IJCAI4
2024 Progressive Point Cloud Denoising with Cross-Stage Cross-Coder Adaptive Edge Graph Convolution Network
abstract
Due to the limitation of collection device and unstable scanning process, point cloud data is usually noisy. This noise deforms the underlying structures of point clouds and inevitably affects downstream tasks such as rendering, reconstruction and classification. In this paper, we propose a Cross-stage Cross-coder Adaptive Edge Graph Convolution Network (C2AENet) to denoise point clouds. Our network uses multiple stages to progressively and iteratively denoise points. To improve the effectiveness, we add connections between two stages and between the encoder and decoder, leading to the cross-stage cross-coder architecture. Additionally, existing graph-based point cloud learning methods tend to capture the local structure. They typically construct a semantic graph based on semantic distance, which may ignore Euclidean neighbors and lead to insufficient geometry perception. Therefore, we introduce a geometric graph and adaptively calculate edge attention based on the local and global structural information of the points. This results in a novel graph convolution module that allows the network to capture richer contextual information and focus on more important parts. Extensive experiments demonstrate that the proposed method is competitive compared with other state-of-the-art methods. The code is available at: https://github.com/chenwuwq/C2AENet.
Hehe Fan, Qiuping Jiang, Chao Huang 0008, Yi Yang 0001
ACM Multimedia4
2024 Uncertainty-aware prototypical learning for anomaly detection in medical images
Chao Huang 0008, Yushu Shi, Bob Zhang 0001, Ke Lyu
Neural Networks1
2024 Video-Based Fall Detection Using Human Pose and Constrained Generative Adversarial Network
abstract
Falls are a major health threat for older people. A timely assistance can reduce the extent of physical injury caused by the falls. Currently, low-cost and convenient video surveillance systems based on ordinary RGB cameras are widely used for improving the safety of people. The fall detection is a research hotspot in intelligent video surveillance. In this work, we propose an unsupervised fall detection method. The proposed method first converts the RGB video frames into human pose images to eliminate the background interferences and focus on human motion and protect privacy. Afterwards, the future pose images are predicted by using the continuous historical human pose images based on a constrained generative adversarial network (GAN). Finally, the prediction errors of the human pose images and the anomaly scores of actual poses calculated by using the traditional hand-crafted features are used to realize the fall detection. As compared to the existing vision-based fall detection methods, the proposed method possesses strong generalization ability, and is robust to environmental interferences and small local occlusions, and effectively protects the privacy, and avoids time-consuming data annotations. In addition, in this work, a new large-scale and comprehensive fall dataset is created and is available for download. We perform extensive experiments on the public benchmark datasets and the proposed dataset. The results demonstrate the validity and superiority of the proposed method.
Lian Wu, Chao Huang 0008, Lunke Fei, Shuping Zhao, Jianchuan Zhao, Zhongwei Cui, Yong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Weakly Supervised Video Anomaly Detection via Self-Guided Temporal Discriminative Transformer
abstract
Weakly supervised video anomaly detection is generally formulated as a multiple instance learning (MIL) problem, where an anomaly detector learns to generate frame-level anomaly scores under the supervision of MIL-based video-level classification. However, most previous works suffer from two drawbacks: 1) they lack ability to model temporal relationships between video segments and 2) they cannot extract sufficient discriminative features to separate normal and anomalous snippets. In this article, we develop a weakly supervised temporal discriminative (WSTD) paradigm, that aims to leverage both temporal relation and feature discrimination to mitigate the above drawbacks. To this end, we propose a transformer-styled temporal feature aggregator (TTFA) and a self-guided discriminative feature encoder (SDFE). Specifically, TTFA captures multiple types of temporal relationships between video snippets from different feature subspaces, while SDFE enhances the discriminative powers of features by clustering normal snippets and maximizing the separability between anomalous snippets and normal centers in embedding space. Experimental results on three public benchmarks indicate that WSTD outperforms state-of-the-art unsupervised and weakly supervised methods, which verifies the superiority of the proposed method.
Chao Huang 0008, Chengliang Liu 0003, Jie Wen 0001, Lian Wu, Yong Xu 0001, Qiuping Jiang, Yaowei Wang 0001
IEEE Trans. Cybern.1
2024 MLFA: Toward Realistic Test Time Adaptive Object Detection by Multi-Level Feature Alignment
abstract
Object detection methods have achieved remarkable performances when the training and testing data satisfy the assumption of i.i.d. However, the training and testing data may be collected from different domains, and the gap between the domains can significantly degrade the detectors. Test Time Adaptive Object Detection (TTA-OD) is a novel online approach that aims to adapt detectors quickly and make predictions during the testing procedure. TTA-OD is more realistic than the existing unsupervised domain adaptation and source-free unsupervised domain adaptation approaches. For example, self-driving cars need to improve their perception of new environments in the TTA-OD paradigm during driving. To address this, we propose a multi-level feature alignment (MLFA) method for TTA-OD, which is able to adapt the model online based on the steaming target domain data. For a more straightforward adaptation, we select informative foreground and background features from image feature maps and capture their distributions using probabilistic models. Our approach includes: i) global-level feature alignment to align all informative feature distributions, thereby encouraging detectors to extract domain-invariant features, and ii) cluster-level feature alignment to match feature distributions for each category cluster across different domains. Through the multi-level alignment, we can prompt detectors to extract domain-invariant features, as well as align the category-specific components of image features from distinct domains. We conduct extensive experiments to verify the effectiveness of our proposed method. Our code is accessible at https://github.com/yaboliudotug/MLFA.
Chao Huang 0008, Yiling Wu, Yong Xu 0001, Xiaochun Cao
IEEE Trans. Image Process.3
2024 Information Recovery-Driven Deep Incomplete Multiview Clustering Network
abstract
Incomplete multiview clustering (IMC) is a hot and emerging topic. It is well known that unavoidable data incompleteness greatly weakens the effective information of multiview data. To date, existing IMC methods usually bypass unavailable views according to prior missing information, which is considered a second-best scheme based on evasion. Other methods that attempt to recover missing information are mostly applicable to specific two-view datasets. To handle these problems, in this article, we propose an information-recovery-driven-deep IMC network, termed as RecFormer. Concretely, a two-stage autoencoder network with self-attention structure is built to synchronously extract high-level semantic representations of multiple views and recover the missing data. Besides, we develop a recurrent graph reconstruction mechanism that cleverly leverages the restored views to promote representation learning and further data reconstruction. Visualization of recovery results are given and sufficient experimental results confirm that our RecFormer has obvious advantages over other top methods.
Chengliang Liu 0003, Jie Wen 0001, Zhihao Wu 0002, Xiaoling Luo 0001, Chao Huang 0008, Yong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 DICNet: Deep Instance-Level Contrastive Network for Double Incomplete Multi-View Multi-Label Classification
abstract
In recent years, multi-view multi-label learning has aroused extensive research enthusiasm. However, multi-view multi-label data in the real world is commonly incomplete due to the uncertain factors of data collection and manual annotation, which means that not only multi-view features are often missing, and label completeness is also difficult to be satisfied. To deal with the double incomplete multi-view multi-label classification problem, we propose a deep instance-level contrastive network, namely DICNet. Different from conventional methods, our DICNet focuses on leveraging deep neural network to exploit the high-level semantic representations of samples rather than shallow-level features. First, we utilize the stacked autoencoders to build an end-to-end multi-view feature extraction framework to learn the view-specific representations of samples. Furthermore, in order to improve the consensus representation ability, we introduce an incomplete instance-level contrastive learning scheme to guide the encoders to better extract the consensus information of multiple views and use a multi-view weighted fusion module to enhance the discrimination of semantic features. Overall, our DICNet is adept in capturing consistent discriminative representations of multi-view multi-label data and avoiding the negative effects of missing views and missing labels. Extensive experiments performed on five datasets validate that our method outperforms other state-of-the-art methods.
Chengliang Liu 0003, Jie Wen 0001, Xiaoling Luo 0001, Chao Huang 0008, Zhihao Wu 0002, Yong Xu 0001
AAAI4
2023 CIGAR: Cross-Modality Graph Reasoning for Domain Adaptive Object Detection
abstract
Unsupervised domain adaptive object detection (UDAOD) aims to learn a detector by generalizing knowledge from a labeled source domain to an unlabeled target domain. Though the existing graph-based methods for UDAOD perform well in some cases, they cannot learn a proper node set for the graph. In addition, these methods build the graph solely based on the visual features and do not consider the linguistic knowledge carried by the semantic prototypes, e.g., dataset labels. To overcome these problems, we propose a cross-modality graph reasoning adaptation (CIGAR) method to take advantage of both visual and linguistic knowledge. Specifically, our method performs cross-modality graph reasoning between the linguistic modality graph and visual modality graphs to enhance their representations. We also propose a discriminative feature selector to find the most discriminative features and take them as the nodes of the visual graph for both efficiency and effectiveness. In addition, we employ the linguistic graph matching loss to regulate the update of linguistic graphs and maintain their semantic representation during the training process. Comprehensive experiments validate the effectiveness of our proposed CIGAR.
Chao Huang 0008, Yaowei Wang 0001, Yong Xu 0001
CVPR3
2023 Highly Confident Local Structure Based Consensus Graph Learning for Incomplete Multi-view Clustering
abstract
Graph-based multi-view clustering has attracted extensive attention because of the powerful clustering-structure representation ability and noise robustness. Considering the reality of a large amount of incomplete data, in this paper, we propose a simple but effective method for incomplete multi-view clustering based on consensus graph learning, termed as HCLS_CGL. Unlike existing methods that utilize graph constructed from raw data to aid in the learning of consistent representation, our method directly learns a consensus graph across views for clustering. Specifically, we design a novel confidence graph and embed it to form a confidence structure driven consensus graph learning model. Our confidence graph is based on an intuitive similar-nearest-neighbor hypothesis, which does not require any additional information and can help the model to obtain a high-quality consensus graph for better clustering. Numerous experiments are performed to confirm the effectiveness of our method.
Jie Wen 0001, Chengliang Liu 0003, Gehui Xu, Zhihao Wu 0002, Chao Huang 0008, Lunke Fei, Yong Xu 0001
CVPR5
2023 Localized and Balanced Efficient Incomplete Multi-view Clustering
abstract
In recent years, many incomplete multi-view clustering methods have been proposed to address the challenging unsupervised clustering issue on the multi-view data with missing views. However, most of the existing works are inapplicable to large-scale clustering task and their clustering results are unstable since these methods have high computational complexities and their results are produced by kmeans rather than their designed learning models. In this paper, we propose a new one-step incomplete multi-view clustering model, called Localized and Balanced Incomplete Multi-view Clustering (LBIMVC), to address these issues. Specifically, LBIMVC develops a new graph regularized incomplete multi-matrix-factorization model to obtain the unique clustering result by learning a consensus probability representation, where each element of the consensus representation can directly reflect the probability of the corresponding sample to the class. In addition, the proposed graph regularized model integrates geometric preserving and consensus representation learning into one term without introducing any extra constraint terms and parameters to explore the structure of data. Moreover, to avoid that samples are over divided into a few clusters, a balanced constraint is introduced to the model. Experimental results on four databases demonstrate that our method not only obtains competitive clustering performance, but also performs faster than some state-of-the-art methods.
Jie Wen 0001, Gehui Xu, Chengliang Liu 0003, Lunke Fei, Chao Huang 0008, Wei Wang 0169, Yong Xu 0001
ACM Multimedia5
2023 Masked Two-channel Decoupling Framework for Incomplete Multi-view Weak Multi-label Learning
abstract
Multi-view learning has become a popular research topic in recent years, but research on the cross-application of classic multi-label classification and multi-view learning is still in its early stages. In this paper, we focus on the complex yet highly realistic task of incomplete multi-view weak multi-label learning and propose a masked two-channel decoupling framework based on deep neural networks to solve this problem. The core innovation of our method lies in decoupling the single-channel view-level representation, which is common in deep multi-view learning methods, into a shared representation and a view-proprietary representation. We also design a cross-channel contrastive loss to enhance the semantic property of the two channels. Additionally, we exploit supervised information to design a label-guided graph regularization loss, helping the extracted embedding features preserve the geometric structure among samples. Inspired by the success of masking mechanisms in image and text analysis, we develop a random fragment masking strategy for vector features to improve the learning ability of encoders. Finally, it is important to emphasize that our model is fully adaptable to arbitrary view and label absences while also performing well on the ideal full data. We have conducted sufficient and convincing experiments to confirm the effectiveness and advancement of our model.
Chengliang Liu 0003, Jie Wen 0001, Chao Huang 0008, Zhihao Wu 0002, Xiaoling Luo 0001, Yong Xu 0001
NeurIPS4
2023 Class-guided human motion prediction via multi-spatial-temporal supervision
Honghu Pan, Lian Wu, Chao Huang 0008, Xiaoling Luo 0001, Yong Xu 0001
Neural Comput. Appl.4
2023 Robust fall detection in video surveillance based on weakly supervised learning
Lian Wu, Chao Huang 0008, Shuping Zhao, Jianchuan Zhao, Zhongwei Cui, Yong Xu 0001, Min Zhang 0005
Neural Networks2
2023 Localized Sparse Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering, which aims to solve the clustering problem on the incomplete multi-view data with partial view missing, has received more and more attention in recent years. Although numerous methods have been developed, most of the methods either cannot flexibly handle the incomplete multi-view data with arbitrary missing views or do not consider the negative factor of information imbalance among views. Moreover, some methods do not fully explore the local structure of all incomplete views. To tackle these problems, this paper proposes a simple but effective method, named localized sparse incomplete multi-view clustering (LSIMVC). Different from the existing methods, LSIMVC intends to learn a sparse and structured consensus latent representation from the incomplete multi-view data by optimizing a sparse regularized and novel graph embedded multi-view matrix factorization model. Specifically, in such a novel model based on the matrix factorization, a norm based sparse constraint is introduced to obtain the sparse low-dimensional individual representations and the sparse consensus representation. Moreover, a novel local graph embedding term is introduced to learn the structured consensus representation. Different from the existing works, our local graph embedding term aggregates the graph embedding task and consensus representation learning task into a concise term. Furthermore, to reduce the imbalance factor of incomplete multi-view learning, an adaptive weighted learning scheme is introduced to LSIMVC. Comprehensive experimental results performed on six incomplete multi-view databases verify that the performance of our LSIMVC is superior to the state-of-the-art IMC approaches.
Chengliang Liu 0003, Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Chao Huang 0008
IEEE Trans. Multim.5
2023 Self-Supervised Attentive Generative Adversarial Networks for Video Anomaly Detection
abstract
Video anomaly detection (VAD) refers to the discrimination of unexpected events in videos. The deep generative model (DGM)-based method learns the regular patterns on normal videos and expects the learned model to yield larger generative errors for abnormal frames. However, DGM cannot always do so, since it usually captures the shared patterns between normal and abnormal events, which results in similar generative errors for them. In this article, we propose a novel self-supervised framework for unsupervised VAD to tackle the above-mentioned problem. To this end, we design a novel self-supervised attentive generative adversarial network (SSAGAN), which is composed of the self-attentive predictor, the vanilla discriminator, and the self-supervised discriminator. On the one hand, the self-attentive predictor can capture the long-term dependences for improving the prediction qualities of normal frames. On the other hand, the predicted frames are fed to the vanilla discriminator and self-supervised discriminator for performing true-false discrimination and self-supervised rotation detection, respectively. Essentially, the role of the self-supervised task is to enable the predictor to encode semantic information into the predicted normal frames via adversarial training, in order for the angles of rotated normal frames can be detected. As a result, our self-supervised framework lessens the generalization ability of the model to abnormal frames, resulting in larger detection errors for abnormal frames. Extensive experimental results indicate that SSAGAN outperforms other state-of-the-art methods, which demonstrates the validity and advancement of SSAGAN.
Chao Huang 0008, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Jian Yang 0003, Yaowei Wang 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Deep Object Detection with Example Attribute Based Prediction Modulation
abstract
Deep object detectors suffer from the gradient contribution imbalance during training. In this paper, we point out that such imbalance can be ascribed to the imbalance in example attributes, e.g., difficulty and shape variation degree. We further propose example attribute based prediction modulation (EAPM) to address it. In EAPM, first, the attribute of an example is defined by the prediction and the corresponding ground truth. Then, a modulating factor w.r.t the example attribute is introduced to modulate the prediction error. Finally, the new prediction and the ground-truth are input into the loss function. Essentially, we adjust the gradients of examples with specific attributes to reweight their contribution on the global gradients. We apply EAPM with focal loss and balanced L1 loss to simultaneously solve the imbalance in classification and localization. The experimental results on MS COCO demonstrate that EAPM can bring substantial improvement for deep object detectors.
Zhihao Wu 0002, Chengliang Liu 0003, Chao Huang 0008, Jie Wen 0001, Yong Xu 0001
ICASSP3
2022 Hierarchical Graph Embedded Pose Regularity Learning via Spatio-Temporal Transformer for Abnormal Behavior Detection
abstract
Abnormal behavior detection in surveillance video is a fundamental task in modern public security. Different from typical pixel-based solutions, pose-based approaches leverage low-dimensional and strongly-structured skeleton feature, which enables the anomaly detector to be immune to complex background noise and obtain higher efficiency. However, existing pose-based methods only utilize the pose of each individual independently while ignore the important interactions between individuals. In this paper, we present a hierarchical graph embedded pose regularity learning framework via spatio-temporal transformer, which leverages the strength of graph representation in encoding strongly-structured skeleton feature. Specifically, skeleton feature is encoded as the hierarchical graph representation, which jointly models the interactions among multiple individuals and the correlations among body joints within the same individual. Furthermore, a novel task-specific spatial-temporal graph transformer is designed to encode the hierarchical spatio-temporal graph embeddings of human skeletons and learn the regular patterns within normal training videos. Experimental results indicate that our method obtains superior performance over state-of-the-art methods on several challenging datasets.
Chao Huang 0008, Zheng Zhang 0006, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001, Yaowei Wang 0001
ACM Multimedia1
2022 Pixel-Level Anomaly Detection via Uncertainty-aware Prototypical Transformer
abstract
Pixel-level visual anomaly detection, which aims to recognize the abnormal areas from images, plays an important role in industrial fault detection and medical diagnosis. However, it is a challenging task due to the following reasons: i) the large variation of anomalies; and ii) the ambiguous boundary between anomalies and their normal surroundings. In this work, we present an uncertainty-aware prototypical transformer (UPformer), which takes into account both the diversity and uncertainty of anomaly to achieve accurate pixel-level visual anomaly detection. To this end, we first design a memory-guided prototype learning transformer encoder to learn and memorize the prototypical representations of anomalies for enabling the model to capture the diversity of anomalies. Additionally, an anomaly detection uncertainty quantizer is designed to learn the distributions of anomaly detection for measuring the anomaly detection uncertainty. Furthermore, an uncertainty-aware transformer decoder is proposed to leverage the detection uncertainties to guide the model to focus on the uncertain areas and generate the final detection results. As a result, our method achieves more accurate anomaly detection by combining the benefits of prototype learning and uncertainty estimation. Experimental results on five datasets indicate that our method achieves state-of-the-art anomaly detection performance.
Chao Huang 0008, Chengliang Liu 0003, Zheng Zhang 0006, Zhihao Wu 0002, Jie Wen 0001, Qiuping Jiang, Yong Xu 0001
ACM Multimedia1
2022 Weakly Supervised Video Anomaly Detection via Transformer-Enabled Temporal Relation Learning
abstract
Weakly supervised video anomaly detection is a challenging problem due to the lack of refined frame-level labels in training videos. Most prior works typically address it with the multiple instance learning paradigm, which divides a video into multiple snippets and trains a snippet classifier to distinguish anomalies from normal snippets via video-level classification loss. However, these solutions are limited in the insufficient representations. In this paper, we propose a novel weakly supervised temporal relation learning framework for anomaly detection, which efficiently explores the temporal relation between snippets and enhances the discriminative powers of features using only video-level labelled videos. To this end, we design a transformer-enabled feature encoder to convert the input task-agnostic features into discriminative task-specific features by mining the semantic similarity and position relation between snippets. As a result, our model can make a more accurate anomaly detection for current video snippet based on the learned discriminative features. Experimental results indicate that the proposed method is superior to existing state-of-the-art approaches, which demonstrates the effectiveness of our model.
Dasheng Zhang, Chao Huang 0008, Chengliang Liu 0003, Yong Xu 0001
IEEE Signal Process. Lett.2
2022 Self-Supervision-Augmented Deep Autoencoder for Unsupervised Visual Anomaly Detection
abstract
Deep autoencoder (AE) has demonstrated promising performances in visual anomaly detection (VAD). Learning normal patterns on normal data, deep AE is expected to yield larger reconstruction errors for anomalous samples, which is utilized as the criterion for detecting anomalies. However, this hypothesis cannot be always tenable since the deep AE usually captures the low-level shared features between normal and abnormal data, which leads to similar reconstruction errors for them. To tackle this problem, we propose a self-supervised representation-augmented deep AE for unsupervised VAD, which can enlarge the gap of anomaly scores between normal and abnormal samples by introducing autoencoding transformation (AT). Essentially, AT is introduced to facilitate AE to learn the high-level visual semantic features of normal images by introducing a self-supervision task (transformation reconstruction). In particular, our model inputs the original and transformed images into the encoder for obtaining latent representations; afterward, they are fed to the decoder for reconstructing both the original image and applied transformation. In this way, our model can utilize both image and transformation reconstruction errors to detect anomaly. Extensive experiments indicate that the proposed method outperforms other state-of-the-art methods, which demonstrates the validity and advancement of our model.
Chao Huang 0008, Zehua Yang, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Jian Yang 0003, Yaowei Wang 0001
IEEE Trans. Cybern.1
2022 Abnormal Event Detection Using Deep Contrastive Learning for Intelligent Video Surveillance System
abstract
The continuous developments of urban and industrial environments have increased the demand for intelligent video surveillance. Deep learning has achieved remarkable performance for anomaly detection in surveillance videos. Previous approaches achieve anomaly detection with a single-pretext task (image reconstruction or prediction) and detect anomalies by larger reconstruction error or poor prediction. However, they cannot fully exploit the discriminative semantics and temporal context information. Moreover, tackling anomaly detection with a single pretext task is suboptimal due to the nonalignment between the pretext task and anomaly detection. In this article, we propose a temporal-aware contrastive network (TAC-Net) to address the abovementioned problems of anomaly detection for intelligence video surveillance. TAC-Net is an unsupervised method that utilizes deep contrastive self-supervised learning to capture the high-level semantic features and tackles anomaly detection with multiple self-supervised tasks. During inference phase, the multiple task losses and contrastive similarity are utilized to calculate the anomaly score. Experimental results show that our method is superior to state-of-the-art approaches on three benchmarks, which demonstrates the validity and advancement of TAC-Net.
Chao Huang 0008, Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Yaowei Wang 0001
IEEE Trans. Ind. Informatics1
2022 Unsupervised Decomposition and Correction Network for Low-Light Image Enhancement
abstract
Vision-based intelligent driving assistance systems and transportation systems can be improved by enhancing the visibility of the scenes captured in extremely challenging conditions. In particular, many low-image image enhancement (LIE) algorithms have been proposed to facilitate such applications in low-light conditions. While deep learning-based methods have achieved substantial success in this field, most of them require paired training data, which is difficult to be collected. This paper advocates a novel Unsupervised Decomposition and Correction Network (UDCN) for LIE without depending on paired data for training. Inspired by the Retinex model, our method first decomposes images into illumination and reflectance components with an image decomposition network (IDN). Then, the decomposed illumination is processed by an illumination correction network (ICN) and fused with the reflectance to generate a primary enhanced result. In contrast with fully supervised learning approaches, UDCN is an unsupervised one which is trained only with low-light images and corresponding histogram equalized (HE) counterparts (can be derived from the low-light image itself) as input. Both the decomposition and correction networks are optimized under the guidance of hybrid no-reference quality-aware losses and inter-consistency constraints between the low-light image and its HE counterpart. In addition, we also utilize an unsupervised noise removal network (NRN) to remove the noise previously hidden in the darkness for further improving the primary result. Qualitative and quantitative comparison results are reported to demonstrate the efficacy of UDCN and its superiority over several representative alternatives in the literature. The results and code will be made public available athttps://github.com/myd945/UDCN.
Qiuping Jiang, Yudong Mao, Runmin Cong, Wenqi Ren, Chao Huang 0008, Feng Shao 0001
IEEE Trans. Intell. Transp. Syst.5
2021 Inter-layer correlation-based adaptive bit allocation for enhancement layer in scalable high efficiency video coding
Zongju Peng, Dongrong Jiang, Chao Huang 0008, Gangyi Jiang, Mei Yu 0001
Signal Process. Image Commun.4
2021 Online Learning-Based Multi-Stage Complexity Control for Live Video Coding
abstract
High Efficiency Video Coding (HEVC) can significantly improve the compression efficiency in comparison with the preceding H.264/Advanced Video Coding (AVC) but at the cost of extremely high computational complexity. Hence, it is challenging to realize live video applications on low-delay and power-constrained devices, such as the smart mobile devices. In this article, we propose an online learning-based multi-stage complexity control method for live video coding. The proposed method consists of three stages: multi-accuracy Coding Unit (CU) decision, multi-stage complexity allocation, and Coding Tree Unit (CTU) level complexity control. Consequently, the encoding complexity can be accurately controlled to correspond with the computing capability of the video-capable device by replacing the traditional brute-force search with the proposed algorithm, which properly determines the optimal CU size. Specifically, the multi-accuracy CU decision model is obtained by an online learning approach to accommodate the different characteristics of input videos. In addition, multi-stage complexity allocation is implemented to reasonably allocate the complexity budgets to each coding level. In order to achieve a good trade-off between complexity control and rate distortion (RD) performance, the CTU-level complexity control is proposed to select the optimal accuracy of the CU decision model. The experimental results show that the proposed algorithm can accurately control the coding complexity from 100% to 40%. Furthermore, the proposed algorithm outperforms the state-of-the-art algorithms in terms of both accuracy of complexity control and RD performance.
Chao Huang 0008, Zongju Peng, Yong Xu 0001, Qiuping Jiang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho
IEEE Trans. Image Process.1
2019 Encoding Complexity Control for Live Video Applications: An Interpretable Machine Learning Approach
abstract
In this paper, we propose an interpretable machine learning-based complexity control method for efficiently im-plementing HEVC on live video applications with different computing capacities and limited powers. Specifically, a complexity allocation method is designed to reasonably assign the complexity resources. Then, a multi-accuracy Coding Unit (CU) decision model is obtained by interpret-ably adjusting the parameters to efficiently and flexibly achieve a tradeoff between encoding complexity and rate distortion performance. Finally, a coding tree unit-level complexity control method is proposed to select appropri-ate accuracy of the CU decision model for making the en-coding complexity approach the target. The experimental results show that the proposed method outperforms state-of-the-art methods in terms of accuracy and encoding efficiency.
Chao Huang 0008, Zongju Peng, Qiuping Jiang, Gangyi Jiang
ICME1
2019 Multiple classifier-based fast coding unit partition for intra coding in future video coding
Zongju Peng, Chao Huang 0008, Gangyi Jiang, Mei Yu 0001
Signal Process. Image Commun.2