EDBT 2026 Demo / reviewers in the wild / expert
Changqing Zhang 0002
dblp:78/2668-2
· DBLP profile ↗
135ranked-venue papers
16as first author
62since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 85 · 14 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 67 · 7 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized DomainsabstractAlthough multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically evaluated on a small selection of public datasets, a limited scope that inadequately represents the complexity and diversity of real-world scenarios, potentially leading to biased evaluations. This issue presents a twofold challenge. On one hand, models may overfit to the biases of specific datasets, hindering their generalization to broader practical applications. On the other hand, the absence of a unified evaluation standard makes fair and objective comparisons between different fusion methods difficult. Consequently, a truly universal and high-performance fusion model has yet to emerge. To address these challenges, we have developed a large-scale, domain-adaptive benchmark for multimodal evaluation. This benchmark integrates over 30 datasets, encompassing 15 modalities and 20 predictive tasks across key application domains. To complement this, we have also developed an open-source, unified, and automated evaluation pipeline that includes standardized implementations of state-of-the-art models and diverse fusion paradigms. Leveraging this platform, we have conducted large-scale experiments, successfully establishing new performance baselines across multiple tasks. This work provides the academic community with a crucial platform for rigorous and reproducible assessment of multimodal models, aiming to propel the field of multimodal artificial intelligence to new heights. Leyan Xue, Changqing Zhang 0002, Kecheng Xue, Zongbo Han |
AAAI | 2 |
| 2026 | Retrieval-augmented prompt for out-of-distribution detection
Ruisong Han, Guofeng Zhang 0031, Mingyue Cheng 0004, Changqing Zhang 0002, Zongbo Han |
Pattern Recognit. | 5 |
| 2026 | Enhancing Reliability in Medical Image Classification of Imperfect ViewsabstractThe fusion of multi-view medical images through deep neural networks is essential for boosting diagnostic precision in the field of medical image analysis. However, the reliability of these diagnostic results is often compromised by imperfections in image views, manifested as noise, artifacts, and data deficits arising from inconsistent diagnostic frequencies. These issues introduce a significant risk when merging medical views in a clinical setting. To address these problems, we introduce the Reliability-Enhanced Multi-view Network (REMNet), a novel framework designed to tackle two critical challenges: 1) reducing misclassification and uncertainty from imperfect view integration, and 2) improving the reliability and interpretability of multi-view medical image predictions. Specifically, REMNet merges information from multiple views into a coherent evidence framework and incorporates a Dirichlet prior within our predictive model to more accurately estimate confidence in predictions. Coupled with a robust fusion strategy and a precise confidence calibration process, REMNet consolidates the diverse strengths of various medical imaging views, reduces the impact of view imperfections, and enhances the reliability of medical imaging diagnostics. The superiority of REMNet is validated through comprehensive theoretical analysis and empirical experiments on multi-view medical image datasets across different modalities. Wei Liu 0303, Yufei Chen 0002, Xiaodong Yue 0002, Changqing Zhang 0002, Shaorong Xie |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Uncertainty-Aware Multi-View Graph ClusteringabstractMulti-view clustering has achieved advanced progress over the years, which typically integrates multi-view information to learn discriminative common representations or a unified clustering distribution for clustering. However, existing methods either simply regard each view as equally important or assign fixed weights to each view, which are insufficient to dynamically assess the sample quality variations caused by noise in multi-view data. To address this issue, by effectively modeling the uncertainty of different samples across different views, this paper proposes a novel uncertainty-aware multi-view graph clustering network, termed UMGC-Net, which achieves trusted multi-view clustering in an unsupervised manner. Specifically, by measuring the clustering distribution entropy, an uncertainty-guided common feature learning mechanism is proposed to estimate the uncertainty for each sample of each view, thus learning multi-view features friendly to clustering. Besides, a cross-view trusted distribution fusion module is designed to obtain robust clustering distribution by exploring the trusted consistency among multi-view clustering distributions based on uncertainty. Finally, experimental results on four popular multi-view datasets validate the superior performance of the proposed UMGC-Net. Bo Peng 0007, Shaobo Bai, Jianjun Lei 0001, Changqing Zhang 0002, Nam Ling |
IEEE Trans. Multim. | 5 |
| 2025 | Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation ModelabstractVision-language foundation models have exhibited remarkable success across a multitude of downstream tasks due to their scalability on extensive image-text paired data. However, these models also display significant limitations when applied to downstream tasks, such as fine-grained image classification, as a result of ``decision shortcuts'' that hinder their generalization capabilities. In this work, we find that the CLIP model possesses a rich set of features, encompassing both desired invariant causal features and undesired decision shortcuts. Moreover, the underperformance of CLIP on downstream tasks originates from its inability to effectively utilize pre-trained features in accordance with specific task requirements. To address this challenge, we propose a simple yet effective method, Spurious Feature Eraser (SEraser), to alleviate the decision shortcuts by erasing the spurious features. Specifically, we introduce a test-time prompt tuning paradigm that optimizes a learnable prompt, thereby compelling the model to exploit invariant features while disregarding decision shortcuts during the inference phase. The proposed method effectively alleviates excessive dependence on potentially misleading spurious information. We conduct comparative analysis of the proposed method against various approaches which validates the significant superiority. Huan Ma 0006, Changqing Zhang 0002, Peilin Zhao, Baoyuan Wu, Long-Kai Huang, Qinghua Hu, Bingzhe Wu |
AAAI | 3 |
| 2025 | Exploring Task-Level Optimal Prompts for Visual In-Context LearningabstractWith the development of Vision Foundation Models (VFMs) in recent years, Visual In-Context Learning (VICL) has become a better choice compared to modifying models in most scenarios. Different from retraining or fine-tuning models, VICL does not require modifications to the model's weights and architecture, and only needs a prompt with demonstrations to teach VFM how to solve tasks. Currently, significant computational cost for finding optimal prompts for every test sample hinders the deployment of VICL, as determining which demonstrations to use for constructing the prompt is very costly. In this paper, however, we find a counterintuitive phenomenon that most test samples actually achieve optimal performance under the same prompts, and searching for sample-level prompts only costs much time but results in completely identical prompts actually. Therefore, we propose task-level prompting to reduce the cost of searching for prompts during the inference stage and introduce two time-saving yet effective task-level prompt search strategies accordingly. Extensive experimental results show that our proposed method can identify near-optimal prompts and reach the best VICL performance with a minimal cost that prior work has never achieved. Huan Ma 0006, Changqing Zhang 0002 |
AAAI | 3 |
| 2025 | Bridging the Vision-Brain Gap with an Uncertainty-Aware Blur PriorabstractCan our brain signals faithfully reflect the original visual stimuli, even including high-frequency details? Although human perceptual and cognitive capacities enable us to process and remember visual information, these abilities are constrained by several factors, such as limited attentional resources and finite capacity of visual memory. When visual stimuli are processed by human visual system into brain signals, some information is inevitably lost, leading to a discrepancy known as the System GAP. Additionally, perceptual and cognitive dynamics, along with technical noise in signal acquisition, degrade the fidelity of brain signals relative to the visual stimuli, known as the Random GAP. When encoded brain representations are directly aligned with the corresponding pretrained image features, the System GAP and Random GAP between paired data challenge the model, requiring it to bridge these gaps. However, due to limited paired data, these gaps are difficult for the model to learn, leading to overfitting and poor generalization to new data. To address these GAPs, we propose a simple yet effective approach called the Uncertainty-Aware Blur Prior (UBP). It estimates the uncertainty within the paired data, reflecting the mismatch between brain signals and visual stimuli. Based on uncertainty, UBP dynamically blurs the high-frequency details of the original images, reducing the impact of mismatch and improving alignment. Our method achieves a top-1 accuracy of 50.9% and a top-5 accuracy of 79.7% on the zero-shot brain-to-image retrieval task, surpassing previous state-of-the-art methods. Code is available at https://github.com/HaitaoWuTJU/Uncertainty-Aware-Blur-Prior. Qing Li 0018, Changqing Zhang 0002, Xiaomin Ying |
CVPR | 3 |
| 2025 | COME: Test-time Adaption by Conservatively Minimizing EntropyabstractMachine learning models must continuously self-adjust themselves for novel data distribution in the open world. As the predominant principle, entropy minimization (EM) has been proven to be a simple yet effective cornerstone in existing test-time adaption (TTA) methods. While unfortunately its fatal limitation (i.e., overconfidence) tends to result in model collapse. For this issue, we propose to \textbf{\texttt{Co}}nservatively \textbf{\texttt{M}}inimize the \textbf{\texttt{E}}ntropy (\texttt{COME}), which is a simple drop-in replacement of traditional EM to elegantly address the limitation. In essence, \texttt{COME} explicitly models the uncertainty by characterizing a Dirichlet prior distribution over model predictions during TTA. By doing so, \texttt{COME} naturally regularizes the model to favor conservative confidence on unreliable samples. Theoretically, we provide a preliminary analysis to reveal the ability of \texttt{COME} in enhancing the optimization stability by introducing a data-adaptive lower bound on the entropy. Empirically, our method achieves state-of-the-art performance on commonly used benchmarks, showing significant improvements in terms of classification accuracy and uncertainty estimation under various settings including standard, life-long and open-world TTA, i.e., up to $34.5\%$ improvement on accuracy and $15.1\%$ on false positive rate. Our code is available at: \href{https://github.com/BlueWhaleLab/COME}{https://github.com/BlueWhaleLab/COME}. Yatao Bian, Xinke Kong, Peilin Zhao, Changqing Zhang 0002 |
ICLR | 5 |
| 2025 | Multimodal Data Augmentation for Dynamic FusionabstractCompared to single modality, multimodal data can provide more comprehensive information for recognition tasks. As the quality of multimodal data varies in open environments, to make full use of the value of each modality and mitigate the impact of low-quality modality data, uncertainty-aware dynamic multimodal fusion has emerged as a promising learning paradigm. Although such methods are widely used, the robustness of these models is still significantly affected by the lack of samples with dynamically varying multimodal quality in training sets. This lackness limits the models’ ability of uncertainty estimation and dynamic multimodal fusion in open environments. In this study, we propose a data augmentation method towards dynamic multimodal fusion. First, through adversarial learning, we generate model-aware high-uncertainty samples, overcoming the limitations of manually predefined low-quality samples. Second, we mix the generated low-quality modalities with original high-quality modality samples to construct multimodal data with dynamically varying quality, thereby enhancing the scenario diversity of multimodal data. Our method not only integrates easily with existing dynamic fusion methods, but also demonstrates state-of-the-art performance in both efficacy and reliability through comprehensive experiments on various benchmark datasets. Haining Li, Changqing Zhang 0002 |
IJCNN | 3 |
| 2025 | A Theoretical Proof of Dynamic Multimodal Fusion Exacerbates Modality GreedyabstractIn recent years, numerous studies have proposed uncertainty-guided multimodal learning to adapt to dynamic relationships between different modalities. Dynamic multimodal fusion enables more dominant modalities to receive greater weight during the fusion process, thereby preventing the influence of spurious features from less reliable modalities on decision-making. However, there is No free lunch. We observe that the introduction of dynamic fusion during training exacerbates the model's tendency toward Greedy (a phenomenon known to induce decision shortcuts in multimodal learning). This results in a model that does not fully take advantage of the lower quality modalities. In this paper, we provide a theoretical analysis showing that dynamic fusion intensifies Greedy, and we present experimental results that support this observation. In summary, this paper explains the Greedy risk in dynamic multimodal learning from both theoretical and experimental perspectives, serving as a cautionary reminder for researchers when employing dynamic multimodal learning. Our code is available at https://github.com/d-xr/GreedyDynFusion. Xiaorui Ding, Huan Ma 0006, Changqing Zhang 0002 |
ACM Multimedia | 3 |
| 2025 | Gradient-Aware Revitalization of Non-Effective Samples in Medical Image SegmentationabstractDeep learning-based medical image segmentation, with its precise lesion localization capabilities, serves as a core component in multimedia medical applications and intelligent diagnostic assistance systems. Recent innovations in network architectures significantly improve segmentation performance. However, the Non-Effective Samples (NES) on model optimization receive little attention. These samples are characterized by minimal gradient variations in loss during training and exhibit a slight contribution to model optimization. They encompass well-segmented samples with near-zero loss values and challenging samples with consistently high loss values. Especially, when NES accumulate, the model will fall into an optimization trap, causing the optimization to stagnate. To address this issue, we propose a lightweight plug-and-play Gradient-Aware Sample Selection and Reactivation Strategy (GA-SRS) that efficiently identifies and revitalizes the training potential of NES. Firstly, GA-SRS filters NES out based on the historical training information of samples and the variations in the loss gradient during the training process. Then, GA-SRS revitalizes the training values of these samples through strong data augmentation. Extensive experiments on four public datasets and three general models demonstrate the effectiveness of GA-SRS. For example, GA-SRS helps improve the IoU metric of U-KAN from 67.20% to 71.24% on the BUSI and from 81.20% to 82.51% on the ISIC dataset, achieving state-of-the-art experimental results. Shiying Lin, Qinghua Lin, Jiawei Wu 0001, Changqing Zhang 0002 |
ACM Multimedia | 6 |
| 2025 | DOTA: Distributional Test-time Adaptation of Vision-Language ModelsabstractVision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when significant distribution gaps exist between training and test data, while fine-tuning for diverse scenarios is often costly. Cache-based test-time adapters offer an efficient alternative by storing representative test samples to guide subsequent classifications. Yet, these methods typically employ naive cache management with limited capacity, leading to severe catastrophic forgetting when samples are inevitably dropped during updates. In this paper, we propose DOTA (DistributiOnal Test-time Adaptation), a simple yet effective method addressing this limitation. Crucially, instead of merely memorizing individual test samples, DOTA continuously estimates the underlying distribution of the test data stream. Test-time posterior probabilities are then computed using these dynamically estimated distributions via Bayes' theorem for adaptation. This distribution-centric approach enables the model to continually learn and adapt to the deployment environment. Extensive experiments validate that DOTA significantly mitigates forgetting and achieves state-of-the-art performance compared to existing methods. Zongbo Han, Jialong Yang, Junfan Li, Qianli Xu, Zheng Shou 0001, Changqing Zhang 0002 |
NeurIPS | 7 |
| 2025 | Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationabstractExisting methods to enhance the reasoning capability of large language models predominantly rely on supervised fine-tuning (SFT) followed by reinforcement learning (RL) on reasoning-specific data. These approaches critically depend on external supervisions--such as labeled reasoning traces, verified golden answers, or pre-trained reward models. In this work, we propose Entropy Minimized Policy Optimization (EMPO), which makes an early attempt at fully unsupervised LLM reasoning incentivization. By minimizing the semantic entropy of LLMs on unlabeled questions, EMPO achieves competitive performance compared to supervised counterparts. Specifically, without any supervised signals, EMPO boosts the accuracy of Qwen2.5-Math-7B Base from 33.7\% to 51.6\% on math benchmarks and improves the accuracy of Qwen2.5-7B Base from 32.1\% to 50.1\% on MMLU-Pro. Primary analysis are also provided to interpret the effectiveness of EMPO. Code is available at https://github.com/QingyangZhang/EMPO. Changqing Zhang 0002, Peilin Zhao, Yatao Bian |
NeurIPS | 3 |
| 2025 | CARE: A Calibration-Aware Cross-Institutional Collaboration Framework for Medical Image Classification
Yafei Yang, Qing Li 0018, Changqing Zhang 0002 |
PRCV (13) | 3 |
| 2025 | Trustworthy deep learning for encrypted traffic classification
Yanbei Liu, Changqing Zhang 0002, Wanjin Shan |
Soft Comput. | 3 |
| 2025 | Smooth Multiple Kernel k-Means via Underlying Graph FilteringabstractClustering has attracted more and more attention as one of the most fundamental techniques in the field of unsupervised learning. To deal with nonlinear problems, clustering methods have been extended to the kernel version. As a traditional kernel clustering algorithm, multiple kernel k-means (MKKM) aims to learn clustering results from a consensus kernel obtained by combining a set of predefined kernels optimally. However, we observe that the existing MKKM algorithm and its variants insufficiently consider the noise that existed in kernel space and the underlying structure of kernelized data points. To this end, we propose a novel smooth MKKM via underlying graph filtering (SMKKM-UGF) to learn the smooth representations of kernelized data points through their nearby nodes in the underlying graph. In particular, different from the common graph filter, we jointly update the graph filter while learning the smooth kernel, so that the graph filter can be guaranteed to adapt to the updating kernel space constantly. Besides, an iterative algorithm with proven convergence is designed to solve the resultant optimization problem. Extensive experiments have been performed on numerous benchmark datasets, whose results prove the superiority of the proposed SMKKM-UGF compared to the other state-of-the-art clustering methods. The demo code of this work is publicly available at https://github.com/wqyang23/SMKKM-UGF.git. Wenqi Yang, Chang Tang, Xinwang Liu 0002, Guanghui Yue 0001, Yuanyuan Liu 0004, Changqing Zhang 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | ID-like Prompt Learning for Few-Shot Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection methods often exploit auxiliary outliers to train model identifying OOD samples, especially discovering challenging outliers from auxiliary outliers dataset to improve OOD detection. However, they may still face limitations in effectively distinguishing between the most challenging OOD samples that are much like in-distribution (ID) data, i.e., ID-like samples. To this end, we propose a novel OOD detection framework that discovers ID-like outliers using CLIP [32]from the vicinity space of the ID samples, thus helping to identify these most challenging OOD samples. Then a prompt learning framework is proposed that utilizes the identified ID-like outliers to further leverage the capabilities of CLIP for OOD detection. Benefiting from the powerful CLIP, we only need a small number of ID samples to learn the prompts of the model without exposing other auxiliary outlier datasets. By focusing on the most challenging ID-like OOD samples and elegantly exploiting the capabilities of CLIP, our method achieves superior few-shot learning performance on various real-world image datasets (e.g., in 4-shot OOD detection on the ImageNet-1k dataset, our method reduces the average FPR95 by 12.16% and improves the average AUROC by 2.76%, compared to state-of-the-art methods). Code is available at https://github.com/ycfate/ID-like Yichen Bai, Zongbo Han, Bing Cao 0002, Xiaoheng Jiang, Qinghua Hu, Changqing Zhang 0002 |
CVPR | 6 |
| 2024 | Test-time Adaptation against Multi-modal Reliability BiasabstractTest-time adaptation (TTA) has emerged as a new paradigm for reconciling distribution shifts across domains without accessing source data. However, existing TTA methods mainly concentrate on uni-modal tasks, overlooking the complexity of multi-modal scenarios.
In this paper, we delve into the multi-modal test-time adaptation and reveal a new challenge named reliability bias. Different from the definition of traditional distribution shifts, reliability bias refers to the information discrepancies across different modalities derived from intra-modal distribution shifts. To solve the challenge, we propose a novel method, dubbed REliable fusion and robust ADaptation (READ). On the one hand, unlike the existing TTA paradigm that mainly repurposes the normalization layers, READ employs a new paradigm that modulates the attention between modalities in a self-adaptive way, supporting reliable fusion against reliability bias. On the other hand, READ adopts a novel objective function for robust multi-modal adaptation, where the contributions of confident predictions could be amplified and the negative impacts of noisy predictions could be mitigated. Moreover, we introduce two new benchmarks to facilitate comprehensive evaluations of multi-modal TTA under reliability bias. Extensive experiments on the benchmarks verify the effectiveness of our method against multi-modal reliability bias. The code and benchmarks are available at https://github.com/XLearning-SCU/2024-ICLR-READ. Mouxing Yang, Yunfan Li 0003, Changqing Zhang 0002, Peng Hu 0002, Xi Peng 0001 |
ICLR | 3 |
| 2024 | Predictive Dynamic FusionabstractMultimodal fusion is crucial in joint decision-making systems for rendering holistic judgments. Since multimodal data changes in open environments, dynamic fusion has emerged and achieved remarkable progress in numerous applications. However, most existing dynamic multimodal fusion methods lack theoretical guarantees and easily fall into suboptimal problems, yielding unreliability and instability. To address this issue, we propose a Predictive Dynamic Fusion (PDF) framework for multimodal learning. We proceed to reveal the multimodal fusion from a generalization perspective and theoretically derive the predictable Collaborative Belief (Co-Belief) with Mono- and Holo-Confidence, which provably reduces the upper bound of generalization error. Accordingly, we further propose a relative calibration strategy to calibrate the predicted Co-Belief for potential uncertainty. Extensive experiments on multiple benchmarks confirm our superiority. Our code is available at https://github.com/Yinan-Xia/PDF. Bing Cao 0002, Yinan Xia, Changqing Zhang 0002, Qinghua Hu |
ICML | 4 |
| 2024 | Uncertainty-aware Dynamic Re-weighting Robust ClassificationabstractIn recent years, deep neural networks (DNNs) have achieved significant success in a variety of computer vision tasks. However, data collected from the wild inevitably suffer from noise, impairing the performance of neural networks. Moreover, due to the diversity of noise, it is impossible to predetermine the types of noise in test time, making traditional methods for handling single noise inadequate. In order to train reliable networks for noisy data, we present a novel framework to systematically address complex noise issues in the real world. Specifically, we propose an uncertainty-aware dynamic re-weighting classification framework. Focusing on the quality and difficulty of samples, our method categorizes samples into easy-to-learn, challenging and unclassifiable samples. Introducing uncertainty to distinguish among these categories, our method dynamically adjusts the weights of samples to prevent the network from overfitting noisy data. Finally, we validate the proposed method on both synthetic and real-world datasets, demonstrating advanced performance in classification tasks. Wenshu Ge, Huan Ma 0006, Peizhuo Sheng, Changqing Zhang 0002 |
IJCNN | 4 |
| 2024 | Can I Trust You? Rethinking Calibration with Controllable Confidence RankingabstractWhile modeling uncertainty is becoming one of the mainstays in machine learning, the problem of evaluating the estimated confidence remains challenging. For the confidence estimation in classification, the main difficulty lies in lacking sample-level confidence. Existing measures either conduct an evaluation based on a set of predictions (Expected Calibration Error: ECE) or use 0/1 as an approximation of confidence (Negative Log-Likelihood: NLL). In this study, we reveal that these measures are inadequate in evaluating reasonable calibration, and propose to evaluate the predictive confidence for modern neural networks using a complementary internal measure to exploit the sample-level confidence. Specifically, based on the principle "information is the eliminated uncertainty", we propose a general framework to investigate whether (even well-calibrated) models estimate the confidence of a sample that matches the information it contains with controllable-confidence-ranking (CCR), and then devise a novel regularization strategy stabilizing the sample-level confidence estimation, producing more reasonable and sharper confidence. We conduct extensive experiments to validate the effectiveness and potential of our method. Wenshu Ge, Huan Ma 0006, Changqing Zhang 0002 |
IJCNN | 3 |
| 2024 | Test-Time Dynamic Image FusionabstractThe inherent challenge of image fusion lies in capturing the correlation of multi-source images and comprehensively integrating effective information from different sources. Most existing techniques fail to perform dynamic image fusion while notably lacking theoretical guarantees, leading to potential deployment risks in this field. Is it possible to conduct dynamic image fusion with a clear theoretical justification? In this paper, we give our solution from a generalization perspective. We proceed to reveal the generalized form of image fusion and derive a new test-time dynamic image fusion paradigm. It provably reduces the upper bound of generalization error. Specifically, we decompose the fused image into multiple components corresponding to its source data. The decomposed components represent the effective information from the source data, thus the gap between them reflects the \textit{Relative Dominability} (RD) of the uni-source data in constructing the fusion image. Theoretically, we prove that the key to reducing generalization error hinges on the negative correlation between the RD-based fusion weight and the uni-source reconstruction loss. Intuitively, RD dynamically highlights the dominant regions of each source and can be naturally converted to the corresponding fusion weight, achieving robust results. Extensive experiments and discussions with in-depth analysis on multiple benchmarks confirm our findings and superiority. Our code is available at https://github.com/Yinan-Xia/TTD. Bing Cao 0002, Yinan Xia, Changqing Zhang 0002, Qinghua Hu |
NeurIPS | 4 |
| 2024 | Interactive Deep Clustering via Value MiningabstractIn the absence of class priors, recent deep clustering methods resort to data augmentation and pseudo-labeling strategies to generate supervision signals. Though achieved remarkable success, existing works struggle to discriminate hard samples at cluster boundaries, mining which is particularly challenging due to their unreliable cluster assignments. To break such a performance bottleneck, we propose incorporating user interaction to facilitate clustering instead of exhaustively mining semantics from the data itself. To be exact, we present Interactive Deep Clustering (IDC), a plug-and-play method designed to boost the performance of pre-trained clustering models with minimal interaction overhead. More specifically, IDC first quantitatively evaluates sample values based on hardness, representativeness, and diversity, where the representativeness avoids selecting outliers and the diversity prevents the selected samples from collapsing into a small number of clusters. IDC then queries the cluster affiliations of high-value samples in a user-friendly manner. Finally, it utilizes the user feedback to finetune the pre-trained clustering model. Extensive experiments demonstrate that IDC could remarkably improve the performance of various pre-trained clustering models, at the expense of low user interaction costs. The code could be accessed at pengxi.me. Peng Hu 0002, Changqing Zhang 0002, Yunfan Li 0003, Xi Peng 0001 |
NeurIPS | 3 |
| 2024 | Out-Of-Distribution Detection with Diversification (Provably)abstractOut-of-distribution (OOD) detection is crucial for ensuring reliable deployment of machine learning models. Recent advancements focus on utilizing easily accessible auxiliary outliers (e.g., data from the web or other datasets) in training. However, we experimentally reveal that these methods still struggle to generalize their detection capabilities to unknown OOD data, due to the limited diversity of the auxiliary outliers collected. Therefore, we thoroughly examine this problem from the generalization perspective and demonstrate that a more diverse set of auxiliary outliers is essential for enhancing the detection capabilities. However, in practice, it is difficult and costly to collect sufficiently diverse auxiliary outlier data. Therefore, we propose a simple yet practical approach with a theoretical guarantee, termed Diversity-induced Mixup for OOD detection (diverseMix), which enhances the diversity of auxiliary outlier set for training in an efficient way. Extensive experiments show that diverseMix achieves superior performance on commonly used and recent challenging large-scale benchmarks, which further confirm the importance of the diversity of auxiliary outliers. Haiyun Yao, Zongbo Han, Huazhu Fu, Xi Peng 0001, Qinghua Hu, Changqing Zhang 0002 |
NeurIPS | 6 |
| 2024 | The Best of Both Worlds: On the Dilemma of Out-of-distribution DetectionabstractOut-of-distribution (OOD) detection is essential for model trustworthiness which aims to sensitively identity semantic OOD samples and robustly generalize for covariate-shifted OOD samples. However, we discover that the superior OOD detection performance of state-of-the-art methods is achieved by secretly sacrificing the OOD generalization ability. The classification accuracy frequently collapses catastrophically when even slight noise is encountered. Such a phenomenon violates the motivation of trustworthiness and significantly limits the model's deployment in the real world. What is the hidden reason behind such a limitation? In this work, we theoretically demystify the "\textit{sensitive-robust}" dilemma that lies in previous OOD detection methods. Consequently, a theory-inspired algorithm is induced to overcome such a dilemma. By decoupling the uncertainty learning objective from a Bayesian perspective, the conflict between OOD detection and OOD generalization is naturally harmonized and a dual-optimized performance could be expected. Empirical studies show that our method achieves superior performance on commonly used benchmarks. To our best knowledge, this work is the first principled OOD detection method that achieves state-of-the-art OOD detection performance without sacrificing OOD generalization ability. Our code is available at https://github.com/QingyangZhang/DUL. Qiuxuan Feng, Joey Tianyi Zhou, Yatao Bian, Qinghua Hu, Changqing Zhang 0002 |
NeurIPS | 6 |
| 2024 | Rethinking the Reliability of Post-hoc Calibration Methods Under Subpopulation Shift
Huan Ma 0006, Changqing Zhang 0002, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
PRICAI (2) | 3 |
| 2024 | A principled framework for explainable multimodal disentanglement
Zongbo Han, Tao Luo 0014, Huazhu Fu, Qinghua Hu, Joey Tianyi Zhou, Changqing Zhang 0002 |
Inf. Sci. | 6 |
| 2024 | Confidence-aware multi-modality learning for eye disease screening
Ke Zou, Tian Lin 0002, Zongbo Han, Meng Wang 0001, Xuedong Yuan, Haoyu Chen 0002, Changqing Zhang 0002, Xiaojing Shen, Huazhu Fu |
Medical Image Anal. | 7 |
| 2024 | Semantic Invariant Multi-View Clustering With Fully Incomplete InformationabstractRobust multi-view learning with incomplete information has received significant attention due to issues such as incomplete correspondences and incomplete instances that commonly affect real-world multi-view applications. Existing approaches heavily rely on paired samples to realign or impute defective ones, but such preconditions cannot always be satisfied in practice due to the complexity of data collection and transmission. To address this problem, we present a novel framework called SeMantic Invariance LEarning (SMILE) for multi-view clustering with incomplete information that does not require any paired samples. To be specific, we discover the existence of invariant semantic distribution across different views, which enables SMILE to alleviate the cross-view discrepancy to learn consensus semantics without requiring any paired samples. The resulting consensus semantics remains unaffected by cross-view distribution shifts, making them useful for realigning/imputing defective instances and forming clusters. We demonstrate the effectiveness of SMILE through extensive comparison experiments with 13 state-of-the-art baselines on five benchmarks. Our approach improves the clustering accuracy of NoisyMNIST from 19.3%/23.2% to 82.7%/69.0% when the correspondences/instances are fully incomplete. We will release the code after acceptance. Pengxin Zeng, Mouxing Yang, Yiding Lu, Changqing Zhang 0002, Peng Hu 0002, Xi Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Evidential Pseudo-Label Ensemble for semi-supervised classification
Changqing Zhang 0002, Huan Ma 0006 |
Pattern Recognit. Lett. | 2 |
| 2024 | Causal Evidence Learning for Trusted Open Set Recognition Under Covariate ShiftabstractTrusted open set recognition aims to classify known classes and reject unknown ones, as well as outputs an uncertainty estimate to measure the reliability of recognition results, thus extending the application scenarios of traditional open set recognition methods to risk-sensitive fields. Current methods assume that the covariate distribution of the known classes remains constant during training and testing. However, due to the common occurrence of covariate shift in practical applications, existing methods often suffer from limited generalization. To this end, a causal evidence learning framework, highlighted by the controllable Evidential Uncertainty Guided Adversarial Data Augmentation (EUG-ADA) and Causal Adversarial Disentanglement (CausalAD) strategies, is proposed to support trusted open set recognition under covariate shift. Specifically, EUG-ADA generates high-quality augmentation samples to increase training data diversity, guided by controllable evidential uncertainty and constrained by semantic consistency. Moreover, it is complemented by the CausalAD, which learns causal representations through causal intervention, mitigating the risk of misrecognition of unknown classes caused by the model’s reliance on shortcuts for prediction. The combined effect of EUG-ADA and CausalAD enables the model to learn more generalized and robust causal evidence for trusted open set recognition. Finally, extensive experimental results on both real-world and synthetic data validate the effectiveness of the proposed method, demonstrating that it improves not only open set recognition performance under covariate shift but also the reliability of uncertainty estimates. The code is released onhttps://github.com/ScorpioBao/CEL-OSR. Qingsen Bao, Lei Chen 0011, Feng Zhang 0052, Jun Wang 0024, Changqing Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Stabilizing Multispectral Pedestrian Detection With Evidential Hybrid FusionabstractMultispectral pedestrian detection is an important task due to its critical role in a wide spectrum of applications. Basically, the complementary information from color and thermal images could provide a more accurate and reliable pedestrian detection result. However, multimodal data usually suffer from the issue of dynamic change or corruption for some modalities. At the same time, as a safety-critical task, how to produce a stable and reliable detection result is also a key challenge. To address these challenges, we propose a stable multispectral pedestrian detection (SMPD) algorithm, providing a new paradigm for multispectral detection by dynamically integrating different modalities at an evidence level. Specifically, we introduce the Dirichlet distribution to characterize the distribution of the class probabilities, parameterized with evidence from different modalities. Then, multi-branch fusion, based on Dempster-Shafer theory, can integrate these pieces of evidence to obtain the detection result. In addition, a Plug-and-Play module, termed modal enhancement module, is introduced to enhance cross-modality interaction. This is an end-to-end framework, which can induce accurate detection and uncertainty estimation, and then endows the model with both reliability and robustness against noise or corruption. Extensive experimental results demonstrate the efficiency of our algorithm compared with state-of-the-art methods. Qing Li 0018, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001, Huazhu Fu, Lei Chen 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Pick-and-Place Transform Learning for Fast Multi-View ClusteringabstractTo manipulate large-scale data, anchor-based multi-view clustering methods have grown in popularity owing to their linear complexity in terms of the number of samples. However, these existing approaches pay less attention to two aspects. 1) They target at learning a shared affinity matrix by using the local information from every single view, yet ignoring the global information from all views, which may weaken the ability to capture complementary information. 2) They do not consider the removal of feature redundancy, which may affect the ability to depict the real sample relationships. To this end, we propose a novel fast multi-view clustering method via pick-and-place transform learning named PPTL, which could capture insightful global features to characterize the sample relationships quickly. Specifically, PPTL first concatenates all the views along the feature direction to produce a global matrix. Considering the redundancy of the global matrix, we design a pick-and-place transform with ℓ2,p-norm regularization to abandon the poor features and consequently construct a compact global representation matrix. Thus, by conducting anchor-based subspace clustering on the compact global representation matrix, PPTL can learn a consensus skinny affinity matrix with a discriminative clustering structure. Numerous experiments performed on small-scale to large-scale datasets demonstrate that our method is not only faster but also achieves superior clustering performance over state-of-the-art methods across a majority of the datasets. Qiangqiang Shen, Yongyong Chen, Changqing Zhang 0002, Yonghong Tian 0001, Yongsheng Liang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Autoencoder-Based Collaborative Attention GAN for Multi-Modal Image SynthesisabstractMulti-modal images are required in a wide range of practical scenarios, from clinical diagnosis to public security. However, certain modalities may be incomplete or unavailable because of the restricted imaging conditions, which commonly leads to decision bias in many real-world applications. Despite the significant advancement of existing image synthesis techniques, learning complementary information from multi-modal inputs remains challenging. To address this problem, we propose an autoencoder-based collaborative attention generative adversarial network (ACA-GAN) that uses available multi-modal images to generate the missing ones. The collaborative attention mechanism deploys a single-modal attention module and a multi-modal attention module to effectively extract complementary information from multiple available modalities. Considering the significant modal gap, we further developed an autoencoder network to extract the self-representation of target modality, guiding the generative model to fuse target-specific information from multiple modalities. This considerably improves cross-modal consistency with the desired modality, thereby greatly enhancing the image synthesis performance. Quantitative and qualitative comparisons for various multi-modal image synthesis tasks highlight the superiority of our approach over several prior methods by demonstrating more precise and realistic results. Bing Cao 0002, Haifang Cao, Pengfei Zhu 0001, Changqing Zhang 0002, Qinghua Hu |
IEEE Trans. Multim. | 5 |
| 2024 | Autoencoder in Autoencoder NetworksabstractModeling complex correlations on multiview data is still challenging, especially for high-dimensional features with possible noise. To address this issue, we propose a novel unsupervised multiview representation learning (UMRL) algorithm, termed autoencoder in autoencoder networks (AE2-Nets). The proposed framework effectively encodes information from high-dimensional heterogeneous data into a compact and informative representation with the proposed bidirectional encoding strategy. Specifically, the proposed AE2-Nets conduct encoding in two directions: the inner-AE-networks extract view-specific intrinsic information (forward encoding), while the outer-AE-networks integrate this view-specific intrinsic information from different views into a latent representation (backward encoding). For the nested architecture, we further provide a probabilistic explanation and extension from hierarchical variational autoencoder. The forward-backward strategy flexibly addresses high-dimensional (noisy) features within each view and encodes complementarity across multiple views in a unified framework. Extensive results on benchmark datasets validate the advantages compared to the state-of-the-art algorithms. Changqing Zhang 0002, Zongbo Han, Yeqing Liu, Huazhu Fu, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Safe Multi-View Deep ClassificationabstractMulti-view deep classification expects to obtain better classification performance than using a single view. However, due to the uncertainty and inconsistency of data sources, adding data views does not necessarily lead to the performance improvements in multi-view classification. How to avoid worsening classification performance when adding views is crucial for multi-view deep learning but rarely studied. To tackle this limitation, in this paper, we reformulate the multi-view classification problem from the perspective of safe learning and thereby propose a Safe Multi-view Deep Classification (SMDC) method, which can guarantee that the classification performance does not deteriorate when fusing multiple views. In the SMDC method, we dynamically integrate multiple views and estimate the inherent uncertainties among multiple views with different root causes based on evidence theory. Through minimizing the uncertainties, SMDC promotes the evidences from data views for correct classification, and in the meantime excludes the incorrect evidences to produce the safe multi-view classification results. Furthermore, we theoretically prove that in the safe multi-view classification, adding data views will certainly not increase the empirical risk of classification. The experiments on various kinds of multi-view datasets validate that the proposed SMDC method can achieve precise and safe classification results. Wei Liu 0303, Yufei Chen 0002, Xiaodong Yue 0002, Changqing Zhang 0002, Shaorong Xie |
AAAI | 4 |
| 2023 | Exploring and Exploiting Uncertainty for Incomplete Multi-View ClassificationabstractClassifying incomplete multi-view data is inevitable since arbitrary view missing widely exists in real-world applications. Although great progress has been achieved, existing incomplete multi-view methods are still difficult to obtain a trustworthy prediction due to the relatively high uncertainty nature of missing views. First, the missing view is of high uncertainty, and thus it is not reasonable to provide a single deterministic imputation. Second, the quality of the imputed data itself is of high uncertainty. To explore and exploit the uncertainty, we propose an Uncertainty-induced Incomplete Multi-View Data Classification (UIMC) model to classify the incomplete multi-view data under a stable and reliable framework. We construct a distribution and sample multiple times to characterize the uncertainty of missing views, and adaptively utilize them according to the sampling quality. Accordingly, the proposed method realizes more perceivable imputation and controllable fusion. Specifically, we model each missing data with a distribution conditioning on the available views and thus introducing uncertainty. Then an evidence-based fusion strategy is employed to guarantee the trustworthy integration of the imputed views. Extensive experiments are conducted on multiple benchmark data sets and our method establishes a state-of-the-art performance in terms of both performance and trustworthiness. Mengyao Xie, Zongbo Han, Changqing Zhang 0002, Yichen Bai, Qinghua Hu |
CVPR | 3 |
| 2023 | Graph Matching with Bi-level Noisy CorrespondenceabstractIn this paper, we study a novel and widely existing problem in graph matching (GM), namely, Bi-level Noisy Correspondence (BNC), which refers to node-level noisy correspondence (NNC) and edge-level noisy correspondence (ENC). In brief, on the one hand, due to the poor recognizability and viewpoint differences between images, it is inevitable to inaccurately annotate some keypoints with offset and confusion, leading to the mismatch between two associated nodes, i.e., NNC. On the other hand, the noisy node-to-node correspondence will further contaminate the edge-to-edge correspondence, thus leading to ENC. For the BNC challenge, we propose a novel method termed Contrastive Matching with Momentum Distillation. Specifically, the proposed method is with a robust quadratic contrastive loss which enjoys the following merits: i) better exploring the node-to-node and edge-to-edge correlations through a GM customized quadratic contrastive learning paradigm; ii) adaptively penalizing the noisy assignments based on the confidence estimated by the momentum teacher. Extensive experiments on three real-world datasets show the robustness of our model compared with 12 competitive baselines. The code is available at https://github.com/XLearning-SCU/2023-ICCV-COMMON. Yijie Lin 0001, Mouxing Yang, Jun Yu 0002, Peng Hu 0002, Changqing Zhang 0002, Xi Peng 0001 |
ICCV | 5 |
| 2023 | Calibrating Multimodal LearningabstractMultimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, we identify current multimodal classification methods suffer from unreliable predictive confidence that tend to rely on partial modalities when estimating confidence. Specifically, we find that the confidence estimated by current models could even increase when some modalities are corrupted. To address the issue, we introduce an intuitive principle for multimodal learning, i.e., the confidence should not increase when one modality is removed. Accordingly, we propose a novel regularization technique, i.e., Calibrating Multimodal Learning (CML) regularization, to calibrate the predictive confidence of previous methods. This technique could be flexibly equipped by existing models and improve the performance in terms of confidence calibration, classification accuracy, and model robustness. Huan Ma 0006, Changqing Zhang 0002, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
ICML | 3 |
| 2023 | dugMatting: Decomposed-Uncertainty-Guided MattingabstractCutting out an object and estimating its opacity mask, known as image matting, is a key task in image and video editing. Due to the highly ill-posed issue, additional inputs, typically user-defined trimaps or scribbles, are usually needed to reduce the uncertainty. Although effective, it is either time consuming or only suitable for experienced users who know where to place the strokes. In this work, we propose a decomposed-uncertainty-guided matting (dugMatting) algorithm, which explores the explicitly decomposed uncertainties to efficiently and effectively improve the results. Basing on the characteristic of these uncertainties, the epistemic uncertainty is reduced in the process of guiding interaction (which introduces prior knowledge), while the aleatoric uncertainty is reduced in modeling data distribution (which introduces statistics for both data and possible noise). The proposed matting framework relieves the requirement for users to determine the interaction areas by using simple and efficient labeling. Extensively quantitative and qualitative results validate that the proposed method significantly improves the original matting algorithms in terms of both efficiency and efficacy. Jiawei Wu 0001, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Joey Tianyi Zhou |
ICML | 2 |
| 2023 | Provable Dynamic Fusion for Low-Quality Multimodal DataabstractThe inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal data, dynamic multimodal fusion emerges as a promising learning paradigm. Despite its widespread use, theoretical justifications in this field are still notably lacking. Can we design a provably robust multimodal fusion method? This paper provides theoretical understandings to answer this question under a most popular multimodal fusion framework from the generalization perspective. We proceed to reveal that several uncertainty estimation solutions are naturally available to achieve robust multimodal fusion. Then a novel multimodal fusion framework termed Quality-aware Multimodal Fusion (QMF) is proposed, which can improve the performance in terms of classification accuracy and model robustness. Extensive experimental results on multiple benchmarks can support our findings. Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Joey Tianyi Zhou, Xi Peng 0001 |
ICML | 3 |
| 2023 | Fairness-guided Few-shot Prompting for Large Language ModelsabstractLarge language models have demonstrated surprising ability to perform in-context learning, i.e., these models can be directly applied to solve numerous downstream tasks by conditioning on a prompt constructed by a few input-output examples. However, prior research has shown that in-context learning can suffer from high instability due to variations in training examples, example order, and prompt formats. Therefore, the construction of an appropriate prompt is essential for improving the performance of in-context learning. In this paper, we revisit this problem from the view of predictive bias. Specifically, we introduce a metric to evaluate the predictive bias of a fixed prompt against labels or a given attributes. Then we empirically show that prompts with higher bias always lead to unsatisfactory predictive quality. Based on this observation, we propose a novel search strategy based on the greedy search to identify the near-optimal prompt for improving the performance of in-context learning. We perform comprehensive experiments with state-of-the-art mainstream models such as GPT-3 on various downstream tasks. Our results indicate that our method can enhance the model's in-context learning performance in an effective and interpretable manner. Huan Ma 0006, Changqing Zhang 0002, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang 0013, Huazhu Fu, Qinghua Hu, Bingzhe Wu |
NeurIPS | 2 |
| 2023 | Trusted Multi-View Classification With Dynamic Evidential FusionabstractExisting multi-view classification algorithms focus on promoting accuracy by exploiting different views, typically integrating them into common representations for follow-up tasks. Although effective, it is also crucial to ensure the reliability of both the multi-view integration and the final decision, especially for noisy, corrupted and out-of-distribution data. Dynamically assessing the trustworthiness of each view for different samples could provide reliable integration. This can be achieved through uncertainty estimation. With this in mind, we propose a novel multi-view classification algorithm, termed trusted multi-view classification (TMC), providing a new paradigm for multi-view learning by dynamically integrating different views at an evidence level. The proposed TMC can promote classification reliability by considering evidence from each view. Specifically, we introduce the variational Dirichlet to characterize the distribution of the class probabilities, parameterized with evidence from different views and integrated with the Dempster-Shafer theory. The unified learning framework induces accurate uncertainty and accordingly endows the model with both reliability and robustness against possible noise or corruption. Both theoretical and experimental results validate the effectiveness of the proposed model in accuracy, robustness and trustworthiness. Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Multiview Clustering by Consensus Spectral Rotation FusionabstractMultiview clustering (MVC) aims to partition data into different groups by taking full advantage of the complementary information from multiple views. Most existing MVC methods fuse information of multiple views at the raw data level. They may suffer from performance degradation due to the redundant information contained in the raw data. Graph learning-based methods often heavily depend on one specific graph construction, which limits their practical applications. Moreover, they often require a computational complexity ofO(n3) because of matrix inversion or eigenvalue decomposition for each iterative computation. In this paper, we propose a consensus spectral rotation fusion (CSRF) method to learn a fused affinity matrix for MVC at the spectral embedding feature level. Specifically, we first introduce a CSRF model to learn a consensus low-dimensional embedding, which explores the complementary and consistent information across multiple views. We develop an alternating iterative optimization algorithm to solve the CSRF optimization problem, where a computational complexity ofO(n2) is required during each iterative computation. Then, the sparsity policy is introduced to design two different graph construction schemes, which are effectively integrated with the CSRF model. Finally, a multiview fused affinity matrix is constructed from the consensus low-dimensional embedding in spectral embedding space. We analyze the convergence of the alternating iterative optimization algorithm and provide an extension of CSRF for incomplete MVC. Extensive experiments on multiview datasets demonstrate the effectiveness and efficiency of the proposed CSRF method. Jie Chen 0065, Hua Mao 0001, Dezhong Peng, Changqing Zhang 0002, Xi Peng 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | SGU-Net: Shape-Guided Ultralight Network for Abdominal Image SegmentationabstractConvolutional neural networks (CNNs) have achieved significant success in medical image segmentation. However, they also suffer from the requirement of a large number of parameters, leading to a difficulty of deploying CNNs to low-source hardwares, e.g., embedded systems and mobile devices. Although some compacted or small memory-hungry models have been reported, most of them may cause degradation in segmentation accuracy. To address this issue, we propose a shape-guided ultralight network (SGU-Net) with extremely low computational costs. The proposed SGU-Net includes two main contributions: it first presents an ultralight convolution that is able to implement double separable convolutions simultaneously, i.e., asymmetric convolution and depthwise separable convolution. The proposed ultralight convolution not only effectively reduces the number of parameters but also enhances the robustness of SGU-Net. Secondly, our SGU-Net employs an additional adversarial shape-constraint to let the network learn shape representation of targets, which can significantly improve the segmentation accuracy for abdomen medical images using self-supervision. The SGU-Net is extensively tested on four public benchmark datasets, LiTS, CHAOS, NIH-TCIA and 3Dircbdb. Experimental results show that SGU-Net achieves higher segmentation accuracy using lower memory costs, and outperforms state-of-the-art networks. Moreover, we apply our ultralight convolution into a 3D volume segmentation network, which obtains a comparable performance with fewer parameters and memory usage. Tao Lei 0003, Xiaogang Du, Huazhu Fu, Changqing Zhang 0002, Asoke K. Nandi |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Structural Attention Graph Neural Network for Diagnosis and Prediction of COVID-19 SeverityabstractWith rapid worldwide spread of Coronavirus Disease 2019 (COVID-19), jointly identifying severe COVID-19 cases from mild ones and predicting the conversion time (from mild to severe) is essential to optimize the workflow and reduce the clinician's workload. In this study, we propose a novel framework for COVID-19 diagnosis, termed as Structural Attention Graph Neural Network (SAGNN), which can combine the multi-source information including features extracted from chest CT, latent lung structural distribution, and non-imaging patient information to conduct diagnosis of COVID-19 severity and predict the conversion time from mild to severe. Specifically, we first construct a graph to incorporate structural information of the lung and adopt graph attention network to iteratively update representations of lung segments. To distinguish different infection degrees of left and right lungs, we further introduce a structural attention mechanism. Finally, we introduce demographic information and develop a multi-task learning framework to jointly perform both tasks of classification and regression. Experiments are conducted on a real dataset with 1687 chest CT scans, which includes 1328 mild cases and 359 severe cases. Experimental results show that our method achieves the best classification (e.g., 86.86% in terms of Area Under Curve) and regression (e.g., 0.58 in terms of Correlation Coefficient) performance, compared with other comparison methods. Yanbei Liu, Henan Li, Tao Luo 0010, Changqing Zhang 0002, Zhitao Xiao, Ying Wei 0009, Yaozong Gao, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Confidence-Aware Fusion Using Dempster-Shafer Theory for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection is an important and valuable task in many applications, which could provide a more accurate and reliable pedestrian detection result by using the complementary visual information from color and thermal images. However, it faces two open and difficult challenges: 1) how to effectively and dynamically integrate multispectral information according to the confidence of different modalities, and 2) how to produce a reliable prediction result. In this paper, we propose a novel confidence-aware multispectral pedestrian detection (CMPD) method, which flexibly learns the multispectral representation while simultaneously producing a reliable result with confidence estimation. Specifically, a dense fusion strategy is first proposed to extract the multilevel multispectral representation at the feature level. Then, an additional confidence subnetwork is utilized to dynamically estimate the detection confidence for each modality. Finally, Dempster's combination rule is introduced to fuse the results of different branches according to the rectified confidence. Our proposed CMPD method not only effectively integrates multimodal information but also provides a reliable prediction. Extensive experimental results demonstrate the efficiency of our algorithm compared with state-of-the-art methods. Qing Li 0018, Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Pengfei Zhu 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal ClassificationabstractIntegration of heterogeneous and high-dimensional data (e.g., multiomics) is becoming increasingly important. Existing multimodal classification algorithms mainly focus on improving performance by exploiting the complementarity from different modalities. However, conventional approaches are basically weak in providing trustworthy multimodal fusion, especially for safety-critical applications (e.g., medical diagnosis). For this issue, we propose a novel trustworthy multimodal classification algorithm termed Multimodal Dynamics, which dynamically evaluates both the feature-level and modality-level informativeness for different samples and thus trustworthily integrates multiple modalities. Specifically, a sparse gating is introduced to capture the information variation of each within-modality feature and the true class probability is employed to assess the classification confidence of each modality. Then a transparent fusion algorithm based on the dynamical informativeness estimation strategy is induced. To the best of our knowledge, this is the first work to jointly model both feature and modality variation for different samples to provide trustworthy fusion in multi-modal classification. Extensive experiments are conducted on multimodal medical classification datasets. In these experiments, superior performance and trustworthiness of our algorithm are clearly validated compared to the state-of-the-art methods. Zongbo Han, Fan Yang 0081, Junzhou Huang, Changqing Zhang 0002, Jianhua Yao 0001 |
CVPR | 4 |
| 2022 | Trustworthy Long-Tailed ClassificationabstractClassification on long-tailed distributed data is a challenging problem, which suffers from serious class-imbalance and accordingly unpromising performance es-pecially on tail classes. Recently, the ensembling based methods achieve the state-of-the-art performance and show great potential. However, there are two limitations for cur-rent methods. First, their predictions are not trustworthy for failure-sensitive applications. This is especially harmful for the tail classes where the wrong predictions is basically fre-quent. Second, they assign unified numbers of experts to all samples, which is redundant for easy samples with excessive computational cost. To address these issues, we propose a Trustworthy Long-tailed Classification (TLC) method to jointly conduct classification and uncertainty estimation to identify hard samples in a multi-expert framework. Our TLC obtains the evidence-based uncertainty (EvU) and ev-idence for each expert, and then combines these uncer-tainties and evidences under the Dempster-Shafer Evidence Theory (DST). Moreover, we propose a dynamic expert en-gagement to reduce the number of engaged experts for easy samples and achieve efficiency while maintaining promising performances. Finally, we conduct comprehensive ex-periments on the tasks of classification, tail detection, OOD detection and failure prediction. The experimental results show that the proposed TLC outperforms existing methods and is trustworthy with reliable uncertainty. Bolian Li, Zongbo Han, Haining Li, Huazhu Fu, Changqing Zhang 0002 |
CVPR | 5 |
| 2022 | UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware MixupabstractSubpopulation shift widely exists in many real-world machine learning applications, referring to the training and test distributions containing the same subpopulation groups but varying in subpopulation frequencies. Importance reweighting is a normal way to handle the subpopulation shift issue by imposing constant or adaptive sampling weights on each sample in the training dataset. However, some recent studies have recognized that most of these approaches fail to improve the performance over empirical risk minimization especially when applied to over-parameterized neural networks. In this work, we propose a simple yet practical framework, called uncertainty-aware mixup (UMIX), to mitigate the overfitting issue in over-parameterized models by reweighting the ''mixed'' samples according to the sample uncertainty. The training-trajectories-based uncertainty estimation is equipped in the proposed UMIX for each sample to flexibly characterize the subpopulation distribution. We also provide insightful theoretical analysis to verify that UMIX achieves better generalization bounds over prior works. Further, we conduct extensive empirical studies across a wide range of tasks to validate the effectiveness of our method both qualitatively and quantitatively. Code is available at https://github.com/TencentAILabHealthcare/UMIX. Zongbo Han, Fan Yang 0081, Liu Liu 0014, Lanqing Li, Yatao Bian, Peilin Zhao, Bingzhe Wu, Changqing Zhang 0002, Jianhua Yao 0001 |
NeurIPS | 9 |
| 2022 | Multi-modal sequence learning for Alzheimer's disease progression prediction with incomplete variable-length longitudinal data
Lei Xu 0028, Chunming He, Jun Wang 0024, Changqing Zhang 0002, Feiping Nie 0001, Lei Chen 0011 |
Medical Image Anal. | 5 |
| 2022 | Ball $k$k-Means: Fast Adaptive Clustering With No BoundsabstractThis paper presents a novel accelerated exact k-means called as "Ball k-means" by using the ball to describe each cluster, which focus on reducing the point-centroid distance computation. The "Ball k-means" can exactly find its neighbor clusters for each cluster, resulting distance computations only between a point and its neighbor clusters' centroids instead of all centroids. What's more, each cluster can be divided into "stable area" and "active area", and the latter one is further divided into some exact "annular area". The assignment of the points in the "stable area" is not changed while the points in each "annular area" will be adjusted within a few neighbor clusters. There are no upper or lower bounds in the whole process. Moreover, ball k-means uses ball clusters and neighbor searching along with multiple novel stratagems for reducing centroid distance computations. In comparison with the current state-of-the art accelerated exact bounded methods, the Yinyang algorithm and the Exponion algorithm, as well as other top-of-the-line tree-based and bounded methods, the ball k-means attains both higher performance and performs fewer distance calculations, especially for large-k problems. The faster speed, no extra parameters and simpler design of "Ball k-means" make it an all-around replacement of the naive k-means. Shuyin Xia, Daowan Peng, Deyu Meng, Changqing Zhang 0002, Guoyin Wang 0001, Elisabeth Giem, Wei Wei 0006, Zizhong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Deep Partial Multi-View LearningabstractAlthough multi-view learning has made significant progress over the past few decades, it is still challenging due to the difficulty in modeling complex correlations among different views, especially under the context of view missing. To address the challenge, we propose a novel framework termed Cross Partial Multi-View Networks (CPM-Nets), which aims to fully and flexibly take advantage of multiple partial views. We first provide a formal definition of completeness and versatility for multi-view representation and then theoretically prove the versatility of the learned latent representations. For completeness, the task of learning latent multi-view representation is specifically translated to a degradation process by mimicking data transmission, such that the optimal tradeoff between consistency and complementarity across different views can be implicitly achieved. Equipped with adversarial strategy, our model stably imputes missing views, encoding information from all views for each sample to be encoded into latent representation to further enhance the completeness. Furthermore, a nonparametric classification loss is introduced to produce structured representations and prevent overfitting, which endows the algorithm with promising generalization under view-missing cases. Extensive experimental results validate the effectiveness of our algorithm over existing state of the arts for classification, representation learning and data imputation. Changqing Zhang 0002, Yajie Cui, Zongbo Han, Joey Tianyi Zhou, Huazhu Fu, Qinghua Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Deep-LIFT: Deep Label-Specific Feature Learning for Image AnnotationabstractImage annotation aims to jointly predict multiple tags for an image. Although significant progress has been achieved, existing approaches usually overlook aligning specific labels and their corresponding regions due to the weak supervised information (i.e., "bag of labels" for regions), thus failing to explicitly exploit the discrimination from different classes. In this article, we propose the deep label-specific feature (Deep-LIFT) learning model to build the explicit and exact correspondence between the label and the local visual region, which improves the effectiveness of feature learning and enhances the interpretability of the model itself. Deep-LIFT extracts features for each label by aligning each label and its region. Specifically, Deep-LIFTs are achieved through learning multiple correlation maps between image convolutional features and label embeddings. Moreover, we construct two variant graph convolutional networks (GCNs) to further capture the interdependency among labels. Empirical studies on benchmark datasets validate that the proposed model achieves superior performance on multilabel classification over other existing state-of-the-art methods. Junbing Li, Changqing Zhang 0002, Joey Tianyi Zhou, Huazhu Fu, Shuyin Xia, Qinghua Hu |
IEEE Trans. Cybern. | 2 |
| 2022 | One-Step Multiview Subspace Segmentation via Joint Skinny Tensor Learning and Latent ClusteringabstractMultiview subspace clustering (MSC) has attracted growing attention due to the extensive value in various applications, such as natural language processing, face recognition, and time-series analysis. In this article, we are devoted to address two crucial issues in MSC: 1) high computational cost and 2) cumbersome multistage clustering. Existing MSC approaches, including tensor singular value decomposition (t-SVD)-MSC that has achieved promising performance, generally utilize the dataset itself as the dictionary and regard representation learning and clustering process as two separate parts, thus leading to the high computational overhead and unsatisfactory clustering performance. To remedy these two issues, we propose a novel MSC model called joint skinny tensor learning and latent clustering (JSTC), which can learn high-order skinny tensor representations and corresponding latent clustering assignments simultaneously. Through such a joint optimization strategy, the multiview complementary information and latent clustering structure can be exploited thoroughly to improve the clustering performance. An alternating direction minimization algorithm, which owns low computational complexity and can be run in parallel when solving several key subproblems, is carefully designed to optimize the JSTC model. Such a nice property makes our JSTC an appealing solution for large-scale MSC problems. We conduct extensive experiments on ten popular datasets and compare our JSTC with 12 competitors. Five commonly used metrics, including four external measures (NMI, ACC, F-score, and RI) and one internal metric (SI), are adopted to evaluate the clustering quality. The experimental results with the Wilcoxon statistical test demonstrate the superiority of the proposed method in both clustering performance and operational efficiency. Yongqiang Tang, Yuan Xie 0006, Changqing Zhang 0002, Zhizhong Zhang 0001, Wensheng Zhang 0002 |
IEEE Trans. Cybern. | 3 |
| 2021 | Uncertainty-Aware Multi-View Representation LearningabstractLearning from different data views by exploring the underlying complementary information among them can endow the representation with stronger expressive ability. However, high-dimensional features tend to contain noise, and furthermore, quality of data usually varies for different samples (even for different views), i.e., one view may be informative for one sample but not the case for another. Therefore, it is quite challenging to integrate multi-view noisy data under unsupervised setting. Traditional multi-view methods either simply treat each view with equal importance or tune the weights of different views to fixed values, which are insufficient to capture the dynamic noise in multi-view data. In this work, we devise a novel unsupervised multi-view learning approach, termed as Dynamic Uncertainty-Aware Networks (DUA-Nets). Guided by the uncertainty of data estimated from the generation perspective, intrinsic information from multiple views is integrated to obtain noise-free representations. Under the help of uncertainty estimation, DUA-Nets weigh each view of individual sample according to data quality so that the high-quality samples (or views) can be fully exploited while the effects from the noisy samples (or views) will be alleviated. Our model achieves superior performance in extensive experiments and shows the robustness to noisy data. Zongbo Han, Changqing Zhang 0002, Qinghua Hu |
AAAI | 3 |
| 2021 | Multi-View Information-Bottleneck Representation LearningabstractIn real-world applications, clustering or classification can usually be improved by fusing information from different views. Therefore, unsupervised representation learning on multi-view data becomes a compelling topic in machine learning. In this paper, we propose a novel and flexible unsupervised multi-view representation learning model termed Collaborative Multi-View Information Bottleneck Networks (CMIB-Nets), which comprehensively explores the common latent structure and the view-specific intrinsic information, and discards the superfluous information in the data significantly improving the generalization capability of the model. Specifically, our proposed model relies on the information bottleneck principle to integrate the shared representation among different views and the view-specific representation of each view, prompting the multi-view complete representation and flexibly balancing the complementarity and consistency among multiple views. We conduct extensive experiments (including clustering analysis, robustness experiment, and ablation study) on real-world datasets, which empirically show promising generalization ability and robustness compared to state-of-the-arts. Zhibin Wan, Changqing Zhang 0002, Pengfei Zhu 0001, Qinghua Hu |
AAAI | 2 |
| 2021 | Trusted Multi-View Classification
Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou |
ICLR | 2 |
| 2021 | Cross-View Equivariant Auto-EncoderabstractUnsupervised representation learning on multi-view data (multiple types of features or modalities) becomes a compelling topic in machine learning. Most existing methods focus on directly projecting different views into a common space to explore the consistency across different views. Al-though simple, the underlying relationships among different views are not guaranteed during the learning process. In this paper, we propose a novel unsupervised multi-view representation learning model termed as Cross-View Equivariant Auto-Encoder (CVE-AE), which jointly conducts data re-construction with view-specific autoencoder for information preservation within each view, and transformation reconstruction with transformation decoder for correlations preservation across different views. Accordingly, the generalization ability of our model is promoted due to the preserved intra-view intrinsic information and underlying inter-view relationships. We conduct extensive experiments on real-world datasets, and the proposed model achieves superior performance over state-of-the-art unsupervised representation learning methods. Zhibin Wan, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Pengfei Zhu 0001, Qinghua Hu |
ICME | 2 |
| 2021 | Trustworthy Multimodal Regression with Mixture of Normal-inverse Gamma DistributionsabstractMultimodal regression is a fundamental task, which integrates the information from different sources to improve the performance of follow-up applications. However, existing methods mainly focus on improving the performance and often ignore the confidence of prediction for diverse situations. In this study, we are devoted to trustworthy multimodal regression which is critical in cost-sensitive domains. To this end, we introduce a novel Mixture of Normal-Inverse Gamma distributions (MoNIG) algorithm, which efficiently estimates uncertainty in principle for adaptive integration of different modalities and produces a trustworthy regression result. Our model can be dynamically aware of uncertainty for each modality, and also robust for corrupted modalities. Furthermore, the proposed MoNIG ensures explicitly representation of (modality-specific/global) epistemic and aleatoric uncertainties, respectively. Experimental results on both synthetic and different real-world data demonstrate the effectiveness and trustworthiness of our method on various multimodal regression tasks (e.g., temperature prediction for superconductivity, relative location prediction for CT slices, and multimodal sentiment analysis). Huan Ma 0006, Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
NeurIPS | 3 |
| 2021 | Incomplete multi-modal representation learning for Alzheimer's disease diagnosis
Yanbei Liu, Lianxi Fan, Changqing Zhang 0002, Tao Zhou 0002, Zhitao Xiao, Lei Geng, Dinggang Shen |
Medical Image Anal. | 3 |
| 2021 | Deep Spectral Representation Learning From Multi-View DataabstractMulti-view representation learning (MvRL) aims to learn a consensus representation from diverse sources or domains to facilitate downstream tasks such as clustering, retrieval, and classification. Due to the limited representative capacity of the adopted shallow models, most existing MvRL methods may yield unsatisfactory results, especially when the labels of data are unavailable. To enjoy the representative capacity of deep learning, this paper proposes a novel multi-view unsupervised representation learning method, termed as Multi-view Laplacian Network (MvLNet), which could be the first deep version of the multi-view spectral representation learning method. Note that, such an attempt is nontrivial because simply combining Laplacian embedding (i.e., spectral representation) with neural networks will lead to trivial solutions. To solve this problem, MvLNet enforces an orthogonal constraint and reformulates it as a layer with the help of Cholesky decomposition. The orthogonal layer is stacked on the embedding network so that a common space could be learned for consensus representation. Compared with numerous recent-proposed approaches, extensive experiments on seven challenging datasets demonstrate the effectiveness of our method in three multi-view tasks including clustering, recognition, and retrieval. The source code could be found at www.pengxi.me. Zhenyu Huang 0005, Joey Tianyi Zhou, Hongyuan Zhu 0002, Changqing Zhang 0002, Jiancheng Lv 0001, Xi Peng 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | SPL-MLL: Selecting Predictable Landmarks for Multi-label Learning
Junbing Li, Changqing Zhang 0002, Pengfei Zhu 0001, Baoyuan Wu, Lei Chen 0011, Qinghua Hu |
ECCV (9) | 2 |
| 2020 | Multi-Scale Cross-Modal Spatial Attention Fusion for Multi-label Image Recognition
Junbing Li, Changqing Zhang 0002, Xueman Wang |
ICANN (1) | 2 |
| 2020 | Progressive Point To Set Metric Learning For Semi-Supervised Few-Shot ClassificationabstractFew-shot learning aims to learn models that can generalize to unseen tasks from very few annotated samples of available tasks. The performance of few-shot learning is greatly affected by the number of samples per class. The massive unlabeled data can help to boost the performance of few shot learning models. In this paper, we propose a novel progressive point to set metric learning (PPSML) model for semisupervised few-shot classification. The distance metric is defined for an image of the query set to a class of the support set by point to set distance. A self-training strategy is designed to select the samples locally or globally with high confidence and use these samples to progressively update the point to set distance. Experiments on benchmark datasets show that our proposed PPSML significantly improves the accuracy of few shot classification and outperforms the state-of-the-art semisupervised few-shot learning methods. Pengfei Zhu 0001, Mingqi Gu, Changqing Zhang 0002, Qinghua Hu |
ICIP | 4 |
| 2020 | SAAN: Semantic Attention Adaptation Network for Face Super-ResolutionabstractFace super-resolution (Face SR) is a sub-domain of SR that reconstructs high-resolution face images from low-resolution ones. The prior knowledge of face is widely used for recovering more realistic facial details, which will increase the complexity of the network and introduce additional knowledge extraction procession both in the training and evaluating stage. To address the above issues, we propose to combine face semantic prior extraction and face SR with the attention adaptation model and design a Semantic Attention Adaptation Network (SAAN) for face SR. Specifically, we train the face semantic parsing network and face SR network jointly, by adopting the semantic attention adaptation (SAA) model to transfer the ability of extracting face prior knowledge to the SR network. Then our SR network can work independently in the testing stage without using the prior knowledge extraction network. To generate realistic face images, we also utilize GAN loss to enrich the texture with more details (i.e. SAAN-G). Extensive experiments on the benchmark dataset illustrate that our SAAN and SAAN-G improve the state-of-the-art both on quality and efficiency. Changqing Zhang 0002 |
ICME | 2 |
| 2020 | M2 Net: Multi-modal Multi-channel Network for Overall Survival Time Prediction of Brain Tumor Patients
Tao Zhou 0002, Huazhu Fu, Yu Zhang 0009, Changqing Zhang 0002, Xiankai Lu, Jianbing Shen, Ling Shao 0001 |
MICCAI (2) | 4 |
| 2020 | Tensorized Multi-view Subspace Representation Learning
Changqing Zhang 0002, Huazhu Fu, Jing Wang 0023, Wen Li 0001, Xiaochun Cao, Qinghua Hu |
Int. J. Comput. Vis. | 1 |
| 2020 | Multi-modal latent space inducing ensemble SVM classifier for early dementia diagnosis with neuroimaging data
Tao Zhou 0002, Kim-Han Thung, Mingxia Liu 0001, Feng Shi 0001, Changqing Zhang 0002, Dinggang Shen |
Medical Image Anal. | 5 |
| 2020 | Generalized Latent Multi-View Subspace ClusteringabstractSubspace clustering is an effective method that has been successfully applied to many applications. Here, we propose a novel subspace clustering model for multi-view data using a latent representation termed Latent Multi-View Subspace Clustering (LMSC). Unlike most existing single-view subspace clustering methods, which directly reconstruct data points using original features, our method explores underlying complementary information from multiple views and simultaneously seeks the underlying latent representation. Using the complementarity of multiple views, the latent representation depicts data more comprehensively than each individual view, accordingly making subspace representation more accurate and robust. We proposed two LMSC formulations: linear LMSC (lLMSC), based on linear correlations between latent representation and each view, and generalized LMSC (gLMSC), based on neural networks to handle general relationships. The proposed method can be efficiently optimized under the Augmented Lagrangian Multiplier with Alternating Direction Minimization (ALM-ADM) framework. Extensive experiments on diverse datasets demonstrate the effectiveness of the proposed method. Changqing Zhang 0002, Huazhu Fu, Qinghua Hu, Xiaochun Cao, Yuan Xie 0006, Dacheng Tao, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Unsupervised Video Action Clustering via Motion-Scene Interaction ConstraintabstractIn the past few years, scene contextual information has been increasingly used for action understanding with promising results. However, unsupervised video action clustering using context has been less explored, and existing clustering methods cannot achieve satisfactory performances. In this paper, we propose a novel unsupervised video action clustering method by using the motion-scene interaction constraint (MSIC). The proposed method takes the unique static scene and dynamic motion characteristics of video action into account, and develops a contextual interaction constraint model under a self-representation subspace clustering framework. First, the complementarity of multi-view subspace representation in each context is explored by single-view and multi-view constraints. Afterward, the context-constrained affinity matrix is calculated and the MSIC is introduced to mutually regularize the disagreement of subspace representation in scene and motion. Finally, by jointly constraining the complementarity of multi-views and the consistency of multi-contexts, an overall objective function is constructed to guarantee the video action clustering result. The experiments on four video benchmark datasets (Weizmann, KTH, UCFsports, and Olympic) demonstrate that the proposed method outperforms the state-of-the-art methods. Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Changqing Zhang 0002, Tat-Seng Chua, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Hybrid Noise-Oriented Multilabel LearningabstractFor real-world applications, multilabel learning usually suffers from unsatisfactory training data. Typically, features may be corrupted or class labels may be noisy or both. Ignoring noise in the learning process tends to result in an unreasonable model and, thus, inaccurate prediction. Most existing methods only consider either feature noise or label noise in multilabel learning. In this paper, we propose a unified robust multilabel learning framework for data with hybrid noise, that is, both feature noise and label noise. The proposed method, hybrid noise-oriented multilabel learning (HNOML), is simple but rather robust for noisy data. HNOML simultaneously addresses feature and label noise by bi-sparsity regularization bridged with label enrichment. Specifically, the label enrichment matrix explores the underlying correlation among different classes which improves the noisy labeling. Bridged with the enriching label matrix, the structured sparsity is imposed to jointly handle the corrupted features and noisy labeling. We utilize the alternating direction method (ADM) to efficiently solve our problem. Experimental results on several benchmark datasets demonstrate the advantages of our method over the state-of-the-art ones. Changqing Zhang 0002, Ziwei Yu, Huazhu Fu, Pengfei Zhu 0001, Lei Chen 0011, Qinghua Hu |
IEEE Trans. Cybern. | 1 |
| 2020 | Multiview Latent Space Learning With Feature Redundancy MinimizationabstractMultiview learning has received extensive research interest and has demonstrated promising results in recent years. Despite the progress made, there are two significant challenges within multiview learning. First, some of the existing methods directly use original features to reconstruct data points without considering the issue of feature redundancy. Second, existing methods cannot fully exploit the complementary information across multiple views and meanwhile preserve the view-specific properties; therefore, the degraded learning performance will be generated. To address the above issues, we propose a novel multiview latent space learning framework with feature redundancy minimization. We aim to learn a latent space to mitigate the feature redundancy and use the learned representation to reconstruct every original data point. More specifically, we first project the original features from multiple views onto a latent space, and then learn a shared dictionary and view-specific dictionaries to, respectively, exploit the correlations across multiple views as well as preserve the view-specific properties. Furthermore, the Hilbert-Schmidt independence criterion is adopted as a diversity constraint to explore the complementarity of multiview representations, which further ensures the diversity from multiple views and preserves the local structure of the data in each view. Experimental results on six public datasets have demonstrated the effectiveness of our multiview learning approach against other state-of-the-art methods. Tao Zhou 0002, Changqing Zhang 0002, Chen Gong 0002, Harish Bhaskar, Jie Yang 0002 |
IEEE Trans. Cybern. | 2 |
| 2020 | Dual Shared-Specific Multiview Subspace ClusteringabstractMultiview subspace clustering has received significant attention as the availability of diverse of multidomain and multiview real-world data has rapidly increased in the recent years. Boosting the performance of multiview clustering algorithms is challenged by two major factors. First, since original features from multiview data are highly redundant, reconstruction based on these attributes inevitably results in inferior performance. Second, since each view of such multiview data may contain unique knowledge as against the others, it remains a challenge to exploit complimentary information across multiple views while simultaneously investigating the uniqueness of each view. In this paper, we present a novel dual shared-specific multiview subspace clustering (DSS-MSC) approach that simultaneously learns the correlations between shared information across multiple views and also utilizes view-specific information to depict specific property for each independent view. Further, we formulate a dual learning framework to capture shared-specific information into the dimensional reduction and self-representation processes, which strengthens the ability of our approach to exploit shared information while preserving view-specific property effectively. The experimental results on several benchmark datasets have demonstrated the effectiveness of the proposed approach against other state-of-the-art techniques. Tao Zhou 0002, Changqing Zhang 0002, Xi Peng 0001, Harish Bhaskar, Jie Yang 0002 |
IEEE Trans. Cybern. | 2 |
| 2020 | Diagnosis of Coronavirus Disease 2019 (COVID-19) With Structured Latent Multi-View Representation LearningabstractRecently, the outbreak of Coronavirus Disease 2019 (COVID-19) has spread rapidly across the world. Due to the large number of infected patients and heavy labor for doctors, computer-aided diagnosis with machine learning algorithm is urgently needed, and could largely reduce the efforts of clinicians and accelerate the diagnosis process. Chest computed tomography (CT) has been recognized as an informative tool for diagnosis of the disease. In this study, we propose to conduct the diagnosis of COVID-19 with a series of features extracted from CT images. To fully explore multiple features describing CT images from different views, a unified latent representation is learned which can completely encode information from different aspects of features and is endowed with promising class structure for separability. Specifically, the completeness is guaranteed with a group of backward neural networks (each for one type of features), while by using class labels the representation is enforced to be compact within COVID-19/community-acquired pneumonia (CAP) and also a large margin is guaranteed between different types of pneumonia. In this way, our model can well avoid overfitting compared to the case of directly projecting high-dimensional features into classes. Extensive experimental results show that the proposed method outperforms all comparison methods, and rather stable performances are observed when varying the number of training data. Hengyuan Kang, Liming Xia, Fuhua Yan, Zhibin Wan, Feng Shi 0001, Huan Yuan, Huiting Jiang, Dijia Wu, He Sui, Changqing Zhang 0002, Dinggang Shen |
IEEE Trans. Medical Imaging | 10 |
| 2019 | Facial Emotion Distribution Learning by Exploiting Low-Rank Label Correlations LocallyabstractEmotion recognition from facial expressions is an interesting and challenging problem and has attracted much attention in recent years. Substantial previous research has only been able to address the ambiguity of “what describes the expression”, which assumes that each facial expression is associated with one or more predefined affective labels while ignoring the fact that multiple emotions always have different intensities in a single picture. Therefore, to depict facial expressions more accurately, this paper adopts a label distribution learning approach for emotion recognition that can address the ambiguity of “how to describe the expression” and proposes an emotion distribution learning method that exploits label correlations locally. Moreover, a local low-rank structure is employed to capture the local label correlations implicitly. Experiments on benchmark facial expression datasets demonstrate that our method can better address the emotion distribution recognition problem than state-of-the-art methods. Xiuyi Jia, Weiwei Li 0001, Changqing Zhang 0002, Zechao Li |
CVPR | 4 |
| 2019 | AE2-Nets: Autoencoder in Autoencoder NetworksabstractLearning on data represented with multiple views (e.g., multiple types of descriptors or modalities) is a rapidly growing direction in machine learning and computer vision. Although effectiveness achieved, most existing algorithms usually focus on classification or clustering tasks. Differently, in this paper, we focus on unsupervised representation learning and propose a novel framework termed Autoencoder in Autoencoder Networks (AE2-Nets), which integrates information from heterogeneous sources into an intact representation by the nested autoencoder framework. The proposed method has the following merits: (1) our model jointly performs view-specific representation learning (with the inner autoencoder networks) and multi-view information encoding (with the outer autoencoder networks) in a unified framework; (2) due to the degradation process from the latent representation to each single view, our model flexibly balances the complementarity and consistence among multiple views. The proposed model is efficiently solved by the alternating direction method (ADM), and demonstrates the effectiveness compared with state-of-the-art algorithms. Changqing Zhang 0002, Yeqing Liu, Huazhu Fu |
CVPR | 1 |
| 2019 | Reciprocal Multi-Layer Subspace Learning for Multi-View ClusteringabstractMulti-view clustering is a long-standing important research topic, however, remains challenging when handling high-dimensional data and simultaneously exploring the consistency and complementarity of different views. In this work, we present a novel Reciprocal Multi-layer Subspace Learning (RMSL) algorithm for multi-view clustering, which is composed of two main components: Hierarchical Self-Representative Layers (HSRL), and Backward Encoding Networks (BEN). Specifically, HSRL constructs reciprocal multi-layer subspace representations linked with a latent representation to hierarchically recover the underlying low-dimensional subspaces in which the high-dimensional data lie; BEN explores complex relationships among different views and implicitly enforces the subspaces of all views to be consistent with each other and more separable. The latent representation flexibly encodes complementary information from multiple views and depicts data more comprehensively. Our model can be efficiently optimized by an alternating optimization scheme. Extensive experiments on benchmark datasets show the superiority of RMSL over other state-of-the-art clustering methods. Ruihuang Li, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Joey Tianyi Zhou, Qinghua Hu |
ICCV | 2 |
| 2019 | Multi-view Spectral Clustering NetworkabstractMulti-view clustering aims to cluster data from diverse sources or domains, which has drawn considerable attention in recent years. In this paper, we propose a novel multi-view clustering method named multi-view spectral clustering network (MvSCN) which could be the first deep version of multi-view spectral clustering to the best of our knowledge. To deeply cluster multi-view data, MvSCN incorporates the local invariance within every single view and the consistency across different views into a novel objective function, where the local invariance is defined by a deep metric learning network rather than the Euclidean distance adopted by traditional approaches. In addition, we enforce and reformulate an orthogonal constraint as a novel layer stacked on an embedding network for two advantages, i.e. jointly optimizing the neural network and performing matrix decomposition and avoiding trivial solutions. Extensive experiments on four challenging datasets demonstrate the effectiveness of our method compared with 10 state-of-the-art approaches in terms of three evaluation metrics. Zhenyu Huang 0005, Joey Tianyi Zhou, Xi Peng 0001, Changqing Zhang 0002, Hongyuan Zhu 0002, Jiancheng Lv 0001 |
IJCAI | 4 |
| 2019 | Flexible Multi-View Representation Learning for Subspace ClusteringabstractIn recent years, numerous multi-view subspace clustering methods have been proposed to exploit the complementary information from multiple views. Most of them perform data reconstruction within each single view, which makes the subspace representation unpromising and thus can not well identify the underlying relationships among data. In this paper, we propose to conduct subspace clustering based on Flexible Multi-view Representation (FMR) learning, which avoids using partial information for data reconstruction. The latent representation is flexibly constructed by enforcing it to be close to different views, which implicitly makes it more comprehensive and well-adapted to subspace clustering. With the introduction of kernel dependence measure, the latent representation can flexibly encode complementary information from different views and explore nonlinear, high-order correlations among these views. We employ the Alternating Direction Minimization (ADM) method to solve our problem. Empirical studies on real-world datasets show that our method achieves superior clustering performance over other state-of-the-art methods. Ruihuang Li, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001, Zheng Wang 0008 |
IJCAI | 2 |
| 2019 | Attributed Subspace ClusteringabstractExisting methods on representation-based subspace clustering mainly treat all features of data as a whole to learn a single self-representation and get one clustering solution. Real data however are often complex and consist of multiple attributes or sub-features, such as a face image has expressions or genders. Each attribute is distinct and complementary on depicting the data. Failing to explore attributes and capture the complementary information among them may lead to an inaccurate representation. Moreover, a single clustering solution is rather limited to depict data, which can often be interpreted from different aspects and grouped into multiple clusters according to attributes. Therefore, we propose an innovative model called attributed subspace clustering (ASC). It simultaneously learns multiple self-representations on latent representations derived from original data. By utilizing Hilbert Schmidt Independence Criterion as a co-regularizing term, ASC enforces that each self-representation is independent and corresponds to a specific attribute. A more comprehensive self-representation is then established by adding these self-representations. Experiments on several benchmark image datasets have demonstrated the effectiveness of ASC not only in terms of clustering accuracy achieved by the integrated representation, but also the diverse interpretation of data, which is beyond what current approaches can offer. Jing Wang 0023, Linchuan Xu, Feng Tian 0006, Atsushi Suzuki 0002, Changqing Zhang 0002, Kenji Yamanishi |
IJCAI | 5 |
| 2019 | Deep Distillation Metric LearningabstractDue to the emergence of large-scale and high-dimensional data, measuring the similarity between data points becomes challenging. In order to obtain effective representations, metric learning has become one of the most active researches in the field of computer vision and pattern recognition. However, models using trained networks for predictions are often cumbersome and difficult to be deployed. Therefore, in this paper, we propose a novel deep distillation metric learning (DDML) for online teaching in the procedure of learning the distance metric. Specifically, we employ model distillation to transfer the knowledge acquired by the larger model to the smaller model. Unlike the 2-step offline and mutual online manners, we propose to train a powerful teacher model, who transfer the knowledge to a lightweight and generalizable student model and iteratively improved by the feedback from the student model. We show that our method has achieved state-of-the-art results on CUB200-2011 and CARS196 while having advantages in computational efficiency. Jiaxu Han, Changqing Zhang 0002 |
MMAsia | 3 |
| 2019 | CPM-Nets: Cross Partial Multi-View NetworksabstractDespite multi-view learning progressed fast in past decades, it is still challenging due to the difficulty in modeling complex correlation among different views, especially under the context of view missing. To address the challenge, we propose a novel framework termed Cross Partial Multi-View Networks (CPM-Nets). In this framework, we first give a formal definition of completeness and versatility for multi-view representation and then theoretically prove the versatility of the latent representation learned from our algorithm. To achieve the completeness, the task of learning latent multi-view representation is specifically translated to degradation process through mimicking data transmitting, such that the optimal tradeoff between consistence and complementarity across different views could be achieved. In contrast with methods that either complete missing views or group samples according to view-missing patterns, our model fully exploits all samples and all views to produce structured representation for interpretability. Extensive experimental results validate the effectiveness of our algorithm over existing state-of-the-arts. Changqing Zhang 0002, Zongbo Han, Yajie Cui, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
NeurIPS | 1 |
| 2019 | Efficient registration of multi-view point sets by K-means clustering
Jihua Zhu, Zutao Jiang, Georgios Evangelidis 0002, Changqing Zhang 0002, Shanmin Pang, Zhongyu Li 0002 |
Inf. Sci. | 4 |
| 2019 | Multi-view subspace clustering with intactness-aware similarity
Xiaobo Wang 0001, Zhen Lei 0001, Xiaojie Guo 0001, Changqing Zhang 0002, Hailin Shi, Stan Z. Li |
Pattern Recognit. | 4 |
| 2019 | Graph Structure Fusion for Multiview ClusteringabstractMost existing multiview clustering methods take graphs, which are usually predefined independently in each view, as input to uncover data distribution. These methods ignore the correlation of graph structure among multiple views and clustering results highly depend on the quality of predefined affinity graphs. We address the problem of multiview clustering by seamlessly integrating graph structures of different views to fully exploit the geometric property of underlying data structure. The proposed method is based on the assumption that the intrinsic underlying graph structure would assign corresponding connected component in each graph to the same cluster. Different graphs from multiple views are integrated by using the Hadamard product since different views usually together admit the same underlying structure across multiple views. Specifically, these graphs are integrated into a global one and the structure of the global graph is adaptively tuned by a well-designed objective function so that the number of components of the graph is exactly equal to the number of clusters. It is worth noting that we directly obtain cluster indicators from the graph itself without performing further graph-cut or k-means clustering algorithms. Experiments show the proposed method obtains better clustering performance than the state-of-the-art methods. Kun Zhan, Chaoxi Niu, Changlu Chen, Feiping Nie 0001, Changqing Zhang 0002, Yi Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Infant Brain Development Prediction With Latent Partial Multi-View Representation LearningabstractThe early postnatal period witnesses rapid and dynamic brain development. However, the relationship between brain anatomical structure and cognitive ability is still unknown. Currently, there is no explicit model to characterize this relationship in the literature. In this paper, we explore this relationship by investigating the mapping between morphological features of the cerebral cortex and cognitive scores. To this end, we introduce a multi-view multi-task learning approach to intuitively explore complementary information from different time-points and handle the missing data issue in longitudinal studies simultaneously. Accordingly, we establish a novel model, latent partial multi-view representation learning. Our approach regards data from different time-points as different views and constructs a latent representation to capture the complementary information from incomplete time-points. The latent representation explores the complementarity across different time-points and improves the accuracy of prediction. The minimization problem is solved by the alternating direction method of multipliers. Experimental results on both synthetic and real data validate the effectiveness of our proposed algorithm. Changqing Zhang 0002, Ehsan Adeli-Mosabbeb, Zhengwang Wu, Gang Li 0001, Weili Lin, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Adaptive Hypergraph Embedded Semi-Supervised Multi-Label Image AnnotationabstractMultilabel image annotation attracts a lot of research interest due to its practicability in multimedia and computer vision fields, while the need for a large amount of labeled training data to achieve promising performance makes it a challenging task. Fortunately, unlabeled and relevant data are widely available and these data can be used to serve the annotation task. To this end, we propose a novel adaptive hypergraph learning (AHL) method for multilabel image annotation in a semisupervised way, in which both the limited labeled data and abundant unlabeled data are utilized to facilitate the annotation performance. In detail, we seek a multilabel propagation scheme by learning a hypergraph which is used to preserve the local geometric structures of data in a high-order manner. Meanwhile, a feature projection is integrated into AHL to obtain a latent feature space where unlabeled instances can be effectively and robustly assigned with multiple labels. Experiments on six widely used image datasets are conducted to evaluate our model and the results demonstrate that the proposed AHL outperforms other state-of-the-art semisupervised methods. Chang Tang, Xinwang Liu 0002, Pichao Wang, Changqing Zhang 0002, Miaomiao Li 0001, Lizhe Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Learning a Joint Affinity Graph for Multiview Subspace ClusteringabstractWith the ability to exploit the internal structure of data, graph-based models have received a lot of attention and have achieved great success in multiview subspace clustering for multimedia data. Most of the existing methods individually construct an affinity graph for each single view and fuse the result obtained from each single graph. However, the common representation shared by different views and the complementary diversity across these views are not efficiently exploited. In addition, noise and outliers are often mixed in original data, which adversely degenerate the clustering performance of many existing methods. In this paper, we propose addressing these issues by learning a joint affinity graph for multiview subspace clustering based on a low-rank representation with diversity regularization and a rank constraint. Specifically, a low-rank representation model is employed to learn a shared sample representation coefficient matrix to generate the affinity graph. At the same time, we use diversity regularization to learn the optimal weights for each view, which can suppress the redundancy and enhance the diversity among different feature views. In addition, the cluster number is used to promote affinity graph learning by using a rank constraint. The final clustering result is obtained by using normalized cuts on the learned affinity graph. An efficient algorithm based on an augmented Lagrangian multiplier with alternating direction minimization is carefully designed to solve the resulting optimization problem. Extensive experiments on various real-world datasets are conducted, and the results demonstrate well the effectiveness of the proposed algorithm. Chang Tang, Xinzhong Zhu, Xinwang Liu 0002, Miaomiao Li 0001, Pichao Wang, Changqing Zhang 0002, Lizhe Wang 0001 |
IEEE Trans. Multim. | 6 |
| 2018 | Consistent and Specific Multi-View Subspace ClusteringabstractMulti-view clustering has attracted intensive attention due to the effectiveness of exploiting multiple views of data. However, most existing multi-view clustering methods only aim to explore the consistency or enhance the diversity of different views. In this paper, we propose a novel multi-view subspace clustering method (CSMSC), where consistency and specificity are jointly exploited for subspace representation learning. We formulate the multi-view self-representation property using a shared consistent representation and a set of specific representations, which better fits the real-world datasets. Specifically, consistency models the common properties among all views, while specificity captures the inherent difference in each view. In addition, to optimize the non-convex problem, we introduce a convex relaxation and develop an alternating optimization algorithm to recover the corresponding data representations. Experimental evaluations on four benchmark datasets demonstrate that the proposed approach achieves better performance over several state-of-the-arts. Shirui Luo, Changqing Zhang 0002, Wei Zhang 0031, Xiaochun Cao |
AAAI | 2 |
| 2018 | Multi-Layer Multi-View Classification for Alzheimer's Disease DiagnosisabstractIn this paper, we propose a novel multi-view learning method for Alzheimer's Disease (AD) diagnosis, using neuroimaging and genetics data. Generally, there are several major challenges associated with traditional classification methods on multi-source imaging and genetics data. First, the correlation between the extracted imaging features and class labels is generally complex, which often makes the traditional linear models ineffective. Second, medical data may be collected from different sources (i.e., multiple modalities of neuroimaging data, clinical scores or genetics measurements), therefore, how to effectively exploit the complementarity among multiple views is of great importance. In this paper, we propose a Multi-Layer Multi-View Classification (ML-MVC) approach, which regards the multi-view input as the first layer, and constructs a latent representation to explore the complex correlation between the features and class labels. This captures the high-order complementarity among different views, as we exploit the underlying information with a low-rank tensor regularization. Intrinsically, our formulation elegantly explores the nonlinear correlation together with complementarity among different views, and thus improves the accuracy of classification. Finally, the minimization problem is solved by the Alternating Direction Method of Multipliers (ADMM). Experimental results on Alzheimer's Disease Neuroimaging Initiative (ADNI) data sets validate the effectiveness of our proposed method. Changqing Zhang 0002, Ehsan Adeli-Mosabbeb, Tao Zhou 0002, Xiaobo Chen 0001, Dinggang Shen |
AAAI | 1 |
| 2018 | Latent Semantic Aware Multi-View Multi-Label ClassificationabstractFor real-world applications, data are often associated with multiple labels and represented with multiple views. Most existing multi-label learning methods do not sufficiently consider the complementary information among multiple views, leading to unsatisfying performance. To address this issue, we propose a novel approach for multi-view multi-label learning based on matrix factorization to exploit complementarity among different views. Specifically, under the assumption that there exists a common representation across different views, the uncovered latent patterns are enforced to be aligned across different views in kernel spaces. In this way, the latent semantic patterns underlying in data could be well uncovered and this enhances the reasonability of the common representation of multiple views. As a result, the consensus multi-view representation is obtained which encodes the complementarity and consistence of different views in latent semantic space. We provide theoretical guarantee for the strict convexity for our method by properly setting parameters. Empirical evidence shows the clear advantages of our method over the state-of-the-art ones. Changqing Zhang 0002, Ziwei Yu, Qinghua Hu, Pengfei Zhu 0001, Xinwang Liu 0002, Xiaobo Wang 0001 |
AAAI | 1 |
| 2018 | Generalized Multi-view Unsupervised Feature Selection
Yue Liu 0008, Changqing Zhang 0002, Pengfei Zhu 0001, Qinghua Hu |
ICANN (2) | 2 |
| 2018 | Support Vector Metric Learning on Symmetric Positive Definite ManifoldabstractThe manifold of symmetric positive definite (SPD) matrices has drawn significant attention because of its widespread applications. SPD matrices provide compact nonlinear representations of data and form a special type of Riemannian manifold. The direct application of support vector machines on SPD manifold maybe fails due to lack of samples per class. In this paper, we propose a support vector metric learning (SVML) model on SPD manifold. We define a positive definite kernel for point pairs on SPD manifold and transform metric learning on SPD manifold to a point pair classification problem. The metric learning problem can be efficiently solved by standard support vector machines. Compared with classifying points on SPD manifold by support vector machines directly, SVML effectively learns a distance metric for SPD matrices by training a binary support vector machine model. Experiments on video based face recognition, image set classification, and material classification show that SVML outperforms the state-of-the-art metric learning algorithms on SPD manifold. Hao Cheng 0010, Pengfei Zhu 0001, Qilong Wang 0001, Changqing Zhang 0002, Qinghua Hu |
ICME | 4 |
| 2018 | Ensemble of Label Specific Features for Multi-Label ClassificationabstractIn this paper, we focus on multi-label classification which associates one instance with multiple labels. The approach Label-specIfic FeaTures (LIFT) achieves state-of-the-art performance due to the label-specific features. However, the main limitation of LIFT is the poor local optima in k-means used at training stage. For this issue, in this paper, we propose to mitigate the limitation for high classification accuracy with ensemble way and term our approach as Ensemble of Label specIfic FeaTures (ELIFT). Specifically, our approach firstly constructs multiple LIFT classifiers by using multiple training sets generated by bagging strategy. Furthermore, different classifiers are weighted automatically according to the loss of each classifier. Finally, for each new instance, the predicted label vector is obtained by the weighted ensemble classifiers learned. Experiments conducted on five benchmark datasets demonstrate the performance of the proposed method outperforms the state-of-the-art approaches. Xiaoya Wei, Ziwei Yu, Changqing Zhang 0002, Qinghua Hu |
ICME | 3 |
| 2018 | FISH-MML: Fisher-HSIC Multi-View Metric LearningabstractThis work presents a simple yet effective model for multi-view metric learning, which aims to improve the classification of data with multiple views, e.g., multiple modalities or multiple types of features. The intrinsic correlation, different views describing same set of instances, makes it possible and necessary to jointly learn multiple metrics of different views, accordingly, we propose a multi-view metric learning method based on Fisher discriminant analysis (FDA) and Hilbert-Schmidt Independence Criteria (HSIC), termed as Fisher-HSIC Multi-View Metric Learning (FISH-MML). In our approach, the class separability is enforced in the spirit of FDA within each single view, while the consistence among different views is enhanced based on HSIC. Accordingly, both intra-view class separability and inter-view correlation are well addressed in a unified framework. The learned metrics can improve multi-view classification, and experimental results on real-world datasets demonstrate the effectiveness of the proposed method. Changqing Zhang 0002, Yeqing Liu, Yue Liu 0008, Qinghua Hu, Xinwang Liu 0002, Pengfei Zhu 0001 |
IJCAI | 1 |
| 2018 | Towards Generalized and Efficient Metric Learning on Riemannian ManifoldabstractModeling data as points on non-linear Riemannian manifold has attracted increasing attentions in many computer vision tasks, especially visual recognition. Learning an appropriate metric on Riemannian manifold plays a key role in achieving promising performance. For widely used symmetric positive definite (SPD) manifold and Grassmann manifold, most of existing metric learning methods are designed for one manifold, and are not straightforward for the other one. Furthermore, optimizations in previous methods usually rely on computationally expensive iterations. To address above limitations, this paper makes an attempt to propose a generalized and efficient Riemannian manifold metric learning (RMML) method, which can be flexibly adopted to both SPD and Grassmann manifolds. By minimizing the geodesic distance of similar pairs and the interpoint geodesic distance of dissimilar ones on nonlinear manifolds, the proposed RMML is optimized by computing the geodesic mean between inverse of similarity matrix and dissimilarity matrix, benefiting a global closed-form solution and high efficiency. The experiments are conducted on various visual recognition tasks, and the results demonstrate our RMML performs favorably against its counterparts in terms of both accuracy and efficiency. Pengfei Zhu 0001, Hao Cheng 0010, Qinghua Hu, Qilong Wang 0001, Changqing Zhang 0002 |
IJCAI | 5 |
| 2018 | Beyond Similar and Dissimilar Relations : A Kernel Regression Formulation for Metric LearningabstractMost existing metric learning methods focus on learning a similarity or distance measure relying on similar and dissimilar relations between sample pairs. However, pairs of samples cannot be simply identified as similar or dissimilar in many real-world applications, e.g., multi-label learning, label distribution learning or tasks with continuous decision values. To this end, in this paper we propose a novel relation alignment metric learning (RAML) formulation to handle the metric learning problem in those scenarios. Since the relation of two samples can be measured by the difference degree of the decision values, motivated by the consistency of the sample relations in the feature space and decision space, our proposed RAML utilizes the sample relations in the decision space to guide the metric learning in the feature space. Specifically, our RAML method formulates metric learning as a kernel regression problem, which can be efficiently optimized by the standard regression solvers. We carry out several experiments on the single-label classification, multi-label classification, and label distribution learning tasks, to demonstrate that our method achieves favorable performance against the state-of-the-art methods. Pengfei Zhu 0001, Ren Qi, Qinghua Hu, Qilong Wang 0001, Changqing Zhang 0002, Liu Yang 0010 |
IJCAI | 5 |
| 2018 | Latent Subspace Representation for Multiclass Classification
Changqing Zhang 0002, Xiao Wang 0017, Pengfei Zhu 0001, Zheng Wang 0008, Qinghua Hu |
PRICAI (1) | 2 |
| 2018 | Co-regularized unsupervised feature selection
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002 |
Neurocomputing | 4 |
| 2018 | Entropy-based active sparse subspace clustering
Yanbei Liu, Changqing Zhang 0002, Xiao Wang 0017, Shaona Wang, Zhitao Xiao |
Multim. Tools Appl. | 3 |
| 2018 | Multi-view label embedding
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002, Zhizhao Feng |
Pattern Recognit. | 4 |
| 2018 | Multi-label feature selection with missing labels
Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002, Hong Zhao 0002 |
Pattern Recognit. | 4 |
| 2018 | Shape-Preserving Object Depth Control for Stereoscopic ImagesabstractIn the field of 3-D technology, it is interesting as well as meaningful issue to control object depth in 3-D space. Recently, some depth control methods for stereoscopic images have been proposed, which usually employ depth map or directly process color images to implement depth control. There are two main disadvantages for these methods. First, the results of these methods usually suffer from object deformation and holes. Second, these methods are prone to cause undesired object size changing in 3-D space. To address these issues, we propose a shape-preserving object depth control method for stereoscopic images. First, a novel depth mapping model is presented for calculating the ideal coordinates of the key points in depth control, so that the shape of the object can be well preserved. Afterward, the image content-based constraints are used to further preserve the structure of the object and its background. Finally, the warping technology is introduced to deal with images optimally as well as to avoid holes. Experimental results show that the proposed method can control object depth and preserve the shape of the object effectively without sensible background distortion. Jianjun Lei 0001, Bo Peng 0007, Changqing Zhang 0002, Xuguang Mei, Xiaochun Cao, Xiaoting Fan, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Graph Learning for Multiview ClusteringabstractMost existing graph-based clustering methods need a predefined graph and their clustering performance highly depends on the quality of the graph. Aiming to improve the multiview clustering performance, a graph learning-based method is proposed to improve the quality of the graph. Initial graphs are learned from data points of different views, and the initial graphs are further optimized with a rank constraint on the Laplacian matrix. Then, these optimized graphs are integrated into a global graph with a well-designed optimization procedure. The global graph is learned by the optimization procedure with the same rank constraint on its Laplacian matrix. Because of the rank constraint, the cluster indicators are obtained directly by the global graph without performing any graph cut technique and the k-means clustering. Experiments are conducted on several benchmark datasets to verify the effectiveness and superiority of the proposed graph learning-based multiview clustering algorithm comparing to the state-of-the-art methods. Kun Zhan, Changqing Zhang 0002, Junpeng Guan |
IEEE Trans. Cybern. | 2 |
| 2018 | Disc-Aware Ensemble Network for Glaucoma Screening From Fundus ImageabstractGlaucoma is a chronic eye disease that leads to irreversible vision loss. Most of the existing automatic screening methods first segment the main structure and subsequently calculate the clinical measurement for the detection and screening of glaucoma. However, these measurement-based methods rely heavily on the segmentation accuracy and ignore various visual features. In this paper, we introduce a deep learning technique to gain additional image-relevant information and screen glaucoma from the fundus image directly. Specifically, a novel disc-aware ensemble network for automatic glaucoma screening is proposed, which integrates the deep hierarchical context of the global fundus image and the local optic disc region. Four deep streams on different levels and modules are, respectively, considered as global image stream, segmentation-guided network, local disc region stream, and disc polar transformation stream. Finally, the output probabilities of different streams are fused as the final screening result. The experiments on two glaucoma data sets (SCES and new SINDI data sets) show that our method outperforms other state-of-the-art algorithms. Huazhu Fu, Jun Cheng 0003, Yanwu Xu 0001, Changqing Zhang 0002, Damon Wing Kee Wong, Jiang Liu 0001, Xiaochun Cao |
IEEE Trans. Medical Imaging | 4 |
| 2017 | Exclusivity-Consistency Regularized Multi-view Subspace ClusteringabstractMulti-view subspace clustering aims to partition a set of multi-source data into their underlying groups. To boost the performance of multi-view clustering, numerous subspace learning algorithms have been developed in recent years, but with rare exploitation of the representation complementarity between different views as well as the indicator consistency among the representations, let alone considering them simultaneously. In this paper, we propose a novel multi-view subspace clustering model that attempts to harness the complementary information between different representations by introducing a novel position-aware exclusivity term. Meanwhile, a consistency term is employed to make these complementary representations to further have a common indicator. We formulate the above concerns into a unified optimization framework. Experimental results on several benchmark datasets are conducted to reveal the effectiveness of our algorithm over other state-of-the-arts. Xiaobo Wang 0001, Xiaojie Guo 0001, Zhen Lei 0001, Changqing Zhang 0002, Stan Z. Li |
CVPR | 4 |
| 2017 | Latent Multi-view Subspace ClusteringabstractIn this paper, we propose a novel Latent Multi-view Subspace Clustering (LMSC) method, which clusters data points with latent representation and simultaneously explores underlying complementary information from multiple views. Unlike most existing single view subspace clustering methods that reconstruct data points using original features, our method seeks the underlying latent representation and simultaneously performs data reconstruction based on the learned latent representation. With the complementarity of multiple views, the latent representation could depict data themselves more comprehensively than each single view individually, accordingly makes subspace representation more accurate and robust as well. The proposed method is intuitive and can be optimized efficiently by using the Augmented Lagrangian Multiplier with Alternating Direction Minimization (ALM-ADM) algorithm. Extensive experiments on benchmark datasets have validated the effectiveness of our proposed method. Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Pengfei Zhu 0001, Xiaochun Cao |
CVPR | 1 |
| 2017 | Semi-Supervised Multi-view Multi-label Classification Based on Nonnegative Matrix Factorization
Guangxia Wang, Changqing Zhang 0002, Pengfei Zhu 0001, Qinghua Hu |
ICANN (2) | 2 |
| 2017 | Unsupervised feature selection by manifold regularized self-representationabstractUnsupervised feature selection has been proven to be an efficient technique in mitigating the curse of dimensionality. It helps to understand and analyze the prevalent high-dimensional unlabeled data. Recently, the self-similarity property of objects, which assumes that a feature can be represented by the linear combination of its relevant features, has been successfully used in unsupervised feature selection. However, it does not take the geometry structure of the sample space into consideration. In this paper, we propose a novel algorithm termed manifold regularized self-representation(MRSR). To preserve the local spatial structure, we incorporate an effective manifold regularization into the objective function. An iterative reweighted least square (IRLS) algorithm is developed to solve the optimization problem and the convergence is proved. Extensive experimental results on several benchmark datasets validate the effectiveness of the proposed method. Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002 |
ICIP | 5 |
| 2017 | Mixed sparsity regularized multi-view unsupervised feature selectionabstractThe traditional learning machines suffer from the curse of dimensionality because of the data explosion in the areas of multi-media, social network, etc. Feature selection is an effective technique to reduce storage burden and time complexity, and improve generalization ability of the learned models. In real-world applications, the data can be collected from different modalities, or described from multi-views as well. Compared with supervised cases, it is more challenging to reduce the feature dimensionality of multi-view data in unsupervised circumstances. The key difficulty with multiview unsupervised feature selection is how to characterize the multi-view relationships. In this paper, we propose a novel method for multi-view unsupervised feature selection by imposing sparsity on both individual features and views. To exploit the complementary information, we also take the view importance into consideration without introducing explicit view weights. Experiments on benchmark datasets show on the proposed algorithm outperforms other unsupervised feature selection methods. Kennedy W. Wangila, Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002 |
ICIP | 5 |
| 2017 | Independence regularized multi-label ensembleabstractIn this paper, we focus on promoting multi-label learning task with ensemble learning. Compared to traditional single algorithm methods, it has been recognized that ensemble methods could achieve much better performance than each constituent learned model, especially under the conditional independence of different classifiers. Existing multi-label ensemble algorithms mainly focus on creating diverse component learners by employing different mechanisms, mostly using randomization strategies by smart heuristics. Different from most existing methods, in this paper, we propose an ensemble method to learn the basic classifiers which considers the general independence of the different classifiers. Therefore, each learned multi-label classifier is guaranteed to be diverse and complementary. Furthermore, considering the different qualities of these classifiers, a weight vector is learned to balance these classifiers. Experiments on several benchmark datasets well demonstrate that the proposed method outperforms the state-of-the-art methods. Ziwei Yu, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001 |
ICME | 2 |
| 2017 | Multi-view Label Space Dimension Reduction
Pengfei Zhu 0001, Changqing Zhang 0002, Qinghua Hu |
ICONIP (1) | 3 |
| 2017 | LEAF: Latent Extended Attribute Features Discovery for Visual ClassificationabstractTo improve the discrimination of attribute representation, in this paper, we propose to extend the traditional attribute representations via embedding the latent high-order structure between attributes. Specifically, our aim is to construct the Latent Extended Attribute Features (LEAF) for visual classification. Since there only exist weak label for each attribute, we firstly propose a feature selection method to explore the common feature structures across categories. After that, the attribute classifiers are trained based on the selected features. Then, the category specific graph is introduced, which is composed of single attributes and their co-occurrence attribute pairs. This attribute graph is used as the initialized representation of each image. Considering our aim, we should discover the discriminative latent structure between attributes and train the robust category classifiers. To that end, we develop a joint learning objective function which is composed of the high-order representation mining term and the classifier training term. The mining term can both preserve category-specific information and discover the common structure between categories. Based on the discovery representation, the robust visual classifiers could be trained by the classifier term. Finally, an alternating optimization method is designed to seek the optimal solution of our objective function. Experimental results on the challenging datasets demonstrate the advantages of our proposed model over existing work. Hua Zhang 0008, Rui Wang 0032, Changqing Zhang 0002, Xiaochun Cao |
ACM Multimedia | 3 |
| 2017 | Research on axial bearing capacity of rectangular concrete-filled steel tubular columns based on artificial neural networks
Yansheng Du, Changqing Zhang 0002, Xiaochun Cao |
Frontiers Comput. Sci. | 3 |
| 2017 | Unsupervised feature selection via Diversity-induced Self-representation
Yanbei Liu, Changqing Zhang 0002, Jing Wang 0023, Xiao Wang 0017 |
Neurocomputing | 3 |
| 2017 | Subspace clustering guided unsupervised feature selection
Pengfei Zhu 0001, Wencheng Zhu, Qinghua Hu, Changqing Zhang 0002, Wangmeng Zuo |
Pattern Recognit. | 4 |
| 2017 | Salient Object Detection via Weighted Low Rank Matrix RecoveryabstractImage-based salient object detection is a useful and important technique, which can promote the efficiency of several applications such as object detection, image classification/retrieval, object co-segmentation, and content-based image editing. In this letter, we present a novel weighted low-rank matrix recovery (WLRR) model for salient object detection. In order to facilitate efficient salient objects-background separation, a high-level background prior map is estimated by employing the property of the color, location, and boundary connectivity, and then this prior map is ensembled into a weighting matrix which indicates the likelihood that each image region belongs to the background. The final salient object detection task is formulated as the WLRR model with the weighting matrix. Both quantitative and qualitative experimental results on three challenging datasets show competitive results as compared with 24 state-of-the-art methods. Chang Tang, Pichao Wang, Changqing Zhang 0002, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2017 | Flexible Multi-View Dimensionality Co-ReductionabstractDimensionality reduction aims to map the high-dimensional inputs onto a low-dimensional subspace, in which the similar points are close to each other and vice versa. In this paper, we focus on unsupervised dimensionality reduction for the data with multiple views, and propose a novel method, called Multi-view Dimensionality co-Reduction. Our method flexibly exploits the complementarity of multiple views during the dimensionality reduction and respects the similarity relationships between data points across these different views. The kernel matching constraint based on Hilbert-Schmidt Independence Criterion enhances the correlations and penalizes the disagreement of different views. Specifically, our method explores the correlations within each view independently, and maximizes the dependence among different views with kernel matching jointly. Thus, the locality within each view and the consistence between different views are guaranteed in the subspaces corresponding to different views. More importantly, benefiting from the kernel matching, our method need not depend on a common low-dimensional subspace, which is critical to reduce the influence of the unbalanced dimensionalities of multiple views. Specifically, our method explicitly produces individual low-dimensional projections for individual views, which could be applied for new coming data in the out-of-sample manner. Experiments on both clustering and recognition tasks demonstrate the advantages of the proposed method over the state-of-the-art approaches. Changqing Zhang 0002, Huazhu Fu, Qinghua Hu, Pengfei Zhu 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2017 | Depth-Preserving Stereo Image Retargeting Based on Pixel FusionabstractIn this paper, we propose a pixel fusion-based stereo image retargeting method, which could adaptively retarget stereo images with flexible aspect ratios, simultaneously preserving the depth. Retargeting each image independently by the pixel fusion method ignores the disparity relationship between pixels in the image pair and hence will introduce the distortion of disparity. To address this issue, we advocate to extend the single pixel fusion-based way to be applicable for stereo image pair. First, seams are selected based on the energy function, which simultaneously considers the seam selecting and seam matching. Second, a seam-matching-based matching map is proposed to preserve the disparity relationship between image pair. Then, the scaling factors for the left image are assigned considering both the important object and depth preservation. Subsequently, the scaling factors for the right image are obtained according to the proposed matching map. Based on these scaling factors, the stereo image pair is retargeted with pixel fusion. In contrast to removing pixels to resize image, the way of pixel fusion can obtain more smooth results with less depth distortion. Experimental results demonstrate that our method achieves more preferable qualities in both depth and shape preservation for stereo image retargeting. Jianjun Lei 0001, Changqing Zhang 0002, Feng Wu 0001, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 3 |
| 2016 | Coupled Dictionary Learning for Unsupervised Feature SelectionabstractUnsupervised feature selection (UFS) aims to reduce the time complexity and storage burden, as well as improve the generalization performance. Most existing methods convert UFS to supervised learning problem by generating labels with specific techniques (e.g., spectral analysis, matrix factorization and linear predictor). Instead, we proposed a novel coupled analysis-synthesis dictionary learning method, which is free of generating labels. The representation coefficients are used to model the cluster structure and data distribution. Specifically, the synthesis dictionary is used to reconstruct samples, while the analysis dictionary analytically codes the samples and assigns probabilities to the samples. Afterwards, the analysis dictionary is used to select features that can well preserve the data distribution. The effective L2p-norm (0 < p <1) regularization is imposed on the analysis dictionary to get much sparse solution and is more effective in feature selection.We proposed an iterative reweighted least squares algorithm to solve the L2p-norm optimization problem and proved it can converge to a fixed point. Experiments on benchmark datasets validated the effectiveness of the proposed method Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002, Wangmeng Zuo |
AAAI | 3 |
| 2016 | SketchNet: Sketch Classification with Web ImagesabstractIn this study, we present a weakly supervised approach that discovers the discriminative structures of sketch images, given pairs of sketch images and web images. In contrast to traditional approaches that use global appearance features or relay on keypoint features, our aim is to automatically learn the shared latent structures that exist between sketch images and real images, even when there are significant appearance differences across its relevant real images. To accomplish this, we propose a deep convolutional neural network, named SketchNet. We firstly develop a triplet composed of sketch, positive and negative real image as the input of our neural network. To discover the coherent visual structures between the sketch and its positive pairs, we introduce the softmax as the loss function. Then a ranking mechanism is introduced to make the positive pairs obtain a higher score comparing over negative ones to achieve robust representation. Finally, we formalize above-mentioned constrains into the unified objective function, and create an ensemble feature representation to describe the sketch images. Experiments on the TUBerlin sketch benchmark demonstrate the effectiveness of our model and show that deep feature representation brings substantial improvements over other state-of-the-art methods on sketch classification. Hua Zhang 0008, Si Liu 0001, Changqing Zhang 0002, Wenqi Ren, Rui Wang 0032, Xiaochun Cao |
CVPR | 3 |
| 2016 | Multi-view Representative and Informative Induced Active Learning
Huaxi Huang, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001 |
PRICAI | 2 |
| 2016 | Set to Set Visual Tracking
Wencheng Zhu, Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002 |
PRICAI | 4 |
| 2016 | Combining neighborhood separable subspaces for classification via sparsity regularized optimization
Pengfei Zhu 0001, Qinghua Hu, Yahong Han, Changqing Zhang 0002 |
Inf. Sci. | 4 |
| 2016 | Saliency Detection for Stereoscopic Images Based on Depth Confidence Analysis and Multiple Cues FusionabstractStereoscopic perception is an important part of human visual system that allows the brain to perceive depth. However, depth information has not been well explored in existing saliency detection models. In this letter, a novel saliency detection method for stereoscopic images is proposed. First, we propose a measure to evaluate the reliability of depth map, and use it to reduce the influence of poor depth map on saliency detection. Then, the input image is represented as a graph, and the depth information is introduced into graph construction. After that, a new definition of compactness using color and depth cues is put forward to compute the compactness saliency map. In order to compensate the detection errors of compactness saliency when the salient regions have similar appearances with background, foreground saliency map is calculated based on depth-refined foreground seeds' selection (DRSS) mechanism and multiple cues contrast. Finally, these two saliency maps are integrated into a final saliency map through weighted-sum method according to their importance. Experiments on two publicly available stereo data sets demonstrate that the proposed method performs better than other ten state-of-the-art approaches. Runmin Cong, Jianjun Lei 0001, Changqing Zhang 0002, Qingming Huang, Xiaochun Cao, Chunping Hou |
IEEE Signal Process. Lett. | 3 |
| 2016 | Saliency-Aware Nonparametric Foreground Annotation Based on Weakly Labeled DataabstractIn this paper, we focus on annotating the foreground of an image. More precisely, we predict both image-level labels (category labels) and object-level labels (locations) for objects within a target image in a unified framework. Traditional learning-based image annotation approaches are cumbersome, because they need to establish complex mathematical models and be frequently updated as the scale of training data varies considerably. Thus, we advocate the nonparametric method, which has shown potential in numerous applications and turned out to be attractive thanks to its advantages, i.e., lightweight training load and scalability. In particular, we exploit the salient object windows to describe images, which is beneficial to image retrieval and, thus, the subsequent image-level annotation and localization tasks. Our method, namely, saliency-aware nonparametric foreground annotation, is practical to alleviate the full label requirement of training data, and effectively addresses the problem of foreground annotation. The proposed method only relies on retrieval results from the image database, while pretrained object detectors are no longer necessary. Experimental results on the challenging PASCAL VOC 2007 and PASCAL VOC 2008 demonstrate the advance of our method. Xiaochun Cao, Changqing Zhang 0002, Huazhu Fu, Xiaojie Guo 0001, Qi Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Diversity-induced Multi-view Subspace ClusteringabstractIn this paper, we focus on how to boost the multi-view clustering by exploring the complementary information among multi-view features. A multi-view clustering framework, called Diversity-induced Multi-view Subspace Clustering (DiMSC), is proposed for this task. In our method, we extend the existing subspace clustering into the multi-view domain, and utilize the Hilbert Schmidt Independence Criterion (HSIC) as a diversity term to explore the complementarity of multi-view representations, which could be solved efficiently by using the alternating minimizing optimization. Compared to other multi-view clustering methods, the enhanced complementarity reduces the redundancy between the multi-view representations, and improves the accuracy of the clustering results. Experiments on both image and video face clustering well demonstrate that the proposed method outperforms the state-of-the-art methods. Xiaochun Cao, Changqing Zhang 0002, Huazhu Fu, Si Liu 0001, Hua Zhang 0008 |
CVPR | 2 |
| 2015 | Low-Rank Tensor Constrained Multiview Subspace ClusteringabstractIn this paper, we explore the problem of multiview subspace clustering. We introduce a low-rank tensor constraint to explore the complementary information from multiple views and, accordingly, establish a novel method called Low-rank Tensor constrained Multiview Subspace Clustering (LT-MSC). Our method regards the subspace representation matrices of different views as a tensor, which captures dexterously the high order correlations underlying multiview data. Then the tensor is equipped with a low-rank constraint, which models elegantly the cross information among different views, reduces effectually the redundancy of the learned subspace representations, and improves the accuracy of clustering as well. The inference process of the affinity matrix for clustering is formulated as a tensor nuclear norm minimization problem, constrained with an additional L2,1-norm regularizer and some linear equalities. The minimization problem is convex and thus can be solved efficiently by an Augmented Lagrangian Alternating Direction Minimization (AL-ADM) method. Extensive experimental results on four benchmark datasets show the effectiveness of our proposed LT-MSC method. Changqing Zhang 0002, Huazhu Fu, Si Liu 0001, Guangcan Liu, Xiaochun Cao |
ICCV | 1 |
| 2015 | Multi-cue Augmented Face ClusteringabstractFace clustering is an important but challenging task since facial images always have huge variation due to change in facial expressions, head poses and partial occlusions, etc. Moreover, face clustering is actually an unsupervised problem which makes it more difficult to reach an accurate result. Fortunately, there are some cues that can be used to improve clustering performance. In this paper, two types of cues are employed. The first one is pairwise constraints: must-link and cannot-link constraints, which can be extracted from the temporal and spatial knowledge of data. The other is that each face is associated with a series of attributes (i.e, gender) which can contribute discrimination among faces. To take advantage of the above cues, we propose a new algorithm, Multi-cue Augmented Face Clustering (McAFC), which effectively incorporates the cues via graph-guided sparse subspace clustering technique. Specially, facial images from the same individual are encouraged to be connected while faces from different persons are restrained to be connected. Experiments on three face datasets from real-world videos show the improvements of our algorithm over the state-of-the-art methods. Chengju Zhou, Changqing Zhang 0002, Huazhu Fu, Rui Wang 0032, Xiaochun Cao |
ACM Multimedia | 2 |
| 2015 | A flexible framework of adaptive method selection for image saliency detection
Changqing Zhang 0002, Zhiqiang Tao, Xingxing Wei 0001, Xiaochun Cao |
Pattern Recognit. Lett. | 1 |
| 2015 | Structured Saliency Fusion Based on Dempster-Shafer TheoryabstractVisual saliency has been widely used in many applications. However, the performance of an individual saliency detection method varies with the different images. Integrating multiple methods together could compensate this shortcoming, and thus is expected to improve the performance of saliency detection. In this paper, we present an unsupervised Dempster–Shafer Theory (DST) based saliency fusion framework. DST simulates the similar reasoning logic with humans to make decision analysis, and has been proved a suitable method for data fusion. Inspired by this, our framework formalizes the saliency fusion as a statistics inference process, considering the results from several saliency methods to accomplish the fusion task. Furthermore, the proposed framework can flexibly incorporate a variety of inherent structured priors within the images (e.g., clusters and saliency voting) when leveraging the fusion rule of DST. Therefore, it is more close to the fusion mechanism. Experimental results on two benchmark datasets demonstrate the effectiveness and robustness of our framework. Xingxing Wei 0001, Zhiqiang Tao, Changqing Zhang 0002, Xiaochun Cao |
IEEE Signal Process. Lett. | 3 |
| 2015 | Constrained Multi-View Video Face ClusteringabstractIn this paper, we focus on face clustering in videos. To promote the performance of video clustering by multiple intrinsic cues, i.e., pairwise constraints and multiple views, we propose a constrained multi-view video face clustering method under a unified graph-based model. First, unlike most existing video face clustering methods which only employ these constraints in the clustering step, we strengthen the pairwise constraints through the whole video face clustering framework, both in sparse subspace representation and spectral clustering. In the constrained sparse subspace representation, the sparse representation is forced to explore unknown relationships. In the constrained spectral clustering, the constraints are used to guide for learning more reasonable new representations. Second, our method considers both the video face pairwise constraints as well as the multi-view consistence simultaneously. In particular, the graph regularization enforces the pairwise constraints to be respected and the co-regularization penalizes the disagreement among different graphs of multiple views. Experiments on three real-world video benchmark data sets demonstrate the significant improvements of our method over the state-of-the-art methods. Xiaochun Cao, Changqing Zhang 0002, Chengju Zhou, Huazhu Fu, Hassan Foroosh |
IEEE Trans. Image Process. | 2 |
| 2014 | Output Feature Augmented LassoabstractLasso simultaneously conducts variable selection and supervised regression. In this paper, we extend Lasso to multiple output prediction, which belongs to the categories of structured learning. Though structured learning makes use of both input and output simultaneously, the joint feature mapping in current framework of structured learning is usually application-specific. As a result, ad hoc heuristics have to be employed to design different joint feature mapping functions for different applications, which results in the lackness of generalization ability for multiple output prediction. To address this limitation, in this paper, we propose to augment Lasso with output by decoupling the joint feature mapping function of traditional structured learning. The contribution of this paper is three-fold: 1) The augmented Lasso conducts regression and variable selection on both the input and output features, and thus the learned model could fit an output with both the selected input variables and the other correlated outputs. 2) To be more general, we set up nonlinear dependencies among output variables by generalized Lasso. 3) Moreover, the Augmented Lagrangian Method (ALM) with Alternating Direction Minimizing (ADM) strategy is used to find the optimal model parameters. The extensive experimental results demonstrate the effectiveness of the proposed method. Changqing Zhang 0002, Yahong Han, Xiaojie Guo 0001, Xiaochun Cao |
ICDM | 1 |
| 2014 | Video Face Clustering via Constrained Sparse RepresentationabstractIn this paper, we focus on the problem of clustering faces in videos. Different from traditional clustering on a collection of facial images, a video provides some inherent benefits: faces from a face track must belong to the same person and faces from a video frame can not be the same person. These benefits can be used to enhance the clustering performance. More precisely, we convert the above benefits into must-link and cannot-link constraints. These constraints are further effectively incorporated into our novel algorithm, Video Face Clustering via Constrained Sparse Representation (CS-VFC). The CS-VFC utilizes the constraints in two stages, including sparse representation and spectral clustering. Experiments on real-world videos show the improvements of our algorithm over the state-of-the-art methods. Chengju Zhou, Changqing Zhang 0002, Gaotao Shi, Xiaochun Cao |
ICME | 2 |