Xiaozhao Fang

dblp:140/6459 · DBLP profile ↗
← Back
95ranked-venue papers
14as first author
53since 2021 · last 2026
0000-0001-8440-1765ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 10 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Prototype-Based Semantic Consistency Alignment for Domain Adaptive Retrieval
abstract
Domain adaptive retrieval aims to transfer knowledge from a labeled source domain to an unlabeled target domain, enabling effective retrieval while mitigating domain discrepancies. However, existing methods encounter several fundamental limitations: 1) neglecting class-level semantic alignment and excessively pursuing pair-wise sample alignment; 2) lacking either pseudo-label reliability consideration or geometric guidance for assessing label correctness; 3) directly quantizing original features affected by domain shift, undermining the quality of learned hash codes. In view of these limitations, we propose Prototype-based Semantic Consistency Alignment (PSCA), a two-stage framework for effective domain adaptive retrieval. In the first stage, a set of orthogonal prototypes directly establishes class-level semantic connections, maximizing inter-class separability while gathering intra-class samples. During the prototype learning, geometric proximity provides a reliability indicator for semantic consistency alignment through adaptive weighting of pseudo-label confidences. The resulting membership matrix and prototypes facilitate feature reconstruction, ensuring quantization on reconstructed rather than original features, thereby improving subsequent hash coding quality and seamlessly connecting both stages. In the second stage, domain-specific quantization functions process the reconstructed features under mutual approximation constraints, generating unified binary hash codes across domains. Extensive experiments validate PSCA's superior performance across multiple datasets.
Tianle Hu, Weijun Lv, Na Han, Xiaozhao Fang, Jie Wen 0001, Jiaxing Li 0009, Guoxu Zhou
AAAI4
2026 Partial multi-label learning with local reconstruction and indirect guidance
Xiaozhao Fang, Guoxu Zhou, Zhouqiang Qiu, Junqiu Fan
Appl. Intell.5
2026 Central similarity joint-learning for cross-domain retrieval
Tianle Hu, Xiaozhao Fang, Jie Wen 0001, Guoxu Zhou, Shengli Xie 0001
Neural Networks4
2026 An Efficient Regenerated Cross-Modal Hashing: Improving Existing Hash Codes With the Arbitrary Length
abstract
In recent years, numerous hashing techniques have been developed to boost efficient cross-modal retrieval. Once a retrieval model is deployed, the hash code length is fixed to achieve optimal performance. To address different retrieval scenarios while maintaining retrieval accuracy, a common approach is to redesign and retrain the original model with different hash code length. However, this retraining process can increase the training load and may lead to worse results. To tackle these challenges, we present Regenerated Cross-Modal Hashing (RCMH), a novel cross-modal hashing framework designed to improve the quality of existing hash codes and convert them to arbitrary lengths with high efficiency. First, we clip or pad the existing hash codes to initialize them with the target length, under the supervision of the similarity matrix generated by the augmented label information. Second, we introduce a linear-nonlinear competitive reconstruction approach to reduce the semantic gaps and further capture the deeper relationships from linear image features and nonlinear text features. In this way, each pair of samples is compared and selected to obtain reconstructed binary codes that can preserve the modality-specific properties. Finally, to reduce the training costs caused by iterations of variables, the regenerate hashing term is utilized to regenerate final hash codes with the reconstructed binary codes while preserving the information from the existing hash codes without iterative optimization. Notably, RCMH can be integrated with existing state-of-the-art (SOTA) methods with robustness, helping them to adjust the hash code length and achieve better retrieval performance.
Kaihang Jiang, Wai Keung Wong, Xiaozhao Fang, Weijun Sun, Guoxu Zhou, Shengli Xie 0001, Xiaochun Cao
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Partial Multiview Incomplete Multilabel Learning via Uncertainty-Driven Reliable Dynamic Fusion
abstract
Currently, an increasing number of researchers are focusing on partial multiview incomplete multilabel learning. However, many methods generally integrate features from multiple views via an average weighting strategy, which overlooks the potential mismatch between the contribution of each view and their assigned fusion weights and thus generates unreliable fused features. To address this issue, we propose a novel uncertainty-driven reliable dynamic fusion framework for partial multiview incomplete multilabel learning. Unlike existing methods, the proposed uncertainty-driven reliable sample-level dynamic fusion module operates on the principle that samples exhibiting greater uncertainty possess fewer reliable features. This module evaluates the uncertainty of each sample and, in turn, estimates the reliability of features with the uncertainty of sample judgement, thereby obtaining reliable weights to guide the information fusion of multiple views. Furthermore, many existing approaches for handling incomplete multilabel scenarios typically concentrate on the information from annotated labels, neglecting the potential information of unknown tags. To bridge this gap, we incorporate an innovative pseudolabelling strategy that effectively identifies trustworthy pseudolabels that correspond to those unannotated uncertain labels, thereby adding additional supervisory information to assist model training. Moreover, we also devise a feature masking strategy to further augment the encoder's representation learning capabilities. The experimental results across five datasets demonstrate that our method outperforms current state-of-the-art methods.
Jie Wen 0001, Xiaohuan Lu, Chengliang Liu 0003, Xiaozhao Fang, Yong Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Joint asymmetric discrete hashing for cross-modal retrieval
Jiaxing Li 0009, Zuopeng Yang, Xiaozhao Fang, Shengli Xie 0001, Yong Xu 0001
Pattern Recognit.4
2026 Joint low-rank and sparse components extraction for cross-domain recognition
Zhixiang Zeng, Weijun Sun, Xiaozhao Fang, Guoxu Zhou, Shengli Xie 0001
Pattern Recognit.3
2026 Adapting Domain-Aware Knowledge to Vision-Language Model for Zero-Shot Anomaly Detection
abstract
Zero-shot anomaly detection (ZSAD) is a challenging task that aims to detect anomalies in images without any prior knowledge of the anomaly classes. This task is especially difficult because anomalies are rare, diverse, and often manifest differently across domains, making it hard for models to generalize when training data is scarce or unavailable. Recently, vision-language models (VLMs), such as CLIP, have shown great potential in ZSAD, but they often struggle to adapt to unseen domains due to the lack of domain-aware knowledge. To address these challenges, we propose the Domain Adaptation CLIP (DACLIP), a novel approach that adapts domain-aware knowledge to the VLM. Specifically, DACLIP leverages a Domain-Aware Knowledge Adaptation (DAKA) strategy to enhance CLIP for ZSAD across different domains. The DAKA strategy comprises multiple experts that specialize in target domains, enabling the model to dynamically select and combine specialized experts tailored to anomaly characteristics, thus improving its ability to generalize and detect a wide range of anomalies. Furthermore, we introduce learnable domain-aware prompts that are jointly learned by and injected into both the CLIP encoders (visual and text) and the DAKA modules. This dual-pathway learning enables the model to capture domain-specific features at multiple levels of the architecture, allowing for more effective adaptation to new domains and anomaly types. We evaluate our approach on several benchmark datasets spanning industrial and medical domains. Extensive experiments demonstrate that DACLIP consistently outperforms state-of-the-art methods in ZSAD, achieving significant improvements in both image-level and pixel-level anomaly detection tasks.
Zeqi Ma, Xiaozhao Fang, Jie Wen 0001, Guoxu Zhou, Shengli Xie 0001
IEEE Trans. Image Process.2
2026 Fine-Grained Enhancement Convolutional Diffusion Transformer for Unsupervised Anomaly Detection
abstract
Reconstruction-based methods have achieved excellent performance in anomaly detection. Diffusion models are considered highly suitable for anomaly detection tasks due to their strong ability in reconstruction. Nevertheless, diffusion-based models require the reconstruction of noise features, which may lack the capacity for fine-grained feature reconstruction and fail to provide adequate semantic information for reconstruction guidance. To solve the aforementioned problems, this paper proposes a Fine-Grained Enhancement Convolutional Diffusion Transformer Anomaly Detection (FECDTAD) framework for multi-class anomaly detection. The core model of the proposed framework is the Fine-Grained Enhancement Convolutional Denoising Transformer (FECDT), which employs the diffusion transformer paradigm. To enhance fine-grained reconstruction in the diffusion process, the FECDTAD adopts a series of feature information fusion strategies. Specifically, to enhance both fine-grained perception and global understanding, the FECDT model employs a simple feature fusion module to integrate shallow-level and deep-level features extracted from a pre-trained vision transformer. To enhance the capacity for fine-grained feature reconstruction, the FECDT integrates local and global information via a CNN-Transformer architecture. Moreover, to provide guidance for the reconstruction of anomalous areas, semantic information is propagated into the FECDT through a Cross-Attention module. Experimental results demonstrate that the proposed method is effective and can surpass the state-of-the-art methods.
Zeqi Ma, Xiaozhao Fang, Jie Wen 0001, Guoxu Zhou, Yong Xu 0001
IEEE Trans. Image Process.2
2026 Toward Bidirectional Adaptability for Few-Shot Class-Incremental Learning With Forward-Backward Knowledge Transfer
abstract
The development of Deep Neural Networks (DNNs) has enabled AI-driven models to excel in recognizing a limited set of classes within static environments. As AI systems progress, few-shot class-incremental learning (FSCIL) aims to expand their understanding of novel classes from minimal samples while retaining knowledge of previously encountered ones. However, most existing FSCIL models face significant challenges, includinginadequate adaptabilityandcatastrophic forgetting, which hinder their ability to maintain robust forward and backward learning capabilities. To address these issues, this paper proposes a novel Forward-Backward Knowledge Transfer (FBKT) paradigm, which strategically integrates forward distribution adaptation (FDA) and backward semantic alignment (BSA) mechanisms to achieve bidirectional adaptability in knowledge transfer. The FDA mechanism enhances forward adaptability by expanding and reserving the embedding space for new classes using semantic-irrelevant masked images as virtual negative classes, thereby mitigating data overfitting. It also employs self-supervised representation learning to utilize semantic-relevant local embeddings as additional positive samples, fostering class separation and generalization. Meanwhile, the BSA mechanism ensures the semantic consistency of previously learned classes across sessions during class-incremental learning, promoting smoother backward adaptability and reducing model degradation. Extensive experiments conducted on multiple benchmark datasets consistently highlight the superior performance and effectiveness of our FBKT compared to state-of-the-art methods.
Bingzhi Chen, Sudong Cai, Xiaozhao Fang, Mohammed Bennamoun, Shengli Xie 0001
IEEE Trans. Multim.4
2026 Dual Label Association Recovery for Partial Multi-Label Learning
abstract
Partial Multi-Label Learning (PML) deals with a practical scenario where each instance is associated with a set of candidate labels, among which only a subset corresponds to the ground-truth labels while the others are unrelated. Existing PML methods typically employ label association recovery as a structural disambiguation strategy to identify credible labels. However, these methods attempt to recover associations directly from candidate labels, which leads to unreliable disambiguation due to the distorted label structures. To this end, this paper proposes a novel PML method via dual label association recovery (PML-DLAR).The essential strategy is to first eliminate the spurious correlations in label space before recovering dual label associations. Specifically, a biorthogonal transformation is employed to decouple the instance-level and label-level association structures affected by noisy labels. Subsequently, reliable instance-level associations are reconstructed through global geometric structure alignment between feature and pseudo-label spaces. Finally, class specific feature representations are constructed through class prototypes to guide label-level semantic association recovery in label space. Comprehensive experiments validate the superior performance of PML-DLAR over state-of-the-art methods.
Xuhuan Zhu, Xiaozhao Fang, Jie Wen 0001, Jing Zhang 0022, Guoxu Zhou, Shengli Xie 0001
IEEE Trans. Multim.3
2025 Lightweight Contrastive Distilled Hashing for Online Cross-modal Retrieval
abstract
Deep online cross-modal hashing has gained much attention from researchers recently, as its promising applications with low storage requirement, fast retrieval efficiency and cross modality adaptive, etc. However, there still exists some technical hurdles that hinder its applications, e.g., 1) how to extract the coexistent semantic relevance of cross-modal data, 2) how to achieve competitive performance when handling the real time data streams, 3) how to transfer the knowledge learned from offline to online training in a lightweight manner. To address these problems, this paper proposes a lightweight contrastive distilled hashing (LCDH) for cross-modal retrieval, by innovatively bridging the offline and online cross-modal hashing by similarity matrix approximation in a knowledge distillation framework. Specifically, in the teacher network, LCDH first extracts the cross-modal features by CLIP, which are further fed into an attention module for representation enhancement after feature fusion. Then, the output of the attention module is fed into a FC layer to obtain hash codes for aligning the sizes of similarity matrices for online and offline training. In the student network, LCDH extracts the visual and textual features by lightweight models, and then the features are fed into a FC layer to generate binary codes. Finally, by approximating the similarity matrices, the performance of online hashing in the lightweight student network can be enhanced by the supervision of coexistent semantic relevance that is distilled from the teacher network. Experimental results on three widely used datasets demonstrate that LCDH outperforms some state-of-the-art methods.
Jiaxing Li 0009, Zeqi Ma, Kaihang Jiang, Xiaozhao Fang, Jie Wen 0001
AAAI5
2025 Pseudo-Label Reconstruction for Partial Multi-Label Learning
abstract
In Partial Multi-Label Learning (PML), each instance is associated with a candidate label set containing multiple relevant labels along with other false positive labels. Currently, most PML methods directly extract instance correlation from instance features while ignoring the candidate labels, which may contain more discriminative instance-related information. This paper argues that, with a well-designed model, more accurate instance correlation can be mined from the candidate labels to facilitate label disambiguation. To this end, we propose a novel PML method based on pseudo-label reconstruction (PML-PLR). Specifically, we first propose a novel orthogonal candidate label reconstruction method, which jointly optimizes with instance features to extract more consistent instance correlation. Then, we use instance correlation as reconstruction coefficient to reconstruct pseudo-labels. Subsequently, through local manifold learning, the reconstructed pseudo-labels are leveraged to propagate the consistency relationship between labels and instances, thereby improving the accuracy of pseudo-labels. Extensive experiments and analyses demonstrate that the proposed PML-PLR outperforms state-of-the-art methods.
Na Han, Guanbin Li, Hongbo Gao 0001, Xiaozhao Fang
IJCAI7
2025 Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases
abstract
Existing Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, that incorporates three well-established mechanisms, i.e., Modality-driven Heterogeneous Optimization (MHO), Gradient-guided Modality Synergy (GMS), and Distribution-adapted Loss Rescaling (DLR), for comprehensively mitigating language biases from both causal and effectual perspectives. Specifically, MHO employs adaptive learning rates for specific modalities to achieve heterogeneous optimization, thus enhancing robust reasoning capabilities. Additionally, GMS leverages the Pareto optimization method to foster synergistic interactions between modalities and enforce gradient orthogonality to eliminate bias updates, thereby mitigating language biases from the effect side, i.e., shortcut bias. Furthermore, DLR is designed to assign adaptive weights to individual losses to ensure balanced learning across all answer categories, effectively alleviating language biases from the cause side, i.e., imbalance biases within datasets. Extensive experiments on multiple traditional and bias-sensitive benchmarks consistently demonstrate the robustness of CEDO over state-of-the-art competitors.
Huanjia Zhu, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Bingzhi Chen
IJCAI3
2025 Label Prediction Inherited Hashing for Cross-Modal Retrieval: Applying Supervised Hashing to Unsupervised Tasks
abstract
Supervised cross-modal hashing has achieved remarkable progress in retrieving related items across different modalities. However, in practical applications, a significant portion of data remains unlabeled, such as online data on websites, which must be included for effective retrieval. To address this challenge, while maintaining the high accuracy and efficiency of supervised methods, few works have attempted to adapt existing supervised techniques to handle unsupervised tasks through a general modular approach. To this end, we introduce a novel cross-modal hashing method, termed Label Prediction Inherited Hashing (LPIH). Initially, LPIH leverages labeled data to learn high-quality general label functions using supervised methods. Subsequently, it inherits the existing hash codes from existing supervised methods to further refine the pseudo-label information. Finally, LPIH integrates the refined pseudo-label information with the existing hash functions to learn new hash functions specifically tailored for unsupervised tasks. Extensive experimental results on three public datasets demonstrate the superior performance of LPIH compared to state-of-the-art (SOTA) cross-modal hashing methods. Specifically, LPIH achieves an average precision improvement of 5% over SOTA methods, highlighting its effectiveness in bridging the gap between supervised and unsupervised learning in the context of cross-modal retrieval.
Kaihang Jiang, Wai Keung Wong, Jianyang Qin, Xiaozhao Fang, Jie Wen 0001, Bingzhi Chen, Hongbo Gao 0001
ACM Multimedia4
2025 Confidence-Aware With Prototype Alignment for Partial Multi-label Learning
abstract
Label prototype learning has emerged as an effective paradigm in Partial Multi-Label Learning (PML), providing a distinctive framework for modeling structured representations of label semantics while naturally filtering noise through prototype-based label confidence estimation. However, existing prototype-based methods face a critical limitation: class prototypes are the biased estimates due to noisy candidate labels, particularly when positive samples are scarce. To this end, we first propose a mutually class prototype alignment strategy bypassing noise interference by introducing two different transformation matrices, which makes the class prototypes learned by the fuzzy clustering and candidate label set mutually alignment for correcting themselves. Such alignment is also passed on to the fuzzy memberships label in turn. In addition, to eliminate noise interference in the candidate label set during the classifier learning, we use the learned permutation matrix to transform the fuzzy memberships label for learning a label reliability indicator matrix accompanied by the candidate label set. This makes the label reliability indicator matrix absolutely prevent the occurrence of numerical values located in non-label and simultaneously eliminate the introduction of incorrect label as much as possible. The resulting indicator matrix guides a robust multi-label classifier training process, jointly optimizing label confidence and classifier parameters. Extensive experiments demonstrate that our proposed model exhibits significant performance advantages over state-of-the-art PML approaches.
Weijun Lv, Xiaozhao Fang, Xuhuan Zhu, Jie Wen 0001, Guoxu Zhou
NeurIPS3
2025 Consistent coding guided domain adaptation retrieval
Tianle Hu, Yonghao Chen, Weijun Lv, Xiaozhao Fang
Appl. Intell.5
2025 Coarse-to-fine label refinement for domain adaptive retrieval
Tianle Hu, Chuwei Cheng, Junhong Xiao, Weijun Sun, Xiaozhao Fang
Inf. Sci.6
2025 Label enhancement hashing induced by class prototypes for domain adaptive retrieval
Chuwei Cheng, Tianle Hu, Qiyu Deng, Xiaozhao Fang
Multim. Syst.6
2025 Asymmetric semantic preserving hashing for cross-modal retrieval
Qiyu Deng, Chuwei Cheng, Junhong Xiao, Xiaozhao Fang
Multim. Syst.6
2025 Structure center fusion and guidance learning for domain adaptive retrieval
Zejiang Xu, Xiaozhao Fang, Han Na, Weijun Sun
Multim. Syst.4
2025 Fuzzy bifocal disambiguation for partial multi-label learning
Xiaozhao Fang, Yonghao Chen, Shengli Xie 0001, Na Han
Neural Networks1
2025 Multi-view graph clustering with Dually Enhanced Tensor Rank Minimization and Diverse Separation of Inconsistent Information
Weijun Sun, Chaoye Li, Jiakai He, Xiaozhao Fang, Guoxu Zhou, Xiyuan Yang, Kangsheng Wu
Neural Networks4
2025 Joint Intra-view and Inter-view Enhanced Tensor Low-rank Induced Affinity Graph Learning
Weijun Sun, Chaoye Li, Qiaoyun Li, Xiaozhao Fang, Jiakai He
Pattern Recognit.4
2025 Asymmetric and Discrete Self-Representation Enhancement Hashing for Cross-Domain Retrieval
abstract
Due to the characteristics of low storage requirement and high retrieval efficiency, hashing-based retrieval has shown its great potential and has been widely applied for information retrieval. However, retrieval tasks in real-world applications are usually required to handle the data from various domains, leading to the unsatisfactory performances of existing hashing-based methods, as most of them assuming that the retrieval pool and the querying set are similar. Most of the existing works overlooked the self-representation that containing the modality-specific semantic information, in the cross-modal data. To cope with the challenges mentioned above, this paper proposes an asymmetric and discrete self-representation enhancement hashing (ADSEH) for cross-domain retrieval. Specifically, ADSEH aligns the mathematical distribution with domain adaptation for cross-domain data, by exploiting the correlation of minimizing the distribution mismatch to reduce the heterogeneous semantic gaps. Then, ADSEH learns the self-representation which is embedded into the generated hash codes, for enhancing the semantic relevance, improving the quality of hash codes, and boosting the generalization ability of ADSEH. Finally, the heterogeneous semantic gaps are further reduced by the log-likelihood similarity preserving for the cross-domain data. Experimental results demonstrate that ADSEH can outperform some SOTA baseline methods on four widely used datasets.
Jiaxing Li 0009, Xiaozhao Fang, Shengli Xie 0001, Yong Xu 0001
IEEE Trans. Image Process.3
2025 Collaboratively Semantic Alignment and Metric Learning for Cross-Modal Hashing
abstract
Cross-modal retrieval is a promising technique nowadays to find semantically similar instances in other modalities while a query instance is given from one modality. However, there still exists many challenges for reducing heterogeneous modality gap by embedding label information to discrete hash codes effectively, solving the binary optimization when generating unified hash codes and reducing the discrepancy of data distribution efficiently during common space learning. In order to overcome the above-mentioned challenges, we propose a Collaboratively Semantic alignment and Metric learning for cross-modal Hashing (CSMH) in this paper. Specifically, by a kernelization operation, CSMH first extracts the non-linear data features for each modality, which are projected into a latent subspace to align both marginal and conditional distributions simultaneously. Then, a maximum mean discrepancy-based metric strategy is customized to mitigate the distribution discrepancies among features from different modalities. Finally, semantic information obtained from the label similarity matrix, is further incorporated to embed the latent semantic structure into the discriminant subspace. Experimental results of CSMH and baseline methods on four widely-used datasets show that CSMH outperforms some state-of-the-art hashing baseline methods for cross-modal retrieval on efficiency and precision.
Jiaxing Li 0009, Wai Keung Wong, Kaihang Jiang, Xiaozhao Fang, Shengli Xie 0001, Jie Wen 0001
IEEE Trans. Knowl. Data Eng.5
2025 LBF-VQA: Towards Language Bias-Free Visual Question Answering With Multi-Space Collaborative Debiasing
Yishu Liu 0001, Huanjia Zhu, Bingzhi Chen, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001
IEEE Trans. Knowl. Data Eng.4
2025 Toward Robust Semi-Supervised Distribution Alignment Against Label Distribution Shift With Noisy Annotations
abstract
Deep learning-based AI models typically require a large amount of high-quality annotated data to achieve optimal performance. However, thelabel distribution shiftcaused by noisy annotations can lead to perturbations in the classification boundary, reducing the robustness and generalization capabilities of deep learning models. To mitigate this issue, we transform the problem of learning from noisy labels into a semi-supervised learning problem, and propose a novel Semi-Supervised Distribution Alignment (SSDA) framework that strategically integrates noise-robust distribution alignment within a unified semi-supervised learning paradigm for combating noisy labels. By leveraging the similarity distribution between historical predictions, the proposed SSDA approach benefits from a flexible multi-historical regression modeling strategy, which aims to identify high-confidence samples/pairs and recalibrate the label shift through pseudo-labels. Furthermore, our approach employs a comprehensive multi-granularity distribution adaptation strategy, incorporating both instance-wise and class-aware distribution alignment to quantitatively minimize semantic discrepancies across different mixed feature domains. In this way, our SSDA approach ultimately achieves more resilient and generalizable performance against label noise, even in the presence of substantial noise. Extensive experiments conducted on multiple simulated and real-world noisy benchmark datasets consistently demonstrate the superiority and effectiveness of our SSDA method compared to existing state-of-the-art baselines.
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001, Xuelong Li 0001
IEEE Trans. Multim.4
2025 Driving Risk Assessment for Intelligent Vehicles Based on Entropy-Informed Graph Neural Networks and Gaussian Distributions
abstract
This study proposes a novel framework based on an entropy-informed graph neural network (EIGNN) integrated with Gaussian distribution (GD) to assess the driving risk of intelligent vehicles in typical traffic scenarios. Existing research often overlooks comprehensive spatiotemporal modeling of vehicle interaction characteristics and the quantification of uncertainty in dynamic risk assessments. In this work, vehicle speed and acceleration are probabilistically modeled using GD, while entropy theory is introduced to quantify risk uncertainty. A risk assessment model based on graph neural networks (GNNs) is then designed to capture the spatiotemporal dynamics of multivehicle interactions and predict the potential risk levels of driving strategies. The results demonstrate that the framework accurately quantifies collision risks in multivehicle interactions in complex traffic scenarios, with high accuracy and robustness across typical situations such as cruising, cut-ins, lane changes, overtaking, and different density traffic. By thoroughly analyzing traffic risk characteristics and incorporating them into intelligent driving decision-making, this study provides significant technical insights and theoretical support for enhancing the safety and decision-making efficiency of autonomous driving systems.
Hongbo Gao 0001, Chengbo Wang 0001, Runda Niu, Xiaozhao Fang, Jinpeng Chen 0001, Yining Sun, Huiqing Jin, Danwei Wang
IEEE Trans. Neural Networks Learn. Syst.4
2025 Random Online Hashing for Cross-Modal Retrieval
abstract
In the past decades, supervised cross-modal hashing methods have attracted considerable attentions due to their high searching efficiency on large-scale multimedia databases. Many of these methods leverage semantic correlations among heterogeneous modalities by constructing a similarity matrix or building a common semantic space with the collective matrix factorization method. However, the similarity matrix may sacrifice the scalability and cannot preserve more semantic information into hash codes in the existing methods. Meanwhile, the matrix factorization methods cannot embed the main modality-specific information into hash codes. To address these issues, we propose a novel supervised cross-modal hashing method called random online hashing (ROH) in this article. ROH proposes a linear bridging strategy to simplify the pair-wise similarities factorization problem into a linear optimization one. Specifically, a bridging matrix is introduced to establish a bidirectional linear relation between hash codes and labels, which preserves more semantic similarities into hash codes and significantly reduces the semantic distances between hash codes of samples with similar labels. Additionally, a novel maximum eigenvalue direction (MED) embedding method is proposed to identify the direction of maximum eigenvalue for the original features and preserve critical information into modality-specific hash codes. Eventually, to handle real-time data dynamically, an online structure is adopted to solve the problem of dealing with new arrival data chunks without considering pairwise constraints. Extensive experimental results on three benchmark datasets demonstrate that the proposed ROH outperforms several state-of-the-art cross-modal hashing methods.
Kaihang Jiang, Wai Keung Wong, Xiaozhao Fang, Jiaxing Li 0009, Jianyang Qin, Shengli Xie 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Denoising High-Order Graph Clustering
abstract
High-Order Graph (HOG) clustering has received much attention for its advantage of exploiting the rich intrinsic structure of data. However, the construction of HOG involves the generation of a large number of redundant walks, which dilutes the useful walks and thus leads to untrustworthy high-order similarity and, consequently, suboptimal clustering results may be obtained. We formalize this issue as the Weight Explosion (WE) problem. Furthermore, current works rarely focus on exploiting the correlation between multi-order graphs that can capture high-order relations at various levels. In this paper, we first analyze the pattern of redundant walks, also termed as noise, and subsequently propose a novel$h$-length Simple Path Search ($h$-SPS) algorithm to solve the WE problem.$h$-SPS aims to find valid walks to denoise HOG and thus avoids enumerating walks to report the similarity. Regarding the second problem, we propose a multi-order graphs fusion method, which adaptively integrates graphs of varying orders by solving a convex problem. This allows us to capture information across different order levels effectively. Extensive experiments on benchmark datasets demonstrate that our method11https://github.com/YonghaoChen511/DenoHOG can effectively solve the proposed WE problem, while also well exploiting the correlation of multi-order graphs.
Yonghao Chen, Ruibing Chen, Qiaoyun Li, Xiaozhao Fang, Jiaxing Li 0009, Wai Keung Wong
ICDE4
2024 Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image Recognition
abstract
Large-scale pre-trained vision-language models (e.g., CLIP) have shown powerful zero-shot transfer capabilities in image recognition tasks. Recent approaches typically employ supervised fine-tuning methods to adapt CLIP for zero-shot multi-label image recognition tasks. However, obtaining sufficient multi-label annotated image data for training is challenging and not scalable. In this paper, we propose a new language-driven framework for zero-shot multi-label recognition that eliminates the need for annotated images during training. Leveraging the aligned CLIP multi-modal embedding space, our method utilizes language data generated by LLMs to train a cross-modal classifier, which is subsequently transferred to the visual modality. During inference, directly applying the classifier to visual inputs may limit performance due to the modality gap. To address this issue, we introduce a cross-modal mapping method that maps image embeddings to the language modality while retaining crucial visual information. Comprehensive experiments demonstrate that our method outperforms other zero-shot multi-label recognition methods and achieves competitive results compared to few-shot methods.
Jie Wen 0001, Chengliang Liu 0003, Xiaozhao Fang, Yong Xu 0001, Zheng Zhang 0006
ICML4
2024 Stay Focused is All You Need for Adversarial Robustness
Bingzhi Chen, Ruihan Liu, Yishu Liu 0001, Xiaozhao Fang, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006
ACM Multimedia4
2024 Partial Multi-label Learning Based On Near-Far Neighborhood Label Enhancement And Nonlinear Guidance
Na Han, Xiaozhao Fang, Bingzhi Chen, Jie Wen 0001
ACM Multimedia4
2024 SCH: Symmetric Consistent Hashing for cross-modal retrieval
Haomin Ni, Xiaozhao Fang, Peipei Kang, Hongbo Gao 0001, Guoxu Zhou, Shengli Xie 0001
Signal Process.2
2024 CKDH: CLIP-Based Knowledge Distillation Hashing for Cross-Modal Retrieval
abstract
Recently, deep hashing-based cross-modal retrieval has attracted much attention of researchers, due to its advantages of fast retrieval efficiency and low storage overhead, etc. However, the existing deep hashing-based cross-modal retrieval methods typically 1) suffer from inadequately capturing the semantic relevance and coexistent information for cross-modal data, which may result in sub-optimal retrieval performance, 2) require a more comprehensive similarity measurement for cross-modal features to ensure high retrieval accuracy, 3) lack of scalability for lightweight deployment framework. To handle the issues mentioned above, we propose a CLIP-based knowledge distillation hashing (CKDH) for cross-modal retrieval, by referring the research trend of combining traditional methods and modern neural architecture to design lightweight networks based on large language models. Specifically, to effectively help capture the semantic relevance and coexistent information, CLIP is fine-tuned to extract visual features, while a graph attention network is used to enhance textual features extracted by bag-of-words model in the teacher model. Then, for better supervising the training of student model, a more comprehensive similarity measurement is introduced to represent distilled knowledge by jointly preserving the log-likelihood, intra and inter modality similarities. Finally, the student model extracts deep features by a lightweight networks, and generates the hash codes under the supervision of the similarity matrix produced by the teacher model. Experimental results on three widely used datasets demonstrate that CKDH can outperform some state-of-the-art methods, by delivering the best result consistently.
Jiaxing Li 0009, Wai Keung Wong, Xiaozhao Fang, Shengli Xie 0001, Yong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Two-Step Strategy for Domain Adaptation Retrieval
abstract
Conventional hash-based retrieval method rely on the assumption that the query and database are of the identical domain. However, cross-domain problem often occurs in real-world applications, leading to the unsatisfactory performance of existing hashing methods. Recently, some researchers have put forward domain adaptation retrieval (DAR) under the perspective of domain adaptation (DA) and achieved promising results. But the following limitations still exist: 1) a single function is used to handle two challenges, i.e., domain adaptation and hashing, which is not flexible to explore enough underlying information for simultaneously accomplishing these challenges well; 2) non-dominant features in the sample are ignored; 3) the dissimilarity structure of dissimilar samples is not taken into account. To address the above problems, we propose a novel framework named two-step strategy (TSS) for domain adaptation retrieval, which advocates dividing DAR into two steps: DA step and hashing step. A DA function and a hash function are learned to handle the above two challenges, respectively, making the process more reasonable. Additionally, a discriminant semantic fusion loss is proposed to improve the discriminative ability among classes. Unlike other works that focus on discovering dominant features, we exploit the neglected non-dominant features and assign them attention with sinusoidal semantic embedding, actively creating a clear separation between classes. At last, we present an adaptive similarity preserving loss to preserve the similarity structure of the original data in all intra-domain and inter-domain hash codes. Extensive experiments on various datasets demonstrate that the proposed TSS achieves state-of-the-art performance.
Yonghao Chen, Xiaozhao Fang, Peipei Kang, Na Han, Shengli Xie 0001
IEEE Trans. Knowl. Data Eng.2
2024 Two-Stage Asymmetric Similarity Preserving Hashing for Cross-Modal Retrieval
abstract
Hashing-based techniques present appealing solutions for cross-modal retrieval due to its low storage requirements and excellent query efficiency. The majority of cross-modal hashing methods typically adopt equal-length encoding scheme to represent multimodal data and achieve cross-modal similarity search. However, such scheme can be regarded as a relatively strict limitation, because it sacrifices the flexible representation of multimodal data in reality and cannot always guarantee the optimal retrieval performance. To address the challenge, this paper focuses on encoding heterogeneous data with varying hash lengths. To achieve this purpose, we propose a flexible cross-modal hashing approach, named Two-stage Asymmetric Similarity Preserving Hashing, TASPH for short, which can be applied to both unequal-length and equal-length retrieval scenarios. Specifically, in the first stage, TASPH designs a novel discrete asymmetric strategy to learn the modality-specific hash codes with varying lengths, enabling a flexible representation of heterogeneous data. Simultaneously, TASPH utilizes two semantic transformation matrices to establish the semantic correlations between varying hash codes. Different from most of the existing approaches that employ relaxation solutions, TASPH satisfies the discrete constraints without any relaxation. In the second stage, the learned semantic transformation matrices are employed to alleviate cross-modal heterogeneity, which guarantees that TASPH can learn more powerful hash functions to improve the discriminative ability of hash codes. Abundant experiments conducted on three benchmark datasets demonstrate encouraging results compared with the state-of-the-art approaches under different retrieval scenarios.
Junfan Huang, Peipei Kang, Na Han, Yonghao Chen, Xiaozhao Fang, Hongbo Gao 0001, Guoxu Zhou
IEEE Trans. Knowl. Data Eng.5
2024 Dual Noise Elimination and Dynamic Label Correlation Guided Partial Multi-Label Learning
abstract
Partial multi-label learning (PML) needs to address the problem of multi-label learning when the dataset contains redundant information. PML is more challenging compared to traditional multi-label learning, because PML needs not only to perform the multi classification task, but also to reduce the impact of noise information on the model. Existing PML methods suffer from the following problems. (1) Only single source of noise is considered. (2) Some methods ignore the label correlations. To solve the above problems, we proposes a new dual noise elimination and dynamic label correlation guided partial multi-label learning (PML-DNDC). Specifically, the hidden ground-truth label matrix is decomposed into two compressed matrices of instance and classifier, which are used to approximate the candidate label matrix, to eliminate the negative effects of label noise on the model. On one hand, the compressed instance matrix maintains local structural consistency with the original instances, eliminating noise in the feature. On the other hand, dynamic label correlation guidance is designed to help classifier training by dynamically exploring the potential label correlations, which encourages relevant labels to obtain similar classifiers. After extensive experiments and analyses, we conclude that the proposed PML-DNDC is superior to the state-of-the-art methods.
Xiaozhao Fang, Peipei Kang, Yonghao Chen, Yuting Fang, Shengli Xie 0001
IEEE Trans. Multim.2
2024 Efficient Discriminative Hashing for Cross-Modal Retrieval
abstract
Hashing techniques have been extensively studied in cross-modal retrieval due to their advantages in high computational efficiency and low storage cost. However, existing methods unconsciously ignore the complementary information of multimodal data, thus failing to consider learning discriminative hash codes from the perspective of information complementarity while often involving time-consuming training overhead. To tackle the above issues, we propose an efficient discriminative hashing (EDH) with information complementarity consideration. Specifically, we reckon that multimodal features and their corresponding semantic labels describe heterogeneous data viewed from low-and high-level structures, which owns complementarity. To this end, low-level latent representation and high-level semantics representation are simply derived. Then, a joint learning strategy is formulated to simultaneously exploit the above two representations for generating discriminative hash codes, which is quite computationally efficient. Besides, EDH decomposes hash learning into two steps. To obtain powerful hash functions which are conductive to retrieval, a regularization term considering pairwise semantic similarity is introduced into hash functions learning. In addition, an efficient optimization algorithm is designed to solve the optimization problem in EDH. Extensive experiments conducted on benchmark datasets demonstrate the superiority of our EDH in terms of retrieval performance and training efficiency. The source code is available at https://github.com/hjf-hjf/EDH.
Junfan Huang, Peipei Kang, Xiaozhao Fang, Na Han, Shengli Xie 0001, Hongbo Gao 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2023 Partial multi-label learning: exploration of binary ground-truth labels
abstract
Partial multi-label learning (PML) aims to accurately predict multi-labels for unknown instances with noise labels in training set. Accurate identification of ground-truth labels is essential to optimize performance. The accuracy of ground-truth labels affects label correlation capture, which guides classifier learning. However, current methods often rely on approximation of ground-truth labels through intermediate variables, rather than restoring binary ground-truth labels directly. To solve this problem, we propose a partial multi-label learning algorithm by exploring binary ground-truth labels (PML-EBGL). First, the candidate label matrix is decomposed into a ground-truth label matrix and a noise matrix. Then, the noise matrix is constrained using the l1norm for sparsity, and the binary ground-truth label matrix is used to explore label correlation, which is transferred to classifier by using the Laplacian term. Finally, the rotation matrix is added after the classifier, and the instance is projected onto the binary label matrix. Extensive experiments show that PML-EBGL outperforms state-of-the-art methods.
Xiaozhao Fang, Weijun Lv, Peipei Kang
ICME2
2023 MKB: Multi-Kernel Bures Metric for Nighttime Aerial Tracking
Peipei Kang, Qintai Hu, Xiaozhao Fang
PRCV (10)4
2023 Low-rank constraint-based multiple projections learning for cross-domain classification
Weiying Guo, Xiaozhao Fang, Na Han, Shaohua Teng
Knowl. Based Syst.2
2023 Cross-modal hashing with missing labels
Haomin Ni, Peipei Kang, Xiaozhao Fang, Weijun Sun, Shengli Xie 0001, Na Han
Neural Networks4
2023 Low-rank constraint based dual projections learning for dimensionality reduction
Xiaozhao Fang, Weijun Sun, Na Han, Shaohua Teng
Signal Process.2
2022 Semantic-Adversarial Graph Convolutional Network for Zero-Shot Cross-Modal Retrieval
Lunke Fei, Peipei Kang, Xiaozhao Fang, Shaohua Teng
PRICAI (2)5
2022 Intra-class low-rank regularization for supervised and semi-supervised cross-modal retrieval
Peipei Kang, Zehang Lin, Zhenguo Yang, Xiaozhao Fang, Alexander M. Bronstein, Qing Li 0001, Wenyin Liu
Appl. Intell.4
2022 Dynamic Double Classifiers Approximation for Cross-Domain Recognition
abstract
In general, existing cross-domain recognition methods mainly focus on changing the feature representation of data or modifying the classifier parameter and their efficiencies are indicated by the better performance. However, most existing methods do not simultaneously integrate them into a unified optimization objective for further improving the learning efficiency. In this article, we propose a novel cross-domain recognition algorithm framework by integrating both of them. Specifically, we reduce the discrepancies in both the conditional distribution and marginal distribution between different domains in order to learn a new feature representation which pulls the data from different domains closer on the whole. However, the data from different domains but the same class cannot interlace together enough and thus it is not reasonable to mix them for training a single classifier. To this end, we further propose to learn double classifiers on the respective domain and require that they dynamically approximate to each other during learning. This guarantees that we finally learn a suitable classifier from the double classifiers by using the strategy of classifier fusion. The experiments show that the proposed method outperforms over the state-of-the-art methods.
Xiaozhao Fang, Na Han, Guoxu Zhou, Shaohua Teng, Yong Xu 0001, Shengli Xie 0001
IEEE Trans. Cybern.1
2022 Average Approximate Hashing-Based Double Projections Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval has attracted considerable attention for searching in large-scale multimedia databases because of its efficiency and effectiveness. As a powerful tool of data analysis, matrix factorization is commonly used to learn hash codes for cross-modal retrieval, but there are still many shortcomings. First, most of these methods only focus on preserving locality of data but they ignore other factors such as preserving reconstruction residual of data during matrix factorization. Second, the energy loss of data is not considered when the data of cross-modal are projected into a common semantic space. Third, the data of cross-modal are directly projected into a unified semantic space which is not reasonable since the data from different modalities have different properties. This article proposes a novel method called average approximate hashing (AAH) to address these problems by: 1) integrating the locality and residual preservation into a graph embedding framework by using the label information; 2) projecting data from different modalities into different semantic spaces and then making the two spaces approximate to each other so that a unified hash code can be obtained; and 3) introducing a principal component analysis (PCA)-like projection matrix into the graph embedding framework to guarantee that the projected data can preserve the main energy of data. AAH obtains the final hash codes by using an average approximate strategy, that is, using the mean of projected data of different modalities as the hash codes. Experiments on standard databases show that the proposed AAH outperforms several state-of-the-art cross-modal hashing methods.
Xiaozhao Fang, Kaihang Jiang, Na Han, Shaohua Teng, Guoxu Zhou, Shengli Xie 0001
IEEE Trans. Cybern.1
2022 Cross-Domain Recognition via Projective Cross-Reconstruction
abstract
This article proposes a novel data reconstruction method, called projective cross-reconstruction (PCR) for cross-domain recognition. The intrinsic philosophy behind PCR is that the data from different domains but with the same label have a strong correlation and thus they can be reconstructed with each other. To this end, we first rearrange the data of source and target domains to form two new cross-data matrices, and the data with the same label but from different domains can be arranged together. Then, we use two different projection matrices to project the new cross-data into two approximate subspaces and perform the cross-reconstruction without introducing any extra matrix as the reconstruction coefficient matrix. This guarantees that the data from different domains can be interlaced well and the data from different domains but sharing the same label can be aligned together. In doing so, the problem of cross-domain distribution mismatch is solved and a discriminative and transferrable feature representation can be obtained. Moreover, PCR integrates the classifier learning and feature representation learning into a unified framework so that these two tasks can be iteratively improved until the termination condition is met. Extensive experiments on six datasets validate the effectiveness of our proposed PCR, compared with the state-of-the-art methods.
Xiaozhao Fang, Na Han, Weijun Sun, Yong Xu 0001, Shengli Xie 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2021 A Collaboration Multi-Domain Sentiment Classification on Specific Domain and Global Features
abstract
Sentiment classification has been attracting increasing attention with the growth of textual data created on the Internet. Text review data covers a wide range of field, and sentiment classification has been widely known as a highly domain-dependent problem. Unfortunately, the existing methods have achieved good results in the domain with a large number of labeled training data. Some researchers apply classifiers learned from source domain to target domain through transfer learning, which still requires the target domain to have enough unlabeled data to learn the similarity between the domains. In this paper, we propose a collaborative domain-specific and global multi-domain sentiment classification approaches with logistic regression. We train a domain-specific sentiment classifier for each source domain, reconstruct the source domain datasets, and train the global sentiment classifiers. Domain-specific sentiment classifier captures domain-specific sentiment features, and global sentiment classifier captures general sentiment knowledge. Finally, taking the output of the first layer as the input of the second layer, a two-level cross-domain sentiment classification model is constructed by logistic regression. Experimental results on benchmark datasets show that the proposed approach can effectively improve the performance of multi-domain sentiment classification and significantly outperform baseline methods.
Junping He, Shaohua Teng, Lunke Fei, Xiaozhao Fang, Wei Zhang 0005
CSCWD4
2021 Incomplete Multi-View Subspace Clustering with Low-Rank Tensor
abstract
Incomplete multi-view clustering has attracted increasing attentions due to its superiority in partitioning unlabeled multi-view data with missing instances in real application. However, most existing methods cannot fully exploit both the view-specific and cross-view relations among data points and ignore the high-order correlations across all views. To address these issues, we propose a novel Incomplete Multi-view Subspace Clustering with Low-rank Tensor (IMSCLT) method, which could be the first tensor-based incomplete multi-view clustering method to the best of our knowledge. Specifically, the subspace representations with low-rank tensor constraint are employed to exploit both the view-specific and cross-view relations among data points and capture the high-order correlations of multiple views simultaneously. In addition, we devise a novel module which can learn a discriminative similarity graph for multi-view learning task by approximating the inner product of the view-specific and common subspace representations. Augmented Lagrangian alternative direction minimization strategy is adopted to solve the proposed IMSCLT. The experiments on several benchmark datasets demonstrate the effectiveness of IMSCLT.
Jianlun Liu, Shaohua Teng, Wei Zhang 0005, Xiaozhao Fang, Lunke Fei, Zhuxiu Zhang
ICASSP4
2021 A novel consensus learning approach to incomplete multi-view clustering
Jianlun Liu, Shaohua Teng, Lunke Fei, Wei Zhang 0005, Xiaozhao Fang, Zhuxiu Zhang
Pattern Recognit.5
2020 Multiple Projections Learning for Dimensional Reduction
Xiaozhao Fang, Na Han
PDCAT2
2020 Double Relaxed Regression for Image Classification
abstract
This paper addresses two fundamental problems: 1) learning discriminative model parameters and 2) avoiding over-fitting, which often occurs in regression-based classification tasks. We formulate these two problems in terms of relaxing both the strict binary label matrix and graph regularization term into more flexible forms so that the margins between different classes are enlarged as much as possible and the problem of over-fitting is avoided to some extent. This task is accomplished by the proposed double relaxed regression (DRR) method. The convex problem of DRR is solved efficiently with an iterative procedure. Extensive experiments on synthetic and real world image data sets demonstrate the effectiveness of the proposed method in terms of both classification accuracy and running time.
Na Han, Jigang Wu, Xiaozhao Fang, Wai Keung Wong, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Active Transfer Learning
abstract
A major assumption in data mining and machine learning is that the training set and test set come from the same domain. They share the same feature space and have the same distribution. However, in many real-world applications, the training set and test set usually come from different domains. Thus, there might be negative similarities between different domains so that the negative transfer problem caused by negative similarity may happen. In this paper, we propose a novel method named active transfer learning (ATL) to solve the above problem. Specifically, the orthogonal projection matrix and the weight coefficient vector are introduced to extend maximum mean discrepancy (MMD) so that it can minimize MMD and simultaneously eliminate the negative transfer. To find the informative and discriminative subsets from the source domain, we then propose an information diversity term by using the local geometric structure information of the source samples. Besides, by using the label information of source samples, our method can guarantee the selected subsets as discriminative as possible. Finally, to efficiently implement the proposed method, an alternating optimization approach, which is based on the alternating direction method of multipliers (ADMM), is designed to solve the optimization problem. To demonstrate the effectiveness of the proposed ATL model, experiments are conducted on five real-world data sets. The experimental results show the superiority of our method over the state-of-the-art methods.
Zhihao Peng 0002, Wei Zhang 0005, Na Han, Xiaozhao Fang, Peipei Kang, Luyao Teng
IEEE Trans. Circuits Syst. Video Technol.4
2020 Clustering Structure-Induced Robust Multi-View Graph Recovery
abstract
Graph based classification methods have been widely applied in the fields of computer vision and machine learning. The quality of the graph highly affects the performance of these methods. The same object is commonly represented by different features, i.e., multi-view features, which leads to multiple graphs corresponding to different features in multi-view learning. However, what kind of graph is important for the task is unknown in advance. Moreover, existing multi-view learning methods become weak in dealing with noisy graphs when the data is corrupted by the noise. In this paper, we address this problem by observing that the noise of each graph has specific structure. Then, based on this observation we propose a robust multi-view graph recovery (RMGR) method in which the specific structure is used to clean the multiple input noisy graphs and these cleaned graphs are simultaneously aggregated into a consensus graph by adaptively assigning great weighted coefficients for important graphs. To make the consensus graph suit classification, the clustering structure is introduced to restrain the rank of Laplacian matrix of the consensus graph such that the number of its connected components is equal to that of clustering. In doing so, the graph is adaptively adjusted during optimization to more accurately partition data. The optimization problem is solved by proposed the iterative update algorithm. Extensive experiments on synthetic and several benchmark data sets show the effectiveness of the proposed method.
Wai Keung Wong, Na Han, Xiaozhao Fang, Shanhua Zhan, Jie Wen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Group Low-Rank Representation-Based Discriminant Linear Regression
abstract
In this paper, a novel least square regression method, named group low-rank representation-based discriminant linear regression (GLRRDLR), is proposed for multi-class classification. Unlike the conventional linear regression methods, the proposed method aims to learn a more discriminative projection. Specially, two main techniques are adopted to improve the discriminability of the projection. The first approach is to make the transformed samples locate in their own subspace by introducing a group low-rank constraint to the model, such that the distance between samples from the same class can be decreased greatly. The second approach is to simultaneously learn a discriminative target matrix for regression. The extensive experimental results show that the proposed method performs much better than the state-of-the-art methods, which proves the effectiveness of the above two approaches in improving the discriminability of the projection.
Shanhua Zhan, Jigang Wu, Na Han, Jie Wen 0001, Xiaozhao Fang
IEEE Trans. Circuits Syst. Video Technol.5
2020 Projective Double Reconstructions Based Dictionary Learning Algorithm for Cross-Domain Recognition
abstract
Dictionary learning plays a significant role in the field of machine learning. Existing works mainly focus on learning dictionary from a single domain. In this paper, we propose a novel projective double reconstructions (PDR) based dictionary learning algorithm for cross-domain recognition. Owing the distribution discrepancy between different domains, the label information is hard utilized for improving discriminability of dictionary fully. Thus, we propose a more flexible label consistent term and associate it with each dictionary item, which makes the reconstruction coefficients have more discriminability as much as possible. Due to the intrinsic correlation between cross-domain data, the data should be reconstructed with each other. Based on this consideration, we further propose a projective double reconstructions scheme to guarantee that the learned dictionary has the abilities of data itself reconstruction and data crossreconstruction. This also guarantees that the data from different domains can be boosted mutually for obtaining a good data alignment, making the learned dictionary have more transferability. We integrate the double reconstructions, label consistency constraint and classifier learning into a unified objective and its solution can be obtained by proposed optimization algorithm that is more efficient than the conventional l1 optimization based dictionary learning methods. The experiments show that the proposed PDR not only greatly reduces the time complexity for both training and testing, but also outperforms over the stateof- the-art methods.
Na Han, Jigang Wu, Xiaozhao Fang, Shaohua Teng, Guoxu Zhou, Shengli Xie 0001, Xuelong Li 0001
IEEE Trans. Image Process.3
2020 Latent Elastic-Net Transfer Learning
abstract
Subspace learning based transfer learning methods commonly find a common subspace where the discrepancy of the source and target domains is reduced. The final classification is also performed in such subspace. However, the minimum discrepancy does not guarantee the best classification performance and thus the common subspace may be not the best discriminative. In this paper, we propose a latent elastic-net transfer learning (LET) method by simultaneously learning a latent subspace and a discriminative subspace. Specifically, the data from different domains can be well interlaced in the latent subspace by minimizing Maximum Mean Discrepancy (MMD). Since the latent subspace decouples inputs and outputs and, thus a more compact data representation is obtained for discriminative subspace learning. Based on the latent subspace, we further propose a low-rank constraint based matrix elastic-net regression to learn another subspace in which the intrinsic intra-class structure correlations of data from different domains is well captured. In doing so, a better discriminative alignment is guaranteed and thus LET finally learns another discriminative subspace for classification. Experiments on visual domains adaptation tasks show the superiority of the proposed LET method.
Na Han, Jigang Wu, Xiaozhao Fang, Shengli Xie 0001, Shanhua Zhan, Kan Xie 0002, Xuelong Li 0001
IEEE Trans. Image Process.3
2020 Transferable Linear Discriminant Analysis
abstract
Linear discriminant analysis (LDA) has been widely used as the technique of feature exaction. However, LDA may be invalid to address the data from different domains. The reasons are as follows: 1) the distribution discrepancy of data may disturb the linear transformation matrix so that it cannot extract the most discriminative feature and 2) the original design of LDA does not consider the unlabeled data so that the unlabeled data cannot take part in the training process for further improving the performance of LDA. To address these problems, in this brief, we propose a novel transferable LDA (TLDA) method to extend LDA into the scenario in which the data have different probability distributions. The whole learning process of TLDA is driven by the philosophy that the data from the same subspace have a low-rank structure. The matrix rank in TLDA is the key learning criterion to conduct local and global linear transformations for restoring the low-rank structure of data from different distributions and enlarging the distances among different subspaces. In doing so, the variations of distribution discrepancy within the same subspace can be reduced, i.e., data can be aligned well and the maximally separated structure can be achieved for the data from different subspaces. A simple projected subgradient-based method is proposed to optimize the objective of TLDA, and a strict theory proof is provided to guarantee a quick convergence. The experimental evaluation on public data sets demonstrates that our TLDA can achieve better classification performance and outperform the state-of-the-art methods.
Na Han, Jigang Wu, Xiaozhao Fang, Jie Wen 0001, Shanhua Zhan, Shengli Xie 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2019 Error-correcting Ability based Collaborative Multi-Layer Selective Classifier Ensemble Model for Intrusion Detection
abstract
Ensemble classifier, b y combining multiple classifiers, can often achieve better performance than single classifiers in intrusion detection. Although some ensemble methods have been used for intrusion detection, most of them directly fuse detection outputs after multiple classifiers a re generated. It potentially reduce the overall performance and flexibility. Aiming at achieving a high-precision intrusion detection model with good generalization performance and robustness, an error-correcting ability based collaborative multi-layer selective classifier ensemble model is proposed in this paper, named ML-SCEM. In the ML-SCEM, a novel multi-layer structure consisting of 5 continuous layers is designed, each layer of which is equivalent to a binary classification. In each layer, an error-correcting based selective classifier ensemble method(SCEM) is used to select the main classifier and error-correcting components from M preselected base classifiers to generate an ensemble classifier suitable for this layer classification category. Furthermore to improve time efficiency a nd detection performance, the original dataset is divided into 3 parts of TCP, UDP and ICMP according to the network protocol, so that the three parts are collaboratively detected. The performance of the proposed ML-SCEM is evaluated and compared on the NSL-KDD dataset. It achieves accuracy of 97.07%, false positive rate of 1.58% and efficiently detects various types of attacks.
Limin Lu, Shaohua Teng, Wei Zhang 0005, Dongning Liu, Xiaozhao Fang
CSCWD6
2019 Deep Semantic Space with Intra-class Low-rank Constraint for Cross-modal Retrieval
abstract
In this paper, a novel Deep Semantic Space learning model with Intra-class Low-rank constraint (DSSIL) is proposed for cross-modal retrieval, which is composed of two subnetworks for modality-specific representation learning, followed by projection layers for common space mapping. In particular, DSSIL takes into account semantic consistency to fuse the cross-modal data in a high-level common space, and constrains the common representation matrix within the same class to be low-rank, in order to induce the intra-class representations more relevant. More formally, two regularization terms are devised for the two aspects, which have been incorporated into the objective of DSSIL. To optimize the modality-specific subnetworks and the projection layers simultaneously by exploiting the gradient decent directly, we approximate the nonconvex low-rank constraint by minimizing a few smallest singular values of the intra-class matrix with theoretical analysis. Extensive experiments conducted on three public datasets demonstrate the competitive superiority of DSSIL for cross-modal retrieval compared with the state-of-the-art methods.
Peipei Kang, Zehang Lin, Zhenguo Yang, Xiaozhao Fang, Qing Li 0001, Wenyin Liu
ICMR4
2019 Protein fold recognition based on multi-view modeling
abstract
MOTIVATION: Protein fold recognition has attracted increasing attention because it is critical for studies of the 3D structures of proteins and drug design. Researchers have been extensively studying this important task, and several features with high discriminative power have been proposed. However, the development of methods that efficiently combine these features to improve the predictive performance remains a challenging problem. RESULTS: In this study, we proposed two algorithms: MV-fold and MT-fold. MV-fold is a new computational predictor based on the multi-view learning model for fold recognition. Different features of proteins were treated as different views of proteins, including the evolutionary information, secondary structure information and physicochemical properties. These different views constituted the latent space. The ε-dragging technique was employed to enlarge the margins between different protein folds, improving the predictive performance of MV-fold. Then, MV-fold was combined with two template-based methods: HHblits and HMMER. The ensemble method is called MT-fold incorporating the advantages of both discriminative methods and template-based methods. Experimental results on five widely used benchmark datasets (DD, RDD, EDD, TG and LE) showed that the proposed methods outperformed some state-of-the-art methods in this field, indicating that MV-fold and MT-fold are useful computational tools for protein fold recognition and protein homology detection and would be efficient tools for protein sequence analysis. Finally, we constructed an update and rigorous benchmark dataset based on SCOPe (version 2.07) to fairly evaluate the performance of the proposed method, and our method achieved stable performance on this new dataset. This new benchmark dataset will become a widely used benchmark dataset to fairly evaluate the performance of different methods for fold recognition. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ke Yan 0003, Xiaozhao Fang, Yong Xu 0001, Bin Liu 0014
Bioinform.2
2019 Unsupervised feature selection with adaptive residual preserving
Luyao Teng, Zhenye Feng, Xiaozhao Fang, Shaohua Teng, Hua Wang 0002, Peipei Kang, Yanchun Zhang
Neurocomputing3
2019 Unsupervised feature extraction by low-rank and sparsity preserving embedding
Shanhua Zhan, Jigang Wu, Na Han, Jie Wen 0001, Xiaozhao Fang
Neural Networks5
2019 Robust Sparse Linear Discriminant Analysis
abstract
Linear discriminant analysis (LDA) is a very popular supervised feature extraction method and has been extended to different variants. However, classical LDA has the following problems: 1) The obtained discriminant projection does not have good interpretability for features; 2) LDA is sensitive to noise; and 3) LDA is sensitive to the selection of number of projection directions. In this paper, a novel feature extraction method called robust sparse linear discriminant analysis (RSLDA) is proposed to solve the above problems. Specifically, RSLDA adaptively selects the most discriminative features for discriminant analysis by introducing the$l_{2,1}$norm. An orthogonal matrix and a sparse matrix are also simultaneously introduced to guarantee that the extracted features can hold the main energy of the original data and enhance the robustness to noise, and thus RSLDA has the potential to perform better than other discriminant methods. Extensive experiments on six databases demonstrate that the proposed method achieves the competitive performance compared with other state-of-the-art feature extraction methods. Moreover, the proposed method is robust to the noisy data.
Jie Wen 0001, Xiaozhao Fang, Jinrong Cui, Lunke Fei, Ke Yan 0003, Yan Chen 0018, Yong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 Low-Rank Preserving Projection Via Graph Regularized Reconstruction
abstract
Preserving global and local structures during projection learning is very important for feature extraction. Although various methods have been proposed for this goal, they commonly introduce an extra graph regularization term and the corresponding regularization parameter that needs to be tuned. However, tuning the parameter manually not only is time-consuming, but also is difficult to find the optimal value to obtain a satisfactory performance. This greatly limits their applications. Besides, projections learned by many methods do not have good interpretability and their performances are commonly sensitive to the value of the selected feature dimension. To solve the above problems, a novel method named low-rank preserving projection via graph regularized reconstruction (LRPP_GRR) is proposed. In particular, LRPP_GRR imposes the graph constraint on the reconstruction error of data instead of introducing the extra regularization term to capture the local structure of data, which can greatly reduce the complexity of the model. Meanwhile, a low-rank reconstruction term is exploited to preserve the global structure of data. To improve the interpretability of the learned projection, a sparse term with${l_{2,1}}$norm is imposed on the projection. Furthermore, we introduce an orthogonal reconstruction constraint to make the learned projection hold main energy of data, which enables LRPP_GRR to be more flexible in the selection of feature dimension. Extensive experimental results show the proposed method can obtain competitive performance with other state-of-the-art methods.
Jie Wen 0001, Na Han, Xiaozhao Fang, Lunke Fei, Ke Yan 0003, Shanhua Zhan
IEEE Trans. Cybern.3
2019 Flexible Affinity Matrix Learning for Unsupervised and Semisupervised Classification
abstract
In this paper, we propose a unified model called flexible affinity matrix learning (FAML) for unsupervised and semisupervised classification by exploiting both the relationship among data and the clustering structure simultaneously. To capture the relationship among data, we exploit the self-expressiveness property of data to learn a structured matrix in which the structures are induced by different norms. A rank constraint is imposed on the Laplacian matrix of the desired affinity matrix, so that the connected components of data are exactly equal to the cluster number. Thus, the clustering structure is explicit in the learned affinity matrix. By making the estimated affinity matrix approximate the structured matrix during the learning procedure, FAML allows the affinity matrix itself to be adaptively adjusted such that the learned affinity matrix can well capture both the relationship among data and the clustering structure. Thus, FAML has the potential to perform better than other related methods. We derive optimization algorithms to solve the corresponding problems. Extensive unsupervised and semisupervised classification experiments on both synthetic data and real-world benchmark data sets show that the proposed FAML consistently outperforms the state-of-the-art methods.
Xiaozhao Fang, Na Han, Wai Keung Wong, Shaohua Teng, Jigang Wu, Shengli Xie 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 A Collaborative Intrusion Detection Model using a novel optimal weight strategy based on Genetic Algorithm for Ensemble Classifier
abstract
Cybersecurity, especially intrusion detection, is becoming increasingly critical in our daily life. The intrusion detection systems (IDS) have been widely used to prevent disclosure of personal information and detect potentially suspicious attacks. Although many machine learning algorithms have been broadly applied to enhance the performance of IDS, low detection rate and high false alarm rate are still two critical problems. A collaborative and robust intrusion detection model using a novel optimal weight strategy based on Genetic Algorithm (GA) for ensemble classifier is proposed in this paper. Since network data stream can be divided into three categories according to network protocols, detectors are applied in the network protocol separately. All of the detectors can work collaboratively and efficiently. In the proposed model, GA is used to optimize the weight of each base classifier of ensemble classifier. In order to improve features quality, Principal Component Analysis (PCA) is used for dimension reduction and attribute extraction. The NSL-KDD datasets is used to test the effectiveness of the collaborative intrusion detection model. Experimental results show that the proposed model has a higher accuracy and better generalized performance than others in this field.
Shaohua Teng, Luyao Teng, Wei Zhang 0005, Haibin Zhu 0001, Xiaozhao Fang, Lunke Fei
CSCWD6
2018 Adaptive Locality Preserving based Discriminative Regression
abstract
Classical linear regression not only lacks of the flexibility in fitting the label, but also ignores to preserve the intrinsic local geometric structure of data, which leads to overfitting. In this paper, we propose a novel discriminative regression method, called adaptive locality preserving based discriminative regression (ALPDR), to address these problems. Firstly, a locality preserving constraint regularized by the adaptive weight is introduced to preserve the intrinsic geometric structures of data, in which the similar points of the same class are adaptively pulled together by the projection. Secondly, ALPDR directly learns the discriminative target matrix from data based on the given label information, which allows more freedom in label fitting and simultaneously enlarges the margins between different classes. Thirdly, ALPDR imposes a row-sparsity constraint on the projection, which enables the method to adaptively select the most discriminative features from data such that the negative influence of noises and redundant features can be eliminated. Finally, an efficient iterative algorithm is provided to optimize the model. Extensive experiments show that the proposed method outperforms the other state-of-art methods, which proves the effectiveness of the proposed method.
Jie Wen 0001, Lunke Fei, Zhihui Lai 0001, Zheng Zhang 0006, Xiaozhao Fang
ICPR6
2018 Low-rank and sparse embedding for dimensionality reduction
Na Han, Jigang Wu, Yingyi Liang, Xiaozhao Fang, Wai Keung Wong, Shaohua Teng
Neural Networks4
2018 Low-rank representation with adaptive graph regularization
Jie Wen 0001, Xiaozhao Fang, Yong Xu 0001, Chunwei Tian, Lunke Fei
Neural Networks2
2018 Robust Spectral Subspace Clustering Based on Least Square Regression
Zongze Wu 0001, Ming Yin 0002, Xiaozhao Fang, Shengli Xie 0001
Neural Process. Lett.4
2018 Approximate Low-Rank Projection Learning for Feature Extraction
abstract
Feature extraction plays a significant role in pattern recognition. Recently, many representation-based feature extraction methods have been proposed and achieved successes in many applications. As an excellent unsupervised feature extraction method, latent low-rank representation (LatLRR) has shown its power in extracting salient features. However, LatLRR has the following three disadvantages: 1) the dimension of features obtained using LatLRR cannot be reduced, which is not preferred in feature extraction; 2) two low-rank matrices are separately learned so that the overall optimality may not be guaranteed; and 3) LatLRR is an unsupervised method, which by far has not been extended to the supervised scenario. To this end, in this paper, we first propose to use two different matrices to approximate the low-rank projection in LatLRR so that the dimension of obtained features can be reduced, which is more flexible than original LatLRR. Then, we treat the two low-rank matrices in LatLRR as a whole in the process of learning. In this way, they can be boosted mutually so that the obtained projection can extract more discriminative features. Finally, we extend LatLRR to the supervised scenario by integrating feature extraction with the ridge regression. Thus, the process of feature extraction is closely related to the classification so that the extracted features are discriminative. Extensive experiments are conducted on different databases for unsupervised and supervised feature extraction, and very encouraging results are achieved in comparison with many state-of-the-arts methods.
Xiaozhao Fang, Na Han, Jigang Wu, Yong Xu 0001, Jian Yang 0003, Wai Keung Wong, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Robust Latent Subspace Learning for Image Classification
abstract
This paper proposes a novel method, called robust latent subspace learning (RLSL), for image classification. We formulate an RLSL problem as a joint optimization problem over both the latent SL and classification model parameter predication, which simultaneously minimizes: 1) the regression loss between the learned data representation and objective outputs and 2) the reconstruction error between the learned data representation and original inputs. The latent subspace can be used as a bridge that is expected to seamlessly connect the origin visual features and their class labels and hence improve the overall prediction performance. RLSL combines feature learning with classification so that the learned data representation in the latent subspace is more discriminative for classification. To learn a robust latent subspace, we use a sparse item to compensate error, which helps suppress the interference of noise via weakening its response during regression. An efficient optimization algorithm is designed to solve the proposed optimization problem. To validate the effectiveness of the proposed RLSL method, we conduct experiments on diverse databases and encouraging recognition results are achieved compared with many state-of-the-arts methods.
Xiaozhao Fang, Shaohua Teng, Zhihui Lai 0001, Zhaoshui He, Shengli Xie 0001, Wai Keung Wong
IEEE Trans. Neural Networks Learn. Syst.1
2018 Regularized Label Relaxation Linear Regression
abstract
Linear regression (LR) and some of its variants have been widely used for classification problems. Most of these methods assume that during the learning phase, the training samples can be exactly transformed into a strict binary label matrix, which has too little freedom to fit the labels adequately. To address this problem, in this paper, we propose a novel regularized label relaxation LR method, which has the following notable characteristics. First, the proposed method relaxes the strict binary label matrix into a slack variable matrix by introducing a nonnegative label relaxation matrix into LR, which provides more freedom to fit the labels and simultaneously enlarges the margins between different classes as much as possible. Second, the proposed method constructs the class compactness graph based on manifold learning and uses it as the regularization item to avoid the problem of overfitting. The class compactness graph is used to ensure that the samples sharing the same labels can be kept close after they are transformed. Two different algorithms, which are, respectively, based on -norm and -norm loss functions are devised. These two algorithms have compact closed-form solutions in each iteration so that they are easily implemented. Extensive experiments show that these two algorithms outperform the state-of-the-art algorithms in terms of the classification accuracy and running time.
Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zhihui Lai 0001, Wai Keung Wong, Bingwu Fang
IEEE Trans. Neural Networks Learn. Syst.1
2017 Protein fold recognition based on sparse representation based classification
Ke Yan 0003, Yong Xu 0001, Xiaozhao Fang, Chun-Hou Zheng 0001, Bin Liu 0014
Artif. Intell. Medicine3
2017 Orthogonal self-guided similarity preserving projection for classification and clustering
Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zhihui Lai 0001, Shaohua Teng, Lunke Fei
Neural Networks1
2017 Low rank representation with adaptive distance penalty for semi-supervised subspace classification
Lunke Fei, Yong Xu 0001, Xiaozhao Fang, Jian Yang 0003
Pattern Recognit.3
2017 Low-Rank Embedding for Robust Image Feature Extraction
abstract
Robustness to noises, outliers, and corruptions is an important issue in linear dimensionality reduction. Since the sample-specific corruptions and outliers exist, the class-special structure or the local geometric structure is destroyed, and thus, many existing methods, including the popular manifold learning- based linear dimensionality methods, fail to achieve good performance in recognition tasks. In this paper, we focus on the unsupervised robust linear dimensionality reduction on corrupted data by introducing the robust low-rank representation (LRR). Thus, a robust linear dimensionality reduction technique termed low-rank embedding (LRE) is proposed in this paper, which provides a robust image representation to uncover the potential relationship among the images to reduce the negative influence from the occlusion and corruption so as to enhance the algorithm's robustness in image feature extraction. LRE searches the optimal LRR and optimal subspace simultaneously. The model of LRE can be solved by alternatively iterating the argument Lagrangian multiplier method and the eigendecomposition. The theoretical analysis, including convergence analysis and computational complexity, of the algorithms is presented. Experiments on some well-known databases with different corruptions show that LRE is superior to the previous methods of feature extraction, and therefore, it indicates the robustness of the proposed method. The code of this paper can be downloaded from http://www.scholat.com/laizhihui.
Wai Keung Wong, Zhihui Lai 0001, Jiajun Wen 0001, Xiaozhao Fang, Yuwu Lu
IEEE Trans. Image Process.4
2017 Rate of Convergence of the FOCUSS Algorithm
abstract
Focal underdetermined system solver (FOCUSS) is a powerful method for basis selection and sparse representation, where it employs the [Formula: see text]-norm with p ∈ (0,2) to measure the sparsity of solutions. In this paper, we give a systematical analysis on the rate of convergence of the FOCUSS algorithm with respect to p ∈ (0,2) . We prove that the FOCUSS algorithm converges superlinearly for and linearly for usually, but may superlinearly in some very special scenarios. In addition, we verify its rates of convergence with respect to p by numerical experiments.
Kan Xie 0002, Zhaoshui He, Andrzej Cichocki, Xiaozhao Fang
IEEE Trans. Neural Networks Learn. Syst.4
2016 Low-rank representation integrated with principal line distance for contactless palmprint recognition
Lunke Fei, Yong Xu 0001, Bob Zhang 0001, Xiaozhao Fang, Jie Wen 0001
Neurocomputing4
2016 Individualized learning for improving kernel Fisher discriminant analysis
Zizhu Fan, Yong Xu 0001, Xiaozhao Fang, David Zhang 0001
Pattern Recognit.4
2016 Robust Semi-Supervised Subspace Clustering via Non-Negative Low-Rank Representation
abstract
Low-rank representation (LRR) has been successfully applied in exploring the subspace structures of data. However, in previous LRR-based semi-supervised subspace clustering methods, the label information is not used to guide the affinity matrix construction so that the affinity matrix cannot deliver strong discriminant information. Moreover, these methods cannot guarantee an overall optimum since the affinity matrix construction and subspace clustering are often independent steps. In this paper, we propose a robust semi-supervised subspace clustering method based on non-negative LRR (NNLRR) to address these problems. By combining the LRR framework and the Gaussian fields and harmonic functions method in a single optimization problem, the supervision information is explicitly incorporated to guide the affinity matrix construction and the affinity matrix construction and subspace clustering are accomplished in one step to guarantee the overall optimum. The affinity matrix is obtained by seeking a non-negative low-rank matrix that represents each sample as a linear combination of others. We also explicitly impose the sparse constraint on the affinity matrix such that the affinity matrix obtained by NNLRR is non-negative low-rank and sparse. We introduce an efficient linearized alternating direction method with adaptive penalty to solve the corresponding optimization problem. Extensive experimental results demonstrate that NNLRR is effective in semi-supervised subspace clustering and robust to different types of noise than other state-of-the-art methods.
Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zhihui Lai 0001, Wai Keung Wong
IEEE Trans. Cybern.1
2016 Discriminative Transfer Subspace Learning via Low-Rank and Sparse Representation
abstract
In this paper, we address the problem of unsupervised domain transfer learning in which no labels are available in the target domain. We use a transformation matrix to transfer both the source and target data to a common subspace, where each target sample can be represented by a combination of source samples such that the samples from different domains can be well interlaced. In this way, the discrepancy of the source and target domains is reduced. By imposing joint low-rank and sparse constraints on the reconstruction coefficient matrix, the global and local structures of data can be preserved. To enlarge the margins between different classes as much as possible and provide more freedom to diminish the discrepancy, a flexible linear classifier (projection) is obtained by learning a non-negative label relaxation matrix that allows the strict binary label matrix to relax into a slack variable matrix. Our method can avoid a potentially negative transfer by using a sparse matrix to model the noise and, thus, is more robust to different types of noise. We formulate our problem as a constrained low-rankness and sparsity minimization problem and solve it by the inexact augmented Lagrange multiplier method. Extensive experiments on various visual domain adaptation tasks show the superiority of the proposed method over the state-of-the art methods. The MATLAB code of our method will be publicly available at http://www.yongxu.org/lunwen.html.
Yong Xu 0001, Xiaozhao Fang, Xuelong Li 0001, David Zhang 0001
IEEE Trans. Image Process.2
2015 Orthogonal self-guided similarity preserving projections
abstract
In this paper, we propose a novel unsupervised dimensionality reduction (DR) method called orthogonal self-guided similarity preserving projections (OSSPP), which seamlessly integrates the procedures of an adjacency graph learning and DR into a one step. Specifically, OSSPP projects the data into a low-dimensional subspace and simultaneously performs similarity preserving learning by using the similarity preserving regularization term in which the reconstruction coefficients of the projected data are used to encode the similarity structure information. An interesting finding is that the problem to determine the reconstruction coefficients can be converted into a weighted non-negative sparse coding problem without any explicit sparsity constraint. Thus the projections obtained by OSSPP contain natural discriminating information. Experimental results demonstrate that OSSPP outperforms state-of-the-art methods in DR.
Xiaozhao Fang, Yong Xu 0001, Zheng Zhang 0006, Zhihui Lai 0001, LinLin Shen
ICIP1
2015 Noise-free representation based classification and face recognition experiments
Yong Xu 0001, Xiaozhao Fang, Jane You, Yan Chen 0018, Hong Liu 0008
Neurocomputing2
2015 Learning a Nonnegative Sparse Graph for Linear Regression
abstract
Previous graph-based semisupervised learning (G-SSL) methods have the following drawbacks: 1) they usually predefine the graph structure and then use it to perform label prediction, which cannot guarantee an overall optimum and 2) they only focus on the label prediction or the graph structure construction but are not competent in handling new samples. To this end, a novel nonnegative sparse graph (NNSG) learning method was first proposed. Then, both the label prediction and projection learning were integrated into linear regression. Finally, the linear regression and graph structure learning were unified within the same framework to overcome these two drawbacks. Therefore, a novel method, named learning a NNSG for linear regression was presented, in which the linear regression and graph learning were simultaneously performed to guarantee an overall optimum. In the learning process, the label information can be accurately propagated via the graph structure so that the linear regression can learn a discriminative projection to better fit sample labels and accurately classify new samples. An effective algorithm was designed to solve the corresponding optimization problem with fast convergence. Furthermore, NNSG provides a unified perceptiveness for a number of graph-based learning methods and linear regression methods. The experimental results showed that NNSG can obtain very high classification accuracy and greatly outperforms conventional G-SSL methods, especially some conventional graph construction methods.
Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zhihui Lai 0001, Wai Keung Wong
IEEE Trans. Image Process.1
2014 Locality and similarity preserving embedding for feature selection
Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zizhu Fan, Hong Liu 0008, Yan Chen 0018
Neurocomputing1
2014 Modified minimum squared error algorithm for robust classification and face recognition experiments
Yong Xu 0001, Xiaozhao Fang, Qi Zhu 0001, Yan Chen 0018, Jane You, Hong Liu 0008
Neurocomputing2
2014 Enhancing sparsity via full rank decomposition for robust face recognition
Yuwu Lu, Jinrong Cui, Xiaozhao Fang
Neural Comput. Appl.3
2014 Kernel linear regression for face recognition
Yuwu Lu, Xiaozhao Fang, Binglei Xie
Neural Comput. Appl.2
2014 Data Uncertainty in Face Recognition
abstract
The image of a face varies with the illumination, pose, and facial expression, thus we say that a single face image is of high uncertainty for representing the face. In this sense, a face image is just an observation and it should not be considered as the absolutely accurate representation of the face. As more face images from the same person provide more observations of the face, more face images may be useful for reducing the uncertainty of the representation of the face and improving the accuracy of face recognition. However, in a real world face recognition system, a subject usually has only a limited number of available face images and thus there is high uncertainty. In this paper, we attempt to improve the face recognition accuracy by reducing the uncertainty. First, we reduce the uncertainty of the face representation by synthesizing the virtual training samples. Then, we select useful training samples that are similar to the test sample from the set of all the original and synthesized virtual training samples. Moreover, we state a theorem that determines the upper bound of the number of useful training samples. Finally, we devise a representation approach based on the selected useful training samples to perform face recognition. Experimental results on five widely used face databases demonstrate that our proposed approach can not only obtain a high face recognition accuracy, but also has a lower computational complexity than the other state-of-the-art approaches.
Yong Xu 0001, Xiaozhao Fang, Xuelong Li 0001, Jane You, Hong Liu 0008, Shaohua Teng
IEEE Trans. Cybern.2
2012 Combine the clustering algorithm and representation-based algorithm for concurrent classification of test samples
abstract
Sparse representation (SR) is a novel pattern recognition method. The algorithm of SR usually performs well. However, in processing a massive concurrent recognition task, SR has a very high computational cost because every test sample has to seek to an optimal linear combination of all the training samples. To this end, we propose a novel method which can perform well without needing to seek a linear combination of all the training samples for every test sample. Our proposed method can be divided into two steps: the first step of the proposed method uses c-means clustering to categorize the test sets into c subsets and then calculates K nearest neighbors for each class centre from all the training samples. The second step represents test samples located in each subset as a linear combination of the according K nearest neighbors and uses representation result to perform ultimate classification. A large number of experimental results show that the proposed algorithm is promising.
Xiaozhao Fang, Yong Xu 0001
CISDA1