VLDB 2026 Research / reviewers in the wild / expert
Qinghai Zheng
dblp:234/8977
· DBLP profile ↗
46ranked-venue papers
13as first author
43since 2021 · last 2026
0000-0002-8684-1577ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 5 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 20 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pseudo Multi-view K-means ClusteringabstractClustering with k-means is well-established and efficient, but often struggles with complex data distributions because the clustering performance hinges on how well the centroids capture the data distribution, and conventional k-means usually fails to produce representative centroids under such conditions. To address this limitation, we propose Pseudo Multi-view K-means Clustering (PMKC), a novel framework that simulates a multi-view learning paradigm within a single-view setting by generating multiple soft k-means decompositions. Each decomposition can be treated as an individual view and investigates a distinct perspective of the data. Specifically, to encourage complementary structure, we impose an independence constraint among cluster centers, and to integrate these diverse clusterings, we model the soft assignment matrices as a third-order tensor and apply low-rank regularization to extract a shared latent structure. This design not only enhances clustering robustness but also improves the stability and consistency of the final results. Experimental results on several benchmark datasets demonstrate that PMKC achieves superior clustering performance compared to state-of-the-art methods. Jinqian Chen, Jihua Zhu, Haoyu Tang 0002, Qinghai Zheng |
AAAI | 4 |
| 2026 | Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationabstractCurrent Zero-Shot Temporal Action Localization (ZSTAL) methods, whether training-based or training-free ones, still predominantly rely on a single, unified query to localize an entire action. This unified representation is fundamentally ill-suited for complex real-world activities, as it fails to capture their internal compositional structure and adapt to dynamic, multi-stage variations across videos. To address this, we regard ZSTAL as a compositional reasoning task and introduce CASCADE, a Context-Aware Staged Action DEcomposition framework. Inspired by the human cognitive process of perceiving context, decomposing events, and reconstructing instances, CASCADE follows a training-free pipeline. It first perceives the video's context by leveraging a Multimodal Large Language Model (MLLM) to both filter out irrelevant actions and then generate a rich, video-specific caption for each action present in the video. An LLM then decomposes this caption into multiple, temporally ordered stages, which serve as fine-grained queries to guide the MLLM in estimating frame-level confidence scores. Recognizing that this decomposition can fragment a single action, a novel hierarchical merging logic then reconstructs complete instances by intelligently fusing these preliminary temporal segments based on their semantic progression and coherence. Extensive experiments and ablation studies on THUMOS14 and ActivityNet-1.3 show that CASCADE not only sets a new state-of-the-art among training-free methods but, most notably, significantly outperforms all prior training-based approaches on ActivityNet-1.3. Haoyu Tang 0002, Tianyuan Liang, Han Jiang 0012, Qinghai Zheng, Yupeng Hu 0003 |
AAAI | 5 |
| 2026 | Uncertainty-aware multi-instance partial-label learning via evidential deep model
Gaowen Jie, Fumiao Wang, Gaojie Song, Luojun Lin, Yuanlong Yu 0001, Qinghai Zheng |
Neurocomputing | 6 |
| 2026 | Multi-View Clustering via Cross-View Alignment and Anchor-Guided RepresentationabstractDue to its effectiveness and efficiency, anchor based multi-view clustering (MVC) has recently attracted much attention. However, existing anchor-based methods often ignore the balance of anchor distribution and fail to fully leverage the complementary nature of view-specific information. To address these issues, we propose Cross-view Alignment and Anchor-guided Representation(CAAR), a unified framework that jointly optimizes anchor learning, anchor graph construction, and clustering partition. To be specific, CAAR employs an explicit alignment mechanism to preserve both cross-view consistency and view-specific characteristics, and introduces a cluster-aware prior to encourage balanced anchor allocation across semantic clusters. The model is optimized via an alternating minimization algorithm and achieves linear computational complexity. Extensive experiments demonstrate the effectiveness and efficiency of CAAR on six real-life datasets. Gaojie Song, Fumiao Wang, Gaowen Jie, Yuanlong Yu 0001, Qinghai Zheng |
IEEE Signal Process. Lett. | 5 |
| 2026 | From One Comes Two: A Tensorized Graph Learning Framework for Clustering
Qinghai Zheng, Jihua Zhu, Yuanlong Yu 0001, Haoyu Tang 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | Enhanced Residual Tensor Norm Minimization for Multiview Subspace ClusteringabstractThe low-rank tensor constraint is widely used in multiview subspace clustering (MSC) and has demonstrated promising clustering performance on many datasets. The key challenges in most existing low-rank tensor constraint-based methods include: 1) the choice of surrogate functions for the tensor rank and 2) the rotation operation applied to the tensor formed by stacking multiple subspace representations along the third mode. The latter plays a critical role in enhancing clustering performance in multiview settings. In this work, we rethink the low-rank tensor constraint and present the enhanced residual tensor norm (ERTN) for multiview subspace clustering, dubbed ERTN-MSC. To be specific, ERTN employs a novel surrogate for the tensor rank, based on the residual learning of singular values, which facilitates better exploitation of the structural information in multiview data. Furthermore, ERTN applies the tensor-singular value decomposition (t-SVD) on three modes of the tensor constructed by multiple subspaces, which generalizes the rotation operation of tensor and enables comprehensive exploration of both intraview information and interview information of multiview data. An augmented Lagrangian multiplier-based algorithm with a convergence guarantee is designed for optimization. Experiments conducted on several real-world multiview datasets demonstrate the effectiveness and competitiveness of our ERTN-MSC. Qinghai Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video LocalizationabstractZero-shot Natural Language Video Localization (NLVL) aims to automatically generate moments and corresponding pseudo queries from raw videos for the training of the localization model without any manual annotations. Existing approaches typically produce pseudo queries as simple words, which overlook the complexity of queries in real-world scenarios. Considering the powerful text modeling capabilities of large language models (LLMs), leveraging LLMs to generate complete queries that are closer to human descriptions is a potential solution. However, directly integrating LLMs into existing approaches introduces several issues, including insensitivity, isolation, and lack of regulation, which prevent the full exploitation of LLMs to enhance zero-shot NLVL performance. To address these issues, we propose BTDP, an innovative framework for Boundary-aware Temporal Dynamic Pseudo-supervision pairs generation. Our method contains two crucial operations: 1) Boundary Segmentation that identifies both visual boundaries and semantic boundaries to generate the atomic segments and activity descriptions, tackling the issue of insensitivity. 2) Context Aggregation that employs the LLMs with a self-evaluation process to aggregate and summarize global video information for optimized pseudo moment-query pairs, tackling the issue of isolation and lack of regulation. Comprehensive experimental results on the Charades-STA and ActivityNet Captions datasets demonstrate the effectiveness of our BTDP method. Xiongwen Deng, Haoyu Tang 0002, Han Jiang 0012, Qinghai Zheng, Jihua Zhu |
AAAI | 4 |
| 2025 | Towards Stable and Storage-efficient Dataset Distillation: Matching Convexified TrajectoryabstractThe rapid evolution of deep learning and large language models has led to an exponential growth in the demand for training data, prompting the development of Dataset Distillation methods to address the challenges of managing large datasets. Among these, Matching Training Trajectories (MTT) has been a prominent approach, which replicates the training trajectory of an expert network on real data with a synthetic dataset. However, our investigation found that this method suffers from three significant limitations: 1. Instability of expert trajectory generated by Stochastic Gradient Descent (SGD); 2. Low convergence speed of the distillation process; 3. High storage consumption of the expert trajectory. To address these issues, we offer a new perspective on understanding the essence of Dataset Distillation and MTT through a simple transformation of the objective function, and introduce a novel method called Matching Convexified Trajectory (MCT), which aims to provide better guidance for the student trajectory. MCT creates convex combinations of expert trajectories by selecting a few expert models, guiding student networks to converge quickly and stably. This trajectory is not only easier to store, but also enables continuous sampling strategies during the distillation process, ensuring thorough learning and fitting of the entire expert trajectory. The comprehensive experiment of three public datasets verified that MCT is superior to the traditional MTT method. Leon Wenliang Zhong, Haoyu Tang 0002, Qinghai Zheng, Yupeng Hu 0003, Weili Guan |
CVPR | 3 |
| 2025 | Neural Collision Detection for Constrained Grasp Pose Optimization in Cluttered EnvironmentsabstractRobust robotic grasping in cluttered environments presents a significant challenge, as existing methods often neglect the complex interactions between the gripper, objects, and obstacles, leading to collisions and grasping failures. To address this, we propose a framework that integrates collision avoidance as a core constraint within the grasp pose optimization process. Central to this framework is a Neural Collision Detection (NCD) network that takes scene configurations and grasp poses as inputs, producing a collision score that approximates traditional collision detection functions. The NCD network provides critical feedback for refining grasp predictions and demonstrates strong generalization across diverse environments, facilitating efficient collision detection and constrained grasp pose optimization. Additionally, we incorporate frictional force closure, geometric symmetry, and surface alignment as regularization terms within the optimization function, enhancing the physical stability and geometric plausibility of the generated grasps. Extensive experiments conducted in real-world environments show a significant improvement in grasp success rates, with robust generalization to previously unseen objects and scenarios. These results validate the efficacy of our framework, highlighting its potential for enabling reliable robotic manipulation in complex and cluttered environments. Longyuan Lin, Yixin Zhuang, Qinghai Zheng, Yuanlong Yu 0001 |
IROS | 4 |
| 2025 | FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person SearchabstractText-Based Person Search, which aims to retrieve target pedestrian images using natural language descriptions, has garnered significant attention in multimedia research due to its potential in suspect retrieval and missing person identification. While supervised and weakly supervised methods rely on costly annotated training data, unsupervised TBPS eliminates the need for textual descriptions or identity annotations, presenting a more practical paradigm. Current unsupervised TBPS approaches face two primary challenges: 1) Predefined attribute templates for caption generation limit linguistic diversity and real-world adaptability, and 2) Threshold-based sample selection using pre-trained vision-language models (VLMs) introduces noisy pairs due to inadequate pedestrian-specific representation. To address these limitations, we propose FACE, a unified framework featuring Dual-template Caption Generation (DCG) and Adaptive Curriculum Training (ACT). The DCG module generates high-quality captions through complementary flexible-style (natural language) and fixed-style (attribute-enumerated) templates, enhanced by LLM-based noise filtering. The ACT framework progressively refines training through a self-improving loop: initial high-confidence sample selection using VLMs bootstraps the model, while evolving feature representations enable dynamic incorporation of harder samples through curriculum learning. This dual strategy achieves mutual reinforcement between caption quality and model discriminability. Extensive experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets under unsupervised settings demonstrate that our framework achieves the state-of-the-art performance. Xiaoxuan Mu, Haoyu Tang 0002, Han Jiang 0012, Tianyuan Liang, Qinghai Zheng, Jihua Zhu |
ACM Multimedia | 5 |
| 2025 | Geometry-aware triplane diffusion for single shape generation with feature alignment
Hongliang Weng, Qinghai Zheng, Yuanlong Yu 0001, Yixin Zhuang |
Comput. Graph. | 2 |
| 2025 | Trusted Cross-view Completion for incomplete multi-view classification
Peihuan Song, Qinghai Zheng, Yuanlong Yu 0001 |
Neurocomputing | 4 |
| 2025 | Relationship completion for incomplete multi-view clustering
Minghong Wu, Jihua Zhu, Wenbiao Yan, Qinghai Zheng |
Neural Networks | 5 |
| 2025 | Partially multi-view clustering via re-alignment
Wenbiao Yan, Jihua Zhu, Jinqian Chen, Haozhe Cheng, Shunshun Bai, Liang Duan, Qinghai Zheng |
Neural Networks | 7 |
| 2025 | Cross-View Fusion for Multi-View ClusteringabstractMulti-view clustering has attracted significant attention in recent years because it can leverage the consistent and complementary information of multiple views to improve clustering performance. However, effectively fuse the information and balance the consistent and complementary information of multiple views are common challenges faced by multi-view clustering. Most existing multi-view fusion works focus on weighted-sum fusion and concatenating fusion, which unable to fully fuse the underlying information, and not consider balancing the consistent and complementary information of multiple views. To this end, we propose Cross-view Fusion for Multi-view Clustering (CFMVC). Specifically, CFMVC combines deep neural network and graph convolutional network for cross-view information fusion, which fully fuses feature information and structural information of multiple views. In order to balance the consistent and complementary information of multiple views, CFMVC enhances the correlation among the same samples to maximize the consistent information while simultaneously reinforcing the independence among different samples to maximize the complementary information. Experimental results on several multi-view datasets demonstrate the effectiveness of CFMVC for multi-view clustering task. Binqiang Huang, Qinghai Zheng, Yuanlong Yu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Supervised Contrastive Learning With Mixed Samples for Long-Tailed RecognitionabstractIn the domain of signal processing and deep learning, long-tailed data distributions present significant challenges due to the class imbalance in which a few classes contain a large number of samples, while most classes have far fewer. This imbalance hinders the ability of traditional models to effectively learn from minority classes. In this work, we focus on long-tailed supervised contrastive learning and introduce a novel approach termed Mixture-based Supervised Contrastive Learning (MixSCL), which integrates image mixing techniques into the supervised contrastive learning framework. By focusing on intra-class diversity and inter-class separability, our method aims to enhance the global uniformity of feature representations and improve model robustness. Specifically, MixSCL employs dual-stream projection heads designed to optimize separately for original and mixed samples, ensuring that the introduction of mixed samples does not distort the representations of original samples. We conduct extensive evaluations on bench mark datasets including CIFAR-100-LT and ImageNet-LT, which demonstrate that MixSCL achieves superior and more balanced performance in long-tailed scenarios. Peihuan Song, Luojun Lin, Yuanlong Yu 0001, Wenjie Yang 0005, Qinghai Zheng |
IEEE Signal Process. Lett. | 5 |
| 2025 | Non-Decreasing Concave Regularized Minimization for Principal Component AnalysisabstractAs a widely used method in signal processing, Principal Component Analysis (PCA) performs both the compression and the recovery of high dimensional data by leveraging the linear transformations. Considering the robustness of PCA, how to discriminate correct samples and outliers in PCA is a crucial and challenging issue. In this paper, we present a general model, which conducts PCA via a non-decreasing concave regularized minimization and is termed PCA-NCRM for short. Different from most existing PCA methods, which learn the linear transformations by minimizing the recovery errors between the recovered data and the original data in the least squared sense, our model adopts the monotonically non-decreasing concave function to enhance the ability of model in distinguishing correct samples and outliers. To be specific, PCA-NCRM enlarges the attention to samples with smaller recovery errors and diminishes the attention to samples with larger recovery errors at the same time. The proposed minimization problem can be efficiently addressed by employing an iterative re-weighting optimization. Experimental results on several datasets show the effectiveness of our model. Qinghai Zheng, Yixin Zhuang |
IEEE Signal Process. Lett. | 1 |
| 2025 | Graph Variational Multi-View ClusteringabstractMulti-view clustering (MVC) aims to extract consensus information from multi-source data and has developed rapidly. While generative model-based methods perform well by leveraging predefined priors, they often overlook inter-instance relationships, which are essential for high-quality clustering. To address this issue, we propose Graph Variational Multi-view Clustering (GVMVC), which integrates graph information into the generative process. Specifically, we treat both the original multi-view features and graph information from each view as observed data, guiding the learning of latent representations. The key principles of this approach are: 1) enhancing discriminative feature learning through graph integration, and 2) ensuring consistent multi-view learning via graph-based constraints. Extensive experiments show that GVMVC outperforms state-of-the-art methods across various datasets and metrics. Code is available at https://github.com/WenB777/GVMVC.git. Wenbiao Yan, Jihua Zhu, Jinqian Chen, Haozhe Cheng, Qinghai Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Causal Label EnhancementabstractLabel enhancement (LE) is still a challenging task to mitigate the dilemma of the lack of label distribution. Existing LE work typically focuses on primarily formulating a projection between feature space and label distribution space from discriminative model perspective, which preserves the relevance consistency that the sign of recovered label distribution should be consistent with the logical label. Different from previous algorithms, we formulate this problem from a causal perspective and present a novel LE method via the structured causal model (LESCM). Specifically, the proposed LESCM deliberates establishing the causal graph with assuming that label distribution is a cause of feature and logical label, which naturally satisfies the definition of label distribution learning (LDL). With capturing the underlying causal relationships, we can significantly boost the interpretability and identifiability of label enhancement. Meanwhile, except for the relevance consistency, LESCM are encouraged to sustain the order consistency that assigns higher description degree of the recovered label distribution to the positive labels, as compared with the negative labels. Empirically, sufficient experiments on several label distribution learning data sets validate the effectiveness of LESCM. Xinyuan Liu 0001, Jihua Zhu, Qinghai Zheng, Zhongyu Li 0002, Mingchen Zhu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Neighbor-Based Completion for Addressing Incomplete Multiview ClusteringabstractDriven by the complementarity and consistency inherent in multiview data, multiview clustering (MVC) has garnered widespread attention in various domains. Real-world data often encounters the issue of missing information, leading to a surge of interest in the domain of incomplete MVC (IMVC). Despite existing approaches having made significant progress in addressing IMVC, two significant challenges persist: 1) many alignment-based methodologies tend to overlook the topological relationships among instances and 2) the view representations based on completion lack reconstructive properties, casting doubt on their alignment with the actual view representations. In response, we present a novel approach termed neighbor-based completion for addressing IMVC (NBIMVC), which capitalizes on the topological information among instances and the consistent information across views. Specifically, our method uses autoencoders to learn feature representations for each view and leverages nearest-neighbor relationships between unique and complete instances to complete missing features in missing views. Subsequently, we enforce hard negative alignment constraints on complete paired instances in the feature space. Finally, we ensure the consistency of views in the semantic space by employing cluster information and a shared clustering network, which facilitates the final multiview categories output and effectively resolves the IMVC problem. Extensive experimental evaluations validate the efficacy of our proposed method, showcasing comparable or superior performance to existing approaches. Wenbiao Yan, Jihua Zhu, Yiyang Zhou, Jinqian Chen, Haozhe Cheng, Kun Yue, Qinghai Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Watch Your Head: Assembling Projection Heads to Save the Reliability of Federated ModelsabstractFederated learning encounters substantial challenges with heterogeneous data, leading to performance degradation and convergence issues. While considerable progress has been achieved in mitigating such an impact, the reliability aspect of federated models has been largely disregarded. In this study, we conduct extensive experiments to investigate the reliability of both generic and personalized federated models. Our exploration uncovers a significant finding: federated models exhibit unreliability when faced with heterogeneous data, demonstrating poor calibration on in-distribution test data and low uncertainty levels on out-of-distribution data. This unreliability is primarily attributed to the presence of biased projection heads, which introduce miscalibration into the federated models. Inspired by this observation, we propose the "Assembled Projection Heads" (APH) method for enhancing the reliability of federated models. By treating the existing projection head parameters as priors, APH randomly samples multiple initialized parameters of projection heads from the prior and further performs targeted fine-tuning on locally available data under varying learning rates. Such a head ensemble introduces parameter diversity into the deterministic model, eliminating the bias and producing reliable predictions via head averaging. We evaluate the effectiveness of the proposed APH method across three prominent federated benchmarks. Experimental results validate the efficacy of APH in model calibration and uncertainty estimation. Notably, APH can be seamlessly integrated into various federated approaches but only requires less than 30% additional computation cost for 100x inferences within large models. Jinqian Chen, Jihua Zhu, Qinghai Zheng, Zhongyu Li 0002 |
AAAI | 3 |
| 2024 | Incomplete Multi-View Clustering Via Inference and EvaluationabstractMulti-view clustering aims to improve the clustering performance by leveraging information from multiple views. Most existing works assume that all views are complete. However, samples in real-world scenarios cannot be always observed in all views, leading to the challenging problem of Incomplete Multi-View Clustering (IMVC). Although some attempts are made recently, they still suffer from the following two limitations: (1) they usually adopt shallow models, which are unable to sufficiently explore the consistency and complementary of multiple views; (2) they lack of a suitable measurement to evaluate the quality of the recovered data during the learning process. To address the aforementioned limitations, we introduce a novel Incomplete Multi-View Clustering via Inference and Evaluation (IMVC-IE). Specifically, IMVC-IE adopts the contrastive learning strategy on features of different views to excavate the underlying information from existing samples firstly. Subsequently, massive alternative simulated data are inferred for missing views and a novel evaluation strategy is presented to obtain the proper data for missing views completion. Extensive experiments are conducted and verify the effectiveness of our method. Binqiang Huang, Shoujie Lan, Qinghai Zheng, Yuanlong Yu 0001 |
ICASSP | 4 |
| 2024 | Graph-guided imputation-free incomplete multi-view clustering
Shunshun Bai, Qinghai Zheng, Xiaojin Ren, Jihua Zhu |
Expert Syst. Appl. | 2 |
| 2024 | MCoCo: Multi-level Consistency Collaborative multi-view clustering
Yiyang Zhou, Qinghai Zheng, Wenbiao Yan, Jihua Zhu |
Expert Syst. Appl. | 2 |
| 2024 | Graph-Driven deep Multi-View Clustering with self-paced learning
Shunshun Bai, Xiaojin Ren, Qinghai Zheng, Jihua Zhu |
Knowl. Based Syst. | 3 |
| 2024 | Multi-view representation learning with dual-label collaborative guidance
Xiaojin Ren, Shunshun Bai, Qinghai Zheng, Jihua Zhu |
Knowl. Based Syst. | 5 |
| 2024 | Double-level View-correlation Multi-view Subspace Clustering
Shoujie Lan, Qinghai Zheng, Yuanlong Yu 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Multi-view Semantic Consistency based Information Bottleneck for Clustering
Wenbiao Yan, Yiyang Zhou, Qinghai Zheng, Jihua Zhu |
Knowl. Based Syst. | 4 |
| 2024 | Flexible and Parameter-Free Graph Learning for Multi-View Spectral ClusteringabstractWith the extensive use of multi-view data in practice, multi-view spectral clustering has received a lot of attention. In this work, we focus on the following two challenges, namely, how to deal with the partially contradictory graph information among different views and how to conduct clustering without the parameter selection. To this end, we establish a novel graph learning framework, which avoids the linear combination of the partially contradictory graph information among different views and learns a unified graph for clustering without the parameter selection. Specifically, we introduce a flexible graph degeneration with a structured graph constraint to address the aforementioned challenging issues. Besides, our method can be employed to deal with large-scale data by using the bipartite graph. Experimental results show the effectiveness and competitiveness of our method, compared to several state-of-the-art methods. Qinghai Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Twin Reciprocal Completion for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering is an important and challenging task, which has attracted significant attention in recent years. The key objective of incomplete multi-view clustering is to excavate the underlying avaliable consistency of multi-view data, so as to enable the effective reconstruction of missing views for clustering. In this paper, we introduce a completion framework that deeply explores the underlying consistency and effectively completes the missing views. Following that, we propose a novel Twin Reciprocal Completion for Incomplete multi-view clustering, termed TRC-IMC for short. To be specific, TRC-IMC jointly conducts the Completion in Feature space (CF) and the Completion in Subspace (CS) to reciprocally complete the data with missing views. The underlying high-order consistency of multi-view data can be fully explored in both the feature space and subspace to guide the completion process of missing views. Extensive experiments are conducted on eight real-world multi-view datasets, and experimental results indicate the promising performance of our method, compared to several state-of-the-arts. Qinghai Zheng, Haoyu Tang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Label Information Bottleneck for Label EnhancementabstractIn this work, we focus on the challenging problem of Label Enhancement (LE), which aims to exactly recover label distributions from logical labels, and present a novel Label Information Bottleneck (LIB) method for LE. For the recovery process of label distributions, the label irrelevant information contained in the dataset may lead to unsatisfactory recovery performance. To address this limitation, we make efforts to excavate the essential label relevant information to improve the recovery performance. Our method formulates the LE problem as the following two joint processes: 1) learning the representation with the essential label relevant information, 2) recovering label distributions based on the learned representation. The label relevant information can be excavated based on the “bottleneck” formed by the learned representation. Significantly, both the label relevant information about the label assignments and the label relevant information about the label gaps can be explored in our method. Evaluation experiments conducted on several benchmark label distribution learning datasets verify the effectiveness and competitiveness of LIB. Our source codes are available at https://github.com/qinghai-zheng/LIBLE. Qinghai Zheng, Jihua Zhu, Haoyu Tang 0002 |
CVPR | 1 |
| 2023 | Towards Fast and Stable Federated Learning: Confronting Heterogeneity via Knowledge AnchorabstractFederated learning encounters a critical challenge of data heterogeneity, adversely affecting the performance and convergence of the federated model. Various approaches have been proposed to address this issue, yet their effectiveness is still limited. Recent studies have revealed that the federated model suffers severe forgetting in local training, leading to global forgetting and performance degradation. Although the analysis provides valuable insights, a comprehensive understanding of the vulnerable classes and their impact factors is yet to be established. In this paper, we aim to bridge this gap by systematically analyzing the forgetting degree of each class during local training across different communication rounds. Our observations are: (1) Both missing and non-dominant classes suffer similar severe forgetting during local training, while dominant classes show improvement in performance. (2) When dynamically reducing the sample size of a dominant class, catastrophic forgetting occurs abruptly when the proportion of its samples is below a certain threshold, indicating that the local model struggles to leverage a few samples of a specific class effectively to prevent forgetting. Motivated by these findings, we propose a novel and straightforward algorithm called Federated Knowledge Anchor (FedKA). Assuming that all clients have a single shared sample for each class, the knowledge anchor is constructed before each local training stage by extracting shared samples for missing classes and randomly selecting one sample per class for non-dominant classes. The knowledge anchor is then utilized to correct the gradient of each mini-batch towards the direction of preserving the knowledge of the missing and non-dominant classes. Extensive experimental results demonstrate that our proposed FedKA achieves fast and stable convergence, significantly improving accuracy on popular benchmarks. Jinqian Chen, Jihua Zhu, Qinghai Zheng |
ACM Multimedia | 3 |
| 2023 | Semantically consistent multi-view representation learning
Yiyang Zhou, Qinghai Zheng, Shunshun Bai, Jihua Zhu |
Knowl. Based Syst. | 2 |
| 2023 | Graph-Guided Unsupervised Multiview Representation LearningabstractWithout the valuable label information to guide the learning process, it is demanding to fully excavate and integrate the underlying information from different views to learn the unified multi-view representation. This paper focuses on this challenge and presents a novel method, termed Graph-guided Unsupervised Multi-view Representation Learning (GUMRL), taking full advantage of multi-view graph information during the learning process. To be specific, GUMRL jointly conducts the view-specific feature representation learning, which is under the guidance of graph information, and the unified feature representation learning, which fuses the underlying graph information of different views to learn the desired unified multi-view feature representation. Regarding downstream tasks, such as clustering and classification, the classic single-view algorithms can be directly performed on the learned unified multi-view representation. The designed objective function is effectively optimized based on an alternating direction minimization method, and experiments conducted on six real-world multi-view datasets show the effectiveness and competitiveness of our GUMRL, compared to several state-of-the-art methods. Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002, Haoyu Tang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Generalized Label Enhancement With Sample CorrelationsabstractRecently, label distribution learning (LDL) has drawn much attention in machine learning, where LDL model is learned from labelel instances. Different from single-label and multi-label annotations, label distributions describe the instance by multiple labels with different intensities and accommodate to more general scenes. Since most existing machine learning datasets merely provide logical labels, label distributions are unavailable in many real-world applications. To handle this problem, we propose two novel label enhancement methods, i.e., Label Enhancement with Sample Correlations (LESC) and generalized Label Enhancement with Sample Correlations (gLESC). More specifically, LESC employs a low-rank representation of samples in the feature space, and gLESC leverages a tensor multi-rank minimization to further investigate the sample correlations in both the feature space and label space. Benefitting from the sample correlations, the proposed methods can boost the performance of label enhancement. Extensive experiments on 14 benchmark datasets demonstrate the effectiveness and superiority of our methods. Qinghai Zheng, Jihua Zhu, Haoyu Tang 0002, Xinyuan Liu 0001, Zhongyu Li 0002, Huimin Lu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Semi-Supervised Label Distribution Learning with Co-regularization
Xinyuan Liu 0001, Jihua Zhu, Qinghai Zheng, Zhongyu Li 0002 |
Neurocomputing | 3 |
| 2022 | Orthogonal multi-view tensor-based learning for clustering
Shuangxun Ma, Yuehu Liu, Guangcan Liu, Qinghai Zheng, Chi Zhang 0020 |
Neurocomputing | 4 |
| 2022 | Large-Scale Multi-View Clustering via Fast Essential Subspace Representation LearningabstractLarge-scale Multi-View Clustering (LMVC) is a hot research problem in the fields of signal processing and machine learning, and many anchor-based multi-view subspace clustering algorithms are proposed in recent years. However, most existing methods usually concentrate on the issue of reducing the time cost and ignore the exploration of the complementary information during the clustering process. To this end, we propose a Fast Essential Subspace Representation Learning (FESRL) method for large-scale multi-view subspace clustering. Specifically, FESRL introduces the orthogonal transformation to investigate both the complementary and consensus information across multiple views. The essential subspace representation can be learned in a linear time cost. Experiments conducted on several benchmark datasets illustrate the competitiveness of the proposed method. Qinghai Zheng |
IEEE Signal Process. Lett. | 1 |
| 2022 | Collaborative Unsupervised Multi-View Representation LearningabstractIn this paper, we delve into the challenging problem in multi-view learning, namely unsupervised multi-view representation learning, the goal of which is to effectively integrate information from multiple views and learn the unified feature representation with comprehensive information in an unsupervised manner. Despite the progress attained in recent years, it is still a challenging issue since the correlations across multiple views are complex and difficult to model during the learning process, especially in the absence of label information. To address this problem, we introduce a novel method, termed Collaborative Unsupervised Multi-view Representation Learning (CUMRL), which benefits from the high-order view correlations of multi-view data by introducing a collaborative learning strategy. Specifically, the low-rank tensor constraint is employed and plays the role of a bridge, which links the view-specific compact learning and unified representation learning in CUMRL. Experiments demonstrate the effectiveness and competitiveness of the multi-view representation achieved by the proposed method for different learning tasks, compared to several state-of-the-art methods. Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Multi-Level Query Interaction for Temporal Language GroundingabstractUnderstanding what is happening in the surveillance video is important for human-machine interface in transportation systems, where temporal language grounding is one of the key tasks, targeting at localizing the desired moment in an untrimmed video with a given sentence query that is relevant to the moment. This task is challenging due to the following reasons: 1) the requirement of understanding the video contents and query semantics comprehensively, and 2) building the bridge between the cross-modal semantics. To tackle these problems, early methods first sample video clips and then match them with the sentence to find the most relevant one. To reduce the computational complexity associated with video clip sampling, recent methods directly predict the temporal boundaries of the desired moment on the fused features of the sentence and the video frames. However, all the previous methods often learn the word-level or phrase-level features of the sentence, or directly generates the global sentence representation by attention mechanisms or graph network. However, we argue that applying only word-level or phrase-level semantic information and cross-modal interactions is not enough to fully capture the correspondence between the video and the query. To this end, we proposed a novel Multi-level Query Exploration and Interaction (MQEI) model, which explores the semantics in both the word- and phrase-level and captures the multi-level interactions between the video and the query through an attention module. Extensive experiments on two public benchmark datasets ActivityNet Captions and Charades-STA demonstrate that the proposed model can outperform all the state-of-the-art methods consistently. Haoyu Tang 0002, Jihua Zhu, Lin Wang 0026, Qinghai Zheng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Multiview spectral clustering via complementary informationabstractAbstract In this article, multiview spectral clustering via complementary information (MSCC) is proposed, in which both the consensus information and the complementary information are explored for multiview clustering. In contrast to most multiview spectral clustering methods, the proposed MSCC considers the differences among multiple views and constructs a similarity matrix for clustering. Furthermore, a convex relaxation is employed and an algorithm that is based on the augmented Lagrange multiplier is proposed for optimizing the objective function of MSCC. In extensive experiments on five real‐world benchmark datasets, our proposed method outperforms two baselines and has significantly improved to several state‐of‐the‐art multiview clustering methods. Shuangxun Ma, Yuehu Liu, Qinghai Zheng, Yaochen Li, Zhichao Cui |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Multi-view subspace clustering networks with local and global graph information
Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002 |
Neurocomputing | 1 |
| 2021 | Bidirectional loss function for Label Enhancement and distribution learning
Xinyuan Liu 0001, Jihua Zhu, Qinghai Zheng, Zhongyu Li 0002, Ruixin Liu, Jun Wang 0024 |
Knowl. Based Syst. | 3 |
| 2020 | Label Enhancement with Sample Correlations via Low-Rank RepresentationabstractCompared with single-label and multi-label annotations, label distribution describes the instance by multiple labels with different intensities and accommodates to more-general conditions. Nevertheless, label distribution learning is unavailable in many real-world applications because most existing datasets merely provide logical labels. To handle this problem, a novel label enhancement method, Label Enhancement with Sample Correlations via low-rank representation, is proposed in this paper. Unlike most existing methods, a low-rank representation method is employed so as to capture the global relationships of samples and predict implicit label correlation to achieve label enhancement. Extensive experiments on 14 datasets demonstrate that the algorithm accomplishes state-of-the-art results as compared to previous label enhancement baselines. Haoyu Tang 0002, Jihua Zhu, Qinghai Zheng, Jun Wang 0024, Shanmin Pang, Zhongyu Li 0002 |
AAAI | 3 |
| 2020 | Feature concatenation multi-view subspace clustering
Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002, Shanmin Pang, Jun Wang 0024, Yaochen Li |
Neurocomputing | 1 |
| 2020 | Constrained bilinear factorization multi-view subspace clustering
Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002, Shanmin Pang, Xiuyi Jia |
Knowl. Based Syst. | 1 |