VLDB 2026 Research / reviewers in the wild / expert
Xiaofeng Cao 0002
dblp:117/3982-2
· DBLP profile ↗
52ranked-venue papers
11as first author
47since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 11 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 12 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyper-Opinion Vagueness Quantification for Robust Multimodal LearningabstractRobust Multimodal Learning (RML) aims to address the issues of unreliable predictions of multimodal models. Nevertheless, previous RML works often struggle to distinguish between different categories that rely on identical intra-modal cues, making ambiguous predictions. We defined this degree of ``uncertain'' in extracting discriminative features of a multimodal model as vagueness. Neglecting such vagueness, as previous RML works commonly do, will undermine the ability to extract unique semantics of each category in multimodal models, further resulting in worse robustness under disturbances that affect semantic representations. Additionally, this vagueness will lead the parameter updating processes towards unreliable fusion, thus diverting the learning processes of the multimodal model from learning unique features of each category. Based on the above insight, we propose a novel robust multimodal learning approach, termed Hyper-Opinion Quantifying Vagueness (HOQV). Specifically, we first introduce hyper-opinion to capture and quantify the vagueness of multimodal learning in discriminating representations of different categories. Moreover, to mitigate the interference in parameter updating of unreliable representations with high vagueness, we also design the Hyper-Opinion Gradient Modulation to guide the optimization processes. We evaluate our HOQV on six datasets with different disturbances, including noise and adversarial attack, and demonstrate that our proposed method achieves state-of-the-art performance consistently. Disen Hu, Xun Jiang 0001, Xiaofeng Cao 0002, Zheng Wang 0044, Jingkuan Song, Heng Tao Shen, Xing Xu 0001 |
AAAI | 3 |
| 2026 | Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New ThinkingabstractIn this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential relationships among agent decisions. Since each decision influences the subsequent choices, this forms a hierarchical and nested tree-like structure of interdependencies. While modeling tree-like data in Euclidean spaces could cause distortion, which results in a loss of agent decision structure information. Motivated by this, we reconsider model agent behaviors in hyperbolic space and propose the Hyperbolic Multi-Agent Representations (HMAR) method, which projects the agent behaviors into a Poincaré ball and leverages hyperbolic neural networks to learn agent policy representations. Additionally, we designed a contrastive loss function to train this network, minimizing the distance in feature space between different representations of the same agent while maximizing the distance between representations of distinct agents. Experimental results provide empirical evidence for the effectiveness of the HMAR method in cooperative and competitive environments, demonstrating the potential of hyperbolic agent representations for effective decision-making in multi-agent environments. Bohao Qu, Xiaofeng Cao 0002, Bing Li 0001, Menglin Zhang, Tuan-Anh Vu, Di Lin 0002, Qing Guo 0005 |
AAAI | 2 |
| 2026 | TR-DQ: Time-Rotation Diffusion QuantizationabstractDiffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impact of time-steps variation during sampling. At the same time, most current approaches fail to account for significant activations that cannot be eliminated, resulting in substantial performance degradation after quantization. To address these issues, we propose Time-Rotation Diffusion Quantization (TR-DQ), a novel quantization method incorporating time-step and rotation-based optimization. TR-DQ first divides the sampling process based on time-steps and applies a rotation matrix to smooth activations and weights dynamically. For different time-steps, a dedicated hyperparameter is introduced for adaptive timing modeling, which enables dynamic quantization across different time steps. Additionally, we also explore the compression potential of Classifier-Free Guidance (CFG-wise) to establish a foundation for subsequent work. TR-DQ achieves state-of-the-art (SOTA) performance on image generation and video generation tasks and a 1.38-1.89× speedup and 1.97-2.58× memory reduction in inference compared to existing quantization methods. Yihua Shao, Deyang Lin, Minxi Yan, Siyu Chen 0021, Fanhu Zeng, Minwen Liao, Ao Ma 0005, Ziyang Yan, Haozhe Wang 0002, Yan Wang 0068, Zhi Chen 0010, Xiaofeng Cao 0002, Haotong Qin, Hao Tang 0005, Jingcai Guo |
AAAI | 12 |
| 2026 | Positional Relation Contextual Mixing for Imbalanced Classification
Yucheng Jiang, Jiateng Li, Yuan Tian 0017, Jiangchao Yao, Xin Yu 0002, Wei Ye 0001, Xiaofeng Cao 0002 |
Mach. Learn. | 7 |
| 2026 | FOCUS: Frequency-Optimized Conditioning of diffUSion models for mitigating catastrophic forgetting during test-time adaptation
Gabriel Tjio, Jie Zhang 0002, Xulei Yang, Nhat Chung, Xiaofeng Cao 0002, Ivor W. Tsang, Chee Keong Kwoh 0001, Qing Guo 0005 |
Mach. Vis. Appl. | 6 |
| 2026 | Generalized Distribution Aggregation Protocol for Federated Statistical HeterogeneityabstractFederated heterogeneity refers to the disparities in data distributions, model architectures, and communication capabilities across various devices or institutional entities. In real-world scenarios, statistical heterogeneity can often lead to ineffective aggregation, severely impacting generalization performance and resulting in biased or unstable model weights. Theoretically, distributional robustness analysis indicates that the generalization performance of a learning model can be bounded with respect to any heterogeneity distribution. This insight motivates us to reconsider the aggregation strategy in federated statistical heterogeneity scenarios, and we thus propose a new weighting aggregation protocol that considers the generalization bound disagreement of each local model. Specifically, we estimate the upper and lower bounds of the second-order origin moment of the shifted distribution for the current local model, and using these bound disagreements as the aggregation proportions for weights in each communication round. Our experiments demonstrate that this proposed aggregation protocol significantly improves the performance of several representative Federated Learning algorithms on benchmark datasets. Xiaofeng Cao 0002, Ivor W. Tsang, James T. Kwok, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Time-Variant Image Inpainting via Interactive Distribution Transition EstimationabstractIn this work, we focus on a novel and practical task, i.e., Time-vAriant iMage inPainting (TAMP). The aim of TAMP is to restore a damaged target image by leveraging the complementary information from a reference image, where both images capture the same scene but with a significant time gap in between, i.e., time-variant images. Different from conventional reference-guided image inpainting, the reference image under TAMP setup presents significant content distinction to the target image and potentially also suffers from damages. Such an application frequently happens in our daily life to restore a damaged image by referring to another reference image, where there is no guarantee of the reference image's source and quality. In particular, our study finds that even SOTA reference-guided image inpainting methods fail to achieve plausible results due to the chaotic image complementation. To address such an ill-posed problem, we propose a novel Interactive Distribution Transition Estimation (InDiTE) module which interactively complements the time-variant images with appropriate semantics thus facilitate the restoration of damaged regions. To further boost the performance, we propose our TAMP solution, namely Interactive Distribution Transition Estimation-driven Diffusion (InDiTE-Diff), which integrates InDiTE with SOTA diffusion model and conducts latent cross-reference during sampling. Moreover, considering the lack of benchmarks for TAMP task, we newly assembled a dataset, i.e., TAMP-Street, based on existing image and mask datasets. We conduct experiments on the TAMP-Street datasets under two different time-variant image inpainting settings, which show our method consistently outperform SOTA reference-guided image inpainting methods for solving TAMP. Yun Xing 0001, Qing Guo 0005, Yihao Huang 0001, Xiaofeng Cao 0002, Luqi Gong, Di Lin 0002, Ivor W. Tsang, Lei Ma 0003 |
IEEE Trans. Image Process. | 5 |
| 2026 | ViDR-GNN: Vision Implicit Discriminative Reorganization Graph Neural NetworksabstractVision GNNs (ViGs) divide an image into multiple patches, treating these image patches as graph nodes. The image is represented by extracting explicit features from these patches as node features and constructing edge connections based on explicit dependencies. However, this explicit graph structure struggles to accurately capture deeper implicit dependencies. For example, at the node-level, implicit relationships include the intra-group consistency of local and global features belonging to the same semantic group and the inter-group distinction of features belonging to different semantic groups. At the graph-level, implicit relationships manifest in whether global consistency of edge connections can be established in the absence of direct edge connection supervision. These aspects are crucial for improving the accuracy of downstream tasks. Therefore, more effective learning of implicit dependencies in vision graph structures remains an area requiring further research. We designed the Discriminative Feature Reorganization (DFR) module to address implicit dependencies at the node-level. This module constructs a loss function using similarity measures between positive and negative sample feature pairs from adjacent layers of the neural network. By adjusting this loss function, the intra-group consistency and inter-group distinction of node-level local and global features can be enhanced. We also designed the Graph Structure Refinement (GSR) module. This module refines the consistency of graph-level implicit relationships of edge connections through interactive supervision of two graphs learned from adjacent layers of the neural network. Experimental results show that ViDR-GNN achieves significant performance improvements in image classification, object detection, and instance segmentation tasks. Xiaofeng Cao 0002, Li Peng 0004, Lijia Ma, Jielong Yang |
IEEE Trans. Multim. | 2 |
| 2026 | Universal Stabilization for Maximum Entropy Optimization in Reinforcement LearningabstractIn real-world decision-making tasks, it is critical for reinforcement learning (RL) methods to be both stable and robust. Maximum entropy RL methods typically generate a robust policy with entropy augmented reward. While incorporating entropy into the reward offers the benefit of exploration, it presents limited universal applicability and persistent convergence difficulties, such as suboptimal policy stabilization and unstable $Q$ value update. From optimization, we define these two issues as tremulous policy and spiky Q-function, investigating their underlying causes and relationships. Analysis with this, the maximum entropy principle leads to a spiky $Q$ -function update, which ultimately results in a tremulous policy. We thus introduce a beta-symmetric Kullback-Leibler (KL) divergence objective to mitigate such issues under the maximum entropy framework. With this objective function, the tremulous nature of the policy could be controlled with a large beta value. The spiky $Q$ -function could be avoided by annealing the entropy in the target $Q$ value, as the beta-symmetric KL divergence is an upper bound of the original reverse KL divergence. Theoretically, we prove that minimizing our new objective function results in a new policy that presents an improvement in the $Q$ value. Guaranteed by these results, we ultimately derive the optimal policy by iteratively updating the $Q$ value and policy, and we call this method max-entropy stable optimization (MeSO). Experimental results on the Mujoco and Roboschool platforms demonstrate that our algorithm maintains stability while offering better flexibility and overall performance. Xing Chen 0022, Yewen Li, Xiaofeng Cao 0002, Hechang Chen, Hengshuai Yao, Bo An 0001, Yi Chang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Concept Matching with Agent for Out-of-Distribution DetectionabstractThe remarkable achievements of Large Language Models (LLMs) have captivated the attention of both academia and industry, transcending their initial role in dialogue generation. To expand the usage scenarios of LLM, some works enhance the effectiveness and capabilities of the model by introducing more external information, which is called the agent paradigm. Based on this idea, we propose a new method that integrates the agent paradigm into out-of-distribution (OOD) detection task, aiming to improve its robustness and adaptability. Our proposed method, Concept Matching with Agent (CMA), employs neutral prompts as agents to augment the CLIP-based OOD detection process. These agents function as dynamic observers and communication hubs, interacting with both In-distribution (ID) labels and data inputs to form vector triangle relationships. This triangular framework offers a more nuanced approach than the traditional binary relationship, allowing for better separation and identification of ID and OOD inputs. Our extensive experimental results showcase the superior performance of CMA over both zero-shot and training-required methods in a diverse array of real-world scenarios. Yuxiao Lee, Xiaofeng Cao 0002, Jingcai Guo, Wei Ye 0001, Qing Guo 0005, Yi Chang 0001 |
AAAI | 2 |
| 2025 | PhysLight: Accurate rPPG Heart Rate Measurement with Adaptive Video RelightingabstractFacial video-based remote physiological measurement (rPPG) can non-invasively estimate vital signs, such as heart rate (HR), which often faces challenges under varying lighting conditions. We propose the PhysLight framework to enhance the accuracy of rPPG heart rate measurement through adaptive video relighting. Our approach subtly modifies illumination in video frames to improve detection accuracy while maintaining visual quality. The framework includes a GenLightNet to extract ideal lighting priors and a WipeLightNet module to refine poorly lit videos. Extensive evaluations on benchmark datasets show that our method significantly improves HR estimation reliability, outperforming existing baselines and enhancing non-contact physiological monitoring in diverse environments. Menglin Zhang, Xiaoxin Guo, Bohao Qu, Xiaofeng Cao 0002, Shuifa Sun, Qing Guo 0005 |
ICME | 4 |
| 2025 | Analytical Construction on Geometric Architectures: Transitioning from Static to Temporal Link PredictionabstractStatic systems exhibit diverse structural properties, such as hierarchical, scale-free, and isotropic patterns, where different geometric spaces offer unique advantages. Methods combining multiple geometries have proven effective in capturing these characteristics. However, real-world systems often evolve dynamically, introducing significant challenges in modeling their temporal changes. To overcome this limitation, we propose a unified cross-geometric learning framework for dynamic systems, which synergistically integrates Euclidean and hyperbolic spaces, aligning embedding spaces with structural properties through fine-grained substructure modeling. Our framework further incorporates a temporal state aggregation mechanism and an evolution-driven optimization objective, enabling comprehensive and adaptive modeling of both nodal and relational dynamics over time. Extensive experiments on diverse real-world dynamic graph datasets highlight the superiority of our approach in capturing complex structural evolution, surpassing existing methods across multiple metrics. Yadong Sun, Xiaofeng Cao 0002, Ivor W. Tsang, Heng Tao Shen |
ICML | 2 |
| 2025 | TsCA: On the Semantic Consistency Alignment via Conditional Transport for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize novel state-object compositions by leveraging the shared knowledge of their primitive components. Despite considerable progress, effectively calibrating the bias between semantically similar multimodal representations, as well as generalizing pre-trained knowledge to novel compositional contexts, remains an enduring challenge. In this paper, our interest is to revisit the conditional transport (CT) theory and its homology to the visual-semantics interaction in CZSL and further, propose a novel Trisets Consistency Alignment framework (dubbed TsCA) that well-addresses these issues. Concretely, we utilize three distinct yet semantically homologous sets, i.e., patches, primitives, and compositions, to construct pairwise CT costs to minimize their semantic discrepancies. To further ensure the consistency transfer within these sets, we implement a cycle-consistency constraint that refines the learning by guaranteeing the feature consistency of the self-mapping during transport flow, regardless of modality. Moreover, we extend the CT plans to an open-world setting, which enables the model to effectively filter out unfeasible pairs, thereby speeding up the inference as well as increasing the accuracy. Extensive experiments are conducted to verify the effectiveness of the proposed method. The code is available at https://github.com/keepgoingjkg/TsCA. Miaoge Li, Jingcai Guo, Xiaofeng Cao 0002, Zhijie Rao, Song Guo 0001 |
IJCAI | 5 |
| 2025 | Long-tailed Recognition with Model RebalancingabstractLong-tailed recognition is ubiquitous and challenging in deep learning and even in the downstream finetuning of foundation models, since the skew class distribution generally prevents the model generalization to the tail classes. Despite the promise of previous methods from the perspectives of data augmentation, loss rebalancing and decoupled training etc., consistent improvement in the broad scenarios like multi-label long-tailed recognition is difficult. In this study, we dive into the essential model capacity impact under long-tailed context, and propose a novel framework, Model Rebalancing (MORE), which mitigates imbalance by directly rebalancing the model's parameter space. Specifically, MORE introduces a low-rank parameter component to mediate the parameter space allocation guided by a tailored loss and sinusoidal reweighting schedule, but without increasing the overall model complexity or inference costs. Extensive experiments on diverse long-tailed benchmarks, spanning multi-class and multi-label tasks, demonstrate that MORE significantly improves generalization, particularly for tail classes, and effectively complements existing imbalance mitigation methods. These results highlight MORE's potential as a robust plug-and-play module in long-tailed settings. Jiaan Luo, Feng Hong 0004, Qiang Hu 0003, Xiaofeng Cao 0002, Feng Liu 0003, Jiangchao Yao |
NeurIPS | 4 |
| 2025 | Preference-driven Knowledge Distillation for Few-shot Node ClassificationabstractGraph neural networks (GNNs) can efficiently process text-attributed graphs (TAGs) due to their message-passing mechanisms, but their training heavily relies on the human-annotated labels. Moreover, the complex and diverse local topologies of nodes of real-world TAGs make it challenging for a single mechanism to handle. Large language models (LLMs) perform well in zero-/few-shot learning on TAGs but suffer from a scalability challenge. Therefore, we propose a preference-driven knowledge distillation (PKD) framework to synergize the complementary strengths of LLMs and various GNNs for few-shot node classification. Specifically, we develop a GNN-preference-driven node selector that effectively promotes prediction distillation from LLMs to teacher GNNs. To further tackle nodes' intricate local topologies, we develop a node-preference-driven GNN selector that identifies the most suitable teacher GNN for each node, thereby facilitating tailored knowledge distillation from teacher GNNs to the student GNN. Extensive experiments validate the efficacy of our proposed framework in few-shot node classification on real-world TAGs.
Our code can be available at <https://github.com/GEEX-Weixing/PKD>. Chunchun Chen, Rui Fan 0001, Xiaofeng Cao 0002, Sourav Medya, Wei Ye 0001 |
NeurIPS | 4 |
| 2025 | Towards a structured analysis of domain perturbations in Tubular Geometry
Xiaofeng Cao 0002, Yi Chang 0001 |
Pattern Recognit. | 2 |
| 2025 | Lighting is Unreliable: Adversarial Video Relighting Against rPPG Heart Rate Measurement
Menglin Zhang, Xiaoxin Guo, Xiaofeng Cao 0002, Shuifa Sun, Huazhu Fu, Qing Guo 0005 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Distributional Shortest-Path Graph KernelsabstractTraditional shortest-path graph kernels generate for each graph a histogram-like feature map, whose elements represent the number of occurrences of non-isomorphic shortest paths in this graph. The histogram-like feature map does not contain the distributions of the shortest paths within and across graphs, causing inaccurate graph similarities. To this end, we propose a novel graph kernel called the Distributional Shortest-Path (DSP) graph kernel to embrace both types of distribution information. Since the distribution of substructures (e.g., the shortest paths) follows a power law like that of words in natural language, we utilize neural language models to learn each node's distributional shortest-path feature map, encompassing the distributions and dependencies of the shortest paths in each graph. Moreover, we design the Partition Kernel (PK) to capture the dataset-wide distribution information of the shortest paths. PK projects similar (i.e., belonging to the same partition) distributional shortest-path node feature maps to the same point in the Reproducing Kernel Hilbert Space. Finally, Kernel Mean Embedding (KME) is applied to compute graph feature maps and efficiently construct the DSP graph kernel. Empirical experiments demonstrate that DSP outperforms state-of-the-art graph kernels on most benchmark datasets. Wei Ye 0001, Wengang Guo, Shuhao Tang, Xin Sun 0003, Xiaofeng Cao 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | A First-Order Multi-Gradient Algorithm for Multi-Objective Bi-Level OptimizationabstractIn this paper, we study the Multi-Objective Bi-Level Optimization (MOBLO) problem, where the upper-level subproblem is a multi-objective optimization problem and the lower-level subproblem is for scalar optimization. Existing gradient-based MOBLO algorithms need to compute the Hessian matrix, causing the computational inefficient problem. To address this, we propose an efficient first-order multi-gradient method for MOBLO, called FORUM. Specifically, we reformulate MOBLO problems as a constrained multi-objective optimization (MOO) problem via the value-function approach. Then we propose a novel multi-gradient aggregation method to solve the challenging constrained MOO problem. Theoretically, we provide the complexity analysis to show the efficiency of the proposed method and a non-asymptotic convergence result. Empirically, extensive experiments demonstrate the effectiveness and efficiency of the proposed FORUM method in different learning problems. In particular, it achieves state-of-the-art performance on three multi-task learning benchmark datasets. The code is available at https://github.com/Baijiong-Lin/FORUM. Feiyang Ye 0001, Baijiong Lin, Xiaofeng Cao 0002, Yu Zhang 0006, Ivor W. Tsang |
ECAI | 3 |
| 2024 | IRAD: Implicit Representation-driven Image Resampling against Adversarial AttacksabstractWe introduce a novel approach to counter adversarial attacks, namely, image resampling. Image resampling transforms a discrete image into a new one, simulating the process of scene recapturing or rerendering as specified by a geometrical transformation. The underlying rationale behind our idea is that image resampling can alleviate the influence of adversarial perturbations while preserving essential semantic information, thereby conferring an inherent advantage in defending against adversarial attacks. To validate this concept, we present a comprehensive study on leveraging image resampling to defend against adversarial attacks. We have developed basic resampling methods that employ interpolation strategies and coordinate shifting magnitudes. Our analysis reveals that these basic methods can partially mitigate adversarial attacks. However, they come with apparent limitations: the accuracy of clean images noticeably decreases, while the improvement in accuracy on adversarial examples is not substantial.We propose implicit representation-driven image resampling (IRAD) to overcome these limitations. First, we construct an implicit continuous representation that enables us to represent any input image within a continuous coordinate space. Second, we introduce SampleNet, which automatically generates pixel-wise shifts for resampling in response to different inputs. Furthermore, we can extend our approach to the state-of-the-art diffusion-based method, accelerating it with fewer time steps while preserving its defense capability. Extensive experiments demonstrate that our method significantly enhances the adversarial robustness of diverse deep models against various attacks while maintaining high accuracy on clean images. Tianlin Li, Xiaofeng Cao 0002, Ivor W. Tsang, Yang Liu 0003, Qing Guo 0005 |
ICLR | 3 |
| 2024 | Dual Expert Distillation Network for Generalized Zero-Shot Learning
Zhijie Rao, Jingcai Guo, Xiaocheng Lu, Jingming Liang, Jie Zhang 0076, Haozhao Wang, Kang Wei 0004, Xiaofeng Cao 0002 |
IJCAI | 8 |
| 2024 | Deep Hierarchical Graph Alignment Kernels
Shuhao Tang, Xiaofeng Cao 0002, Wei Ye 0001 |
IJCAI | 3 |
| 2024 | MetaRepair: Learning to Repair Deep Neural Networks from Repairing Experiences
Yun Xing 0001, Qing Guo 0005, Xiaofeng Cao 0002, Ivor W. Tsang, Lei Ma 0003 |
ACM Multimedia | 3 |
| 2024 | Geometry Awakening: Cross-Geometry Learning Exhibits Superiority over Individual StructuresabstractRecent research has underscored the efficacy of Graph Neural Networks (GNNs) in modeling diverse geometric structures within graph data. However, real-world graphs typically exhibit geometrically heterogeneous characteristics, rendering the confinement to a single geometric paradigm insufficient for capturing their intricate structural complexities. To address this limitation, we examine the performance of GNNs across various geometries through the lens of knowledge distillation (KD) and introduce a novel cross-geometric framework. This framework encodes graphs by integrating both Euclidean and hyperbolic geometries in a space-mixing fashion. Our approach employs multiple teacher models, each generating hint embeddings that encapsulate distinct geometric properties. We then implement a structure-wise knowledge transfer module that optimally leverages these embeddings within their respective geometric contexts, thereby enhancing the training efficacy of the student model. Additionally, our framework incorporates a geometric optimization network designed to bridge the distributional disparities among these embeddings. Experimental results demonstrate that our model-agnostic framework more effectively captures topological graph knowledge, resulting in superior performance of the student models when compared to traditional KD methodologies. Yadong Sun, Xiaofeng Cao 0002, Yu Wang 0152, Wei Ye 0001, Jingcai Guo, Qing Guo 0005 |
NeurIPS | 2 |
| 2024 | Sharpness-Aware Minimization Activates the Interactive Teaching's Understanding and OptimizationabstractTeaching is a potentially effective approach for understanding interactions among multiple intelligences. Previous explorations have convincingly shown that teaching presents additional opportunities for observation and demonstration within the learning model, such as data distillation and selection. However, the underlying optimization principles and convergence of interactive teaching lack theoretical analysis, and in this regard co-teaching serves as a notable prototype. In this paper, we discuss its role as a reduction of the larger loss landscape derived from Sharpness-Aware Minimization (SAM). Then, we classify it as an iterative parameter estimation process using Expectation-Maximization. The convergence of this typical interactive teaching is achieved by continuously optimizing a variational lower bound on the log marginal likelihood. This lower bound represents the expected value of the log posterior distribution of the latent variables under a scaled, factorized variational distribution. To further enhance interactive teaching's performance, we incorporate SAM's strong generalization information into interactive teaching, referred as Sharpness Reduction Interactive Teaching (SRIT). This integration can be viewed as a novel sequential optimization process. Finally, we validate the performance of our approach through multiple experiments. Xiaofeng Cao 0002, Ivor W. Tsang |
NeurIPS | 2 |
| 2024 | Breaking the curse of dimensional collapse in graph contrastive learning: A whitening perspective
Kai Guo 0003, Yizhen Zheng, Shirui Pan, Xiaofeng Cao 0002, Yi Chang 0001 |
Inf. Sci. | 5 |
| 2024 | Mentored Learning: Improving Generalization and Convergence of Student LearnerabstractStudent learners typically engage in an iterative process of actively updating its hypotheses, like active learning. While this behavior can be advantageous, there is an inherent risk of introducing mistakes through incremental updates including weak initialization, inaccurate or insignificant history states, resulting in expensive convergence cost. In this work, rather than solely monitoring the update of the learner's status, we propose monitoring the disagreement w.r.t. $\mathcal{F}^\mathcal{T}(\cdot)$ between the learner and teacher, and call this new paradigm “Mentored Learning”, which consists of `how to teach' and `how to learn'. By actively incorporating feedback that deviates from the learner's current hypotheses, convergence will be much easier to analyze without strict assumptions on learner's historical status, then deriving tighter generalization bounds on error and label complexity. Formally, we introduce an approximately optimal teaching hypothesis, $h^\mathcal{T}$, incorporating a tighter slack term $\left(1+\mathcal{F}^{\mathcal{T}}(\widehat{h}_t)\right)\Delta_t$ to replace the typical $2\Delta_t$ used in hypothesis pruning. Theoretically, we demonstrate that, guided by this teaching hypothesis, the learner can converge to tighter generalization bounds on error and label complexity compared to non-educated learners who lack guidance from a teacher: 1) the generalization error upper bound can be reduced from $R(h^*)+4\Delta_{T-1}$ to approximately $R(h^{\mathcal{T}})+2\Delta_{T-1}$, and 2) the label complexity upper bound can be decreased from $4 \theta\left(TR(h^{*})+2O(\sqrt{T})\right)$ to approximately $2\theta\left(2TR(h^{\mathcal{T}})+3 O(\sqrt{T})\right)$. To adhere strictly to our assumption, self-improvement of teaching is proposed when $h^\mathcal{T}$ loosely approximates $h^*$. In the context of learning, we further consider two teaching scenarios: instructing a white-box and black-box learner. Experiments validate this teaching concept and demonstrate superior generalization performance compared to fundamental active learning strategies, such as IWAL, IWAL-D, etc. Xiaofeng Cao 0002, Yaming Guo, Heng Tao Shen, Ivor W. Tsang, James T. Kwok |
J. Mach. Learn. Res. | 1 |
| 2024 | Diversifying Policies With Non-Markov Dispersion to Expand the Solution SpaceabstractPolicy diversity, encompassing the variety of policies an agent can adopt, enhances reinforcement learning (RL) success by fostering more robust, adaptable, and innovative problem-solving in the environment. The environment in which standard RL operates is usually modeled with a Markov Decision Process (MDP) as the theoretical foundation. However, in many real-world scenarios, the rewards depend on an agent's history of states and actions leading to a non-MDP. Under the premise of policy diffusion initialization, non-MDPs may have unstructured expanding solution space due to varying historical information and temporal dependencies. This results in solutions having non-equivalent closed forms in non-MDPs. In this paper, deriving diverse solutions for non-MDPs requires policies to break through the boundaries of the current solution space through gradual dispersion. The goal is to expand the solution space, thereby obtaining more diverse policies. Specifically, we first model the sequences of states and actions by a transformer-based method to learn policy embeddings for dispersion in the solution space, since the transformer has advantages in handling sequential data and capturing long-range dependencies for non-MDP. Then, we stack the policy embeddings to construct a dispersion matrix as the policy diversity measure to induce the policy dispersion in the solution space and obtain a set of diverse policies. Finally, we prove that if the dispersion matrix is positive definite, the dispersed embeddings can effectively enlarge the disagreements across policies, yielding a diverse expression for the original policy embedding distribution. Experimental results of both non-MDP and MDP environments show that this dispersion scheme can obtain more expressive diverse policies via expanding the solution space, showing more robust performance than the recent learning baselines. Bohao Qu, Xiaofeng Cao 0002, Yi Chang 0001, Ivor W. Tsang, Yew-Soon Ong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Improving Augmentation Consistency for Graph Contrastive Learning
Weixin Bu, Xiaofeng Cao 0002, Yizhen Zheng, Shirui Pan |
Pattern Recognit. | 2 |
| 2024 | Hyperbolic Uncertainty Aware Semantic SegmentationabstractSemantic segmentation (SS) aims to classify each pixel into one of the pre-defined classes. This task plays an important role in self-driving cars and autonomous drones. In SS, many works have shown that most misclassified pixels are commonly near object boundaries with high uncertainties. However, existing SS loss functions are not tailored to handle these uncertain pixels during training, as these pixels are usually treated equally as confidently classified pixels and cannot be embedded with arbitrary low distortion in Euclidean space, thereby degenerating the performance of SS. To overcome this problem, this paper designs a Hyperbolic Uncertainty Loss (HyperUL), which dynamically highlights the misclassified and high-uncertainty pixels in Hyperbolic space during training via the hyperbolic distances. The proposed HyperUL is model agnostic and can be easily applied to various neural architectures. After employing HyperUL to three recent SS models, the experimental results on Cityscapes, UAVid, and ACDC datasets reveal that the segmentation performance of existing SS models can be consistently improved. Additionally, reliable measurement of model uncertainty plays a key role in real-world applications such as autonomous controls of vehicles and drones. To meet this requirement, we propose the Hyperbolic Uncertainty Estimation method, which is easily implemented by only post-processing the generated Hyperbolic embeddings. By this approach, we can calculate the uncertainty values almost for free. Quantitative and qualitative results on Cityscapes, UAVid, and ACDC datasets verify that our proposed uncertainty estimation method usually outputs more meaningful results compared with popular MC-dropout and ensembling methods. Bike Chen, Wei Peng 0009, Xiaofeng Cao 0002, Juha Röning |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Transductive Reward Inference on GraphabstractIn this study, we present a transductive inference approach on that reward information propagation graph, which enables the effective estimation of rewards for unlabelled data in offline reinforcement learning. Reward inference is the key to learning effective policies in practical scenarios, while direct environmental interactions are either too costly or unethical and the reward functions are rarely accessible, such as in healthcare and robotics. Our research focuses on developing a reward inference method based on the contextual properties of information propagation on graphs that capitalizes on a constrained number of human reward annotations to infer rewards for unlabelled data. We leverage both the available data and limited reward annotations to construct a reward propagation graph, wherein the edge weights incorporate various influential factors pertaining to the rewards. Subsequently, we employ the constructed graph for transductive reward inference, thereby estimating rewards for unlabelled data. Furthermore, we establish the existence of a fixed point during several iterations of the transductive inference process and demonstrate its at least convergence to a local optimum. Empirical evaluations on locomotion and robotic manipulation tasks validate the effectiveness of our approach. The application of our inferred rewards improves the performance in offline reinforcement learning tasks. Bohao Qu, Xiaofeng Cao 0002, Qing Guo 0005, Yi Chang 0001, Ivor W. Tsang, Chengqi Zhang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Enhancing Locally Adaptive Smoothing of Graph Neural Networks Via Laplacian Node DisagreementabstractGraph neural networks (GNNs) are designed to perform inference on data described by graph-structured node features and topology information. From the perspective of graph signal denoising, the typical message passing schemes of GNNs act as a globally uniform smoothing that minimizes disagreements between embeddings of connected nodes. However, the level of smoothing over different regions of the graph should be different, especially for those inter-class regions. This deviation limits the expressiveness of GNNs, and then renders them fragile to over-smoothing, long-range dependencies, and non-homophily settings. In this paper, we find that the node disagreements of initial graph features can present more trustworthy constraints on node embeddings, thereby enhancing the locally adaptive smoothing of GNNs. To spread the inherent disagreements of nodes, we propose the Laplacian node disagreement to jointly measure the initial features and output embeddings. With such a measurement, we then present a new graph signal denoising objective deriving a more effective message passing scheme and further incorporate it into the GNN architecture, named Laplacian node disagreement-based GNN (LND-GNN). Learning from its output node representations, we integrate an auxiliary disagreement constraint into the overall classification loss. Experiments demonstrate the expressive ability of LND-GNN in the downstream semi-supervised node classification task. Yu Wang 0152, Liang Hu 0001, Xiaofeng Cao 0002, Yi Chang 0001, Ivor W. Tsang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Distribution Matching for Machine TeachingabstractMachine teaching is an inverse problem of machine learning that aims at steering the student toward its target hypothesis, in which the teacher has already known the student's learning parameters. Previous studies on machine teaching focused on balancing the teaching risk and cost to find the best teaching examples deriving from the student model. This optimization solver is in general ineffective when the student does not disclose any cue of the learning parameters. To supervise such a teaching scenario, this article presents a distribution matching-based machine teaching strategy via iteratively shrinking the teaching cost in a smooth surrogate, which eliminates boundary perturbations from the version space. Technically, our strategy could be redefined as a cost-controlled optimization process that finds the optimal teaching examples without further exploring the parameter distribution of the student. Then, given any limited teaching cost, the training examples would have a closed-form expression. Theoretical analysis and experiment results demonstrate the effectiveness of this strategy. Xiaofeng Cao 0002, Ivor W. Tsang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Refining Euclidean Obfuscatory Nodes Helps: A Joint-Space Graph Learning Method for Graph Neural NetworksabstractMany graph neural networks (GNNs) are inapplicable when the graph structure representing the node relations is unavailable. Recent studies have shown that this problem can be effectively solved by jointly learning the graph structure and the parameters of GNNs. However, most of these methods learn graphs by using either a Euclidean or hyperbolic metric, which means that the space curvature is assumed to be either constant zero or constant negative. Graph embedding spaces usually have nonconstant curvatures, and thus, such an assumption may produce some obfuscatory nodes, which are improperly embedded and close to multiple categories. In this article, we propose a joint-space graph learning (JSGL) method for GNNs. JSGL learns a graph based on Euclidean embeddings and identifies Euclidean obfuscatory nodes. Then, the graph topology near the identified obfuscatory nodes is refined in hyperbolic space. We also present a theoretical justification of our method for identifying obfuscatory nodes and conduct a series of experiments to test the performance of JSGL. The results show that JSGL outperforms many baseline methods. To obtain more insights, we analyze potential reasons for this superior performance. Zhaogeng Liu, Jielong Yang, Xiaofeng Cao 0002, Muhan Zhang, Hechang Chen, Yi Chang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Out-of-Distribution Generalization of Federated Learning via Implicit Invariant RelationshipsabstractOut-of-distribution generalization is challenging for non-participating clients of federated learning under distribution shifts. A proven strategy is to explore those invariant relationships between input and target variables, working equally well for non-participating clients. However, learning invariant relationships is often in an explicit manner from data, representation, and distribution, which violates the federated principles of privacy-preserving and limited communication. In this paper, we propose FedIIR, which implicitly learns invariant relationships from parameter for out-of-distribution generalization, adhering to the above principles. Specifically, we utilize the prediction disagreement to quantify invariant relationships and implicitly reduce it through inter-client gradient alignment. Theoretically, we demonstrate the range of non-participating clients to which FedIIR is expected to generalize and present the convergence results for FedIIR in the massively distributed with limited communication. Extensive experiments show that FedIIR significantly outperforms relevant baselines in terms of out-of-distribution generalization of federated learning. Yaming Guo, Kai Guo 0003, Xiaofeng Cao 0002, Tieru Wu, Yi Chang 0001 |
ICML | 3 |
| 2023 | Nonparametric Iterative Machine TeachingabstractIn this paper, we consider the problem of Iterative Machine Teaching (IMT), where the teacher provides examples to the learner iteratively such that the learner can achieve fast convergence to a target model. However, existing IMT algorithms are solely based on parameterized families of target models. They mainly focus on convergence in the parameter space, resulting in difficulty when the target models are defined to be functions without dependency on parameters. To address such a limitation, we study a more general task – Nonparametric Iterative Machine Teaching (NIMT), which aims to teach nonparametric target models to learners in an iterative fashion. Unlike parametric IMT that merely operates in the parameter space, we cast NIMT as a functional optimization problem in the function space. To solve it, we propose both random and greedy functional teaching algorithms. We obtain the iterative teaching dimension (ITD) of the random teaching algorithm under proper assumptions, which serves as a uniform upper bound of ITD in NIMT. Further, the greedy teaching algorithm has a significantly lower ITD, which reaches a tighter upper bound of ITD in NIMT. Finally, we verify the correctness of our theoretical findings with extensive experiments in nonparametric scenarios. Xiaofeng Cao 0002, Weiyang Liu, Ivor W. Tsang, James T. Kwok |
ICML | 2 |
| 2023 | Nonparametric Teaching for Multiple LearnersabstractWe study the problem of teaching multiple learners simultaneously in the nonparametric iterative teaching setting, where the teacher iteratively provides examples to the learner for accelerating the acquisition of a target concept. This problem is motivated by the gap between current single-learner teaching setting and the real-world scenario of human instruction where a teacher typically imparts knowledge to multiple students. Under the new problem formulation, we introduce a novel framework -- Multi-learner Nonparametric Teaching (MINT). In MINT, the teacher aims to instruct multiple learners, with each learner focusing on learning a scalar-valued target model. To achieve this, we frame the problem as teaching a vector-valued target model and extend the target model space from a scalar-valued reproducing kernel Hilbert space used in single-learner scenarios to a vector-valued space. Furthermore, we demonstrate that MINT offers significant teaching speed-up over repeated single-learner teaching, particularly when the multiple learners can communicate with each other. Lastly, we conduct extensive experiments to validate the practicality and efficiency of MINT. Xiaofeng Cao 0002, Weiyang Liu, Ivor W. Tsang, James T. Kwok |
NeurIPS | 2 |
| 2023 | Taming over-smoothing representation on heterophilic graphs
Kai Guo 0003, Xiaofeng Cao 0002, Zhining Liu 0002, Yi Chang 0001 |
Inf. Sci. | 2 |
| 2023 | Towards fidelity of graph data augmentation via equivariance
Bai Zhang, Yixing Gao 0001, Linbo Xie, Xiaofeng Cao 0002, Yixiang Shan, Jielong Yang |
Knowl. Based Syst. | 5 |
| 2023 | Data-Efficient Learning via Minimizing Hyperspherical EnergyabstractDeep learning on large-scale data is currently dominant nowadays. The unprecedented scale of data has been arguably one of the most important driving forces behind its success. However, there still exist scenarios where collecting data or labels could be extremely expensive, e.g., medical imaging and robotics. To fill up this gap, this paper considers the problem of data-efficient learning from scratch using a small amount of representative data. First, we characterize this problem by active learning on homeomorphic tubes of spherical manifolds. This naturally generates feasible hypothesis class. With homologous topological properties, we identify an important connection - finding tube manifolds is equivalent to minimizing hyperspherical energy (MHE) in physical geometry. Inspired by this connection, we propose a MHE-based active learning (MHEAL) algorithm, and provide comprehensive theoretical guarantees for MHEAL, covering convergence and generalization analysis. Finally, we demonstrate the empirical performance of MHEAL in a wide range of applications for data-efficient learning, including deep clustering, distribution matching, version space sampling, and deep active learning. Xiaofeng Cao 0002, Weiyang Liu, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | AdaNS: Adaptive negative sampling for unsupervised graph representation learning
Yu Wang 0152, Liang Hu 0001, Wanfu Gao, Xiaofeng Cao 0002, Yi Chang 0001 |
Pattern Recognit. | 4 |
| 2023 | Improving generalization of double low-rank representation using Schatten-p norm
Jiaoyan Zhao, Yongsheng Liang 0001, Shuangyan Yi, Qiangqiang Shen, Xiaofeng Cao 0002 |
Pattern Recognit. | 5 |
| 2022 | Robust active representation via ℓ2, p-norm constraints
Jiaoyan Zhao, Shuangyan Yi, Yongsheng Liang 0001, Wei Liu 0065, Xiaofeng Cao 0002 |
Knowl. Based Syst. | 5 |
| 2022 | Distribution Disagreement via Lorentzian Focal RepresentationabstractError disagreement-based active learning (AL) selects the data that maximally update the error of a classification hypothesis. However, poor human supervision (e.g., few labels, improper classifier parameters) may weaken or clutter this update; moreover, the computational cost of performing a greedy search to estimate the errors using a deep neural network is intolerable. In this paper, a novel disagreement coefficient based on distribution, not error, provides a tighter bound on label complexity, which further guarantees its generalization in hyperbolic space. The focal points derived from the squared Lorentzian distance, present more effective hyperbolic representations on aspherical distribution from geometry, replacing the typical euclidean, kernelized, and Poincaré centroids. Experiments on different deep AL tasks show that, the focal representation adopted in a tree-likeliness splitting, significantly performs better than typical baselines of centroid representations and error disagreement, and state-of-the-art neural network architectures-based AL, dramatically accelerating the learning process. Xiaofeng Cao 0002, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Cold-Start Active Sampling Via γ-TubeabstractActive learning (AL) improves the generalization performance for the current classification hypothesis by querying labels from a pool of unlabeled data. The sampling process is typically assessed by an informative, representative, or diverse evaluation policy. However, the policy, which needs an initial labeled set to start, may degenerate its performance in a cold-start hypothesis. In this article, we first show that typical AL sampling can be equivalently formulated as geometric sampling over minimum enclosing ballsMEB of this article denotes a conceptual geometry over the cluster in generalization analysis. In the SVM community, it is related to hard-margin support vector data description.(MEBs) of clusters. Following the γ -tube structure in geometric clustering, we then divide one MEB covering a cluster into two parts: 1) a γ -tube and 2) a γ -ball. By estimating the error disagreement between sampling in MEB and γ -ball, our theoretical insight reveals that γ -tube can effectively measure the disagreement of hypotheses in original space over MEB and sampling space over γ -ball. To tighten our insight, we present generalization analysis, and the results show that sampling in γ -tube can derive higher probability bound to achieve a nearly zero generalization error. With these analyses, we finally apply the informative sampling policy of AL over γ -tube to present a tube AL (TAL) algorithm against the cold-start sampling issue. As a result, the dependency between the querying process and the evaluation policy of active sampling can be alleviated. Experimental results show that by using the γ -tube structure to deal with cold-start sampling, TAL achieves the superior performance than standard AL evaluation baselines by presenting substantial accuracy improvements. Image edge recognition extends our theoretical results. Xiaofeng Cao 0002, Ivor W. Tsang, Jianliang Xu |
IEEE Trans. Cybern. | 1 |
| 2022 | Shattering Distribution for Active LearningabstractActive learning (AL) aims to maximize the learning performance of the current hypothesis by drawing as few labels as possible from an input distribution. Generally, most existing AL algorithms prune the hypothesis set via querying labels of unlabeled samples and could be deemed as a hypothesis-pruning strategy. However, this process critically depends on the initial hypothesis and its subsequent updates. This article presents a distribution-shattering strategy without an estimation of hypotheses by shattering the number density of the input distribution. For any hypothesis class, we halve the number density of an input distribution to obtain a shattered distribution, which characterizes any hypothesis with a lower bound on VC dimension. Our analysis shows that sampling in a shattered distribution reduces label complexity and error disagreement. With this paradigm guarantee, in an input distribution, a Shattered Distribution-based AL (SDAL) algorithm is derived to continuously split the shattered distribution into a number of representative samples. An empirical evaluation of benchmark data sets further verifies the effectiveness of the halving and querying abilities of SDAL in real-world AL tasks with limited labels. Experiments on active querying with adversarial examples and noisy labels further verify our theoretical insights on the performance disagreement of the hypothesis-pruning and distribution-shattering strategies. Our code is available at https://github.com/XiaofengCao-MachineLearning/Shattering-Distribution-for-Active-Learning. Xiaofeng Cao 0002, Ivor W. Tsang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | High-dimensional cluster boundary detection using directed Markov tree
Xiaofeng Cao 0002 |
Pattern Anal. Appl. | 1 |
| 2020 | A divide-and-conquer approach to geometric sampling for active learning
Xiaofeng Cao 0002 |
Expert Syst. Appl. | 1 |
| 2020 | A structured perspective of volumes on active learning
Xiaofeng Cao 0002 |
Neurocomputing | 1 |
| 2019 | Learning Image-Specific Attributes by Hyperbolic Neighborhood Graph PropagationabstractAs a kind of semantic representation of visual object descriptions, attributes are widely used in various computer vision tasks. In most of existing attribute-based research, class-specific attributes (CSA), which are class-level annotations, are usually adopted due to its low annotation cost for each class instead of each individual image. However, class-specific attributes are usually noisy because of annotation errors and diversity of individual images. Therefore, it is desirable to obtain image-specific attributes (ISA), which are image-level annotations, from the original class-specific attributes. In this paper, we propose to learn image-specific attributes by graph-based attribute propagation. Considering the intrinsic property of hyperbolic geometry that its distance expands exponentially, hyperbolic neighborhood graph (HNG) is constructed to characterize the relationship between samples. Based on HNG, we define neighborhood consistency for each sample to identify inconsistent samples. Subsequently, inconsistent samples are refined based on their neighbors in HNG. Extensive experiments on five benchmark datasets demonstrate the significant superiority of the learned image-specific attributes over the original class-specific attributes in the zero-shot object classification task. Ivor W. Tsang, Xiaofeng Cao 0002, Ruiheng Zhang 0001, Chuancai Liu |
IJCAI | 3 |
| 2019 | BorderShift: toward optimal MeanShift vector for cluster boundary detection in high-dimensional data
Xiaofeng Cao 0002, Baozhi Qiu, Guandong Xu |
Pattern Anal. Appl. | 1 |
| 2019 | Multidimensional Balance-Based Cluster Boundary Detection for High-Dimensional DataabstractThe balance of neighborhood space around a central point is an important concept in cluster analysis. It can be used to effectively detect cluster boundary objects. The existing neighborhood analysis methods focus on the distribution of data, i.e., analyzing the characteristic of the neighborhood space from a single perspective, and could not obtain rich data characteristics. In this paper, we analyze the high-dimensional neighborhood space from multiple perspectives. By simulating each dimension of a data point's k nearest neighbors space ( k NNs) as a lever, we apply the lever principle to compute the balance fulcrum of each dimension after proving its inevitability and uniqueness. Then, we model the distance between the projected coordinate of the data point and the balance fulcrum on each dimension and construct the DHBlan coefficient to measure the balance of the neighborhood space. Based on this theoretical model, we propose a simple yet effective cluster boundary detection algorithm called Lever. Experiments on both low- and high-dimensional data sets validate the effectiveness and efficiency of our proposed algorithm. Xiaofeng Cao 0002, Baozhi Qiu, Zenglin Shi, Guandong Xu, Jianliang Xu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |