VLDB 2026 Research / reviewers in the wild / expert
Xuhui Fan 0001
dblp:117/4874
· DBLP profile ↗
45ranked-venue papers
18as first author
25since 2021 · last 2026
0000-0002-7558-7200ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 18 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyperbolic-Enhanced Mixture-of-Experts Mamba for Sequential RecommendationabstractSequential recommendation has emerged as a fundamental task in various domains, aiming to predict a user's next interaction based on historical behavior. Recent advances in deep sequence models, particularly Transformer-based architectures and the more recent Mamba, have substantially pushed the boundaries of sequential modeling performance. However, existing methods still face two critical challenges. First, many current approaches overlook the hierarchical structures and high-order dependencies among items, typically restricting representation learning to conventional Euclidean spaces, which limits their capacity to capture complex relational information. Second, although Mamba excels at long-range dependency modeling, its reliance on static Feed-Forward Networks (FFNs) hinders its ability to dynamically adapt to evolving user preferences across diverse contexts. To address these limitations, we propose a Hyperbolic-Enhanced Mixture-of-Experts Mamba recommender (HM2Rec) for sequential recommendation. HM2Rec first encodes user-item relationships through hyperbolic graph convolution to exploit hierarchical structure more effectively. Then, a Variational Graph Auto-Encoder (VGAE) is employed to reconstruct node embeddings, improving structural robustness. To further enhance sequential modeling, we integrate Rotary Positional Encoding (RoPE) into Mamba to better capture relative position dependencies, and replace the FFN with Mixture-of-Expert (MOE) module, enabling dynamic and personalized expert selection for each token. Our extensive experiments on four widely-used public datasets demonstrate that HM2Rec outperforms several advanced baseline models. Yuwen Liu 0003, Lianyong Qi, Xingyuan Mao, Weiming Liu 0005, Xuhui Fan 0001, Qiang Ni, Xuyun Zhang, Yang Zhang 0095, Amin Beheshti |
AAAI | 5 |
| 2026 | Federated neural nonparametric point processesabstractTemporal point processes (TPPs) are effective for modeling event occurrences over time but struggle with sparse and uncertain events in federated systems, where privacy is a major concern. To address this, we propose FedPP , a federated neural nonparametric point process model. FedPP integrates neural embeddings into sigmoidal Gaussian Cox processes (SGCPs) on the client side. SGCPs is a flexible and expressive class of TPPs, allowing FedPP to generate highly flexible intensity functions that capture client-specific event dynamics and uncertainties while efficiently summarizing historical records. For global aggregation, FedPP introduces a divergence-based mechanism to communicate the distributions of kernel hyperparameters in SGCPs between the server and clients, while keeping client-specific parameters local to ensure privacy and personalization. FedPP effectively captures event uncertainty and sparsity. Extensive experiments demonstrate its superior performance in federated settings, showing global aggregation with the KL divergence and the Wasserstein distance. Hui Chen 0026, Xuhui Fan 0001, Hengyu Liu 0001, Yaqiong Li, Zhi-Lin Zhao 0001, Feng Zhou 0011, Christopher J. Quinn, Longbing Cao |
Artif. Intell. | 2 |
| 2026 | FigBO: A Generalized Acquisition Function Framework with Look-Ahead Capability for Bayesian OptimizationabstractAbstract Bayesian optimization is a powerful technique for optimizing expensive-to-evaluate black-box functions, consisting of two main components: a surrogate model and an acquisition function. In recent years, myopic acquisition functions have been widely adopted for their simplicity and effectiveness. However, their lack of look-ahead capability limits their performance. To address this limitation, we propose FigBO, a generalized acquisition function that incorporates the future impact of candidate points on global information gain. FigBO is a plug-and-play method that can integrate seamlessly with most existing myopic acquisition functions. Theoretically, we analyze the regret bound and convergence rate of FigBO when combined with the myopic base acquisition function expected improvement (EI), comparing them to those of standard EI. Empirically, extensive experimental results across diverse tasks demonstrate that FigBO achieves state-of-the-art performance and significantly faster convergence compared to existing methods. Hui Chen 0026, Xuhui Fan 0001, Zhangkai Wu, Longbing Cao |
Mach. Learn. | 2 |
| 2026 | Data-Dependent Rectangular Bounding ProcessesabstractStochastic partition processes divide a multi-dimensional space into a number of regions, such that the data within each region exhibit some form of homogeneity. Due to the nature of their partition strategies, partition processes can often create many unnecessary divisions in sparse regions when trying to describe data in dense regions. To avoid this problem we introduce a parsimonious partition model - the Rectangular Bounding Process (RBP) - to efficiently partition multi-dimensional spaces, by employing a bounding strategy to enclose data points within rectangular bounding boxes. The RBP is self-consistent and as such can be directly extended from a finite hypercube to an infinite (unbounded) space. We extend the RBP to establish a data-dependent RBP (data-RBP) to generate bounding boxes only over existing data points in a sequential manner, which can effectively reduce model complexity and enable online learning. To achieve this, we design an alternative way to generate bounding boxes and prove the distributional equivalence between the data-RBP and the RBP when empty boxes are removed. We demonstrate application of the RBP and the data-RBP in three scenarios: regression trees, relational modelling, and random feature construction for online learning. Extensive experimental results validate the performance of the RBP and the data-RBP for both accuracy and efficiency. Xuhui Fan 0001, Bin Li 0015, Prosha Rahman, Scott A. Sisson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | PBDD: A Prompt-Based Learning Approach for Few-Shot Social Media Depression DetectionabstractAutomated detection of depressive moods from social media holds great promise for early mental health intervention, yet existing multimodal approaches typically require large quantities of annotated data and extensive feature engineering, impeding their deployment in real‐world settings where labels are scarce. To address this challenge, we propose prompt‐based depression detection (PBDD), a novel prompt‐based few‐shot learning framework that leverages frozen pretrained language and vision models to identify depression indicators from paired text‐image posts without fine‐tuning. Our method begins with rigorous data cleaning and sampling to construct a high‐quality few‐shot dataset, then encodes text via a masked language model and images via a self‐supervised rotation‐prediction task to capture deep semantic cues. Multimodal representations are seamlessly fused into a unified prompt template containing a [MASK] token, enabling the pre‐trained model to infer depressive states by language completion. Extensive experiments on both large‐scale and 1 % few‐shot subsets demonstrate that PBDD consistently outperforms state‐of‐the‐art baselines, achieving significant gains in accuracy and Macro‐F1. These results validate the effectiveness and scalability of our framework for depression detection under severe label scarcity, offering a practical solution for real‐time mental health monitoring in social media environments. Rui Wang 0034, Heyang Feng, Erik Cambria, Kaize Shi, Xiaohan Yu 0001, Xuhui Fan 0001, Xianxun Zhu |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Navigating Towards Fairness with Data SelectionabstractMachine learning algorithms often struggle to eliminate inherent data biases, particularly those arising from unreliable labels, which poses a significant challenge in ensuring fairness. Existing fairness techniques that address label bias typically involve modifying models and intervening in the training process, but these lack flexibility for large-scale datasets. To address this limitation, we introduce a data selection method designed to efficiently and flexibly mitigate label bias, tailored to more practical needs. Our approach utilizes a zero-shot predictor as a proxy model that simulates training on a clean holdout set. This strategy, supported by peer predictions, ensures the fairness of the proxy model and eliminates the need for an additional holdout set, which is a common requirement in previous methods. Without altering the classifier's architecture, our modality-agnostic method effectively selects appropriate training data and has proven efficient and effective in handling label bias and improving fairness across diverse datasets in experimental evaluations. Yixuan Zhang 0006, Zhidong Li, Yang Wang 0002, Fang Chen 0001, Xuhui Fan 0001, Feng Zhou 0011 |
AAAI | 5 |
| 2025 | Dynamic Spectral Graph Anomaly DetectionabstractGraph anomaly detection is crucial for identifying anomalous nodes within graphs and addressing applications like financial fraud detection and social spam detection. Recent spectral graph neural network methods advance graph anomaly detection by focusing on anomalies that notably affect the distribution of graph spectral energy. Such spectrum-based methods rely on two steps: graph wavelet extraction and feature fusion. However, both steps are hand-designed, capturing incomprehensive anomaly information of wavelet-specific features and resulting in their inconsistent feature fusion. To address these problems, we propose a dynamic spectral graph anomaly detection framework DSGAD to adaptively capture comprehensive anomaly information and perform consistent feature fusion. DSGAD introduces dynamic wavelets, consisting of trainable wavelets to adaptively learn anomalous patterns and capture wavelet-specific features with comprehensive anomaly information. Furthermore, the consistent fusion of wavelet-specific features achieves dynamic fusion by combining wavelet-specific feature extraction with energy difference and channel convolution fusion using location correlation. Experimental results on four datasets substantiate the efficacy of our DSGAD method, surpassing state-of-the-art methods in both homogeneous and heterogeneous graphs. Jianbo Zheng, Chao Yang 0015, Tairui Zhang, Longbing Cao, Bin Jiang 0006, Xuhui Fan 0001, Xiao-Ming Wu 0002, Xianxun Zhu |
AAAI | 6 |
| 2025 | SepDiff: Self-Encoding Parameter Diffusion for Learning Latent SemanticsabstractThe recently proposed Bayesian Flow Networks (BFNs) show great potential in modeling parameter spaces via a diffusion process, offering a unified strategy for handling continuous, discrete data. However, these parameter diffusion models cannot learn high-level semantic representation from the parameter space since common encoders, which encode data into one static representation, can- not capture semantic changes in parameters. This motivates a new direction: learning semantic representations hidden in the param- eter spaces to characterize noisy data. Accordingly, we propose a representation learning framework named SepDiff which operates in the parameter space to obtain parameter-wise latent semantics that exhibit progressive structures. Specifically, SepDiff proposes a self-encoder to learn latent semantics directly from parameters, rather than from observations. The encoder is then integrated into parameter diffusion model, enabling representation learning with various formats of observations. Mutual information terms further promote the disentanglement of latent semantics and capture mean- ingful semantics simultaneously. We illustrate seven representation learning tasks in SepDiff via expanding this parameter diffusion model, and extensive quantitative experimental results demonstrate the superior effectiveness of SepDiff in learning parameter repre- sentation. Zhangkai Wu, Xuhui Fan 0001, Jin Li 0028, Zhi-Lin Zhao 0001, Hui Chen 0026, Longbing Cao |
KDD (2) | 2 |
| 2025 | ProgDiffusion: Progressively Self-encoding Diffusion ModelsabstractLearning low-dimensional semantic representations in diffusion models (DMs) is an open task, since in standard DMs, the dimensions of its intermediate latents are the same as that of the observations and thus are unable to represent low-dimensional semantics. Existing methods address this task either by encoding observations into semantics which makes it difficult to generate samples without observations, or by synthesizing the U-Net's layers of pre-trained DMs into low-dimensional semantics, which is mainly used for downstream tasks rather than using semantics to facilitate the training process. Further, those generated static representations might not be aligned with dynamic timestep-wise intermediate latents. This work introduces a Progressive self-encoded Diffusion model (ProgDiffusion), which simultaneously learns semantic representations and reconstructs observations, does efficient unconditional generation, and produces progressively structured semantic representations. These benefits are gained by a novel self-encoder mechanism which takes the U-Net's upsampling features, intermediate latent and the denoising timestep as conditions to generate time-specific semantic representations, differing from existing work of conditioning on observations only. As a result, the learned intermediate latents are dynamic and mapped to a series of semantic representations that capture their gradual changes. Notably, our proposed encoder operates independently of the observations, making it feasible for unconditional generation as observations are not required. To evaluate ProgDiffusion, we design tasks to visualise the learned progressive semantic representations, in addition to other common tasks, which validate the effectiveness of ProgDiffusion against the state-of-the-art. The code is available at https://github.com/amasawa/ProgDiffusion. Zhangkai Wu, Xuhui Fan 0001, Longbing Cao |
KDD (1) | 2 |
| 2025 | SCoT: Unifying Consistency Models and Rectified Flows via Straight-Consistent TrajectoriesabstractPre-trained diffusion models are commonly used to generate clean data (e.g., images) from random noises, effectively forming pairs of noises and corresponding clean images. Distillation on these pre-trained models can be viewed as the process of constructing advanced trajectories within the pair to accelerate sampling. For instance, consistency model distillation develops consistent projection functions to regulate trajectories, although sampling efficiency remains a concern. Rectified flow method enforces straight trajectories to enable faster sampling, yet relies on numerical ODE solvers, which may introduce approximation errors. In this work, we bridge the gap between the consistency model and the rectified flow method by proposing a Straight-Consistent Trajectories~(SCoT) model. SCoT enjoys the benefits of both approaches for fast sampling, producing trajectories with consistent and straight properties simultaneously. These dual properties are strategically balanced by targeting two critical objectives: (1) regulating the gradient of SCoT's mapping function to a constant and (2) ensuring trajectory consistency. Extensive experimental results demonstrate the effectiveness and efficiency of SCoT. Zhangkai Wu, Xuhui Fan 0001, Longbing Cao |
NeurIPS | 2 |
| 2025 | FedSI: Federated Subnetwork Inference for Efficient Uncertainty QuantificationabstractWhile deep neural networks (DNNs)-based personalized federated learning (PFL) is demanding for addressing data heterogeneity and shows promising performance, existing methods for federated learning (FL) suffer from efficient systematic uncertainty quantification. The Bayesian DNNs-based PFL is usually questioned of either oversimplified model structures or high computational and memory costs. In this article, we introduce FedSI, a novel Bayesian DNNs-based subnetwork inference (SI) PFL framework. FedSI is simple and scalable by leveraging Bayesian methods to incorporate systematic uncertainties effectively. It implements a client-specific SI mechanism, selects network parameters with large variance to be inferred through posterior distributions, and fixes the rest as deterministic ones. FedSI achieves fast and scalable inference while preserving the systematic uncertainties to the fullest extent. Extensive experiments on four different benchmark datasets demonstrate that FedSI outperforms existing Bayesian and non-Bayesian FL baselines in heterogeneous FL scenarios. Hui Chen 0026, Hengyu Liu 0001, Zhangkai Wu, Xuhui Fan 0001, Longbing Cao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | TransFeat-TPP: An Interpretable Deep Covariate Temporal Point ProcessesabstractThe classical temporal point process (TPP) constructs an intensity function by taking the occurrence times into account. Nevertheless, occurrence time may not be the only relevant factor, other contextual data, termed covariates, may also impact the event evolution. Incorporating such covariates into the model is beneficial, while distinguishing their relevance to the event dynamics is of great practical significance. In this work, we propose a Transformer-based covariate temporal point process (TransFeat-TPP) model to improve the interpretability of deep covariate-TPPs while maintaining powerful expressiveness. TransFeat-TPP can effectively model complex relationships between events and covariates, and provide enhanced interpretability by discerning the importance of various covariates. Experimental results on synthetic and real datasets demonstrate improved prediction accuracy and consistently interpretable feature importance when compared to existing deep covariate-TPPs. Our code is available at https://github.com/waystogetthere/TransFeat.git. Zizhuo Meng, Boyu Li 0003, Xuhui Fan 0001, Zhidong Li, Yang Wang 0002, Fang Chen 0001, Feng Zhou 0011 |
ECAI | 3 |
| 2024 | Conditionally-Conjugate Gaussian Process Factor Analysis for Spike Count Data via Data AugmentationabstractGaussian process factor analysis (GPFA) is a latent variable modeling technique commonly used to identify smooth, low-dimensional latent trajectories underlying high-dimensional neural recordings. Specifically, researchers model spiking rates as Gaussian observations, resulting in tractable inference. Recently, GPFA has been extended to model spike count data. However, due to the non-conjugacy of the likelihood, the inference becomes intractable. Prior works rely on either black-box inference techniques, numerical integration or polynomial approximations of the likelihood to handle intractability. To overcome this challenge, we propose a conditionally-conjugate Gaussian process factor analysis (ccGPFA) resulting in both analytically and computationally tractable inference for modeling neural activity from spike count data. In particular, we develop a novel data augmentation based method that renders the model conditionally conjugate. Consequently, our model enjoys the advantage of simple closed-form updates using a variational EM algorithm. Furthermore, due to its conditional conjugacy, we show our model can be readily scaled using sparse Gaussian Processes and accelerated inference via natural gradients. To validate our method, we empirically demonstrate its efficacy through experiments. Yididiya Y. Nadew, Xuhui Fan 0001, Christopher J. Quinn |
ICML | 2 |
| 2024 | Revealing Distribution Discrepancy by Sampling Transfer in Unlabeled DataabstractThere are increasing cases where the class labels of test samples are unavailable, creating a significant need and challenge in measuring the discrepancy between training and test distributions. This distribution discrepancy complicates the assessment of whether the hypothesis selected by an algorithm on training samples remains applicable to test samples. We present a novel approach called Importance Divergence (I-Div) to address the challenge of test label unavailability, enabling distribution discrepancy evaluation using only training samples. I-Div transfers the sampling patterns from the test distribution to the training distribution by estimating density and likelihood ratios. Specifically, the density ratio, informed by the selected hypothesis, is obtained by minimizing the Kullback-Leibler divergence between the actual and estimated input distributions. Simultaneously, the likelihood ratio is adjusted according to the density ratio by reducing the generalization error of the distribution discrepancy as transformed through the two ratios. Experimentally, I-Div accurately quantifies the distribution discrepancy, as evidenced by a wide range of complex data scenarios and tasks. Zhi-Lin Zhao 0001, Longbing Cao, Xuhui Fan 0001, Wei-Shi Zheng 0001 |
NeurIPS | 3 |
| 2024 | Nonstationary Sparse Spectral Permanental ProcessabstractExisting permanental processes often impose constraints on kernel types or stationarity, limiting the model's expressiveness. To overcome these limitations, we propose a novel approach utilizing the sparse spectral representation of nonstationary kernels.
This technique relaxes the constraints on kernel types and stationarity, allowing for more flexible modeling while reducing computational complexity to the linear level.
Additionally, we introduce a deep kernel variant by hierarchically stacking multiple spectral feature mappings, further enhancing the model's expressiveness to capture complex patterns in data. Experimental results on both synthetic and real-world datasets demonstrate the effectiveness of our approach, particularly in scenarios with pronounced data nonstationarity. Additionally, ablation studies are conducted to provide insights into the impact of various hyperparameters on model performance. Zicheng Sun, Yixuan Zhang 0006, Zenan Ling, Xuhui Fan 0001, Feng Zhou 0011 |
NeurIPS | 4 |
| 2023 | Free-Form Variational Inference for Gaussian Process State-Space ModelsabstractGaussian process state-space models (GPSSMs) provide a principled and flexible approach to modeling the dynamics of a latent state, which is observed at discrete-time points via a likelihood model. However, inference in GPSSMs is computationally and statistically challenging due to the large number of latent variables in the model and the strong temporal dependencies between them. In this paper, we propose a new method for inference in Bayesian GPSSMs, which overcomes the drawbacks of previous approaches, namely over-simplified assumptions, and high computational requirements. Our method is based on free-form variational inference via stochastic gradient Hamiltonian Monte Carlo within the inducing-variable formalism. Furthermore, by exploiting our proposed variational distribution, we provide a collapsed extension of our method where the inducing variables are marginalized analytically. We also showcase results when combining our framework with particle MCMC methods. We show that, on six real-world datasets, our approach can learn transition dynamics and latent states more accurately than competing methods. Xuhui Fan 0001, Edwin V. Bonilla, Terence J. O'Kane, Scott A. Sisson |
ICML | 1 |
| 2023 | Dynamic customer segmentation via hierarchical fragmentation-coagulation processes
Ling Luo 0002, Bin Li 0015, Xuhui Fan 0001, Yang Wang 0002, Irena Koprinska, Fang Chen 0001 |
Mach. Learn. | 3 |
| 2023 | Hawkes Processes With Stochastic Exogenous Effects for Continuous-Time Interaction ModellingabstractContinuous-time interaction data is usually generated under time-evolving environment. Hawkes processes (HP) are commonly used mechanisms for the analysis of such data. However, typical model implementations (such as e.g., stochastic block models) assume that the exogenous (background) interaction rate is constant, and so they are limited in their ability to adequately describe any complex time-evolution in the background rate of a process. In this paper, we introduce a stochastic exogenous rate Hawkes process (SE-HP) which is able to learn time variations in the exogenous rate. The model affiliates each node with a piecewise-constant membership distribution with an unknown number of changepoint locations, and allows these distributions to be related to the membership distributions of interacting nodes. The time-varying background rate function is derived through combinations of these membership functions. We introduce a stochastic gradient MCMC algorithm for efficient, scalable inference. The performance of the SE-HP is explored on real world, continuous-time interaction datasets, where we demonstrate that the SE-HP strongly outperforms comparable state-of-the-art methods. Xuhui Fan 0001, Yaqiong Li, Ling Chen 0006, Bin Li 0015, Scott A. Sisson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Smoothing graphons for modelling exchangeable relational data
Yaqiong Li, Xuhui Fan 0001, Ling Chen 0006, Bin Li 0015, Scott A. Sisson |
Mach. Learn. | 2 |
| 2022 | Supervised Categorical Metric Learning With Schatten p-NormsabstractMetric learning has been successful in learning new metrics adapted to numerical datasets. However, its development of categorical data still needs further exploration. In this article, we propose a method, called CPML for categorical projected metric learning, which tries to efficiently (i.e., less computational time and better prediction accuracy) address the problem of metric learning in categorical data. We make use of the value distance metric to represent our data and propose new distances based on this representation. We then show how to efficiently learn new metrics. We also generalize several previous regularizers through the Schatten p -norm and provide a generalization bound for it that complements the standard generalization bound for metric learning. The experimental results show that our method provides state-of-the-art results while being faster. Yaqiong Li, Xuhui Fan 0001, Éric Gaussier |
IEEE Trans. Cybern. | 2 |
| 2021 | Poisson-Randomised DirBN: Large Mutation is Needed in Dirichlet Belief NetworksabstractThe Dirichlet Belief Network (DirBN) was recently proposed as a promising deep generative model to learn interpretable deep latent distributions for objects. However, its current representation capability is limited since its latent distributions across different layers is prone to form similar patterns and can thus hardly use multi-layer structure to form flexible distributions. In this work, we propose Poisson-randomised Dirichlet Belief Networks (Pois-DirBN), which allows large mutations for the latent distributions across layers to enlarge the representation capability. Based on our key idea of inserting Poisson random variables in the layer-wise connection, Pois-DirBN first introduces a component-wise propagation mechanism to enable latent distributions to have large variations across different layers. Then, we develop a layer-wise Gibbs sampling algorithm to infer the latent distributions, leading to a larger number of effective layers compared to DirBN. In addition, we integrate out latent distributions and form a multi-stochastic deep integer network, which provides an alternative view on Pois-DirBN. We apply Pois-DirBN to relational modelling and validate its effectiveness through improved link prediction performance and more interpretable latent distribution visualisations. The code can be downloaded at https://github.com/xuhuifan/Pois_DirBN. Xuhui Fan 0001, Bin Li 0015, Yaqiong Li, Scott A. Sisson |
ICML | 1 |
| 2021 | Bayesian Nonparametric Space Partitions: A SurveyabstractBayesian nonparametric space partition (BNSP) models provide a variety of strategies for partitioning a D-dimensional space into a set of blocks, such that the data within the same block share certain kinds of homogeneity. BNSP models are applicable to many areas, including regression/classification trees, random feature construction, and relational modelling. This survey provides the first comprehensive review of this subject. We explore the current progress of BNSP research through three perspectives: (1) Partition strategies, where we review the various techniques for generating partitions and discuss their theoretical foundation, `self-consistency'; (2) Applications, where we detail the current mainstream usages of BNSP models and identify some potential future applications; and (3) Challenges, where we discuss current unsolved problems and possible avenues for future research. Xuhui Fan 0001, Bin Li 0015, Ling Luo 0002, Scott A. Sisson |
IJCAI | 1 |
| 2021 | Continuous-time edge modelling using non-parametric point processesabstractThe mutually-exciting Hawkes process (ME-HP) is a natural choice to model reciprocity, which is an important attribute of continuous-time edge (dyadic) data. However, existing ways of implementing the ME-HP for such data are either inflexible, as the exogenous (background) rate functions are typically constant and the endogenous (excitation) rate functions are specified parametrically, or inefficient, as inference usually relies on Markov chain Monte Carlo methods with high computational costs. To address these limitations, we discuss various approaches to model design, and develop three variants of non-parametric point processes for continuous-time edge modelling (CTEM). The resulting models are highly adaptable as they generate intensity functions through sigmoidal Gaussian processes, and so provide greater modelling flexibility than parametric forms. The models are implemented via a fast variational inference method enabled by a novel edge modelling construction. The superior performance of the proposed CTEM models is demonstrated through extensive experimental evaluations on four real-world continuous-time edge data sets. Xuhui Fan 0001, Bin Li 0015, Feng Zhou 0011, Scott A. Sisson |
NeurIPS | 1 |
| 2021 | Decoupling Sparsity and Smoothness in Dirichlet Belief Networks
Yaqiong Li, Xuhui Fan 0001, Ling Chen 0006, Bin Li 0015, Scott A. Sisson |
ECML/PKDD (2) | 2 |
| 2021 | Kernelized Sparse Bayesian Matrix FactorizationabstractExtracting low-rank and/or sparse structures using matrix factorization techniques has been extensively studied in the machine learning community. Kernelized matrix factorization (KMF) is a powerful tool to incorporate side information into the low-rank approximation model, which has been applied to solve the problems of data mining, recommender systems, image restoration, and machine vision. However, most existing KMF models rely on specifying the rows and columns of the data matrix through a Gaussian process prior and have to tune manually the rank. There are also computational issues of existing models based on regularization or the Markov chain Monte Carlo. In this article, we develop a hierarchical kernelized sparse Bayesian matrix factorization (KSBMF) model to integrate side information. The KSBMF automatically infers the parameters and latent variables including the reduced rank using the variational Bayesian inference. In addition, the model simultaneously achieves low-rankness through sparse Bayesian learning and columnwise sparsity through an enforced constraint on latent factor matrices. We further connect the KSBMF with the nonlocal image processing framework to develop two algorithms for image denoising and inpainting. Experimental results demonstrate that KSBMF outperforms the state-of-the-art approaches for these image-restoration tasks under various levels of corruption. Caoyuan Li, Hong-Bo Xie, Xuhui Fan 0001, Sabine Van Huffel, Kerrie L. Mengersen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Fragmentation Coagulation Based Mixed Membership Stochastic BlockmodelabstractThe Mixed-Membership Stochastic Blockmodel (MMSB) is proposed as one of the state-of-the-art Bayesian relational methods suitable for learning the complex hidden structure underlying the network data. However, the current formulation of MMSB suffers from the following two issues: (1), the prior information (e.g. entities' community structural information) can not be well embedded in the modelling; (2), community evolution can not be well described in the literature. Therefore, we propose a non-parametric fragmentation coagulation based Mixed Membership Stochastic Blockmodel (fcMMSB). Our model performs entity-based clustering to capture the community information for entities and linkage-based clustering to derive the group information for links simultaneously. Besides, the proposed model infers the network structure and models community evolution, manifested by appearances and disappearances of communities, using the discrete fragmentation coagulation process (DFCP). By integrating the community structure with the group compatibility matrix we derive a generalized version of MMSB. An efficient Gibbs sampling scheme with Polya Gamma (PG) approach is implemented for posterior inference. We validate our model on synthetic and real world data. Xuhui Fan 0001, Marcin Pietrasik, Marek Z. Reformat |
AAAI | 2 |
| 2020 | Online Binary Space Partitioning ForestsabstractThe Binary Space Partitioning-Tree (BSP-Tree) process was recently proposed as an efficient strategy for space partitioning tasks. Because it uses more than one dimension to partition the space, the BSP-Tree process is more efficient and flexible than conventional axis-aligned cut strategies. However, due to its batch learning setting, it is not well suited to large-scale classification and regression problems. In this paper, we develop an online BSP-Forest framework to address this limitation. With the arrival of new data, the resulting online algorithm can simultaneously expand the space coverage and refine the partition structure, with guaranteed universal consistency for classification problems. The effectiveness and competitive performance of the online BSP-Forest is verified via simulations. Xuhui Fan 0001, Bin Li 0015, Scott A. Sisson |
AISTATS | 1 |
| 2020 | Recurrent Dirichlet Belief Networks for interpretable Dynamic Relational Data ModellingabstractThe Dirichlet Belief Network~(DirBN) has been recently proposed as a promising approach in learning interpretable deep latent representations for objects. In this work, we leverage its interpretable modelling architecture and propose a deep dynamic probabilistic framework -- the Recurrent Dirichlet Belief Network~(Recurrent-DBN) -- to study interpretable hidden structures from dynamic relational data. The proposed Recurrent-DBN has the following merits: (1) it infers interpretable and organised hierarchical latent structures for objects within and across time steps; (2) it enables recurrent long-term temporal dependence modelling, which outperforms the one-order Markov descriptions in most of the dynamic probabilistic frameworks; (3) the computational cost scales to the number of positive links only. In addition, we develop a new inference strategy, which first upward-and-backward propagates latent counts and then downward-and-forward samples variables, to enable efficient Gibbs sampling for the Recurrent-DBN. We apply the Recurrent-DBN to dynamic relational data problems. The extensive experiment results on real-world data validate the advantages of the Recurrent-DBN over the state-of-the-art models in interpretable latent structure discovery and improved link prediction performance. Yaqiong Li, Xuhui Fan 0001, Ling Chen 0006, Bin Li 0015, Scott A. Sisson |
IJCAI | 2 |
| 2019 | Binary Space Partitioning ForestabstractThe Binary Space Partitioning (BSP)-Tree process is proposed to produce flexible 2-D partition structures which are originally used as a Bayesian nonparametric prior for relational modelling. It can hardly be applied to other learning tasks such as regression trees because extending the BSP-Tree process to a higher dimensional space is nontrivial. This paper is the first attempt to extend the BSP-Tree process to a d-dimensional ($d>2$) space. We propose to generate a cutting hyperplane, which is assumed to be parallel to $d-2$ dimensions, to cut each node in the d-dimensional BSP-tree. By designing a subtle strategy to sample two free dimensions from d dimensions, the extended BSP-Tree process can inherit the essential self-consistency property from the original version. Based on the extended BSP-Tree process, an ensemble model, which is named the BSP-Forest, is further developed for regression tasks. Thanks to the retained self-consistency property, we can thus significantly reduce the geometric calculations in the inference stage. Compared to its counterpart, the Mondrian Forest, the BSP-Forest can achieve similar performance with fewer cuts due to its flexibility. The BSP-Forest also outperforms other (Bayesian) regression forests on a number of real-world data sets. Xuhui Fan 0001, Bin Li 0015, Scott A. Sisson |
AISTATS | 1 |
| 2019 | Scalable Deep Generative Relational Model with High-Order Node DependenceabstractIn this work, we propose a probabilistic framework for relational data modelling and latent structure exploring. Given the possible feature information for the nodes in a network, our model builds up a deep architecture that can approximate to the possible nonlinear mappings between the nodes' feature information and latent representations. For each node, we incorporate all its neighborhoods' high-order structure information to generate latent representation, such that these latent representations are ``smooth'' in terms of the network. Since the latent representations are generated from Dirichlet distributions, we further develop a data augmentation trick to enable efficient Gibbs sampling for Ber-Poisson likelihood with Dirichlet random variables. Our model can be ready to apply to large sparse network as its computations cost scales to the number of positive links in the networks. The superior performance of our model is demonstrated through improved link prediction performance on a range of real-world datasets. Xuhui Fan 0001, Bin Li 0015, Caoyuan Li, Scott A. Sisson, Ling Chen 0006 |
NeurIPS | 1 |
| 2019 | Hawkes Process with Stochastic Triggering Kernel
Feng Zhou 0011, Yixuan Zhang 0006, Zhidong Li, Xuhui Fan 0001, Yang Wang 0002, Arcot Sowmya, Fang Chen 0001 |
PAKDD (1) | 4 |
| 2019 | Image Denoising Based on Nonlocal Bayesian Singular Value Thresholding and Stein's Unbiased Risk EstimatorabstractSingular value thresholding (SVT)- or nuclear norm minimization (NNM)-based nonlocal image denoising methods often rely on the precise estimation of the noise variance. However, most existing methods either assume that the noise variance is known or require an extra step to estimate it. Under the iterative regularization framework, the error in the noise variance estimate propagates and accumulates with each iteration, ultimately degrading the overall denoising performance. In addition, the essence of these methods is still least squares estimation, which can cause a very high mean-squared error (MSE) and is inadequate for handling missing data or outliers. In order to address these deficiencies, we present a hybrid denoising model based on variational Bayesian inference and Stein's unbiased risk estimator (SURE), which consists of two complementary steps. In the first step, the variational Bayesian SVT performs a low-rank approximation of the nonlocal image patch matrix to simultaneously remove the noise and estimate the noise variance. In the second step, we modify the conventional SURE full-rank SVT and its divergence formulas for rank-reduced eigen-triplets to remove the residual artifacts. The proposed hybrid BSSVT method achieves better performance in recovering the true image compared with state-of-the-art methods. Caoyuan Li, Hong-Bo Xie, Xuhui Fan 0001, Sabine Van Huffel, Scott A. Sisson, Kerrie L. Mengersen |
IEEE Trans. Image Process. | 3 |
| 2018 | The Binary Space Partitioning-Tree ProcessabstractThe Mondrian process represents an elegant and powerful approach for space partition modelling. However, as it restricts the partitions to be axis-aligned, its modelling flexibility is limited. In this work, we propose a self-consistent Binary Space Partitioning (BSP)-Tree process to generalize the Mondrian process. The BSP-Tree process is an almost surely right continuous Markov jump process that allows uniformly distributed oblique cuts in a two-dimensional convex polygon. The BSP-Tree process can also be extended using a non-uniform probability measure to generate direction differentiated cuts. The process is also self-consistent, maintaining distributional invariance under a restricted subdomain. We use Conditional-Sequential Monte Carlo for inference using the tree structure as the high-dimensional variable. The BSP-Tree process’s performance on synthetic data partitioning and relational modelling demonstrates clear inferential improvements over the standard Mondrian process and other related methods. Xuhui Fan 0001, Bin Li 0015, Scott A. Sisson |
AISTATS | 1 |
| 2018 | Rectangular Bounding ProcessabstractStochastic partition models divide a multi-dimensional space into a number of rectangular regions, such that the data within each region exhibit certain types of homogeneity. Due to the nature of their partition strategy, existing partition models may create many unnecessary divisions in sparse regions when trying to describe data in dense regions. To avoid this problem we introduce a new parsimonious partition model -- the Rectangular Bounding Process (RBP) -- to efficiently partition multi-dimensional spaces, by employing a bounding strategy to enclose data points within rectangular bounding boxes. Unlike existing approaches, the RBP possesses several attractive theoretical properties that make it a powerful nonparametric partition prior on a hypercube. In particular, the RBP is self-consistent and as such can be directly extended from a finite hypercube to infinite (unbounded) space. We apply the RBP to regression trees and relational models as a flexible partition prior. The experimental results validate the merit of the RBP {in rich yet parsimonious expressiveness} compared to the state-of-the-art methods. Xuhui Fan 0001, Bin Li 0015, Scott A. Sisson |
NeurIPS | 1 |
| 2018 | Corrosion Prediction on Sewer Networks with Sparse Monitoring Sites: A Case Study
Jianjia Zhang, Bin Li 0015, Xuhui Fan 0001, Yang Wang 0002, Fang Chen 0001 |
PAKDD (1) | 3 |
| 2018 | A Refined MISD Algorithm Based on Gaussian Process Regression
Feng Zhou 0011, Zhidong Li, Xuhui Fan 0001, Yang Wang 0002, Arcot Sowmya, Fang Chen 0001 |
PAKDD (2) | 3 |
| 2017 | Learning Nonparametric Relational Models by Conjugately Incorporating Node Information in a NetworkabstractRelational model learning is useful for numerous practical applications. Many algorithms have been proposed in recent years to tackle this important yet challenging problem. Existing algorithms utilize only binary directional link data to recover hidden network structures. However, there exists far richer and more meaningful information in other parts of a network which one can (and should) exploit. The attributes associated with each node, for instance, contain crucial information to help practitioners understand the underlying relationships in a network. For this reason, in this paper, we propose two models and their solutions, namely the node-information involved mixed-membership model and the node-information involved latent-feature model, in an effort to systematically incorporate additional node information. To effectively achieve this aim, node information is used to generate individual sticks of a stickbreaking process. In this way, not only can we avoid the need to prespecify the number of communities beforehand, the algorithm also encourages that nodes exhibiting similar information have a higher chance of assigning the same community membership. Substantial efforts have been made toward achieving the appropriateness and efficiency of these models, including the use of conjugate priors. We evaluate our framework and its inference algorithms using real-world data sets, which show the generality and effectiveness of our models in capturing implicit network structures. Xuhui Fan 0001, Longbing Cao, Yin Song |
IEEE Trans. Cybern. | 1 |
| 2016 | The Ostomachion ProcessabstractStochastic partition processes for exchangeable graphs produce axis-aligned blocks on a product space. In relational modeling, the resulting blocks uncover the underlying interactions between two sets of entities of the relational data. Although some flexible axis-aligned partition processes, such as the Mondrian process, have been able to capture complex interacting patterns in a hierarchical fashion, they are still in short of capturing dependence between dimensions. To overcome this limitation, we propose the Ostomachion process (OP), which relaxes the cutting direction by allowing for oblique cuts. The partitions generated by an OP are convex polygons that can capture inter-dimensional dependence. The OP also exhibits interesting properties: 1) Along the time line the cutting times can be characterized by a homogeneous Poisson process, and 2) on the partition space the areas of the resulting components comply with a Dirichlet distribution. We can thus control the expected number of cuts and the expected areas of components through hyper-parameters. We adapt the reversible-jump MCMC algorithm for inferring OP partition structures. The experimental results on relational modeling and decision tree classification have validated the merit of the OP. Xuhui Fan 0001, Bin Li 0015, Yi Wang 0041, Yang Wang 0002, Fang Chen 0001 |
AAAI | 1 |
| 2016 | Copula Mixed-Membership Stochastic Blockmodel
Xuhui Fan 0001, Longbing Cao |
IJCAI | 1 |
| 2016 | Bayesian Optimization of Partition Layouts for Mondrian Processes
Yi Wang 0041, Bin Li 0015, Xuhui Fan 0001, Yang Wang 0002, Fang Chen 0001 |
IJCAI | 3 |
| 2015 | A convergence theorem for graph shift-type algorithms
Xuhui Fan 0001, Longbing Cao |
Pattern Recognit. | 1 |
| 2015 | Dynamic Infinite Mixed-Membership Stochastic BlockmodelabstractDirectional and pairwise measurements are often used to model interactions in a social network setting. The mixed-membership stochastic blockmodel (MMSB) was a seminal work in this area, and its ability has been extended. However, models such as MMSB face particular challenges in modeling dynamic networks, for example, with the unknown number of communities. Accordingly, this paper proposes a dynamic infinite mixed-membership stochastic blockmodel, a generalized framework that extends the existing work to potentially infinite communities inside a network in dynamic settings (i.e., networks are observed over time). Additional model parameters are introduced to reflect the degree of persistence among one's memberships at consecutive time stamps. Under this framework, two specific models, namely mixture time variant and mixture time invariant models, are proposed to depict two different time correlation structures. Two effective posterior sampling strategies and their results are presented, respectively, using synthetic and real-world data. Xuhui Fan 0001, Longbing Cao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | A Theoretical Framework of the Graph Shift Algorithm
Xuhui Fan 0001, Longbing Cao |
AAAI | 1 |
| 2012 | Maximum margin clustering on evolutionary dataabstractEvolutionary data, such as topic changing blogs and evolving trading behaviors in capital market, is widely seen in business and social applications. The time factor and intrinsic change embedded in evolutionary data greatly challenge evolutionary clustering. To incorporate the time factor, existing methods mainly regard the evolutionary clustering problem as a linear combination of snapshot cost and temporal cost, and reflect the time factor through the temporal cost. It still faces accuracy and scalability challenge though promising results gotten. This paper proposes a novel evolutionary clustering approach, evolutionary maximum margin clustering (e-MMC), to cluster large-scale evolutionary data from the maximum margin perspective. e-MMC incorporates two frameworks: Data Integration from the data changing perspective and Model Integration corresponding to model adjustment to tackle the time factor and change, with an adaptive label allocation mechanism. Three e-MMC clustering algorithms are proposed based on the two frameworks. Extensive experiments are performed on synthetic data, UCI data and real-world blog data, which confirm that e-MMC outperforms the state-of-the-art clustering algorithms in terms of accuracy, computational cost and scalability. It shows that e-MMC is particularly suitable for clustering large-scale evolving data. Xuhui Fan 0001, Longbing Cao, Xia Cui 0002, Yew-Soon Ong |
CIKM | 1 |
| 2012 | Model the complex dependence structures of financial variables by using canonical vineabstractFinancial variables such as asset returns in the massive market contain various hierarchical and horizontal relationships forming complicated dependence structures. Modeling and mining of these structures is challenging due to their own high structural complexities as well as the stylized facts of the market data. This paper introduces a new canonical vine dependence model to identify the asymmetric and non-linear dependence structures of asset returns without any prior independence assumptions. To simplify the model while maintaining its merit, a partial correlation based method is proposed to optimize the canonical vine. Compared with the original canonical vine, the new model can still maintain the most important dependence but many unimportant nodes are removed to simplify the canonical vine structure. Our model is applied to construct and analyze dependence structures of European stocks as case studies. Its performance is evaluated by measuring portfolio of Value at Risk, a widely used risk management measure. In comparison to a very recent canonical vine model and the 'full' model, our experimental results demonstrate that our model has a much better quality of Value at Risk, providing insightful knowledge for investors to control and reduce the aggregation risk of the portfolio. Wei Wei 0039, Xuhui Fan 0001, Jinyan Li 0001, Longbing Cao |
CIKM | 2 |