Fangzheng Xie

dblp:336/5541 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-2436-9542ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 86% Information retrieval · 14%
Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
bayesian mixture model
0.912025
Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › relational model › random graph model
random dot product graph
0.912025
Statistical Inference of Random Graphs With a Surrogate Likelihood Function · J. Mach. Learn. Res. 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model › gaussian mixture model
sparse gaussian mixture
0.912025
Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025
Data mining
clustering
0.912025
Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025
Data mining › clustering
high-dimensional clustering
0.912025
Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025
Mathematical optimization › statistical estimation
maximum likelihood estimation
0.912025
Statistical Inference of Random Graphs With a Surrogate Likelihood Function · J. Mach. Learn. Res. 2025
Data mining › structured data mining › graph mining
community detection
0.812024
Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024
Data mining › structured data mining › graph mining
multi-layer graph
0.812024
Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024
Data mining
network analysis
0.812024
Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024
Data mining › dimensionality reduction
spectral embedding
0.812024
Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024
Information retrieval › evaluation
statistical significance testing
0.812024
Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior contraction
0.312025
Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 1.7spike-and-slab prior · 1.7spectral estimation · 1.7minimax lower bound · 1.7matrix perturbation theory · 1.7bernstein-von mises theorem · 1.7spectral embedding · 0.8martingale argument · 0.8leave-one-out analysis · 0.8
YearPublicationVenuePosition
2025 Statistical Inference of Random Graphs With a Surrogate Likelihood Function
abstract
Spectral estimators have been broadly applied to statistical network analysis, but they do not incorporate the likelihood information of the network sampling model. This paper proposes a novel surrogate likelihood function for statistical inference of a class of popular network models referred to as random dot product graphs. In contrast to the structurally complicated exact likelihood function, the surrogate likelihood function has a separable structure and is log-concave yet approximates the exact likelihood function well. From the frequentist perspective, we study the maximum surrogate likelihood estimator and establish the accompanying theory. We show its existence, uniqueness, large sample properties, and that it improves upon the baseline spectral estimator with a smaller sum of squared errors. Furthermore, we derive the second-order bias of the proposed estimator and gain insight into why it outperforms some of the existing estimators. A computationally convenient stochastic gradient descent algorithm is designed to find the maximum surrogate likelihood estimator in practice. From the Bayesian perspective, we establish the Bernstein--von Mises theorem of the posterior distribution with the surrogate likelihood function and show that the resulting credible sets have the correct frequentist coverage. The empirical performance of the proposed surrogate-likelihood-based methods is validated through the analyses of simulation examples and two real-world data sets.
Dingbo Wu, Fangzheng Xie
J. Mach. Learn. Res.2
2025 Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions
abstract
We study the sparse high-dimensional Gaussian mixture model when the number of clusters is allowed to grow with the sample size. A minimax lower bound for parameter estimation is established, and we show that a constrained maximum likelihood estimator achieves the minimax lower bound. However, this optimization-based estimator is computationally intractable because the objective function is highly nonconvex and the feasible set involves discrete structures. To address the computational challenge, we propose a computationally tractable Bayesian approach to estimate high-dimensional Gaussian mixtures whose cluster centers exhibit sparsity using a continuous spike-and-slab prior. We further prove that the posterior contraction rate of the proposed Bayesian method is minimax optimal. The mis- clustering rate is obtained as a by-product using tools from matrix perturbation theory. The proposed Bayesian sparse Gaussian mixture model does not require pre-specifying the number of clusters, which can be adaptively estimated. The validity and usefulness of the proposed method is demonstrated through simulation studies and the analysis of a real-world single-cell RNA sequencing data set.
Dapeng Yao, Fangzheng Xie, Yanxun Xu
J. Mach. Learn. Res.2
2024 Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference
abstract
In this paper, we propose to estimate the invariant subspace across heterogeneous multiple networks using a novel bias-corrected joint spectral embedding algorithm. The proposed algorithm recursively calibrates the diagonal bias of the sum of squared network adjacency matrices by leveraging the closed-form bias formula and iteratively updates the subspace estimator using the most recent estimated bias. Correspondingly, we establish a complete recipe for the entrywise subspace estimation theory for the proposed algorithm, including a sharp entrywise subspace perturbation bound and the entrywise eigenvector central limit theorem. Leveraging these results, we settle two multiple network inference problems: the exact community detection in multilayer stochastic block models and the hypothesis testing of the equality of membership profiles in multilayer mixed membership models. Our proof relies on delicate leave-one-out and leave-two-out analyses that are specifically tailored to block-wise symmetric random matrices and a martingale argument that is of fundamental interest for the entrywise eigenvector central limit theorem.
Fangzheng Xie
IEEE Trans. Inf. Theory1