VLDB 2026 Research / reviewers in the wild / expert
Fangzheng Xie
dblp:336/5541
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-2436-9542ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 86% Information retrieval · 14% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
bayesian mixture model |
0.9 | 1 | 2025 | Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › relational model › random graph model
random dot product graph |
0.9 | 1 | 2025 | Statistical Inference of Random Graphs With a Surrogate Likelihood Function · J. Mach. Learn. Res. 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model › gaussian mixture model
sparse gaussian mixture |
0.9 | 1 | 2025 | Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025 |
Data mining
clustering |
0.9 | 1 | 2025 | Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025 |
Data mining › clustering
high-dimensional clustering |
0.9 | 1 | 2025 | Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025 |
Mathematical optimization › statistical estimation
maximum likelihood estimation |
0.9 | 1 | 2025 | Statistical Inference of Random Graphs With a Surrogate Likelihood Function · J. Mach. Learn. Res. 2025 |
Data mining › structured data mining › graph mining
community detection |
0.8 | 1 | 2024 | Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024 |
Data mining › structured data mining › graph mining
multi-layer graph |
0.8 | 1 | 2024 | Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024 |
Data mining
network analysis |
0.8 | 1 | 2024 | Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024 |
Data mining › dimensionality reduction
spectral embedding |
0.8 | 1 | 2024 | Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024 |
Information retrieval › evaluation
statistical significance testing |
0.8 | 1 | 2024 | Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and Inference · IEEE Trans. Inf. Theory 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior contraction |
0.3 | 1 | 2025 | Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions · J. Mach. Learn. Res. 2025 |
Methods — techniques the papers use, named apart from their topics
stochastic gradient descent · 1.7spike-and-slab prior · 1.7spectral estimation · 1.7minimax lower bound · 1.7matrix perturbation theory · 1.7bernstein-von mises theorem · 1.7spectral embedding · 0.8martingale argument · 0.8leave-one-out analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Statistical Inference of Random Graphs With a Surrogate Likelihood FunctionabstractSpectral estimators have been broadly applied to statistical network analysis, but they do not incorporate the likelihood information of the network sampling model. This paper proposes a novel surrogate likelihood function for statistical inference of a class of popular network models referred to as random dot product graphs. In contrast to the structurally complicated exact likelihood function, the surrogate likelihood function has a separable structure and is log-concave yet approximates the exact likelihood function well. From the frequentist perspective, we study the maximum surrogate likelihood estimator and establish the accompanying theory. We show its existence, uniqueness, large sample properties, and that it improves upon the baseline spectral estimator with a smaller sum of squared errors. Furthermore, we derive the second-order bias of the proposed estimator and gain insight into why it outperforms some of the existing estimators. A computationally convenient stochastic gradient descent algorithm is designed to find the maximum surrogate likelihood estimator in practice. From the Bayesian perspective, we establish the Bernstein--von Mises theorem of the posterior distribution with the surrogate likelihood function and show that the resulting credible sets have the correct frequentist coverage. The empirical performance of the proposed surrogate-likelihood-based methods is validated through the analyses of simulation examples and two real-world data sets. Dingbo Wu, Fangzheng Xie |
J. Mach. Learn. Res. | 2 |
| 2025 | Bayesian Sparse Gaussian Mixture Model for Clustering in High DimensionsabstractWe study the sparse high-dimensional Gaussian mixture model when the number of clusters is allowed to grow with the sample size. A minimax lower bound for parameter estimation is established, and we show that a constrained maximum likelihood estimator achieves the minimax lower bound. However, this optimization-based estimator is computationally intractable because the objective function is highly nonconvex and the feasible set involves discrete structures. To address the computational challenge, we propose a computationally tractable Bayesian approach to estimate high-dimensional Gaussian mixtures whose cluster centers exhibit sparsity using a continuous spike-and-slab prior. We further prove that the posterior contraction rate of the proposed Bayesian method is minimax optimal. The mis- clustering rate is obtained as a by-product using tools from matrix perturbation theory. The proposed Bayesian sparse Gaussian mixture model does not require pre-specifying the number of clusters, which can be adaptively estimated. The validity and usefulness of the proposed method is demonstrated through simulation studies and the analysis of a real-world single-cell RNA sequencing data set. Dapeng Yao, Fangzheng Xie, Yanxun Xu |
J. Mach. Learn. Res. | 2 |
| 2024 | Bias-Corrected Joint Spectral Embedding for Multilayer Networks With Invariant Subspace: Entrywise Eigenvector Perturbation and InferenceabstractIn this paper, we propose to estimate the invariant subspace across heterogeneous multiple networks using a novel bias-corrected joint spectral embedding algorithm. The proposed algorithm recursively calibrates the diagonal bias of the sum of squared network adjacency matrices by leveraging the closed-form bias formula and iteratively updates the subspace estimator using the most recent estimated bias. Correspondingly, we establish a complete recipe for the entrywise subspace estimation theory for the proposed algorithm, including a sharp entrywise subspace perturbation bound and the entrywise eigenvector central limit theorem. Leveraging these results, we settle two multiple network inference problems: the exact community detection in multilayer stochastic block models and the hypothesis testing of the equality of membership profiles in multilayer mixed membership models. Our proof relies on delicate leave-one-out and leave-two-out analyses that are specifically tailored to block-wise symmetric random matrices and a martingale argument that is of fundamental interest for the entrywise eigenvector central limit theorem. Fangzheng Xie |
IEEE Trans. Inf. Theory | 1 |