VLDB 2026 Research / reviewers in the wild / expert
Tianyu Xie 0001
dblp:345/3987-1
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0000-8013-978XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 59% Probabilistic and Bayesian machine learning · 31% Kernel, tree and ensemble methods · 5% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.4 | 3 | 2025 | Provable Sample-Efficient Transfer Learning Conditional Diffusion Models via Representation Learning · NeurIPS 2025 Continuous Semi-Implicit Models · ICML 2025 Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
1.6 | 3 | 2024 | Kernel Semi-Implicit Variational Inference · ICML 2024 Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023 ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
diffusion model acceleration |
1.5 | 2 | 2025 | Continuous Semi-Implicit Models · ICML 2025 Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023 |
Bioinformatics and computational biology › phylogenetics
phylogenetic inference |
1.5 | 2 | 2025 | PhyloVAE: Unsupervised Learning of Phylogenetic Trees via Variational Autoencoders · ICLR 2025 ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
semi-implicit variational inference |
1.4 | 2 | 2024 | Kernel Semi-Implicit Variational Inference · ICML 2024 Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.9 | 1 | 2025 | Provable Sample-Efficient Transfer Learning Conditional Diffusion Models via Representation Learning · NeurIPS 2025 |
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow |
0.8 | 1 | 2024 | Reflected Flow Matching · ICML 2024 |
Machine learning › Generative modeling
flow matching |
0.8 | 1 | 2024 | Reflected Flow Matching · ICML 2024 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.8 | 1 | 2024 | Kernel Semi-Implicit Variational Inference · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › divergence measure
kernel stein discrepancy |
0.8 | 1 | 2024 | Kernel Semi-Implicit Variational Inference · ICML 2024 |
Machine learning › Generative modeling
normalizing flow |
0.8 | 1 | 2024 | Reflected Flow Matching · ICML 2024 |
Machine learning › Graph learning
graph neural network |
0.7 | 1 | 2023 | ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023 |
Bioinformatics and computational biology
phylogenetics |
0.7 | 1 | 2023 | ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2025 | PhyloVAE: Unsupervised Learning of Phylogenetic Trees via Variational Autoencoders · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
score-based generative model |
0.2 | 1 | 2024 | Reflected Flow Matching · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference |
0.2 | 1 | 2023 | Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.2 | 1 | 2023 | Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational bayesian inference |
0.2 | 1 | 2023 | ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
variational autoencoder · 1.7autoregressive tree topology generation · 1.7score matching · 1.4sample complexity analysis · 0.9multi-step distillation · 0.9continuous transition kernel · 0.9velocity field matching · 0.8reproducing kernel hilbert space · 0.8reflected flow matching · 0.8boundary constraint · 0.8graph neural network · 0.7autoregressive model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PhyloVAE: Unsupervised Learning of Phylogenetic Trees via Variational AutoencodersabstractLearning informative representations of phylogenetic tree structures is essential for analyzing evolutionary relationships. Classical distance-based methods have been widely used to project phylogenetic trees into Euclidean space, but they are often sensitive to the choice of distance metric and may lack sufficient resolution. In this paper, we introduce *phylogenetic variational autoencoders* (PhyloVAEs), an unsupervised learning framework designed for representation learning and generative modeling of tree topologies. Leveraging an efficient encoding mechanism inspired by autoregressive tree topology generation, we develop a deep latent-variable generative model that facilitates fast, parallelized topology generation. PhyloVAE combines this generative model with a collaborative inference model based on learnable topological features, allowing for high-resolution representations of phylogenetic tree samples. Extensive experiments demonstrate PhyloVAE's robust representation learning capabilities and fast generation of phylogenetic tree topologies. Tianyu Xie 0001, Harry Richman, Jiansi Gao, Frederick A. Matsen IV |
ICLR | 1 |
| 2025 | Continuous Semi-Implicit ModelsabstractSemi-implicit distributions have shown great promise in variational inference and generative modeling.
Hierarchical semi-implicit models, which stack multiple semi-implicit layers, enhance the expressiveness of semi-implicit distributions and can be used to accelerate diffusion models given pretrained score networks.
However, their sequential training often suffers from slow convergence.
In this paper, we introduce CoSIM, a continuous semi-implicit model that extends hierarchical semi-implicit models into a continuous framework.
By incorporating a continuous transition kernel, CoSIM enables efficient, simulation-free training.
Furthermore, we show that CoSIM achieves consistency with a carefully designed transition kernel, offering a novel approach for multistep distillation of generative models at the distributional level.
Extensive experiments on image generation demonstrate that CoSIM performs on par or better than existing diffusion model acceleration methods, achieving superior performance on FD-DINOv2. Longlin Yu, Jiajun Zha, Tianyu Xie 0001, Shueng-Han Gary Chan |
ICML | 4 |
| 2025 | Provable Sample-Efficient Transfer Learning Conditional Diffusion Models via Representation LearningabstractWhile conditional diffusion models have achieved remarkable success in various applications, they require abundant data to train from scratch, which is often infeasible in practice. To address this issue, transfer learning has emerged as an essential paradigm in small data regimes. Despite its empirical success, the theoretical underpinnings of transfer learning conditional diffusion models remain unexplored. In this paper, we take the first step towards understanding the sample efficiency of transfer learning conditional diffusion models through the lens of representation learning. Inspired by practical training procedures, we assume that there exists a low-dimensional representation of conditions shared across all tasks. Our analysis shows that with a well-learned representation from source tasks, the sample complexity of target tasks can be reduced substantially. Numerical experiments are also conducted to verify our results. Tianyu Xie 0001, Shiyue Zhang 0002 |
NeurIPS | 2 |
| 2025 | ARTreeFormer: A faster attention-based autoregressive model for phylogenetic inferenceabstractProbabilistic modeling over the combinatorially large space of tree topologies remains a central challenge in phylogenetic inference. Previous approaches often necessitate pre-sampled tree topologies, limiting their modeling capability to a subset of the entire tree space. A recent advancement is ARTree, a deep autoregressive model that offers unrestricted distributions for tree topologies. However, its reliance on repetitive tree traversals and inefficient local message passing for computing topological node representations may hamper the scalability to large datasets. This paper proposes ARTreeFormer, a novel approach that harnesses fixed-point iteration and attention mechanisms to accelerate ARTree. By introducing a fixed-point iteration algorithm for computing the topological node embeddings, ARTreeFormer allows for fast vectorized computation, especially on CUDA devices. This, together with an attention-based global message passing scheme, significantly improves the computation speed of ARTree while maintaining great approximation performance. We demonstrate the effectiveness and efficiency of our method on a benchmark of challenging real data phylogenetic inference problems. Tianyu Xie 0001, Yicong Mao |
PLoS Comput. Biol. | 1 |
| 2024 | Reflected Flow MatchingabstractContinuous normalizing flows (CNFs) learn an ordinary differential equation to transform prior samples into data. Flow matching (FM) has recently emerged as a simulation-free approach for training CNFs by regressing a velocity model towards the conditional velocity field. However, on constrained domains, the learned velocity model may lead to undesirable flows that result in highly unnatural samples, e.g., oversaturated images, due to both flow matching error and simulation error. To address this, we add a boundary constraint term to CNFs, which leads to reflected CNFs that keep trajectories within the constrained domains. We propose reflected flow matching (RFM) to train the velocity model in reflected CNFs by matching the conditional velocity fields in a simulation-free manner, similar to the vanilla FM. Moreover, the analytical form of conditional velocity fields in RFM avoids potentially biased approximations, making it superior to existing score-based generative models on constrained domains. We demonstrate that RFM achieves comparable or better results on standard image benchmarks and produces high-quality class-conditioned samples under high guidance weight. Tianyu Xie 0001, Yu Zhu 0004, Longlin Yu, Shiyue Zhang 0002 |
ICML | 1 |
| 2024 | Kernel Semi-Implicit Variational InferenceabstractSemi-implicit variational inference (SIVI) extends traditional variational families with semi-implicit distributions defined in a hierarchical manner. Due to the intractable densities of semi-implicit distributions, classical SIVI often resorts to surrogates of evidence lower bound (ELBO) that would introduce biases for training. A recent advancement in SIVI, named SIVI-SM, utilizes an alternative score matching objective made tractable via a minimax formulation, albeit requiring an additional lower-level optimization. In this paper, we propose kernel SIVI (KSIVI), a variant of SIVI-SM that eliminates the need for the lower-level optimization through kernel tricks. Specifically, we show that when optimizing over a reproducing kernel Hilbert space (RKHS), the lower-level problem has an explicit solution. This way, the upper-level objective becomes the kernel Stein discrepancy (KSD), which is readily computable for stochastic gradient descent due to the hierarchical structure of semi-implicit variational distributions. An upper bound for the variance of the Monte Carlo gradient estimators of the KSD objective is derived, which allows us to establish novel convergence guarantees of KSIVI. We demonstrate the effectiveness and efficiency of KSIVI on both synthetic distributions and a variety of real data Bayesian inference tasks. Longlin Yu, Tianyu Xie 0001, Shiyue Zhang 0002 |
ICML | 3 |
| 2023 | ARTree: A Deep Autoregressive Model for Phylogenetic InferenceabstractDesigning flexible probabilistic models over tree topologies is important for developing efficient phylogenetic inference methods. To do that, previous works often leverage the similarity of tree topologies via hand-engineered heuristic features which would require domain expertise and may suffer from limited approximation capability. In this paper, we propose a deep autoregressive model for phylogenetic inference based on graph neural networks (GNNs), called ARTree. By decomposing a tree topology into a sequence of leaf node addition operations and modeling the involved conditional distributions based on learnable topological features via GNNs, ARTree can provide a rich family of distributions over tree topologies that have simple sampling algorithms, without using heuristic features. We demonstrate the effectiveness and efficiency of our method on a benchmark of challenging real data tree topology density estimation and variational Bayesian phylogenetic inference problems. Tianyu Xie 0001 |
NeurIPS | 1 |
| 2023 | Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model AccelerationabstractSemi-implicit variational inference (SIVI) has been introduced to expand the analytical variational families by defining expressive semi-implicit distributions in a hierarchical manner. However, the single-layer architecture commonly used in current SIVI methods can be insufficient when the target posterior has complicated structures. In this paper, we propose hierarchical semi-implicit variational inference, called HSIVI, which generalizes SIVI to allow more expressive multi-layer construction of semi-implicit distributions. By introducing auxiliary distributions that interpolate between a simple base distribution and the target distribution, the conditional layers can be trained by progressively matching these auxiliary distributions one layer after another. Moreover, given pre-trained score networks, HSIVI can be used to accelerate the sampling process of diffusion models with the score matching objective. We show that HSIVI significantly enhances the expressiveness of SIVI on several Bayesian inference problems with complicated target distributions. When used for diffusion model acceleration, we show that HSIVI can produce high quality samples comparable to or better than the existing fast diffusion model based samplers with a small number of function evaluations on various datasets. Longlin Yu, Tianyu Xie 0001, Yu Zhu 0004 |
NeurIPS | 2 |