Tianyu Xie 0001

dblp:345/3987-1 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0000-8013-978XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 59% Probabilistic and Bayesian machine learning · 31% Kernel, tree and ensemble methods · 5%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.432025
Provable Sample-Efficient Transfer Learning Conditional Diffusion Models via Representation Learning · NeurIPS 2025
Continuous Semi-Implicit Models · ICML 2025
Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
1.632024
Kernel Semi-Implicit Variational Inference · ICML 2024
Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023
ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
diffusion model acceleration
1.522025
Continuous Semi-Implicit Models · ICML 2025
Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023
Bioinformatics and computational biology › phylogenetics
phylogenetic inference
1.522025
PhyloVAE: Unsupervised Learning of Phylogenetic Trees via Variational Autoencoders · ICLR 2025
ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
semi-implicit variational inference
1.422024
Kernel Semi-Implicit Variational Inference · ICML 2024
Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.912025
Provable Sample-Efficient Transfer Learning Conditional Diffusion Models via Representation Learning · NeurIPS 2025
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow
0.812024
Reflected Flow Matching · ICML 2024
Machine learning › Generative modeling
flow matching
0.812024
Reflected Flow Matching · ICML 2024
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.812024
Kernel Semi-Implicit Variational Inference · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › divergence measure
kernel stein discrepancy
0.812024
Kernel Semi-Implicit Variational Inference · ICML 2024
Machine learning › Generative modeling
normalizing flow
0.812024
Reflected Flow Matching · ICML 2024
Machine learning › Graph learning
graph neural network
0.712023
ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023
Bioinformatics and computational biology
phylogenetics
0.712023
ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023
Machine learning › Generative modeling
variational autoencoder
0.312025
PhyloVAE: Unsupervised Learning of Phylogenetic Trees via Variational Autoencoders · ICLR 2025
Machine learning › Generative modeling › diffusion model
score-based generative model
0.212024
Reflected Flow Matching · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.212023
Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.212023
Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational bayesian inference
0.212023
ARTree: A Deep Autoregressive Model for Phylogenetic Inference · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

variational autoencoder · 1.7autoregressive tree topology generation · 1.7score matching · 1.4sample complexity analysis · 0.9multi-step distillation · 0.9continuous transition kernel · 0.9velocity field matching · 0.8reproducing kernel hilbert space · 0.8reflected flow matching · 0.8boundary constraint · 0.8graph neural network · 0.7autoregressive model · 0.7
YearPublicationVenuePosition
2025 PhyloVAE: Unsupervised Learning of Phylogenetic Trees via Variational Autoencoders
abstract
Learning informative representations of phylogenetic tree structures is essential for analyzing evolutionary relationships. Classical distance-based methods have been widely used to project phylogenetic trees into Euclidean space, but they are often sensitive to the choice of distance metric and may lack sufficient resolution. In this paper, we introduce *phylogenetic variational autoencoders* (PhyloVAEs), an unsupervised learning framework designed for representation learning and generative modeling of tree topologies. Leveraging an efficient encoding mechanism inspired by autoregressive tree topology generation, we develop a deep latent-variable generative model that facilitates fast, parallelized topology generation. PhyloVAE combines this generative model with a collaborative inference model based on learnable topological features, allowing for high-resolution representations of phylogenetic tree samples. Extensive experiments demonstrate PhyloVAE's robust representation learning capabilities and fast generation of phylogenetic tree topologies.
Tianyu Xie 0001, Harry Richman, Jiansi Gao, Frederick A. Matsen IV
ICLR1
2025 Continuous Semi-Implicit Models
abstract
Semi-implicit distributions have shown great promise in variational inference and generative modeling. Hierarchical semi-implicit models, which stack multiple semi-implicit layers, enhance the expressiveness of semi-implicit distributions and can be used to accelerate diffusion models given pretrained score networks. However, their sequential training often suffers from slow convergence. In this paper, we introduce CoSIM, a continuous semi-implicit model that extends hierarchical semi-implicit models into a continuous framework. By incorporating a continuous transition kernel, CoSIM enables efficient, simulation-free training. Furthermore, we show that CoSIM achieves consistency with a carefully designed transition kernel, offering a novel approach for multistep distillation of generative models at the distributional level. Extensive experiments on image generation demonstrate that CoSIM performs on par or better than existing diffusion model acceleration methods, achieving superior performance on FD-DINOv2.
Longlin Yu, Jiajun Zha, Tianyu Xie 0001, Shueng-Han Gary Chan
ICML4
2025 Provable Sample-Efficient Transfer Learning Conditional Diffusion Models via Representation Learning
abstract
While conditional diffusion models have achieved remarkable success in various applications, they require abundant data to train from scratch, which is often infeasible in practice. To address this issue, transfer learning has emerged as an essential paradigm in small data regimes. Despite its empirical success, the theoretical underpinnings of transfer learning conditional diffusion models remain unexplored. In this paper, we take the first step towards understanding the sample efficiency of transfer learning conditional diffusion models through the lens of representation learning. Inspired by practical training procedures, we assume that there exists a low-dimensional representation of conditions shared across all tasks. Our analysis shows that with a well-learned representation from source tasks, the sample complexity of target tasks can be reduced substantially. Numerical experiments are also conducted to verify our results.
Tianyu Xie 0001, Shiyue Zhang 0002
NeurIPS2
2025 ARTreeFormer: A faster attention-based autoregressive model for phylogenetic inference
abstract
Probabilistic modeling over the combinatorially large space of tree topologies remains a central challenge in phylogenetic inference. Previous approaches often necessitate pre-sampled tree topologies, limiting their modeling capability to a subset of the entire tree space. A recent advancement is ARTree, a deep autoregressive model that offers unrestricted distributions for tree topologies. However, its reliance on repetitive tree traversals and inefficient local message passing for computing topological node representations may hamper the scalability to large datasets. This paper proposes ARTreeFormer, a novel approach that harnesses fixed-point iteration and attention mechanisms to accelerate ARTree. By introducing a fixed-point iteration algorithm for computing the topological node embeddings, ARTreeFormer allows for fast vectorized computation, especially on CUDA devices. This, together with an attention-based global message passing scheme, significantly improves the computation speed of ARTree while maintaining great approximation performance. We demonstrate the effectiveness and efficiency of our method on a benchmark of challenging real data phylogenetic inference problems.
Tianyu Xie 0001, Yicong Mao
PLoS Comput. Biol.1
2024 Reflected Flow Matching
abstract
Continuous normalizing flows (CNFs) learn an ordinary differential equation to transform prior samples into data. Flow matching (FM) has recently emerged as a simulation-free approach for training CNFs by regressing a velocity model towards the conditional velocity field. However, on constrained domains, the learned velocity model may lead to undesirable flows that result in highly unnatural samples, e.g., oversaturated images, due to both flow matching error and simulation error. To address this, we add a boundary constraint term to CNFs, which leads to reflected CNFs that keep trajectories within the constrained domains. We propose reflected flow matching (RFM) to train the velocity model in reflected CNFs by matching the conditional velocity fields in a simulation-free manner, similar to the vanilla FM. Moreover, the analytical form of conditional velocity fields in RFM avoids potentially biased approximations, making it superior to existing score-based generative models on constrained domains. We demonstrate that RFM achieves comparable or better results on standard image benchmarks and produces high-quality class-conditioned samples under high guidance weight.
Tianyu Xie 0001, Yu Zhu 0004, Longlin Yu, Shiyue Zhang 0002
ICML1
2024 Kernel Semi-Implicit Variational Inference
abstract
Semi-implicit variational inference (SIVI) extends traditional variational families with semi-implicit distributions defined in a hierarchical manner. Due to the intractable densities of semi-implicit distributions, classical SIVI often resorts to surrogates of evidence lower bound (ELBO) that would introduce biases for training. A recent advancement in SIVI, named SIVI-SM, utilizes an alternative score matching objective made tractable via a minimax formulation, albeit requiring an additional lower-level optimization. In this paper, we propose kernel SIVI (KSIVI), a variant of SIVI-SM that eliminates the need for the lower-level optimization through kernel tricks. Specifically, we show that when optimizing over a reproducing kernel Hilbert space (RKHS), the lower-level problem has an explicit solution. This way, the upper-level objective becomes the kernel Stein discrepancy (KSD), which is readily computable for stochastic gradient descent due to the hierarchical structure of semi-implicit variational distributions. An upper bound for the variance of the Monte Carlo gradient estimators of the KSD objective is derived, which allows us to establish novel convergence guarantees of KSIVI. We demonstrate the effectiveness and efficiency of KSIVI on both synthetic distributions and a variety of real data Bayesian inference tasks.
Longlin Yu, Tianyu Xie 0001, Shiyue Zhang 0002
ICML3
2023 ARTree: A Deep Autoregressive Model for Phylogenetic Inference
abstract
Designing flexible probabilistic models over tree topologies is important for developing efficient phylogenetic inference methods. To do that, previous works often leverage the similarity of tree topologies via hand-engineered heuristic features which would require domain expertise and may suffer from limited approximation capability. In this paper, we propose a deep autoregressive model for phylogenetic inference based on graph neural networks (GNNs), called ARTree. By decomposing a tree topology into a sequence of leaf node addition operations and modeling the involved conditional distributions based on learnable topological features via GNNs, ARTree can provide a rich family of distributions over tree topologies that have simple sampling algorithms, without using heuristic features. We demonstrate the effectiveness and efficiency of our method on a benchmark of challenging real data tree topology density estimation and variational Bayesian phylogenetic inference problems.
Tianyu Xie 0001
NeurIPS1
2023 Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration
abstract
Semi-implicit variational inference (SIVI) has been introduced to expand the analytical variational families by defining expressive semi-implicit distributions in a hierarchical manner. However, the single-layer architecture commonly used in current SIVI methods can be insufficient when the target posterior has complicated structures. In this paper, we propose hierarchical semi-implicit variational inference, called HSIVI, which generalizes SIVI to allow more expressive multi-layer construction of semi-implicit distributions. By introducing auxiliary distributions that interpolate between a simple base distribution and the target distribution, the conditional layers can be trained by progressively matching these auxiliary distributions one layer after another. Moreover, given pre-trained score networks, HSIVI can be used to accelerate the sampling process of diffusion models with the score matching objective. We show that HSIVI significantly enhances the expressiveness of SIVI on several Bayesian inference problems with complicated target distributions. When used for diffusion model acceleration, we show that HSIVI can produce high quality samples comparable to or better than the existing fast diffusion model based samplers with a small number of function evaluations on various datasets.
Longlin Yu, Tianyu Xie 0001, Yu Zhu 0004
NeurIPS2