Shiwei Lan

dblp:144/4462 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-9167-3715ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 91% Representation and self-supervised learning · 9%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 89% Bioinformatics and computational biology · 11%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
1.422025
Solving and Learning Partial Differential Equations with Variational Q-Exponential Processes · NeurIPS 2025
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models · J. Mach. Learn. Res. 2022
Machine learning › Representation and self-supervised learning
latent representation
0.912025
Bayesian Regularization of Latent Representation · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.912025
Bayesian Regularization of Latent Representation · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.912025
Bayesian Regularization of Latent Representation · ICLR 2025
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
0.912025
Solving and Learning Partial Differential Equations with Variational Q-Exponential Processes · NeurIPS 2025
Computational science and engineering › scientific machine learning
physics-informed machine learning
0.912025
Solving and Learning Partial Differential Equations with Variational Q-Exponential Processes · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.822022
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models · J. Mach. Learn. Res. 2022
Spherical Hamiltonian Monte Carlo for Constrained Target Distributions · ICML 2014
Machine learning › Probabilistic and Bayesian machine learning
bayesian regularization
0.712023
Bayesian Learning via Q-Exponential Process · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.612022
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › covariance function
nonstationary covariance
0.612022
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior contraction
0.612022
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
spatio-temporal gaussian process
0.612022
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
hamiltonian monte carlo
0.422014
Spherical Hamiltonian Monte Carlo for Constrained Target Distributions · ICML 2014
Wormhole Hamiltonian Monte Carlo · AAAI 2014
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.422014
Spherical Hamiltonian Monte Carlo for Constrained Target Distributions · ICML 2014
Wormhole Hamiltonian Monte Carlo · AAAI 2014
Bioinformatics and computational biology › phylogenetics
phylodynamics
0.212015
An efficient Bayesian inference framework for coalescent-based nonparametric phylodynamics · Bioinform. 2015
Image and video processing
image reconstruction
0.212023
Bayesian Learning via Q-Exponential Process · NeurIPS 2023
Image and video processing › image restoration
inverse problem
0.212023
Bayesian Learning via Q-Exponential Process · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › sampling
multimodal sampling
0.212014
Wormhole Hamiltonian Monte Carlo · AAAI 2014

Methods — techniques the papers use, named apart from their topics

gaussian process · 2.2uncertainty quantification · 1.7sparse variational inference · 1.7q-exponential process · 1.7elliptic contour distribution · 1.3besov process · 1.3variational inference · 0.9bayesian regularization · 0.9kronecker sum structure · 0.6hierarchical bayesian model · 0.6hamiltonian monte carlo · 0.2gaussian process prior · 0.2bayesian inference · 0.2
YearPublicationVenuePosition
2025 Bayesian Regularization of Latent Representation
abstract
The effectiveness of statistical and machine learning methods depends on how well data features are characterized. Developing informative and interpretable latent representations with controlled complexity is essential for visualizing data structure and for facilitating efficient model building through dimensionality reduction. Latent variable models, such as Gaussian Process Latent Variable Models (GP-LVM), have become popular for learning complex, nonlinear representations as alternatives to Principal Component Analysis (PCA). In this paper, we propose a novel class of latent variable models based on the recently introduced Q-exponential process (QEP), which generalizes GP-LVM with a tunable complexity parameter, $q>0$. Our approach, the \emph{Q-exponential Process Latent Variable Model (QEP-LVM)}, subsumes GP-LVM as a special case when $q=2$, offering greater flexibility in managing representation complexity while enhancing interpretability. To ensure scalability, we incorporate sparse variational inference within a Bayesian training framework. We establish connections between QEP-LVM and probabilistic PCA, demonstrating its superior performance through experiments on datasets such as the Swiss roll, oil flow, and handwritten digits.
Chukwudi Paul Obite, Zhi Chang, Keyan Wu, Shiwei Lan
ICLR4
2025 Solving and Learning Partial Differential Equations with Variational Q-Exponential Processes
abstract
Solving and learning partial differential equations (PDEs) lies at the core of physics-informed machine learning. Traditional numerical methods, such as finite difference and finite element approaches, are rooted in domain-specific techniques and often lack scalability. Recent advances have introduced neural networks and Gaussian processes (GPs) as flexible tools for automating PDE solving and incorporating physical knowledge into learning frameworks. While GPs offer tractable predictive distributions and a principled probabilistic foundation, they may be suboptimal in capturing complex behaviors such as sharp transitions or non-smooth dynamics. To address this limitation, we propose the use of the Q-exponential process (Q-EP), a recently developed generalization of GPs designed to better handle data with abrupt changes and to more accurately model derivative information. We advocate for Q-EP as a superior alternative to GPs in solving PDEs and associated inverse problems. Leveraging sparse variational inference, our method enables principled uncertainty quantification -- a capability not naturally afforded by neural network-based approaches. Through a series of experiments, including the Eikonal equation, Burgers’ equation, and an inverse Darcy flow problem, we demonstrate that the variational Q-EP method consistently yields more accurate solutions while providing meaningful uncertainty estimates.
Guangting Yu, Shiwei Lan
NeurIPS2
2025 A cancelable multi-biometric system based on the feature-level fusion of fingerprint and finger vein
Xueshuang Li, Chuanxian Xin, Shiwei Lan, Zhenhuan Hu
Multim. Tools Appl.5
2023 Bayesian Learning via Q-Exponential Process
abstract
Regularization is one of the most fundamental topics in optimization, statistics and machine learning. To get sparsity in estimating a parameter $u\in\mathbb{R}^d$, an $\ell_q$ penalty term, $\Vert u\Vert_q$, is usually added to the objective function. What is the probabilistic distribution corresponding to such $\ell_q$ penalty? What is the \emph{correct} stochastic process corresponding to $\Vert u\Vert_q$ when we model functions $u\in L^q$? This is important for statistically modeling high-dimensional objects such as images, with penalty to preserve certainty properties, e.g. edges in the image. In this work, we generalize the $q$-exponential distribution (with density proportional to) $\exp{(- \frac{1}{2}|u|^q)}$ to a stochastic process named \emph{$Q$-exponential (Q-EP) process} that corresponds to the $L_q$ regularization of functions. The key step is to specify consistent multivariate $q$-exponential distributions by choosing from a large family of elliptic contour distributions. The work is closely related to Besov process which is usually defined in terms of series. Q-EP can be regarded as a definition of Besov process with explicit probabilistic formulation, direct control on the correlation strength, and tractable prediction formula. From the Bayesian perspective, Q-EP provides a flexible prior on functions with sharper penalty ($q<2$) than the commonly used Gaussian process (GP, $q=2$). We compare GP, Besov and Q-EP in modeling functional data, reconstructing images and solving inverse problems and demonstrate the advantage of our proposed methodology.
Michael O'Connor, Shiwei Lan
NeurIPS3
2022 Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models
abstract
A large number of scientific studies involve high-dimensional spatiotemporal data with complicated relationships. In this paper, we focus on a type of space-time interaction named temporal evolution of spatial dependence (TESD), which is a zero time-lag spatiotemporal covariance. For this purpose, we propose a novel Bayesian nonparametric method based on non-stationary spatiotemporal Gaussian process (STGP). The classic STGP has a covariance kernel separable in space and time, failed to characterize TESD. More recent works on non-separable STGP treat location and time together as a joint variable, which is unnecessarily inefficient. We generalize STGP (gSTGP) to introduce time-dependence to the spatial kernel by varying its eigenvalues over time in the Mercer's representation. The resulting non-stationary non-separable covariance model bares a quasi Kronecker sum structure. Finally, a hierarchical Bayesian model for the joint covariance is proposed to allow for full flexibility in learning TESD. A simulation study and a longitudinal neuroimaging analysis on Alzheimer's patients demonstrate that the proposed methodology is (statistically) effective and (computationally) efficient in characterizing TESD. Theoretic properties of gSTGP including posterior contraction (for covariance) are also studied.
Shiwei Lan
J. Mach. Learn. Res.1
2020 Nonparametric Fisher Geometry with Application to Density Estimation
abstract
It is well known that the Fisher information induces a Riemannian geometry on parametric families of probability density functions. Following recent work, we consider the nonparametric generalization of the Fisher geometry. The resulting nonparametric Fisher geometry is shown to be equivalent to a familiar, albeit infinite-dimensional, geometric object—the sphere. By shifting focus away from density functions and toward square-root density functions, one may calculate theoretical quantities of interest with ease. More importantly, the sphere of square-root densities is much more computationally tractable. As discussed here, this insight leads to a novel Bayesian nonparametric density estimation model.
Andrew Holbrook, Shiwei Lan, Jeffrey Streets, Babak Shahbaba
UAI2
2015 An efficient Bayesian inference framework for coalescent-based nonparametric phylodynamics
abstract
MOTIVATION: The field of phylodynamics focuses on the problem of reconstructing population size dynamics over time using current genetic samples taken from the population of interest. This technique has been extensively used in many areas of biology but is particularly useful for studying the spread of quickly evolving infectious diseases agents, e.g. influenza virus. Phylodynamic inference uses a coalescent model that defines a probability density for the genealogy of randomly sampled individuals from the population. When we assume that such a genealogy is known, the coalescent model, equipped with a Gaussian process prior on population size trajectory, allows for nonparametric Bayesian estimation of population size dynamics. Although this approach is quite powerful, large datasets collected during infectious disease surveillance challenge the state-of-the-art of Bayesian phylodynamics and demand inferential methods with relatively low computational cost. RESULTS: To satisfy this demand, we provide a computationally efficient Bayesian inference framework based on Hamiltonian Monte Carlo for coalescent process models. Moreover, we show that by splitting the Hamiltonian function, we can further improve the efficiency of this approach. Using several simulated and real datasets, we show that our method provides accurate estimates of population size dynamics and is substantially faster than alternative methods based on elliptical slice sampler and Metropolis-adjusted Langevin algorithm. AVAILABILITY AND IMPLEMENTATION: The R code for all simulation studies and real data analysis conducted in this article are publicly available at http://www.ics.uci.edu/∼slan/lanzi/CODES.html and in the R package phylodyn available at https://github.com/mdkarcher/phylodyn. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shiwei Lan, Julia A. Palacios, Michael D. Karcher, Volodymyr M. Minin, Babak Shahbaba
Bioinform.1
2014 Wormhole Hamiltonian Monte Carlo
abstract
In machine learning and statistics, probabilistic inference involving multimodal distributions is quite difficult. This is especially true in high dimensional problems, where most existing algorithms cannot easily move from one mode to another. To address this issue, we propose a novel Bayesian inference approach based on Markov Chain Monte Carlo. Our method can effectively sample from multimodal distributions, especially when the dimension is high and the modes are isolated. To this end, it exploits and modifies the Riemannian geometric properties of the target distribution to create \emph{wormholes} connecting modes in order to facilitate moving between them. Further, our proposed method uses the regeneration technique in order to adapt the algorithm by identifying new modes and updating the network of wormholes without affecting the stationary distribution. To find new modes, as opposed to rediscovering those previously identified, we employ a novel mode searching algorithm that explores a \emph{residual energy} function obtained by subtracting an approximate Gaussian mixture density (based on previously discovered modes) from the target density function.
Shiwei Lan, Jeffrey Streets, Babak Shahbaba
AAAI1
2014 Spherical Hamiltonian Monte Carlo for Constrained Target Distributions
abstract
Statistical models with constrained probability distributions are abundant in machine learning. Some examples include regression models with norm constraints (e.g., Lasso), probit models, many copula models, and Latent Dirichlet Allocation (LDA) models. Bayesian inference involving probability distributions confined to constrained domains could be quite challenging for commonly used sampling algorithms. For such problems, we propose a novel Markov Chain Monte Carlo (MCMC) method that provides a general and computationally efficient framework for handling boundary conditions. Our method first maps the D-dimensional constrained domain of parameters to the unit ball \bf B_0^D(1), then augments it to the D-dimensional sphere \bf S^D such that the original boundary corresponds to the equator of \bf S^D. This way, our method handles the constraints implicitly by moving freely on sphere generating proposals that remain within boundaries when mapped back to the original space. To improve the computational efficiency of our algorithm, we divide the dynamics into several parts such that the resulting split dynamics has a partial analytical solution as a geodesic flow on the sphere. We apply our method to several examples including truncated Gaussian, Bayesian Lasso, Bayesian bridge regression, and a copula model for identifying synchrony among multiple neurons. Our results show that the proposed method can provide a natural and efficient framework for handling several types of constraints on target distributions.
Shiwei Lan, Bo Zhou 0003, Babak Shahbaba
ICML1
2014 A Semiparametric Bayesian Model for Detecting Synchrony Among Multiple Neurons
abstract
We propose a scalable semiparametric Bayesian model to capture dependencies among multiple neurons by detecting their cofiring (possibly with some lag time) patterns over time. After discretizing time so there is at most one spike at each interval, the resulting sequence of 1s (spike) and 0s (silence) for each neuron is modeled using the logistic function of a continuous latent variable with a gaussian process prior. For multiple neurons, the corresponding marginal distributions are coupled to their joint probability distribution using a parametric copula model. The advantages of our approach are as follows. The nonparametric component (i.e., the gaussian process model) provides a flexible framework for modeling the underlying firing rates, and the parametric component (i.e., the copula model) allows us to make inferences regarding both contemporaneous and lagged relationships among neurons. Using the copula model, we construct multivariate probabilistic models by separating the modeling of univariate marginal distributions from the modeling of a dependence structure among variables. Our method is easy to implement using a computationally efficient sampling algorithm that can be easily extended to high-dimensional problems. Using simulated data, we show that our approach could correctly capture temporal dependencies in firing rates and identify synchronous neurons. We also apply our model to spike train data obtained from prefrontal cortical areas.
Babak Shahbaba, Bo Zhou 0003, Shiwei Lan, Hernando C. Ombao, David Moorman, Sam Behseta
Neural Comput.3