Guang Lin 0001

dblp:12/7586-1 · DBLP profile ↗
← Back
31ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0002-0976-1987ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring Non-Convex Discrete Energy Landscapes: An Efficient Langevin-Like Sampler with Replica Exchange
abstract
Gradient-based Discrete Samplers (GDSs) are effective for sampling discrete energy landscapes. However, they often stagnate in complex, non-convex settings. To improve exploration, we introduce the Discrete Replica EXchangE Langevin (DREXEL) sampler and its variant with Adjusted Metropolis (DREAM). These samplers use two GDSs at different temperatures and step sizes: one focuses on local exploitation, while the other explores broader energy landscapes. When energy differences are significant, sample swaps occur, governed by a mechanism tailored for discrete sampling to ensure detailed balance. Theoretically, we prove that the proposed samplers satisfy detailed balance and converge to the target distribution under mild conditions. Experiments across 2d synthetic simulations, sampling from Ising models and restricted Boltzmann machines, and training deep energy-based models further confirm their efficiency in exploring non-convex discrete energy landscapes.
Haoyang Zheng, Hengrong Du, Ruqi Zhang, Guang Lin 0001
AAAI4
2025 LLM Safety Alignment is Divergence Estimation in Disguise
abstract
We present a theoretical framework showing that popular LLM alignment methods—including RLHF and its variants—can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less-preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance–refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.
Rajdeep Haldar, Guang Lin 0001, Yue Xing 0002, Qifan Song
NeurIPS3
2025 High-quality three-dimensional cartoon avatar reconstruction with Gaussian splatting
MinHyuk Jang, Jong Wook Kim, Youngdong Jang, Donghyun Kim 0006, Wonseok Roh, Inyong Hwang, Guang Lin 0001, Sangpil Kim
Eng. Appl. Artif. Intell.7
2025 A self-adaptive energy-based learning rate for stochastic gradient descent via Vector Auxiliary Variable method
Jiahao Zhang 0002, Christian Moya, Guang Lin 0001
Eng. Appl. Artif. Intell.3
2025 Conformalized prediction of post-fault voltage trajectories using pre-trained and finetuned attention-driven neural operators
Amirhossein Mollaali, Gabriel Zufferey, Gonzalo E. Constante-Flores, Christian Moya, Can Li 0010, Meng Yue 0001, Guang Lin 0001
Neural Networks7
2024 Federated X-armed Bandit
abstract
This work establishes the first framework of federated X-armed bandit, where different clients face heterogeneous local objective functions defined on the same domain and are required to collaboratively figure out the global optimum. We propose the first federated algorithm for such problems, named Fed-PNE. By utilizing the topological structure of the global objective inside the hierarchical partitioning and the weak smoothness property, our algorithm achieves sublinear cumulative regret with respect to both the number of clients and the evaluation budget. Meanwhile, it only requires logarithmic communications between the central server and clients, protecting the client privacy. Experimental results on synthetic functions and real datasets validate the advantages of Fed-PNE over various centralized and federated baseline algorithms.
Qifan Song, Jean Honorio, Guang Lin 0001
AAAI4
2024 Accelerating Approximate Thompson Sampling with Underdamped Langevin Monte Carlo
abstract
Approximate Thompson sampling with Langevin Monte Carlo broadens its reach from Gaussian posterior sampling to encompass more general smooth posteriors. However, it still encounters scalability issues in high-dimensional problems when demanding high accuracy. To address this, we propose an approximate Thompson sampling strategy, utilizing underdamped Langevin Monte Carlo, where the latter is the go-to workhorse for simulations of high-dimensional posteriors. Based on the standard smoothness and log-concavity conditions, we study the accelerated posterior concentration and sampling using a specific potential function. This design improves the sample complexity for realizing logarithmic regrets from $\mathcal{\tilde O}(d)$ to $\mathcal{\tilde O}(\sqrt{d})$. The scalability and robustness of our algorithm are also empirically validated through synthetic experiments in high-dimensional bandit problems.
Haoyang Zheng, Wei Deng 0002, Christian Moya, Guang Lin 0001
AISTATS4
2024 Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics
abstract
Replica exchange stochastic gradient Langevin dynamics (reSGLD) is an effective sampler for non-convex learning in large-scale datasets. However, the simulation may encounter stagnation issues when the high-temperature chain delves too deeply into the distribution tails. To tackle this issue, we propose reflected reSGLD (r2SGLD): an algorithm tailored for constrained non-convex exploration by utilizing reflection steps within a bounded domain. Theoretically, we observe that reducing the diameter of the domain enhances mixing rates, exhibiting a *quadratic* behavior. Empirically, we test its performance through extensive experiments, including identifying dynamical systems with physical constraints, simulations of constrained multi-modal distributions, and image classification tasks. The theoretical and empirical findings highlight the crucial role of constrained exploration in improving the simulation efficiency.
Haoyang Zheng, Hengrong Du, Qi Feng 0005, Wei Deng 0002, Guang Lin 0001
ICML5
2024 On Convergence of Federated Averaging Langevin Dynamics
abstract
We propose a federated averaging Langevin algorithm (FA-LD) for uncertainty quantification and mean predictions with distributed clients. In particular, we generalize beyond normal posterior distributions and consider a general class of models. We develop theoretical guarantees for FA-LD for strongly log-concave distributions with non-i.i.d data and study how the injected noise and the stochastic-gradient noise, the heterogeneity of data, and the varying learning rates affect the convergence. Such an analysis sheds light on the optimal choice of local updates to minimize the communication cost. Important to our approach is that the communication efficiency does not deteriorate with the injected noise in the Langevin algorithms. In addition, we examine in our FA-LD algorithm both independent and correlated noise used over different clients. We observe that there is a trade-off between the pairs among communication, accuracy, and data privacy. As local devices may become inactive in federated networks, we also show convergence results based on different averaging schemes where only partial device updates are available. In such a case, we discover an additional bias that does not decay to zero.
Wei Deng 0002, Qian Zhang 0067, Yi-An Ma, Zhao Song 0002, Guang Lin 0001
UAI5
2023 Non-reversible Parallel Tempering for Deep Posterior Approximation
abstract
Parallel tempering (PT), also known as replica exchange, is the go-to workhorse for simulations of multi-modal distributions. The key to the success of PT is to adopt efficient swap schemes. The popular deterministic even-odd (DEO) scheme exploits the non-reversibility property and has successfully reduced the communication cost from quadratic to linear given the sufficiently many chains. However, such an innovation largely disappears in big data due to the limited chains and few bias-corrected swaps. To handle this issue, we generalize the DEO scheme to promote non-reversibility and propose a few solutions to tackle the underlying bias caused by the geometric stopping time. Notably, in big data scenarios, we obtain a nearly linear communication cost based on the optimal window size. In addition, we also adopt stochastic gradient descent (SGD) with large and constant learning rates as exploration kernels. Such a user-friendly nature enables us to conduct approximation tasks for complex posteriors without much tuning costs.
Wei Deng 0002, Qian Zhang 0067, Qi Feng 0005, Faming Liang, Guang Lin 0001
AAAI5
2023 Learning the dynamical response of nonlinear non-autonomous dynamical systems with deep operator neural networks
Guang Lin 0001, Christian Moya, Zecheng Zhang
Eng. Appl. Artif. Intell.1
2023 DeepONet-grid-UQ: A trustworthy deep operator framework for predicting the power grid's post-fault trajectories
Christian Moya, Shiqi Zhang 0006, Guang Lin 0001, Meng Yue 0001
Neurocomputing3
2023 DAE-PINN: a physics-informed neural network model for simulating differential algebraic equations with application to power networks
Christian Moya, Guang Lin 0001
Neural Comput. Appl.2
2022 Glassoformer: A Query-Sparse Transformer for Post-Fault Power Grid Voltage Prediction
abstract
We propose GLassoformer, a novel and efficient transformer architecture leveraging group Lasso regularization to reduce the number of queries of the standard self-attention mechanism. Due to the sparsified queries, GLassoformer is more computationally efficient than the standard transformers. On the power grid post-fault voltage prediction task, GLasso-former shows remarkably better prediction than many existing benchmark algorithms in terms of accuracy and stability.
Yunling Zheng, Carson Hu, Guang Lin 0001, Meng Yue 0001, Bao Wang 0001, Jack Xin
ICASSP3
2022 Interacting Contour Stochastic Gradient Langevin Dynamics
Wei Deng 0002, Siqi Liang 0005, Botao Hao, Guang Lin 0001, Faming Liang
ICLR4
2022 Learning-PDE-Based Approximate Optimal Control for an MHD System With Uncertainty Quantification
abstract
Handling uncertainty is one of the most important challenges in real physical systems. In this article, we study an approximate optimal magnetic control strategy for a one-dimensional (1-D) magnetohydrodynamic (MHD) system with uncertainty quantification within the learning framework of the underlying MHD model. First, the MHD flow system is modeled by coupled partial differential equations (PDEs) wherein the Reynolds number is not deterministic but random. Then, the optimal magnetic control problem is formulated and reduced to a parameter selection problem with stochastic PDE constraints by means of the control parameterization method. Significantly different from the conventional sensitivity analysis and adjoint methods, a polynomial chaos expansion (PCE) based on a multifidelity model is developed to construct the underlying PDEs model by polynomial functions. Thus, the relationship between the objective function and the control sequence is derived explicitly, which, in turn, transforms the optimal parameter selection problem with stochastic PDE constraints into a typical algebraic optimization problem that can be easily solved by the existing nonlinear programming algorithm. Numerical simulations are illustrated to demonstrate the high performance of our proposed method.
Tehuan Chen, Guang Lin 0001, Chao Xu 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Accelerating Convergence of Replica Exchange Stochastic Gradient MCMC via Variance Reduction
Wei Deng 0002, Qi Feng 0005, Georgios Karagiannis, Guang Lin 0001, Faming Liang
ICLR4
2021 DeepLight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving
abstract
Click-through rate (CTR) prediction is a crucial task in recommender systems and online advertising. The embedding-based neural networks have been proposed to learn both explicit feature interactions through a shallow component and deep feature interactions by a deep neural network (DNN) component. These sophisticated models, however, slow down the prediction inference by at least hundreds of times. To address the issue of significantly increased serving latency and high memory usage for real-time serving in production, this paper presents DeepLight: a framework to accelerate the CTR predictions in three aspects: 1) accelerate the model inference via explicitly searching informative feature interactions in the shallow component; 2) prune redundant parameters at the inter-layer level in the DNN component; 3) prune the dense embedding vectors to make them sparse in the embedding matrix. By combining the above efforts, the proposed approach accelerates the model inference by 46X on Criteo dataset and 27X on Avazu dataset without any loss on the prediction accuracy. This paves the way for successfully deploying complicated embedding-based neural networks in real-world serving systems.
Wei Deng 0002, Junwei Pan, Tian Zhou 0006, Deguang Kong, Aaron Flores 0001, Guang Lin 0001
WSDM6
2021 Binary classification of floor vibrations for human activity detection based on dynamic mode decomposition
Guang Lin 0001, Qinfang Qian, Chao Xu 0001
Neurocomputing2
2021 An integrated framework for building trustworthy data-driven epidemiological models: Application to the COVID-19 outbreak in New York City
abstract
Epidemiological models can provide the dynamic evolution of a pandemic but they are based on many assumptions and parameters that have to be adjusted over the time the pandemic lasts. However, often the available data are not sufficient to identify the model parameters and hence infer the unobserved dynamics. Here, we develop a general framework for building a trustworthy data-driven epidemiological model, consisting of a workflow that integrates data acquisition and event timeline, model development, identifiability analysis, sensitivity analysis, model calibration, model robustness analysis, and projection with uncertainties in different scenarios. In particular, we apply this framework to propose a modified susceptible-exposed-infectious-recovered (SEIR) model, including new compartments and model vaccination in order to project the transmission dynamics of COVID-19 in New York City (NYC). We find that we can uniquely estimate the model parameters and accurately project the daily new infection cases, hospitalizations, and deaths, in agreement with the available data from NYC's government's website. In addition, we employ the calibrated data-driven model to study the effects of vaccination and timing of reopening indoor dining in NYC.
Joan Ponce, Zhen Zhang 0029, Guang Lin 0001, George Em Karniadakis
PLoS Comput. Biol.4
2020 Non-convex Learning via Replica Exchange Stochastic Gradient MCMC
abstract
Replica exchange Monte Carlo (reMC), also known as parallel tempering, is an important technique for accelerating the convergence of the conventional Markov Chain Monte Carlo (MCMC) algorithms. However, such a method requires the evaluation of the energy function based on the full dataset and is not scalable to big data. The naïve implementation of reMC in mini-batch settings introduces large biases, which cannot be directly extended to the stochastic gradient MCMC (SGMCMC), the standard sampling method for simulating from deep neural networks (DNNs). In this paper, we propose an adaptive replica exchange SGMCMC (reSGMCMC) to automatically correct the bias and study the corresponding properties. The analysis implies an acceleration-accuracy trade-off in the numerical discretization of a Markov jump process in a stochastic environment. Empirically, we test the algorithm through extensive experiments on various setups and obtain the state-of-the-art results on CIFAR10, CIFAR100, and SVHN in both supervised learning and semi-supervised learning tasks.
Wei Deng 0002, Qi Feng 0005, Liyao Gao, Faming Liang, Guang Lin 0001
ICML5
2020 A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal Distributions
abstract
We propose an adaptively weighted stochastic gradient Langevin dynamics algorithm (SGLD), so-called contour stochastic gradient Langevin dynamics (CSGLD), for Bayesian learning in big data statistics. The proposed algorithm is essentially a scalable dynamic importance sampler, which automatically flattens the target distribution such that the simulation for a multi-modal distribution can be greatly facilitated. Theoretically, we prove a stability condition and establish the asymptotic convergence of the self-adapting parameter to a unique fixed-point, regardless of the non-convexity of the original energy function; we also present an error analysis for the weighted averaging estimators. Empirically, the CSGLD algorithm is tested on multiple benchmark datasets including CIFAR10 and CIFAR100. The numerical results indicate its superiority over the existing state-of-the-art algorithms in training deep neural networks.
Wei Deng 0002, Guang Lin 0001, Faming Liang
NeurIPS2
2020 Robust weighted SVD-type latent factor models for rating prediction
Yiqi Gu, Mengjiao Peng, Guang Lin 0001
Expert Syst. Appl.4
2020 Latent transformations neural network for object view synthesis
Sangpil Kim, Nick Winovich, Hyung-Gun Chi, Guang Lin 0001, Karthik Ramani
Vis. Comput.4
2019 An Adaptive Empirical Bayesian Method for Sparse Deep Learning
abstract
We propose a novel adaptive empirical Bayesian (AEB) method for sparse deep learning, where the sparsity is ensured via a class of self-adaptive spike-and-slab priors. The proposed method works by alternatively sampling from an adaptive hierarchical posterior distribution using stochastic gradient Markov Chain Monte Carlo (MCMC) and smoothly optimizing the hyperparameters using stochastic approximation (SA). The convergence of the proposed method to the asymptotically correct distribution is established under mild conditions. Empirical applications of the proposed method lead to the state-of-the-art performance on MNIST and Fashion MNIST with shallow convolutional neural networks (CNN) and the state-of-the-art compression performance on CIFAR10 with Residual Networks. The proposed method also improves resistance to adversarial attacks.
Wei Deng 0002, Faming Liang, Guang Lin 0001
NeurIPS4
2018 Local Feature Sufficiency Exploration for Predicting Security-Constrained Generation Dispatch in Multi-area Power Systems
abstract
Deriving generation dispatch is essential for efficient and secure operation of electric power systems. This is usually achieved by solving a security-constrained optimal power flow (SCOPF) problem, which is by nature non-convex, usually nonlinear and thus computationally intensive. The state-of-the-art optimization approaches are not able to solve this problem for large-scale power systems within the power system operation time window (usually 5 minutes). In this work, we developed supervised learning approaches to determine security-constrained generation dispatch within a much shorter time window. More importantly, the physical constraint of only accessing to local measurements and other information in most utilities' real-time operation can not be ignored for the predictive models. The feasibility and accuracy of utilizing only local features (measurements and grid information in one area) to predict optimal local generation dispatch (dispatch of all generators in the corresponding area) in multi-area power systems has been explored. The results showed optimal local generation dispatch can be predicted with local features with high accuracy, which is comparable to the results obtained with global features.
Yixuan Sun, Xiaoyuan Fan, Qiuhua Huang, Xinya Li 0002, Renke Huang, Tianzhixi Yin, Guang Lin 0001
ICMLA7
2017 Visualization of Time-Varying Weather Ensembles across Multiple Resolutions
abstract
Uncertainty quantification in climate ensembles is an important topic for the domain scientists, especially for decision making in the real-world scenarios. With powerful computers, simulations now produce time-varying and multi-resolution ensemble data sets. It is of extreme importance to understand the model sensitivity given the input parameters such that more computation power can be allocated to the parameters with higher influence on the output. Also, when ensemble data is produced at different resolutions, understanding the accuracy of different resolutions helps the total time required to produce a desired quality solution with improved storage and computation cost. In this work, we propose to tackle these non-trivial problems on the Weather Research and Forecasting (WRF) model output. We employ a moment independent sensitivity measure to quantify and analyze parameter sensitivity across spatial regions and time domain. A comparison of clustering structures across three resolutions enables the users to investigate the sensitivity variation over the spatial regions of the five input parameters. The temporal trend in the sensitivity values is explored via an MDS view linked with a line chart for interactive brushing. The spatial and temporal views are connected to provide a full exploration system for complete spatio-temporal sensitivity analysis. To analyze the accuracy across varying resolutions, we formulate a Bayesian approach to identify which regions are better predicted at which resolutions compared to the observed precipitation. This information is aggregated over the time domain and finally encoded in an output image through a custom color map that guides the domain experts towards an adaptive grid implementation given a cost model. Users can select and further analyze the spatial and temporal error patterns for multi-resolution accuracy analysis via brushing and linking on the produced image. In this work, we collaborate with a domain expert whose feedback shows the effectiveness of our proposed exploration work-flow.
Ayan Biswas 0001, Guang Lin 0001, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2017 Multi-Resolution Climate Ensemble Parameter Analysis with Nested Parallel Coordinates Plots
abstract
Due to the uncertain nature of weather prediction, climate simulations are usually performed multiple times with different spatial resolutions. The outputs of simulations are multi-resolution spatial temporal ensembles. Each simulation run uses a unique set of values for multiple convective parameters. Distinct parameter settings from different simulation runs in different resolutions constitute a multi-resolution high-dimensional parameter space. Understanding the correlation between the different convective parameters, and establishing a connection between the parameter settings and the ensemble outputs are crucial to domain scientists. The multi-resolution high-dimensional parameter space, however, presents a unique challenge to the existing correlation visualization techniques. We present Nested Parallel Coordinates Plot (NPCP), a new type of parallel coordinates plots that enables visualization of intra-resolution and inter-resolution parameter correlations. With flexible user control, NPCP integrates superimposition, juxtaposition and explicit encodings in a single view for comparative data visualization and analysis. We develop an integrated visual analytics system to help domain scientists understand the connection between multi-resolution convective parameters and the large spatial temporal ensembles. Our system presents intricate climate ensembles with a comprehensive overview and on-demand geographic details. We demonstrate NPCP, along with the climate ensemble visualization system, based on real-world use-cases from our collaborators in computational and predictive science.
Junpeng Wang 0001, Han-Wei Shen, Guang Lin 0001
IEEE Trans. Vis. Comput. Graph.4
2016 The stabilization of BAM neural networks with time-varying delays in the leakage terms via sampled-data control
Li Li 0046, Yongqing Yang, Guang Lin 0001
Neural Comput. Appl.3
2013 Exploring Cloud Computing for Large-Scale Scientific Applications
abstract
This paper explores cloud computing for large-scale data intensive scientific applications. Cloud computing is attractive because it provides hardware and software resources on-demand, which relieves the burden of acquiring and maintaining a huge amount of resources that may be used only once by a scientific application. However, unlike typical commercial applications that often just requires a moderate amount of ordinary resources, large-scale scientific applications often need to process enormous amount of data in the terabyte or even petabyte range and require special high performance hardware with low latency connections to complete computation in a reasonable amount of time. To address these challenges, we build an infrastructure that can dynamically select high performance computing hardware across institutions and dynamically adapt the computation to the selected resources to achieve high performance. We have also demonstrated the effectiveness of our infrastructure by building a system biology application and an uncertainty quantification application for carbon sequestration, which can efficiently utilize data and computation resources across several institutions.
Guang Lin 0001, Binh Han, Jian Yin 0002, Ian Gorton
SERVICES1
2013 Hybrid parallel computing of minimum action method
Xiaoliang Wan, Guang Lin 0001
Parallel Comput.2