Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sulin Liu

dblp:192/1289 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 56% Probabilistic and Bayesian machine learning · 15% Learning paradigms · 12%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model
0.912025
Think while You Generate: Discrete Diffusion with Planned Denoising · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Think while You Generate: Discrete Diffusion with Planned Denoising · ICLR 2025
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.912025
Think while You Generate: Discrete Diffusion with Planned Denoising · ICLR 2025
Machine learning › Generative modeling › generative model
discrete generative model
0.812024
Generative Marginalization Models · ICML 2024
Machine learning › Generative modeling
energy-based model
0.812024
Generative Marginalization Models · ICML 2024
Machine learning › Learning paradigms
multi-task learning
0.622017
Distributed Multi-Task Relationship Learning · KDD 2017
Adaptive Group Sparse Multi-task Learning via Trace Lasso · IJCAI 2017
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
amortized inference
0.412020
Task-Agnostic Amortized Inference of Gaussian Process Hyperparameters · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.412020
Task-Agnostic Amortized Inference of Gaussian Process Hyperparameters · NeurIPS 2020
Machine learning › Trustworthy machine learning
adversarial machine learning
0.312018
Data Poisoning Attacks on Multi-Task Relationship Learning · AAAI 2018
Machine learning › Trustworthy machine learning › robustness
data poisoning
0.312018
Data Poisoning Attacks on Multi-Task Relationship Learning · AAAI 2018
Machine learning › Efficient and distributed learning › distributed training
distributed multi-task learning
0.312017
Distributed Multi-Task Relationship Learning · KDD 2017
Machine learning › Efficient and distributed learning
distributed training
0.312017
Distributed Multi-Task Relationship Learning · KDD 2017
Machine learning › Learning paradigms › multi-task learning
task relationship modeling
0.312017
Distributed Multi-Task Relationship Learning · KDD 2017
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
marginal inference
0.212024
Generative Marginalization Models · ICML 2024

Methods — techniques the papers use, named apart from their topics

neural network · 1.2planner-denoiser framework · 0.9mask diffusion · 0.9marginalization self-consistency · 0.8self-attention · 0.4bayesian quadrature · 0.4bayesian optimization · 0.4optimality conditions · 0.3implicit gradient · 0.3bi-level optimization · 0.3primal-dual optimization · 0.3parameter server · 0.3alternating optimization · 0.3
YearPublicationVenuePosition
2025 Think while You Generate: Discrete Diffusion with Planned Denoising
abstract
Discrete diffusion has achieved state-of-the-art performance, outperforming or approaching autoregressive models on standard benchmarks. In this work, we introduce *Discrete Diffusion with Planned Denoising* (DDPD), a novel framework that separates the generation process into two models: a planner and a denoiser. At inference time, the planner selects which positions to denoise next by identifying the most corrupted positions in need of denoising, including both initially corrupted and those requiring additional refinement. This plan-and-denoise approach enables more efficient reconstruction during generation by iteratively identifying and denoising corruptions in the optimal order. DDPD outperforms traditional denoiser-only mask diffusion methods, achieving superior results on language modeling benchmarks such as *text8*, *OpenWebText*, and token-based generation on *ImageNet 256 × 256*. Notably, in language modeling, DDPD significantly reduces the performance gap between diffusion-based and autoregressive methods in terms of generative perplexity. Code is available at [github.com/liusulin/DDPD](https://github.com/liusulin/DDPD).
Sulin Liu, Juno Nam, Hannes Stärk, Tommi S. Jaakkola, Rafael Gómez-Bombarelli
ICLR1
2024 Generative Marginalization Models
abstract
We introduce marginalization models (MAMs), a new family of generative models for high-dimensional discrete data. They offer scalable and flexible generative modeling by explicitly modeling all induced marginal distributions. Marginalization models enable fast approximation of arbitrary marginal probabilities with a single forward pass of the neural network, which overcomes a major limitation of arbitrary marginal inference models, such as any-order autoregressive models. MAMs also address the scalability bottleneck encountered in training any-order generative models for high-dimensional problems under the context of energy-based training, where the goal is to match the learned distribution to a given desired probability (specified by an unnormalized log-probability function such as energy or reward function). We propose scalable methods for learning the marginals, grounded in the concept of "marginalization self-consistency". We demonstrate the effectiveness of the proposed model on a variety of discrete data distributions, including images, text, physical systems, and molecules, for maximum likelihood and energy-based training settings. MAMs achieve orders of magnitude speedup in evaluating the marginal probabilities on both settings. For energy-based training tasks, MAMs enable any-order generative modeling of high-dimensional problems beyond the scale of previous methods. Code is available at github.com/PrincetonLIPS/MaM.
Sulin Liu, Peter J. Ramadge, Ryan P. Adams
ICML1
2023 Sparse Bayesian optimization
abstract
Bayesian optimization (BO) is a powerful approach to sample-efficient optimization of black-box objective functions. However, the application of BO to areas such as recommendation systems often requires taking the interpretability and simplicity of the configurations into consideration, a setting that has not been previously studied in the BO literature. To make BO applicable in this setting, we present several regularization-based approaches that allow us to discover sparse and more interpretable configurations. We propose a novel differentiable relaxation based on homotopy continuation that makes it possible to target sparsity by working directly with $L_0$ regularization. We identify failure modes for regularized BO and develop a hyperparameter-free method, sparsity exploring Bayesian optimization (SEBO) that seeks to simultaneously maximize a target objective and sparsity. SEBO and methods based on fixed regularization are evaluated on synthetic and real-world problems, and we show that we are able to efficiently optimize for sparsity.
Sulin Liu, David Eriksson, Benjamin Letham, Eytan Bakshy
AISTATS1
2020 Revisiting the Landscape of Matrix Factorization
abstract
Prior work has shown that low-rank matrix factorization has infinitely many critical points, each of which is either a global minimum or a (strict) saddle point. We revisit this problem and provide simple, intuitive proofs of a set of extended results for low-rank and general-rank problems. We couple our investigation with a known invariant manifold M0 of gradient flow. This restriction admits a uniform negative upper bound on the least eigenvalue of the Hessian map at all strict saddles in M0. The bound depends on the size of the nonzero singular values and the separation between distinct singular values of the matrix to be factorized.
Hossein Valavi, Sulin Liu, Peter J. Ramadge
AISTATS2
2020 Task-Agnostic Amortized Inference of Gaussian Process Hyperparameters
abstract
Gaussian processes (GPs) are flexible priors for modeling functions. However, their success depends on the kernel accurately reflecting the properties of the data. One of the appeals of the GP framework is that the marginal likelihood of the kernel hyperparameters is often available in closed form, enabling optimization and sampling procedures to fit these hyperparameters to data. Unfortunately, point-wise evaluation of the marginal likelihood is expensive due to the need to solve a linear system; searching or sampling the space of hyperparameters thus often dominates the practical cost of using GPs. We introduce an approach to the identification of kernel hyperparameters in GP regression and related problems that sidesteps the need for costly marginal likelihoods. Our strategy is to "amortize" inference over hyperparameters by training a single neural network, which consumes a set of regression data and produces an estimate of the kernel function, useful across different tasks. To accommodate the varying dimension and cardinality of different regression problems, we use a hierarchical self-attention-based neural network that produces estimates of the hyperparameters which are invariant to the order of the input data points and data dimensions. We show that a single neural model trained on synthetic data is able to generalize directly to several different unseen real-world GP use cases. Our experiments demonstrate that the estimated hyperparameters are comparable in quality to those from the conventional model selection procedures, while being much faster to obtain, significantly accelerating GP regression and its related applications such as Bayesian optimization and Bayesian quadrature. The code and pre-trained model are available at https://github.com/PrincetonLIPS/AHGP.
Sulin Liu, Xingyuan Sun, Peter J. Ramadge, Ryan P. Adams
NeurIPS1
2018 Data Poisoning Attacks on Multi-Task Relationship Learning
abstract
Multi-task learning (MTL) is a machine learning paradigm that improves the performance of each task by exploiting useful information contained in multiple related tasks. However, the relatedness of tasks can be exploited by attackers to launch data poisoning attacks, which has been demonstrated a big threat to single-task learning. In this paper, we provide the first study on the vulnerability of MTL. Specifically, we focus on multi-task relationship learning (MTRL) models, a popular subclass of MTL models where task relationships are quantized and are learned directly from training data. We formulate the problem of computing optimal poisoning attacks on MTRL as a bilevel program that is adaptive to arbitrary choice of target tasks and attacking tasks. We propose an efficient algorithm called PATOM for computing optimal attack strategies. PATOM leverages the optimality conditions of the subproblem of MTRL to compute the implicit gradients of the upper level objective function. Experimental results on real-world datasets show that MTRL models are very sensitive to poisoning attacks and the attacker can significantly degrade the performance of target tasks, by either directly poisoning the target tasks or indirectly poisoning the related tasks exploiting the task relatedness. We also found that the tasks being attacked are always strongly correlated, which provides a clue for defending against such attacks.
Mengchen Zhao, Bo An 0001, Yaodong Yu, Sulin Liu, Sinno Jialin Pan
AAAI4
2017 Adaptive Group Sparse Multi-task Learning via Trace Lasso
abstract
In multi-task learning (MTL), tasks are learned jointly so that information among related tasks is shared and utilized to help improve generalization for each individual task. A major challenge in MTL is how to selectively choose what to share among tasks. Ideally, only related tasks should share information with each other. In this paper, we propose a new MTL method that can adaptively group correlated tasks into clusters and share information among the correlated tasks only. Our method is based on the assumption that each task parameter is a linear combination of other tasks' and the coefficients of the linear combination are active only if there is relatedness between the two tasks. Through introducing trace Lasso penalty on these coefficients, our method is able to adaptively select the subset of coefficients with respect to the tasks that are correlated to the task. Our model frees the process of determining task clustering structure as used in the literature. Efficient optimization methods based on alternating direction method of multipliers (ADMM) is developed to solve the problem. Experimental results on both synthetic and real-world datasets demonstrate the effectiveness of our method in terms of clustering related tasks and generalization performance.
Sulin Liu, Sinno Jialin Pan
IJCAI1
2017 Distributed Multi-Task Relationship Learning
abstract
Multi-task learning aims to learn multiple tasks jointly by exploiting their relatedness to improve the generalization performance for each task. Traditionally, to perform multi-task learning, one needs to centralize data from all the tasks to a single machine. However, in many real-world applications, data of different tasks may be geo-distributed over different local machines. Due to heavy communication caused by transmitting the data and the issue of data privacy and security, it is impossible to send data of different task to a master machine to perform multi-task learning. Therefore, in this paper, we propose a distributed multi-task learning framework that simultaneously learns predictive models for each task as well as task relationships between tasks alternatingly in the parameter server paradigm. In our framework, we first offer a general dual form for a family of regularized multi-task relationship learning methods. Subsequently, we propose a communication-efficient primal-dual distributed optimization algorithm to solve the dual problem by carefully designing local subproblems to make the dual problem decomposable. Moreover, we provide a theoretical convergence analysis for the proposed algorithm, which is specific for distributed multi-task relationship learning. We conduct extensive experiments on both synthetic and real-world datasets to evaluate our proposed framework in terms of effectiveness and convergence.
Sulin Liu, Sinno Jialin Pan, Qirong Ho
KDD1
2017 Communication-Efficient Distributed Primal-Dual Algorithm for Saddle Point Problem
Yaodong Yu, Sulin Liu, Sinno Jialin Pan
UAI2