EDBT 2026 Demo / reviewers in the wild / expert
Lantao Yu
dblp:186/7892
· DBLP profile ↗
33ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0003-0569-952XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Blind Image Super-Resolution with Local and Global Dual-GuidanceabstractBlind image super-resolution (BISR) aims to recover the high-resolution image from its degraded low-resolution version with unknown degradation. Recent research on BISR has demonstrated impressive results by using convolutional neural network (CNN) based techniques. However, these methods suffer from limited receptive fields brought by CNN. In addition, they have limited adaptivity to different frequency components of natural images. To address those drawbacks, we propose a deep local and global dual-guidance degradation-adaptive BISR network that exploits global information by involving Fourier coefficients in the degradation representation and image reconstruction process. Additionally, we propose a novel frequency component enhancement module that explicitly decomposes images into multiple frequency bands and assign different weights for each band so as to construct high-quality image. Experimental results demonstrate the superior performance of our method. Yajun Qiu, Shuyuan Zhu, Lantao Yu, Bing Zeng 0001 |
MMSP | 3 |
| 2025 | A Data Perspective on Enhanced Identity Preservation for Diffusion PersonalizationabstractLarge text-to-image models have revolutionized the ability to generate imagery using natural language. However, particularly unique or personal visual concepts, such as pets and furniture, will not be captured by the original model. This has led to interest in how to personalize a text-to-image model. Despite significant progress, this task remains a formidable challenge, particularly in preserving the subject's identity. Most researchers attempt to address this issue by modifying model architectures. These methods are capable of keeping the subject structure and color but fail to preserve identity details. Towards this issue, our approach takes a data-centric perspective. We introduce a novel regularization dataset generation strategy on both the text and image level. This strategy enables the model to preserve fine details of the desired subjects, such as text and logos. Our method is architecture-agnostic and can be flexibly applied on various text-to-image models. We show on established benchmarks that our data-centric approach forms the new state of the art in terms of identity preservation and text alignment. Xingzhe He, Zhiwen Cao, Nicholas I. Kolkin, Lantao Yu, Kun Wan 0001, Helge Rhodin, Ratheesh Kalarot |
WACV | 4 |
| 2024 | DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D VisionabstractWe have witnessed significant progress in deep learning-based 3D vision, ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However, existing scene-level datasets for deep learning-based 3D vision, limited to ei-ther synthetic environments or a narrow selection of real-world scenes, are quite insufficient. This insufficiency not only hinders a comprehensive benchmark of existing methods but also caps what could be explored in deep learning-based 3D analysis. To address this critical gap, we present DL3DV-10K, a large-scale scene dataset, featuring 51.2 million frames from 10,510 videos captured from 65 types of point- of-interest (POI) locations, covering both bounded and unbounded scenes, with different levels of reflection, transparency, and lighting. We conducted a comprehensive benchmark of recent NVS methods on DL3DV-10K, which revealed valuable insights for future research in NVS. In addition, we have obtained encouraging results in a pilot study to learn generalizable NeRF from DL3DV-10K, which manifests the necessity of a large-scale scene-level dataset to forge a path toward a foundation model for learning 3D representation. Our DL3DV-10K dataset, benchmark results, and models will be publicly accessible. Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan 0001, Lantao Yu, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua 0001, Tianyi Zhang 0001, Bedrich Benes, Aniket Bera |
CVPR | 7 |
| 2024 | I-Matting: Improved Trimap-Free Image MattingabstractImage matting has become an essential functionality of image capturing and editing tools. While trimap and scribble-based techniques have shown notable success in these applications, generating high-quality alpha mattes without trimap inputs remains challenging. Existing trimap-free methods divide the task into coarse semantic mask prediction and detailed matte prediction, and an optimization is formulated by balancing these two tasks. However, emphasizing the optimization of the coarse mask leads to inaccurate matte, and emphasizing the optimization of the detailed matte leads to degraded semantic integrity or background artifacts. In this paper, we propose an improved trimap-free training strategy (I-Matting) that effectively ensures semantic integrity, removes background artifacts, and improves local details. First, we introduce two discriminators to distinguish the matting outputs versus the ground truths, which boosts the semantic without hurting the matte prediction. Second, a novel patch-rank module is proposed to improve the matting accuracy by leveraging high-resolution inputs, without hurting the semantic integrity. Meanwhile, the accuracy gain produced by I-Matting is not at the expense of any additional cost in the inference. Extensive experiments show that our method significantly outperforms existing approaches. Zichuan Liu, Mingyuan Wu, Lantao Yu, Klara Nahrstedt |
ICME | 4 |
| 2023 | Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution MatchingabstractOffline imitation learning (IL) promises the ability to learn performant policies from pre-collected demonstrations without interactions with the environment. However, imitating behaviors fully offline typically requires numerous expert data. To tackle this issue, we study the setting where we have limited expert data and supplementary suboptimal data. In this case, a well-known issue is the distribution shift between the learned policy and the behavior policy that collects the offline data. Prior works mitigate this issue by regularizing the KL divergence between the stationary state-action distributions of the learned policy and the behavior policy. We argue that such constraints based on exact distribution matching can be overly conservative and hamper policy learning, especially when the imperfect offline data is highly suboptimal. To resolve this issue, we present RelaxDICE, which employs an asymmetrically-relaxed f-divergence for explicit support regularization. Specifically, instead of driving the learned policy to exactly match the behavior policy, we impose little penalty whenever the density ratio between their stationary state-action distributions is upper bounded by a constant. Note that such formulation leads to a nested min-max optimization problem, which causes instability in practice. RelaxDICE addresses this challenge by supporting a closed-form solution for the inner maximization problem. Extensive empirical study shows that our method significantly outperforms the best prior offline IL method in six standard continuous control environments with over 30% performance gain on average, across 22 settings where the imperfect dataset is highly suboptimal. Lantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger, Stefano Ermon |
AAAI | 1 |
| 2022 | GeoDiff: A Geometric Diffusion Model for Molecular Conformation Generation
Minkai Xu, Lantao Yu, Yang Song 0011, Chence Shi, Stefano Ermon, Jian Tang 0005 |
ICLR | 2 |
| 2022 | A General Recipe for Likelihood-free Bayesian OptimizationabstractThe acquisition function, a critical component in Bayesian optimization (BO), can often be written as the expectation of a utility function under a surrogate model. However, to ensure that acquisition functions are tractable to optimize, restrictions must be placed on the surrogate model and utility function. To extend BO to a broader class of models and utilities, we propose likelihood-free BO (LFBO), an approach based on likelihood-free inference. LFBO directly models the acquisition function without having to separately perform inference with a probabilistic surrogate model. We show that computing the acquisition function in LFBO can be reduced to optimizing a weighted classification problem, which extends an existing likelihood-free density ratio estimation method related to probability of improvement (PI). By choosing the utility function for expected improvement (EI), LFBO outperforms the aforementioned method, as well as various state-of-the-art black-box optimization methods on several real-world optimization problems. LFBO can also leverage composite structures of the objective function, which further improves its regret by several orders of magnitude. Jiaming Song, Lantao Yu, Willie Neiswanger, Stefano Ermon |
ICML | 2 |
| 2022 | Generalizing Bayesian Optimization with Decision-theoretic EntropiesabstractBayesian optimization (BO) is a popular method for efficiently inferring optima of an expensive black-box function via a sequence of queries. Existing information-theoretic BO procedures aim to make queries that most reduce the uncertainty about optima, where the uncertainty is captured by Shannon entropy. However, an optimal measure of uncertainty would, ideally, factor in how we intend to use the inferred quantity in some downstream procedure. In this paper, we instead consider a generalization of Shannon entropy from work in statistical decision theory (DeGroot 1962, Rao 1984), which contains a broad class of uncertainty measures parameterized by a problem-specific loss function corresponding to a downstream task. We first show that special cases of this entropy lead to popular acquisition functions used in BO procedures such as knowledge gradient, expected improvement, and entropy search. We then show how alternative choices for the loss yield a flexible family of acquisition functions that can be customized for use in novel optimization settings. Additionally, we develop gradient-based methods to efficiently optimize our proposed family of acquisition functions, and demonstrate strong empirical performance on a diverse set of sequential decision making tasks, including variants of top-$k$ optimization, multi-level set estimation, and sequence search. Willie Neiswanger, Lantao Yu, Shengjia Zhao, Chenlin Meng, Stefano Ermon |
NeurIPS | 2 |
| 2022 | Fast and High-Quality Blind Multi-Spectral Image PansharpeningabstractBlind pansharpening addresses the problem of generating a high spatial-resolution multi-spectral (HRMS) image given a low spatial-resolution multi-spectral (LRMS) image with the guidance of its associated spatially misaligned high spatial-resolution panchromatic (PAN) image without parametric side information. In this article, we propose a fast approach to blind pansharpening and achieve the state-of-the-art image reconstruction quality. Typical blind pansharpening algorithms are often computationally intensive since the blur kernel and the target HRMS image are often computed using iterative solvers and in an alternating fashion. To achieve fast blind pansharpening, we decouple the solution of the blur kernel and of the HRMS image. First, we estimate the blur kernel by computing the kernel coefficients with minimum total generalized variation that blur a downsampled version of the PAN image to approximate a linear combination of the LRMS image channels. Then, we estimate each channel of the HRMS image using local Laplacian prior (LLP) to regularize the relationship between each HRMS channel and the PAN image. Solving the HRMS image is accelerated by both parallelizing across the channels and by fast numerical algorithms for each channel. Due to the fast scheme and the powerful priors we used on the blur kernel coefficients (total generalized variation) and on the cross-channel relationship (LLP), numerical experiments demonstrate that our algorithm outperforms the state-of-the-art model-based counterparts in terms of both computational time and reconstruction quality of the HRMS images. Lantao Yu, Dehong Liu, Hassan Mansour, Petros Boufounos |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | A Discrete Scheme for Computing Image's Weighted Gaussian CurvatureabstractWeighted Gaussian curvature is an important smoothness measurement for images. However, its conventional computation scheme has low performance, low accuracy and requires that the input image must be second order differentiable. To tackle these three issues, we propose a novel discrete computation scheme for the weighted Gaussian curvature. Our scheme does not require the second order differentiability. Moreover, our scheme is more accurate, has smaller support region and computationally more efficient than the conventional schemes. Therefore, our scheme holds promise for a large range of applications where the weighted Gaussian curvature is needed, for example, image smoothing, cartoon texture decomposition, optical flow estimation, etc. Yuanhao Gong, Wenming Tang, Lebin Zhou, Lantao Yu, Guoping Qiu |
ICIP | 4 |
| 2021 | Quarter Laplacian Filter For Edge Aware Image ProcessingabstractThis paper presents a quarter Laplacian filter that can preserve corners and edges during image smoothing. Its support region is $2\times 2$, which is smaller than the $3\times 3$ support region of the classical Laplacian filter. Thus, it is more local. Moreover, this filter can be implemented via the classical box filter, leading to high performance for real time applications. Finally, we show its edge preserving property in several image processing tasks, including image smoothing, texture enhancement, and low-light image enhancement. The proposed filter can be adopted in a wide range of image processing applications. Yuanhao Gong, Wenming Tang, Lebin Zhou, Lantao Yu, Guoping Qiu |
ICIP | 4 |
| 2021 | Pseudo-Spherical Contrastive DivergenceabstractEnergy-based models (EBMs) offer flexible distribution parametrization. However, due to the intractable partition function, they are typically trained via contrastive divergence for maximum likelihood estimation. In this paper, we propose pseudo-spherical contrastive divergence (PS-CD) to generalize maximum likelihood learning of EBMs. PS-CD is derived from the maximization of a family of strictly proper homogeneous scoring rules, which avoids the computation of the intractable partition function and provides a generalized family of learning objectives that include contrastive divergence as a special case. Moreover, PS-CD allows us to flexibly choose various learning objectives to train EBMs without additional computational cost or variational minimax optimization. Theoretical analysis on the proposed method and extensive experiments on both synthetic data and commonly used image datasets demonstrate the effectiveness and modeling flexibility of PS-CD, as well as its robustness to data contamination, thus showing its superiority over maximum likelihood and $f$-EBMs. Lantao Yu, Jiaming Song, Yang Song 0011, Stefano Ermon |
NeurIPS | 1 |
| 2021 | Multi-agent Imitation Learning with Copulas
Hongwei Wang 0004, Lantao Yu, Zhangjie Cao, Stefano Ermon |
ECML/PKDD (1) | 2 |
| 2020 | Infomax Neural Joint Source-Channel Coding via Adversarial Bit FlipabstractAlthough Shannon theory states that it is asymptotically optimal to separate the source and channel coding as two independent processes, in many practical communication scenarios this decomposition is limited by the finite bit-length and computational power for decoding. Recently, neural joint source-channel coding (NECST) (Choi et al. 2018) is proposed to sidestep this problem. While it leverages the advancements of amortized inference and deep learning (Kingma and Welling 2013; Grover and Ermon 2018) to improve the encoding and decoding process, it still cannot always achieve compelling results in terms of compression and error correction performance due to the limited robustness of its learned coding networks. In this paper, motivated by the inherent connections between neural joint source-channel coding and discrete representation learning, we propose a novel regularization method called Infomax Adversarial-Bit-Flip (IABF) to improve the stability and robustness of the neural joint source-channel coding scheme. More specifically, on the encoder side, we propose to explicitly maximize the mutual information between the codeword and data; while on the decoder side, the amortized reconstruction is regularized within an adversarial framework. Extensive experiments conducted on various real-world datasets evidence that our IABF can achieve state-of-the-art performances on both compression and error correction benchmarks and outperform the baselines by a significant margin. Yuxuan Song 0002, Minkai Xu, Lantao Yu, Hao Zhou 0012, Shuo Shao 0001, Yong Yu 0001 |
AAAI | 3 |
| 2020 | Improving Maximum Likelihood Training for Text Generation with Density Ratio EstimationabstractAutoregressive neural sequence generative models trained by Maximum Likelihood Estimation suffer the exposure bias problem in practical finite sample scenarios. The crux is that the number of training samples for Maximum Likelihood Estimation is usually limited and the input data distributions are different at training and inference stages. Many methods have been proposed to solve the above problem, which relies on sampling from the non-stationary model distribution and suffers from high variance or biased estimations. In this paper, we propose $\psi$-MLE, a new training scheme for autoregressive sequence generative models, which is effective and stable when operating at large sample space encountered in text generation. We derive our algorithm from a new perspective of self-augmentation and introduce bias correction with density ratio estimation. Extensive experimental results on synthetic data and real-world text generation tasks demonstrate that our method stably outperforms Maximum Likelihood Estimation and other state-of-the-art sequence generative models in terms of both quality and diversity. Yuxuan Song 0002, Ning Miao, Hao Zhou 0012, Lantao Yu, Mingxuan Wang, Lei Li 0005 |
AISTATS | 4 |
| 2020 | Improving Unsupervised Domain Adaptation with Variational Information BottleneckabstractDomain adaptation aims to leverage the supervision signal of source domain to obtain an accurate model for target domain, where the labels are not available. To leverage and adapt the label information from source domain, most existing methods employ a feature extracting function and match the marginal distributions of source and target domains in a shared feature space. In this paper, from the perspective of information theory, we show that representation matching is actually an insufficient constraint on the feature space for obtaining a model with good generalization performance in target domain. We then propose variational bottleneck domain adaptation (VBDA), a new domain adaptation method which improves feature transferability by explicitly enforcing the feature extractor to ignore the task-irrelevant factors and focus on the information that is essential to the task of interest for both source and target domains. Extensive experimental results demonstrate that VBDA significantly outperforms state-of-the-art methods across three domain adaptation benchmark datasets. Yuxuan Song 0002, Lantao Yu, Zhangjie Cao, Zhiming Zhou 0001, Jian Shen 0003, Shuo Shao 0001, Weinan Zhang 0001, Yong Yu 0001 |
ECAI | 2 |
| 2020 | Blind Multi-Spectral Image Pan-SharpeningabstractWe address the problem of sharpening low spatial-resolution multi-spectral (MS) images with their associated misaligned high spatial-resolution panchromatic (PAN) image, based on priors on the spatial blur kernel and on the cross-channel relationship. In particular, we formulate the blind pan-sharpening problem within a multi-convex optimization framework using total generalized variation for the blur kernel and local Laplacian prior for the cross-channel relationship. The problem is solved by the alternating direction method of multipliers (ADMM), which alternately updates the blur kernel and sharpens intermediate MS images. Numerical experiments demonstrate that our approach is more robust to large misalignment errors and yields better super resolved MS images compared to state-of-the-art optimization-based and deep-learning-based algorithms. Lantao Yu, Dehong Liu, Hassan Mansour, Petros Boufounos, Yanting Ma |
ICASSP | 1 |
| 2020 | Training Deep Energy-Based Models with f-Divergence MinimizationabstractDeep energy-based models (EBMs) are very flexible in distribution parametrization but computationally challenging because of the intractable partition function. They are typically trained via maximum likelihood, using contrastive divergence to approximate the gradient of the KL divergence between data and model distribution. While KL divergence has many desirable properties, other f-divergences have shown advantages in training implicit density generative models such as generative adversarial networks. In this paper, we propose a general variational framework termed f-EBM to train EBMs using any desired f-divergence. We introduce a corresponding optimization algorithm and prove its local convergence property with non-linear dynamical systems theory. Experimental results demonstrate the superiority of f-EBM over contrastive divergence, as well as the benefits of training EBMs using f-divergences other than KL. Lantao Yu, Yang Song 0011, Jiaming Song, Stefano Ermon |
ICML | 1 |
| 2020 | Autoregressive Score MatchingabstractAutoregressive models use chain rule to define a joint probability distribution as a product of conditionals. These conditionals need to be normalized, imposing constraints on the functional families that can be used. To increase flexibility, we propose autoregressive conditional score models (AR-CSM) where we parameterize the joint distribution in terms of the derivatives of univariate log-conditionals (scores), which need not be normalized. To train AR-CSM, we introduce a new divergence between distributions named Composite Score Matching (CSM). For AR-CSM models, this divergence between data and model distributions can be computed and optimized efficiently, requiring no expensive sampling or adversarial training. Compared to previous score matching algorithms, our method is more scalable to high dimensional data and more stable to optimize. We show with extensive experimental results that it can be applied to density estimation on synthetic data, image generation, image denoising, and training latent variable models with implicit encoders. Chenlin Meng, Lantao Yu, Yang Song 0011, Jiaming Song, Stefano Ermon |
NeurIPS | 2 |
| 2020 | MOPO: Model-based Offline Policy OptimizationabstractOffline reinforcement learning (RL) refers to the problem of learning policies entirely from a batch of previously collected data. This problem setting is compelling, because it offers the promise of utilizing large, diverse, previously collected datasets to acquire policies without any costly or dangerous active exploration, but it is also exceptionally difficult, due to the distributional shift between the offline training data and the learned policy. While there has been significant progress in model-free offline RL, the most successful prior methods constrain the policy to the support of the data, precluding generalization to new states. In this paper, we observe that an existing model-based RL algorithm on its own already produces significant gains in the offline setting, as compared to model-free approaches, despite not being designed for this setting. However, although many standard model-based RL methods already estimate the uncertainty of their model, they do not by themselves provide a mechanism to avoid the issues associated with distributional shift in the offline setting. We therefore propose to modify existing model-based RL methods to address these issues by casting offline model-based RL into a penalized MDP framework. We theoretically show that, by using this penalized MDP, we are maximizing a lower bound of the return in the true MDP. Based on our theoretical results, we propose a new model-based offline RL algorithm that applies the variance of a Lipschitz-regularized model as a penalty to the reward function. We find that this algorithm outperforms both standard model-based RL methods and existing state-of-the-art model-free offline RL approaches on existing offline RL benchmarks, as well as two challenging continuous control tasks that require generalizing from data collected for a different task. Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou 0001, Sergey Levine, Chelsea Finn, Tengyu Ma 0001 |
NeurIPS | 3 |
| 2019 | Deep Reinforcement Learning for Green Security Games with Real-Time InformationabstractGreen Security Games (GSGs) have been proposed and applied to optimize patrols conducted by law enforcement agencies in green security domains such as combating poaching, illegal logging and overfishing. However, real-time information such as footprints and agents’ subsequent actions upon receiving the information, e.g., rangers following the footprints to chase the poacher, have been neglected in previous work. To fill the gap, we first propose a new game model GSG-I which augments GSGs with sequential movement and the vital element of real-time information. Second, we design a novel deep reinforcement learning-based algorithm, DeDOL, to compute a patrolling strategy that adapts to the real-time information against a best-responding attacker. DeDOL is built upon the double oracle framework and the policy-space response oracle, solving a restricted game and iteratively adding best response strategies to it through training deep Q-networks. Exploring the game structure, DeDOL uses domain-specific heuristic strategies as initial strategies and constructs several local modes for efficient and parallelized training. To our knowledge, this is the first attempt to use Deep Q-Learning for security games. Zheyuan Shi, Lantao Yu, Yi Wu 0013, Lucas Joppa, Fei Fang 0001 |
AAAI | 3 |
| 2019 | Single Image Interpolation Exploiting Semi-local SimilarityabstractThis paper explores the modeling and exploitation of semi-local similarity in natural images to address the ill-posed nature of image interpolation. Our approach distinguishes itself from prior approaches by direct and careful use of semi-local similar patches to interpolate each individual patch. Our work uses a simple, parallelizable algorithm without the need to solve complicated optimization problems. Experimental results demonstrate that our interpolated images achieve significantly higher objective and subjective quality compared with those from state-of-the-art algorithms. Lantao Yu, Michael T. Orchard |
ICASSP | 1 |
| 2019 | When Spatially-Variant Filtering Meets Low-Rank Regularization: Exploiting Non-Local Similarity for Single Image InterpolationabstractThis paper combines spatially-variant filtering and non-local low-rank regularization (NLR) to exploit non-local similarity in natural images in addressing the problem of image interpolation. We propose to build a carefully designed spatially-variant, non-local filtering scheme to generate a reliable estimate of the interpolated image and utilize NLR to refine the estimation. Our work uses a simple, parallelizable algorithm without the need to solve complicated optimization problems. Experiment results demonstrate that our algorithm significantly improves PSNR and SSIM of the interpolated images compared with state-of-the-art algorithms. Lantao Yu, Michael T. Orchard |
ICIP | 1 |
| 2019 | Accurate Edge Location Identification Based on Location-Directed Image ModelingabstractThis paper introduces a new approach for determining accurate locations of edges in natural images based on a location-directed image modeling framework. The inability to identify accurate locations follows from the inability to characterize the continuous location variation of edges. We exploit the linear phase-location characteristics in a complex-valued image representation, which guarantees even minuscule location variation be captured by changes in phases. In each band of complex-valued coefficients, the fields of phases represent the locations of edges with particular resolution and along particular direction, and thereby the representation offers a nonparametric framework for gathering multiresolution and multi-directional pieces of evidence about an edge's locations to jointly locate the edge. Our approach quantifies an edge's spatial shift that is less than 0.01 pixel and demonstrates its periodic movement within 0.35-pixel range in a series of natural images. This remarkable performance verifies our framework's ability in identifying accurate locations of edges and opens the door to unveil imperceptible phenomena that are not previously detected. Lantao Yu, Michael T. Orchard |
ICIP | 1 |
| 2019 | CoT: Cooperative Training for Generative Modeling of Discrete DataabstractIn this paper, we study the generative models of sequential discrete data. To tackle the exposure bias problem inherent in maximum likelihood estimation (MLE), generative adversarial networks (GANs) are introduced to penalize the unrealistic generated samples. To exploit the supervision signal from the discriminator, most previous models leverage REINFORCE to address the non-differentiable problem of sequential discrete data. However, because of the unstable property of the training signal during the dynamic process of adversarial training, the effectiveness of REINFORCE, in this case, is hardly guaranteed. To deal with such a problem, we propose a novel approach called Cooperative Training (CoT) to improve the training of sequence generative models. CoT transforms the min-max game of GANs into a joint maximization framework and manages to explicitly estimate and optimize Jensen-Shannon divergence. Moreover, CoT works without the necessity of pre-training via MLE, which is crucial to the success of previous methods. In the experiments, compared to existing state-of-the-art methods, CoT shows superior or at least competitive performance on sample quality, diversity, as well as training stability. Sidi Lu, Lantao Yu, Siyuan Feng 0007, Yaoming Zhu, Weinan Zhang 0001 |
ICML | 2 |
| 2019 | Multi-Agent Adversarial Inverse Reinforcement LearningabstractReinforcement learning agents are prone to undesired behaviors due to reward mis-specification. Finding a set of reward functions to properly guide agent behaviors is particularly challenging in multi-agent scenarios. Inverse reinforcement learning provides a framework to automatically acquire suitable reward functions from expert demonstrations. Its extension to multi-agent settings, however, is difficult due to the more complex notions of rational behaviors. In this paper, we propose MA-AIRL, a new framework for multi-agent inverse reinforcement learning, which is effective and scalable for Markov games with high-dimensional state-action space and unknown dynamics. We derive our algorithm based on a new solution concept and maximum pseudolikelihood estimation within an adversarial reward learning framework. In the experiments, we demonstrate that MA-AIRL can recover reward functions that are highly correlated with the ground truth rewards, while significantly outperforms prior methods in terms of policy imitation. Lantao Yu, Jiaming Song, Stefano Ermon |
ICML | 1 |
| 2019 | Lipschitz Generative Adversarial NetsabstractIn this paper we show that generative adversarial networks (GANs) without restriction on the discriminative function space commonly suffer from the problem that the gradient produced by the discriminator is uninformative to guide the generator. By contrast, Wasserstein GAN (WGAN), where the discriminative function is restricted to 1-Lipschitz, does not suffer from such a gradient uninformativeness problem. We further show in the paper that the model with a compact dual form of Wasserstein distance, where the Lipschitz condition is relaxed, may also theoretically suffer from this issue. This implies the importance of Lipschitz condition and motivates us to study the general formulation of GANs with Lipschitz constraint, which leads to a new family of GANs that we call Lipschitz GANs (LGANs). We show that LGANs guarantee the existence and uniqueness of the optimal discriminative function as well as the existence of a unique Nash equilibrium. We prove that LGANs are generally capable of eliminating the gradient uninformativeness problem. According to our empirical analysis, LGANs are more stable and generate consistently higher quality samples compared with WGAN. Zhiming Zhou 0001, Jiadong Liang, Yuxuan Song 0002, Lantao Yu, Hongwei Wang 0004, Weinan Zhang 0001, Yong Yu 0001, Zhihua Zhang 0004 |
ICML | 4 |
| 2019 | Meta-Inverse Reinforcement Learning with Probabilistic Context VariablesabstractReinforcement learning demands a reward function, which is often difficult to provide or design in real world applications. While inverse reinforcement learning (IRL) holds promise for automatically learning reward functions from demonstrations, several major challenges remain. First, existing IRL methods learn reward functions from scratch, requiring large numbers of demonstrations to correctly infer the reward for each task the agent may need to perform. Second, and more subtly, existing methods typically assume demonstrations for one, isolated behavior or task, while in practice, it is significantly more natural and scalable to provide datasets of heterogeneous behaviors. To this end, we propose a deep latent variable model that is capable of learning rewards from unstructured, multi-task demonstration data, and critically, use this experience to infer robust rewards for new, structurally-similar tasks from a single demonstration. Our experiments on multiple continuous control tasks demonstrate the effectiveness of our approach compared to state-of-the-art imitation and inverse reinforcement learning methods. Lantao Yu, Tianhe Yu, Chelsea Finn, Stefano Ermon |
NeurIPS | 1 |
| 2018 | Exploiting Data and Human Knowledge for Predicting Wildlife PoachingabstractPoaching continues to be a significant threat to the conservation of wildlife and the associated ecosystem. Estimating and predicting where the poachers have committed or would commit crimes is essential to more effective allocation of patrolling resources. The real-world data in this domain is often sparse, noisy and incomplete, consisting of a small number of positive data (poaching signs), a large number of negative data with label uncertainty, and an even larger number of unlabeled data. Fortunately, domain experts such as rangers can provide complementary information about poaching activity patterns. However, this kind of human knowledge has rarely been used in previous approaches. Swaminathan Gurumurthy, Lantao Yu, Chenyan Zhang, Yongchao Jin, Fei Fang 0001 |
COMPASS | 2 |
| 2018 | Location-Directed Image Modeling and its Application to Image InterpolationabstractThis paper explores the development and the use of a complex-valued, multi-resolution image representation to model and exploit local image structures in image-processing applications. Our approach distinguishes itself from prior approaches by constructing and exploiting the direct relationship between the locations of local structures and representation coefficients. Coefficients of our representation have magnitudes that measure the image energy within a specific region in both space and frequency, and phases that carry information about the distance of that energy from a local reference position. Our work proposes to model relationships, both across spatial regions and across different frequency bands, among the field of coefficient magnitudes, and among the field of coefficient phases. To illustrate the advantages of modeling these relationships, we present an algorithm for interpolating a natural image by a factor of two, both horizontally and vertically. Relationships among magnitudes and phases of available bands of coefficients are exploited to estimate local edge parameters (e.g. location, orientation, sharpness) that provide information about higher-frequency coefficients that are not available in the original image. Our work produces PSNR results that are competitive with state-of-the-art single image interpolation algorithms around edges and preserves both edge sharpness and contour smoothness. Lantao Yu, Michael T. Orchard |
ICIP | 1 |
| 2017 | SeqGAN: Sequence Generative Adversarial Nets with Policy GradientabstractAs a new way of training generative models, Generative Adversarial Net (GAN) that uses a discriminative model to guide the training of the generative model has enjoyed considerable success in generating real-valued data. However, it has limitations when the goal is for generating sequences of discrete tokens. A major reason lies in that the discrete outputs from the generative model make it difficult to pass the gradient update from the discriminative model to the generative model. Also, the discriminative model can only assess a complete sequence, while for a partially generated sequence, it is non-trivial to balance its current score and the future one once the entire sequence has been generated. In this paper, we propose a sequence generation framework, called SeqGAN, to solve the problems. Modeling the data generator as a stochastic policy in reinforcement learning (RL), SeqGAN bypasses the generator differentiation problem by directly performing gradient policy update. The RL reward signal comes from the GAN discriminator judged on a complete sequence, and is passed back to the intermediate state-action steps using Monte Carlo search. Extensive experiments on synthetic data and real-world tasks demonstrate significant improvements over strong baselines. Lantao Yu, Weinan Zhang 0001, Jun Wang 0012, Yong Yu 0001 |
AAAI | 1 |
| 2017 | Dynamic Attention Deep Model for Article Recommendation by Learning Human Editors' DemonstrationabstractAs aggregators, online news portals face great challenges in continuously selecting a pool of candidate articles to be shown to their users. Typically, those candidate articles are recommended manually by platform editors from a much larger pool of articles aggregated from multiple sources. Such a hand-pick process is labor intensive and time-consuming. In this paper, we study the editor article selection behavior and propose a learning by demonstration system to automatically select a subset of articles from the large pool. Our data analysis shows that (i) editors' selection criteria are non-explicit, which are less based only on the keywords or topics, but more depend on the quality and attractiveness of the writing from the candidate article, which is hard to capture based on traditional bag-of-words article representation. And (ii) editors' article selection behaviors are dynamic: articles with different data distribution come into the pool everyday and the editors' preference varies, which are driven by some underlying periodic or occasional patterns. To address such problems, we propose a meta-attention model across multiple deep neural nets to (i) automatically catch the editors' underlying selection criteria via the automatic representation learning of each article and its interaction with the meta data and (ii) adaptively capture the change of such criteria via a hybrid attention model. The attention model strategically incorporates multiple prediction models, which are trained in previous days. The system has been deployed in a commercial article feed platform. A 9-day A/B testing has demonstrated the consistent superiority of our proposed model over several strong baselines. Xuejian Wang, Lantao Yu, Kan Ren, Guanyu Tao, Weinan Zhang 0001, Yong Yu 0001, Jun Wang 0012 |
KDD | 2 |
| 2017 | IRGAN: A Minimax Game for Unifying Generative and Discriminative Information Retrieval ModelsabstractThis paper provides a unified account of two schools of thinking in information retrieval modelling: the generative retrieval focusing on predicting relevant documents given a query, and the discriminative retrieval focusing on predicting relevancy given a query-document pair. We propose a game theoretical minimax game to iteratively optimise both models. On one hand, the discriminative model, aiming to mine signals from labelled and unlabelled data, provides guidance to train the generative model towards fitting the underlying relevance distribution over documents given the query. On the other hand, the generative model, acting as an attacker to the current discriminative model, generates difficult examples for the discriminative model in an adversarial way by minimising its discrimination objective. With the competition between these two models, we show that the unified framework takes advantage of both schools of thinking: (i) the generative model learns to fit the relevance distribution over documents via the signals from the discriminative model, and (ii) the discriminative model is able to exploit the unlabelled data selected by the generative model to achieve a better estimation for document ranking. Our experimental results have demonstrated significant performance gains as much as 23.96% on [email protected] and 15.50% on MAP over strong baselines in a variety of applications including web search, item recommendation, and question answering. Jun Wang 0012, Lantao Yu, Weinan Zhang 0001, Benyou Wang, Peng Zhang 0002, Dell Zhang |
SIGIR | 2 |