Jonathan Ho

dblp:80/8677 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
13since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Integrating Equity in Public Sector Data-Driven Decision Making: Exploring the Desired Futures of Underserved Stakeholders
abstract
Public sector agencies aim to innovate not just for efficiency but also to enhance equity. Despite the growing adoption of data-driven decision-making systems in the public sector, efforts to integrate equity as a primary goal often fall short. This typically arises from inadequate early-stage involvement of underserved stakeholders and prevalent misunderstandings concerning the authentic meaning of equity from these stakeholders' perspectives. Our research seeks to address this gap by actively involving undersevered stakeholders in the process of envisioning the integration of equity within public sector data-driven decisions, particularly in the context of a building department in a Northeastern mid-sized U.S. city. Applying a speed dating method with storyboards, we explore diverse equity-centric futures within the realm of local business development, a domain where small businesses, particularly women-and minority-owned businesses, historically confront inequitable distribution of public services. We explored three essential aspects of equity: monitoring equity, resource allocation prioritization, as well as information and equity. Our findings illuminate the complexities of integrating equity into data-driven decisions, offering nuanced insights about the needs of stakeholders. We found that attempts to monitor and incorporate equity goals into public sector decision-making can unexpectedly backfire, inadvertently sparking community apprehension and potentially exacerbating existing inequities. Small business owners, including those identifying as women-and minority-owned, advocated against the use of demographic-based data in equity-focused data-driven decision-making in the public sector, instead emphasizing factors such as community needs, application complexity, and uncertainties inherent in small businesses. Drawing from these insights, we propose design implications to assist designers of public sector data-driven decision-making systems to better accommodate equity considerations.
Seyun Kim, Jonathan Ho, Yinan Li 0008, Bonnie Fan, Willa Yunqi Yang, Jessie Ramey, Sarah E. Fox, Haiyi Zhu, John Zimmerman, Motahhare Eslami
Proc. ACM Hum. Comput. Interact.2
2023 On Distillation of Guided Diffusion Models
abstract
Classifier-free guided diffusion models have recently been shown to be highly effective at high-resolution image generation, and they have been widely used in large-scale diffusion frameworks including DALL.E 2, Stable Diffusion and Imagen. However, a downside of classifier-free guided diffusion models is that they are computationally expensive at inference time since they require evaluating two diffusion models, a class-conditional model and an unconditional model, tens to hundreds of times. To deal with this limitation, we propose an approach to distilling classifier-free guided diffusion models into models that are fast to sample from: Given a pre-trained classifier-free guided model, we first learn a single model to match the output of the combined conditional and unconditional models, and then we progressively distill that model to a diffusion model that requires much fewer sampling steps. For standard diffusion models trained on the pixel-space, our approach is able to generate images visually comparable to that of the original model using as few as 4 sampling steps on ImageNet$64\times 64$and CIFAR-10, achieving FID/IS scores comparable to that of the original model while being up to 256 times faster to sample from. For diffusion models trained on the latent-space (e.g., Stable Diffusion), our approach is able to generate high-fidelity images using as few as 1 to 4 denoising steps, accelerating inference by at least 10-fold compared to existing methods on ImageNet$256\times 256$and LAION datasets. We further demonstrate the effectiveness of our approach on text-guided image editing and inpainting, where our distilled model is able to generate high-quality results using as few as 2–4 denoising steps.
Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma, Stefano Ermon, Jonathan Ho, Tim Salimans
CVPR6
2023 Discrete Predictor-Corrector Diffusion Models for Image Synthesis
José Lezama, Tim Salimans, Lu Jiang 0004, Huiwen Chang, Jonathan Ho, Irfan A. Essa
ICLR5
2023 Novel View Synthesis with Diffusion Models
Daniel Watson, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, Mohammad Norouzi 0002
ICLR4
2023 Image Super-Resolution via Iterative Refinement
abstract
We present SR3, an approach to image Super-Resolution via Repeated Refinement. SR3 adapts denoising diffusion probabilistic models (Ho et al. 2020), (Sohl-Dickstein et al. 2015) to image-to-image translation, and performs super-resolution through a stochastic iterative denoising process. Output images are initialized with pure Gaussian noise and iteratively refined using a U-Net architecture that is trained on denoising at various noise levels, conditioned on a low-resolution input image. SR3 exhibits strong performance on super-resolution tasks at different magnification factors, on faces and natural images. We conduct human evaluation on a standard 8× face super-resolution task on CelebA-HQ for which SR3 achieves a fool rate close to 50%, suggesting photo-realistic outputs, while GAN baselines do not exceed a fool rate of 34%. We evaluate SR3 on a 4× super-resolution task on ImageNet, where SR3 outperforms baselines in human evaluation and classification accuracy of a ResNet-50 classifier trained on high-resolution images. We further show the effectiveness of SR3 in cascaded image generation, where a generative model is chained with super-resolution models to synthesize high-resolution images with competitive FID scores on the class-conditional 256×256 ImageNet generation challenge.
Chitwan Saharia, Jonathan Ho, Tim Salimans, David J. Fleet, Mohammad Norouzi 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Progressive Distillation for Fast Sampling of Diffusion Models
Tim Salimans, Jonathan Ho
ICLR2
2022 Learning Fast Samplers for Diffusion Models by Differentiating Through Sample Quality
Daniel Watson, Jonathan Ho, Mohammad Norouzi 0002
ICLR3
2022 Video Diffusion Models
abstract
Generating temporally coherent high fidelity video is an important milestone in generative modeling research. We make progress towards this milestone by proposing a diffusion model for video generation that shows very promising initial results. Our model is a natural extension of the standard image diffusion architecture, and it enables jointly training from image and video data, which we find to reduce the variance of minibatch gradients and speed up optimization. To generate long and higher resolution videos we introduce a new conditional sampling technique for spatial and temporal video extension that performs better than previously proposed methods. We present the first results on a large text-conditioned video generation task, as well as state-of-the-art results on established benchmarks for video prediction and unconditional video generation. Supplementary material is available at https://video-diffusion.github.io/.
Jonathan Ho, Tim Salimans, Alexey A. Gritsenko, Mohammad Norouzi 0002, David J. Fleet
NeurIPS1
2022 Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
abstract
We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image generation. Our key discovery is that generic large language models (e.g., T5), pretrained on text-only corpora, are surprisingly effective at encoding text for image synthesis: increasing the size of the language model in Imagen boosts both sample fidelity and image-text alignment much more than increasing the size of the image diffusion model. Imagen achieves a new state-of-the-art FID score of 7.27 on the COCO dataset, without ever training on COCO, and human raters find Imagen samples to be on par with the COCO data itself in image-text alignment. To assess text-to-image models in greater depth, we introduce DrawBench, a comprehensive and challenging benchmark for text-to-image models. With DrawBench, we compare Imagen with recent methods including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2, and find that human raters prefer Imagen over other models in side-by-side comparisons, both in terms of sample quality and image-text alignment.
Chitwan Saharia, Saurabh Saxena, Lala Li, Jay Whang, Remi Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi 0002
NeurIPS11
2022 Cascaded Diffusion Models for High Fidelity Image Generation
abstract
We show that cascaded diffusion models are capable of generating high fidelity images on the class-conditional ImageNet generation benchmark, without any assistance from auxiliary image classifiers to boost sample quality. A cascaded diffusion model comprises a pipeline of multiple diffusion models that generate images of increasing resolution, beginning with a standard diffusion model at the lowest resolution, followed by one or more super-resolution diffusion models that successively upsample the image and add higher resolution details. We find that the sample quality of a cascading pipeline relies crucially on conditioning augmentation, our proposed method of data augmentation of the lower resolution conditioning inputs to the super-resolution models. Our experiments show that conditioning augmentation prevents compounding error during sampling in a cascaded model, helping us to train cascading pipelines achieving FID scores of 1.48 at 64x64, 3.52 at 128x128 and 4.88 at 256x256 resolutions, outperforming BigGAN-deep, and classification accuracy scores of 63.02% (top-1) and 84.06% (top-5) at 256x256, outperforming VQ-VAE-2.
Jonathan Ho, Chitwan Saharia, David J. Fleet, Mohammad Norouzi 0002, Tim Salimans
J. Mach. Learn. Res.1
2021 Understanding and Segmenting Human Demonstrations into Reusable Compliant Primitives
abstract
Hard coded robotic manipulation skills work well in known, predictable and repeatable situations. Human environments, however, are better described as dynamic, chaotic, uncertain or unstructured. Therefore, plans relying on preprogrammed trajectories are bound to fail in these settings. In order to increase robustness to uncertainty and avoid coding new skills from scratch, we can make flexible plans that execute existing autonomous primitives based on the sensed state of the environment. A key challenge of this approach is finding the sequence of primitives required to perform the desired task. This work uses a variation of a Hidden Markov Model (HMM) with an augmented particle filter to find the primitive sequence using only a reduced number of human demonstrations. The algorithm was tested on 40 demonstrations of two different manipulation tasks involving six primitives. It was seeded with a single manually labelled demonstration of each task and was able to automatically label the other 38 demonstration sequences with an average success of 81.5%. The results show improved convergence and a 9% increase in accuracy over other versions of the algorithm.
Elena Galbally, Jonathan Ho, Oussama Khatib
IROS2
2021 Structured Denoising Diffusion Models in Discrete State-Spaces
abstract
Denoising diffusion probabilistic models (DDPMs) [Ho et al. 2021] have shown impressive results on image and waveform generation in continuous state spaces. Here, we introduce Discrete Denoising Diffusion Probabilistic Models (D3PMs), diffusion-like generative models for discrete data that generalize the multinomial diffusion model of Hoogeboom et al. [2021], by going beyond corruption processes with uniform transition probabilities. This includes corruption with transition matrices that mimic Gaussian kernels in continuous space, matrices based on nearest neighbors in embedding space, and matrices that introduce absorbing states. The third allows us to draw a connection between diffusion models and autoregressive and mask-based generative models. We show that the choice of transition matrix is an important design decision that leads to improved results in image and text domains. We also introduce a new loss function that combines the variational lower bound with an auxiliary cross entropy loss. For text, this model class achieves strong results on character-level text generation while scaling to large vocabularies on LM1B. On the image dataset CIFAR-10, our models approach the sample quality and exceed the log-likelihood of the continuous-space DDPM model.
Jacob Austin, Daniel D. Johnson 0001, Jonathan Ho, Daniel Tarlow, Rianne van den Berg
NeurIPS3
2021 On Density Estimation with Diffusion Models
Diederik P. Kingma, Tim Salimans, Ben Poole, Jonathan Ho
NeurIPS4
2020 Denoising Diffusion Probabilistic Models
abstract
We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN.
Jonathan Ho, Ajay Jain, Pieter Abbeel
NeurIPS1
2019 Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design
abstract
Flow-based generative models are powerful exact likelihood models with efficient sampling and inference. Despite their computational efficiency, flow-based models generally have much worse density modeling performance compared to state-of-the-art autoregressive models. In this paper, we investigate and improve upon three limiting design choices employed by flow-based models in prior work: the use of uniform noise for dequantization, the use of inexpressive affine flows, and the use of purely convolutional conditioning networks in coupling layers. Based on our findings, we propose Flow++, a new flow-based model that is now the state-of-the-art non-autoregressive model for unconditional density estimation on standard image benchmarks. Our work has begun to close the significant performance gap that has so far existed between autoregressive models and flow-based models.
Jonathan Ho, Xi Chen 0022, Aravind Srinivas, Yan Duan, Pieter Abbeel
ICML1
2019 Bit-Swap: Recursive Bits-Back Coding for Lossless Compression with Hierarchical Latent Variables
abstract
The bits-back argument suggests that latent variable models can be turned into lossless compression schemes. Translating the bits-back argument into efficient and practical lossless compression schemes for general latent variable models, however, is still an open problem. Bits-Back with Asymmetric Numeral Systems (BB-ANS), recently proposed by Townsend et al,. 2019, makes bits-back coding practically feasible for latent variable models with one latent layer, but it is inefficient for hierarchical latent variable models. In this paper we propose Bit-Swap, a new compression scheme that generalizes BB-ANS and achieves strictly better compression rates for hierarchical latent variable models with Markov chain structure. Through experiments we verify that Bit-Swap results in lossless compression rates that are empirically superior to existing techniques.
Friso H. Kingma, Pieter Abbeel, Jonathan Ho
ICML3
2019 Compression with Flows via Local Bits-Back Coding
abstract
Likelihood-based generative models are the backbones of lossless compression due to the guaranteed existence of codes with lengths close to negative log likelihood. However, there is no guaranteed existence of computationally efficient codes that achieve these lengths, and coding algorithms must be hand-tailored to specific types of generative models to ensure computational efficiency. Such coding algorithms are known for autoregressive models and variational autoencoders, but not for general types of flow models. To fill in this gap, we introduce local bits-back coding, a new compression technique for flow models. We present efficient algorithms that instantiate our technique for many popular types of flows, and we demonstrate that our algorithms closely achieve theoretical codelengths for state-of-the-art flow models on high-dimensional data.
Jonathan Ho, Evan Lohn, Pieter Abbeel
NeurIPS1
2018 Meta Learning Shared Hierarchies
Kevin Frans, Jonathan Ho, Xi Chen 0022, Pieter Abbeel, John Schulman
ICLR (Poster)2
2018 Evolved Policy Gradients
abstract
We propose a metalearning approach for learning gradient-based reinforcement learning (RL) algorithms. The idea is to evolve a differentiable loss function, such that an agent, which optimizes its policy to minimize this loss, will achieve high rewards. The loss is parametrized via temporal convolutions over the agent's experience. Because this loss is highly flexible in its ability to take into account the agent's history, it enables fast task learning. Empirical results show that our evolved policy gradient algorithm (EPG) achieves faster learning on several randomized environments compared to an off-the-shelf policy gradient method. We also demonstrate that EPG's learned loss can generalize to out-of-distribution test time tasks, and exhibits qualitatively different behavior from other popular metalearning algorithms.
Rein Houthooft, Phillip Isola, Bradly C. Stadie, Filip Wolski, Jonathan Ho, Pieter Abbeel
NeurIPS6
2017 One-Shot Imitation Learning
abstract
Imitation learning has been commonly applied to solve different tasks in isolation. This usually requires either careful feature engineering, or a significant number of samples. This is far from what we desire: ideally, robots should be able to learn from very few demonstrations of any given task, and instantly generalize to new situations of the same task, without requiring task-specific engineering. In this paper, we propose a meta-learning framework for achieving such capability, which we call one-shot imitation learning. Specifically, we consider the setting where there is a very large (maybe infinite) set of tasks, and each task has many instantiations. For example, a task could be to stack all blocks on a table into a single tower, another task could be to place all blocks on a table into two-block towers, etc. In each case, different instances of the task would consist of different sets of blocks with different initial states. At training time, our algorithm is presented with pairs of demonstrations for a subset of all tasks. A neural net is trained that takes as input one demonstration and the current state (which initially is the initial state of the other demonstration of the pair), and outputs an action with the goal that the resulting sequence of states and actions matches as closely as possible with the second demonstration. At test time, a demonstration of a single instance of a new task is presented, and the neural net is expected to perform well on new instances of this new task. Our experiments show that the use of soft attention allows the model to generalize to conditions and tasks unseen in the training data. We anticipate that by training this model on a much greater variety of tasks and settings, we will obtain a general system that can turn any demonstrations into robust policies that can accomplish an overwhelming variety of tasks.
Yan Duan, Marcin Andrychowicz, Bradly C. Stadie, Jonathan Ho, Jonas Schneider 0002, Ilya Sutskever, Pieter Abbeel, Wojciech Zaremba
NIPS4
2016 Model-Free Imitation Learning with Policy Optimization
abstract
In imitation learning, an agent learns how to behave in an environment with an unknown cost function by mimicking expert demonstrations. Existing imitation learning algorithms typically involve solving a sequence of planning or reinforcement learning problems. Such algorithms are therefore not directly applicable to large, high-dimensional environments, and their performance can significantly degrade if the planning problems are not solved to optimality. Under the apprenticeship learning formalism, we develop alternative model-free algorithms for finding a parameterized stochastic policy that performs at least as well as an expert policy on an unknown cost function, based on sample trajectories from the expert. Our approach, based on policy gradients, scales to large continuous environments with guaranteed convergence to local minima.
Jonathan Ho, Jayesh K. Gupta, Stefano Ermon
ICML1
2016 Generative Adversarial Imitation Learning
abstract
Consider learning a policy from example expert behavior, without interaction with the expert or access to a reinforcement signal. One approach is to recover the expert's cost function with inverse reinforcement learning, then extract a policy from that cost function with reinforcement learning. This approach is indirect and can be slow. We propose a new general framework for directly extracting a policy from data as if it were obtained by reinforcement learning following inverse reinforcement learning. We show that a certain instantiation of our framework draws an analogy between imitation learning and generative adversarial networks, from which we derive a model-free imitation learning algorithm that obtains significant performance gains over existing model-free methods in imitating complex behaviors in large, high-dimensional environments.
Jonathan Ho, Stefano Ermon
NIPS1
2013 Tracking deformable objects with point clouds
abstract
We introduce an algorithm for tracking deformable objects from a sequence of point clouds. The proposed tracking algorithm is based on a probabilistic generative model that incorporates observations of the point cloud and the physical properties of the tracked object and its environment. We propose a modified expectation maximization algorithm to perform maximum a posteriori estimation to update the state estimate at each time step. Our modification makes it practical to perform the inference through calls to a physics simulation engine. This is significant because (i) it allows for the use of highly optimized physics simulation engines for the core computations of our tracking algorithm, and (ii) it makes it possible to naturally, and efficiently, account for physical constraints imposed by collisions, grasping actions, and material properties in the observation updates. Even in the presence of the relatively large occlusions that occur during manipulation tasks, our algorithm is able to robustly track a variety of types of deformable objects, including ones that are one-dimensional, such as ropes; two-dimensional, such as cloth; and three-dimensional, such as sponges. Our implementation can track these objects in real time.
John Schulman, Alex X. Lee, Jonathan Ho, Pieter Abbeel
ICRA3
2013 Learning from Demonstrations Through the Use of Non-rigid Registration
John Schulman, Jonathan Ho, Cameron Lee, Pieter Abbeel
ISRR2