EDBT 2026 Demo / reviewers in the wild / expert
Jiatao Gu
dblp:164/5848
· DBLP profile ↗
66ranked-venue papers
20as first author
37since 2021 · last 2025
0000-0003-3578-2711ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 19 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free SimulationabstractFaithfully reconstructing textured shapes and physical properties from videos presents an intriguing yet challenging problem. Significant efforts have been dedicated to advancing such a system identification problem in this area. Previous methods often rely on heavy optimization pipelines with a differentiable simulator and renderer to estimate physical parameters. However, these approaches frequently necessitate extensive hyperparameter tuning for each scene and involve a costly optimization process, which limits both their practicality and generalizability. In this work, we propose a novel framework, Vid2Sim, a generalizable video-based approach for recovering geometry and physical properties through a mesh-free reduced simulation based on Linear Blend Skinning (LBS), offering high computational efficiency and versatile representation capability. Specifically, Vid2Sim first reconstructs the observed configuration of the physical system from video using a feed-forward neural network trained to capture physical world knowledge. A lightweight optimization pipeline then refines the estimated appearance, geometry, and physical properties to closely align with video observations within just a few minutes. Additionally, after the reconstruction, Vid2Sim enables high-quality, mesh-free simulation with high efficiency. Extensive experiments demonstrate that our method achieves superior accuracy and efficiency in reconstructing geometry and physical properties from video data. Chuhao Chen 0005, Zhiyang Dou, Chen Wang 0049, Yiming Huang 0011, Anjun Chen, Qiao Feng 0001, Jiatao Gu, Lingjie Liu |
CVPR | 7 |
| 2025 | World-consistent Video Diffusion with Explicit 3D ModelingabstractRecent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly generating 3D-consistent content. To address this, we propose World-consistent Video Diffusion (WVD), a novel framework that incorporates explicit 3D supervision using XYZ images, which encode global 3D coordinates for each image pixel. More specifically, we train a diffusion transformer to learn the joint distribution of RGB and XYZ frames. This approach supports multi-task adaptability via a flexible inpainting strategy. For example, WVD can estimate XYZ frames from ground-truth RGB or generate novel RGB frames using XYZ projections along a specified camera trajectory. In doing so, WVD unifies tasks like singleimage-to-3D generation, multi-view stereo, and camera-controlled video generation. Our approach demonstrates competitive performance across multiple benchmarks, providing a scalable solution for 3D-consistent video and image generation with a single pretrained model. Our project website is at https://zqh0253.github.io/wvd. Qihang Zhang, Shuangfei Zhai, Miguel Ángel Bautista Martin, Kevin Miao, Alexander Toshev, Joshua M. Susskind, Jiatao Gu |
CVPR | 7 |
| 2025 | Denoising Autoregressive Transformers for Scalable Text-to-Image GenerationabstractDiffusion models have become the dominant approach for visual generation. They are trained by denoising a Markovian process which gradually adds noise to the input. We argue that the Markovian property limits the model’s ability to fully utilize the generation trajectory, leading to inefficiencies during training and inference. In this paper, we propose DART, a transformer-based model that unifies autoregressive (AR) and diffusion within a non-Markovian framework. DART iteratively denoises image patches spatially and spectrally using an AR model that has the same architecture as standard language models. DART does not rely on image quantization, which enables more effective image modeling while maintaining flexibility. Furthermore, DART seamlessly trains with both text and image data in a unified model. Our approach demonstrates competitive performance on class-conditioned and text-to-image generation tasks, offering a scalable, efficient alternative to traditional diffusion models. Through this unified framework, DART sets a new benchmark for scalable, high-quality image synthesis. Jiatao Gu, Yizhe Zhang 0002, Qihang Zhang, Dinghuai Zhang, Navdeep Jaitly, Joshua M. Susskind, Shuangfei Zhai |
ICLR | 1 |
| 2025 | Normalizing Flows are Capable Generative ModelsabstractNormalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relatively little attention in recent years. In this work, we demonstrate that NFs are more powerful than previously believed. We present TarFlow: a simple and scalable architecture that enables highly performant NF models. TarFlow can be thought of as a Transformer-based variant of Masked Autoregressive Flows (MAFs): it consists of a stack of autoregressive Transformer blocks on image patches, alternating the autoregression direction between layers. TarFlow is straightforward to train end-to-end, and capable of directly modeling and generating pixels. We also propose three key techniques to improve sample quality: Gaussian noise augmentation during training, a post training denoising procedure, and an effective guidance method for both class-conditional and unconditional settings. Putting these together, TarFlow sets new state-of-the-art results on likelihood estimation for images, beating the previous best methods by a large margin, and generates samples with quality and diversity comparable to diffusion models, for the first time with a stand-alone NF model. We make our code available at https://github.com/apple/ml-tarflow. Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Ángel Bautista 0001, Navdeep Jaitly, Joshua M. Susskind |
ICML | 5 |
| 2025 | TADA: Improved Diffusion Sampling with Training-free Augmented DynAmicsabstractDiffusion models have demonstrated exceptional capabilities in generating high-fidelity images but typically suffer from inefficient sampling.
Many solver designs and noise scheduling strategies have been proposed to dramatically improve sampling speeds. In this paper, we introduce a new sampling method that is up to 186\% faster than the current state of the art solver for comparative FID on ImageNet512. This new sampling method is training-free and uses an ordinary differential equation (ODE) solver.
The key to our method resides in using higher-dimensional initial noise, allowing to produce more detailed samples with less function evaluations from existing pretrained diffusion models. In addition, by design our solver allows to control the level of detail through a simple hyper-parameter at no extra computational cost. We present how our approach leverages momentum dynamics by establishing a fundamental equivalence between momentum diffusion models and conventional diffusion models with respect to their training paradigms. Moreover, we observe the use of higher-dimensional noise naturally exhibits characteristics similar to stochastic differential equations (SDEs). Finally, we demonstrate strong performances on a set of representative pretrained diffusion models, including EDM, EDM2, and Stable-Diffusion 3, which cover models in both pixel and latent spaces, as well as class and text conditional settings. The code is available at https://github.com/apple/ml-tada. Tianrong Chen, Huangjie Zheng, David Berthelot, Jiatao Gu, Joshua M. Susskind, Shuangfei Zhai |
NeurIPS | 4 |
| 2025 | STARFlow: Scaling Latent Normalizing Flows for High-resolution Image SynthesisabstractWe present STARFlow, a scalable generative model based on normalizing flows that achieves strong performance on high-resolution image synthesis.
STARFlow's main building block is Transformer Autoregressive Flow (TARFlow), which combines normalizing flows with Autoregressive Transformer architectures and has recently achieved impressive results in image modeling. In this work, we first establish the theoretical universality of TARFlow for modeling continuous distributions. Building on this foundation, we introduce a set of architectural and algorithmic innovations that significantly enhance the scalability: (1) a deep-shallow design where a deep Transformer block captures most of the model’s capacity, followed by a few shallow Transformer blocks that are computationally cheap yet contribute non-negligibly, (2) learning in the latent space of pretrained autoencoders, which proves far more effective than modeling pixels directly, and (3) a novel guidance algorithm that substantially improves sample quality. Crucially, our model remains a single, end-to-end normalizing flow, allowing exact maximum likelihood training in continuous space without discretization. STARFlow achieves competitive results in both class- and text-conditional image generation, with sample quality approaching that of state-of-the-art diffusion models. To our knowledge, this is the **first** successful demonstration of normalizing flows at this scale and resolution. Code and weights available at https://github.com/apple/ml-starflow. Jiatao Gu, Tianrong Chen, David Berthelot, Huangjie Zheng, Ruixiang Zhang, Laurent Dinh, Miguel Ángel Bautista 0001, Joshua M. Susskind, Shuangfei Zhai |
NeurIPS | 1 |
| 2025 | PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video GenerationabstractExisting video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for physics-grounded image-to-video generation with physical parameters and force control. At its core is a generative physics network that learns the distribution of physical dynamics across four materials (elastic, sand, plasticine, and rigid) via a diffusion model conditioned on physics parameters and applied forces. We represent physical dynamics as 3D point trajectories and train on a large-scale synthetic dataset of 550K animations generated by physics simulators. We enhance the diffusion model with a novel spatiotemporal attention block that emulates particle interactions and incorporates physics-based constraints during training to enforce physical plausibility. Experiments show that PhysCtrl generates realistic, physics-grounded motion trajectories which, when used to drive image-to-video models, yield high-fidelity, controllable videos that outperform existing methods in both visual quality and physical plausibility. Our code, model and data will be made publicly available upon publication. Chen Wang 0049, Chuhao Chen 0005, Yiming Huang 0011, Zhiyang Dou, Yuan Liu 0025, Jiatao Gu, Lingjie Liu |
NeurIPS | 6 |
| 2025 | Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsabstractAutoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and single-pass decoding, while central to their success, also inspires the exploration of a design space that could offer new axes of modeling flexibility.
In this work, we explore an alternative paradigm, shifting language modeling from a discrete token space to a continuous latent space.
We propose a novel framework that employs transformer-based autoregressive normalizing flows to model these continuous representations.
This approach unlocks substantial flexibility, enabling the construction of models that can capture global bi-directional context through stacked, alternating-direction autoregressive transformations, support block-wise generation with flexible token patch sizes, and facilitate a hierarchical multi-pass generation process.
We further propose new mixture-based coupling transformations designed to capture complex dependencies within the latent space shaped by discrete data, and demonstrate theoretical connections to conventional discrete autoregressive models.
Extensive experiments on language modeling benchmarks demonstrate strong likelihood performance and highlight the flexible modeling capabilities inherent in our framework. Ruixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang 0002, Huangjie Zheng, Tianrong Chen, Miguel Ángel Bautista 0001, Joshua M. Susskind, Navdeep Jaitly |
NeurIPS | 3 |
| 2025 | PhysHMR: Learning Humanoid Control Policies from Vision for Physically Plausible Human Motion ReconstructionabstractReconstructing physically plausible human motion from monocular videos remains a challenging problem in computer vision and graphics. Existing methods primarily focus on kinematics-based pose estimation, often leading to unrealistic results due to the lack of physical constraints. To address such artifacts, prior methods have typically relied on physics-based post-processing following the initial kinematics-based motion estimation. However, this two-stage design introduces error accumulation, ultimately limiting the overall reconstruction quality. In this paper, we present PhysHMR, a unified framework that directly learns a visual-to-action policy for humanoid control in a physics-based simulator, enabling motion reconstruction that is both physically grounded and visually aligned with the input video. A key component of our approach is the pixel-as-ray strategy, which lifts 2D keypoints into 3D spatial rays and transforms them into global space. These rays are incorporated as policy inputs, providing robust global pose guidance without depending on noisy 3D root predictions. This soft global grounding, combined with local visual features from a pretrained encoder, allows the policy to reason over both detailed pose and global positioning. To overcome the sample inefficiency of reinforcement learning, we further introduce a distillation scheme that transfers motion knowledge from a mocap-trained expert to the vision-conditioned policy, which is then refined using physically motivated reinforcement learning rewards. Extensive experiments demonstrate that PhysHMR produces high-fidelity, physically plausible motion across diverse scenarios, outperforming prior approaches in both visual accuracy and physical realism. Qiao Feng 0001, Yiming Huang 0011, Yufu Wang, Jiatao Gu, Lingjie Liu |
SIGGRAPH Asia | 4 |
| 2025 | Diffusion Models for 3D Generation: A SurveyabstractDenoising diffusion models have demonstrated tremendous success in modeling data distributions and synthesizing high-quality samples. In the 2D image domain, they have become the state-of-the-art and are capable of generating photo-realistic images with high controllability. More recently, researchers have begun to explore how to utilize diffusion models to generate 3D data, as doing so has more potential in real-world applications. This requires careful design choices in two key ways: identifying a suitable 3D representation and determining how to apply the diffusion process. In this survey, we provide the first comprehensive review of diffusion models for manipulating 3D content, including 3D generation, reconstruction, and 3D-aware image synthesis. We classify existing methods into three major categories: 2D space diffusion with pretrained models, 2D space diffusion without pretrained models, and 3D space diffusion. We also summarize popular datasets used for 3D generation with diffusion models. Along with this survey, we maintain a repository https://github.com/cwchenwang/awesome-3d-diffusion to track the latest relevant papers and codebases. Finally, we pose current challenges for diffusion models for 3D generation, and suggest future research directions. Chen Wang 0049, Hao-Yang Peng, Ying-Tian Liu, Jiatao Gu, Shi-Min Hu 0001 |
Comput. Vis. Media | 4 |
| 2025 | GECO : Fast Generative Image-to-3D Within One SECOndabstractRecent advancements in single-image 3D generation have produced two main categories of methods: reconstruction-based and generative methods. Reconstruction-based methods are efficient but lack uncertainty handling, leading to blurry artifacts in unseen regions. Generative approaches that based on score distillation (Poole et al. 2023), (Wang et al. 2024) are slow due to scene-specific optimization. Other methods, like InstantMesh (Xu et al. 2024), use a two-stage process - generating multi-view images with a diffusion model and then reconstructing 3D - which is inefficient due to multiple denoising steps of the diffusion model. To overcome these limitations, we introduce GECO, a feed-forward method for fast and high-quality single-image-to-3D generation within one second on a single GPU. Our approach resolves uncertainty and inefficiency issues through a two-stage distillation process. In the first stage, we distill a multi-step diffusion model (Shi et al. 2023) into a one-step model using score distillation for single-image-to-multi-view synthesis. To mitigate the synthesis quality degradation caused by the one-step model, we introduce a second distillation stage to learn to predict high-quality 3D from imperfect multi-view generated images by performing distillation directly on 3D representations. Experiments demonstrate that GECO offers significant speed improvements and comparable reconstruction quality compared to prior two-stage methods. Chen Wang 0049, Jiatao Gu, Xiaoxiao Long, Yuan Liu 0025, Lingjie Liu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Control3Diff: Learning Controllable 3D Diffusion Models from Single-view ImagesabstractDiffusion models have recently become the de-facto approach for generative modeling in the 2D domain. However, extending diffusion models to 3D is challenging, due to the difficulties in acquiring 3D ground truth data for training. On the other hand, 3D GANs that integrate implicit 3D representations into GANs have shown remarkable 3D-aware generation when trained only on single-view image datasets. However, 3D GANs do not provide straightforward ways to precisely control image synthesis. To address these challenges, We present Control3Diff, a 3D diffusion model that combines the strengths of diffusion models and 3D GANs for versatile controllable 3D-aware image synthesis for single-view datasets. Control3Diff explicitly models the underlying latent distribution (optionally conditioned on external inputs), thus enabling direct control during the diffusion process. Moreover, our approach is general and applicable to any types of controlling inputs, allowing us to train it with the same diffusion objective without any auxiliary supervision. We validate the efficacy of Control3Diff on standard image generation benchmarks including FFHQ, AFHQ, and ShapeNet, using various conditioning inputs such as images, sketches, and text prompts. Jiatao Gu, Qingzhe Gao, Shuangfei Zhai, Baoquan Chen, Lingjie Liu, Joshua M. Susskind |
3DV | 1 |
| 2024 | Diffusion Models Without AttentionabstractIn recent advancements in high-fidelity image generation, Denoising Diffusion Probabilistic Models (DDPMs) have emerged as a key player. However, their application at high resolutions presents significant computational challenges. Current methods, such as patchifying, expedite processes in UNet and Transformer architectures but at the expense of rep-resentational capacity. Addressing this, we introduce the Dif-fusion State Space Model (DIFFUSSM), an architecture that supplants attention mechanisms with a more scalable state space model backbone. This approach effectively handles higher resolutions without resorting to global compression, thus preserving detailed image representation throughout the diffusion process. Our focus on FLOP-efficient architectures in diffusion training marks a significant step forward. Comprehensive evaluations on both ImageNet and LSUN datasets at two resolutions demonstrate that DiffuSSMs are on par or even outperform existing diffusion models with attention modules in FID and Inception Score metrics while significantly reducing total FLOP usage. Jing Nathan Yan, Jiatao Gu, Alexander M. Rush |
CVPR | 2 |
| 2024 | Generative Modeling with Phase Stochastic BridgeabstractDiffusion models (DMs) represent state-of-the-art generative models for continuous inputs. DMs work by constructing a Stochastic Differential Equation (SDE) in the input space (ie, position space), and using a neural network to reverse it. In this work, we introduce a novel generative modeling framework grounded in \textbf{phase space dynamics}, where a phase space is defined as {an augmented space encompassing both position and velocity.} Leveraging insights from Stochastic Optimal Control, we construct a path measure in the phase space that enables efficient sampling. {In contrast to DMs, our framework demonstrates the capability to generate realistic data points at an early stage of dynamics propagation.} This early prediction sets the stage for efficient data generation by leveraging additional velocity information along the trajectory. On standard image generation benchmarks, our model yields favorable performance over baselines in the regime of small Number of Function Evaluations (NFEs). Furthermore, our approach rivals the performance of diffusion models equipped with efficient sampling techniques, underscoring its potential as a new tool generative modeling. Tianrong Chen, Jiatao Gu, Laurent Dinh, Evangelos A. Theodorou, Joshua M. Susskind, Shuangfei Zhai |
ICLR | 2 |
| 2024 | Matryoshka Diffusion ModelsabstractDiffusion models are the de-facto approach for generating high-quality images and videos, but learning high-dimensional models remains a formidable task due to computational and optimization challenges. Existing methods often resort to training cascaded models in pixel space, or using a downsampled latent space of a separately trained auto-encoder. In this paper, we introduce Matryoshka Diffusion (MDM), an end-to-end framework for high-resolution image and video synthesis. We propose a diffusion process that denoises inputs at multiple resolutions jointly and uses a NestedUNet architecture where features and parameters for small-scale inputs are nested within those of large scales. In addition, MDM enables a progressive training schedule from lower to higher resolutions, which leads to significant improvements in optimization for high-resolution generation. We demonstrate the effectiveness of our approach on various benchmarks, including class-conditioned image generation, high-resolution text-to-image, and text-to-video applications. Remarkably, we can train a single pixel-space model at resolutions of up to 1024x1024 pixels, demonstrating strong zero-shot generalization using the CC12M dataset, which contains only 12 million images. Code and pre-trained checkpoints are released at https://github.com/apple/ml-mdm. Jiatao Gu, Shuangfei Zhai, Yizhe Zhang 0002, Joshua M. Susskind, Navdeep Jaitly |
ICLR | 1 |
| 2024 | Data-free Distillation of Diffusion Models with BootstrappingabstractDiffusion models have demonstrated great potential for generating diverse images. However, their performance often suffers from slow generation due to iterative denoising. Knowledge distillation has been recently proposed as a remedy which can reduce the number of inference steps to one or a few, without significant quality degradation. However, existing distillation methods either require significant amounts of offline computation for generating synthetic training data from the teacher model, or need to perform expensive online learning with the help of real data. In this work, we present a novel technique called BOOT, that overcomes these limitations with an efficient data-free distillation algorithm. The core idea is to learn a time-conditioned model that predicts the output of a pre-trained diffusion model teacher given any time-step. Such a model can be efficiently trained based on bootstrapping from two consecutive sampled steps. Furthermore, our method can be easily adapted to large-scale text-to-image diffusion models, which are challenging for previous methods given the fact that the training sets are often large and difficult to access. We demonstrate the effectiveness of our approach on several benchmark datasets in the DDIM setting, achieving comparable generation quality while being orders of magnitude faster than the diffusion teacher. The text-to-image results show that the proposed approach is able to handle highly complex distributions, shedding light on more efficient generative modeling. Jiatao Gu, Chen Wang 0049, Shuangfei Zhai, Yizhe Zhang 0002, Lingjie Liu, Joshua M. Susskind |
ICML | 1 |
| 2024 | Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent ModelingabstractDiffusion models have emerged as a powerful tool for generating high-quality images from textual descriptions. Despite their successes, these models often exhibit limited diversity in the sampled images, particularly when sampling with a high classifier-free guidance weight. To address this issue, we present Kaleido, a novel approach that enhances the diversity of samples by incorporating autoregressive latent priors. Kaleido integrates an autoregressive language model that encodes the original caption and generates latent variables, serving as abstract and intermediary representations for guiding and facilitating the image generation process.
In this paper, we explore a variety of discrete latent representations, including textual descriptions, detection bounding boxes, object blobs, and visual tokens. These representations diversify and enrich the input conditions to the diffusion models, enabling more diverse outputs.
Our experimental results demonstrate that Kaleido effectively broadens the diversity of the generated image samples from a given textual description while maintaining high image quality. Furthermore, we show that Kaleido adheres closely to the guidance provided by the generated latent variables, demonstrating its capability to effectively control and direct the image generation process. Jiatao Gu, Ying Shen 0006, Shuangfei Zhai, Yizhe Zhang 0002, Navdeep Jaitly, Joshua M. Susskind |
NeurIPS | 1 |
| 2024 | ProLiF: Progressively-connected Light Field network for efficient view synthesis
Peng Wang 0099, Yuan Liu 0025, Guying Lin, Jiatao Gu, Lingjie Liu, Taku Komura, Wenping Wang 0001 |
Comput. Graph. | 4 |
| 2023 | Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstructionabstract3D-aware image synthesis encompasses a variety of tasks, such as scene generation and novel view synthesis from images. Despite numerous task-specific methods, developing a comprehensive model remains challenging. In this paper, we present SSDNeRF, a unified approach that employs an expressive diffusion model to learn a generalizable prior of neural radiance fields (NeRF) from multi-view images of diverse objects. Previous studies have used two-stage approaches that rely on pretrained NeRFs as real data to train diffusion models. In contrast, we propose a new single-stage training paradigm with an end-to-end objective that jointly optimizes a NeRF auto-decoder and a latent diffusion model, enabling simultaneous 3D reconstruction and prior learning, even from sparsely available views. At test time, we can directly sample the diffusion prior for unconditional generation, or combine it with arbitrary observations of unseen objects for NeRF reconstruction. SSDNeRF demonstrates robust results comparable to or better than leading task-specific methods in unconditional generation and single/sparse-view 3D reconstruction.6 Hansheng Chen 0001, Jiatao Gu, Anpei Chen, Wei Tian 0001, Zhuowen Tu, Lingjie Liu, Hao Su 0001 |
ICCV | 2 |
| 2023 | MAST: Masked Augmentation Subspace Training for Generalizable Self-Supervised Priors
Chen Huang 0001, Hanlin Goh, Jiatao Gu, Joshua M. Susskind |
ICLR | 3 |
| 2023 | f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang 0002, Miguel Ángel Bautista 0001, Joshua M. Susskind |
ICLR | 1 |
| 2023 | Diffusion Probabilistic Fields
Peiye Zhuang, Samira Abnar, Jiatao Gu, Alexander G. Schwing, Joshua M. Susskind, Miguel Ángel Bautista 0001 |
ICLR | 3 |
| 2023 | NerfDiff: Single-image View Synthesis with NeRF-guided Distillation from 3D-aware DiffusionabstractNovel view synthesis from a single image requires inferring occluded regions of objects and scenes whilst simultaneously maintaining semantic and physical consistency with the input. Existing approaches condition neural radiance fields (NeRF) on local image features, projecting points to the input image plane, and aggregating 2D features to perform volume rendering. However, under severe occlusion, this projection fails to resolve uncertainty, resulting in blurry renderings that lack details. In this work, we propose NerfDiff, which addresses this issue by distilling the knowledge of a 3D-aware conditional diffusion model (CDM) into NeRF through synthesizing and refining a set of virtual views at test-time. We further propose a novel NeRF-guided distillation algorithm that simultaneously generates 3D consistent virtual views from the CDM samples, and finetunes the NeRF based on the improved virtual views. Our approach significantly outperforms existing NeRF-based and geometry-free approaches on challenging datasets including ShapeNet, ABO, and Clevr3D. Jiatao Gu, Alex Trevithick, Kai-En Lin, Joshua M. Susskind, Christian Theobalt, Lingjie Liu, Ravi Ramamoorthi |
ICML | 1 |
| 2023 | Stabilizing Transformer Training by Preventing Attention Entropy CollapseabstractTraining stability is of great importance to Transformers. In this work, we investigate the training dynamics of Transformers by examining the evolution of the attention layers. In particular, we track the attention entropy for each attention head during the course of training, which is a proxy for model sharpness. We identify a common pattern across different architectures and tasks, where low attention entropy is accompanied by high training instability, which can take the form of oscillating loss or divergence. We denote the pathologically low attention entropy, corresponding to highly concentrated attention scores, as $\textit{entropy collapse}$. As a remedy, we propose $\sigma$Reparam, a simple and efficient solution where we reparametrize all linear layers with spectral normalization and an additional learned scalar. We demonstrate that $\sigma$Reparam successfully prevents entropy collapse in the attention layers, promoting more stable training. Additionally, we prove a tight lower bound of the attention entropy, which decreases exponentially fast with the spectral norm of the attention logits, providing additional motivation for our approach. We conduct experiments with $\sigma$Reparam on image classification, image self-supervised learning, machine translation, speech recognition, and language modeling tasks. We show that $\sigma$Reparam provides stability and robustness with respect to the choice of hyperparameters, going so far as enabling training (a) a Vision Transformer to competitive performance without warmup, weight decay, layer normalization or adaptive optimizers; (b) deep architectures in machine translation and (c) speech recognition to competitive performance without warmup and adaptive optimizers. Code is available at https://github.com/apple/ml-sigma-reparam. Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang 0002, Jiatao Gu, Joshua M. Susskind |
ICML | 7 |
| 2023 | PLANNER: Generating Diversified Paragraph via Latent Language Diffusion ModelabstractAutoregressive models for text sometimes generate repetitive and low-quality output because errors accumulate during the steps of generation. This issue is often attributed to exposure bias -- the difference between how a model is trained, and how it is used during inference. Denoising diffusion models provide an alternative approach in which a model can revisit and revise its output. However, they can be computationally expensive and prior efforts on text have led to models that produce less fluent output compared to autoregressive models, especially for longer text and paragraphs. In this paper, we propose PLANNER, a model that combines latent semantic diffusion with autoregressive generation, to generate fluent text while exercising global control over paragraphs. The model achieves this by combining an autoregressive "decoding" module with a "planning" module that uses latent diffusion to generate semantic paragraph embeddings in a coarse-to-fine manner. The proposed method is evaluated on various conditional generation tasks, and results on semantic generation, text completion and summarization show its effectiveness in generating high-quality long-form text in an efficient manner. Yizhe Zhang 0002, Jiatao Gu, Zhuofeng Wu 0001, Shuangfei Zhai, Joshua M. Susskind, Navdeep Jaitly |
NeurIPS | 2 |
| 2022 | Direct Speech-to-Speech Translation With Discrete UnitsabstractAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, Wei-Ning Hsu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ann Lee 0001, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Yun Tang 0002, Juan Pino 0001, Wei-Ning Hsu |
ACL (1) | 4 |
| 2022 | Unified Speech-Text Pre-training for Speech Translation and RecognitionabstractYun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, Juan Pino. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yun Tang 0002, Hongyu Gong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li 0003, Abdel-rahman Mohamed, Michael Auli, Juan Pino 0001 |
ACL (1) | 6 |
| 2022 | StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image Synthesis
Jiatao Gu, Lingjie Liu, Peng Wang 0099, Christian Theobalt |
ICLR | 1 |
| 2022 | data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageabstractWhile the general idea of self-supervised learning is identical across modalities, the actual algorithms and objectives differ widely because they were developed with a single modality in mind. To get us closer to general self-supervised learning, we present data2vec, a framework that uses the same learning method for either speech, NLP or computer vision. The core idea is to predict latent representations of the full input data based on a masked view of the input in a self-distillation setup using a standard Transformer architecture. Instead of predicting modality-specific targets such as words, visual tokens or units of human speech which are local in nature, data2vec predicts contextualized latent representations that contain information from the entire input. Experiments on the major benchmarks of speech recognition, image classification, and natural language understanding demonstrate a new state of the art or competitive performance to predominant approaches. Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, Michael Auli |
ICML | 5 |
| 2022 | Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino 0001, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, Ann Lee 0001 |
INTERSPEECH | 6 |
| 2022 | IDPG: An Instance-Dependent Prompt Generation MethodabstractZhuofeng Wu, Sinong Wang, Jiatao Gu, Rui Hou, Yuxiao Dong, V.G.Vinod Vydiswaran, Hao Ma. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zhuofeng Wu 0001, Sinong Wang, Jiatao Gu, Yuxiao Dong, V. G. Vinod Vydiswaran, Hao Ma 0001 |
NAACL-HLT | 3 |
| 2022 | Textless Speech-to-Speech Translation on Real DataabstractAnn Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, Wei-Ning Hsu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ann Lee 0001, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino 0001, Jiatao Gu, Wei-Ning Hsu |
NAACL-HLT | 10 |
| 2021 | Multilingual Neural Machine Translation with Deep Encoder and Multiple Shallow DecodersabstractXiang Kong, Adithya Renduchintala, James Cross, Yuqing Tang, Jiatao Gu, Xian Li. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Xiang Kong, Adithya Renduchintala, James Cross 0003, Jiatao Gu, Xian Li 0003 |
EACL | 5 |
| 2021 | The Source-Target Domain Mismatch Problem in Machine TranslationabstractJiajun Shen, Peng-Jen Chen, Matthew Le, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc’Aurelio Ranzato. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Peng-Jen Chen, Matt Le 0001, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc'Aurelio Ranzato |
EACL | 5 |
| 2021 | CoVoST 2 and Massively Multilingual Speech Translation
Changhan Wang, Anne Wu, Jiatao Gu, Juan Pino 0001 |
Interspeech | 3 |
| 2021 | Volume Rendering of Neural Implicit SurfacesabstractNeural volume rendering became increasingly popular recently due to its success in synthesizing novel views of a scene from a sparse set of input images. So far, the geometry learned by neural volume rendering techniques was modeled using a generic density function. Furthermore, the geometry itself was extracted using an arbitrary level set of the density function leading to a noisy, often low fidelity reconstruction.The goal of this paper is to improve geometry representation and reconstruction in neural volume rendering. We achieve that by modeling the volume density as a function of the geometry. This is in contrast to previous work modeling the geometry as a function of the volume density. In more detail, we define the volume density function as Laplace's cumulative distribution function (CDF) applied to a signed distance function (SDF) representation. This simple density representation has three benefits: (i) it provides a useful inductive bias to the geometry learned in the neural volume rendering process; (ii) it facilitates a bound on the opacity approximation error, leading to an accurate sampling of the viewing ray. Accurate sampling is important to provide a precise coupling of geometry and radiance; and (iii) it allows efficient unsupervised disentanglement of shape and appearance in volume rendering.Applying this new density representation to challenging scene multiview datasets produced high quality geometry reconstructions, outperforming relevant baselines. Furthermore, switching shape and appearance between scenes is possible due to the disentanglement of the two. Lior Yariv, Jiatao Gu, Yoni Kasten, Yaron Lipman |
NeurIPS | 2 |
| 2021 | Neural actor: neural free-view synthesis of human actors with pose controlabstractWe propose Neural Actor (NA), a new method for high-quality synthesis of humans from arbitrary viewpoints and under arbitrary controllable poses. Our method is developed upon recent neural scene representation and rendering works which learn representations of geometry and appearance from only 2D images. While existing works demonstrated compelling rendering of static scenes and playback of dynamic scenes, photo-realistic reconstruction and rendering of humans with neural implicit methods, in particular under user-controlled novel poses, is still difficult. To address this problem, we utilize a coarse body model as a proxy to unwarp the surrounding 3D space into a canonical pose. A neural radiance field learns pose-dependent geometric deformations and pose- and view-dependent appearance effects in the canonical space from multi-view video input. To synthesize novel views of high-fidelity dynamic geometry and appearance, NA leverages 2D texture maps defined on the body model as latent variables for predicting residual deformations and the dynamic appearance. Experiments demonstrate that our method achieves better quality than the state-of-the-arts on playback as well as novel pose synthesis, and can even generalize well to new poses that starkly differ from the training poses. Furthermore, our method also supports shape control on the free-view synthesis of human actors. Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, Christian Theobalt |
ACM Trans. Graph. | 5 |
| 2020 | Neural Machine Translation with Byte-Level SubwordsabstractAlmost all existing machine translation models are built on top of character-based vocabularies: characters, subwords or words. Rare characters from noisy text or character-rich languages such as Japanese and Chinese however can unnecessarily take up vocabulary slots and limit its compactness. Representing text at the level of bytes and using the 256 byte set as vocabulary is a potential solution to this issue. High computational cost has however prevented it from being widely deployed or used in practice. In this paper, we investigate byte-level subwords, specifically byte-level BPE (BBPE), which is compacter than character vocabulary and has no out-of-vocabulary tokens, but is more efficient than using pure bytes only is. We claim that contextualizing BBPE embeddings is necessary, which can be implemented by a convolutional or recurrent layer. Our experiments show that BBPE has comparable performance to BPE while its size is only 1/8 of that for BPE. In the multilingual setting, BBPE maximizes vocabulary sharing across many languages and achieves better translation quality. Moreover, we show that BBPE enables transferring models between languages with non-overlapping character sets. Changhan Wang, Kyunghyun Cho, Jiatao Gu |
AAAI | 3 |
| 2020 | Addressing Posterior Collapse with Mutual Information for Improved Variational Neural Machine TranslationabstractThis paper proposes a simple and effective approach to address the problem of posterior collapse in conditional variational autoencoders (CVAEs).It thus improves performance of machine translation models that use noisy or monolingual data, as well as in conventional settings.Extending Transformer and conditional VAEs, our proposed latent variable model measurably prevents posterior collapse by (1) using a modified evidence lower bound (ELBO) objective which promotes mutual information between the latent variable and the target, and (2) guiding the latent variable with an auxiliary bag-of-words prediction task.As a result, the proposed model yields improved translation quality compared to existing variational NMT models on WMT Ro↔En and De↔En.With latent variables being effectively utilized, our model demonstrates improved robustness over non-latent Transformer in handling uncertainty: exploiting noisy source-side monolingual data (up to +3.2 BLEU), and training with weakly aligned web-mined parallel data (up to +4.7 BLEU). Arya McCarthy, Xian Li 0003, Jiatao Gu |
ACL | 3 |
| 2020 | Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech TranslationabstractWe introduce dual-decoder Transformer, a new model architecture that jointly performs automatic speech recognition (ASR) and multilingual speech translation (ST).Our models are based on the original Transformer architecture (Vaswani et al., 2017) but consist of two decoders, each responsible for one task (ASR or ST).Our major contribution lies in how these decoders interact with each other: one decoder can attend to different information sources from the other via a dual-attention mechanism.We propose two variants of these architectures corresponding to two different levels of dependencies between the decoders, called the parallel and cross dual-decoder Transformers, respectively.Extensive experiments on the MuST-C dataset show that our models outperform the previously-reported highest translation performance in the multilingual settings, and outperform as well bilingual one-to-one results.Furthermore, our parallel models demonstrate no trade-off between ASR and ST compared to the vanilla multi-task architecture.Our code and pre-trained models are available at https:// Hang Le 0001, Juan Pino 0001, Changhan Wang, Jiatao Gu, Didier Schwab, Laurent Besacier |
COLING | 4 |
| 2020 | PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding
Saining Xie, Jiatao Gu, Demi Guo, Charles R. Qi, Leonidas J. Guibas, Or Litany |
ECCV (3) | 2 |
| 2020 | Depth-Adaptive Transformer
Maha Elbayad, Jiatao Gu, Edouard Grave, Michael Auli |
ICLR | 2 |
| 2020 | Revisiting Self-Training for Neural Sequence Generation
Junxian He, Jiatao Gu, Marc'Aurelio Ranzato |
ICLR | 2 |
| 2020 | Monotonic Multihead Attention
Xutai Ma, Juan Pino 0001, James Cross 0003, Liezl Puzon, Jiatao Gu |
ICLR | 5 |
| 2020 | Understanding Knowledge Distillation in Non-autoregressive Machine Translation
Chunting Zhou, Jiatao Gu, Graham Neubig |
ICLR | 2 |
| 2020 | Non-autoregressive Machine Translation with Disentangled Context TransformerabstractState-of-the-art neural machine translation models generate a translation from left to right and every step is conditioned on the previously generated tokens. The sequential nature of this generation process causes fundamental latency in inference since we cannot generate multiple tokens in each sentence in parallel. We propose an attention-masking based model, called Disentangled Context (DisCo) transformer, that simultaneously generates all tokens given different contexts. The DisCo transformer is trained to predict every output token given an arbitrary subset of the other reference tokens. We also develop the parallel easy-first inference algorithm, which iteratively refines every token in parallel and reduces the number of required iterations. Our extensive experiments on 7 translation directions with varying data sizes demonstrate that our model achieves competitive, if not better, performance compared to the state of the art in non-autoregressive machine translation while significantly reducing decoding time on average. Jungo Kasai, James Cross 0003, Marjan Ghazvininejad, Jiatao Gu |
ICML | 4 |
| 2020 | Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech TranslationabstractTransfer learning from high-resource languages is known to be an efficient way to improve end-to-end automatic speech recognition (ASR) for low-resource languages.Pre-trained or jointly trained encoder-decoder models, however, do not share the language modeling (decoder) for the same language, which is likely to be inefficient for distant target languages.We introduce speech-to-text translation (ST) as an auxiliary task to incorporate additional knowledge of the target language and enable transferring from that target language.Specifically, we first translate high-resource ASR transcripts into a target lowresource language, with which a ST model is trained.Both ST and target ASR share the same attention-based encoderdecoder architecture and vocabulary.The former task then provides a fully pre-trained model for the latter, bringing up to 24.6% word error rate (WER) reduction to the baseline (direct transfer from high-resource ASR).We show that training ST with human translations is not necessary.ST trained with machine translation (MT) pseudo-labels brings consistent gains.It can even outperform those using human labels when transferred to target ASR by leveraging only 500K MT examples.Even with pseudo-labels from low-resource MT (200K examples), ST-enhanced transfer brings up to 8.9% WER reduction to direct transfer. Changhan Wang, Juan Pino 0001, Jiatao Gu |
INTERSPEECH | 3 |
| 2020 | Self-Supervised Representations Improve End-to-End Speech TranslationabstractEnd-to-end speech-to-text translation can provide a simpler and smaller system but is facing the challenge of data scarcity.Pre-training methods can leverage unlabeled data and have been shown to be effective on data-scarce settings.In this work, we explore whether self-supervised pre-trained speech representations can benefit the speech translation task in both highand low-resource settings, whether they can transfer well to other languages, and whether they can be effectively combined with other common methods that help improve low-resource end-to-end speech translation such as using a pre-trained highresource speech recognition system.We demonstrate that selfsupervised pre-trained features can consistently improve the translation performance, and cross-lingual transfer allows to extend to a variety of languages without or with little tuning. Anne Wu, Changhan Wang, Juan Pino 0001, Jiatao Gu |
INTERSPEECH | 4 |
| 2020 | CoVoST: A Diverse Multilingual Speech-To-Text Translation CorpusabstractSpoken language translation has recently witnessed a resurgence in popularity, thanks to the development of end-to-end models and the creation of new corpora, such as Augmented LibriSpeech and MuST-C. Existing datasets involve language pairs with English as a source language, involve very specific domains or are low resource. We introduce CoVoST, a multilingual speech-to-text translation corpus from 11 languages into English, diversified with over 11,000 speakers and over 60 accents. We describe the dataset creation methodology and provide empirical evidence of the quality of the data. We also provide initial benchmarks, including, to our knowledge, the first end-to-end many-to-one multilingual models for spoken language translation. CoVoST is released under CC0 license and free to use. We also provide additional evaluation data derived from Tatoeba under CC licenses. Changhan Wang, Juan Pino 0001, Anne Wu, Jiatao Gu |
LREC | 4 |
| 2020 | Neural Sparse Voxel FieldsabstractPhoto-realistic free-viewpoint rendering of real-world scenes using classical computer graphics techniques is challenging, because it requires the difficult step of capturing detailed appearance and geometry models. Recent studies have demonstrated promising results by learning scene representations that implicitly encodes both geometry and appearance without 3D supervision. However, existing approaches in practice often show blurry renderings caused by the limited network capacity or the difficulty in finding accurate intersections of camera rays with the scene geometry. Synthesizing high-resolution imagery from these representations often requires time-consuming optical ray marching. In this work, we introduce Neural Sparse Voxel Fields (NSVF), a new neural scene representation for fast and high-quality free-viewpoint rendering. The NSVF defines a series of voxel-bounded implicit fields organized in a sparse voxel octree to model local properties in each cell. We progressively learn the underlying voxel structures with a differentiable ray-marching operation from only a set of posed RGB images. With the sparse voxel octree structure, rendering novel views at inference time can be accelerated by skipping the voxels without relevant scene content. Our method is over 10 times faster than the state-of-the-art while achieving higher quality results. Furthermore, by utilizing an explicit sparse voxel representation, our method can be easily applied to scene editing and scene composition. we also demonstrate various kinds of challenging tasks, including multi-object learning, free-viewpoint rendering of a moving human, and large-scale scene rendering. Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, Christian Theobalt |
NeurIPS | 2 |
| 2020 | Cross-lingual Retrieval for Iterative Self-Supervised TrainingabstractRecent studies have demonstrated the cross-lingual alignment ability of multilingual pretrained language models. In this work, we found that the cross-lingual alignment can be further improved by training seq2seq models on sentence pairs mined using their own encoder outputs. We utilized these findings to develop a new approach --- cross-lingual retrieval for iterative self-supervised training (CRISS), where mining and training processes are applied iteratively, improving cross-lingual alignment and translation ability at the same time. Using this method, we achieved state-of-the-art unsupervised machine translation results on 9 language directions with an average improvement of 2.4 BLEU, and on the Tatoeba sentence retrieval task in the XTREME benchmark on 16 languages with an average improvement of 21.5% in absolute accuracy. Furthermore, CRISS also brings an additional 1.8 BLEU improvement on average compared to mBART, when finetuned on supervised machine translation downstream tasks. Chau Tran, Xian Li 0003, Jiatao Gu |
NeurIPS | 4 |
| 2020 | Multilingual Denoising Pre-training for Neural Machine TranslationabstractThis paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART—a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using the BART objective (Lewis et al., 2019 ). mBART is the first method for pre-training a complete sequence-to-sequence model by denoising full texts in multiple languages, whereas previous approaches have focused only on the encoder, decoder, or reconstructing parts of the text. Pre-training a complete model allows it to be directly fine-tuned for supervised (both sentence-level and document-level) and unsupervised machine translation, with no task- specific modifications. We demonstrate that adding mBART initialization produces performance gains in all but the highest-resource settings, including up to 12 BLEU points for low resource MT and over 5 BLEU points for many document-level and unsupervised models. We also show that it enables transfer to language pairs with no bi-text or that were not in the pre-training corpus, and present extensive analysis of which factors contribute the most to effective pre-training. 1 Yinhan Liu, Jiatao Gu, Naman Goyal 0001, Xian Li 0003, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, Luke Zettlemoyer |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | Improved Zero-shot Neural Machine Translation via Ignoring Spurious CorrelationsabstractZero-shot translation, translating between language pairs on which a Neural Machine Translation (NMT) system has never been trained, is an emergent property when training the system in multilingual settings.However, naïve training for zero-shot NMT easily fails, and is sensitive to hyper-parameter setting.The performance typically lags far behind the more conventional pivot-based approach which translates twice using a third language as a pivot.In this work, we address the degeneracy problem due to capturing spurious correlations by quantitatively analyzing the mutual information between language IDs of the source and decoded sentences.Inspired by this analysis, we propose to use two simple but effective approaches: (1) decoder pre-training; (2) backtranslation.These methods show significant improvement (4 ∼ 22 BLEU points) over the vanilla zero-shot translation on three challenging multilingual datasets, and achieve similar or better results than the pivot-based approach. Jiatao Gu, Yong Wang 0032, Kyunghyun Cho, Victor O. K. Li |
ACL (1) | 1 |
| 2019 | Levenshtein TransformerabstractModern neural sequence generation models are built to either generate tokens step-by-step from scratch or (iteratively) modify a sequence of tokens bounded by a fixed length. In this work, we develop Levenshtein Transformer, a new partially autoregressive model devised for more flexible and amenable sequence generation. Unlike previous approaches, the basic operations of our model are insertion and deletion. The combination of them facilitates not only generation but also sequence refinement allowing dynamic length changes. We also propose a set of new training techniques dedicated at them, effectively exploiting one as the other's learning signal thanks to their complementary nature. Experiments applying the proposed model achieve comparable or even better performance with much-improved efficiency on both generation (e.g. machine translation, text summarization) and refinement tasks (e.g. automatic post-editing). We further confirm the flexibility of our model by showing a Levenshtein Transformer trained by machine translation can straightforwardly be used for automatic post-editing. Jiatao Gu, Changhan Wang |
NeurIPS | 1 |
| 2019 | Insertion-based Decoding with Automatically Inferred Generation OrderabstractConventional neural autoregressive decoding commonly assumes a fixed left-to-right generation order, which may be sub-optimal. In this work, we propose a novel decoding algorithm— InDIGO—which supports flexible sequence generation in arbitrary orders through insertion operations. We extend Transformer, a state-of-the-art sequence generation model, to efficiently implement the proposed approach, enabling it to be trained with either a pre-defined generation order or adaptive orders obtained from beam-search. Experiments on four real-world tasks, including word order recovery, machine translation, image caption, and code generation, demonstrate that our algorithm can generate sequences following arbitrary orders, while achieving competitive or even better performance compared with the conventional left-to-right generation. The generated sequences show that InDIGO adopts adaptive generation orders based on input information. Jiatao Gu, Qi Liu 0049, Kyunghyun Cho |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Real-Time Traffic Speed Estimation With Graph Convolutional Generative AutoencoderabstractReal-time traffic speed estimation is an essential component of intelligent transportation system (ITS) technologies. It is the foundation of modern transportation control and management applications. However, the existing traffic speed acquisition systems can only provide real-time speed measurements of a small number of roads with stationary speed sensors and crowdsourcing vehicles. How to utilize this information to provide traffic speed maps for transportation networks is becoming a key problem in ITSs. In this paper, we present a novel deep-learning model called graph convolutional generative autoencoder to fully address the real-time traffic speed estimation problem. The proposed model incorporates the recent development in deep-learning techniques to extract the spatial correlation of the transportation network from the input incomplete historical data. To evaluate the proposed speed estimation technique, we conduct comprehensive case studies on a real-world transportation network and vehicular traces. The simulation results demonstrate that the proposed technique can notably outperform existing traffic speed estimation and deep-learning techniques. In addition, the impact of dataset properties and control parameters is investigated. James Jian Qiao Yu, Jiatao Gu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Online Vehicle Routing With Neural Combinatorial Optimization and Deep Reinforcement LearningabstractOnline vehicle routing is an important task of the modern transportation service provider. Contributed by the ever-increasing real-time demand on the transportation system, especially small-parcel last-mile delivery requests, vehicle route generation is becoming more computationally complex than before. The existing routing algorithms are mostly based on mathematical programming, which requires huge computation time in city-size transportation networks. To develop routes with minimal time, in this paper, we propose a novel deep reinforcement learning-based neural combinatorial optimization strategy. Specifically, we transform the online routing problem to a vehicle tour generation problem, and propose a structural graph embedded pointer network to develop these tours iteratively. Furthermore, since constructing supervised training data for the neural network is impractical due to the high computation complexity, we propose a deep reinforcement learning mechanism with an unsupervised auxiliary network to train the model parameters. A multisampling scheme is also devised to further improve the system performance. Since the parameter training process is offline, the proposed strategy can achieve a superior online route generation speed. To assess the proposed strategy, we conduct comprehensive case studies with a real-world transportation network. The simulation results show that the proposed strategy can significantly outperform conventional strategies with limited computation time in both static and dynamic logistic systems. In addition, the influence of control parameters on the system performance is investigated. James Jian Qiao Yu, Wen Yu 0001, Jiatao Gu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Neural Machine Translation with Gumbel-Greedy DecodingabstractPrevious neural machine translation models used some heuristic search algorithms (e.g., beam search) in order to avoid solving the maximum a posteriori problem over translation sentences at test phase. In this paper, we propose the \textit{Gumbel-Greedy Decoding} which trains a generative network to predict translation under a trained model. We solve such a problem using the Gumbel-Softmax reparameterization, which makes our generative network differentiable and trainable through standard stochastic gradient methods. We empirically demonstrate that our proposed model is effective for generating sequences of discrete words. Jiatao Gu, Daniel Jiwoong Im, Victor O. K. Li |
AAAI | 1 |
| 2018 | Search Engine Guided Neural Machine TranslationabstractIn this paper, we extend an attention-based neural machine translation (NMT) model by allowing it to access an entire training set of parallel sentence pairs even after training. The proposed approach consists of two stages. In the first stage –retrieval stage–, an off-the-shelf, black-box search engine is used to retrieve a small subset of sentence pairs from a training set given a source sentence. These pairs are further filtered based on a fuzzy matching score based on edit distance. In the second stage–translation stage–, a novel translation model, called search engine guided NMT (SEG-NMT), seamlessly uses both the source sentence and a set of retrieved sentence pairs to perform the translation. Empirical evaluation on three language pairs (En-Fr, En-De, and En-Es) shows that the proposed approach significantly outperforms the baseline approach and the improvement is more significant when more relevant sentence pairs were retrieved. Jiatao Gu, Yong Wang 0032, Kyunghyun Cho, Victor O. K. Li |
AAAI | 1 |
| 2018 | Meta-Learning for Low-Resource Neural Machine TranslationabstractIn this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm (MAML, Finn et al., 2017) for lowresource neural machine translation (NMT).We frame low-resource translation as a metalearning problem, and we learn to adapt to low-resource languages based on multilingual high-resource language tasks.We use the universal lexical representation (Gu et al., 2018b) to overcome the input-output mismatch across different languages.We evaluate the proposed meta-learning strategy using eighteen European languages (Bg, Cs, Da, De, El, Es, Et, Fr, Hu, It, Lt, Nl, Pl, Pt, Sk, Sl, Sv and Ru) as source tasks and five diverse languages (Ro, Lv, Fi, Tr and Ko) as target tasks.We show that the proposed approach significantly outperforms the multilingual, transfer learning based approach (Zoph et al., 2016) and enables us to train a competitive NMT system with only a fraction of training examples.For instance, the proposed approach can achieve as high as 22.04 BLEU on Romanian-English WMT'16 by seeing only 16,000 translated words (⇠ 600 parallel sentences). Jiatao Gu, Yong Wang 0032, Yun Chen 0007, Victor O. K. Li, Kyunghyun Cho |
EMNLP | 1 |
| 2018 | Non-Autoregressive Neural Machine Translation
Jiatao Gu, James Bradbury 0002, Caiming Xiong, Victor O. K. Li, Richard Socher |
ICLR (Poster) | 1 |
| 2018 | Universal Neural Machine Translation for Extremely Low Resource LanguagesabstractJiatao Gu, Hany Hassan, Jacob Devlin, Victor O.K. Li. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jiatao Gu, Hany Hassan, Jacob Devlin, Victor O. K. Li |
NAACL-HLT | 1 |
| 2017 | Learning to Translate in Real-time with Neural Machine TranslationabstractTranslating in real-time, a.k.a.simultaneous translation, outputs translation words before the input sentence ends, which is a challenging problem for conventional machine translation methods.We propose a neural machine translation (NMT) framework for simultaneous translation in which an agent learns to make decisions on when to translate from the interaction with a pre-trained NMT environment.To trade off quality and delay, we extensively explore various targets for delay and design a method for beam-search applicable in the simultaneous MT setting.Experiments against state-of-the-art baselines on two language pairs demonstrate the efficacy of the proposed framework both quantitatively and qualitatively. 1 Jiatao Gu, Graham Neubig, Kyunghyun Cho, Victor O. K. Li |
EACL (1) | 1 |
| 2017 | Trainable Greedy Decoding for Neural Machine TranslationabstractRecent research in neural machine translation has largely focused on two aspects; neural network architectures and end-toend learning algorithms.The problem of decoding, however, has received relatively little attention from the research community.In this paper, we solely focus on the problem of decoding given a trained neural machine translation model.Instead of trying to build a new decoding algorithm for any specific decoding objective, we propose the idea of trainable decoding algorithm in which we train a decoding algorithm to find a translation that maximizes an arbitrary decoding objective.More specifically, we design an actor that observes and manipulates the hidden state of the neural machine translation decoder and propose to train it using a variant of deterministic policy gradient.We extensively evaluate the proposed algorithm using four language pairs and two decoding objectives, and show that we can indeed train a trainable greedy decoder that generates a better translation (in terms of a target decoding objective) with minimal computational overhead. Jiatao Gu, Kyunghyun Cho, Victor O. K. Li |
EMNLP | 1 |
| 2017 | Deep Learning Model to Estimate Air Pollution Using M-BP to Fill in Missing Proxy Urban DataabstractAir quality has deteriorated rapidly in Hong Kong and China in the past two decades, with NO2and PM2.5levels frequently exceeding WHO safety guidelines. While poor air quality has clear public health impacts, there are very limited air quality monitoring (AQM) stations, severely constraining evidence-based air quality decision-making, leading to severe criticisms about the utility of the current official Air Quality Health Index to the public. Since air pollution is highly location-dependent, a city-wide deployment of traditional, highly sophisticated air quality monitors would be prohibitively expensive. In this paper, we propose a deep learning model to estimate air pollution throughout the city, utilizing the readily available urban data as proxy data. As with many big data driven approaches, the proxy data may be sparse/missing. We propose the M-BP algorithm to recover/fill in such missing data. Our results show that the proposed model gives better estimates compared with existing big data approaches. Victor O. K. Li, Jacqueline C. K. Lam, Yun Chen 0007, Jiatao Gu |
GLOBECOM | 4 |
| 2016 | Incorporating Copying Mechanism in Sequence-to-Sequence LearningabstractWe address an important problem in sequence-to-sequence (Seq2Seq) learning referred to as copying, in which certain segments in the input sequence are selectively replicated in the output sequence. A similar phenomenon is observable in human language communication. For example, humans tend to repeat entity names or even long phrases in conversation. The challenge with regard to copying in Seq2Seq is that new machinery is needed to decide when to perform the operation. In this paper, we incorporate copying into neural network-based Seq2Seq learning and propose a new model called CopyNet with encoder-decoder structure. CopyNet can nicely integrate the regular way of word generation in the decoder with the new copying mechanism which can choose sub-sequences in the input sequence and put them at proper places in the output sequence. Our empirical study on both synthetic data sets and real world data sets demonstrates the efficacy of CopyNet. For example, CopyNet can outperform regular RNN-based model with remarkable margins on text summarization tasks. Jiatao Gu, Zhengdong Lu, Hang Li 0001, Victor O. K. Li |
ACL (1) | 1 |