EDBT 2026 Demo / reviewers in the wild / expert
Yizhe Zhang 0002
dblp:132/4966-2
· DBLP profile ↗
66ranked-venue papers
14as first author
27since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 14 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ChipChat: Low-Latency Cascaded Conversational Agent in MLXabstractThe emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves subsecond response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents. Tatiana Likhomanenko, Luke Carlson, He Bai 0002, Zijin Gu, Han Tran, Zakaria Aldeneh, Yizhe Zhang 0002, Ruixiang Zhang, Huangjie Zheng, Navdeep Jaitly |
ASRU | 7 |
| 2025 | OpenHands: An Open Platform for AI Software Developers as Generalist AgentsabstractSoftware is one of the most powerful tools that we humans have at our disposal; it allows a skilled programmer to interact with the world in complex and profound ways. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that interact with and effect change in their surrounding environments. In this paper, we introduce OpenHands, a platform for the development of powerful and flexible AI agents that interact with the world in similar ways to a human developer: by writing code, interacting with a command line, and browsing the web. We describe how the platform allows for the implementation of new agents, utilization of various LLMs, safe interaction with sandboxed environments for code execution, and incorporation of evaluation benchmarks. Based on our currently incorporated benchmarks, we perform an evaluation of agents over 13 challenging tasks, including software engineering (e.g., SWE-Bench) and web browsing (e.g., WebArena), amongst others. Released under the permissive MIT license, OpenHands is a community project spanning academia and industry with more than 2K contributions from over 186 contributors in less than six months of development, and will improve going forward. Xingyao Wang 0002, Boxuan Li, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Yueqi Song, Bowen Li 0002, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang 0002, Binyuan Hui, Junyang Lin |
ICLR | 18 |
| 2025 | Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsabstractDiffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, current DLMs have been studied at a smaller scale compared to their AR counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. Given the prevalence of open-source AR language models, we propose adapting these models to build text diffusion models. We demonstrate connections between AR and diffusion modeling objectives and introduce a simple continual pre-training approach for training diffusion models. Through systematic evaluation on language modeling, reasoning, and commonsense benchmarks, we show that we can convert AR models ranging from 127M to 7B parameters (GPT2 and LLaMA) into diffusion models DiffuGPT and DiffuLLaMA, using less than 200B tokens for training. Our experimental results reveal that these models outperform earlier DLMs and are competitive with their AR counterparts. We release a suite of DLMs (127M-355M-7B) capable of generating fluent text, performing in-context learning, filling in the middle without prompt re-ordering, and following instructions. Shansan Gong, Shivam Agarwal, Yizhe Zhang 0002, Jiacheng Ye, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han 0001, Hao Peng 0009, Lingpeng Kong |
ICLR | 3 |
| 2025 | Denoising Autoregressive Transformers for Scalable Text-to-Image GenerationabstractDiffusion models have become the dominant approach for visual generation. They are trained by denoising a Markovian process which gradually adds noise to the input. We argue that the Markovian property limits the model’s ability to fully utilize the generation trajectory, leading to inefficiencies during training and inference. In this paper, we propose DART, a transformer-based model that unifies autoregressive (AR) and diffusion within a non-Markovian framework. DART iteratively denoises image patches spatially and spectrally using an AR model that has the same architecture as standard language models. DART does not rely on image quantization, which enables more effective image modeling while maintaining flexibility. Furthermore, DART seamlessly trains with both text and image data in a unified model. Our approach demonstrates competitive performance on class-conditioned and text-to-image generation tasks, offering a scalable, efficient alternative to traditional diffusion models. Through this unified framework, DART sets a new benchmark for scalable, high-quality image synthesis. Jiatao Gu, Yizhe Zhang 0002, Qihang Zhang, Dinghuai Zhang, Navdeep Jaitly, Joshua M. Susskind, Shuangfei Zhai |
ICLR | 3 |
| 2025 | Training Software Engineering Agents and Verifiers with SWE-GymabstractWe present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popular SWE-Bench Verified and Lite test sets. We also experiment with inference-time scaling through verifiers trained on agent trajectories sampled from SWE-Gym. When combined with our fine-tuned SWE agents, we achieve 32.0% and 26.0% on SWE-Bench Verified and Lite, respectively, reflecting a new state-of-the-art for open-weight SWE agents. To facilitate further research, we publicly release SWE-Gym, models, and agent trajectories. Xingyao Wang 0002, Graham Neubig, Navdeep Jaitly, Heng Ji 0001, Alane Suhr, Yizhe Zhang 0002 |
ICML | 7 |
| 2025 | Target Concrete Score Matching: A Holistic Framework for Discrete DiffusionabstractDiscrete diffusion is a promising framework for modeling and generating discrete data. In this work, we present Target Concrete Score Matching (TCSM), a novel and versatile objective for training and fine-tuning discrete diffusion models. TCSM provides a general framework with broad applicability. It supports pre-training discrete diffusion models directly from data samples, and many existing discrete diffusion approaches naturally emerge as special cases of our more general TCSM framework. Furthermore, the same TCSM objective extends to post-training of discrete diffusion models, including fine-tuning using reward functions or preference data, and distillation of knowledge from pre-trained autoregressive models. These new capabilities stem from the core idea of TCSM, estimating the concrete score of the target distribution, which resides in the original (clean) data space. This allows seamless integration with reward functions and pre-trained models, which inherently only operate in the clean data space rather than the noisy intermediate spaces of diffusion processes. Our experiments on language modeling tasks demonstrate that TCSM matches or surpasses current methods. Additionally, TCSM is versatile, applicable to both pre-training and post-training scenarios, offering greater flexibility and sample efficiency. Ruixiang Zhang, Shuangfei Zhai, Yizhe Zhang 0002, James Thornton, Zijing Ou, Joshua M. Susskind, Navdeep Jaitly |
ICML | 3 |
| 2025 | Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsabstractAutoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and single-pass decoding, while central to their success, also inspires the exploration of a design space that could offer new axes of modeling flexibility.
In this work, we explore an alternative paradigm, shifting language modeling from a discrete token space to a continuous latent space.
We propose a novel framework that employs transformer-based autoregressive normalizing flows to model these continuous representations.
This approach unlocks substantial flexibility, enabling the construction of models that can capture global bi-directional context through stacked, alternating-direction autoregressive transformations, support block-wise generation with flexible token patch sizes, and facilitate a hierarchical multi-pass generation process.
We further propose new mixture-based coupling transformations designed to capture complex dependencies within the latent space shaped by discrete data, and demonstrate theoretical connections to conventional discrete autoregressive models.
Extensive experiments on language modeling benchmarks demonstrate strong likelihood performance and highlight the flexible modeling capabilities inherent in our framework. Ruixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang 0002, Huangjie Zheng, Tianrong Chen, Miguel Ángel Bautista 0001, Joshua M. Susskind, Navdeep Jaitly |
NeurIPS | 4 |
| 2024 | Probing the Multi-turn Planning Capabilities of LLMs via 20 Question GamesabstractLarge language models (LLMs) are effective at answering questions that are clearly asked.However, when faced with ambiguous queries they can act unpredictably and produce incorrect outputs.This underscores the need for the development of intelligent agents capable of asking clarification questions to resolve ambiguities effectively.This capability requires complex understanding, state tracking, reasoning and planning over multiple conversational turns.However, directly measuring this can be challenging.In this paper, we offer a surrogate problem which assesses an LLMs's capability to deduce an entity unknown to itself, but revealed to a judge, by asking the judge a series of queries.This entity-deducing game can serve as an evaluation framework to probe the conversational reasoning and planning capabilities of language models.We systematically evaluate various LLMs and discover significant differences in their performance on this task.We find that strong LLMs like GPT-4 outperform human players by a large margin.We further employ Behavior Cloning (BC) to examine whether a weaker model is capable of imitating a stronger model and generalizing to data or domains, using only the demonstrations from a stronger model.We finally propose to use Reinforcement Learning to enhance reasoning and planning capacity of Vicuna models through episodes of game playing, which lead to significant performance improvement.We hope that this problem offers insights into how autonomous agents could be trained to behave more intelligently in ambiguous circumstances. Get MeshuggaThe band Play their song Do you mean Meshugana, the story set, or Meshuggah, the band?Do you want to know their information, or to play their album?Do you want something popular or unique from them?I am thinking of a movie that I cannot remember the name.Can you help me?Nope, it's like 10 years ago.I think so!Is this movie a drama movie? Yizhe Zhang 0002, Jiarui Lu, Navdeep Jaitly |
ACL (1) | 1 |
| 2024 | Rephrasing the Web: A Recipe for Compute and Data-Efficient Language ModelingabstractPratyush Maini, Skyler Seto, Richard Bai, David Grangier, Yizhe Zhang, Navdeep Jaitly. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Pratyush Maini, Skyler Seto, He Bai 0002, David Grangier, Yizhe Zhang 0002, Navdeep Jaitly |
ACL (1) | 5 |
| 2024 | Matryoshka Diffusion ModelsabstractDiffusion models are the de-facto approach for generating high-quality images and videos, but learning high-dimensional models remains a formidable task due to computational and optimization challenges. Existing methods often resort to training cascaded models in pixel space, or using a downsampled latent space of a separately trained auto-encoder. In this paper, we introduce Matryoshka Diffusion (MDM), an end-to-end framework for high-resolution image and video synthesis. We propose a diffusion process that denoises inputs at multiple resolutions jointly and uses a NestedUNet architecture where features and parameters for small-scale inputs are nested within those of large scales. In addition, MDM enables a progressive training schedule from lower to higher resolutions, which leads to significant improvements in optimization for high-resolution generation. We demonstrate the effectiveness of our approach on various benchmarks, including class-conditioned image generation, high-resolution text-to-image, and text-to-video applications. Remarkably, we can train a single pixel-space model at resolutions of up to 1024x1024 pixels, demonstrating strong zero-shot generalization using the CC12M dataset, which contains only 12 million images. Code and pre-trained checkpoints are released at https://github.com/apple/ml-mdm. Jiatao Gu, Shuangfei Zhai, Yizhe Zhang 0002, Joshua M. Susskind, Navdeep Jaitly |
ICLR | 3 |
| 2024 | Data-free Distillation of Diffusion Models with BootstrappingabstractDiffusion models have demonstrated great potential for generating diverse images. However, their performance often suffers from slow generation due to iterative denoising. Knowledge distillation has been recently proposed as a remedy which can reduce the number of inference steps to one or a few, without significant quality degradation. However, existing distillation methods either require significant amounts of offline computation for generating synthetic training data from the teacher model, or need to perform expensive online learning with the help of real data. In this work, we present a novel technique called BOOT, that overcomes these limitations with an efficient data-free distillation algorithm. The core idea is to learn a time-conditioned model that predicts the output of a pre-trained diffusion model teacher given any time-step. Such a model can be efficiently trained based on bootstrapping from two consecutive sampled steps. Furthermore, our method can be easily adapted to large-scale text-to-image diffusion models, which are challenging for previous methods given the fact that the training sets are often large and difficult to access. We demonstrate the effectiveness of our approach on several benchmark datasets in the DDIM setting, achieving comparable generation quality while being orders of magnitude faster than the diffusion teacher. The text-to-image results show that the proposed approach is able to handle highly complex distributions, shedding light on more efficient generative modeling. Jiatao Gu, Chen Wang 0049, Shuangfei Zhai, Yizhe Zhang 0002, Lingjie Liu, Joshua M. Susskind |
ICML | 4 |
| 2024 | Executable Code Actions Elicit Better LLM AgentsabstractLarge Language Model (LLM) agents, capable of performing a broad range of actions, such as invoking tools and controlling robots, show great potential in tackling real-world challenges. LLM agents are typically prompted to produce actions by generating JSON or text in a pre-defined format, which is usually limited by constrained action space (e.g., the scope of pre-defined tools) and restricted flexibility (e.g., inability to compose multiple tools). This work proposes to use executable Python code to consolidate LLM agents’ actions into a unified action space (CodeAct). Integrated with a Python interpreter, CodeAct can execute code actions and dynamically revise prior actions or emit new actions upon new observations through multi-turn interactions. Our extensive analysis of 17 LLMs on API-Bank and a newly curated benchmark shows that CodeAct outperforms widely used alternatives (up to 20% higher success rate). The encouraging performance of CodeAct motivates us to build an open-source LLM agent that interacts with environments by executing interpretable code and collaborates with users using natural language. To this end, we collect an instruction-tuning dataset CodeActInstruct that consists of 7k multi-turn interactions using CodeAct. We show that it can be used with existing data to improve models in agent-oriented tasks without compromising their general capability. CodeActAgent, finetuned from Llama2 and Mistral, is integrated with Python interpreter and uniquely tailored to perform sophisticated tasks (e.g., model training) using existing libraries and autonomously self-debug. Xingyao Wang 0002, Yangyi Chen, Lifan Yuan, Yizhe Zhang 0002, Yunzhu Li, Hao Peng 0009, Heng Ji 0001 |
ICML | 4 |
| 2024 | Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent ModelingabstractDiffusion models have emerged as a powerful tool for generating high-quality images from textual descriptions. Despite their successes, these models often exhibit limited diversity in the sampled images, particularly when sampling with a high classifier-free guidance weight. To address this issue, we present Kaleido, a novel approach that enhances the diversity of samples by incorporating autoregressive latent priors. Kaleido integrates an autoregressive language model that encodes the original caption and generates latent variables, serving as abstract and intermediary representations for guiding and facilitating the image generation process.
In this paper, we explore a variety of discrete latent representations, including textual descriptions, detection bounding boxes, object blobs, and visual tokens. These representations diversify and enrich the input conditions to the diffusion models, enabling more diverse outputs.
Our experimental results demonstrate that Kaleido effectively broadens the diversity of the generated image samples from a given textual description while maintaining high image quality. Furthermore, we show that Kaleido adheres closely to the guidance provided by the generated latent variables, demonstrating its capability to effectively control and direct the image generation process. Jiatao Gu, Ying Shen 0006, Shuangfei Zhai, Yizhe Zhang 0002, Navdeep Jaitly, Joshua M. Susskind |
NeurIPS | 4 |
| 2023 | Towards More Efficient Insertion Transformer with Fractional Positional EncodingabstractAuto-regressive neural sequence models have been shown to be effective across text generation tasks.However, their left-to-right decoding order prevents generation from being parallelized.Insertion Transformer (Stern et al., 2019) is an attractive alternative that allows outputting multiple tokens in a single generation step.Nevertheless, due to the incompatibility between absolute positional encoding and insertion-based generation schemes, it needs to refresh the encoding of every token in the generated partial hypothesis at each step, which could be costly.We design a novel reusable positional encoding scheme for Insertion Transformers called Fractional Positional Encoding (FPE), which allows reusing representations calculated in previous steps.Empirical studies on various text generation tasks demonstrate the effectiveness of FPE, which leads to floating-point operation reduction and latency improvements on batched decoding. Zhisong Zhang, Yizhe Zhang 0002, William B. Dolan |
EACL | 2 |
| 2023 | Interactive Text GenerationabstractFelix Faltings, Michel Galley, Kianté Brantley, Baolin Peng, Weixin Cai, Yizhe Zhang, Jianfeng Gao, Bill Dolan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Felix Faltings, Michel Galley, Kianté Brantley, Baolin Peng, Weixin Cai, Yizhe Zhang 0002, Jianfeng Gao 0001, William B. Dolan |
EMNLP | 6 |
| 2023 | f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang 0002, Miguel Ángel Bautista 0001, Joshua M. Susskind |
ICLR | 3 |
| 2023 | Stabilizing Transformer Training by Preventing Attention Entropy CollapseabstractTraining stability is of great importance to Transformers. In this work, we investigate the training dynamics of Transformers by examining the evolution of the attention layers. In particular, we track the attention entropy for each attention head during the course of training, which is a proxy for model sharpness. We identify a common pattern across different architectures and tasks, where low attention entropy is accompanied by high training instability, which can take the form of oscillating loss or divergence. We denote the pathologically low attention entropy, corresponding to highly concentrated attention scores, as $\textit{entropy collapse}$. As a remedy, we propose $\sigma$Reparam, a simple and efficient solution where we reparametrize all linear layers with spectral normalization and an additional learned scalar. We demonstrate that $\sigma$Reparam successfully prevents entropy collapse in the attention layers, promoting more stable training. Additionally, we prove a tight lower bound of the attention entropy, which decreases exponentially fast with the spectral norm of the attention logits, providing additional motivation for our approach. We conduct experiments with $\sigma$Reparam on image classification, image self-supervised learning, machine translation, speech recognition, and language modeling tasks. We show that $\sigma$Reparam provides stability and robustness with respect to the choice of hyperparameters, going so far as enabling training (a) a Vision Transformer to competitive performance without warmup, weight decay, layer normalization or adaptive optimizers; (b) deep architectures in machine translation and (c) speech recognition to competitive performance without warmup and adaptive optimizers. Code is available at https://github.com/apple/ml-sigma-reparam. Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang 0002, Jiatao Gu, Joshua M. Susskind |
ICML | 6 |
| 2023 | PLANNER: Generating Diversified Paragraph via Latent Language Diffusion ModelabstractAutoregressive models for text sometimes generate repetitive and low-quality output because errors accumulate during the steps of generation. This issue is often attributed to exposure bias -- the difference between how a model is trained, and how it is used during inference. Denoising diffusion models provide an alternative approach in which a model can revisit and revise its output. However, they can be computationally expensive and prior efforts on text have led to models that produce less fluent output compared to autoregressive models, especially for longer text and paragraphs. In this paper, we propose PLANNER, a model that combines latent semantic diffusion with autoregressive generation, to generate fluent text while exercising global control over paragraphs. The model achieves this by combining an autoregressive "decoding" module with a "planning" module that uses latent diffusion to generate semantic paragraph embeddings in a coarse-to-fine manner. The proposed method is evaluated on various conditional generation tasks, and results on semantic generation, text completion and summarization show its effectiveness in generating high-quality long-form text in an efficient manner. Yizhe Zhang 0002, Jiatao Gu, Zhuofeng Wu 0001, Shuangfei Zhai, Joshua M. Susskind, Navdeep Jaitly |
NeurIPS | 1 |
| 2023 | Learning Deep Generative Clustering via Mutual Information MaximizationabstractDeep clustering refers to joint representation learning and clustering using deep neural networks. Existing methods can be mainly categorized into two types: discriminative and generative methods. The former learns representations for clustering with discriminative mechanisms directly, and the latter estimate the latent distribution of each cluster for generating data points and then infers cluster assignments. Although generative methods have the advantage of estimating the latent distributions of clusters, their performances still significantly fall behind discriminative methods. In this work, we argue that this performance gap might be partly due to the overlap of data distribution of different clusters. In fact, there is little guarantee of generative methods to separate the distributions of different clusters in the data space. To tackle these problems, we theoretically prove that mutual information maximization promotes the separation of different clusters in the data space, which provides a theoretical justification for deep generative clustering with mutual information maximization. Our theoretical analysis directly leads to a model which integrates a hierarchical generative adversarial network and mutual information maximization. Moreover, we further propose three techniques and empirically show their effects to stabilize and enhance the model. The proposed approach notably outperforms other generative models for deep clustering on public benchmarks. Xiaojiang Yang, Junchi Yan, Yu Cheng 0001, Yizhe Zhang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | RetGen: A Joint Framework for Retrieval and Grounded Text Generation ModelingabstractRecent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded generation models appear to offer remedies, but their training typically relies on rarely-available parallel data where information-relevant documents are provided for context. We propose a framework that alleviates this data constraint by jointly training a grounded generator and document retriever on the language model signal. The model learns to reward retrieval of the documents with the highest utility in generation, and attentively combines them using a Mixture-of-Experts (MoE) ensemble to generate follow-on text. We demonstrate that both generator and retriever can take advantage of this joint training and work synergistically to produce more informative and relevant text in both prose and dialogue generation. Yizhe Zhang 0002, Xiang Gao 0011, Yuwei Fang, Chris Brockett, Michel Galley, Jianfeng Gao 0001, William B. Dolan |
AAAI | 1 |
| 2022 | A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationabstractTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, Bill Dolan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Tianyu Liu 0001, Yizhe Zhang 0002, Chris Brockett, Zhifang Sui, Weizhu Chen, William B. Dolan |
ACL (1) | 2 |
| 2022 | Linearizing Transformer with Key-Value MemoryabstractEfficient transformer variants with linear time complexity have been developed to mitigate the quadratic computational overhead of the vanilla transformer.Among them are lowrank projection methods such as Linformer and kernel-based Transformers.Despite their unique merits, they usually suffer from a performance drop comparing with the vanilla transformer on many sequence generation tasks, and often fail to obtain computation gain when the generation is short.We propose Mem-Sizer, an approach towards closing the performance gap while improving the efficiency even with short generation.It projects the source sequences into lower dimension representations like Linformer, while enjoying efficient recurrent-style incremental computation similar to kernel-based transformers.This yields linear computation time and constant memory complexity at inference time.Mem-Sizer also employs a lightweight multi-head mechanism which renders the computation as light as a single-head model.We demonstrate that MemSizer provides an improved balance between efficiency and accuracy over the vanilla transformer and other efficient transformer variants in three typical sequence generation tasks, including machine translation, abstractive text summarization, and language modeling.Our code is released at https: //github.com/jcyk/memsizer Yizhe Zhang 0002, Deng Cai 0002 |
EMNLP | 1 |
| 2021 | Data Augmentation for Abstractive Query-Focused Multi-Document SummarizationabstractThe progress in Query-focused Multi-Document Summarization (QMDS) has been limited by the lack of sufficient largescale high-quality training datasets. We present two QMDS training datasets, which we construct using two data augmentation methods: (1) transferring the commonly used single-document CNN/Daily Mail summarization dataset to create the QMDSCNN dataset, and (2) mining search-query logs to create the QMDSIR dataset. These two datasets have complementary properties, i.e., QMDSCNN has real summaries but queries are simulated, while QMDSIR has real queries but simulated summaries. To cover both these real summary and query aspects, we build abstractive end-to-end neural network models on the combined datasets that yield new state-of-the-art transfer results on DUC datasets. We also introduce new hierarchical encoders that enable a more efficient encoding of the query together with multiple documents. Empirical results demonstrate that our data augmentation and encoding methods outperform baseline models on automatic metrics, as well as on human evaluations along multiple attributes. Ramakanth Pasunuru, Asli Celikyilmaz, Michel Galley, Chenyan Xiong, Yizhe Zhang 0002, Mohit Bansal, Jianfeng Gao 0001 |
AAAI | 5 |
| 2021 | A Controllable Model of Grounded Response GenerationabstractCurrent end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines. Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan |
AAAI | 4 |
| 2021 | Contrastive Multi-document Question GenerationabstractWoon Sang Cho, Yizhe Zhang, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao, Mengdi Wang, Bill Dolan. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Woon Sang Cho, Yizhe Zhang 0002, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao 0001, Mengdi Wang 0001, William B. Dolan |
EACL | 2 |
| 2021 | Finetuning Pretrained Transformers into RNNsabstractJungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, Noah A. Smith. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jungo Kasai, Hao Peng 0009, Yizhe Zhang 0002, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas 0002, Weizhu Chen, Noah A. Smith |
EMNLP (1) | 3 |
| 2021 | Contextualized Perturbation for Textual Adversarial AttackabstractDianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dianqi Li, Yizhe Zhang 0002, Hao Peng 0009, Liqun Chen 0001, Chris Brockett, Ming-Ting Sun, William B. Dolan |
NAACL-HLT | 2 |
| 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement LearningabstractReinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions. Liqun Chen 0001, Ke Bai 0001, Chenyang Tao, Yizhe Zhang 0002, Guoyin Wang 0002, Wenlin Wang, Ricardo Henao, Lawrence Carin |
AAAI | 4 |
| 2020 | Complementary Auxiliary Classifiers for Label-Conditional Text GenerationabstractLearning to generate text with a given label is a challenging task because natural language sentences are highly variable and ambiguous. It renders difficulties in trade-off between sentence quality and label fidelity. In this paper, we present CARA to alleviate the issue, where two auxiliary classifiers work simultaneously to ensure that (1) the encoder learns disentangled features and (2) the generator produces label-related sentences. Two practical techniques are further proposed to improve the performance, including annealing the learning signal from the auxiliary classifier, and enhancing the encoder with pre-trained language models. To establish a comprehensive benchmark fostering future research, we consider a suite of four datasets, and systematically reproduce three representative methods. CARA shows consistent improvement over the previous methods on the task of label-conditional text generation, and achieves state-of-the-art on the task of attribute transfer. Yuan Li 0032, Chunyuan Li, Yizhe Zhang 0002, Xiujun Li, Guoqing Zheng, Lawrence Carin, Jianfeng Gao 0001 |
AAAI | 3 |
| 2020 | Contrastively Smoothed Class Alignment for Unsupervised Domain Adaptation
Shuyang Dai, Yu Cheng 0001, Yizhe Zhang 0002, Zhe Gan, Jingjing Liu 0001, Lawrence Carin |
ACCV (4) | 3 |
| 2020 | Improving Disentangled Text Representation Learning with Information-Theoretic GuidanceabstractPengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang, Yitong Li, Lawrence Carin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang 0002, Yitong Li 0001, Lawrence Carin |
ACL | 5 |
| 2020 | INSET: Sentence Infilling with INter-SEntential TransformerabstractMissing sentence generation (or sentence infilling) fosters a wide range of applications in natural language generation, such as document auto-completion and meeting note expansion.This task asks the model to generate intermediate missing sentences that can syntactically and semantically bridge the surrounding context.Solving the sentence infilling task requires techniques in natural language processing ranging from understanding to discourselevel planning to generation.In this paper, we propose a framework to decouple the challenge and address these three aspects respectively, leveraging the power of existing largescale pre-trained models such as BERT and GPT-2.We empirically demonstrate the effectiveness of our model in learning a sentence representation for generation and further generating a missing sentence that fits the context. Yizhe Zhang 0002, Oussama Elachqar, Yu Cheng 0001 |
ACL | 2 |
| 2020 | Advancing weakly supervised cross-domain alignment with optimal transport
Siyang Yuan, Ke Bai 0001, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Chunyuan Li, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin |
BMVC | 4 |
| 2020 | Dialogue Response Ranking Training with Large-Scale Human Feedback DataabstractExisting open-domain dialog models are generally trained to minimize the perplexity of target human responses.However, some human replies are more engaging than others, spawning more followup interactions.Current conversational models are increasingly capable of producing turns that are context-relevant, but in order to produce compelling agents, these models need to be able to predict and optimize for turns that are genuinely engaging.We leverage social media feedback data (number of replies and upvotes) to build a large-scale training dataset for feedback prediction.To alleviate possible distortion between the feedback and engagingness, we convert the ranking problem to a comparison of response pairs which involve few confounding factors.We trained DIALOGRPT, a set of GPT-2 based models on 133M pairs of human feedback data and the resulting ranker outperformed several baselines.Particularly, our ranker outperforms the conventional dialog perplexity baseline with a large margin on predicting Reddit feedback.We finally combine the feedback prediction models and a human-like scoring model to rank the machine-generated dialog responses.Crowd-sourced human evaluation shows that our ranking method correlates better with real human preferences than baseline models. 1 Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, William B. Dolan |
EMNLP (1) | 2 |
| 2020 | Optimus: Organizing Sentences via Pre-trained Modeling of a Latent SpaceabstractWhen trained effectively, the Variational Autoencoder (VAE) (Kingma and Welling, 2013;Bowman et al., 2016) can be both a powerful generative model and an effective representation learning framework for natural language.In this paper, we propose the first large-scale language VAE model OPTIMUS 1 .A universal latent embedding space for sentences is first pre-trained on large text corpus, and then fine-tuned for various language generation and understanding tasks.Compared with GPT-2, OPTIMUS enables guided language generation from an abstract level using the latent vectors.Compared with BERT, OPTIMUS can generalize better on low-resource language understanding tasks due to the smooth latent space structure.Extensive experimental results on a wide range of language tasks demonstrate the effectiveness of OPTIMUS.It achieves new state-of-the-art on VAE language modeling benchmarks.Encoder Chunyuan Li, Xiang Gao 0011, Yuan Li 0032, Baolin Peng, Xiujun Li, Yizhe Zhang 0002, Jianfeng Gao 0001 |
EMNLP (1) | 6 |
| 2020 | Improving Text Generation with Student-Forcing Optimal TransportabstractJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu, Yuhchen Lin, Liqun Chen, Yizhe Zhang, Chenyang Tao, Ruiyi Zhang, Wenlin Wang, Dinghan Shen, Qian Yang, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jianqiao Li, Chunyuan Li, Guoyin Wang 0002, Hao Fu 0002, Yuh-Chen Lin, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Ruiyi Zhang 0002, Wenlin Wang, Dinghan Shen, Qian Yang 0003, Lawrence Carin |
EMNLP (1) | 7 |
| 2020 | POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-trainingabstractLarge-scale pre-trained language models, such as BERT and GPT-2, have achieved excellent performance in language representation learning and free-form text generation.However, these models cannot be directly employed to generate text under specified lexical constraints.To address this challenge, we present POINTER 1 , a simple yet novel insertion-based approach for hard-constrained text generation.The proposed method operates by progressively inserting new tokens between existing tokens in a parallel manner.This procedure is recursively applied until a sequence is completed.The resulting coarse-to-fine hierarchy makes the generation process intuitive and interpretable.We pre-train our model with the proposed progressive insertion-based objective on a 12GB Wikipedia dataset, and finetune it on downstream hard-constrained generation tasks.Non-autoregressive decoding yields an empirically logarithmic time complexity during inference time.Experimental results on both News and Yelp datasets demonstrate that POINTER achieves state-of-the-art performance on constrained text generation.We released the pre-trained models and the source code to facilitate future research 2 . Yizhe Zhang 0002, Guoyin Wang 0002, Chunyuan Li, Zhe Gan, Chris Brockett, William B. Dolan |
EMNLP (1) | 1 |
| 2020 | Adaptive Correlated Monte Carlo for Contextual Categorical Sequence Generation
Xinjie Fan, Yizhe Zhang 0002, Zhendong Wang 0005, Mingyuan Zhou |
ICLR | 2 |
| 2020 | Datasets and Benchmarks for Task-Oriented Log Dialogue Ranking Task
Xinnuo Xu, Yizhe Zhang 0002, Lars Liden |
INTERSPEECH | 2 |
| 2020 | Contextual Re-Ranking with Behavior Aware TransformersabstractIn this work, we focus on the contextual document ranking task, which deals with the challenge of user interaction modeling for conversational search. Given a history of user feedback behaviors, such as issuing a query, clicking a document, and skipping a document, we propose to introduce behavior awareness to a neural ranker, resulting in a Hierarchical Behavior Aware Transformers (HBA-Transformers) model. The hierarchy is composed of an intra-behavior attention layer and an inter-behavior attention layer to let the system effectively distinguish and model different user behaviors. Our extensive experiments on the AOL session dataset demonstrate that the hierarchical behavior aware architecture is more powerful than a simple combination of history behaviors. Besides, we analyze the conversational property of queries. We show that coherent sessions tend to be more conversational and thus are more demanding in terms of considering history user behaviors. Chen Qu 0001, Chenyan Xiong, Yizhe Zhang 0002, Corby Rosset, W. Bruce Croft, Paul N. Bennett |
SIGIR | 3 |
| 2019 | Improving Textual Network Embedding with Global Attention via Optimal TransportabstractLiqun Chen, Guoyin Wang, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang, Wenlin Wang, Yizhe Zhang, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Liqun Chen 0001, Guoyin Wang 0002, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang 0001, Wenlin Wang, Yizhe Zhang 0002, Lawrence Carin |
ACL (1) | 8 |
| 2019 | Towards Generating Long and Coherent Text with Multi-Level Latent Variable ModelsabstractVariational autoencoders (VAEs) have received much attention recently as an end-toend architecture for text generation with latent variables.However, previous works typically focus on synthesizing relatively short sentences (up to 20 words), and the posterior collapse issue has been widely identified in text-VAEs.In this paper, we propose to leverage several multi-level structures to learn a VAE model for generating long, and coherent text.In particular, a hierarchy of stochastic layers between the encoder and decoder networks is employed to abstract more informative and semantic-rich latent codes.Besides, we utilize a multi-level decoder structure to capture the coherent long-term structure inherent in long-form texts, by generating intermediate sentence representations as highlevel plan vectors.Extensive experimental results demonstrate that the proposed multi-level VAE model produces more coherent and less repetitive long text compared to baselines as well as can mitigate the posterior-collapse issue. Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang 0002, Liqun Chen 0001, Xin Wang 0061, Jianfeng Gao 0001, Lawrence Carin |
ACL (1) | 3 |
| 2019 | Structuring Latent Spaces for Stylized Response GenerationabstractXiang Gao, Yizhe Zhang, Sungjin Lee, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Domain Adaptive Text Style TransferabstractDianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, Ming-Ting Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dianqi Li, Yizhe Zhang 0002, Zhe Gan, Yu Cheng 0001, Chris Brockett, William B. Dolan, Ming-Ting Sun |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Improving Sequence-to-Sequence Learning via Optimal Transport
Liqun Chen 0001, Yizhe Zhang 0002, Ruiyi Zhang 0002, Chenyang Tao, Zhe Gan, Bai Li 0001, Dinghan Shen, Changyou Chen, Lawrence Carin |
ICLR (Poster) | 2 |
| 2019 | Unsupervised Dialogue Spectrum Generation for Log Dialogue RankingabstractAlthough the data-driven approaches of some recent bot building platforms make it possible for a wide range of users to easily create dialogue systems, those platforms don't offer tools for quickly identifying which log dialogues contain problems.This is important since corrections to log dialogues provide a means to improve performance after deployment.A log dialogue ranker, which ranks problematic dialogues higher, is an essential tool due to the sheer volume of log dialogues that could be generated.However, training a ranker typically requires labelling a substantial amount of data, which is not feasible for most users.In this paper, we present a novel unsupervised approach for dialogue ranking using GANs and release a corpus of labelled dialogues for evaluation and comparison with supervised methods.The evaluation result shows that our method compares favorably to supervised methods without any labelled data. Xinnuo Xu, Yizhe Zhang 0002, Lars Liden |
SIGdial | 2 |
| 2019 | A convergence analysis for a class of practical variance-reduction stochastic gradient MCMC
Changyou Chen, Wenlin Wang, Yizhe Zhang 0002, Qinliang Su, Lawrence Carin |
Sci. China Inf. Sci. | 3 |
| 2018 | Deconvolutional Latent-Variable Model for Text Sequence MatchingabstractA latent-variable model is introduced for text matching, inferring sentence representations by jointly optimizing generative and discriminative objectives. To alleviate typical optimization challenges in latent-variable models for text, we employ deconvolutional networks as the sequence decoder (generator), providing learned latent codes with more semantic information and better generalization. Our model, trained in an unsupervised manner, yields stronger empirical predictive performance than a decoder based on Long Short-Term Memory (LSTM), with less parameters and considerably faster training. Further, we apply it to text sequence-matching problems. The proposed model significantly outperforms several strong sentence-encoding baselines, especially in the semi-supervised setting. Dinghan Shen, Yizhe Zhang 0002, Ricardo Henao, Qinliang Su, Lawrence Carin |
AAAI | 2 |
| 2018 | Zero-Shot Learning via Class-Conditioned Deep Generative ModelsabstractWe present a deep generative model for Zero-Shot Learning (ZSL). Unlike most existing methods for this problem, that represent each class as a point (via a semantic embedding), we represent each seen/unseen class using a class-specific latent-space distribution, conditioned on class attributes. We use these latent-space distributions as a prior for a supervised variational autoencoder (VAE), which also facilitates learning highly discriminative feature representations for the inputs. The entire framework is learned end-to-end using only the seen-class training data. At test time, the label for an unseen-class test input is the class that maximizes the VAE lower bound. We further extend the model to a (i) semi-supervised/transductive setting by leveraging unlabeled unseen-class data via an unsupervised learning module, and (ii) few-shot learning where we also have a small number of labeled inputs from the unseen classes. We compare our model with several state-of-the-art methods through a comprehensive set of experiments on a variety of benchmark data sets. Wenlin Wang, Yunchen Pu, Vinay Kumar Verma, Kai Fan 0002, Yizhe Zhang 0002, Changyou Chen, Piyush Rai, Lawrence Carin |
AAAI | 5 |
| 2018 | Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling MechanismsabstractDinghan Shen, Guoyin Wang, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang, Chunyuan Li, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Dinghan Shen, Guoyin Wang 0002, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang 0002, Chunyuan Li, Ricardo Henao, Lawrence Carin |
ACL (1) | 6 |
| 2018 | Joint Embedding of Words and Labels for Text ClassificationabstractGuoyin Wang, Chunyuan Li, Wenlin Wang, Yizhe Zhang, Dinghan Shen, Xinyuan Zhang, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Guoyin Wang 0002, Chunyuan Li, Wenlin Wang, Yizhe Zhang 0002, Dinghan Shen, Xinyuan Zhang 0001, Ricardo Henao, Lawrence Carin |
ACL (1) | 4 |
| 2018 | JointGAN: Multi-Domain Joint Distribution Learning with Generative Adversarial NetsabstractA new generative adversarial network is developed for joint distribution matching.Distinct from most existing approaches, that only learn conditional distributions, the proposed model aims to learn a joint distribution of multiple random variables (domains). This is achieved by learning to sample from conditional distributions between the domains, while simultaneously learning to sample from the marginals of each individual domain.The proposed framework consists of multiple generators and a single softmax-based critic, all jointly trained via adversarial learning.From a simple noise source, the proposed framework allows synthesis of draws from the marginals, conditional draws given observations from a subset of random variables, or complete draws from the full joint distribution. Most examples considered are for joint analysis of two domains, with examples for three domains also presented. Yunchen Pu, Shuyang Dai, Zhe Gan, Weiyao Wang 0002, Guoyin Wang 0002, Yizhe Zhang 0002, Ricardo Henao, Lawrence Carin |
ICML | 6 |
| 2018 | Adversarial Text Generation via Feature-Mover's DistanceabstractGenerative adversarial networks (GANs) have achieved significant success in generating real-valued data. However, the discrete nature of text hinders the application of GAN to text-generation tasks. Instead of using the standard GAN objective, we propose to improve text-generation GAN via a novel approach inspired by optimal transport. Specifically, we consider matching the latent feature distributions of real and synthetic sentences using a novel metric, termed the feature-mover's distance (FMD). This formulation leads to a highly discriminative critic and easy-to-optimize objective, overcoming the mode-collapsing and brittle-training problems in existing methods. Extensive experiments are conducted on a variety of tasks to evaluate the proposed model empirically, including unconditional text generation, style transfer from non-parallel text, and unsupervised cipher cracking. The proposed model yields superior performance, demonstrating wide applicability and effectiveness. Liqun Chen 0001, Shuyang Dai, Chenyang Tao, Zhe Gan, Dinghan Shen, Yizhe Zhang 0002, Guoyin Wang 0002, Ruiyi Zhang 0002, Lawrence Carin |
NeurIPS | 7 |
| 2018 | Generating Informative and Diverse Conversational Responses via Adversarial Information MaximizationabstractResponses generated by neural conversational models tend to lack informativeness and diversity. We present Adversarial Information Maximization (AIM), an adversarial learning framework that addresses these two related but distinct problems. To foster response diversity, we leverage adversarial training that allows distributional matching of synthetic and real responses. To improve informativeness, our framework explicitly optimizes a variational lower bound on pairwise mutual information between query and response. Empirical results from automatic and human evaluations demonstrate that our methods significantly boost informativeness and diversity. Yizhe Zhang 0002, Michel Galley, Jianfeng Gao 0001, Zhe Gan, Xiujun Li, Chris Brockett, William B. Dolan |
NeurIPS | 1 |
| 2017 | Stochastic Gradient Monomial Gamma SamplerabstractScaling Markov Chain Monte Carlo (MCMC) to estimate posterior distributions from large datasets has been made possible as a result of advances in stochastic gradient techniques. Despite their success, mixing performance of existing methods when sampling from multimodal distributions can be less efficient with insufficient Monte Carlo samples; this is evidenced by slow convergence and insufficient exploration of posterior distributions. We propose a generalized framework to improve the sampling efficiency of stochastic gradient MCMC, by leveraging a generalized kinetics that delivers superior stationary mixing, especially in multimodal distributions, and propose several techniques to overcome the practical issues. We show that the proposed approach is better at exploring a complicated multimodal posterior distribution, and demonstrate improvements over other stochastic gradient MCMC methods on various applications. Yizhe Zhang 0002, Changyou Chen, Zhe Gan, Ricardo Henao, Lawrence Carin |
ICML | 1 |
| 2017 | Adversarial Feature Matching for Text GenerationabstractThe Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We employ a long short-term memory network as generator, and a convolutional network as discriminator. Instead of using the standard objective of GAN, we propose matching the high-dimensional latent feature distributions of real and synthetic sentences, via a kernelized discrepancy metric. This eases adversarial training by alleviating the mode-collapsing problem. Our experiments show superior performance in quantitative evaluation, and demonstrate that our model can generate realistic-looking sentences. Yizhe Zhang 0002, Zhe Gan, Kai Fan 0002, Zhi Chen 0009, Ricardo Henao, Dinghan Shen, Lawrence Carin |
ICML | 1 |
| 2017 | Triangle Generative Adversarial NetworksabstractA Triangle Generative Adversarial Network ($\Delta$-GAN) is developed for semi-supervised cross-domain joint distribution matching, where the training data consists of samples from each domain, and supervision of domain correspondence is provided by only a few paired samples. $\Delta$-GAN consists of four neural networks, two generators and two discriminators. The generators are designed to learn the two-way conditional distributions between the two domains, while the discriminators implicitly define a ternary discriminative function, which is trained to distinguish real data pairs and two kinds of fake data pairs. The generators and discriminators are trained together using adversarial learning. Under mild assumptions, in theory the joint distributions characterized by the two generators concentrate to the data distribution. In experiments, three different kinds of domain pairs are considered, image-label, image-image and image-attribute pairs. Experiments on semi-supervised image classification, image-to-image translation and attribute-based image generation demonstrate the superiority of the proposed approach. Zhe Gan, Liqun Chen 0001, Weiyao Wang 0002, Yunchen Pu, Yizhe Zhang 0002, Hao Liu 0015, Chunyuan Li, Lawrence Carin |
NIPS | 5 |
| 2017 | Deconvolutional Paragraph Representation LearningabstractLearning latent representations from long text sequences is an important first step in many natural language processing applications. Recurrent Neural Networks (RNNs) have become a cornerstone for this challenging task. However, the quality of sentences during RNN-based decoding (reconstruction) decreases with the length of the text. We propose a sequence-to-sequence, purely convolutional and deconvolutional autoencoding framework that is free of the above issue, while also being computationally efficient. The proposed method is simple, easy to implement and can be leveraged as a building block for many applications. We show empirically that compared to RNNs, our framework is better at reconstructing and correcting long paragraphs. Quantitative evaluation on semi-supervised text classification and summarization tasks demonstrate the potential for better utilization of long unlabeled text data. Yizhe Zhang 0002, Dinghan Shen, Guoyin Wang 0002, Zhe Gan, Ricardo Henao, Lawrence Carin |
NIPS | 1 |
| 2016 | Learning a Hybrid Architecture for Sequence Regression and AnnotationabstractWhen learning a hidden Markov model (HMM), sequential observations can often be complemented by real-valued summary response variables generated from the path of hidden states. Such settings arise in numerous domains, including many applications in biology, like motif discovery and genome annotation. In this paper, we present a flexible framework for jointly modeling both latent sequence features and the functional mapping that relates the summary response variables to the hidden state sequence. The algorithm is compatible with a rich set of mapping functions. Results show that the availability of additional continuous response variables can simultaneously improve the annotation of the sequential observations and yield good prediction performance in both synthetic data and real-world datasets. Yizhe Zhang 0002, Ricardo Henao, Lawrence Carin, Jianling Zhong, Alexander J. Hartemink |
AAAI | 1 |
| 2016 | Triply Stochastic Variational Inference for Non-linear Beta Process Factor AnalysisabstractWe propose a non-linear extension to factor analysis with beta process priors for improved data representation ability. This non-linear Beta Process Factor Analysis (nBPFA) allows data to be represented as a non-linear transformation of a standard sparse factor decomposition. We develop a scalable variational inference framework, which builds upon the ideas of the variational auto-encoder, by allowing latent variables of the model to be sparse. Our framework can be readily used for real-valued, binary and count data. We show theoretically and with experiments that our training scheme, with additive or multiplicative noise on observations, improves performance and prevents overfitting. We benchmark our algorithms on image, text and collaborative filtering datasets. We demonstrate faster convergence rates and competitive performance compared to standard gradient-based approaches. Kai Fan 0002, Yizhe Zhang 0002, Ricardo Henao, Katherine A. Heller |
ICDM | 2 |
| 2016 | Dynamic Poisson Factor AnalysisabstractWe introduce a novel dynamic model for discrete time-series data, in which the temporal sampling may be nonuniform. The model is specified by constructing a hierarchy of Poisson factor analysis blocks, one for the transitions between latent states and the other for the emissions between latent states and observations. Latent variables are binary and linked to Poisson factor analysis via Bernoulli-Poisson specifications. The model is derived for count data but can be readily modified for binary observations. We derive efficient inference via Markov chain Monte Carlo, that scales with the number of non-zeros in the data and latent binary states, yielding significant acceleration compared to related models. Experimental results on benchmark data show the proposed model achieves state-of-the-art predictive performance. Additional experiments on microbiome data demonstrate applicability of the proposed model to interesting problems in computational biology where interpretability is of utmost importance. Yizhe Zhang 0002, Lawrence David, Ricardo Henao, Lawrence Carin |
ICDM | 1 |
| 2016 | Bayesian Dictionary Learning with Gaussian Processes and Sigmoid Belief Networks
Yizhe Zhang 0002, Ricardo Henao, Chunyuan Li, Lawrence Carin |
IJCAI | 1 |
| 2016 | Stochastic Gradient MCMC with Stale GradientsabstractStochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, where stochastic gradients are computed based on some outdated parameters, yielding what are termed stale gradients. While stale gradients could be directly used in SG-MCMC, their impact on convergence properties has not been well studied. In this paper we develop theory to show that while the bias and MSE of an SG-MCMC algorithm depend on the staleness of stochastic gradients, its estimation variance (relative to the expected estimate, based on a prescribed number of samples) is independent of it. In a simple Bayesian distributed system with SG-MCMC, where stale gradients are computed asynchronously by a set of workers, our theory indicates a linear speedup on the decrease of estimation variance w.r.t. the number of workers. Experiments on synthetic data and deep neural networks validate our theory, demonstrating the effectiveness and scalability of SG-MCMC with stale gradients. Changyou Chen, Nan Ding 0002, Chunyuan Li, Yizhe Zhang 0002, Lawrence Carin |
NIPS | 4 |
| 2016 | Towards Unifying Hamiltonian Monte Carlo and Slice SamplingabstractWe unify slice sampling and Hamiltonian Monte Carlo (HMC) sampling, demonstrating their connection via the Hamiltonian-Jacobi equation from Hamiltonian mechanics. This insight enables extension of HMC and slice sampling to a broader family of samplers, called Monomial Gamma Samplers (MGS). We provide a theoretical analysis of the mixing performance of such samplers, proving that in the limit of a single parameter, the MGS draws decorrelated samples from the desired target distribution. We further show that as this parameter tends toward this limit, performance gains are achieved at a cost of increasing numerical difficulty and some practical convergence issues. Our theoretical results are validated with synthetic data and real-world applications. Yizhe Zhang 0002, Xiangyu Wang 0006, Changyou Chen, Ricardo Henao, Kai Fan 0002, Lawrence Carin |
NIPS | 1 |
| 2016 | Laplacian Hamiltonian Monte Carlo
Yizhe Zhang 0002, Changyou Chen, Ricardo Henao, Lawrence Carin |
ECML/PKDD (1) | 1 |
| 2002 | Model-based statistical sensor fusion for unexploded ordnance detectionabstractDetection and remediation of unexploded ordnance (UXO) represents a major challenge on closed, closing, and transferred military ranges as well as on active installations. The detection problem is exacerbated by the fact that on sites contaminated with UXO, extensive surface and sub-surface clutter and shrapnel is also present. Traditional methods used for UXO remediation have difficulty distinguishing buried UXO from these anthropic clutter items as well as from naturally occurring magnetic geologic noise, and thus incur prohibitively high false alarm rates. The reduction of the false alarm rate has proven to be the greatest challenge for UXO remediation. In this paper, sensor fusion techniques are applied to field data from magnetometer and electromagnetic induction (EMI) sensors in order to determine to what degree such an approach results in false alarm mitigation. The adoption of a model consisting of multiple non-colocated dipoles is shown to improve our ability to predict measured signatures. The results indicate that performance can be improved by limiting the processing bandwidth to those frequencies that are the most robust to naturally occurring geological noise. Leslie M. Collins, Yizhe Zhang 0002, Lawrence Carin |
IGARSS | 2 |