Yong Cheng 0003

dblp:34/6276-3 · DBLP profile ↗
← Back
24ranked-venue papers
10as first author
9since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 10 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Language Model Beats Diffusion - Tokenizer is key to visual generation
abstract
While Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to discrete tokens appropriate for LLM learning. In this paper, we introduce \modelname{}, a video tokenizer designed to generate concise and expressive tokens for both videos and images using a common token vocabulary. Equipped with this new tokenizer, we show that LLMs outperform diffusion models on standard image and video generation benchmarks including ImageNet and Kinetics. In addition, we demonstrate that our tokenizer surpasses the previously top-performing video tokenizer on two more tasks: (1) video compression comparable to the next-generation video codec (VCC) according to human evaluations, and (2) learning effective representations for action recognition tasks.
Lijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng 0003, Agrim Gupta, Xiuye Gu, Alex Hauptmann 0001, Boqing Gong, Ming-Hsuan Yang 0001, Irfan A. Essa, David A. Ross, Lu Jiang 0004
ICLR7
2024 VideoPoet: A Large Language Model for Zero-Shot Video Generation
abstract
We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004
ICML14
2023 Improving Multilingual and Code-Switching ASR Using Large Language Model Generated Text
abstract
We investigate using large language models (LLMs) to generate text-only training data for improving multilingual and code-switching automatic speech recognition (ASR) through a text injection method. In a multilingual setup or a low-resource scenario such as code-switching, we propose to generate text data using the state-of-the-art PaLM 2. To better match the generated text data with specific tasks, we use prompt tuning to adapt PaLM 2 to generate domain-relevant multilingual or code-switched text data for text injection. We can achieve significant improvements in Word Error Rate (WER) in both multilingual and code-switching scenarios. The multilingual experiment shows a $6.2 \%$ relative WER reduction on average, i.e., from $11.25 \%$ to $10.55 \%$, compared to a baseline without text injection. The improvement is up to $23.1 \%$ improvement for certain languages. While in the code-switching scenario, we use English-only prompts to generate Mandarin-English code-switching text and achieve a 3.6% relative WER reduction for a code-switching test set, as well as WER reductions in both English and Mandarin monolingual scenarios, $5.3 \%$ and $8.5 \%$ relative, respectively. Our findings demonstrate that leveraging LLMs for text generation and then injection benefits multilingual or code-switching ASR tasks.
Tara N. Sainath, Bo Li 0028, Yu Zhang 0033, Yong Cheng 0003, Frederick Liu
ASRU5
2023 MAGVIT: Masked Generative Video Transformer
abstract
We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method for masked video token modeling to facilitate multi-task learning. We conduct extensive experiments to demonstrate the quality, efficiency, and flexibility of MAGVIT. Our experiments show that (i) MAGVIT performs favorably against state-of-the-art approaches and establishes the best-published FVD on three video generation benchmarks, including the challenging Kinetics-600. (ii) MAGVIT outperforms existing methods in inference time by two orders of magnitude against diffusion models and by 60x against autoregressive models. (iii) A single MAGVIT model supports ten diverse generation tasks and generalizes across videos from different visual domains. The source code and trained models will be released to the public at https://magvit.cs.cmu.edu.
Lijun Yu, Yong Cheng 0003, Kihyuk Sohn, José Lezama, Han Zhang 0010, Huiwen Chang, Alex Hauptmann 0001, Ming-Hsuan Yang 0001, Yuan Hao, Irfan A. Essa, Lu Jiang 0004
CVPR2
2023 Mu2SLAM: Multitask, Multilingual Speech and Language Models
abstract
We present Mu$^2$SLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition (ASR), Automatic Speech Translation (AST) and Machine Translation (MT), in over 100 languages. By leveraging a quantized representation of speech as a target, Mu$^2$SLAM trains the speech-text models with a sequence-to-sequence masked denoising objective similar to T5 on the decoder and a masked language modeling objective (MLM) on the encoder, for both unlabeled speech and text, while utilizing the supervised tasks to improve cross-lingual and cross-modal representation alignment within the model. On CoVoST AST, Mu$^2$SLAM establishes a new state-of-the-art for models trained on public datasets, improving on xx-en translation over the previous best by 1.9 BLEU points and on en-xx translation by 1.1 BLEU points. On Voxpopuli ASR, our model matches the performance of an mSLAM model fine-tuned with an RNN-T decoder, despite using a relatively weaker Transformer decoder. On text understanding tasks, our model improves by more than 6% over mSLAM on XNLI, getting closer to the performance of mT5 models of comparable capacity on XNLI and TydiQA, paving the way towards a single model for all speech and text understanding tasks.
Yong Cheng 0003, Yu Zhang 0033, Melvin Johnson, Wolfgang Macherey, Ankur Bapna
ICML1
2023 SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
abstract
In this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM's vocabulary. The resulting tokens capture both the rich semantic meaning and the fine-grained details needed for visual reconstruction, effectively translating the visual content into a language comprehensible to the LLM, and empowering it to perform a wide array of multimodal tasks. Our approach is validated through in-context learning experiments with frozen PaLM 2 and GPT 3.5 on a diverse set of image understanding and generation tasks. Our method marks the first successful attempt to enable a frozen LLM to generate image content while surpassing state-of-the-art performance in image understanding tasks, under the same setting, by over 25%.
Lijun Yu, Yong Cheng 0003, Zhiruo Wang 0001, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan A. Essa, Yonatan Bisk, Ming-Hsuan Yang 0001, Kevin Murphy 0002, Alex Hauptmann 0001, Lu Jiang 0004
NeurIPS2
2022 Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine Translation
abstract
Multilingual neural machine translation models are trained to maximize the likelihood of a mix of examples drawn from multiple language pairs.The dominant inductive bias applied to these models is a shared vocabulary and a shared set of parameters across languages; the inputs and labels corresponding to examples drawn from different language pairs might still reside in distinct subspaces.In this paper, we introduce multilingual crossover encoder-decoder (mXEncDec) to fuse language pairs at an instance level.Our approach interpolates instances from different language pairs into joint 'crossover examples' in order to encourage sharing input and output spaces across languages.To ensure better fusion of examples in multilingual settings, we propose several techniques to improve example interpolation across dissimilar languages under heavy data imbalance.Experiments on a large-scale WMT multilingual dataset demonstrate that our approach significantly improves quality on English-to-Many, Many-to-English and zero-shot translation tasks (from +0.5 BLEU up to +5.5 BLEU points).Results on code-switching sets demonstrate the capability of our approach to improve model generalization to out-of-distribution multilingual examples.We also conduct qualitative and quantitative representation comparisons to analyze the advantages of our approach at the representation level.
Yong Cheng 0003, Ankur Bapna, Orhan Firat, Yuan Cao 0007, Pidong Wang, Wolfgang Macherey
ACL (1)1
2022 Examining Scaling and Transfer of Language Model Architectures for Machine Translation
abstract
Natural language understanding and generation models follow one of the two dominant architectural paradigms: language models (LMs) that process concatenated sequences in a single stack of layers, and encoder-decoder models (EncDec) that utilize separate layer stacks for input and output processing. In machine translation, EncDec has long been the favoured approach, but with few studies investigating the performance of LMs. In this work, we thoroughly examine the role of several architectural design choices on the performance of LMs on bilingual, (massively) multilingual and zero-shot translation tasks, under systematic variations of data conditions and model sizes. Our results show that: (i) Different LMs have different scaling properties, where architectural differences often have a significant impact on model performance at small scales, but the performance gap narrows as the number of parameters increases, (ii) Several design choices, including causal masking and language-modeling objectives for the source sequence, have detrimental effects on translation quality, and (iii) When paired with full-visible masking for source sequences, LMs could perform on par with EncDec on supervised bilingual and multilingual translation tasks, and improve greatly on zero-shot directions by facilitating the reduction of off-target translations.
Biao Zhang 0006, Behrooz Ghorbani, Ankur Bapna, Yong Cheng 0003, Xavier Garcia, Jonathan Shen, Orhan Firat
ICML4
2021 Self-supervised and Supervised Joint Training for Resource-rich Machine Translation
abstract
Self-supervised pre-training of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains on resource-rich NMT. In this paper, we propose a joint training approach, F2-XEnDec, to combine self-supervised and supervised learning to optimize NMT models. To exploit complementary self-supervised signals for supervised learning, NMT models are trained on examples that are interbred from monolingual and parallel sentences through a new process called crossover encoder-decoder. Experiments on two resource-rich translation benchmarks, WMT’14 English-German and WMT’14 English-French, demonstrate that our approach achieves substantial improvements over several strong baseline methods and obtains a new state of the art of 46.19 BLEU on English-French when incorporating back translation. Results also show that our approach is capable of improving model robustness to input perturbations such as code-switching noise which frequently appears on the social media.
Yong Cheng 0003, Lu Jiang 0004, Wolfgang Macherey
ICML1
2020 AdvAug: Robust Adversarial Augmentation for Neural Machine Translation
abstract
In this paper, we propose a new adversarial augmentation method for Neural Machine Translation (NMT).The main idea is to minimize the vicinal risk over virtual sentences sampled from two vicinity distributions, of which the crucial one is a novel vicinity distribution for adversarial sentences that describes a smooth interpolated embedding space centered around observed training sentence pairs.We then discuss our approach, AdvAug, to train NMT models using the embeddings of virtual sentences in sequence-tosequence learning.Experiments on Chinese-English, English-French, and English-German translation benchmarks show that AdvAug achieves significant improvements over the Transformer (up to 4.9 BLEU points), and substantially outperforms other data augmentation techniques (e.g.back-translation) without using extra corpora.
Yong Cheng 0003, Lu Jiang 0004, Wolfgang Macherey, Jacob Eisenstein
ACL1
2019 Robust Neural Machine Translation with Doubly Adversarial Inputs
abstract
Neural machine translation (NMT) often suffers from the vulnerability to noisy perturbations in the input.We propose an approach to improving the robustness of NMT models, which consists of two parts: (1) attack the translation model with adversarial source examples; (2) defend the translation model with adversarial target inputs to improve its robustness against the adversarial source inputs.For the generation of adversarial inputs, we propose a gradient-based method to craft adversarial examples informed by the translation loss over the clean inputs.Experimental results on Chinese-English and English-German translation tasks demonstrate that our approach achieves significant improvements (2.8 and 1.6 BLEU points) over Transformer on standard clean benchmarks as well as exhibiting higher robustness on noisy data.
Yong Cheng 0003, Lu Jiang 0004, Wolfgang Macherey
ACL (1)1
2019 Reducing Word Omission Errors in Neural Machine Translation: A Contrastive Learning Approach
abstract
While neural machine translation (NMT) has achieved remarkable success, NMT systems are prone to make word omission errors.In this work, we propose a contrastive learning approach to reducing word omission errors in NMT.The basic idea is to enable the NMT model to assign a higher probability to a ground-truth translation and a lower probability to an erroneous translation, which is automatically constructed from the ground-truth translation by omitting words.We design different types of negative examples depending on the number of omitted words, word frequency, and part of speech.Experiments on Chinese-to-English, German-to-English, and Russian-to-English translation tasks show that our approach is effective in reducing word omission errors and achieves better translation performance than three baseline methods.
Zonghan Yang, Yong Cheng 0003, Yang Liu 0005, Maosong Sun 0001
ACL (1)2
2019 An End-to-End Generative Architecture for Paraphrase Generation
abstract
Qian Yang, Zhouyuan Huo, Dinghan Shen, Yong Cheng, Wenlin Wang, Guoyin Wang, Lawrence Carin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qian Yang 0003, Zhouyuan Huo, Dinghan Shen, Yong Cheng 0003, Wenlin Wang, Guoyin Wang 0002, Lawrence Carin
EMNLP/IJCNLP (1)4
2018 Towards Robust Neural Machine Translation
abstract
Small perturbations in the input can severely distort intermediate representations and thus impact translation quality of neural machine translation (NMT) models.In this paper, we propose to improve the robustness of NMT models with adversarial stability training.The basic idea is to make both the encoder and decoder in NMT models robust against input perturbations by enabling them to behave similarly for the original input and its perturbed counterpart.Experimental results on Chinese-English, English-German and English-French translation tasks show that our approaches can not only achieve significant improvements over strong NMT systems but also improve the robustness of NMT models.
Yong Cheng 0003, Zhaopeng Tu, Fandong Meng, Junjie Zhai, Yang Liu 0005
ACL (1)1
2018 Neural Machine Translation with Key-Value Memory-Augmented Attention
abstract
Although attention-based Neural Machine Translation (NMT) has achieved remarkable progress in recent years, it still suffers from issues of repeating and dropping translations. To alleviate these issues, we propose a novel key-value memory-augmented attention model for NMT, called KVMEMATT. Specifically, we maintain a timely updated keymemory to keep track of attention history and a fixed value-memory to store the representation of source sentence throughout the whole translation process. Via nontrivial transformations and iterative interactions between the two memories, the decoder focuses on more appropriate source word(s) for predicting the next target word at each decoding step, therefore can improve the adequacy of translations. Experimental results on Chinese)English and WMT17 German,English translation tasks demonstrate the superiority of the proposed model.
Fandong Meng, Zhaopeng Tu, Yong Cheng 0003, Junjie Zhai, Yuekui Yang
IJCAI3
2017 Maximum Reconstruction Estimation for Generative Latent-Variable Models
Yong Cheng 0003, Yang Liu 0005, Wei Xu 0005
AAAI1
2017 A Teacher-Student Framework for Zero-Resource Neural Machine Translation
abstract
While end-to-end neural machine translation (NMT) has made remarkable progress recently, it still suffers from the data scarcity problem for low-resource language pairs and domains.In this paper, we propose a method for zero-resource NMT by assuming that parallel sentences have close probabilities of generating a sentence in a third language.Based on the assumption, our method is able to train a source-to-target NMT model ("student") without parallel corpora available guided by an existing pivot-to-target NMT model ("teacher") on a source-pivot parallel corpus.Experimental results show that the proposed method significantly improves over a baseline pivot-based model by +3.0 BLEU points across various language pairs.
Yun Chen 0007, Yang Liu 0005, Yong Cheng 0003, Victor O. K. Li
ACL (1)3
2017 Joint Training for Pivot-based Neural Machine Translation
abstract
While recent neural machine translation approaches have delivered state-of-the-art performance for resource-rich language pairs, they suffer from the data scarcity problem for resource-scarce language pairs. Although this problem can be alleviated by exploiting a pivot language to bridge the source and target languages, the source-to-pivot and pivot-to-target translation models are usually independently trained. In this work, we introduce a joint training algorithm for pivot-based neural machine translation. We propose three methods to connect the two models and enable them to interact with each other during training. Experiments on Europarl and WMT corpora show that joint training of source-to-pivot and pivot-to-target models leads to significant improvements over independent training across various languages.
Yong Cheng 0003, Qian Yang 0003, Yang Liu 0005, Maosong Sun 0001, Wei Xu 0005
IJCAI1
2017 Maximum Expected Likelihood Estimation for Zero-resource Neural Machine Translation
abstract
While neural machine translation (NMT) has made remarkable progress in translating a handful of high-resource language pairs recently, parallel corpora are not always available for many zero-resource language pairs. To deal with this problem, we propose an approach to zero-resource NMT via maximum expected likelihood estimation. The basic idea is to maximize the expectation with respect to a pivot-to-source translation model for the intended source-to-target model on a pivot-target parallel corpus. To approximate the expectation, we propose two methods to connect the pivot-to-source and source-to-target models. Experiments on two zero-resource language pairs show that the proposed approach yields substantial gains over baseline methods. We also observe that when trained jointly with the source-to-target model, the pivot-to-source translation model also obtains improvements over independent training.
Yong Cheng 0003, Yang Liu 0005
IJCAI2
2017 Neural Parse Combination
Liner Yang, Maosong Sun 0001, Yong Cheng 0003, Zhenghao Liu 0001, Huan-Bo Luan, Yang Liu 0005
J. Comput. Sci. Technol.3
2016 Semi-Supervised Learning for Neural Machine Translation
abstract
While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource languages, it is appealing to exploit monolingual corpora to improve NMT. We propose a semi-supervised approach for training NMT models on the concatenation of labeled (parallel corpora) and unlabeled (monolingual corpora) data. The central idea is to reconstruct the monolingual corpora using an autoencoder, in which the source-to-target and target-to-source translation models serve as the encoder and decoder, respectively. Our approach can not only exploit the monolingual corpora of the target language, but also of the source language. Experiments on the Chinese-English dataset show that our approach achieves significant improvements over state-of-the-art SMT and NMT systems.
Yong Cheng 0003, Wei Xu 0005, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005
ACL (1)1
2016 Minimum Risk Training for Neural Machine Translation
abstract
We propose minimum risk training for end-to-end neural machine translation.Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily differentiable.Experiments show that our approach achieves significant improvements over maximum likelihood estimation on a state-of-the-art neural machine translation system across various languages pairs.Transparent to architectures, our approach can be applied to more neural networks and potentially benefit more NLP tasks.
Shiqi Shen, Yong Cheng 0003, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005
ACL (1)2
2016 Agreement-Based Joint Training for Bidirectional Attention-Based Neural Machine Translation
Yong Cheng 0003, Shiqi Shen, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005
IJCAI1
2014 Query Lattice for Translation Retrieval
Meiping Dong, Yong Cheng 0003, Yang Liu 0005, Jia Xu 0004, Maosong Sun 0001, Tatsuya Izuha
COLING2