EDBT 2026 Demo / reviewers in the wild / expert
Chunting Zhou
dblp:161/2679
· DBLP profile ↗
26ranked-venue papers
7as first author
17since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 17 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Byte Latent Transformer: Patches Scale Better Than TokensabstractArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini Iyer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez 0001, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer 0001 |
ACL (1) | 7 |
| 2025 | Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal ModelabstractWe introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data.
Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences.
We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks.
Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens.
By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches.
We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds. Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, Omer Levy |
ICLR | 1 |
| 2025 | CAT: Content-Adaptive Image TokenizationabstractMost existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity and introducing unnecessary computate overhead for simpler images. To address this, we propose Content-Adaptive Tokenizer (CAT), which dynamically adjusts representation capacity based on the image content and encodes simpler images into fewer tokens. We design (1) a caption-based evaluation system that leverages LLMs to predict content complexity and determine the optimal compression ratio for an image, and (2) a novel nested VAE architecture that performs variable-rate compression in a single model.
Trained on images with varying complexity, CAT achieves an average of 15% reduction in rFID across seven detail-rich datasets containing text, humans, and complex textures. On natural image datasets like ImageNet and COCO, it reduces token usage by 18% while maintaining high-fidelity reconstructions. We further evaluate CAT on two downstream tasks. For image classification, CAT consistently improves top-1 accuracy across five datasets spanning diverse domains. For image generation, it boosts training throughput by 23% on ImageNet, leading to more efficient learning and improved FIDs over fixed-token baselines. Junhong Shen, Kushal Tirumala, Michihiro Yasunaga, Ishan Misra, Luke Zettlemoyer, Lili Yu, Chunting Zhou |
NeurIPS | 7 |
| 2025 | LMFusion: Adapting Pretrained Language Models for Multimodal GenerationabstractWe present LMFusion, a framework for empowering pretrained text-only large language models (LLMs) with multimodal generative capabilities, enabling them to understand and generate both text and images in arbitrary sequences. LMFusion leverages existing Llama-3's weights for processing texts autoregressively while introducing additional and parallel transformer modules for processing images with diffusion. During training, the data from each modality is routed to its dedicated modules: modality-specific feedforward layers, query-key-value projections, and normalization layers process each modality independently, while the shared self-attention layers allow interactions across text and image features. By freezing the text-specific modules and only training the image-specific modules, LMFusion preserves the language capabilities of text-only LLMs while developing strong visual understanding and generation abilities. Compared to methods that pretrain multimodal generative models from scratch, our experiments demonstrate that, LMFusion improves image understanding by 20% and image generation by 3.6% using only 50% of the FLOPs while maintaining Llama-3's language capabilities. We also demonstrate that this framework can adapt existing vision-language models with multimodal generation ability. Overall, this framework not only leverages existing computational investments in text-only LLMs but also enables the parallel development of language and vision capabilities, presenting a promising direction for efficient multimodal model development. Xiaochuang Han, Chunting Zhou, Weixin Liang, Xi Victoria Lin, Luke Zettlemoyer, Lili Yu |
NeurIPS | 3 |
| 2024 | Instruction-tuned Language Models are Better Knowledge LearnersabstractZhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, Srini Iyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhengbao Jiang, Zhiqing Sun, Pedro Rodríguez 0001, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Scott Yih, Srinivasan Iyer 0001 |
ACL (1) | 5 |
| 2024 | Self-Alignment with Instruction BacktranslationabstractWe present a scalable method to build a high quality instruction following language model by automatically labelling human-written text with corresponding instructions. Our approach, named instruction backtranslation, starts with a language model finetuned on a small amount of seed data, and a given web corpus. The seed model is used to construct training examples by generating instruction prompts for web documents (self-augmentation), and then selecting high quality examples from among these candidates (self-curation). This data is then used to finetune a stronger model. Finetuning LLaMa on two iterations of our approach yields a model that outperforms all other LLaMa-based models on the Alpaca leaderboard not relying on distillation data, demonstrating highly effective self-alignment. Xian Li 0003, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis |
ICLR | 3 |
| 2024 | In-Context Pretraining: Language Modeling Beyond Document BoundariesabstractLanguage models are currently trained to predict tokens given document prefixes, enabling them to zero shot long form generation and prompting-style tasks which can be reduced to document completion. We instead present IN-CONTEXT PRETRAINING, a new approach where language models are trained on a sequence of related documents, thereby explicitly encouraging them to read and reason across document boundaries. Our approach builds on the fact that current pipelines train by concatenating random sets of shorter documents to create longer context windows; this improves efficiency even though the prior documents provide no signal for predicting the next document. Given this fact, we can do IN-CONTEXT PRETRAINING by simply changing the document ordering so that each context contains related documents, and directly applying existing pretraining pipelines. However, this document sorting problem is challenging. There are billions of documents and we would like the sort to maximize contextual similarity for every document without repeating any data. To do this, we introduce approximate algorithms for finding related documents with efficient nearest neighbor search and constructing coherent batches with a graph cover algorithm. Our experiments show IN-CONTEXT PRETRAINING offers a scalable and simple approach to significantly enhance LM performance: we see notable improvements in tasks that require more complex contextual reasoning, including in-context learning (+8%), reading comprehension (+15%), faithfulness to previous contexts (+16%), long-context reasoning (+5%), and retrieval augmentation (+9%). Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Scott Yih, Mike Lewis |
ICLR | 4 |
| 2024 | MART: Improving LLM Safety with Multi-round Automatic Red-TeamingabstractSuyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, Yuning Mao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Suyu Ge, Chunting Zhou, Madian Khabsa, Yi-Chia Wang, Qifan Wang 0001, Jiawei Han 0001, Yuning Mao |
NAACL-HLT | 2 |
| 2024 | Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context LengthabstractThe quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce MEGALODON, an neural architecture for efficient sequence modeling with unlimited context length. MEGALODON inherits the architecture of MEGA (exponential moving average with gated attention), and further introduces multiple technical components to improve its capability and stability, including complex exponential moving average (CEMA), timestep normalization layer, normalized attention mechanism and pre-norm with two-hop residual configuration. In a controlled head-to-head comparison with LLAMA2, MEGALODON achieves better efficiency than Transformer in the scale of 7 billion parameters and 2 trillion training tokens. MEGALODON reaches a training loss of 1.70, landing mid-way between LLAMA2-7B (1.75) and LLAMA2-13B (1.67). This result is robust throughout a wide range of benchmarks, where MEGALODON consistently outperforms Transformers across different tasks, domains, and modalities. Xuezhe Ma, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang 0025, Jonathan May, Luke Zettlemoyer, Omer Levy, Chunting Zhou |
NeurIPS | 10 |
| 2023 | Training Trajectories of Language Models Across ScalesabstractMengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen 0001, Luke Zettlemoyer, Veselin Stoyanov |
ACL (1) | 3 |
| 2023 | Look-back Decoding for Open-Ended Text GenerationabstractGiven a prefix (context), open-ended generation aims to decode texts that are coherent, which do not abruptly drift from previous topics, and informative, which do not suffer from undesired repetitions.In this paper, we propose Look-back , an improved decoding algorithm that leverages the Kullback-Leibler divergence to track the distribution distance between current and historical decoding steps.Thus Lookback can automatically predict potential repetitive phrase and topic drift, and remove tokens that may cause the failure modes, restricting the next token probability distribution within a plausible distance to the history.We perform decoding experiments on document continuation and story generation, and demonstrate that Look-back is able to generate more fluent and coherent text, outperforming other strong decoding methods significantly in both automatic and human evaluations 1 . Chunting Zhou, Asli Celikyilmaz, Xuezhe Ma |
EMNLP | 2 |
| 2023 | Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, Luke Zettlemoyer |
ICLR | 2 |
| 2023 | LIMA: Less Is More for AlignmentabstractLarge language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and reinforcement learning, to better align to end tasks and user preferences.
We measure the relative importance of these two stages by training LIMA, a 65B parameter LLaMa language model fine-tuned with the standard supervised loss on only 1,000 carefully curated prompts and responses, without any reinforcement learning or human preference modeling.
LIMA demonstrates remarkably strong performance, learning to follow specific response formats from only a handful of examples in the training data, including complex queries that range from planning trip itineraries to speculating about alternate history.
Moreover, the model tends to generalize well to unseen tasks that did not appear in the training data.
In a controlled human study, responses from LIMA are either equivalent or strictly preferred to GPT-4 in 43\% of cases; this statistic is as high as 58\% when compared to Bard and 65\% versus DaVinci003, which was trained with human feedback.
Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output. Chunting Zhou, Puxin Xu, Srinivasan Iyer 0001, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Lili Yu, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy |
NeurIPS | 1 |
| 2022 | Towards a Unified View of Parameter-Efficient Transfer Learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, Graham Neubig |
ICLR | 2 |
| 2021 | Distributionally Robust Multilingual Machine TranslationabstractMultilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, potentially improving both the accuracy and the memoryefficiency of deployed models.However, the heavy data imbalance between languages hinders the model from performing uniformly across language pairs.In this paper, we propose a new learning objective for MNMT based on distributionally robust optimization, which minimizes the worst-case expected loss over the set of language pairs.We further show how to practically optimize this objective for large translation corpora using an iterated best response scheme, which is both effective and incurs negligible additional computational cost compared to standard empirical risk minimization.We perform extensive experiments on three sets of languages from two datasets and show that our method consistently outperforms strong baseline methods in terms of average and per-language performance under both many-to-one and one-to-many translation settings.1 Chunting Zhou, Daniel Levy 0002, Xian Li 0003, Marjan Ghazvininejad, Graham Neubig |
EMNLP (1) | 1 |
| 2021 | Examining and Combating Spurious Features under Distribution ShiftabstractA central goal of machine learning is to learn robust representations that capture the fundamental relationship between inputs and output labels. However, minimizing training errors over finite or biased datasets results in models latching on to spurious correlations between the training input/output pairs that are not fundamental to the problem at hand. In this paper, we define and analyze robust and spurious representations using the information-theoretic concept of minimal sufficient statistics. We prove that even when there is only bias of the input distribution (i.e. covariate shift), models can still pick up spurious features from their training data. Group distributionally robust optimization (DRO) provides an effective tool to alleviate covariate shift by minimizing the worst-case training losses over a set of pre-defined groups. Inspired by our analysis, we demonstrate that group DRO can fail when groups do not directly account for various spurious correlations that occur in the data. To address this, we further propose to minimize the worst-case losses over a more flexible set of distributions that are defined on the joint distribution of groups and instances, instead of treating each group as a whole at optimization time. Through extensive experiments on one image and two language tasks, we show that our model is significantly more robust than comparable baselines under various partitions. Chunting Zhou, Xuezhe Ma, Paul Michel, Graham Neubig |
ICML | 1 |
| 2021 | Luna: Linear Unified Nested AttentionabstractThe quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Luna, a linear unified nested attention mechanism that approximates softmax attention with two nested linear attention functions, yielding only linear (as opposed to quadratic) time and space complexity. Specifically, with the first attention function, Luna packs the input sequence into a sequence of fixed length. Then, the packed sequence is unpacked using the second attention function. As compared to a more traditional attention mechanism, Luna introduces an additional sequence with a fixed length as input and an additional corresponding output, which allows Luna to perform attention operation linearly, while also storing adequate contextual information. We perform extensive evaluations on three benchmarks of sequence modeling tasks: long-context sequence modelling, neural machine translation and masked language modeling for large-scale pretraining. Competitive or even better experimental results demonstrate both the effectiveness and efficiency of Luna compared to a variety of strong baseline methods including the full-rank attention and other efficient sparse and dense attention methods. Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma 0001, Luke Zettlemoyer |
NeurIPS | 4 |
| 2020 | Understanding Knowledge Distillation in Non-autoregressive Machine Translation
Chunting Zhou, Jiatao Gu, Graham Neubig |
ICLR | 1 |
| 2019 | FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative FlowabstractXuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xuezhe Ma, Chunting Zhou, Xian Li 0003, Graham Neubig, Eduard H. Hovy |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Handling Syntactic Divergence in Low-resource Machine TranslationabstractChunting Zhou, Xuezhe Ma, Junjie Hu, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunting Zhou, Xuezhe Ma, Junjie Hu 0001, Graham Neubig |
EMNLP/IJCNLP (1) | 1 |
| 2019 | MAE: Mutual Posterior-Divergence Regularization for Variational AutoEncoders
Xuezhe Ma, Chunting Zhou, Eduard H. Hovy |
ICLR (Poster) | 2 |
| 2018 | StructVAE: Tree-structured Latent Variable Models for Semi-supervised Semantic ParsingabstractSemantic parsing is the task of transducing natural language (NL) utterances into formal meaning representations (MRs), commonly represented as tree structures.Annotating NL utterances with their corresponding MRs is expensive and timeconsuming, and thus the limited availability of labeled data often becomes the bottleneck of data-driven, supervised models.We introduce STRUCTVAE, a variational auto-encoding model for semisupervised semantic parsing, which learns both from limited amounts of parallel data, and readily-available unlabeled NL utterances.STRUCTVAE models latent MRs not observed in the unlabeled data as treestructured latent variables.Experiments on semantic parsing on the ATIS domain and Python code generation show that with extra unlabeled data, STRUCTVAE outperforms strong supervised models. 1 Chunting Zhou, Junxian He, Graham Neubig |
ACL (1) | 2 |
| 2018 | Adapting Word Embeddings to New Languages with Morphological and Phonological Subword RepresentationsabstractMuch work in Natural Language Processing (NLP) has been for resource-rich languages, making generalization to new, less-resourced languages challenging.We present two approaches for improving generalization to lowresourced languages by adapting continuous word representations using linguistically motivated subword units: phonemes, morphemes and graphemes.Our method requires neither parallel corpora nor bilingual dictionaries and provides a significant gain in performance over previous methods relying on these resources.We demonstrate the effectiveness of our approaches on Named Entity Recognition for four languages, namely Uyghur, Turkish, Bengali and Hindi, of which Uyghur and Bengali are low resource languages, and also perform experiments on Machine Translation.Exploiting subwords with transfer learning gives us a boost of +15.2 NER F1 for Uyghur and +9.7 F1 for Bengali.We also show improvements in the monolingual setting where we achieve (avg.)+3 F1 and (avg.)+1.35 BLEU. Aditi Chaudhary, Chunting Zhou, Lori S. Levin, Graham Neubig, David R. Mortensen, Jaime G. Carbonell |
EMNLP | 2 |
| 2017 | Multi-space Variational Encoder-Decoders for Semi-supervised Labeled Sequence TransductionabstractLabeled sequence transduction is a task of transforming one sequence into another sequence that satisfies desiderata specified by a set of labels.In this paper we propose multi-space variational encoderdecoders, a new model for labeled sequence transduction with semi-supervised learning.The generative model can use neural networks to handle both discrete and continuous latent variables to exploit various features of data.Experiments show that our model provides not only a powerful supervised framework but also can effectively take advantage of the unlabeled data.On the SIGMORPHON morphological inflection benchmark, our model outperforms single-model state-ofart results by a large margin for the majority of languages.1 Chunting Zhou, Graham Neubig |
ACL (1) | 1 |
| 2015 | Optimized Periodical Charging in Large-Scale Deployed WSNsabstractRestricted by finite battery energy, traditional wireless sensor networks (WSNs) can only maintain for a limited period of time, resulting in serious performance bottleneck in long-term deployment of WSN. Fortunately, the advancement in the wireless energy transfer technology provides a potential to free WSNs from limited energy supply and remain perpetual operational. A mobile charger called wireless charging vehicle (WCV) is employed to periodically charge each sensor node and keep its energy level above the minimum threshold. Aiming at maximizing the ratio of the WCV's vocation time over the cycle time as well as guaranteeing the perpetual operation of networks, we proposes a feasible and optimal solution to this issue within the context of a real-time large-scale deployed WSN. Zhenquan Qin, Bingxian Lu, Chunting Zhou, Lei Wang 0005, Ming Zhu 0001, Lei Shu 0001 |
GLOBECOM | 3 |
| 2015 | Efficient Methods for Multi-label Classification
Chonglin Sun, Chunting Zhou, Bo Jin 0001, Francis C. M. Lau 0001 |
PAKDD (1) | 2 |