VLDB 2026 Research / reviewers in the wild / expert
Ben Athiwaratkun
dblp:166/1659
· DBLP profile ↗
20ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0002-2009-496XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Instruction-tuned LLMs to Million-token Contexts via Hierarchical Synthetic Data GenerationabstractLarge Language Models (LLMs) struggle with long-context reasoning, not only due to the quadratic scaling of computational complexity with sequence length but also because of the scarcity and expense of annotating long-context data. There has been barely any open-source work that systematically ablates long-context data, nor is there any openly available instruction tuning dataset with contexts surpassing 100K tokens. To bridge this gap, we introduce a novel post-training synthetic data generation strategy designed to efficiently extend the context window of LLMs while preserving their general task performance. Our approach scalably extends to arbitrarily long context lengths, unconstrained by the length of available real-world data, which effectively addresses the scarcity of raw long-context data.
Through a step-by-step rotary position embedding (RoPE) scaling training strategy, we demonstrate that our model, with a context length of up to 1M tokens, performs well on the RULER benchmark and InfiniteBench and maintains robust performance on general language tasks. Linda He, Maurice Weber, Shang Zhu, Ben Athiwaratkun, Ce Zhang 0001 |
ICLR | 5 |
| 2025 | Training-Free Activation Sparsity in Large Language ModelsabstractActivation sparsity can enable practical inference speedups in large language models (LLMs) by reducing the compute and memory-movement required for matrix multiplications during the forward pass.
However, existing methods face limitations that inhibit widespread adoption. Some approaches are tailored towards older models with ReLU-based sparsity, while others require extensive continued pre-training on up to hundreds of billions of tokens.
This paper describes TEAL (**T**raining-Fre**e** **A**ctivation Sparsity in **L**LMs), a simple training-free method that applies magnitude-based activation sparsity to hidden states throughout the entire model. TEAL achieves 40-50\% model-wide sparsity with minimal performance degradation across Llama-2, Llama-3, and Mistral families, with sizes varying from 7B to 70B. We improve existing sparse kernels and demonstrate wall-clock decoding speed-ups of up to 1.53× and 1.8× at 40\% and 50\% model-wide sparsity. TEAL is compatible with weight quantization, enabling further efficiency gains. James Liu, Pragaash Ponnusamy, Tianle Cai, Ben Athiwaratkun |
ICLR | 6 |
| 2025 | Mixture-of-Agents Enhances Large Language Model CapabilitiesabstractRecent advances in large language models (LLMs) demonstrate substantial capabilities in natural language understanding and generation tasks. With the growing number of LLMs, how to harness the collective expertise of multiple LLMs is an exciting open direction. Toward this goal, we propose a new approach that leverages the collective strengths of multiple LLMs through a Mixture-of-Agents (MoA) methodology. In our approach, we construct a layered MoA architecture wherein each layer comprises multiple LLM agents. Each agent takes all the outputs from agents in the previous layer as auxiliary information in generating its response. MoA models achieves state-of-art performance on AlpacaEval 2.0, Arena-Hard, MT-Bench, and FLASK, surpassing GPT-4 Omni. For example, our MoA using only open-source LLMs achieves a score of 65.1% on AlpacaEval 2.0 compared to 57.5% by GPT-4 Omni. Ben Athiwaratkun, Ce Zhang 0001, James Zou 0001 |
ICLR | 3 |
| 2025 | Improving Model Alignment Through Collective Intelligence of Open-Source ModelsabstractBuilding helpful and harmless large language models (LLMs) requires effective model alignment approach based on human instructions and feedback, which necessitates high-quality human-labeled data. Constructing such datasets is often expensive and hard to scale, and may face potential limitations on diversity and generalization. To address these challenges, we introduce Mixture of Agents Alignment (MoAA), that leverages the collective strengths of various language models to provide high-quality data for model alignment. By employing MoAA, we enhance both supervised fine-tuning and preference optimization, leading to improved performance compared to using a single model alone to generate alignment data (e.g. using GPT-4o alone). Evaluation results show that our approach can improve win rate of LLaMA-3.1-8B-Instruct from 19.5 to 48.3 on Arena-Hard and from 22.33 to 57.23 on AlpacaEval2, highlighting a promising direction for model alignment through this new scalable and diverse synthetic data recipe. Furthermore, we demonstrate that MoAA enables a self-improvement pipeline, where models fine-tuned on MoA-generated data surpass their own initial capabilities, providing evidence that our approach can push the frontier of open-source LLMs without reliance on stronger external supervision. Data and code will be released. Roy Xie, Shang Zhu, Ben Athiwaratkun, Bhuwan Dhingra, Shuaiwen Song, Ce Zhang 0001, James Zou 0001 |
ICML | 5 |
| 2025 | Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication OverlappingabstractLarge language model inference is both memory-intensive and time-consuming, often requiring distributed algorithms to efficiently scale. Various model parallelism strategies are used in multi-gpu training and inference to partition computation across multiple devices, reducing memory load and computation time. However, using model parallelism necessitates communication of information between GPUs, which has been a major bottleneck and limits the gains obtained by scaling up the number of devices. We introduce Ladder Residual, a simple architectural modification applicable to all residual-based models that enables straightforward overlapping that effectively hides the latency of communication. Our insight is that in addition to systems optimization, one can also redesign the model architecture to decouple communication from computation. While Ladder Residual can allow communication-computation decoupling in conventional parallelism patterns, we focus on Tensor Parallelism in this paper, which is particularly bottlenecked by its heavy communication. For a Transformer model with 70B parameters, applying Ladder Residual to all its layers can achieve 29% end-to-end wall clock speed up at inference time with TP sharding over 8 devices. We refer the resulting Transformer model as the Ladder Transformer. We train a 1B and 3B Ladder Transformer from scratch and observe comparable performance to a standard dense transformer baseline. We also show that it is possible to convert parts of the Llama-3.1 8B model to our Ladder Residual architecture with minimal accuracy degradation by only retraining for 3B tokens. Muru Zhang, Zhongzhu Zhou, William Brandon, Jonathan Ragan-Kelley, Shuaiwen Song, Ben Athiwaratkun, Tri Dao |
ICML | 9 |
| 2025 | Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for VerificationabstractVerifiers can improve language model (LM) capabilities by providing feedback or selecting the best response from a pool of generated candidates. Currently, high-quality verifiers are either unscalable (e.g., humans) or limited in utility (e.g., tools like Lean for formal proofs). While LM judges and reward models have become broadly useful as general-purpose verifiers, a significant performance gap remains between them and oracle verifiers. To help close this gap, we introduce Weaver, a framework for designing a strong verifier by combining multiple weak, imperfect verifiers. First we find that weighted ensembles of verifiers, which typically require learning from labeled data, significantly outperform unweighted combinations due to differences in the verifiers. To reduce the dependency on labeled data, Weaver leverages weak supervision to estimate each verifier’s accuracy and combines their outputs into a unified score that better reflects true response quality. However, directly applying weak supervision algorithms poses several challenges, including inconsistent verifier output formats and handling low-quality verifiers. Weaver addresses these challenges by using dataset statistics to normalize outputs and filter specific verifiers. We study the effectiveness of Weaver in repeated sampling settings, where a model generates multiple candidate responses at test time and a verifier is used to select the correct one. Our evaluations demonstrate that Weaver significantly improves the pass@1 performance across several reasoning and math tasks, achieving o3-mini level accuracy with Llama 3.3 70B Instruct (a much cheaper non-reasoning model) as the generator, and an ensemble of smaller judge and reward models as the verifiers (86.2% average). This gain mirrors the jump achieved between GPT-4o and o3-mini (69.0% vs. 86.7%), which required extensive finetuning and post-training interventions. To make Weaver more efficient, we train a compact 400M cross-encoder using Weaver's combined output scores. This distilled model retains 98.7% of Weaver's full accuracy while reducing verification compute by up to 99.97%. Jon Saad-Falcon, Estefany Kelly Buchanan, Mayee F. Chen, Tzu-Heng Huang, Brendan McLaughlin, Tanvir Bhathal, Shang Zhu, Ben Athiwaratkun, Frederic Sala, Scott W. Linderman, Azalia Mirhoseini, Christopher Ré |
NeurIPS | 8 |
| 2024 | Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning StrategiesabstractA diverse array of reasoning strategies has been proposed to elicit the capabilities of large language models.However, in this paper, we point out that traditional evaluations which focus solely on performance metrics miss a key factor: the increased effectiveness due to additional compute.By overlooking this aspect, a skewed view of strategy efficiency is often presented.This paper introduces a framework that incorporates the compute budget into the evaluation, providing a more informative comparison that takes into account both performance metrics and computational cost.In this budgetaware perspective, we find that complex reasoning strategies often don't surpass simpler baselines purely due to algorithmic ingenuity, but rather due to the larger computational resources allocated.When we provide a simple baseline like chain-of-thought self-consistency with comparable compute resources, it frequently outperforms reasoning strategies proposed in the literature.In this scale-aware perspective, we find that unlike self-consistency, certain strategies such as multi-agent debate or Reflexion can become worse if more compute budget is utilized. Siddhartha Jain 0001, Dejiao Zhang, Baishakhi Ray, Ben Athiwaratkun |
EMNLP | 6 |
| 2024 | Bifurcated Attention for Single-Context Large-Batch SamplingabstractIn our study, we present bifurcated attention, a method developed for language model inference in single-context batch sampling contexts. This approach aims to reduce redundant memory IO costs, a significant factor in latency for high batch sizes and long context lengths. Bifurcated attention achieves this by dividing the attention mechanism during incremental decoding into two distinct GEMM operations, focusing on the KV cache from prefill and the decoding process. This method ensures precise computation and maintains the usual computational load (FLOPs) of standard attention mechanisms, but with reduced memory IO. Bifurcated attention is also compatible with multi-query attention mechanism known for reduced memory IO for KV cache, further enabling higher batch size and context length. The resulting efficiency leads to lower latency, improving suitability for real-time applications, e.g., enabling massively-parallel answer generation without substantially increasing latency, enhancing performance when integrated with post-processing techniques such as reranking. Ben Athiwaratkun, Sujan K. Gonugondla, Sanjay Krishna Gouda, Haifeng Qian, Hantian Ding, Qing Sun 0013, Jun Wang 0022, Jiacheng Guo, Liangfu Chen, Parminder Bhatia, Ramesh Nallapati, Sudipta Sengupta, Bing Xiang |
ICML | 1 |
| 2024 | RedPajama: an Open Dataset for Training Large Language ModelsabstractLarge language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset composition and filtering remain largely elusive. Many of the top-performing models lack transparency in their dataset curation and model development processes, posing an obstacle to the development of fully open language models. In this paper, we identify three core data-related challenges that must be addressed to advance open-source language models. These include (1) transparency in model development, including the data curation process, (2) access to large quantities of high-quality data, and (3) availability of artifacts and metadata for dataset curation and analysis. To address these challenges, we release RedPajama-V1, an open reproduction of the LLaMA training dataset. In addition, we release RedPajama-V2, a massive web-only dataset consisting of raw, unfiltered text data together with quality signals and metadata.Together, the RedPajama datasets comprise over 100 trillion tokens spanning multiple domains and with their quality signals facilitate the filtering of data, aiming to inspire the development of numerous new datasets. To date, these datasets have already been used in the training of strong language models used in production, such as Snowflake Arctic, Salesforce's XGen and AI2's OLMo. To provide insight into the quality of RedPajama, we present a series of analyses and ablation studies with decoder-only language models with up to 1.6B parameters. Our findings demonstrate how quality signals for web data can be effectively leveraged to curate high-quality subsets of the dataset, underscoring the potential of RedPajama to advance the development of transparent and high-performing language models at scale. Maurice Weber, Daniel Y. Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, Ce Zhang 0001 |
NeurIPS | 11 |
| 2023 | Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati |
ICLR | 1 |
| 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical StudyabstractML-powered code generation aims to assist developers to write code in a more productive manner by intelligently generating code blocks based on natural language prompts. Recently, large pretrained deep learning models have pushed the boundary of code generation and achieved impressive performance. However, the huge number of model parameters poses a significant challenge to their adoption in a typical software development environment, where a developer might use a standard laptop or mid-size server to develop code. Such large models cost significant resources in terms of memory, latency, dollars, as well as carbon footprint. Xiaokai Wei, Sujan K. Gonugondla, Shiqi Wang 0002, Wasi Uddin Ahmad, Baishakhi Ray, Haifeng Qian, Xiaopeng Li 0002, Zijian Wang 0002, Qing Sun 0013, Ben Athiwaratkun, Mingyue Shang, Murali Krishna Ramanathan, Parminder Bhatia, Bing Xiang |
ESEC/SIGSOFT FSE | 12 |
| 2021 | Generative Context Pair Selection for Multi-hop Question AnsweringabstractDheeru Dua, Cicero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner, Sameer Singh. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Dheeru Dua, Cícero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner 0001, Sameer Singh 0001 |
EMNLP (1) | 4 |
| 2021 | Structured Prediction as Translation between Augmented Natural Languages
Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, Stefano Soatto |
ICLR | 2 |
| 2020 | Augmented Natural Language for Generative Sequence LabelingabstractWe propose a generative framework for joint sequence labeling and sentence-level classification.Our model performs multiple sequence labeling tasks at once using a single, shared natural language output space.Unlike prior discriminative methods, our model naturally incorporates label semantics and shares knowledge across tasks.Our framework is general purpose, performing well on fewshot, low-resource, and high-resource tasks.We demonstrate these advantages on popular named entity recognition, slot labeling, and intent classification benchmarks.We set a new state-of-the-art for few-shot slot labeling, improving substantially upon the previous 5-shot (75.0%!90.9%) and 1-shot (70.4% !81.0%) state-of-the-art results.Furthermore, our model generates large improvements (46.27% !63.83%) in low-resource slot labeling over a BERT baseline by incorporating label semantics.We also maintain competitive results on high-resource tasks, performing within two points of the state-of-theart on all tasks and setting a new state-of-theart on the SNIPS dataset. Ben Athiwaratkun, Cícero Nogueira dos Santos, Jason Krone, Bing Xiang |
EMNLP (1) | 1 |
| 2019 | There Are Many Consistent Explanations of Unlabeled Data: Why You Should Average
Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, Andrew Gordon Wilson |
ICLR (Poster) | 1 |
| 2018 | Probabilistic FastText for Multi-Sense Word EmbeddingsabstractWe introduce Probabilistic FastText, a new model for word embeddings that can capture multiple word senses, sub-word structure, and uncertainty information.In particular, we represent each word with a Gaussian mixture density, where the mean of a mixture component is given by the sum of n-grams.This representation allows the model to share statistical strength across sub-word structures (e.g.Latin roots), producing accurate representations of rare, misspelt, or even unseen words.Moreover, each component of the mixture can capture a different word sense.Probabilistic FastText outperforms both FASTTEXT, which has no probabilistic model, and dictionary-level probabilistic embeddings, which do not incorporate subword structures, on several word-similarity benchmarks, including English RareWord and foreign language datasets.We also achieve state-ofart performance on benchmarks that measure ability to discern different meanings.Thus, the proposed model is the first to achieve multi-sense representations while having enriched semantics on rare words. Ben Athiwaratkun, Andrew Gordon Wilson, Anima Anandkumar |
ACL (1) | 1 |
| 2018 | Hierarchical Density Order Embeddings
Ben Athiwaratkun, Andrew Gordon Wilson |
ICLR (Poster) | 1 |
| 2018 | Adversarial Deep Averaging Networks for Cross-Lingual Sentiment ClassificationabstractIn recent years great success has been achieved in sentiment classification for English, thanks in part to the availability of copious annotated resources. Unfortunately, most languages do not enjoy such an abundance of labeled data. To tackle the sentiment classification problem in low-resource languages without adequate annotated data, we propose an Adversarial Deep Averaging Network (ADAN 1 ) to transfer the knowledge learned from labeled data on a resource-rich source language to low-resource languages where only unlabeled data exist. ADAN has two discriminative branches: a sentiment classifier and an adversarial language discriminator. Both branches take input from a shared feature extractor to learn hidden representations that are simultaneously indicative for the classification task and invariant across languages. Experiments on Chinese and Arabic sentiment classification demonstrate that ADAN significantly outperforms state-of-the-art systems. Xilun Chen 0002, Yu Sun 0020, Ben Athiwaratkun, Claire Cardie, Kilian Q. Weinberger |
Trans. Assoc. Comput. Linguistics | 3 |
| 2017 | Multimodal Word DistributionsabstractWord embeddings provide point representations of words containing useful semantic information.We introduce multimodal word distributions formed from Gaussian mixtures, for multiple word meanings, entailment, and rich uncertainty information.To learn these distributions, we propose an energy-based max-margin objective.We show that the resulting approach captures uniquely expressive semantic information, and outperforms alternatives, such as word2vec skip-grams, and Gaussian embeddings, on benchmark datasets such as word similarity and entailment. Ben Athiwaratkun, Andrew Gordon Wilson |
ACL (1) | 1 |
| 2017 | Malware classification with LSTM and GRU language models and a character-level CNNabstractMalicious software, or malware, continues to be a problem for computer users, corporations, and governments. Previous research [1] has explored training file-based, malware classifiers using a two-stage approach. In the first stage, a malware language model is used to learn the feature representation which is then input to a second stage malware classifier. In Pascanu et al. [1], the language model is either a standard recurrent neural network (RNN) or an echo state network (ESN). In this work, we propose several new malware classification architectures which include a long short-term memory (LSTM) language model and a gated recurrent unit (GRU) language model. We also propose using an attention mechanism similar to [12] from the machine translation literature, in addition to temporal max pooling used in [1], as an alternative way to construct the file representation from neural features. Finally, we propose a new single-stage malware classifier based on a character-level convolutional neural network (CNN). Results show that the LSTM with temporal max pooling and logistic regression offers a 31.3% improvement in the true positive rate compared to the best system in [1] at a false positive rate of 1%. Ben Athiwaratkun, Jack W. Stokes |
ICASSP | 1 |