VLDB 2026 Research / reviewers in the wild / expert
Akshita Bhagia
dblp:321/0726
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 61% Efficient and distributed learning · 17% Deep learning architectures and training · 8% | |
| Network and information security
1 paper |
Authentication and access control · 100% |
Topics — the 22 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › neural language model
mixture-of-experts language model |
1.7 | 2 | 2025 | FlexOLMo: Open Language Models for Flexible Data Use · NeurIPS 2025 OLMoE: Open Mixture-of-Experts Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning
distributed training |
0.9 | 1 | 2025 | FlexOLMo: Open Language Models for Flexible Data Use · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training
pretraining data selection |
0.9 | 1 | 2025 | DataDecide: How to Predict Best Pretraining Data with Small Experiments · ICML 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | DataDecide: How to Predict Best Pretraining Data with Small Experiments · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression
sparse training |
0.9 | 1 | 2025 | OLMoE: Open Mixture-of-Experts Language Models · ICLR 2025 |
Natural language and speech › Information extraction and text analysis › corpus linguistics
corpus analysis |
0.8 | 1 | 2024 | What's In My Big Data? · ICLR 2024 |
Natural language and speech › Language models and text generation › large language model evaluation
data contamination |
0.8 | 1 | 2024 | What's In My Big Data? · ICLR 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | Paloma: A Benchmark for Evaluating Language Model Fit · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model
open language model development |
0.8 | 1 | 2024 | OLMo: Accelerating the Science of Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model training
pretraining corpus construction |
0.8 | 1 | 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › pre-training
pretraining data |
0.8 | 1 | 2024 | Paloma: A Benchmark for Evaluating Language Model Fit · NeurIPS 2024 |
Natural language and speech › Language models and text generation
training data analysis |
0.8 | 1 | 2024 | What's In My Big Data? · ICLR 2024 |
Natural language and speech › Language models and text generation
instruction tuning |
0.7 | 1 | 2023 | HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot Generalisation · ACL (1) 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.7 | 1 | 2023 | HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot Generalisation · ACL (1) 2023 |
Natural language and speech › Language models and text generation › in-context learning
few-shot prompting |
0.6 | 1 | 2022 | Continued Pretraining for Better Zero- and Few-Shot Promptability · EMNLP 2022 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.6 | 1 | 2022 | Continued Pretraining for Better Zero- and Few-Shot Promptability · EMNLP 2022 |
Natural language and speech › Language models and text generation
prompt tuning |
0.6 | 1 | 2022 | Continued Pretraining for Better Zero- and Few-Shot Promptability · EMNLP 2022 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.3 | 1 | 2025 | DataDecide: How to Predict Best Pretraining Data with Small Experiments · ICML 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2025 | OLMoE: Open Mixture-of-Experts Language Models · ICLR 2025 |
Authentication and access control › access control
data access control |
0.3 | 1 | 2025 | FlexOLMo: Open Language Models for Flexible Data Use · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.2 | 1 | 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024 |
Information retrieval › evaluation › benchmark
benchmark construction |
0.2 | 1 | 2024 | Paloma: A Benchmark for Evaluating Language Model Fit · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
nonparametric routing · 1.7model merging · 1.7mixture of experts · 1.7perplexity analysis · 1.5sparse mixture-of-experts · 0.9scaling law extrapolation · 0.9routing analysis · 0.9controlled pretraining experiments · 0.9benchmark prediction · 0.9corpus curation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OLMoE: Open Mixture-of-Experts Language ModelsabstractWe introduce OLMoE, a fully open, state-of-the-art language model leveraging sparse Mixture-of-Experts (MoE). OLMoE-1B-7B has 7 billion (B) parameters but uses only 1B per input token. We pretrain it on 5 trillion tokens and further adapt it to create OLMoE-1B-7B-Instruct. Our models outperform all available models with similar active parameters, even surpassing larger ones like Llama2-13B-Chat and DeepSeekMoE-16B. We present novel findings on MoE training, define and analyze new routing properties showing high specialization in our model, and open-source all our work: model weights, training data, code, and logs. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Pete Walsh 0001, Oyvind Tafjord, Nathan Lambert 0001, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, Dave Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi |
ICLR | 13 |
| 2025 | DataDecide: How to Predict Best Pretraining Data with Small ExperimentsabstractBecause large language models are expensive to pretrain on different datasets, using smaller-scale experiments to decide on data is crucial for reducing costs. Which benchmarks and methods of making decisions from observed performance at small scale most accurately predict the datasets that yield the best large models? To empower open exploration of this question, we release models, data, and evaluations in DataDecide—the most extensive open suite of models over differences in data and scale. We conduct controlled pretraining experiments across 25 corpora with differing sources, deduplication, and filtering up to 100B tokens, model sizes up to 1B parameters, and 3 random seeds. We find that the ranking of models at a single, small size (e.g., 150M parameters) is a strong baseline for predicting best models at our larger target scale (1B) ($\tilde$ 80% of comparisons correct). No scaling law methods among 8 baselines exceed the compute-decision frontier of single-scale predictions, but DataDecide can measure improvement in future scaling laws. We also identify that using continuous likelihood metrics as proxies in small experiments makes benchmarks including MMLU, ARC, HellaSwag, MBPP, and HumanEval $>$ 80% predictable at the target 1B scale with just 0.01% of the compute. Ian Magnusson, Nguyen Tai, Ben Bogin, David Heineman, Jena D. Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu 0010, Dirk Groeneveld, Oyvind Tafjord, Noah A. Smith, Pang Wei Koh, Jesse Dodge |
ICML | 7 |
| 2025 | FlexOLMo: Open Language Models for Flexible Data UseabstractWe introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference. Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min |
NeurIPS | 2 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 4 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 3 |
| 2024 | What's In My Big Data?abstractLarge text corpora are the backbone of language models.
However, we have a limited understanding of the content of these corpora, including general statistics, quality, social factors, and inclusion of evaluation data (contamination).
In this work, we propose What's In My Big Data? (WIMBD), a platform and a set of sixteen analyses that allow us to reveal and compare the contents of large text corpora. WIMBD builds on two basic capabilities---count and search---*at scale*, which allows us to analyze more than 35 terabytes on a standard compute node.
We apply WIMBD to ten different corpora used to train popular language models, including *C4*, *The Pile*, and *RedPajama*.
Our analysis uncovers several surprising and previously undocumented findings about these corpora, including the high prevalence of duplicate, synthetic, and low-quality content, personally identifiable information, toxic language, and benchmark contamination.
For instance, we find that about 50% of the documents in *RedPajama* and *LAION-2B-en* are duplicates. In addition, several datasets used for benchmarking models trained on such corpora are contaminated with respect to important benchmarks, including the Winograd Schema Challenge and parts of GLUE and SuperGLUE.
We open-source WIMBD's code and artifacts to provide a standard set of evaluations for new text-based corpora and to encourage more analyses and transparency around them. Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh 0001, Dirk Groeneveld, Luca Soldaini, Sameer Singh 0001, Hannaneh Hajishirzi, Noah A. Smith, Jesse Dodge |
ICLR | 2 |
| 2024 | Paloma: A Benchmark for Evaluating Language Model FitabstractEvaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Pete Walsh 0001, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson 0001, Jesse Dodge |
NeurIPS | 2 |
| 2023 | HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot GeneralisationabstractRecent NLP models have shown the remarkable ability to effectively generalise 'zero-shot' to new tasks using only natural language instructions as guidance.However, many of these approaches suffer from high computational costs due to their reliance on concatenating lengthy instructions with every input example, resulting in costly reprocessing of the instruction.To avoid this, we introduce Hypernetworks for INstruction Tuning (HINT), which convert task instructions and examples into parameter-efficient modules inserted into an underlying model using a pretrained text encoder, eliminating the need to include instructions in the model input.The hypernetwork in HINT also produces an encoded instruction, which we concatenate with encoded inputs during decoding to further improve performance.HINT models outperform strong state-of-theart baselines by over 10% when controlling for compute (measured in FLOPs).By converting instructions into modules, HINT models can effectively disregard the length of instructions and few-shot example inputs in terms of compute usage.As a result, HINT can enhance its performance by up to 25% by incorporating additional few-shot data, while utilizing only up to 5% more compute.This combines the strengths of parameter-efficient fine-tuning and in-context learning.We release our code publicly 1 . Hamish Ivison, Akshita Bhagia, Yizhong Wang, Hannaneh Hajishirzi, Matthew E. Peters |
ACL (1) | 2 |
| 2022 | Continued Pretraining for Better Zero- and Few-Shot PromptabilityabstractRecently introduced language model prompting methods can achieve high accuracy in zeroand few-shot settings while requiring few to no learned task-specific parameters.Nevertheless, these methods still often trail behind full model finetuning.In this work, we investigate if a dedicated continued pretraining stage could improve "promptability", i.e., zero-shot performance with natural language prompts or few-shot performance with prompt tuning.We reveal settings where existing continued pretraining methods lack promptability.We also identify current methodological gaps, which we fill with thorough large-scale experiments.We demonstrate that a simple recipe, continued pretraining that incorporates a trainable prompt during multi-task learning, leads to improved promptability in both zero-and fewshot settings compared to existing methods, up to 31% relative.On the other hand, we find that continued pretraining using MAML-style metalearning, a method that directly optimizes fewshot promptability, yields subpar performance.We validate our findings with two prompt tuning methods, and, based on our results, we provide concrete recommendations to optimize promptability for different use cases. Zhaofeng Wu, Robert L. Logan IV, Pete Walsh 0001, Akshita Bhagia, Dirk Groeneveld, Sameer Singh 0001, Iz Beltagy |
EMNLP | 4 |