VLDB 2026 Research / reviewers in the wild / expert
Jacob Morrison
dblp:306/7813
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0001-8592-4744ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 70% Efficient and distributed learning · 23% Information extraction and text analysis · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 60% Environmental and earth informatics · 40% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › neural language model
mixture-of-experts language model |
1.7 | 2 | 2025 | FlexOLMo: Open Language Models for Flexible Data Use · NeurIPS 2025 OLMoE: Open Mixture-of-Experts Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning
distributed training |
0.9 | 1 | 2025 | FlexOLMo: Open Language Models for Flexible Data Use · NeurIPS 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.9 | 1 | 2025 | SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature · EMNLP 2025 |
Machine learning › Efficient and distributed learning › model compression
sparse training |
0.9 | 1 | 2025 | OLMoE: Open Mixture-of-Experts Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model
open language model development |
0.8 | 1 | 2024 | OLMo: Accelerating the Science of Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model training
pretraining corpus construction |
0.8 | 1 | 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024 |
Bioinformatics and computational biology › sequence analysis › sequencing read preprocessing
duplicate marking |
0.7 | 1 | 2023 | Dupsifter: a lightweight duplicate marking tool for whole genome bisulfite sequencing · Bioinform. 2023 |
Bioinformatics and computational biology
sequence analysis |
0.7 | 1 | 2023 | Dupsifter: a lightweight duplicate marking tool for whole genome bisulfite sequencing · Bioinform. 2023 |
Natural language and speech › Information extraction and text analysis › document understanding
scientific document understanding |
0.3 | 1 | 2025 | SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature · EMNLP 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2025 | OLMoE: Open Mixture-of-Experts Language Models · ICLR 2025 |
Authentication and access control › access control
data access control |
0.3 | 1 | 2025 | FlexOLMo: Open Language Models for Flexible Data Use · NeurIPS 2025 |
Cloud and datacenter computing › resource management
datacenter resource management |
0.3 | 1 | 2025 | Holistically Evaluating the Environmental Impact of Creating Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.2 | 1 | 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024 |
Methods — techniques the papers use, named apart from their topics
power measurement · 2.6life cycle assessment · 2.6nonparametric routing · 1.7model merging · 1.7mixture of experts · 1.7sparse mixture-of-experts · 0.9routing analysis · 0.9instruction tuning · 0.9dataset construction · 0.9corpus curation · 0.8streaming alignment · 0.7duplicate marking · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific LiteratureabstractDavid Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, Shannon Zejiang Shen, Doug Downey, Hannaneh Hajishirzi, Arman Cohan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Dave Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh 0001, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, Shannon Shen 0001, Doug Downey, Hannaneh Hajishirzi, Arman Cohan |
EMNLP | 3 |
| 2025 | Holistically Evaluating the Environmental Impact of Creating Language ModelsabstractAs the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the power consumption and carbon emissions from the final training runs for their latest models, there is comparatively little transparency into the impact of model development, hardware manufacturing, and total water usage throughout. In this work, we estimate the real-world environmental impact of developing a series of language models, ranging from 20 million to 13 billion active parameters, trained on up to 5.6 trillion tokens each. When accounting for hardware manufacturing, model development, and our final training runs, we find that our series of models released **493 metric tons** of carbon emissions, equivalent to powering about 98 homes in the United States for one year, and consumed **2.769 million liters of water**, equivalent to about 24.5 years of water usage by a person in the United States, even though our data center is extremely water-efficient. We measure and report the environmental impact of our model development; to the best of our knowledge we are the first to do so for LLMs, and we find that model development, the impact of which is generally not disclosed by most model developers, amounted to **~50%** of that of training. By looking at detailed time series data for power consumption, we also find that power usage throughout training is not consistent, fluctuating between ~15% and ~85% of our hardware's maximum power draw, with negative implications for grid-scale planning as demand continues to grow. We close with a discussion on the continued difficulty of estimating the environmental impact of AI systems, and key takeaways for model developers and the public at large. Jacob Morrison, Clara Na, Jared Fernandez, Tim Dettmers, Emma Strubell, Jesse Dodge |
ICLR | 1 |
| 2025 | OLMoE: Open Mixture-of-Experts Language ModelsabstractWe introduce OLMoE, a fully open, state-of-the-art language model leveraging sparse Mixture-of-Experts (MoE). OLMoE-1B-7B has 7 billion (B) parameters but uses only 1B per input token. We pretrain it on 5 trillion tokens and further adapt it to create OLMoE-1B-7B-Instruct. Our models outperform all available models with similar active parameters, even surpassing larger ones like Llama2-13B-Chat and DeepSeekMoE-16B. We present novel findings on MoE training, define and analyze new routing properties showing high specialization in our model, and open-source all our work: model weights, training data, code, and logs. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Pete Walsh 0001, Oyvind Tafjord, Nathan Lambert 0001, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, Dave Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi |
ICLR | 5 |
| 2025 | FlexOLMo: Open Language Models for Flexible Data UseabstractWe introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference. Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min |
NeurIPS | 5 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 22 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 18 |
| 2023 | Dupsifter: a lightweight duplicate marking tool for whole genome bisulfite sequencingabstractSUMMARY: In whole genome sequencing data, polymerase chain reaction amplification results in duplicate DNA fragments coming from the same location in the genome. The process of preparing a whole genome bisulfite sequencing (WGBS) library, on the other hand, can create two DNA fragments from the same location that should not be considered duplicates. Currently, only one WGBS-aware duplicate marking tool exists. However, it only works with the output from a single tool, does not accept streaming input or output, and requires a substantial amount of memory relative to the input size. Dupsifter provides an aligner-agnostic duplicate marking tool that is lightweight, has streaming capabilities, and is memory efficient. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are freely available at https://github.com/huishenlab/dupsifter under the MIT license. Dupsifter is implemented in C and is supported on macOS and Linux. Jacob Morrison, Wanding Zhou, Benjamin K. Johnson |
Bioinform. | 1 |
| 2022 | Bidimensional Leaderboards: Generate and Evaluate Language Hand in HandabstractJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Lavinia Dunagan, Jacob Morrison, Alexander R. Fabbri, Yejin Choi 0001, Noah A. Smith |
NAACL-HLT | 5 |
| 2022 | Transparent Human Evaluation for Image CaptioningabstractJungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras 0001, Yejin Choi 0001, Noah A. Smith |
NAACL-HLT | 4 |