VLDB 2026 Research / reviewers in the wild / expert
Emma Strubell
dblp:153/2253
· DBLP profile ↗
31ranked-venue papers
5as first author
23since 2021 · last 2025
0000-0003-2798-0726ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Energy Considerations of Large Language Model Inference and Efficiency OptimizationsabstractJared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk, Sasha Luccioni, Emma Strubell. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk, Sasha Luccioni, Emma Strubell |
ACL (1) | 6 |
| 2025 | Holistically Evaluating the Environmental Impact of Creating Language ModelsabstractAs the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the power consumption and carbon emissions from the final training runs for their latest models, there is comparatively little transparency into the impact of model development, hardware manufacturing, and total water usage throughout. In this work, we estimate the real-world environmental impact of developing a series of language models, ranging from 20 million to 13 billion active parameters, trained on up to 5.6 trillion tokens each. When accounting for hardware manufacturing, model development, and our final training runs, we find that our series of models released **493 metric tons** of carbon emissions, equivalent to powering about 98 homes in the United States for one year, and consumed **2.769 million liters of water**, equivalent to about 24.5 years of water usage by a person in the United States, even though our data center is extremely water-efficient. We measure and report the environmental impact of our model development; to the best of our knowledge we are the first to do so for LLMs, and we find that model development, the impact of which is generally not disclosed by most model developers, amounted to **~50%** of that of training. By looking at detailed time series data for power consumption, we also find that power usage throughout training is not consistent, fluctuating between ~15% and ~85% of our hardware's maximum power draw, with negative implications for grid-scale planning as demand continues to grow. We close with a discussion on the continued difficulty of estimating the environmental impact of AI systems, and key takeaways for model developers and the public at large. Jacob Morrison, Clara Na, Jared Fernandez, Tim Dettmers, Emma Strubell, Jesse Dodge |
ICLR | 5 |
| 2025 | On-device Streaming Discrete Speech Units
Kwanghee Choi, Masao Someki, Emma Strubell, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsabstractLarge language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve its scalability with an efficient gradient projection strategy called LoGra that leverages the gradient structure in backpropagation. We then provide a theoretical motivation of gradient projection approaches to influence functions to promote trust in the data valuation process. Lastly, we lower the barrier to implementing data valuation systems by introducing LogIX, a software package that can transform existing training code into data valuation code with minimal effort. In our data valuation experiments, LoGra achieves competitive accuracy against more expensive baselines while showing up to 6,500x improvement in throughput and 5x reduction in GPU memory usage when applied to Llama3-8B-Instruct and the 1B-token dataset. Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff G. Schneider, Eduard H. Hovy, Roger B. Grosse, Eric P. Xing |
NeurIPS | 8 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 32 |
| 2024 | AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data FiltersabstractLi Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren Klein, Jesse Dodge. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge |
ACL (1) | 4 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 26 |
| 2024 | Scalable Data Ablation Approximations for Language Models through Modular Training and MergingabstractTraining data compositions for Large Language Models (LLMs) can significantly affect their downstream performance.However, a thorough data ablation study exploring large sets of candidate data mixtures is typically prohibitively expensive since the full effect is seen only after training the models; this can lead practitioners to settle for sub-optimal data mixtures.We propose an efficient method for approximating data ablations which trains individual models on subsets of a training corpus and reuses them across evaluations of combinations of subsets.In continued pre-training experiments, we find that, given an arbitrary evaluation set, the perplexity score of a single model trained on a candidate set of data is strongly correlated with perplexity scores of parameter averages of models trained on distinct partitions of that data.From this finding, we posit that researchers and practitioners can conduct inexpensive simulations of data ablations by maintaining a pool of models that were each trained on partitions of a large training corpus, and assessing candidate data mixtures by evaluating parameter averages of combinations of these models.This approach allows for substantial improvements in amortized training efficiency -scaling only linearly with respect to new data -by enabling reuse of previous training computation, opening new avenues for improving model performance through rigorous, incremental data assessment and mixing. Clara Na, Ian Magnusson, Ananya Harsh Jha, Tom Sherborne, Emma Strubell, Jesse Dodge, Pradeep Dasigi |
EMNLP | 5 |
| 2023 | To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question AnsweringabstractRecent advances in open-domain question answering (ODQA) have demonstrated impressive accuracy on general-purpose domains like Wikipedia.While some work has been investigating how well ODQA models perform when tested for out-of-domain (OOD) generalization, these studies have been conducted only under conservative shifts in data distribution and typically focus on a single component (i.e., retriever or reader) rather than an end-to-end system.This work proposes a more realistic endto-end domain shift evaluation setting covering five diverse domains.We not only find that endto-end models fail to generalize but that high retrieval scores often still yield poor answer prediction accuracy.To address these failures, we investigate several interventions, in the form of data augmentations, for improving model adaption and use our evaluation set to elucidate the relationship between the efficacy of an intervention scheme and the particular type of dataset shifts we consider.We propose a generalizability test that estimates the type of shift in a target dataset without training a model in the target domain and that the type of shift is predictive of which data augmentation schemes will be effective for domain adaption.Overall, we find that these interventions increase end-to-end performance by up to ∼24 points.* *This work was done while authors were at Google.Average F1 over all target datasets Average F1 over target datasets with specific shifts Dheeru Dua, Emma Strubell, Sameer Singh 0001, Patrick Verga |
ACL (1) | 2 |
| 2023 | Annotating Mentions Alone Enables Efficient Domain Adaptation for Coreference ResolutionabstractAlthough recent neural models for coreference resolution have led to substantial improvements on benchmark datasets, transferring these models to new target domains containing out-of-vocabulary spans and requiring differing annotation schemes remains challenging.Typical approaches involve continued training on annotated target-domain data, but obtaining annotations is costly and time-consuming.We show that annotating mentions alone is nearly twice as fast as annotating full coreference chains.Accordingly, we propose a method for efficiently adapting coreference models, which includes a high-precision mention detection objective and requires annotating only mentions in the target domain.Extensive evaluation across three English coreference datasets: CoNLL-2012 (news/conversation), i2b2/VA (medical notes), and previously unstudied child welfare notes, reveals that our approach facilitates annotation-efficient transfer and results in a 7-14% improvement in average F1 without increasing annotator time 1 . Nupoor Gandhi, Anjalie Field, Emma Strubell |
ACL (1) | 3 |
| 2023 | The Framework Tax: Disparities Between Inference Efficiency in NLP Research and DeploymentabstractIncreased focus on the computational efficiency of NLP systems has motivated the design of efficient model architectures and improvements to underlying hardware accelerators.However, the resulting increases in computational throughput and reductions in floating point operations have not directly translated to improvements in wall-clock inference latency.We demonstrate that these discrepancies can be largely attributed to bottlenecks introduced by deep learning frameworks.We denote this phenomenon as the framework tax, and observe that the disparity is growing as hardware speed increases over time.In this work, we examine this phenomenon through a series of case studies analyzing the effects of model design decisions, framework paradigms, and hardware platforms on total model latency. Jared Fernandez, Jacob Kahn, Clara Na, Yonatan Bisk, Emma Strubell |
EMNLP | 5 |
| 2023 | Understanding the Effect of Model Compression on Social Bias in Large Language ModelsabstractLarge Language Models (LLMs) trained with self-supervision on vast corpora of web text fit to the social biases of that text.Without intervention, these social biases persist in the model's predictions in downstream tasks, leading to representational harm.Many strategies have been proposed to mitigate the effects of inappropriate social biases learned during pretraining.Simultaneously, methods for model compression have become increasingly popular to reduce the computational burden of LLMs.Despite the popularity and need for both approaches, little work has been done to explore the interplay between these two.We perform a carefully controlled study of the impact of model compression via quantization and knowledge distillation on measures of social bias in LLMs.Longer pretraining and larger models led to higher social bias, and quantization showed a regularizer effect with its best trade-off around 20% of the original pretraining time. 1 Gustavo Gonçalves, Emma Strubell |
EMNLP | 2 |
| 2023 | To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language ProcessingabstractNLP is in a period of disruptive change that is impacting our methodologies, funding sources, and public perception.In this work, we seek to understand how to shape our future by better understanding our past.We study factors that shape NLP as a field, including culture, incentives, and infrastructure by conducting longform interviews with 26 NLP researchers of varying seniority, research area, institution, and social identity.Our interviewees identify cyclical patterns in the field, as well as new shifts without historical parallel, including changes in benchmark culture and software infrastructure.We complement this discussion with quantitative analysis of citation, authorship, and language use in the ACL Anthology over time.We conclude by discussing shared visions, concerns, and hopes for the future of NLP.We hope that this study of our field's past and present can prompt informed discussion of our community's implicit norms and more deliberate action to consciously shape the future. Sireesh Gururaja, Amanda Bertsch, Clara Na, David Gray Widder, Emma Strubell |
EMNLP | 5 |
| 2023 | DSI++: Updating Transformer Memory with New DocumentsabstractSanket Mehta, Jai Gupta, Yi Tay, Mostafa Dehghani, Vinh Tran, Jinfeng Rao, Marc Najork, Emma Strubell, Donald Metzler. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sanket Vaibhav Mehta, Jai Gupta 0001, Yi Tay, Mostafa Dehghani 0001, Vinh Q. Tran 0002, Jinfeng Rao, Marc Najork, Emma Strubell, Donald Metzler |
EMNLP | 8 |
| 2023 | Making Scalable Meta Learning PracticalabstractDespite its flexibility to learn diverse inductive biases in machine learning programs, meta learning (i.e.,\ learning to learn) has long been recognized to suffer from poor scalability due to its tremendous compute/memory costs, training instability, and a lack of efficient distributed training support. In this work, we focus on making scalable meta learning practical by introducing SAMA, which combines advances in both implicit differentiation algorithms and systems. Specifically, SAMA is designed to flexibly support a broad range of adaptive optimizers in the base level of meta learning programs, while reducing computational burden by avoiding explicit computation of second-order gradient information, and exploiting efficient distributed training techniques implemented for first-order gradients. Evaluated on multiple large-scale meta learning benchmarks, SAMA showcases up to 1.7/4.8x increase in throughput and 2.0/3.8x decrease in memory consumption respectively on single-/multi-GPU setups compared to other baseline meta learning algorithms. Furthermore, we show that SAMA-based data optimization leads to consistent improvements in text classification accuracy with BERT and RoBERTa large language models, and achieves state-of-the-art results in both small- and large-scale data pruning on image classification tasks, demonstrating the practical applicability of scalable meta learning across language and vision domains. Sang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger, Pengtao Xie, Emma Strubell, Eric P. Xing |
NeurIPS | 6 |
| 2023 | An Empirical Investigation of the Role of Pre-training in Lifelong LearningabstractThe lifelong learning paradigm in machine learning is an attractive alternative to the more prominent isolated learning scheme not only due to its resemblance to biological learning but also its potential to reduce energy waste by obviating excessive model re-training. A key challenge to this paradigm is the phenomenon of catastrophic forgetting. With the increasing popularity and success of pre-trained models in machine learning, we pose the question: What role does pre-training play in lifelong learning, specifically with respect to catastrophic forgetting? We investigate existing methods in the context of large, pre-trained models and evaluate their performance on a variety of text and image classification tasks, including a large-scale study using a novel data set of 15 diverse NLP tasks. Across all settings, we observe that generic pre-training implicitly alleviates the effects of catastrophic forgetting when learning multiple tasks sequentially compared to randomly initialized models. We then further investigate why pre-training alleviates forgetting in this setting. We study this phenomenon by analyzing the loss landscape, finding that pre-trained weights appear to ease forgetting by leading to wider minima. Based on this insight, we propose jointly optimizing for current task loss and loss basin sharpness to explicitly encourage wider basins during sequential fine-tuning. We show that this optimization approach outperforms several state-of-the-art task-sequential continual learning algorithms across multiple settings, occasionally even without retaining a memory that scales in size with the number of tasks. Sanket Vaibhav Mehta, Darshan Patil, Sarath Chandar, Emma Strubell |
J. Mach. Learn. Res. | 4 |
| 2023 | Efficient Methods for Natural Language Processing: A SurveyabstractAbstract Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows. Such resources include data, time, storage, or energy, all of which are naturally limited and unevenly distributed. This motivates research into efficient methods that require fewer resources to achieve similar results. This survey synthesizes and relates current methods and findings in efficient NLP. We aim to provide both guidance for conducting NLP under limited resources, and point towards promising research directions for developing more efficient methods. Marcos V. Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Manuel R. Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, Pedro Henrique Martins, André F. T. Martins, Jessica Zosa Forde, Peter A. Milder, Edwin Simpson, Noam Slonim, Jesse Dodge, Emma Strubell, Niranjan Balasubramanian, Leon Derczynski, Iryna Gurevych, Roy Schwartz 0001 |
Trans. Assoc. Comput. Linguistics | 18 |
| 2022 | Improving Compositional Generalization with Self-Training for Data-to-Text GenerationabstractSanket Vaibhav Mehta, Jinfeng Rao, Yi Tay, Mihir Kale, Ankur Parikh, Emma Strubell. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Sanket Vaibhav Mehta, Jinfeng Rao, Yi Tay, Mihir Kale, Ankur Parikh, Emma Strubell |
ACL (1) | 6 |
| 2022 | Bridging Fairness and Environmental Sustainability in Natural Language ProcessingabstractFairness and environmental impact are important research directions for the sustainable development of artificial intelligence.However, while each topic is an active research area in natural language processing (NLP), there is a surprising lack of research on the interplay between the two fields.This lacuna is highly problematic, since there is increasing evidence that an exclusive focus on fairness can actually hinder environmental sustainability, and vice versa.In this work, we shed light on this crucial intersection in NLP by (1) investigating the efficiency of current fairness approaches through surveying example methods for reducing unfair stereotypical bias from the literature, and(2) evaluating a common technique to reduce energy consumption (and thus environmental impact) of English NLP models, knowledge distillation (KD), for its impact on fairness.In this case study, we evaluate the effect of important KD factors, including layer and dimensionality reduction, with respect to: (a) performance on the distillation task (natural language inference and semantic similarity prediction), and (b) multiple measures and dimensions of stereotypical bias (e.g., gender bias measured via the Word Embedding Association Test).Our results lead us to clarify current assumptions regarding the effect of KD on unfair bias: contrary to other findings, we show that KD can actually decrease model fairness. Marius Hessenthaler, Emma Strubell, Dirk Hovy, Anne Lauscher |
EMNLP | 2 |
| 2022 | Transfer Learning from Semantic Role Labeling to Event Argument Extraction with Template-based Slot QueryingabstractIn this work, we investigate transfer learning from semantic role labeling (SRL) to event argument extraction (EAE), considering their similar argument structures.We view the extraction task as a role querying problem, unifying various methods into a single framework.There are key discrepancies on role labels and distant arguments between semantic role and event argument annotations.To mitigate these discrepancies, we specify natural language-like queries to tackle the label mismatch problem and devise argument augmentation to recover distant arguments.We show that SRL annotations can serve as a valuable resource for EAE, and a template-based slot querying strategy is especially effective for facilitating the transfer.In extensive evaluations on two English EAE benchmarks, our proposed model obtains impressive zero-shot results by leveraging SRL annotations, reaching nearly 80% of the fullysupervised scores.It further provides benefits in low-resource cases, where few EAE annotations are available.Moreover, we show that our approach generalizes to cross-domain and multilingual scenarios. Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP | 2 |
| 2022 | A Survey of Active Learning for Natural Language ProcessingabstractIn this work, we provide a literature review of active learning (AL) for its applications in natural language processing (NLP).In addition to a fine-grained categorization of query strategies, we also investigate several other important aspects of applying AL to NLP problems.These include AL for structured prediction tasks, annotation cost, model learning (especially with deep neural models), and starting and stopping AL.Finally, we conclude with a discussion of related topics and future directions. Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP | 2 |
| 2021 | WiFiMod: Transformer-based Indoor Human Mobility Modeling using Passive SensingabstractModeling human mobility has a wide range of applications from urban planning to simulations of disease spread. It is well known that humans spend 80% of their time indoors but modeling indoor human mobility is challenging due to three main reasons: (i) the absence of easily acquirable, reliable, low-cost indoor mobility datasets, (ii) high prediction space in modeling the frequent indoor mobility, and (iii) multi-scalar periodicity and correlations in mobility. To deal with all these challenges, we propose WiFiMod, a Transformer-based, data-driven approach that models indoor human mobility at multiple spatial scales using WiFi system logs. WiFiMod takes as input enterprise WiFi system logs to extract human mobility trajectories from smartphone digital traces. Next, for each extracted trajectory, we identify the mobility features at multiple spatial scales, macro and micro, to design a multi-modal embedding Transformer that predicts user mobility for several hours to an entire day across multiple spatial granularities. Multi-modal embedding captures the mobility periodicity and correlations across various scales while Transformers capture long term mobility dependencies boosting model prediction performance. This approach significantly reduces the prediction space by first predicting macro mobility, then modeling indoor scale mobility, micro mobility, conditioned on the estimated macro mobility distribution, thereby using the topological constraint of the macro-scale. Experimental results show that WiFiMod achieves a prediction accuracy of at least 10% points higher than the current state-of-art models. Additionally, we present 3 real-world applications of WiFiMod - (i) predict high density hot pockets and space utilization for policy making decisions for COVID19 or ILI, (ii) generate a realistic simulation of indoor mobility data to simulate spread of diseases, (iii) design personal assistants. Amee Trivedi, Kate Silverstein, Emma Strubell, Prashant J. Shenoy, Mohit Iyyer |
COMPASS | 3 |
| 2021 | On the Benefit of Syntactic Supervision for Cross-lingual Transfer in Semantic Role LabelingabstractAlthough recent developments in neural architectures and pre-trained representations have greatly increased state-of-the-art model performance on fully-supervised semantic role labeling (SRL), the task remains challenging for languages where supervised SRL training data are not abundant.Cross-lingual learning can improve performance in this setting by transferring knowledge from high-resource languages to low-resource ones.Moreover, we hypothesize that annotations of syntactic dependencies can be leveraged to further facilitate cross-lingual transfer.In this work, we perform an empirical exploration of the helpfulness of syntactic supervision for crosslingual SRL within a simple multitask learning scheme.With comprehensive evaluations across ten languages (in addition to English) and three SRL benchmark datasets, including both dependency-and span-based SRL, we show the effectiveness of syntactic supervision in low-resource scenarios.Experiments Target Languages SRL Style Same Frames?Compatible Roles?Main SRL Setting EWT/UPB † ( §3.2) de,fr,it,es,pt,fi Dependency-based Yes Yes Zero-shot EWT/FiPB ( §3.3) fi Dependency-based No Yes Semi-supervised CoNLL-2009 ( §3.4) cs,zh,es,ca Dependency-based No No Semi-supervised OntoNotes ( §3.5) zh,ar Span-based No Yes Semi-supervised Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP (1) | 2 |
| 2020 | Energy and Policy Considerations for Modern Deep Learning ResearchabstractThe field of artificial intelligence has experienced a dramatic methodological shift towards large neural networks trained on plentiful data. This shift has been fueled by recent advances in hardware and techniques enabling remarkable levels of computation, resulting in impressive advances in AI across many applications. However, the massive computation required to obtain these exciting results is costly both financially, due to the price of specialized hardware and electricity or cloud compute time, and to the environment, as a result of non-renewable energy used to fuel modern tensor processing hardware. In a paper published this year at ACL, we brought this issue to the attention of NLP researchers by quantifying the approximate financial and environmental costs of training and tuning neural network models for NLP (Strubell, Ganesh, and McCallum 2019). In this extended abstract, we briefly summarize our findings in NLP, incorporating updated estimates and broader information from recent related publications, and provide actionable recommendations to reduce costs and improve equity in the machine learning and artificial intelligence community. Emma Strubell, Ananya Ganesh, Andrew McCallum |
AAAI | 1 |
| 2019 | Energy and Policy Considerations for Deep Learning in NLPabstractRecent progress in hardware and methodology for training neural networks has ushered in a new generation of large networks trained on abundant data.These models have obtained notable gains in accuracy across many NLP tasks.However, these accuracy improvements depend on the availability of exceptionally large computational resources that necessitate similarly substantial energy consumption.As a result these models are costly to train and develop, both financially, due to the cost of hardware and electricity or cloud compute time, and environmentally, due to the carbon footprint required to fuel modern tensor processing hardware.In this paper we bring this issue to the attention of NLP researchers by quantifying the approximate financial and environmental costs of training a variety of recently successful neural network models for NLP.Based on these findings, we propose actionable recommendations to reduce costs and improve equity in NLP research and practice. Emma Strubell, Ananya Ganesh, Andrew McCallum |
ACL (1) | 1 |
| 2018 | Multi-Task Learning For Parsing The Alexa Meaning Representation LanguageabstractThe Alexa Meaning Representation Language (AMRL) is a compositional graph-based semantic representation that includes fine-grained types, properties, actions, and roles and can represent a wide variety of spoken language. AMRL increases the ability of virtual assistants to represent more complex requests, including logical and conditional statements as well as ones with nested clauses. Due to this representational capacity, the acquisition of large scale data resources is challenging, which limits the accuracy of resulting models. This paper has two primary contributions. First, we develop a linearization of AMRL graphs along with a deep multi-task model that predicts fine-grained types, properties, and intents. Second, we show how to jointly train a model that predicts an existing representation for spoken language understanding (SLU) along with the linearized AMRL parse. The resulting model, which leverages learned embeddings from both tasks, is able to predict the AMRL representation more accurately than other approaches, decreasing the error rates in the full parse by 3.56% absolute and reducing the amount of natively annotated data needed to train accurate parsing models. Vittorio Perera, Tagyoung Chung, Thomas Kollar, Emma Strubell |
AAAI | 4 |
| 2018 | Linguistically-Informed Self-Attention for Semantic Role LabelingabstractCurrent state-of-the-art semantic role labeling (SRL) uses a deep neural network with no explicit linguistic features.However, prior work has shown that gold syntax trees can dramatically improve SRL decoding, suggesting the possibility of increased accuracy from explicit modeling of syntax.In this work, we present linguistically-informed self-attention (LISA): a neural network model that combines multi-head self-attention with multi-task learning across dependency parsing, part-ofspeech tagging, predicate detection and SRL.Unlike previous models which require significant pre-processing to prepare linguistic features, LISA can incorporate syntax using merely raw tokens as input, encoding the sequence only once to simultaneously perform parsing, predicate detection and role labeling for all predicates.Syntax is incorporated by training one attention head to attend to syntactic parents for each token.Moreover, if a high-quality syntactic parse is already available, it can be beneficially injected at test time without re-training our SRL model.In experiments on CoNLL-2005 SRL, LISA achieves new state-of-the-art performance for a model using predicted predicates and standard word embeddings, attaining 2.5 F1 absolute higher than the previous state-of-the-art on newswire and more than 3.5 F1 on outof-domain data, nearly 10% reduction in error.On ConLL-2012 English SRL we also show an improvement of more than 2.5 F1.LISA also out-performs the state-of-the-art with contextually-encoded (ELMo) word representations, by nearly 1.0 F1 on news and more than 2.0 F1 on out-of-domain text. Emma Strubell, Patrick Verga, Daniel Andor, David Weiss 0001, Andrew McCallum |
EMNLP | 1 |
| 2018 | Simultaneously Self-Attending to All Mentions for Full-Abstract Biological Relation ExtractionabstractPatrick Verga, Emma Strubell, Andrew McCallum. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Patrick Verga, Emma Strubell, Andrew McCallum |
NAACL-HLT | 2 |
| 2017 | Fast and Accurate Entity Recognition with Iterated Dilated ConvolutionsabstractToday when many practitioners run basic NLP on the entire web and large-volume traffic, faster methods are paramount to saving time and energy costs.Recent advances in GPU hardware have led to the emergence of bi-directional LSTMs as a standard method for obtaining pertoken vector representations serving as input to labeling tasks such as NER (often followed by prediction in a linear-chain CRF).Though expressive and accurate, these models fail to fully exploit GPU parallelism, limiting their computational efficiency.This paper proposes a faster alternative to Bi-LSTMs for NER: Iterated Dilated Convolutional Neural Networks (ID-CNNs), which have better capacity than traditional CNNs for large context and structured prediction.Unlike LSTMs whose sequential processing on sentences of length N requires O(N ) time even in the face of parallelism, ID-CNNs permit fixed-depth convolutions to run in parallel across entire documents.We describe a distinct combination of network structure, parameter sharing and training procedures that enable dramatic 14-20x testtime speedups while retaining accuracy comparable to the Bi-LSTM-CRF.Moreover, ID-CNNs trained to aggregate context from the entire document are even more accurate while maintaining 8x faster test time speeds. Emma Strubell, Patrick Verga, David Belanger 0002, Andrew McCallum |
EMNLP | 1 |
| 2016 | Multilingual Relation Extraction using Compositional Universal SchemaabstractPatrick Verga, David Belanger, Emma Strubell, Benjamin Roth, Andrew McCallum. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Patrick Verga, David Belanger 0002, Emma Strubell, Benjamin Roth 0001, Andrew McCallum |
HLT-NAACL | 3 |
| 2015 | Learning Dynamic Feature Selection for Fast Sequential PredictionabstractEmma Strubell, Luke Vilnis, Kate Silverstein, Andrew McCallum. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Emma Strubell, Luke Vilnis, Kate Silverstein, Andrew McCallum |
ACL (1) | 1 |