Li Lucy

dblp:200/8869 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 36% Language models and text generation · 27% Probabilistic and Bayesian machine learning · 21%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
data filtering
0.812024
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters · ACL (1) 2024
Natural language and speech › Language models and text generation › large language model training
pretraining corpus construction
0.812024
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024
Machine learning › Representation and self-supervised learning › pre-training
pretraining data
0.812024
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters · ACL (1) 2024
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation
0.612022
Discovering Differences in the Representation of People using Contextualized Semantic Axes · EMNLP 2022
Natural language and speech › Information extraction and text analysis › distributional semantics
semantic axes
0.612022
Discovering Differences in the Representation of People using Contextualized Semantic Axes · EMNLP 2022
Natural language and speech › Language models and text generation › large language model training
language model pretraining
0.212024
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024
Computational social science and digital humanities › AI ethics
social bias analysis
0.212022
Discovering Differences in the Representation of People using Contextualized Semantic Axes · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

self-description extraction · 1.5filter analysis · 1.5contextualized semantic axes · 1.1BERT embeddings · 1.1corpus curation · 0.8
YearPublicationVenuePosition
2025 DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
abstract
Sami Baral, Li Lucy, Ryan Knight, Alice Ng, Luca Soldaini, Neil Heffernan, Kyle Lo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sami Baral, Li Lucy, Ryan Knight, Alice Ng, Luca Soldaini, Neil T. Heffernan, Kyle Lo
NAACL (Long Papers)2
2024 AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
abstract
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren Klein, Jesse Dodge. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge
ACL (1)1
2024 Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
abstract
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo
ACL (1)14
2024 "One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
abstract
Li Lucy, Su Lin Blodgett, Milad Shokouhi, Hanna Wallach, Alexandra Olteanu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Li Lucy, Su Lin Blodgett, Milad Shokouhi, Hanna M. Wallach, Alexandra Olteanu
NAACL-HLT1
2022 Discovering Differences in the Representation of People using Contextualized Semantic Axes
abstract
A common paradigm for identifying semantic differences across social and temporal contexts is the use of static word embeddings and their distances.In particular, past work has compared embeddings against "semantic axes" that represent two opposing concepts.We extend this paradigm to BERT embeddings, and construct contextualized axes that mitigate the pitfall where antonyms have neighboring representations.We validate and demonstrate these axes on two people-centric datasets: occupations from Wikipedia, and multi-platform discussions in extremist, men's communities over fourteen years.In both studies, contextualized semantic axes can characterize differences among instances of the same word type.In the latter study, we show that references to women and the contexts around them have become more detestable over time.
Li Lucy, Divya Tadimeti, David Bamman
EMNLP1
2021 Characterizing English Variation across Social Media Communities with BERT
abstract
Abstract Much previous work characterizing language variation across Internet social groups has focused on the types of words used by these groups. We extend this type of study by employing BERT to characterize variation in the senses of words as well, analyzing two months of English comments in 474 Reddit communities. The specificity of different sense clusters to a community, combined with the specificity of a community’s unique word types, is used to identify cases where a social group’s language deviates from the norm. We validate our metrics using user-created glossaries and draw on sociolinguistic theories to connect language variation with trends in community behavior. We find that communities with highly distinctive language are medium-sized, and their loyal and highly engaged users interact in dense networks.
Li Lucy, David Bamman
Trans. Assoc. Comput. Linguistics1