VLDB 2026 Research / reviewers in the wild / expert
Li Lucy
dblp:200/8869
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Representation and self-supervised learning · 36% Language models and text generation · 27% Probabilistic and Bayesian machine learning · 21% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
data filtering |
0.8 | 1 | 2024 | AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model training
pretraining corpus construction |
0.8 | 1 | 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › pre-training
pretraining data |
0.8 | 1 | 2024 | AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation |
0.6 | 1 | 2022 | Discovering Differences in the Representation of People using Contextualized Semantic Axes · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis › distributional semantics
semantic axes |
0.6 | 1 | 2022 | Discovering Differences in the Representation of People using Contextualized Semantic Axes · EMNLP 2022 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.2 | 1 | 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research · ACL (1) 2024 |
Computational social science and digital humanities › AI ethics
social bias analysis |
0.2 | 1 | 2022 | Discovering Differences in the Representation of People using Contextualized Semantic Axes · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
self-description extraction · 1.5filter analysis · 1.5contextualized semantic axes · 1.1BERT embeddings · 1.1corpus curation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math ImagesabstractSami Baral, Li Lucy, Ryan Knight, Alice Ng, Luca Soldaini, Neil Heffernan, Kyle Lo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sami Baral, Li Lucy, Ryan Knight, Alice Ng, Luca Soldaini, Neil T. Heffernan, Kyle Lo |
NAACL (Long Papers) | 2 |
| 2024 | AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data FiltersabstractLi Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren Klein, Jesse Dodge. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge |
ACL (1) | 1 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 14 |
| 2024 | "One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System BehaviorsabstractLi Lucy, Su Lin Blodgett, Milad Shokouhi, Hanna Wallach, Alexandra Olteanu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Li Lucy, Su Lin Blodgett, Milad Shokouhi, Hanna M. Wallach, Alexandra Olteanu |
NAACL-HLT | 1 |
| 2022 | Discovering Differences in the Representation of People using Contextualized Semantic AxesabstractA common paradigm for identifying semantic differences across social and temporal contexts is the use of static word embeddings and their distances.In particular, past work has compared embeddings against "semantic axes" that represent two opposing concepts.We extend this paradigm to BERT embeddings, and construct contextualized axes that mitigate the pitfall where antonyms have neighboring representations.We validate and demonstrate these axes on two people-centric datasets: occupations from Wikipedia, and multi-platform discussions in extremist, men's communities over fourteen years.In both studies, contextualized semantic axes can characterize differences among instances of the same word type.In the latter study, we show that references to women and the contexts around them have become more detestable over time. Li Lucy, Divya Tadimeti, David Bamman |
EMNLP | 1 |
| 2021 | Characterizing English Variation across Social Media Communities with BERTabstractAbstract Much previous work characterizing language variation across Internet social groups has focused on the types of words used by these groups. We extend this type of study by employing BERT to characterize variation in the senses of words as well, analyzing two months of English comments in 474 Reddit communities. The specificity of different sense clusters to a community, combined with the specificity of a community’s unique word types, is used to identify cases where a social group’s language deviates from the norm. We validate our metrics using user-created glossaries and draw on sociolinguistic theories to connect language variation with trends in community behavior. We find that communities with highly distinctive language are medium-sized, and their loyal and highly engaged users interact in dense networks. Li Lucy, David Bamman |
Trans. Assoc. Comput. Linguistics | 1 |