David M. Mimno

dblp:39/5487 · also David Mimno · DBLP profile ↗
← Back
44ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0001-7510-9404ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 10 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Large Language Models in Qualitative Research: Uses, Tensions, and Intentions
abstract
CHI ’25, Yokohama, Japan
Hope Schroeder, Marianne Aubin Le Quéré, Casey Randazzo, David M. Mimno, Sarita Yardi Schoenebeck
CHI4
2024 Automate or Assist? The Role of Computational Models in Identifying Gendered Discourse in US Capital Trial Transcripts
abstract
The language used by US courtroom actors in criminal trials has long been studied for biases. However, systematic studies for bias in high-stakes court trials have been difficult, due to the nuanced nature of bias and the legal expertise required. Large language models offer the possibility to automate annotation. But validating the computational approach requires both an understanding of how automated methods fit in existing annotation workflows and what they really offer. We present a case study of adding a computational model to a complex and high-stakes problem: identifying gender-biased language in US capital trials for women defendants. Our team of experienced death-penalty lawyers and NLP technologists pursue a three-phase study: first annotating manually, then training and evaluating computational models, and finally comparing expert annotations to model predictions. Unlike many typical NLP tasks, annotating for gender bias in months-long capital trials is complicated, with many individual judgment calls. Contrary to standard arguments for automation that are based on efficiency and scalability, legal experts find the computational models most useful in providing opportunities to reflect on their own bias in annotation and to build consensus on annotation rules. This experience suggests that seeking to replace experts with computational models for complex annotation is both unrealistic and undesirable. Rather, computational models offer valuable opportunities to assist the legal experts in annotation-based studies.
Andrea W. Wen-Yi, Kathryn Adamson, Nathalie Greenfield, Rachel Goldberg, Sandra Babcock, David M. Mimno, Allison Koenecke
AIES (1)6
2024 Sensemaking about Contraceptive Methods across Online Platforms
abstract
Selecting a birth control method is a complex healthcare decision. While birth control methods provide important benefits, they can also cause unpredictable side effects and be stigmatized, leading many people to seek additional information online, where they can privately find reviews, advice, hypotheses, and experiences of other birth control users. However, the relationships between their healthcare concerns, sensemaking activities, and online settings are not well understood. We gather texts about birth control shared on Twitter and Reddit—popular communities with different affordances, moderation, and audiences—to study where and how birth control is discussed online. Using a combination of topic modeling and hand annotation, we identify and characterize the dominant sensemaking practices across these platforms, and we create lexica to draw comparisons across birth control methods and side effects. We use these to measure variations from survey reports of side effect experiences, highlighting topics that social media users discuss more than expected online. Our findings characterize how online platforms are used to make sense of difficult healthcare choices, including analyzing risks, calculating timing and dosages, hypothesizing about causes of side effects, and storytelling about painful experiences. We contribute both to understanding unmet needs of birth control users and to exploring context-specific patterns in social media discussions.
LeAnn McDowall, Maria Antoniak, David M. Mimno
ICWSM3
2024 A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity
abstract
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, Daphne Ippolito. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David M. Mimno, Daphne Ippolito
NAACL-HLT10
2023 Modeling Legal Reasoning: LM Annotation at the Edge of Human Agreement
abstract
Generative language models (LMs) are increasingly used for document class-prediction tasks and promise enormous improvements in cost and efficiency.Existing research often examines simple classification tasks, but the capability of LMs to classify on complex or specialized tasks is less well understood.We consider a highly complex task that is challenging even for humans: the classification of legal reasoning according to jurisprudential philosophy.Using a novel dataset of historical United States Supreme Court opinions annotated by a team of domain experts, we systematically test the performance of a variety of LMs.We find that generative models perform poorly when given instructions (i.e.prompts) equal to the instructions presented to human annotators through our codebook.Our strongest results derive from fine-tuning models on the annotated dataset; the best performing model is an in-domain model, LEGAL-BERT.We apply predictions from this fine-tuned model to study historical trends in jurisprudence, an exercise that both aligns with prominent qualitative historical accounts and points to areas of possible refinement in those accounts.Our findings generally sound a note of caution in the use of generative LMs on complex tasks without finetuning and point to the continued relevance of human annotation-intensive classification methods.
Rosamond Elizabeth Thalken, Edward H. Stiglitz, David M. Mimno, Matthew Wilkens
EMNLP3
2023 Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token Embeddings
abstract
Cross-lingual transfer learning is an important property of multilingual large language models (LLMs).But how do LLMs represent relationships between languages?Every language model has an input layer that maps tokens to vectors.This ubiquitous layer of language models is often overlooked.We find that similarities between these input embeddings are highly interpretable and that the geometry of these embeddings differs between model families.In one case (XLM-RoBERTa), embeddings encode language: tokens in different writing systems can be linearly separated with an average of 99.2% accuracy.Another family (mT5) represents cross-lingual semantic similarity: the 50 nearest neighbors for any token represent an average of 7.61 writing systems, and are frequently translations.This result is surprising given that there is no explicit parallel crosslingual training corpora and no explicit incentive for translations in pre-training objectives.Our research opens the door for investigations in 1) The effect of pre-training and model architectures on representations of languages and 2) The applications of cross-lingual representations embedded in language models.
Andrea W. Wen-Yi, David M. Mimno
EMNLP2
2023 Data Similarity is Not Enough to Explain Language Model Performance
abstract
Large language models achieve high performance on many but not all downstream tasks.The interaction between pretraining data and task data is commonly assumed to determine this variance: a task with data that is more similar to a model's pretraining data is assumed to be easier for that model.We test whether distributional and example-specific similarity measures (embedding-, token-and model-based) correlate with language model performance through a large-scale comparison of the Pile and C4 pretraining datasets with downstream benchmarks.Similarity correlates with performance for multilingual datasets, but in other benchmarks, we surprisingly find that similarity metrics are not correlated with accuracy or even each other.This suggests that the relationship between pretraining data and downstream tasks is more complex than often assumed.
Gregory Yauney, Emily Reif, David M. Mimno
EMNLP3
2021 Bad Seeds: Evaluating Lexical Methods for Bias Measurement
abstract
Maria Antoniak, David Mimno. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Maria Antoniak, David M. Mimno
ACL/IJCNLP (1)2
2021 Comparing Text Representations: A Theory-Driven Approach
abstract
Much of the progress in contemporary NLP has come from learning representations, such as masked language model (MLM) contextual embeddings, that turn challenging problems into simple classification tasks.But how do we quantify and explain this effect?We adapt general tools from computational learning theory to fit the specific characteristics of text datasets and present a method to evaluate the compatibility between representations and tasks.Even though many tasks can be easily solved with simple bag-of-words (BOW) representations, BOW does poorly on hard natural language inference tasks.For one such task we find that BOW cannot distinguish between real and randomized labelings, while pre-trained MLM representations show 72x greater distinction between real and random labelings than BOW.This method provides a calibrated, quantitative measure of the difficulty of a classificationbased NLP task, enabling comparisons between representations without requiring empirical evaluations that may be sensitive to initializations and hyperparameters.The method provides a fresh perspective on the patterns in a dataset and the alignment of those patterns with specific labels.
Gregory Yauney, David M. Mimno
EMNLP (1)2
2021 On-the-fly Rectification for Robust Large-Vocabulary Topic Inference
abstract
Across many data domains, co-occurrence statistics about the joint appearance of objects are powerfully informative. By transforming unsupervised learning problems into decompositions of co-occurrence statistics, spectral algorithms provide transparent and efficient algorithms for posterior inference such as latent topic analysis and community detection. As object vocabularies grow, however, it becomes rapidly more expensive to store and run inference algorithms on co-occurrence statistics. Rectifying co-occurrence, the key process to uphold model assumptions, becomes increasingly more vital in the presence of rare terms, but current techniques cannot scale to large vocabularies. We propose novel methods that simultaneously compress and rectify co-occurrence statistics, scaling gracefully with the size of vocabulary and the dimension of latent space. We also present new algorithms learning latent variables from the compressed statistics, and verify that our methods perform comparably to previous approaches on both textual and non-textual data.
Moontae Lee, Sungjun Cho, David M. Mimno, David Bindel
ICML4
2021 Tags, Borders, and Catalogs: Social Re-Working of Genre on LibraryThing
abstract
Through a computational reading of the online book reviewing community LibraryThing, we examine the dynamics of a collaborative tagging system and learn how its users refine and redefine literary genres. LibraryThing tags are overlapping and multi-dimensional, created in a shared space by thousands of users, including readers, bookstore owners, and librarians. A common understanding of genre is that it relates to the content of books, but this resource allows us to view genre as an intersection of user communities and reader values and interests. We explore different methods of computational genre measurement within the open space of user-created tags. We measure overlap between books, tags, and users, and we also measure the homogeneity of communities associated with genre tags and correlate this homogeneity with reviewing behavior.Finally, by analyzing the text of reviews, we identify the thematic signatures of genres on LibraryThing, revealing similarities and differences between them. These measurements are intended to elucidate the genre conceptions of the users, not, as in prior work, to normalize the tags or enforce a hierarchy. We find that LibraryThing users make sense of genre through a variety of values and expectations, many of which fall outside common definitions and understandings of genre.
Maria Antoniak, Melanie Walsh, David M. Mimno
Proc. ACM Hum. Comput. Interact.3
2020 Prior-aware Composition Inference for Spectral Topic Models
abstract
Spectral algorithms operate on matrices or tensors of word co-occurrence to learn latent topics. These approaches remove the dependence on the original documents and produce substantial gains in efficiency with provable inference, but at a cost: the models can no longer infer any information about individual documents. Thresholded Linear Inverse is developed to learn document-specific topic compositions, but its linear characteristics limit the inference quality without considering any prior information on topic distributions. We propose two novel estimation methods that respect previously unclear prior structures of spectral topic models. Experiments on a variety of synthetic to real collections demonstrate that our Prior-Aware Dual Decomposition outperforms the baseline method, whereas our Prior-Aware Manifold Iteration performs even better on short realistic data.
Moontae Lee, David Bindel, David M. Mimno
AISTATS3
2020 Domain-Specific Lexical Grounding in Noisy Visual-Textual Documents
abstract
Images can give us insights into the contextual meanings of words, but current imagetext grounding approaches require detailed annotations.Such granular annotation is rare, expensive, and unavailable in most domainspecific contexts.In contrast, unlabeled multiimage, multi-sentence documents are abundant.Can lexical grounding be learned from such documents, even though they have significant lexical and visual overlap?Working with a case study dataset of real estate listings, we demonstrate the challenge of distinguishing highly correlated grounded terms, such as "kitchen" and "bedroom", and introduce metrics to assess this document similarity.We present a simple unsupervised clusteringbased method that increases precision and recall beyond object detection and image tagging baselines when evaluated on labeled subsets of the dataset.The proposed method is particularly effective for local contextual meanings of a word, for example associating "granite" with countertops in the real estate dataset and with rocky landscapes in a Wikipedia dataset."kitchen" (18.4% labeled true) E: 72.9 AUC W: 52.7 AUC R: 21.1 AUC "outdoor" (16.9% labeled true) E: 68.5 AUC W: 20.0 AUC R: 13.2 AUC "washer" (1.6% labeled true)
Gregory Yauney, Jack Hessel, David M. Mimno
EMNLP (1)3
2019 Unsupervised Discovery of Multimodal Links in Multi-image, Multi-sentence Documents
abstract
Jack Hessel, Lillian Lee, David Mimno. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jack Hessel, Lillian Lee, David M. Mimno
EMNLP/IJCNLP (1)3
2019 Practical Correlated Topic Modeling and Analysis via the Rectified Anchor Word Algorithm
abstract
Moontae Lee, Sungjun Cho, David Bindel, David Mimno. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Moontae Lee, Sungjun Cho, David Bindel, David M. Mimno
EMNLP/IJCNLP (1)4
2019 Narrative Paths and Negotiation of Power in Birth Stories
abstract
Birth stories have become increasingly common on the internet, but they have received little attention as a computational dataset. These unsolicited, publicly posted stories provide rich descriptions of decisions, emotions, and relationships during a common but sometimes traumatic medical experience. These personal details can be illuminating for medical practitioners, and due to their shared structures, birth stories are also an ideal testing ground for narrative analysis techniques. We present an analysis of 2,847 birth stories from an online forum and demonstrate the utility of these stories for computational work. We discover clear sentiment, topic and persona-based patterns that both model the expected narrative event sequences of birth stories and highlight diverging pathways and exceptions to narrative norms. The authors' motivation to publicly post these personal stories can be a way to regain power after a surveilled and disempowering experience, and we explore power relationships between the personas in the stories, showing that these dynamics can vary with the type of birth (e.g., medicated vs unmedicated). Finally, birth stories exist in a space that is both public and deeply personal. This liminality poses a challenge for analysis and presentation, and we discuss tradeoffs and ethical practices for this collection. WARNING: This paper includes detailed narratives of pregnancy and birth.
Maria Antoniak, David M. Mimno, Karen Levy
Proc. ACM Hum. Comput. Interact.2
2019 Boosted negative sampling by quadratically constrained entropy maximization
Taygun Kekeç, David M. Mimno, David M. J. Tax
Pattern Recognit. Lett.2
2018 Authorless Topic Models: Biasing Models Away from Known Structure
abstract
Most previous work in unsupervised semantic modeling in the presence of metadata has assumed that our goal is to make latent dimensions more correlated with metadata, but in practice the exact opposite is often true. Some users want topic models that highlight differences between, for example, authors, but others seek more subtle connections across authors. We introduce three metrics for identifying topics that are highly correlated with metadata, and demonstrate that this problem affects between 30 and 50% of the topics in models trained on two real-world collections, regardless of the size of the model. We find that we can predict which words cause this phenomenon and that by selectively subsampling these words we dramatically reduce topic-metadata correlation, improve topic stability, and maintain or even improve model quality.
Laure Thompson, David M. Mimno
COLING2
2018 Quantifying the Visual Concreteness of Words and Topics in Multimodal Datasets
abstract
Jack Hessel, David Mimno, Lillian Lee. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Jack Hessel, David M. Mimno, Lillian Lee
NAACL-HLT2
2018 Evaluating the Stability of Embedding-based Word Similarities
abstract
Word embeddings are increasingly being used as a tool to study word associations in specific corpora. However, it is unclear whether such embeddings reflect enduring properties of language or if they are sensitive to inconsequential variations in the source documents. We find that nearest-neighbor distances are highly sensitive to small changes in the training corpus for a variety of algorithms. For all methods, including specific documents in the training set can result in substantial variations. We show that these effects are more prominent for smaller training corpora. We recommend that users never rely on single embedding models for distance calculations, but rather average over multiple bootstrap samples, especially for small corpora.
Maria Antoniak, David M. Mimno
Trans. Assoc. Comput. Linguistics2
2017 The strange geometry of skip-gram with negative sampling
abstract
Despite their ubiquity, word embeddings trained with skip-gram negative sampling (SGNS) remain poorly understood.We find that vector positions are not simply determined by semantic similarity, but rather occupy a narrow cone, diametrically opposed to the context vectors.We show that this geometric concentration depends on the ratio of positive to negative examples, and that it is neither theoretically nor empirically inherent in related embedding algorithms.
David M. Mimno, Laure Thompson
EMNLP1
2017 Quantifying the Effects of Text Duplication on Semantic Models
abstract
Duplicate documents are a pervasive problem in text datasets and can have a strong effect on unsupervised models.Methods to remove duplicate texts are typically heuristic or very expensive, so it is vital to know when and why they are needed.We measure the sensitivity of two latent semantic methods to the presence of different levels of document repetition.By artificially creating different forms of duplicate text we confirm several hypotheses about how repeated text impacts models.While a small amount of duplication is tolerable, substantial over-representation of subsets of the text may overwhelm meaningful topical patterns.
Alexandra Schofield, Laure Thompson, David M. Mimno
EMNLP3
2017 Cats and Captions vs. Creators and the Clock: Comparing Multimodal Content to Context in Predicting Relative Popularity
abstract
The content of today's social media is becoming more and more rich, increasingly mixing text, images, videos, and audio. It is an intriguing research question to model the interplay between these different modes in attracting user attention and engagement. But in order to pursue this study of multimodal content, we must also account for context: timing effects, community preferences, and social factors (e.g., which authors are already popular) also affect the amount of feedback and reaction that social-media posts receive. In this work, we separate out the influence of these non-content factors in several ways. First, we focus on ranking pairs of submissions posted to the same community in quick succession, e.g., within 30 seconds; this framing encourages models to focus on time-agnostic and community-specific content features. Within that setting, we determine the relative performance of author vs. content features. We find that victory usually belongs to "cats and captions," as visual and textual features together tend to outperform identity-based features. Moreover, our experiments show that when considered in isolation, simple unigram text features and deep neural network visual features yield the highest accuracy individually, and that the combination of the two modalities generally leads to the best accuracies overall.
Jack Hessel, Lillian Lee, David M. Mimno
WWW3
2017 Comparing grounded theory and topic modeling: Extreme divergence or unlikely convergence?
abstract
Researchers in information science and related areas have developed various methods for analyzing textual data, such as survey responses. This article describes the application of analysis methods from two distinct fields, one method from interpretive social science and one method from statistical machine learning, to the same survey data. The results show that the two analyses produce some similar and some complementary insights about the phenomenon of interest, in this case, nonuse of social media. We compare both the processes of conducting these analyses and the results they produce to derive insights about each method's unique advantages and drawbacks, as well as the broader roles that these methods play in the respective fields where they are often used. These insights allow us to make more informed decisions about the tradeoffs in choosing different methods for analyzing textual data. Furthermore, this comparison suggests ways that such methods might be combined in novel and compelling ways.
Eric P. S. Baumer, David M. Mimno, Shion Guha, Emily Quan, Geri Gay
J. Assoc. Inf. Sci. Technol.2
2016 Machine Learning and Grounded Theory Method: Convergence, Divergence, and Combination
abstract
Grounded Theory Method (GTM) and Machine Learning (ML) are often considered to be quite different. In this note, we explore unexpected convergences between these methods. We propose new research directions that can further clarify the relationships between these methods, and that can use those relationships to strengthen our ability to describe our phenomena and develop stronger hybrid theories.
Michael J. Muller, Shion Guha, Eric P. S. Baumer, David M. Mimno, N. Sadat Shami
GROUP4
2016 Beyond Exchangeability: The Chinese Voting Process
abstract
Many online communities present user-contributed responses, such as reviews of products and answers to questions. User-provided helpfulness votes can highlight the most useful responses, but voting is a social process that can gain momentum based on the popularity of responses and the polarity of existing votes. We propose the Chinese Voting Process (CVP) which models the evolution of helpfulness votes as a self-reinforcing process dependent on position and presentation biases. We evaluate this model on Amazon product reviews and more than 80 StackExchange forums, measuring the intrinsic quality of individual responses and behavioral coefficients of different communities.
Moontae Lee, Seok Hyun Jin, David M. Mimno
NIPS3
2016 Comparing Apples to Apple: The Effects of Stemmers on Topic Models
abstract
Rule-based stemmers such as the Porter stemmer are frequently used to preprocess English corpora for topic modeling. In this work, we train and evaluate topic models on a variety of corpora using several different stemming algorithms. We examine several different quantitative measures of the resulting models, including likelihood, coherence, model stability, and entropy. Despite their frequent use in topic modeling, we find that stemmers produce no meaningful improvement in likelihood and coherence and in fact can degrade topic stability.
Alexandra Schofield, David M. Mimno
Trans. Assoc. Comput. Linguistics2
2015 Evaluation methods for unsupervised word embeddings
abstract
We present a comprehensive study of evaluation methods for unsupervised embedding techniques that obtain meaningful representations of words from text.Different evaluations result in different orderings of embedding methods, calling into question the common assumption that there is one single optimal vector representation.We present new evaluation techniques that directly compare embeddings with respect to specific queries.These methods reduce bias, provide greater insight, and allow us to solicit data-driven relevance judgments rapidly and accurately through crowdsourcing.
Tobias Schnabel, Igor Labutov, David M. Mimno, Thorsten Joachims
EMNLP3
2015 Robust Spectral Inference for Joint Stochastic Matrix Factorization
abstract
Spectral inference provides fast algorithms and provable optimality for latent topic analysis. But for real data these algorithms require additional ad-hoc heuristics, and even then often produce unusable results. We explain this poor performance by casting the problem of topic inference in the framework of Joint Stochastic Matrix Factorization (JSMF) and showing that previous methods violate the theoretical conditions necessary for a good solution to exist. We then propose a novel rectification method that learns high quality topics and their interactions even on small, noisy data. This method achieves results comparable to probabilistic techniques in several domains while maintaining scalability and provable optimality.
Moontae Lee, David Bindel, David M. Mimno
NIPS3
2014 Low-dimensional Embeddings for Interpretable Anchor-based Topic Inference
abstract
The anchor words algorithm performs provably efficient topic model inference by finding an approximate convex hull in a high-dimensional word co-occurrence space.However, the existing greedy algorithm often selects poor anchor words, reducing topic quality and interpretability.Rather than finding an approximate convex hull in a high-dimensional space, we propose to find an exact convex hull in a visualizable 2-or 3-dimensional space.Such low-dimensional embeddings both improve topics and clearly show users why the algorithm selects certain words.
David M. Mimno, Moontae Lee
EMNLP1
2013 A Practical Algorithm for Topic Modeling with Provable Guarantees
abstract
Topic models provide a useful method for dimensionality reduction and exploratory data analysis in large text corpora. Most approaches to topic model learning have been based on a maximum likelihood objective. Efficient algorithms exist that attempt to approximate this objective, but they have no provable guarantees. Recently, algorithms have been introduced that provide provable bounds, but these algorithms are not practical because they are inefficient and not robust to violations of model assumptions. In this paper we present an algorithm for learning topic models that is both provable and practical. The algorithm produces results comparable to the best MCMC implementations while running orders of magnitude faster.
Sanjeev Arora, Rong Ge 0001, Yoni Halpern, David M. Mimno, Ankur Moitra, David A. Sontag, Michael Zhu
ICML (2)4
2012 Sparse stochastic inference for latent Dirichlet allocation
David M. Mimno, Matthew Hoffman 0001, David M. Blei
ICML1
2012 Scalable Inference of Overlapping Communities
abstract
We develop a scalable algorithm for posterior inference of overlapping communities in large networks. Our algorithm is based on stochastic variational inference in the mixed-membership stochastic blockmodel. It naturally interleaves subsampling the network with estimating its community structure. We apply our algorithm on ten large, real-world networks with up to 60,000 nodes. It converges several orders of magnitude faster than the state-of-the-art algorithm for MMSB, finds hundreds of communities in large real-world networks, and detects the true communities in 280 benchmark networks with equal or better accuracy compared to other scalable algorithms.
Prem Gopalan, David M. Mimno, Sean Gerrish, Michael J. Freedman, David M. Blei
NIPS2
2011 Bayesian Checking for Topic Models
David M. Mimno, David M. Blei
EMNLP1
2011 Optimizing Semantic Coherence in Topic Models
David M. Mimno, Hanna M. Wallach, Edmund M. Talley, Miriam Leenders, Andrew McCallum
EMNLP1
2011 Reconstructing Pompeian Households
David M. Mimno
UAI1
2009 Polylingual Topic Models
David M. Mimno, Hanna M. Wallach, Jason Naradowsky, David A. Smith, Andrew McCallum
EMNLP1
2009 Evaluation methods for topic models
abstract
A natural evaluation metric for statistical topic models is the probability of held-out documents given a trained model. While exact computation of this probability is intractable, several estimators for this probability have been used in the topic modeling literature, including the harmonic mean method and empirical likelihood method. In this paper, we demonstrate experimentally that commonly-used methods are unlikely to accurately estimate the probability of held-out documents, and propose two alternative methods that are both accurate and efficient.
Hanna M. Wallach, Iain Murray 0001, Ruslan Salakhutdinov, David M. Mimno
ICML4
2009 Efficient methods for topic model inference on streaming document collections
abstract
Topic models provide a powerful tool for analyzing large text collections by representing high dimensional data in a low dimensional subspace. Fitting a topic model given a set of training documents requires approximate inference techniques that are computationally expensive. With today's large-scale, constantly expanding document collections, it is useful to be able to infer topic distributions for new documents without retraining the model. In this paper, we empirically evaluate the performance of several methods for topic inference in previously unseen documents, including methods based on Gibbs sampling, variational inference, and a new method inspired by text classification. The classification-based inference method produces results similar to iterative inference methods, but requires only a single matrix multiplication. In addition to these inference methods, we present SparseLDA, an algorithm and data structure for evaluating Gibbs sampling distributions. Empirical results indicate that SparseLDA can be approximately 20 times faster than traditional LDA and provide twice the speedup of previously published fast sampling methods, while also using substantially less memory.
Limin Yao, David M. Mimno, Andrew McCallum
KDD2
2009 Rethinking LDA: Why Priors Matter
abstract
Implementations of topic models typically use symmetric Dirichlet priors with fixed concentration parameters, with the implicit assumption that such smoothing parameters" have little practical effect. In this paper, we explore several classes of structured priors for topic models. We find that an asymmetric Dirichlet prior over the document-topic distributions has substantial advantages over a symmetric prior, while an asymmetric prior over the topic-word distributions provides no real benefit. Approximation of this prior structure through simple, efficient hyperparameter optimization steps is sufficient to achieve these performance gains. The prior structure we advocate substantially increases the robustness of topic models to variations in the number of topics and to the highly skewed word frequency distributions common in natural language. Since this prior structure can be implemented using efficient algorithms that add negligible cost beyond standard inference techniques, we recommend it as a new standard for topic modeling."
Hanna M. Wallach, David M. Mimno, Andrew McCallum
NIPS2
2008 InterNano: e-Science for the Nanomanufacturing Community
abstract
As network-enabled scholarship produces huge quantities of formal and informal research outputs in a variety of formats and varying levels of access, it is "enhanced" science that will facilitate the discovery, selection, and analysis of information that are a necessary part of the scientific research cycle particularly among interdisciplinary research communities. The national nanomanufacturing network (NNN) has developed a Web service, internano, to support the information needs of the nanomanufacturing community within this network-enabled research environment. Based on the clearinghouse concept, InterNano brings together heterogeneous resources and research outputs with networking and computational tools to enable discovery, facilitate selection of information, and encourage collaboration. With seventeen content features and services, internano's goal is not just to make information work within this community more efficient, but also to enable researchers to act on their discoveries within the same virtual framework.
Rebecca Reznik-Zellen, Bob Stevens, Michael Thorn, Jeff Morse, Mark D. Smucker, James Allan 0001, David M. Mimno, Andrew McCallum, Mark Tuominen
eScience7
2008 Topic Models Conditioned on Arbitrary Features with Dirichlet-multinomial Regression
David M. Mimno, Andrew McCallum
UAI1
2007 Mixtures of hierarchical topics with Pachinko allocation
abstract
The four-level pachinko allocation model (PAM) (Li & McCallum, 2006) represents correlations among topics using a DAG structure. It does not, however, represent a nested hierarchy of topics, with some topical word distributions representing the vocabulary that is shared among several more specific topics. This paper presents hierarchical PAM---an enhancement that explicitly represents a topic hierarchy. This model can be seen as combining the advantages of hLDA's topical hierarchy representation with PAM's ability to mix multiple leaves of the topic hierarchy. Experimental results show improvements in likelihood of held-out documents, as well as mutual information between automatically-discovered topics and humangenerated categories such as journals.
David M. Mimno, Wei Li 0010, Andrew McCallum
ICML1
2007 Expertise modeling for matching papers with reviewers
abstract
An essential part of an expert-finding task, such as matching reviewers to submitted papers, is the ability to model the expertise of a person based on documents. We evaluate several measures of the association between an author in an existing collection of research papers and a previously unseen document. We compare two language model based approaches with a novel topic model, Author-Persona-Topic (APT). In this model, each author can write under one or more "personas," which are represented as independent distributions over hidden topics. Examples of previous papers written by prospective reviewers are gathered from the Rexa database, which extracts and disambiguates author mentions from documents gathered from the web. We evaluate the models using a reviewer matching task based on human relevance judgments determining how well the expertise of proposed reviewers matches a submission. We find that the APT topic model outperforms the other models.
David M. Mimno, Andrew McCallum
KDD1