VLDB 2026 Research / reviewers in the wild / expert
David M. Mimno
dblp:39/5487 · also David Mimno
· DBLP profile ↗
44ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0001-7510-9404ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 10 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models in Qualitative Research: Uses, Tensions, and IntentionsabstractCHI ’25, Yokohama, Japan Hope Schroeder, Marianne Aubin Le Quéré, Casey Randazzo, David M. Mimno, Sarita Yardi Schoenebeck |
CHI | 4 |
| 2024 | Automate or Assist? The Role of Computational Models in Identifying Gendered Discourse in US Capital Trial TranscriptsabstractThe language used by US courtroom actors in criminal trials has long been studied for biases. However, systematic studies for bias in high-stakes court trials have been difficult, due to the nuanced nature of bias and the legal expertise required. Large language models offer the possibility to automate annotation. But validating the computational approach requires both an understanding of how automated methods fit in existing annotation workflows and what they really offer. We present a case study of adding a computational model to a complex and high-stakes problem: identifying gender-biased language in US capital trials for women defendants. Our team of experienced death-penalty lawyers and NLP technologists pursue a three-phase study: first annotating manually, then training and evaluating computational models, and finally comparing expert annotations to model predictions. Unlike many typical NLP tasks, annotating for gender bias in months-long capital trials is complicated, with many individual judgment calls. Contrary to standard arguments for automation that are based on efficiency and scalability, legal experts find the computational models most useful in providing opportunities to reflect on their own bias in annotation and to build consensus on annotation rules. This experience suggests that seeking to replace experts with computational models for complex annotation is both unrealistic and undesirable. Rather, computational models offer valuable opportunities to assist the legal experts in annotation-based studies. Andrea W. Wen-Yi, Kathryn Adamson, Nathalie Greenfield, Rachel Goldberg, Sandra Babcock, David M. Mimno, Allison Koenecke |
AIES (1) | 6 |
| 2024 | Sensemaking about Contraceptive Methods across Online PlatformsabstractSelecting a birth control method is a complex healthcare decision. While birth control methods provide important benefits, they can also cause unpredictable side effects and be stigmatized, leading many people to seek additional information online, where they can privately find reviews, advice, hypotheses, and experiences of other birth control users. However, the relationships between their healthcare concerns, sensemaking activities, and online settings are not well understood. We gather texts about birth control shared on Twitter and Reddit—popular communities with different affordances, moderation, and audiences—to study where and how birth control is discussed online. Using a combination of topic modeling and hand annotation, we identify and characterize the dominant sensemaking practices across these platforms, and we create lexica to draw comparisons across birth control methods and side effects. We use these to measure variations from survey reports of side effect experiences, highlighting topics that social media users discuss more than expected online. Our findings characterize how online platforms are used to make sense of difficult healthcare choices, including analyzing risks, calculating timing and dosages, hypothesizing about causes of side effects, and storytelling about painful experiences. We contribute both to understanding unmet needs of birth control users and to exploring context-specific patterns in social media discussions. LeAnn McDowall, Maria Antoniak, David M. Mimno |
ICWSM | 3 |
| 2024 | A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & ToxicityabstractShayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, Daphne Ippolito. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David M. Mimno, Daphne Ippolito |
NAACL-HLT | 10 |
| 2023 | Modeling Legal Reasoning: LM Annotation at the Edge of Human AgreementabstractGenerative language models (LMs) are increasingly used for document class-prediction tasks and promise enormous improvements in cost and efficiency.Existing research often examines simple classification tasks, but the capability of LMs to classify on complex or specialized tasks is less well understood.We consider a highly complex task that is challenging even for humans: the classification of legal reasoning according to jurisprudential philosophy.Using a novel dataset of historical United States Supreme Court opinions annotated by a team of domain experts, we systematically test the performance of a variety of LMs.We find that generative models perform poorly when given instructions (i.e.prompts) equal to the instructions presented to human annotators through our codebook.Our strongest results derive from fine-tuning models on the annotated dataset; the best performing model is an in-domain model, LEGAL-BERT.We apply predictions from this fine-tuned model to study historical trends in jurisprudence, an exercise that both aligns with prominent qualitative historical accounts and points to areas of possible refinement in those accounts.Our findings generally sound a note of caution in the use of generative LMs on complex tasks without finetuning and point to the continued relevance of human annotation-intensive classification methods. Rosamond Elizabeth Thalken, Edward H. Stiglitz, David M. Mimno, Matthew Wilkens |
EMNLP | 3 |
| 2023 | Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token EmbeddingsabstractCross-lingual transfer learning is an important property of multilingual large language models (LLMs).But how do LLMs represent relationships between languages?Every language model has an input layer that maps tokens to vectors.This ubiquitous layer of language models is often overlooked.We find that similarities between these input embeddings are highly interpretable and that the geometry of these embeddings differs between model families.In one case (XLM-RoBERTa), embeddings encode language: tokens in different writing systems can be linearly separated with an average of 99.2% accuracy.Another family (mT5) represents cross-lingual semantic similarity: the 50 nearest neighbors for any token represent an average of 7.61 writing systems, and are frequently translations.This result is surprising given that there is no explicit parallel crosslingual training corpora and no explicit incentive for translations in pre-training objectives.Our research opens the door for investigations in 1) The effect of pre-training and model architectures on representations of languages and 2) The applications of cross-lingual representations embedded in language models. Andrea W. Wen-Yi, David M. Mimno |
EMNLP | 2 |
| 2023 | Data Similarity is Not Enough to Explain Language Model PerformanceabstractLarge language models achieve high performance on many but not all downstream tasks.The interaction between pretraining data and task data is commonly assumed to determine this variance: a task with data that is more similar to a model's pretraining data is assumed to be easier for that model.We test whether distributional and example-specific similarity measures (embedding-, token-and model-based) correlate with language model performance through a large-scale comparison of the Pile and C4 pretraining datasets with downstream benchmarks.Similarity correlates with performance for multilingual datasets, but in other benchmarks, we surprisingly find that similarity metrics are not correlated with accuracy or even each other.This suggests that the relationship between pretraining data and downstream tasks is more complex than often assumed. Gregory Yauney, Emily Reif, David M. Mimno |
EMNLP | 3 |
| 2021 | Bad Seeds: Evaluating Lexical Methods for Bias MeasurementabstractMaria Antoniak, David Mimno. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Maria Antoniak, David M. Mimno |
ACL/IJCNLP (1) | 2 |
| 2021 | Comparing Text Representations: A Theory-Driven ApproachabstractMuch of the progress in contemporary NLP has come from learning representations, such as masked language model (MLM) contextual embeddings, that turn challenging problems into simple classification tasks.But how do we quantify and explain this effect?We adapt general tools from computational learning theory to fit the specific characteristics of text datasets and present a method to evaluate the compatibility between representations and tasks.Even though many tasks can be easily solved with simple bag-of-words (BOW) representations, BOW does poorly on hard natural language inference tasks.For one such task we find that BOW cannot distinguish between real and randomized labelings, while pre-trained MLM representations show 72x greater distinction between real and random labelings than BOW.This method provides a calibrated, quantitative measure of the difficulty of a classificationbased NLP task, enabling comparisons between representations without requiring empirical evaluations that may be sensitive to initializations and hyperparameters.The method provides a fresh perspective on the patterns in a dataset and the alignment of those patterns with specific labels. Gregory Yauney, David M. Mimno |
EMNLP (1) | 2 |
| 2021 | On-the-fly Rectification for Robust Large-Vocabulary Topic InferenceabstractAcross many data domains, co-occurrence statistics about the joint appearance of objects are powerfully informative. By transforming unsupervised learning problems into decompositions of co-occurrence statistics, spectral algorithms provide transparent and efficient algorithms for posterior inference such as latent topic analysis and community detection. As object vocabularies grow, however, it becomes rapidly more expensive to store and run inference algorithms on co-occurrence statistics. Rectifying co-occurrence, the key process to uphold model assumptions, becomes increasingly more vital in the presence of rare terms, but current techniques cannot scale to large vocabularies. We propose novel methods that simultaneously compress and rectify co-occurrence statistics, scaling gracefully with the size of vocabulary and the dimension of latent space. We also present new algorithms learning latent variables from the compressed statistics, and verify that our methods perform comparably to previous approaches on both textual and non-textual data. Moontae Lee, Sungjun Cho, David M. Mimno, David Bindel |
ICML | 4 |
| 2021 | Tags, Borders, and Catalogs: Social Re-Working of Genre on LibraryThingabstractThrough a computational reading of the online book reviewing community LibraryThing, we examine the dynamics of a collaborative tagging system and learn how its users refine and redefine literary genres. LibraryThing tags are overlapping and multi-dimensional, created in a shared space by thousands of users, including readers, bookstore owners, and librarians. A common understanding of genre is that it relates to the content of books, but this resource allows us to view genre as an intersection of user communities and reader values and interests. We explore different methods of computational genre measurement within the open space of user-created tags. We measure overlap between books, tags, and users, and we also measure the homogeneity of communities associated with genre tags and correlate this homogeneity with reviewing behavior.Finally, by analyzing the text of reviews, we identify the thematic signatures of genres on LibraryThing, revealing similarities and differences between them. These measurements are intended to elucidate the genre conceptions of the users, not, as in prior work, to normalize the tags or enforce a hierarchy. We find that LibraryThing users make sense of genre through a variety of values and expectations, many of which fall outside common definitions and understandings of genre. Maria Antoniak, Melanie Walsh, David M. Mimno |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2020 | Prior-aware Composition Inference for Spectral Topic ModelsabstractSpectral algorithms operate on matrices or tensors of word co-occurrence to learn latent topics. These approaches remove the dependence on the original documents and produce substantial gains in efficiency with provable inference, but at a cost: the models can no longer infer any information about individual documents. Thresholded Linear Inverse is developed to learn document-specific topic compositions, but its linear characteristics limit the inference quality without considering any prior information on topic distributions. We propose two novel estimation methods that respect previously unclear prior structures of spectral topic models. Experiments on a variety of synthetic to real collections demonstrate that our Prior-Aware Dual Decomposition outperforms the baseline method, whereas our Prior-Aware Manifold Iteration performs even better on short realistic data. Moontae Lee, David Bindel, David M. Mimno |
AISTATS | 3 |
| 2020 | Domain-Specific Lexical Grounding in Noisy Visual-Textual DocumentsabstractImages can give us insights into the contextual meanings of words, but current imagetext grounding approaches require detailed annotations.Such granular annotation is rare, expensive, and unavailable in most domainspecific contexts.In contrast, unlabeled multiimage, multi-sentence documents are abundant.Can lexical grounding be learned from such documents, even though they have significant lexical and visual overlap?Working with a case study dataset of real estate listings, we demonstrate the challenge of distinguishing highly correlated grounded terms, such as "kitchen" and "bedroom", and introduce metrics to assess this document similarity.We present a simple unsupervised clusteringbased method that increases precision and recall beyond object detection and image tagging baselines when evaluated on labeled subsets of the dataset.The proposed method is particularly effective for local contextual meanings of a word, for example associating "granite" with countertops in the real estate dataset and with rocky landscapes in a Wikipedia dataset."kitchen" (18.4% labeled true) E: 72.9 AUC W: 52.7 AUC R: 21.1 AUC "outdoor" (16.9% labeled true) E: 68.5 AUC W: 20.0 AUC R: 13.2 AUC "washer" (1.6% labeled true) Gregory Yauney, Jack Hessel, David M. Mimno |
EMNLP (1) | 3 |
| 2019 | Unsupervised Discovery of Multimodal Links in Multi-image, Multi-sentence DocumentsabstractJack Hessel, Lillian Lee, David Mimno. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jack Hessel, Lillian Lee, David M. Mimno |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Practical Correlated Topic Modeling and Analysis via the Rectified Anchor Word AlgorithmabstractMoontae Lee, Sungjun Cho, David Bindel, David Mimno. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Moontae Lee, Sungjun Cho, David Bindel, David M. Mimno |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Narrative Paths and Negotiation of Power in Birth StoriesabstractBirth stories have become increasingly common on the internet, but they have received little attention as a computational dataset. These unsolicited, publicly posted stories provide rich descriptions of decisions, emotions, and relationships during a common but sometimes traumatic medical experience. These personal details can be illuminating for medical practitioners, and due to their shared structures, birth stories are also an ideal testing ground for narrative analysis techniques. We present an analysis of 2,847 birth stories from an online forum and demonstrate the utility of these stories for computational work. We discover clear sentiment, topic and persona-based patterns that both model the expected narrative event sequences of birth stories and highlight diverging pathways and exceptions to narrative norms. The authors' motivation to publicly post these personal stories can be a way to regain power after a surveilled and disempowering experience, and we explore power relationships between the personas in the stories, showing that these dynamics can vary with the type of birth (e.g., medicated vs unmedicated). Finally, birth stories exist in a space that is both public and deeply personal. This liminality poses a challenge for analysis and presentation, and we discuss tradeoffs and ethical practices for this collection. WARNING: This paper includes detailed narratives of pregnancy and birth. Maria Antoniak, David M. Mimno, Karen Levy |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2019 | Boosted negative sampling by quadratically constrained entropy maximization
Taygun Kekeç, David M. Mimno, David M. J. Tax |
Pattern Recognit. Lett. | 2 |
| 2018 | Authorless Topic Models: Biasing Models Away from Known StructureabstractMost previous work in unsupervised semantic modeling in the presence of metadata has assumed that our goal is to make latent dimensions more correlated with metadata, but in practice the exact opposite is often true. Some users want topic models that highlight differences between, for example, authors, but others seek more subtle connections across authors. We introduce three metrics for identifying topics that are highly correlated with metadata, and demonstrate that this problem affects between 30 and 50% of the topics in models trained on two real-world collections, regardless of the size of the model. We find that we can predict which words cause this phenomenon and that by selectively subsampling these words we dramatically reduce topic-metadata correlation, improve topic stability, and maintain or even improve model quality. Laure Thompson, David M. Mimno |
COLING | 2 |
| 2018 | Quantifying the Visual Concreteness of Words and Topics in Multimodal DatasetsabstractJack Hessel, David Mimno, Lillian Lee. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jack Hessel, David M. Mimno, Lillian Lee |
NAACL-HLT | 2 |
| 2018 | Evaluating the Stability of Embedding-based Word SimilaritiesabstractWord embeddings are increasingly being used as a tool to study word associations in specific corpora. However, it is unclear whether such embeddings reflect enduring properties of language or if they are sensitive to inconsequential variations in the source documents. We find that nearest-neighbor distances are highly sensitive to small changes in the training corpus for a variety of algorithms. For all methods, including specific documents in the training set can result in substantial variations. We show that these effects are more prominent for smaller training corpora. We recommend that users never rely on single embedding models for distance calculations, but rather average over multiple bootstrap samples, especially for small corpora. Maria Antoniak, David M. Mimno |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | The strange geometry of skip-gram with negative samplingabstractDespite their ubiquity, word embeddings trained with skip-gram negative sampling (SGNS) remain poorly understood.We find that vector positions are not simply determined by semantic similarity, but rather occupy a narrow cone, diametrically opposed to the context vectors.We show that this geometric concentration depends on the ratio of positive to negative examples, and that it is neither theoretically nor empirically inherent in related embedding algorithms. David M. Mimno, Laure Thompson |
EMNLP | 1 |
| 2017 | Quantifying the Effects of Text Duplication on Semantic ModelsabstractDuplicate documents are a pervasive problem in text datasets and can have a strong effect on unsupervised models.Methods to remove duplicate texts are typically heuristic or very expensive, so it is vital to know when and why they are needed.We measure the sensitivity of two latent semantic methods to the presence of different levels of document repetition.By artificially creating different forms of duplicate text we confirm several hypotheses about how repeated text impacts models.While a small amount of duplication is tolerable, substantial over-representation of subsets of the text may overwhelm meaningful topical patterns. Alexandra Schofield, Laure Thompson, David M. Mimno |
EMNLP | 3 |
| 2017 | Cats and Captions vs. Creators and the Clock: Comparing Multimodal Content to Context in Predicting Relative PopularityabstractThe content of today's social media is becoming more and more rich, increasingly mixing text, images, videos, and audio. It is an intriguing research question to model the interplay between these different modes in attracting user attention and engagement. But in order to pursue this study of multimodal content, we must also account for context: timing effects, community preferences, and social factors (e.g., which authors are already popular) also affect the amount of feedback and reaction that social-media posts receive. In this work, we separate out the influence of these non-content factors in several ways. First, we focus on ranking pairs of submissions posted to the same community in quick succession, e.g., within 30 seconds; this framing encourages models to focus on time-agnostic and community-specific content features. Within that setting, we determine the relative performance of author vs. content features. We find that victory usually belongs to "cats and captions," as visual and textual features together tend to outperform identity-based features. Moreover, our experiments show that when considered in isolation, simple unigram text features and deep neural network visual features yield the highest accuracy individually, and that the combination of the two modalities generally leads to the best accuracies overall. Jack Hessel, Lillian Lee, David M. Mimno |
WWW | 3 |
| 2017 | Comparing grounded theory and topic modeling: Extreme divergence or unlikely convergence?abstractResearchers in information science and related areas have developed various methods for analyzing textual data, such as survey responses. This article describes the application of analysis methods from two distinct fields, one method from interpretive social science and one method from statistical machine learning, to the same survey data. The results show that the two analyses produce some similar and some complementary insights about the phenomenon of interest, in this case, nonuse of social media. We compare both the processes of conducting these analyses and the results they produce to derive insights about each method's unique advantages and drawbacks, as well as the broader roles that these methods play in the respective fields where they are often used. These insights allow us to make more informed decisions about the tradeoffs in choosing different methods for analyzing textual data. Furthermore, this comparison suggests ways that such methods might be combined in novel and compelling ways. Eric P. S. Baumer, David M. Mimno, Shion Guha, Emily Quan, Geri Gay |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Machine Learning and Grounded Theory Method: Convergence, Divergence, and CombinationabstractGrounded Theory Method (GTM) and Machine Learning (ML) are often considered to be quite different. In this note, we explore unexpected convergences between these methods. We propose new research directions that can further clarify the relationships between these methods, and that can use those relationships to strengthen our ability to describe our phenomena and develop stronger hybrid theories. Michael J. Muller, Shion Guha, Eric P. S. Baumer, David M. Mimno, N. Sadat Shami |
GROUP | 4 |
| 2016 | Beyond Exchangeability: The Chinese Voting ProcessabstractMany online communities present user-contributed responses, such as reviews of products and answers to questions. User-provided helpfulness votes can highlight the most useful responses, but voting is a social process that can gain momentum based on the popularity of responses and the polarity of existing votes. We propose the Chinese Voting Process (CVP) which models the evolution of helpfulness votes as a self-reinforcing process dependent on position and presentation biases. We evaluate this model on Amazon product reviews and more than 80 StackExchange forums, measuring the intrinsic quality of individual responses and behavioral coefficients of different communities. Moontae Lee, Seok Hyun Jin, David M. Mimno |
NIPS | 3 |
| 2016 | Comparing Apples to Apple: The Effects of Stemmers on Topic ModelsabstractRule-based stemmers such as the Porter stemmer are frequently used to preprocess English corpora for topic modeling. In this work, we train and evaluate topic models on a variety of corpora using several different stemming algorithms. We examine several different quantitative measures of the resulting models, including likelihood, coherence, model stability, and entropy. Despite their frequent use in topic modeling, we find that stemmers produce no meaningful improvement in likelihood and coherence and in fact can degrade topic stability. Alexandra Schofield, David M. Mimno |
Trans. Assoc. Comput. Linguistics | 2 |
| 2015 | Evaluation methods for unsupervised word embeddingsabstractWe present a comprehensive study of evaluation methods for unsupervised embedding techniques that obtain meaningful representations of words from text.Different evaluations result in different orderings of embedding methods, calling into question the common assumption that there is one single optimal vector representation.We present new evaluation techniques that directly compare embeddings with respect to specific queries.These methods reduce bias, provide greater insight, and allow us to solicit data-driven relevance judgments rapidly and accurately through crowdsourcing. Tobias Schnabel, Igor Labutov, David M. Mimno, Thorsten Joachims |
EMNLP | 3 |
| 2015 | Robust Spectral Inference for Joint Stochastic Matrix FactorizationabstractSpectral inference provides fast algorithms and provable optimality for latent topic analysis. But for real data these algorithms require additional ad-hoc heuristics, and even then often produce unusable results. We explain this poor performance by casting the problem of topic inference in the framework of Joint Stochastic Matrix Factorization (JSMF) and showing that previous methods violate the theoretical conditions necessary for a good solution to exist. We then propose a novel rectification method that learns high quality topics and their interactions even on small, noisy data. This method achieves results comparable to probabilistic techniques in several domains while maintaining scalability and provable optimality. Moontae Lee, David Bindel, David M. Mimno |
NIPS | 3 |
| 2014 | Low-dimensional Embeddings for Interpretable Anchor-based Topic InferenceabstractThe anchor words algorithm performs provably efficient topic model inference by finding an approximate convex hull in a high-dimensional word co-occurrence space.However, the existing greedy algorithm often selects poor anchor words, reducing topic quality and interpretability.Rather than finding an approximate convex hull in a high-dimensional space, we propose to find an exact convex hull in a visualizable 2-or 3-dimensional space.Such low-dimensional embeddings both improve topics and clearly show users why the algorithm selects certain words. David M. Mimno, Moontae Lee |
EMNLP | 1 |
| 2013 | A Practical Algorithm for Topic Modeling with Provable GuaranteesabstractTopic models provide a useful method for dimensionality reduction and exploratory data analysis in large text corpora. Most approaches to topic model learning have been based on a maximum likelihood objective. Efficient algorithms exist that attempt to approximate this objective, but they have no provable guarantees. Recently, algorithms have been introduced that provide provable bounds, but these algorithms are not practical because they are inefficient and not robust to violations of model assumptions. In this paper we present an algorithm for learning topic models that is both provable and practical. The algorithm produces results comparable to the best MCMC implementations while running orders of magnitude faster. Sanjeev Arora, Rong Ge 0001, Yoni Halpern, David M. Mimno, Ankur Moitra, David A. Sontag, Michael Zhu |
ICML (2) | 4 |
| 2012 | Sparse stochastic inference for latent Dirichlet allocation
David M. Mimno, Matthew Hoffman 0001, David M. Blei |
ICML | 1 |
| 2012 | Scalable Inference of Overlapping CommunitiesabstractWe develop a scalable algorithm for posterior inference of overlapping communities in large networks. Our algorithm is based on stochastic variational inference in the mixed-membership stochastic blockmodel. It naturally interleaves subsampling the network with estimating its community structure. We apply our algorithm on ten large, real-world networks with up to 60,000 nodes. It converges several orders of magnitude faster than the state-of-the-art algorithm for MMSB, finds hundreds of communities in large real-world networks, and detects the true communities in 280 benchmark networks with equal or better accuracy compared to other scalable algorithms. Prem Gopalan, David M. Mimno, Sean Gerrish, Michael J. Freedman, David M. Blei |
NIPS | 2 |
| 2011 | Bayesian Checking for Topic Models
David M. Mimno, David M. Blei |
EMNLP | 1 |
| 2011 | Optimizing Semantic Coherence in Topic Models
David M. Mimno, Hanna M. Wallach, Edmund M. Talley, Miriam Leenders, Andrew McCallum |
EMNLP | 1 |
| 2011 | Reconstructing Pompeian Households
David M. Mimno |
UAI | 1 |
| 2009 | Polylingual Topic Models
David M. Mimno, Hanna M. Wallach, Jason Naradowsky, David A. Smith, Andrew McCallum |
EMNLP | 1 |
| 2009 | Evaluation methods for topic modelsabstractA natural evaluation metric for statistical topic models is the probability of held-out documents given a trained model. While exact computation of this probability is intractable, several estimators for this probability have been used in the topic modeling literature, including the harmonic mean method and empirical likelihood method. In this paper, we demonstrate experimentally that commonly-used methods are unlikely to accurately estimate the probability of held-out documents, and propose two alternative methods that are both accurate and efficient. Hanna M. Wallach, Iain Murray 0001, Ruslan Salakhutdinov, David M. Mimno |
ICML | 4 |
| 2009 | Efficient methods for topic model inference on streaming document collectionsabstractTopic models provide a powerful tool for analyzing large text collections by representing high dimensional data in a low dimensional subspace. Fitting a topic model given a set of training documents requires approximate inference techniques that are computationally expensive. With today's large-scale, constantly expanding document collections, it is useful to be able to infer topic distributions for new documents without retraining the model. In this paper, we empirically evaluate the performance of several methods for topic inference in previously unseen documents, including methods based on Gibbs sampling, variational inference, and a new method inspired by text classification. The classification-based inference method produces results similar to iterative inference methods, but requires only a single matrix multiplication. In addition to these inference methods, we present SparseLDA, an algorithm and data structure for evaluating Gibbs sampling distributions. Empirical results indicate that SparseLDA can be approximately 20 times faster than traditional LDA and provide twice the speedup of previously published fast sampling methods, while also using substantially less memory. Limin Yao, David M. Mimno, Andrew McCallum |
KDD | 2 |
| 2009 | Rethinking LDA: Why Priors MatterabstractImplementations of topic models typically use symmetric Dirichlet priors with fixed concentration parameters, with the implicit assumption that such smoothing parameters" have little practical effect. In this paper, we explore several classes of structured priors for topic models. We find that an asymmetric Dirichlet prior over the document-topic distributions has substantial advantages over a symmetric prior, while an asymmetric prior over the topic-word distributions provides no real benefit. Approximation of this prior structure through simple, efficient hyperparameter optimization steps is sufficient to achieve these performance gains. The prior structure we advocate substantially increases the robustness of topic models to variations in the number of topics and to the highly skewed word frequency distributions common in natural language. Since this prior structure can be implemented using efficient algorithms that add negligible cost beyond standard inference techniques, we recommend it as a new standard for topic modeling." Hanna M. Wallach, David M. Mimno, Andrew McCallum |
NIPS | 2 |
| 2008 | InterNano: e-Science for the Nanomanufacturing CommunityabstractAs network-enabled scholarship produces huge quantities of formal and informal research outputs in a variety of formats and varying levels of access, it is "enhanced" science that will facilitate the discovery, selection, and analysis of information that are a necessary part of the scientific research cycle particularly among interdisciplinary research communities. The national nanomanufacturing network (NNN) has developed a Web service, internano, to support the information needs of the nanomanufacturing community within this network-enabled research environment. Based on the clearinghouse concept, InterNano brings together heterogeneous resources and research outputs with networking and computational tools to enable discovery, facilitate selection of information, and encourage collaboration. With seventeen content features and services, internano's goal is not just to make information work within this community more efficient, but also to enable researchers to act on their discoveries within the same virtual framework. Rebecca Reznik-Zellen, Bob Stevens, Michael Thorn, Jeff Morse, Mark D. Smucker, James Allan 0001, David M. Mimno, Andrew McCallum, Mark Tuominen |
eScience | 7 |
| 2008 | Topic Models Conditioned on Arbitrary Features with Dirichlet-multinomial Regression
David M. Mimno, Andrew McCallum |
UAI | 1 |
| 2007 | Mixtures of hierarchical topics with Pachinko allocationabstractThe four-level pachinko allocation model (PAM) (Li & McCallum, 2006) represents correlations among topics using a DAG structure. It does not, however, represent a nested hierarchy of topics, with some topical word distributions representing the vocabulary that is shared among several more specific topics. This paper presents hierarchical PAM---an enhancement that explicitly represents a topic hierarchy. This model can be seen as combining the advantages of hLDA's topical hierarchy representation with PAM's ability to mix multiple leaves of the topic hierarchy. Experimental results show improvements in likelihood of held-out documents, as well as mutual information between automatically-discovered topics and humangenerated categories such as journals. David M. Mimno, Wei Li 0010, Andrew McCallum |
ICML | 1 |
| 2007 | Expertise modeling for matching papers with reviewersabstractAn essential part of an expert-finding task, such as matching reviewers to submitted papers, is the ability to model the expertise of a person based on documents. We evaluate several measures of the association between an author in an existing collection of research papers and a previously unseen document. We compare two language model based approaches with a novel topic model, Author-Persona-Topic (APT). In this model, each author can write under one or more "personas," which are represented as independent distributions over hidden topics. Examples of previous papers written by prospective reviewers are gathered from the Rexa database, which extracts and disambiguates author mentions from documents gathered from the web. We evaluate the models using a reviewer matching task based on human relevance judgments determining how well the expertise of proposed reviewers matches a submission. We find that the APT topic model outperforms the other models. David M. Mimno, Andrew McCallum |
KDD | 1 |