VLDB 2026 Research / reviewers in the wild / expert
Michael J. Paul
dblp:00/9714
· DBLP profile ↗
29ranked-venue papers
13as first author
0since 2021 · last 2020
0000-0002-9149-7539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 10 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Representation and self-supervised learning · 39% Information extraction and text analysis · 30% Transfer learning and domain adaptation · 15% | |
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 54% Data mining · 46% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Computational social science and digital humanities · 76% Medical and health informatics · 24% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › word representation
word embedding |
1.2 | 3 | 2020 | Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries · ACL 2020 Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation Models · ACL (1) 2019 A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
topic model |
1.1 | 5 | 2019 | Evaluating Topic Quality with Posterior Variability · EMNLP/IJCNLP (1) 2019 Diagnosing and Improving Topic Models by Analyzing Posterior Variability · AAAI 2018 Mixed Membership Markov Models for Unsupervised Conversation Modeling · EMNLP-CoNLL 2012 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
cross-lingual word embedding |
0.8 | 2 | 2020 | Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries · ACL 2020 A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity · ACL (1) 2019 |
Natural language and speech › Machine translation
bilingual lexicon induction |
0.4 | 1 | 2020 | Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries · ACL 2020 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.4 | 1 | 2019 | Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation Models · ACL (1) 2019 |
Machine learning › Transfer learning and domain adaptation › test-time adaptation
temporal adaptation |
0.4 | 1 | 2019 | Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation Models · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
text classification |
0.4 | 1 | 2019 | Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation Models · ACL (1) 2019 |
Graph algorithms and graph theory › graph clustering
network modularity |
0.4 | 1 | 2019 | A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity · ACL (1) 2019 |
Data mining › text mining
topic model |
0.3 | 2 | 2018 | Collective Supervision of Topic Models for Predicting Surveys with Social Media · AAAI 2016 Diagnosing and Improving Topic Models by Analyzing Posterior Variability · AAAI 2018 |
Information retrieval › information seeking
health information seeking |
0.2 | 1 | 2015 | Diagnoses, Decisions, and Outcomes: Web Search as Decision Support for Cancer · WWW 2015 |
Information retrieval › query understanding
query analysis |
0.2 | 1 | 2015 | Diagnoses, Decisions, and Outcomes: Web Search as Decision Support for Cancer · WWW 2015 |
Information retrieval › user behavior
search behavior analysis |
0.2 | 1 | 2015 | Diagnoses, Decisions, and Outcomes: Web Search as Decision Support for Cancer · WWW 2015 |
Natural language and speech › Question answering and dialogue systems
conversational modeling |
0.1 | 1 | 2012 | Mixed Membership Markov Models for Unsupervised Conversation Modeling · EMNLP-CoNLL 2012 |
Data mining › text mining › topic modeling
latent dirichlet allocation |
0.1 | 1 | 2012 | Factorial LDA: Sparse Multi-Dimensional Text Models · NIPS 2012 |
Data mining › text mining › topic modeling
sparse topic model |
0.1 | 1 | 2012 | Factorial LDA: Sparse Multi-Dimensional Text Models · NIPS 2012 |
Information retrieval
text analysis |
0.1 | 1 | 2012 | Factorial LDA: Sparse Multi-Dimensional Text Models · NIPS 2012 |
Data mining › text mining
topic modeling |
0.1 | 1 | 2012 | Factorial LDA: Sparse Multi-Dimensional Text Models · NIPS 2012 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
bayesian mixture model |
0.1 | 1 | 2010 | A Two-Dimensional Topic-Aspect Model for Discovering Multi-Faceted Topics · AAAI 2010 |
Natural language and speech › Language models and text generation › text summarization
opinion summarization |
0.1 | 1 | 2010 | Summarizing Contrastive Viewpoints in Opinionated Text · EMNLP 2010 |
Information retrieval
retrieval models |
0.1 | 1 | 2018 | Diagnosing and Improving Topic Models by Analyzing Posterior Variability · AAAI 2018 |
Computational social science and digital humanities › cultural analysis
cross-cultural analysis |
0.1 | 1 | 2009 | Cross-Cultural Analysis of Blogs and Forums with Mixed-Collection Topic Models · EMNLP 2009 |
Computational social science and digital humanities
text analysis |
0.0 | 1 | 2010 | A Two-Dimensional Topic-Aspect Model for Discovering Multi-Faceted Topics · AAAI 2010 |
Methods — techniques the papers use, named apart from their topics
intrinsic evaluation metric · 0.8latent dirichlet allocation · 0.7bayesian inference · 0.7collective supervision · 0.5timeline annotation · 0.4query log analysis · 0.4retrofitting · 0.4linear projection · 0.4posterior variability · 0.4neural classification model · 0.4domain adaptation · 0.4bayesian topic model · 0.4variational inference · 0.3structured word priors · 0.1mixed membership markov models · 0.1bayesian mixture model · 0.1mixed-collection topic models · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to DictionariesabstractCross-lingual word embeddings (CLWE) are often evaluated on bilingual lexicon induction (BLI).Recent CLWE methods use linear projections, which underfit the training dictionary, to generalize on BLI.However, underfitting can hinder generalization to other downstream tasks that rely on words from the training dictionary.We address this limitation by retrofitting CLWE to the training dictionary, which pulls training translation pairs closer in the embedding space and overfits the training dictionary.This simple post-processing step often improves accuracy on two downstream tasks, despite lowering BLI test accuracy.We also retrofit to both the training dictionary and a synthetic dictionary induced from CLWE, which sometimes generalizes even better on downstream tasks.Our results confirm the importance of fully exploiting the training dictionary in downstream tasks and explains why BLI is a flawed CLWE evaluation. Mozhi Zhang, Yoshinari Fujinuma, Michael J. Paul, Jordan L. Boyd-Graber |
ACL | 3 |
| 2020 | Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech RecognitionabstractExisting research on fairness evaluation of document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. In this work, we assemble and publish a multilingual Twitter corpus for the task of hate speech detection with inferred four author demographic factors: age, country, gender and race/ethnicity. The corpus covers five languages: English, Italian, Polish, Portuguese and Spanish. We evaluate the inferred demographic labels with a crowdsourcing platform, Figure Eight. To examine factors that can cause biases, we take an empirical analysis of demographic predictability on the English corpus. We measure the performance of four popular document classifiers and evaluate the fairness and bias of the baseline classifiers on the author-level demographic attributes. Xiaolei Huang 0002, Linzi Xing, Franck Dernoncourt, Michael J. Paul |
LREC | 4 |
| 2020 | An Empirical Study on Crosslingual Transfer in Probabilistic Topic ModelsabstractProbabilistic topic modeling is a common first step in crosslingual tasks to enable knowledge transfer and extract multilingual features. Although many multilingual topic models have been developed, their assumptions about the training corpus are quite varied, and it is not clear how well the different models can be utilized under various training conditions. In this article, the knowledge transfer mechanisms behind different multilingual topic models are systematically studied, and through a broad set of experiments with four models on ten languages, we provide empirical insights that can inform the selection and future development of multilingual topic models. Shudong Hao, Michael J. Paul |
Comput. Linguistics | 2 |
| 2019 | A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph ModularityabstractCross-lingual word embeddings encode the meaning of words from different languages into a shared low-dimensional space. An important requirement for many downstream tasks is that word similarity should be independent of language - i.e., word vectors within one language should not be more similar to each other than to words in another language. We measure this characteristic using modularity, a network measurement that measures the strength of clusters in a graph. Modularity has a moderate to strong correlation with three downstream tasks, even though modularity is based only on the structure of embeddings and does not require any external resources. We show through experiments that modularity can serve as an intrinsic validation metric to improve unsupervised cross-lingual word embeddings, particularly on distant language pairs in low-resource settings. Yoshinari Fujinuma, Jordan L. Boyd-Graber, Michael J. Paul |
ACL (1) | 3 |
| 2019 | Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation ModelsabstractLanguage usage can change across periods of time, but document classifiers models are usually trained and tested on corpora spanning multiple years without considering temporal variations.This paper describes two complementary ways to adapt classifiers to shifts across time.First, we show that diachronic word embeddings, which were originally developed to study language change, can also improve document classification, and we show a simple method for constructing this type of embedding.Second, we propose a time-driven neural classification model inspired by methods for domain adaptation.Experiments on six corpora show how these methods can make classifiers more robust over time. Xiaolei Huang 0002, Michael J. Paul |
ACL (1) | 2 |
| 2019 | Evaluating Topic Quality with Posterior VariabilityabstractLinzi Xing, Michael J. Paul, Giuseppe Carenini. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Linzi Xing, Michael J. Paul, Giuseppe Carenini |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Diagnosing and Improving Topic Models by Analyzing Posterior VariabilityabstractBayesian inference methods for probabilistic topic models can quantify uncertainty in the parameters, which has primarily been used to increase the robustness of parameter estimates. In this work, we explore other rich information that can be obtained by analyzing the posterior distributions in topic models. Experimenting with latent Dirichlet allocation on two datasets, we propose ideas incorporating information about the posterior distributions at the topic level and at the word level. At the topic level, we propose a metric called topic stability that measures the variability of the topic parameters under the posterior. We show that this metric is correlated with human judgments of topic quality as well as with the consistency of topics appearing across multiple models. At the word level, we experiment with different methods for adjusting individual word probabilities within topics based on their uncertainty. Humans prefer words ranked by our adjusted estimates nearly twice as often when compared to the traditional approach. Finally, we describe how the ideas presented in this work could potentially applied to other predictive or exploratory models in future work. Linzi Xing, Michael J. Paul |
AAAI | 2 |
| 2018 | Learning Multilingual Topics from Incomparable CorporaabstractMultilingual topic models enable crosslingual tasks by extracting consistent topics from multilingual corpora. Most models require parallel or comparable training corpora, which limits their ability to generalize. In this paper, we first demystify the knowledge transfer mechanism behind multilingual topic models by defining an alternative but equivalent formulation. Based on this analysis, we then relax the assumption of training data required by most existing models, creating a model that only requires a dictionary for training. Experiments show that our new method effectively learns coherent multilingual topics from partially and fully incomparable corpora with limited amounts of dictionary resources. Shudong Hao, Michael J. Paul |
COLING | 2 |
| 2018 | Lessons from the Bible on Modern Topics: Low-Resource Multilingual Topic Model EvaluationabstractShudong Hao, Jordan Boyd-Graber, Michael J. Paul. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shudong Hao, Jordan L. Boyd-Graber, Michael J. Paul |
NAACL-HLT | 3 |
| 2017 | Feature Selection as Causal Inference: Experiments with Text ClassificationabstractThis paper proposes a matching technique for learning causal associations between word features and class labels in document classification. The goal is to identify more meaningful and generalizable features than with only correlational approaches. Experiments with sentiment classification show that the proposed method identifies interpretable word associations with sentiment and improves classification performance in a majority of cases. The proposed feature selection method is particularly effective when applied to out-of-domain data. Michael J. Paul |
CoNLL | 1 |
| 2016 | Collective Supervision of Topic Models for Predicting Surveys with Social MediaabstractThis paper considers survey prediction from social media. We use topic models to correlate social media messages with survey outcomes and to provide an interpretable representation of the data. Rather than rely on fully unsupervised topic models, we use existing aggregated survey data to inform the inferred topics, a class of topic model supervision referred to as collective supervision. We introduce and explore a variety of topic model variants and provide an empirical analysis, with conclusions of the most effective models for this task. Adrian Benton, Michael J. Paul, Braden Hancock, Mark Dredze |
AAAI | 2 |
| 2016 | Search and Breast Cancer: On Episodic Shifts of Attention over Life Histories of an IllnessabstractWe seek to understand the evolving needs of people who are faced with a life-changing medical diagnosis based on analyses of queries extracted from an anonymized search query log. Focusing on breast cancer, we manually tag a set of Web searchers as showing patterns of search behavior consistent with someone grappling with the screening, diagnosis, and treatment of breast cancer. We build and apply probabilistic classifiers to detect these searchers from multiple sessions and to identify the timing of diagnosis using temporal and statistical features. We explore the changes in information seeking over time before and after an inferred diagnosis of breast cancer by aligning multiple searchers by the estimated time of diagnosis. We employ the classifier to automatically identify 1,700 candidate searchers with an estimated 90% precision, and we predict the day of diagnosis within 15 days with an 88% accuracy. We show that the geographic and demographic attributes of searchers identified with high probability are strongly correlated with ground truth of reported incidence rates. We then analyze the content of queries over time for inferred cancer patients, using a detailed ontology of cancer-related search terms. The analysis reveals the rich temporal structure of the evolving queries of people likely diagnosed with breast cancer. Finally, we focus on subtypes of illness based on inferred stages of cancer and show clinically relevant dynamics of information seeking based on the dominant stage expressed by searchers. Michael J. Paul, Ryen W. White, Eric Horvitz |
ACM Trans. Web | 1 |
| 2015 | Diagnoses, Decisions, and Outcomes: Web Search as Decision Support for CancerabstractPeople diagnosed with a serious illness often turn to the Web for their rising information needs, especially when decisions are required. We analyze the search and browsing behavior of searchers who show a surge of interest in prostate cancer. Prostate cancer is the most common serious cancer in men and is a leading cause of cancer-related death. Diagnoses of prostate cancer typically involve reflection and decision making about treatment based on assessments of preferences and outcomes. We annotated timelines of treatment-related queries from nearly 300 searchers with tags indicating different phases of treatment, including decision making, preparation, and recovery. Using this corpus, we present a variety of analyses toward the goal of understanding search and decision making about treatments. We characterize search queries and the content of accessed pages for different treatment phases, model search behavior during the decision-making phase, and create an aggregate alignment of treatment timelines illustrated with a variety of visualizations. The experiments provide insights about how people who are engaged in intensive searches about prostate cancer over an extended period of time pursue and access information from the Web. Michael J. Paul, Ryen W. White, Eric Horvitz |
WWW | 1 |
| 2015 | Combining Search, Social Media, and Traditional Data Sources to Improve Influenza SurveillanceabstractWe present a machine learning-based methodology capable of providing real-time ("nowcast") and forecast estimates of influenza activity in the US by leveraging data from multiple data sources including: Google searches, Twitter microblogs, nearly real-time hospital visit records, and data from a participatory surveillance system. Our main contribution consists of combining multiple influenza-like illnesses (ILI) activity estimates, generated independently with each data source, into a single prediction of ILI utilizing machine learning ensemble approaches. Our methodology exploits the information in each data source and produces accurate weekly ILI predictions for up to four weeks ahead of the release of CDC's ILI reports. We evaluate the predictive ability of our ensemble approach during the 2013-2014 (retrospective) and 2014-2015 (live) flu seasons for each of the four weekly time horizons. Our ensemble approach demonstrates several advantages: (1) our ensemble method's predictions outperform every prediction using each data source independently, (2) our methodology can produce predictions one week ahead of GFT's real-time estimates with comparable accuracy, and (3) our two and three week forecast estimates have comparable accuracy to real-time predictions using an autoregressive model. Moreover, our results show that considerable insight is gained from incorporating disparate data streams, in the form of social media and crowd sourced data, into influenza predictions in all time horizons. Mauricio Santillana, André T. Nguyen, Mark Dredze, Michael J. Paul, Elaine O. Nsoesie, John S. Brownstein |
PLoS Comput. Biol. | 4 |
| 2015 | SPRITE: Generalizing Topic Models with Structured PriorsabstractWe introduce Sprite, a family of topic models that incorporates structure into model priors as a function of underlying components. The structured priors can be constrained to model topic hierarchies, factorizations, correlations, and supervision, allowing Sprite to be tailored to particular settings. We demonstrate this flexibility by constructing a Sprite-based model to jointly infer topic hierarchies and author perspective, which we apply to corpora of political debates and online reviews. We show that the model learns intuitive topics, outperforming several other topic models at predictive tasks. Michael J. Paul, Mark Dredze |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | A large-scale quantitative analysis of latent factors and sentiment in online doctor reviewsabstractOnline physician reviews are a massive and potentially rich source of information capturing patient sentiment regarding healthcare. We analyze a corpus comprising nearly 60,000 such reviews with a state-of-the-art probabilistic model of text. We describe a probabilistic generative model that captures latent sentiment across aspects of care (eg, interpersonal manner). We target specific aspects by leveraging a small set of manually annotated reviews. We perform regression analysis to assess whether model output improves correlation with state-level measures of healthcare. We report both qualitative and quantitative results. Model output correlates with state-level measures of quality healthcare, including patient likelihood of visiting their primary care physician within 14 days of discharge (p=0.03), and using the proposed model better predicts this outcome (p=0.10). We find similar results for healthcare expenditure. Generative models of text can recover important information from online physician reviews, facilitating large-scale analyses of such reviews. Byron C. Wallace, Michael J. Paul, Urmimala Sarkar, Thomas A. Trikalinos, Mark Dredze |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Separating Fact from Fear: Tracking Flu Infections on Twitter
Alex Lamb, Michael J. Paul, Mark Dredze |
HLT-NAACL | 2 |
| 2013 | Drug Extraction from the Web: Summarizing Drug Experiences with Multi-Dimensional Topic Models
Michael J. Paul, Mark Dredze |
HLT-NAACL | 1 |
| 2012 | Twitter as a Source for Learning about Patient Safety Events
Ralph Passarella, Atul Nakhasi, Sarah G. Bell, Michael J. Paul, Peter Pronovost, Mark Dredze |
AMIA | 4 |
| 2012 | Mixed Membership Markov Models for Unsupervised Conversation Modeling
Michael J. Paul |
EMNLP-CoNLL | 1 |
| 2012 | Implicitly Intersecting Weighted Automata using Dual Decomposition
Michael J. Paul, Jason Eisner |
HLT-NAACL | 1 |
| 2012 | Factorial LDA: Sparse Multi-Dimensional Text ModelsabstractMulti-dimensional latent variable models can capture the many latent factors in a text corpus, such as topic, author perspective and sentiment. We introduce factorial LDA, a multi-dimensional latent variable model in which a document is influenced by K different factors, and each word token depends on a K-dimensional vector of latent variables. Our model incorporates structured word priors and learns a sparse product of factors. Experiments on research abstracts show that our model can learn latent factors such as research topic, scientific discipline, and focus (e.g. methods vs. applications.) Our modeling improvements reduce test perplexity and improve human interpretability of the discovered factors. Michael J. Paul, Mark Dredze |
NIPS | 1 |
| 2011 | You Are What You Tweet: Analyzing Twitter for Public Health
Michael J. Paul, Mark Dredze |
ICWSM | 1 |
| 2011 | Hierarchical Bayesian Models for Latent Attribute Detection in Social Media
Delip Rao, Michael J. Paul, Clayton Fink, David Yarowsky, Tim Oates 0001, Glen A. Coppersmith |
ICWSM | 2 |
| 2011 | Modeling reciprocity in social interactions with probabilistic latent space modelsabstractAbstract Reciprocity is a pervasive concept that plays an important role in governing people's behavior, judgments, and thus their social interactions. In this paper we present an analysis of the concept of reciprocity as expressed in English and a way to model it. At a larger structural level the reciprocity model will induce representations and clusters of relations between interpersonal verbs. In particular, we introduce an algorithm that semi-automatically discovers patterns encoding reciprocity based on a set of simple yet effective pronoun templates. Using the most frequently occurring patterns we queried the web and extracted 13,443 reciprocal instances, which represent a broad-coverage resource. Unsupervised clustering procedures are performed to generate meaningful semantic clusters of reciprocal instances. We also present several extensions (along with observations) to these models that incorporate meta-attributes like the verbs' affective value, identify gender differences between participants, consider the textual context of the instances, and automatically discover verbs with certain presuppositions. The pattern discovery procedure yields an accuracy of 97 per cent, while the clustering procedures – clustering with pairwise membership and clustering with transitions – indicate accuracies of 91 per cent and 64 per cent, respectively. Our affective value clustering can predict an unknown verb's affective value (positive, negative, or neutral) with 51 per cent accuracy, while it can discriminate between positive and negative values with 68 per cent accuracy. The presupposition discovery procedure yields an accuracy of 97 per cent. Roxana Girju, Michael J. Paul |
Nat. Lang. Eng. | 2 |
| 2010 | A Two-Dimensional Topic-Aspect Model for Discovering Multi-Faceted TopicsabstractThis paper presents the Topic-Aspect Model (TAM), a Bayesian mixture model which jointly discovers topics and aspects. We broadly define an aspect of a document as a characteristic that spans the document, such as an underlying theme or perspective. Unlike previous models which cluster words by topic or aspect, our model can generate token assignments in both of these dimensions, rather than assuming words come from only one of two orthogonal models. We present two applications of the model. First, we model a corpus of computational linguistics abstracts, and find that the scientific topics identified in the data tend to include both a computational aspect and a linguistic aspect. For example, the computational aspect of GRAMMAR emphasizes parsing, whereas the linguistic aspect focuses on formal languages. Secondly, we show that the model can capture different viewpoints on a variety of topics in a corpus of editorials about the Israeli-Palestinian conflict. We show both qualitative and quantitative improvements in TAM over two other state-of-the-art topic models. Michael J. Paul, Roxana Girju |
AAAI | 1 |
| 2010 | Summarizing Contrastive Viewpoints in Opinionated Text
Michael J. Paul, ChengXiang Zhai, Roxana Girju |
EMNLP | 1 |
| 2009 | Mining the Web for Reciprocal Relationships
Michael J. Paul, Roxana Girju |
CoNLL | 1 |
| 2009 | Cross-Cultural Analysis of Blogs and Forums with Mixed-Collection Topic Models
Michael J. Paul, Roxana Girju |
EMNLP | 1 |