EDBT 2026 Demo / reviewers in the wild / expert
Mark Dredze
dblp:31/5468
· DBLP profile ↗
96ranked-venue papers
14as first author
17since 2021 · last 2026
0000-0002-0422-2474ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 11 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-authorHuman-computer interaction and ubiquitous computing · 12 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10Databases, data management, data science and information retrieval · 7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain Generalizable AI Guardrails with Augmented Policy TrainingabstractMinqian Liu, Ioana Baldini, David Rabinowitz, David S Rosenberg, Sebastian Gehrmann, Mark Dredze. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Minqian Liu, Ioana Baldini, David Rabinowitz, David S. Rosenberg, Sebastian Gehrmann, Mark Dredze |
ACL (1) | 6 |
| 2025 | Making FETCH! Happen: Finding Emergent Dog Whistles Through Common HabitatsabstractWARNING: This paper contains content that maybe upsetting or offensive to some readers.Dog whistles are coded expressions with dual meanings: one intended for the general public (outgroup) and another that conveys a specific message to an intended audience (ingroup).Often, these expressions are used to convey controversial political opinions while maintaining plausible deniability and slip by content moderation filters.Identification of dog whistles relies on curated lexicons, which have trouble keeping up to date.We introduce FETCH!, a task for finding emergent dog whistles in massive social media corpora.We find that stateof-the-art systems fail to achieve meaningful results across three distinct social media case studies.We present EarShot, a strong baseline system that combines the strengths of vector databases and Large Language Models (LLMs) to efficiently and effectively identify new dog whistles.1 Scenario Filtering Model Keyword Extraction Model Threshold Max n-grams Prec DPR F 0.5 Balanced RoBERTA R4 KeyBERT 50 1 14.81 1.65 5.70 LLaMa 13B KeyBERT 100 1 11.43 1.65 5.22 Synthetic RoBERTA R4 KeyBERT 800 1 14.86 13.45 14.55 LLaMA 8B KeyBERT 400 1 19.13 7.14 14.32 Kuleen Sasse, Carlos Alejandro Aguirre, Isabel Cachola, Sharon Levy, Mark Dredze |
ACL (1) | 5 |
| 2025 | LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity RecognitionabstractIn-context learning (ICL) enables large language models (LLMs) to perform new tasks using only a few demonstrations.In Named Entity Recognition (NER), demonstrations are typically selected based on semantic similarity to the test instance, ignoring training labels and resulting in suboptimal performance.We introduce DEER, a new method that leverages training labels through token-level statistics to improve ICL performance.DEER first enhances example selection with a label-guided, token-based retriever that prioritizes tokens most informative for entity recognition.It then prompts the LLM to revisit error-prone tokens, which are also identified using label statistics, and make targeted corrections.Evaluated on five NER datasets using four different LLMs, DEER consistently outperforms existing ICL methods and approaches the performance of supervised fine-tuning.Further analysis shows its effectiveness on both seen and unseen entities and its robustness in low-resource settings.1 Fan Bai 0006, Hamid Hassanzadeh, Ardavan Saeedi, Mark Dredze |
EMNLP | 4 |
| 2025 | Evaluating the Evaluators: Are readability metrics good measures of readability?abstractPlain Language Summarization (PLS) aims to distill complex documents into accessible summaries for non-expert audiences.In this paper, we conduct a thorough survey of PLS literature, and identify that the current standard practice for readability evaluation is to use traditional readability metrics, such as Flesch-Kincaid Grade Level (FKGL).However, despite proven utility in other fields, these metrics have not been compared to human readability judgments in PLS.We evaluate 8 readability metrics and show that most correlate poorly with human judgments, including the most popular metric, FKGL.We then show that Language Models (LMs) are better judges of readability, with the best-performing model achieving a Pearson correlation of 0.56 with human judgments.Extending our analysis to PLS datasets, which contain summaries aimed at non-expert audiences, we find that LMs better capture deeper measures of readability, such as required background knowledge, and lead to different conclusions than the traditional metrics.Based on these findings, we offer recommendations for best practices in the evaluation of plain language summaries.We release our analysis code and survey data. Isabel Cachola, Daniel Khashabi, Mark Dredze |
EMNLP | 3 |
| 2025 | DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text GenerationabstractThe decompose-then-verify strategy for verification of Large Language Model (LLM) generations decomposes claims that are then independently verified.Decontextualization augments text (claims) to ensure it can be verified outside of the original context, enabling reliable verification.While decomposition and decontextualization have been explored independently, their interactions in a complete system have not been investigated.Their conflicting purposes can create tensions: decomposition isolates atomic facts while decontextualization inserts relevant information.Furthermore, a decontextualized subclaim presents a challenge to the verification step: what part of the augmented text should be verified as it now contains multiple atomic facts?We conduct an evaluation of different decomposition, decontextualization, and verification strategies and find that the choice of strategy matters in the resulting factuality scores.Additionally, we introduce DNDSCORE, a decontextualization aware verification method that validates subclaims in the context of contextual information. Miriam Wanner, Benjamin Van Durme, Mark Dredze |
EMNLP | 3 |
| 2025 | RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language ModelsabstractBang An, Shiyue Zhang, Mark Dredze. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shiyue Zhang 0001, Mark Dredze |
NAACL (Long Papers) | 3 |
| 2025 | Benchmarking Large Language Models on Answering and Explaining Challenging Medical QuestionsabstractHanjie Chen, Zhouxiang Fang, Yash Singla, Mark Dredze. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhouxiang Fang, Yash Singla, Mark Dredze |
NAACL (Long Papers) | 4 |
| 2024 | Academics Can Contribute to Domain-Specialized Language ModelsabstractMark Dredze, Genta Indra Winata, Prabhanjan Kambadur, Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, David S Rosenberg, Sebastian Gehrmann. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Mark Dredze, Genta Indra Winata, Prabhanjan Kambadur, Ozan Irsoy, Steven Lu 0003, Vadim Dabravolski, David S. Rosenberg, Sebastian Gehrmann |
EMNLP | 1 |
| 2024 | Do LLMs Plan Like Human Writers? Comparing Journalist Coverage of Press Releases with LLMsabstractJournalists engage in multiple steps in the news writing process that depend on human creativity, like exploring different "angles" (i.e. the specific perspectives a reporter takes).These can potentially be aided by large language models (LLMs).By affecting planning decisions, such interventions can have an outsize impact on creative output.We advocate a careful approach to evaluating these interventions to ensure alignment with human values.In a case study of journalistic coverage of press releases, we assemble a large dataset of 250k press releases 1 and 650k articles covering them. 2 We develop methods to identify news articles that challenge and contextualize press releases.Finally, we evaluate suggestions made by LLMs for these articles and compare these with decisions made by human journalists.Our findings are three-fold: (1) Human-written news articles that challenge and contextualize press releases more take more creative angles and use more informational sources.(2) LLMs align better with humans when recommending angles, compared with informational sources.(3) Both the angles and sources LLMs suggest are significantly less creative than humans. Alexander Spangher, Nanyun Peng 0001, Sebastian Gehrmann, Mark Dredze |
EMNLP | 4 |
| 2023 | MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-EntropiesabstractShiyue Zhang, Shijie Wu, Ozan Irsoy, Steven Lu, Mohit Bansal, Mark Dredze, David Rosenberg. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shiyue Zhang 0001, Ozan Irsoy, Steven Lu 0003, Mohit Bansal, Mark Dredze, David S. Rosenberg |
ACL (1) | 6 |
| 2022 | Updated Headline Generation: Creating Updated Summaries for Evolving News StoriesabstractWe propose the task of updated headline generation, in which a system generates a headline for an updated article, considering both the previous article and headline.The system must identify the novel information in the article update, and modify the existing headline accordingly.We create data for this task using the NewsEdits corpus (Spangher and May, 2021) by automatically identifying contiguous article versions that are likely to require a substantive headline update.We find that models conditioned on the prior headline and body revisions produce headlines judged by humans to be as factual as gold headlines while making fewer unnecessary edits compared to a standard headline generation model.Our experiments establish benchmarks for this new contextual summarization task. Sheena Panthaplackel, Adrian Benton, Mark Dredze |
ACL (1) | 3 |
| 2022 | Bernice: A Multilingual Pre-trained Encoder for TwitterabstractThe language of Twitter differs significantly from that of other domains commonly included in large language model training.While tweets are typically multilingual and contain informal language, including emoji and hashtags, most pre-trained language models for Twitter are either monolingual, adapted from other domains rather than trained exclusively on Twitter, or are trained on a limited amount of in-domain Twitter data.We introduce Bernice, the first multilingual RoBERTa language model trained from scratch on 2.5 billion tweets with a custom tweet-focused tokenizer.We evaluate on a variety of monolingual and multilingual Twitter benchmarks, finding that our model consistently exceeds or matches the performance of a variety of models adapted to social media data as well as strong multilingual baselines, despite being trained on less data overall.We posit that it is more efficient compute-and data-wise to train completely on in-domain data with a specialized domain-specific tokenizer. Alexandra DeLucia, Aaron Mueller, Carlos Alejandro Aguirre, Philip Resnik, Mark Dredze |
EMNLP | 6 |
| 2021 | Gender and Racial Fairness in Depression Research using Social MediaabstractMultiple studies have demonstrated that behavior on internet-based social media platforms can be indicative of an individual's mental health status.The widespread availability of such data has spurred interest in mental health research from a computational lens.While previous research has raised concerns about possible biases in models produced from this data, no study has quantified how these biases actually manifest themselves with respect to different demographic groups, such as gender and racial/ethnic groups.Here, we analyze the fairness of depression classifiers trained on Twitter data with respect to gender and racial demographic groups.We find that model performance systematically differs for underrepresented groups and that these discrepancies cannot be fully explained by trivial data representation issues.Our study concludes with recommendations on how to avoid these biases in future research. Carlos Alejandro Aguirre, Keith Harrigian, Mark Dredze |
EACL | 3 |
| 2021 | Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information ExtractionabstractMahsa Yarmohammadi, Shijie Wu, Marc Marone, Haoran Xu, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, Benjamin Van Durme. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Mahsa Yarmohammadi, Marc Marone, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, Benjamin Van Durme |
EMNLP (1) | 12 |
| 2021 | Proxy Model Explanations for Time Series RNNsabstractWhile machine learning models can produce accurate predictions of complex real-world phenomena, domain experts may be unwilling to trust such a prediction without an explanation of the model’s behavior. This concern has motivated widespread research and produced many methods for interpreting black-box models. Many such methods explain predictions one-by-one, which can be slow and inconsistent across a large dataset, and ill-suited for time series applications. We introduce a proxy model approach that is fast to train, faithful to the original model, and globally consistent in its explanations. We compare our approach to several previous methods and find both that methods disagree with one another and that our approach improves over existing methods in an application to political event forecasting. Zach Wood-Doughty, Isabel Cachola, Mark Dredze |
ICMLA | 3 |
| 2021 | Fine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic ModelingabstractNeural topic models can augment or replace bag-of-words inputs with the learned representations of deep pre-trained transformer-based word prediction models.One added benefit when using representations from multilingual models is that they facilitate zero-shot polylingual topic modeling.However, while it has been widely observed that pre-trained embeddings should be fine-tuned to a given task, it is not immediately clear what supervision should look like for an unsupervised task such as topic modeling.Thus, we propose several methods for fine-tuning encoders to improve both monolingual and zero-shot polylingual neural topic modeling.We consider fine-tuning on auxiliary tasks, constructing a new topic classification task, integrating the topic classification objective directly into topic model training, and continued pre-training.We find that fine-tuning encoder representations on topic classification and integrating the topic classification task directly into topic modeling improves topic quality, and that fine-tuning encoder representations on any task is the most important factor for facilitating cross-lingual transfer. Aaron Mueller, Mark Dredze |
NAACL-HLT | 2 |
| 2021 | Demographic Representation and Collective Storytelling in the Me Too Twitter Hashtag Activism MovementabstractThe #MeToo movement on Twitter has drawn attention to the pervasive nature of sexual harassment and violence. While #MeToo has been praised for providing support for self-disclosures of harassment or violence and shifting societal response, it has also been criticized for exemplifying how women of color have been discounted for their historical contributions to and excluded from feminist movements. Through an analysis of over 600,000 tweets from over 256,000 unique users, we examine online #MeToo conversations across gender and racial/ethnic identities and the topics that each demographic emphasized. We found that tweets authored by white women were overrepresented in the movement compared to other demographics, aligning with criticism of unequal representation. We found that intersected identities contributed differing narratives to frame the movement, co-opted the movement to raise visibility in parallel ongoing movements, employed the same hashtags both critically and supportively, and revived and created new hashtags in response to pivotal moments. Notably, tweets authored by black women often expressed emotional support and were critical about differential treatment in the justice system and by police. In comparison, tweets authored by white women and men often highlighted sexual harassment and violence by public figures and weaved in more general political discussions. We discuss the implications of this work for digital activism research and design, including suggestions to raise visibility by those who were under-represented in this hashtag activism movement. Aaron Mueller, Zach Wood-Doughty, Silvio Amir, Mark Dredze, Alicia L. Nobles |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2020 | Sources of Transfer in Multilingual Named Entity RecognitionabstractNamed-entities are inherently multilingual, and annotations in any given language may be limited.This motivates us to consider polyglot named-entity recognition (NER), where one model is trained using annotated data drawn from more than one language.However, a straightforward implementation of this simple idea does not always work in practice: naive training of NER models using annotated data drawn from multiple languages consistently underperforms models trained on monolingual data alone, despite having access to more training data.The starting point of this paper is a simple solution to this problem, in which polyglot models are fine-tuned on monolingual data to consistently and significantly outperform their monolingual counterparts.To explain this phenomena, we explore the sources of multilingual transfer in polyglot NER models and examine the weight structure of polyglot models compared to their monolingual counterparts.We find that polyglot models efficiently share many parameters across languages and that fine-tuning may utilize a large number of those parameters. David Mueller, Nicholas Andrews, Mark Dredze |
ACL | 3 |
| 2020 | Clinical Concept Linking with Contextualized Neural RepresentationsabstractIn traditional approaches to entity linking, linking decisions are based on three sources of information -the similarity of the mention string to an entity's name, the similarity of the context of the document to the entity, and broader information about the knowledge base (KB).In some domains, there is little contextual information present in the KB and thus we rely more heavily on mention string similarity.We consider one example of this, concept linking, which seeks to link mentions of medical concepts to a medical concept ontology.We propose an approach to concept linking that leverages recent work in contextualized neural models, such as ELMo (Peters et al., 2018), which create a token representation that integrates the surrounding context of the mention and concept name.We find a neural ranking approach paired with contextualized embeddings provides gains over a competitive baseline (Leaman et al., 2013).Additionally, we find that a pre-training step using synonyms from the ontology offers a useful initialization for the ranker. Elliot Schumacher, Andriy Mulyar, Mark Dredze |
ACL | 3 |
| 2020 | Crowd-Diagnosis: When the Public Turns to Social Media to Obtain Clinical Diagnoses
Alicia L. Nobles, Eric C. Leas, Mark Dredze, Christopher A. Longhurst, Davey Smith, John W. Ayers |
AMIA | 3 |
| 2020 | Do Explicit Alignments Robustly Improve Multilingual Encoders?abstractMultilingual BERT (Devlin et al., 2019, mBERT), XLM-RoBERTa (Conneau et al., 2019, XLMR) and other unsupervised multilingual encoders can effectively learn crosslingual representation.Explicit alignment objectives based on bitexts like Europarl or Mul-tiUN have been shown to further improve these representations.However, word-level alignments are often suboptimal and such bitexts are unavailable for many languages.In this paper, we propose a new contrastive alignment objective that can better utilize such signal, and examine whether these previous alignment methods can be adapted to noisier sources of aligned data: a randomly sampled 1 million pair subset of the OPUS collection.Additionally, rather than report results on a single dataset with a single model run, we report the mean and standard derivation of multiple runs with different seeds, on four datasets and tasks.Our more extensive analysis finds that, while our new objective outperforms previous work, overall these methods do not improve performance with a more robust evaluation framework.Furthermore, the gains from using a better underlying model eclipse any benefits from alignment training.These negative results dictate more care in evaluating these methods and suggest limitations in applying explicit alignment objectives. Mark Dredze |
EMNLP (1) | 2 |
| 2020 | Examining Peer-to-Peer and Patient-Provider Interactions on a Social Media Community Facilitating Ask the Doctor Services
Alicia L. Nobles, Eric C. Leas, Mark Dredze, John W. Ayers |
ICWSM | 3 |
| 2020 | Aligning Public Feedback to Requests for Comments on Regulations.gov
Manya Wadhwa, Silvio Amir, Mark Dredze |
ICWSM | 3 |
| 2019 | Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERTabstractShijie Wu, Mark Dredze. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mark Dredze |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Identifying vulnerable older adult populations by contextualizing geriatric syndrome information in clinical notes of electronic health recordsabstractOBJECTIVE: Geriatric syndromes such as functional disability and lack of social support are often not encoded in electronic health records (EHRs), thus obscuring the identification of vulnerable older adults in need of additional medical and social services. In this study, we automatically identify vulnerable older adult patients with geriatric syndrome based on clinical notes extracted from an EHR system, and demonstrate how contextual information can improve the process. MATERIALS AND METHODS: We propose a novel end-to-end neural architecture to identify sentences that contain geriatric syndromes. Our model learns a representation of the sentence and augments it with contextual information: surrounding sentences, the entire clinical document, and the diagnosis codes associated with the document. We trained our system on annotated notes from 85 patients, tuned the model on another 50 patients, and evaluated its performance on the rest, 50 patients. RESULTS: Contextual information improved classification, with the most effective context coming from the surrounding sentences. At sentence level, our best performing model achieved a micro-F1 of 0.605, significantly outperforming context-free baselines. At patient level, our best model achieved a micro-F1 of 0.843. DISCUSSION: Our solution can be used to expand the identification of vulnerable older adults with geriatric syndromes. Since functional and social factors are often not captured by diagnosis codes in EHRs, the automatic identification of the geriatric syndrome can reduce disparities by ensuring consistent care across the older adult population. CONCLUSION: EHR free-text can be used to identify vulnerable older adults with a range of geriatric syndromes. Tao Chen 0008, Mark Dredze, Jonathan P. Weiner, Hadi Kharrazi |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | Challenges of Using Text Classifiers for Causal InferenceabstractCausal understanding is essential for many kinds of decision-making, but causal inference from observational data has typically only been applied to structured, low-dimensional datasets. While text classifiers produce low-dimensional outputs, their use in causal inference has not previously been studied. To facilitate causal analyses based on language data, we consider the role that text classifiers can play in causal inference through established modeling mechanisms from the causality literature on missing data and measurement error. We demonstrate how to conduct causal analyses using text classifiers on simulated and Yelp data, and discuss the opportunities and challenges of future work that uses text data in causal inference. Zach Wood-Doughty, Ilya Shpitser, Mark Dredze |
EMNLP | 3 |
| 2018 | Deep Dirichlet Multinomial RegressionabstractAdrian Benton, Mark Dredze. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Adrian Benton, Mark Dredze |
NAACL-HLT | 2 |
| 2017 | Bayesian Modeling of Lexical Resources for Low-Resource SettingsabstractLexical resources such as dictionaries and gazetteers are often used as auxiliary data for tasks such as part-of-speech induction and named-entity recognition.However, discriminative training with lexical features requires annotated data to reliably estimate the lexical feature weights and may result in overfitting the lexical features at the expense of features which generalize better.In this paper, we investigate a more robust approach: we stipulate that the lexicon is the result of an assumed generative process.Practically, this means that we may treat the lexical resources as observations under the proposed generative model.The lexical resources provide training data for the generative model without requiring separate data to estimate lexical feature weights.We evaluate the proposed approach in two settings: part-of-speech induction and lowresource named-entity recognition. Nicholas Andrews, Mark Dredze, Benjamin Van Durme, Jason Eisner |
ACL (1) | 2 |
| 2017 | Leveraging side information for speaker identification with the Enron conversational telephone speech collectionabstractSpeaker identification experiments typically focus on acoustic signals, but conversational speech often occurs in settings where additional useful side information may be available. This paper introduces a new distributable speaker identification test collection based on recorded telephone calls of Enron energy traders. Experiments with these recordings demonstrate that social network features and recording channel metadata can be used to reduce error rates in speaker identification below that achieved using acoustic evidence alone. Social network features from the parallel Enron email collection (37 of the 41 speakers in the telephone recordings sent or received emails in the collection) improve speaker identification, as do social network features computed using lightly supervised techniques to estimate a social network from more than one thousand unlabeled recordings. Ning Gao 0006, Gregory Sell, Douglas W. Oard, Mark Dredze |
ASRU | 4 |
| 2017 | Support for Interactive Identification of Mentioned Entities in Conversational SpeechabstractSearching conversational speech poses several new challenges, among which is how the searcher will make sense of what they find. This paper describes our initial experiments with a freely available collection of Enron telephone conversations. Our goal is to help the user make sense of search results by finding information about mentioned people, places and organizations. Because automated entity recognition is not yet sufficiently accurate on conversational telephone speech, we ask the user to transcribe just the name, and to indicate where in the recording it was heard. We then seek to link that mention to other mentions of the same entity in a variety of sources (in our experiments, in email and in Wikipedia). We cast this as an entity linking problem, and achieve promising results by utilizing social network features to help compensate for the limited accuracy of automatic transcription for this challenging content. Ning Gao 0006, Douglas W. Oard, Mark Dredze |
SIGIR | 3 |
| 2017 | Person entity linking in email with NIL detectionabstractFor each specific mention of an entity found in a text, the goal of entity linking is to determine whether the referenced entity is present in an existing knowledge base, and if so to determine which KB entity is the correct referent. Entity linking has been well explored for dissemination‐oriented sources such as news stories, blogs, and microblog posts, but the limited work to date on “conversational” sources such as email or text chat has not yet attempted to determine when the referent entity is not in the knowledge base (a task known as “NIL detection”). This article presents a supervised machine learning system for linking named mentions of people in email messages to a collection‐specific knowledge base, and that is also capable of NIL detection. This system learns from manually annotated training examples to leverage a rich set of features. The entity linking accuracy for entities present in the knowledge base is substantially and significantly better than the best previously reported results on the Enron email collection, comparable accuracy is reported for the challenging NIL detection task, and these results are for the first time replicated on a second email collection from a different source with comparable results. Ning Gao 0006, Mark Dredze, Douglas W. Oard |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Collective Supervision of Topic Models for Predicting Surveys with Social MediaabstractThis paper considers survey prediction from social media. We use topic models to correlate social media messages with survey outcomes and to provide an interpretable representation of the data. Rather than rely on fully unsupervised topic models, we use existing aggregated survey data to inform the inferred topics, a class of topic model supervision referred to as collective supervision. We introduce and explore a variety of topic model variants and provide an empirical analysis, with conclusions of the most effective models for this task. Adrian Benton, Michael J. Paul, Braden Hancock, Mark Dredze |
AAAI | 4 |
| 2016 | Discovering Shifts to Suicidal Ideation from Mental Health Content in Social MediaabstractHistory of mental illness is a major factor behind suicide risk and ideation. However research efforts toward characterizing and forecasting this risk is limited due to the paucity of information regarding suicide ideation, exacerbated by the stigma of mental illness. This paper fills gaps in the literature by developing a statistical methodology to infer which individuals could undergo transitions from mental health discourse to suicidal ideation. We utilize semi-anonymous support communities on Reddit as unobtrusive data sources to infer the likelihood of these shifts. We develop language and interactional measures for this purpose, as well as a propensity score matching based statistical approach. Our approach allows us to derive distinct markers of shifts to suicidal ideation. These markers can be modeled in a prediction framework to identify individuals likely to engage in suicidal ideation in the future. We discuss societal and ethical implications of this research. Munmun De Choudhury, Emre Kiciman, Mark Dredze, Glen A. Coppersmith, Mrinal Kumar 0004 |
CHI | 3 |
| 2016 | Geolocation for Twitter: Timing MattersabstractAutomated geolocation of social media messages can benefit a variety of downstream applications.However, these geolocation systems are typically evaluated without attention to how changes in time impact geolocation.Since different people, in different locations write messages at different times, these factors can significantly vary the performance of a geolocation system over time.We demonstrate cyclical temporal effects on geolocation accuracy in Twitter, as well as rapid drops as test data moves beyond the time period of training data.We show that temporal drift can effectively be countered with even modest online model updates. Mark Dredze, Miles Osborne, Prabhanjan Kambadur |
HLT-NAACL | 1 |
| 2016 | Embedding Lexical Features via Low-Rank TensorsabstractMo Yu, Mark Dredze, Raman Arora, Matthew R. Gormley. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Mo Yu, Mark Dredze, Raman Arora, Matthew R. Gormley |
HLT-NAACL | 2 |
| 2015 | Improved Relation Extraction with Feature-Rich Compositional Embedding ModelsabstractCompositional embedding models build a representation (or embedding) for a linguistic structure based on its component word embeddings.We propose a Feature-rich Compositional Embedding Model (FCM) for relation extraction that is expressive, generalizes to new domains, and is easy-to-implement.The key idea is to combine both (unlexicalized) handcrafted features with learned word embeddings.The model is able to directly tackle the difficulties met by traditional compositional embeddings models, such as handling arbitrary types of sentence annotations and utilizing global information for composition.We test the proposed model on two relation extraction tasks, and demonstrate that our model outperforms both previous compositional models and traditional feature rich models on the ACE 2005 relation extraction task, and the SemEval 2010 relation classification task.The combination of our model and a loglinear classifier with hand-crafted features gives state-of-the-art results.We made our implementation available for general use 1 . Matthew R. Gormley, Mo Yu, Mark Dredze |
EMNLP | 3 |
| 2015 | Named Entity Recognition for Chinese Social Media with Jointly Trained EmbeddingsabstractWe consider the task of named entity recognition for Chinese social media. The long line of work in Chinese NER has fo-cused on formal domains, and NER for social media has been largely restricted to English. We present a new corpus of Weibo messages annotated for both name and nominal mentions. Additionally, we evaluate three types of neural embeddings for representing Chinese text. Finally, we propose a joint training objective for the embeddings that makes use of both (NER) labeled and unlabeled raw text. Our meth-ods yield a 9 % improvement over a state-of-the-art baseline. 1 Nanyun Peng 0001, Mark Dredze |
EMNLP | 2 |
| 2015 | Entity Linking for Spoken LanguageabstractResearch on entity linking has considered a broad range of text, including newswire, blogs and web documents in multiple languages.However, the problem of entity linking for spoken language remains unexplored.Spoken language obtained from automatic speech recognition systems poses different types of challenges for entity linking; transcription errors can distort the context, and named entities tend to have high error rates.We propose features to mitigate these errors and evaluate the impact of ASR errors on entity linking using a new corpus of entity linked broadcast news transcripts. Adrian Benton, Mark Dredze |
HLT-NAACL | 2 |
| 2015 | A Concrete Chinese NLP PipelineabstractNanyun Peng, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Nanyun Peng 0001, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze |
HLT-NAACL | 11 |
| 2015 | Predicate Argument Alignment using a Global Coherence ModelabstractWe present a joint model for predicate argument alignment.We leverage multiple sources of semantic information, including temporal ordering constraints between events.These are combined in a max-margin framework to find a globally consistent view of entities and events across multiple documents, which leads to improvements over a very strong local baseline. Travis Wolfe, Mark Dredze, Benjamin Van Durme |
HLT-NAACL | 2 |
| 2015 | Combining Word Embeddings and Feature Embeddings for Fine-grained Relation ExtractionabstractCompositional embedding models build a rep-resentation for a linguistic structure based on its component word embeddings. While re-cent work has combined these word embed-dings with hand crafted features for improved performance, it was restricted to a small num-ber of features due to model complexity, thus limiting its applicability. We propose a new model that conjoins features and word em-beddings while maintaing a small number of parameters by learning feature embeddings jointly with the parameters of a compositional model. The result is a method that can scale to more features and more labels, while avoiding overfitting. We demonstrate that our model at-tains state-of-the-art results on ACE and ERE fine-grained relation extraction. 1 Mo Yu, Matthew R. Gormley, Mark Dredze |
HLT-NAACL | 3 |
| 2015 | Combining Search, Social Media, and Traditional Data Sources to Improve Influenza SurveillanceabstractWe present a machine learning-based methodology capable of providing real-time ("nowcast") and forecast estimates of influenza activity in the US by leveraging data from multiple data sources including: Google searches, Twitter microblogs, nearly real-time hospital visit records, and data from a participatory surveillance system. Our main contribution consists of combining multiple influenza-like illnesses (ILI) activity estimates, generated independently with each data source, into a single prediction of ILI utilizing machine learning ensemble approaches. Our methodology exploits the information in each data source and produces accurate weekly ILI predictions for up to four weeks ahead of the release of CDC's ILI reports. We evaluate the predictive ability of our ensemble approach during the 2013-2014 (retrospective) and 2014-2015 (live) flu seasons for each of the four weekly time horizons. Our ensemble approach demonstrates several advantages: (1) our ensemble method's predictions outperform every prediction using each data source independently, (2) our methodology can produce predictions one week ahead of GFT's real-time estimates with comparable accuracy, and (3) our two and three week forecast estimates have comparable accuracy to real-time predictions using an autoregressive model. Moreover, our results show that considerable insight is gained from incorporating disparate data streams, in the form of social media and crowd sourced data, into influenza predictions in all time horizons. Mauricio Santillana, André T. Nguyen, Mark Dredze, Michael J. Paul, Elaine O. Nsoesie, John S. Brownstein |
PLoS Comput. Biol. | 3 |
| 2015 | Approximation-Aware Dependency Parsing by Belief PropagationabstractWe show how to train the fast dependency parser of Smith and Eisner (2008) for improved accuracy. This parser can consider higher-order interactions among edges while retaining O( n3) runtime. It outputs the parse with maximum expected recall—but for speed, this expectation is taken under a posterior distribution that is constructed only approximately, using loopy belief propagation through structured factors. We show how to adjust the model parameters to compensate for the errors introduced by this approximation, by following the gradient of the actual loss on training data. We find this gradient by back-propagation. That is, we treat the entire parser (approximations and all) as a differentiable circuit, as others have done for loopy CRFs (Domke, 2010; Stoyanov et al., 2011; Domke, 2011; Stoyanov and Eisner, 2012). The resulting parser obtains higher accuracy with fewer iterations of belief propagation than one trained by conditional log-likelihood. Matthew R. Gormley, Mark Dredze, Jason Eisner |
Trans. Assoc. Comput. Linguistics | 2 |
| 2015 | SPRITE: Generalizing Topic Models with Structured PriorsabstractWe introduce Sprite, a family of topic models that incorporates structure into model priors as a function of underlying components. The structured priors can be constrained to model topic hierarchies, factorizations, correlations, and supervision, allowing Sprite to be tailored to particular settings. We demonstrate this flexibility by constructing a Sprite-based model to jointly infer topic hierarchies and author perspective, which we apply to corpora of political debates and online reviews. We show that the model learns intuitive topics, outperforming several other topic models at predictive tasks. Michael J. Paul, Mark Dredze |
Trans. Assoc. Comput. Linguistics | 2 |
| 2015 | Learning Composition Models for Phrase EmbeddingsabstractLexical embeddings can serve as useful representations for words for a variety of NLP tasks, but learning embeddings for phrases can be challenging. While separate embeddings are learned for each word, this is infeasible for every phrase. We construct phrase embeddings by learning how to compose word embeddings using features that capture phrase structure and context. We propose efficient unsupervised and task-specific learning objectives that scale our model to large datasets. We demonstrate improvements on both language modeling and several phrase semantic similarity tasks with various phrase lengths. We make the implementation of our model and the datasets available for general use. Mo Yu, Mark Dredze |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Robust Entity Clustering via Phylogenetic InferenceabstractEntity clustering must determine when two named-entity mentions refer to the same entity. Typical approaches use a pipeline ar-chitecture that clusters the mentions using fixed or learned measures of name and con-text similarity. In this paper, we propose a model for cross-document coreference res-olution that achieves robustness by learn-ing similarity from unlabeled data. The generative process assumes that each entity mention arises from copying and option-ally mutating an earlier name from a sim-ilar context. Clustering the mentions into entities depends on recovering this copying tree jointly with estimating models of the mutation process and parent selection pro-cess. We present a block Gibbs sampler for posterior inference and an empirical evalu-ation on several datasets. 1 Nicholas Andrews, Jason Eisner, Mark Dredze |
ACL (1) | 3 |
| 2014 | Low-Resource Semantic Role LabelingabstractWe explore the extent to which highresource manual annotations such as treebanks are necessary for the task of semantic role labeling (SRL).We examine how performance changes without syntactic supervision, comparing both joint and pipelined methods to induce latent syntax.This work highlights a new application of unsupervised grammar induction and demonstrates several approaches to SRL in the absence of supervised syntax.Our best models obtain competitive results in the high-resource setting and state-ofthe-art results in the low resource setting, reaching 72.48% F1 averaged across languages.We release our code for this work along with a larger toolkit for specifying arbitrary graphical structure.1 Matthew R. Gormley, Margaret Mitchell, Benjamin Van Durme, Mark Dredze |
ACL (1) | 4 |
| 2014 | Measuring Post Traumatic Stress Disorder in Twitter
Glen A. Coppersmith, Craig Harman, Mark Dredze |
ICWSM | 3 |
| 2014 | Facebook, Twitter and Google Plus for Breaking News: Is There a Winner?
Miles Osborne, Mark Dredze |
ICWSM | 2 |
| 2014 | A large-scale quantitative analysis of latent factors and sentiment in online doctor reviewsabstractOnline physician reviews are a massive and potentially rich source of information capturing patient sentiment regarding healthcare. We analyze a corpus comprising nearly 60,000 such reviews with a state-of-the-art probabilistic model of text. We describe a probabilistic generative model that captures latent sentiment across aspects of care (eg, interpersonal manner). We target specific aspects by leveraging a small set of manually annotated reviews. We perform regression analysis to assess whether model output improves correlation with state-level measures of healthcare. We report both qualitative and quantitative results. Model output correlates with state-level measures of quality healthcare, including patient likelihood of visiting their primary care physician within 14 days of discharge (p=0.03), and using the proposed model better predicts this outcome (p=0.10). We find similar results for healthcare expenditure. Generative models of text can recover important information from online physician reviews, facilitating large-scale analyses of such reviews. Byron C. Wallace, Michael J. Paul, Urmimala Sarkar, Thomas A. Trikalinos, Mark Dredze |
J. Am. Medical Informatics Assoc. | 5 |
| 2013 | Broadly Improving User Classification via Communication-Based Name and Location Clustering on Twitter
Shane Bergsma, Mark Dredze, Benjamin Van Durme, Theresa Wilson, David Yarowsky |
HLT-NAACL | 2 |
| 2013 | What's in a Domain? Multi-Domain Learning for Multi-Attribute Data
Mahesh Joshi, Mark Dredze, William W. Cohen, Carolyn P. Rosé |
HLT-NAACL | 2 |
| 2013 | Separating Fact from Fear: Tracking Flu Infections on Twitter
Alex Lamb, Michael J. Paul, Mark Dredze |
HLT-NAACL | 3 |
| 2013 | Drug Extraction from the Web: Summarizing Drug Experiences with Multi-Dimensional Topic Models
Michael J. Paul, Mark Dredze |
HLT-NAACL | 2 |
| 2013 | Topic Models and Metadata for Visualizing Text Corpora
Justin Snyder, Rebecca Knowles, Mark Dredze, Matthew R. Gormley, Travis Wolfe |
HLT-NAACL | 3 |
| 2013 | Adaptive regularization of weight vectors
Koby Crammer, Alex Kulesza, Mark Dredze |
Mach. Learn. | 3 |
| 2012 | Fast Syntactic Analysis for Statistical Language Modeling via Substructure Sharing and Uptraining
Ariya Rastrow, Mark Dredze, Sanjeev Khudanpur |
ACL (1) | 2 |
| 2012 | Twitter as a Source for Learning about Patient Safety Events
Ralph Passarella, Atul Nakhasi, Sarah G. Bell, Michael J. Paul, Peter Pronovost, Mark Dredze |
AMIA | 6 |
| 2012 | Name Phylogeny: A Generative Model of String Variation
Nicholas Andrews, Jason Eisner, Mark Dredze |
EMNLP-CoNLL | 3 |
| 2012 | Multi-Domain Learning: When Do Domains Matter?
Mahesh Joshi, Mark Dredze, William W. Cohen, Carolyn P. Rosé |
EMNLP-CoNLL | 2 |
| 2012 | New ℌ∞ bounds for the recursive least squares algorithm exploiting input structureabstractThe recursive least squares (RLS) algorithm is well known and has been widely used for many years. Most analyses of RLS have assumed statistical properties of the data or the noise process, but recent robust ℌ∞analyses have been used to bound the ratio of the performance of the algorithm to the total noise. In this paper, we provide an additive analysis bounding the difference between performance and noise. Our analysis provides additional convergence guarantees in general, and particular benefits for structured input data. We illustrate the analysis using human speech and white noise. Koby Crammer, Alex Kulesza, Mark Dredze |
ICASSP | 3 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 12 |
| 2012 | Efficient Structured Language Modeling for Speech RecognitionabstractThe structured language model (SLM) of [1] was one of the first to successfully integrate syntactic structure into language models. We extend the SLM framework in two new directions. First, we propose a new syntactic hierarchical interpolation that improves over previous approaches. Second, we develop a general information-theoretic algorithm for pruning the underlying Jelinek-Mercer interpolated LM used in [1], which substantially reduces the size of the LM, enabling us to train on large data. When combined with hill-climbing [2] the SLM is an accurate model, space-efficient and fast for rescoring large speech lattices. Experimental results on broadcast news demonstrate that the SLM outperforms a large 4-gram LM. 1. Ariya Rastrow, Mark Dredze, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2012 | Shared Components Topic Models
Matthew R. Gormley, Mark Dredze, Benjamin Van Durme, Jason Eisner |
HLT-NAACL | 2 |
| 2012 | Entity Clustering Across Languages
Spence Green, Nicholas Andrews, Matthew R. Gormley, Mark Dredze, Christopher D. Manning |
HLT-NAACL | 4 |
| 2012 | Factorial LDA: Sparse Multi-Dimensional Text ModelsabstractMulti-dimensional latent variable models can capture the many latent factors in a text corpus, such as topic, author perspective and sentiment. We introduce factorial LDA, a multi-dimensional latent variable model in which a document is influenced by K different factors, and each word token depends on a K-dimensional vector of latent variables. Our model incorporates structured word priors and learns a sparse product of factors. Experiments on research abstracts show that our model can learn latent factors such as research topic, scientific discipline, and focus (e.g. methods vs. applications.) Our modeling improvements reduce test perplexity and improve human interpretability of the discovered factors. Michael J. Paul, Mark Dredze |
NIPS | 2 |
| 2012 | Confidence-Weighted Linear Classification for Text Categorization
Koby Crammer, Mark Dredze, Fernando Pereira 0003 |
J. Mach. Learn. Res. | 2 |
| 2011 | Learning Sub-Word Units for Open Vocabulary Speech Recognition
Carolina Parada, Mark Dredze, Abhinav Sethy, Ariya Rastrow |
ACL | 2 |
| 2011 | Estimating document frequencies in a speech corpusabstractInverse Document Frequency (IDF) is an important quantity in many applications, including Information Retrieval. IDF is defined in terms of document frequency, df (w), the number of documents that mention w at least once. This quantity is relatively easy to compute over textual documents, but spoken documents are more challenging. This paper considers two baselines: (1) an estimate based on the 1-best ASR output and (2) an estimate based on expected term frequencies computed from the lattice. We improve over these baselines by taking advantage of repetition. Whatever the document is about is likely to be repeated, unlike ASR errors, which tend to be more random (Poisson). In addition, we find it helpful to consider an ensemble of language models. There is an opportunity for the ensemble to reduce noise, assuming that the errors across language models are relatively uncorrelated. The opportunity for improvement is larger when WER is high. This paper considers a pairing task application that could benefit from improved estimates of df. The pairing task inputs conversational sides from the English Fisher corpus and outputs estimates of which sides were from the same conversation. Better estimates of df lead to better performance on this task. Damianos Karakos, Mark Dredze, Kenneth Church 0001, Aren Jansen, Sanjeev Khudanpur |
ASRU | 2 |
| 2011 | Efficient discriminative training of long-span language modelsabstractLong-span language models, such as those involving syntactic dependencies, produce more coherent text than their n-gram counterparts. However, evaluating the large number of sentence-hypotheses in a packed representation such as an ASR lattice is intractable under such long-span models both during decoding and discriminative training. The accepted compromise is to rescore only the N-best hypotheses in the lattice using the long-span LM. We present discriminative hill climbing, an efficient and effective discriminative training procedure for long-span LMs based on a hill climbing rescoring algorithm [1]. We empirically demonstrate significant computational savings as well as error-rate reduction over N-best training methods in a state of the art ASR system for Broadcast News transcription. Ariya Rastrow, Mark Dredze, Sanjeev Khudanpur |
ASRU | 2 |
| 2011 | Adapting n-gram maximum entropy language models with conditional entropy regularizationabstractAccurate estimates of language model parameters are critical for building quality text generation systems, such as automatic speech recognition. However, text training data for a domain of interest is often unavailable. Instead, we use semi-supervised model adaptation; parameters are estimated using both unlabeled in-domain data (raw speech audio) and labeled out of domain data (text.) In this work, we present a new semi-supervised language model adaptation procedure for Maximum Entropy models with n-gram features. We augment the conventional maximum likelihood training criterion on out-of-domain text data with an additional term to minimize conditional entropy on in-domain audio. Additionally, we demonstrate how to compute conditional entropy efficiently on speech lattices using first- and second-order expectation semirings. We demonstrate improvements in terms of word error rate over other adaptation techniques when adapting a maximum entropy language model from broadcast news to MIT lectures. Ariya Rastrow, Mark Dredze, Sanjeev Khudanpur |
ASRU | 2 |
| 2011 | Hill climbing on speech lattices: A new rescoring frameworkabstractWe describe a new approach for rescoring speech lattices - with long-span language models or wide-context acoustic models - that does not entail computationally intensive lattice expansion or limited rescoring of only an N-best list. We view the set of word-sequences in a lattice as a discrete space equipped with the edit-distance metric, and develop a hill climbing technique to start with, say, the 1-best hypothesis under the lattice-generating model(s) and iteratively search a local neighborhood for the highest-scoring hypothesis under the rescoring model(s); such neighborhoods are efficiently constructed via finite state techniques. We demonstrate empirically that to achieve the same reduction in error rate using a better estimated, higher order language model, our technique evaluates fewer utterance-length hypotheses than conventional N-best rescoring by two orders of magnitude. For the same number of hypotheses evaluated, our technique results in a significantly lower error rate. Ariya Rastrow, Markus Dreyer, Abhinav Sethy, Sanjeev Khudanpur, Bhuvana Ramabhadran, Mark Dredze |
ICASSP | 6 |
| 2011 | You Are What You Tweet: Analyzing Twitter for Public Health
Michael J. Paul, Mark Dredze |
ICWSM | 2 |
| 2011 | OOV Sensitive Named-Entity Recognition in SpeechabstractNamed Entity Recognition (NER), an information extraction task, is typically applied to spoken documents by cascading a large vocabulary continuous speech recognizer (LVCSR) and a named entity tagger. Recognizing named entities in automatically decoded speech is difficult since LVCSR errors can confuse the tagger. This is especially true of out-of-vocabulary (OOV) words, which are often named entities and always produce transcription errors. In this work, we improve speech NER by including features indicative of OOVs based on a OOV detector, allowing for the identification of regions of speech containing named entities, even if they are incorrectly transcribed. We construct a new speech NER data set and demonstrate significant improvements for this task. Carolina Parada, Mark Dredze, Frederick Jelinek |
INTERSPEECH | 2 |
| 2010 | Entity Disambiguation for Knowledge Base Population
Mark Dredze, Paul McNamee, Delip Rao, Adam Gerber, Tim Finin |
COLING | 1 |
| 2010 | NLP on Spoken Documents Without ASR
Mark Dredze, Aren Jansen, Glen A. Coppersmith, Kenneth Church 0001 |
EMNLP | 1 |
| 2010 | We're Not in Kansas Anymore: Detecting Domain Changes in Streams
Mark Dredze, Tim Oates 0001, Christine D. Piatko |
EMNLP | 1 |
| 2010 | A spoken term detection framework for recovering out-of-vocabulary words using the webabstractVocabulary restrictions in large vocabulary continuous speech recognition (LVCSR) systems mean that out-of-vocabulary (OOV) words are lost in the output. However, OOV words tend to be information rich terms (often named entities) and their omission from the transcript negatively affects both usability and downstream NLP technologies, such as machine translation or knowledge distillation. We propose a novel approach to OOV recovery that uses a spoken term detection (STD) framework. Given an identified OOV region in the LVCSR output, we recover the uttered OOVs by utilizing contextual information and the vast and constantly updated vocabulary on the Web. Discovered words are integrated into system output, recovering up to 40 % of OOVs and resulting in a reduction in system error. Index Terms: language modeling, data selection, spoken term detection, oov detection Carolina Parada, Abhinav Sethy, Mark Dredze, Frederick Jelinek |
INTERSPEECH | 3 |
| 2010 | Contextual Information Improves OOV Detection in Speech
Carolina Parada, Mark Dredze, Denis Filimonov, Frederick Jelinek |
HLT-NAACL | 2 |
| 2010 | Multi-domain learning by confidence-weighted parameter combination
Mark Dredze, Alex Kulesza, Koby Crammer |
Mach. Learn. | 1 |
| 2009 | Multi-Class Confidence Weighted Algorithms
Koby Crammer, Mark Dredze, Alex Kulesza |
EMNLP | 2 |
| 2009 | Suggesting Email View Filters for Triage and Search
Mark Dredze, Bill N. Schilit, Peter Norvig |
IJCAI | 1 |
| 2009 | Adaptive Regularization of Weight VectorsabstractWe present AROW, a new online learning algorithm that combines several properties of successful : large margin training, confidence weighting, and the capacity to handle non-separable data. AROW performs adaptive regularization of the prediction function upon seeing each new instance, allowing it to perform especially well in the presence of label noise. We derive a mistake bound, similar in form to the second order perceptron bound, which does not assume separability. We also relate our algorithm to recent confidence-weighted online learning techniques and empirically show that AROW achieves state-of-the-art performance and notable robustness in the case of non-separable data. Koby Crammer, Alex Kulesza, Mark Dredze |
NIPS | 3 |
| 2008 | Intelligent Email: Aiding Users with AI
Mark Dredze, Hanna M. Wallach, Danny Puller, Tova Brooks, Josh Carroll, Joshua Magarick, John Blitzer, Fernando Pereira 0003 |
AAAI | 1 |
| 2008 | Reading the Markets: Forecasting Public Opinion of Political Candidates by News Analysis
Kevin Lerman, Ari Gilder, Mark Dredze, Fernando Pereira 0003 |
COLING | 3 |
| 2008 | Online Methods for Multi-Domain Learning and Adaptation
Mark Dredze, Koby Crammer |
EMNLP | 1 |
| 2008 | Confidence-weighted linear classificationabstractWe introduce confidence-weighted linear classifiers, which add parameter confidence information to linear classifiers. Online learners in this setting update both classifier parameters and the estimate of their confidence. The particular online algorithms we study here maintain a Gaussian distribution over parameter vectors and update the mean and covariance of the distribution with each instance. Empirical evaluation on a range of NLP tasks show that our algorithm improves over other state of the art online and batch methods, learns faster in the online setting, and lends itself to better classifier combination after parallel training. Mark Dredze, Koby Crammer, Fernando Pereira 0003 |
ICML | 1 |
| 2008 | Intelligent email: reply and attachment predictionabstractWe present two prediction problems under the rubric of Intelligent Email that are designed to support enhanced email interfaces that relieve the stress of email overload. Reply prediction alerts users when an email requires a response and facilitates email response management. Attachment prediction alerts users when they are about to send an email missing an attachment or triggers a document recommendation system, which can catch missing attachment emails before they are sent. Both problems use the same underlying email classification system and task specific features. Each task is evaluated for both single-user and cross-user settings. Mark Dredze, Tova Brooks, Josh Carroll, Joshua Magarick, John Blitzer, Fernando Pereira 0003 |
IUI | 1 |
| 2008 | Generating summary keywords for emails using topicsabstractEmail summary keywords, used to concisely represent the gist of an email, can help users manage and prioritize large numbers of messages. We develop an unsupervised learning framework for selecting summary keywords from emails using latent representations of the underlying topics in a user's mailbox. This approach selects words that describe each message in the context of existing topics rather than simply selecting keywords based on a single message in isolation. We present and compare four methods for selecting summary keywords based on two well-known models for inferring latent topics: latent semantic analysis and latent Dirichlet allocation. The quality of the summary keywords is assessed by generating summaries for emails from twelve users in the Enron corpus. The summary keywords are then used in place of entire messages in two proxy tasks: automated foldering and recipient prediction. We also evaluate the extent to which summary keywords enhance the information already available in a typical email user interface by repeating the same tasks using email subject lines. Mark Dredze, Hanna M. Wallach, Danny Puller, Fernando Pereira 0003 |
IUI | 1 |
| 2008 | Exact Convex Confidence-Weighted LearningabstractConfidence-weighted (CW) learning [6], an online learning method for linear classifiers, maintains a Gaussian distributions over weight vectors, with a covariance matrix that represents uncertainty about weights and correlations. Confidence constraints ensure that a weight vector drawn from the hypothesis distribution correctly classifies examples with a specified probability. Within this framework, we derive a new convex form of the constraint and analyze it in the mistake bound model. Empirical evaluation with both synthetic and text data shows our version of CW learning achieves lower cumulative and out-of-sample errors than commonly used first-order and second-order online methods. Koby Crammer, Mark Dredze, Fernando Pereira 0003 |
NIPS | 2 |
| 2007 | Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification
John Blitzer, Mark Dredze, Fernando Pereira 0003 |
ACL | 2 |
| 2007 | Frustratingly Hard Domain Adaptation for Dependency Parsing
Mark Dredze, John Blitzer, Partha P. Talukdar, Kuzman Ganchev, João Graça, Fernando Pereira 0003 |
EMNLP-CoNLL | 1 |
| 2006 | Activity-Centric Email: A Machine Learning Approach
Nicholas Kushmerick, Tessa A. Lau, Mark Dredze, Rinat Khoussainov |
AAAI | 3 |
| 2006 | Automatically classifying emails into activitiesabstractEmail-based activity management systems promise to give users better tools for managing increasing volumes of email, by organizing email according to a user's activities. Current activity management systems do not automatically classify incoming messages by the activity to which they belong, instead relying on simple heuristics (such as message threads), or asking the user to manually classify incoming messages as belonging to an activity. This paper presents several algorithms for automatically recognizing emails as part of an ongoing activity. Our baseline methods are the use of message reply-to threads to determine activity membership and a naïve Bayes classifier. Our SimSubset and SimOverlap algorithms compare the people involved in an activity against the recipients of each incoming message. Our SimContent algorithm uses IRR (a variant of latent semantic indexing) to classify emails into activities using similarity based on message contents. An empirical evaluation shows that each of these methods provide a significant improvement to the baseline methods. In addition, we show that a combined approach that votes the predictions of the individual methods performs better than each individual method alone. Mark Dredze, Tessa A. Lau, Nicholas Kushmerick |
IUI | 1 |
| 2003 | Beyond broadcastabstractThe work presented in this paper takes a novel approach to the task of providing information to viewers of broadcast news. Instead of considering the broadcast news as the end product, this work uses it as a starting point to dynamically build an information space for the user to explore. This information space is designed to satisfy the users information needs, by containing more breadth, depth, and points of view than the original broadcast story. The architecture and current implementation are discussed, and preliminary results from the analysis of some its components are presented Kevin Livingston, Mark Dredze, Kristian J. Hammond, Lawrence Birnbaum |
IUI | 2 |
| 2003 | Beyond broadcast: a demoabstractThis research discusses a method for delivering just-in-time information to television viewers to provide more depth and more breadth to television broadcasts. A novel aspect of this research is that it uses broadcast news as a starting point for gathering information regarding specific stories, as opposed to considering the broadcast version to be the end of the viewers exploration. This work is implemented in Cronkite, a system that provides viewers with expanded coverage of broadcast news stories. Kevin Livingston, Mark Dredze, Kristian J. Hammond, Lawrence Birnbaum |
IUI | 2 |