Saif M. Mohammad

dblp:58/380 · DBLP profile ↗
← Back
65ranked-venue papers
30as first author
23since 2021 · last 2026
0000-0003-2716-7516ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 28 first-author · 23 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence
abstract
Warmth (W) (often further broken down into Trust (T) and Sociability (S)) and Competence (C) are central dimensions along which people evaluate individuals and social groups (Fiske, 2018).While these constructs are well established in social psychology, they are only starting to get attention in NLP research through word-level lexicons, which do not fully capture their contextual expression in larger text units and discourse.In this work, we introduce Warmth and Competence Sentences (W&C-Sent), the first sentence-level dataset annotated for warmth and competence.The dataset includes over 1,600 English sentence-target pairs annotated along three dimensions: trust and sociability (components of warmth), and competence 1 .The sentences in W&C-Sent are social media posts that express attitudes and opinions about specific individuals or social groups (the targets of our annotations).We describe the data collection, annotation, and quality-control procedures in detail, and evaluate a range of large language models (LLMs) on their ability to identify trust, sociability, and competence in text.W&C-Sent provides a new resource for analyzing warmth and competence in language and supports future research at the intersection of NLP and computational social science.
Mutaz Ayesh, Saif M. Mohammad, Nedjma Ousidhoum
ACL (1)2
2026 DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
abstract
Lung-Hao Lee, Liang-Chih Yu, Natalia V Loukachevitch, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng, Jin Wang, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lung-Hao Lee, Liang-Chih Yu, Natalia V. Loukachevitch, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng 0001, Jin Wang 0008, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad
ACL (1)16
2026 From Trial by Fire to Sleep like a Baby: A Lexicon of Anxiety Associations for 20K English Multi-Word Expressions
Saif M. Mohammad
LREC1
2026 I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
Shijia Zhou, Saif M. Mohammad, Barbara Plank, Diego Frassinelli
LREC2
2025 Words of Warmth: Trust and Sociability Norms for over 26k English Words
abstract
Social psychologists have shown that Warmth (W) and Competence (C) are the primary dimensions along which we assess other people and groups.These dimensions impact various aspects of our lives from social competence and emotion regulation to success in the work place and how we view the world.More recent work has started to explore how these dimensions develop, why they have developed, and what they constitute.Of particular note, is the finding that warmth has two distinct components: Trust (T) and Sociability (S).In this work, we introduce Words of Warmth, the first large-scale repository of manually derived word-warmth (as well as word-trust and word-sociability) associations for over 26k English words.We show that the associations are highly reliable.We use the lexicons to study the rate at which children acquire WCTS words with age.Finally, we show that the lexicon enables a wide variety of bias and stereotype research through case studies on various target entities.Words of Warmth is freely available at: http://saifmohammad.com/warmth.html
Saif M. Mohammad
ACL (1)1
2025 BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages
abstract
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava 0001, Aura Cristina Udrea, Lilian Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou 0019, Saif M. Mohammad
ACL (1)48
2025 Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
abstract
Language is a form of symbolic capital that affects people's lives in many ways (Bourdieu, 1977(Bourdieu, , 1991)).As a powerful means of communication, it reflects identities, cultures, traditions, and societies more broadly.Therefore, data in a given language should be regarded as more than just a collection of tokens.Rigorous data collection and labeling practices are essential for developing more human-centered and socially aware technologies.Although there has been growing interest in under-resourced languages within the NLP community, work in this area faces unique challenges, such as data scarcity and limited access to qualified annotators.In this paper, we collect feedback from individuals directly involved in and impacted by NLP artefacts for medium-and low-resource languages.We conduct both quantitative and qualitative analyses of their responses and highlight key issues related to: (1) data quality, including linguistic and cultural appropriateness; and (2) the ethics of common annotation practices, such as the misuse of participatory research.Based on these findings, we make several recommendations for creating high-quality language artefacts that reflect the cultural milieu of their speakers, while also respecting the dignity and labor of data workers.
Nedjma Ousidhoum, Meriem Beloucif, Saif M. Mohammad
ACL (1)3
2025 The Nature of NLP: Analyzing Contributions in NLP Papers
abstract
Natural Language Processing (NLP) is an established and dynamic field.Despite this, what constitutes NLP research remains debated.In this work, we address the question by quantitatively examining NLP research papers.We propose a taxonomy of research contributions and introduce NLPContributions, a dataset of nearly 2k NLP research paper abstracts, carefully annotated to identify scientific contributions and classify their types according to this taxonomy.We also introduce a novel task of automatically identifying contribution statements and classifying their types from research papers.We present experimental results for this task and apply our model to ∼29k NLP research papers to analyze their contributions, aiding in the understanding of the nature of NLP research.We show that NLP research has taken a winding path -with the focus on language and human-centric studies being prominent in the 1970s and 80s, tapering off in the 1990s and 2000s, and starting to rise again since the late 2010s.Alongside this revival, we observe a steady rise in dataset and methodological contributions since the 1990s, such that today, on average, individual NLP papers contribute in more ways than ever before.Our dataset and analyses offer a powerful lens for tracing research trends and offer potential for generating informed, datadriven literature surveys. 1
Aniket Pramanick, Yufang Hou 0001, Saif M. Mohammad, Iryna Gurevych
ACL (1)3
2025 Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields
abstract
This study examines the tendency to cite older work across 20 fields of study over 43 years (1980–2023). We put NLP’s propensity to cite older work in the context of these 20 other fields to analyze whether NLP shows similar temporal citation patterns to them over time or whether differences can be observed. Our analysis, based on a dataset of ~240 million papers, reveals a broader scientific trend: many fields have markedly declined in citing older works (e.g., psychology, computer science). The trend is strongest in NLP and ML research (-12.8% and -5.5% in citation age from previous peaks). Our results suggest that citing more recent works is not directly driven by the growth in publication rates (-3.4% across fields; -5.2% in humanities; -5.5% in formal sciences) — even when controlling for an increase in the volume of papers. Our findings raise questions about the scientific community’s engagement with past literature, particularly for NLP, and the potential consequences of neglecting older but relevant research. The data and a demo showcasing our results are publicly available.
Jan Philip Wahle, Terry Ruas, Mohamed Abdalla 0001, Bela Gipp, Saif M. Mohammad
COLING5
2024 Emotion Granularity from Text: An Aggregate-Level Indicator of Mental Health
abstract
We are united in how emotions are central to shaping our experiences; yet, individuals differ greatly in how we each identify, categorize, and express emotions.In psychology, variation in the ability of individuals to differentiate between emotion concepts is called emotion granularity (determined through self-reports of one's emotions).High emotion granularity has been linked with better mental and physical health; whereas low emotion granularity has been linked with maladaptive emotion regulation strategies and poor health outcomes.In this work, we propose computational measures of emotion granularity derived from temporallyordered speaker utterances in social media (in lieu of self-reports that suffer from various biases).We then investigate the effectiveness of such text-derived measures of emotion granularity in functioning as markers of various mental health conditions (MHCs).We establish baseline measures of emotion granularity derived from textual utterances, and show that, at an aggregate level, emotion granularities are significantly lower for people self-reporting as having an MHC than for the control population.This paves the way towards a better understanding of the MHCs, and specifically the role emotions play in our well-being.
Krishnapriya Vishnubhotla, Daniela Teodorescu, Mallory J. Feldman, Kristen A. Lindquist, Saif M. Mohammad
EMNLP5
2023 The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research
abstract
Mohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aurélie Névéol, Fanny Ducel, Saif Mohammad, Karen Fort. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mohamed Abdalla 0001, Jan Philip Wahle, Terry Ruas, Aurélie Névéol, Fanny Ducel, Saif M. Mohammad, Karën Fort
ACL (1)6
2023 Forgotten Knowledge: Examining the Citational Amnesia in NLP
abstract
Citing papers is the primary method through which modern scientific writing discusses and builds on past work.Collectively, citing a diverse set of papers (in time and area of study) is an indicator of how widely the community is reading.Yet there is little work looking at broad temporal patterns of citation.This work, systematically and empirically examines: How far back in time do we tend to go to cite papers?How has that changed over time, and what factors correlate with this citational attention/amnesia?We chose NLP as our domain of interest, and analyzed ∼71.5K papers to show and quantify several key trends in citation.Notably, ∼62% of cited papers are from the immediate five years prior to publication, whereas only ∼17% are more than ten years old.Furthermore, we show that the median age and age diversity of cited papers was steadily increasing from 1990 to 2014, but since then the trend has reversed, and current NLP papers have an all-time low temporal citation diversity.Finally, we show that unlike the 1990s, the highly cited papers in the last decade were also papers with the least citation diversity; likely contributing to the intense (and arguably harmful) recency focus.Code, data, and a demo are available at the project homepage.1 2
Janvijay Singh, Mukund Rungta, Diyi Yang, Saif M. Mohammad
ACL (1)4
2023 What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study
abstract
The degree of semantic relatedness of two units of language has long been considered fundamental to understanding meaning.Additionally, automatically determining relatedness has many applications such as question answering and summarization.However, prior NLP work has largely focused on semantic similarity, a subset of relatedness, because of a lack of relatedness datasets.In this paper, we introduce a dataset for Semantic Textual Relatedness, STR-2022, that has 5,500 English sentence pairs manually annotated using a comparative annotation framework, resulting in fine-grained scores.We show that human intuition regarding relatedness of sentence pairs is highly reliable, with a repeat annotation correlation of 0.84.We use the dataset to explore questions on what makes sentences semantically related.We also show the utility of STR-2022 for evaluating automatic methods of sentence representation and for various downstream NLP tasks.Our dataset, data statement, and annotation questionnaire can be found at:
Mohamed Abdalla 0001, Krishnapriya Vishnubhotla, Saif M. Mohammad
EACL3
2023 AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages
abstract
Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum, David Ifeoluwa Adelani, Seid Muhie Yimam, Ibrahim Said Ahmad, Meriem Beloucif, Saif M. Mohammad, Sebastian Ruder, Oumaima Hourrane, Alípio Mário Jorge, Pavel Brazdil, Felermino D. M. A. Ali, Davis David, Salomey Osei, Bello Shehu Bello, Falalu Ibrahim Lawan, Tajuddeen Rabiu Gwadabe, Samuel Rutunda, Tadesse Destaw Belay, Wendimu Baye Messelle, Hailu Beshada Balcha, Sisay Adugna Chala, Hagos Tesfahun Gebremichael, Bernard Opoku, Stephen Arthur
EMNLP9
2023 A Diachronic Analysis of Paradigm Shifts in NLP Research: When, How, and Why?
abstract
Understanding the fundamental concepts and trends in a scientific field is crucial for keeping abreast of its continuous advancement.In this study, we propose a systematic framework for analyzing the evolution of research topics in a scientific field using causal discovery and inference techniques.We define three variables to encompass diverse facets of the evolution of research topics within NLP and utilize a causal discovery algorithm to unveil the causal connections among these variables using observational data.Subsequently, we leverage this structure to measure the intensity of these relationships.By conducting extensive experiments on the ACL Anthology corpus, we demonstrate that our framework effectively uncovers evolutionary trends and the underlying causes for a wide range of NLP research topics.Specifically, we show that tasks and methods are primary drivers of research in NLP, with datasets following, while metrics have minimal impact. 1
Aniket Pramanick, Yufang Hou 0001, Saif M. Mohammad, Iryna Gurevych
EMNLP3
2023 Language and Mental Health: Measures of Emotion Dynamics from Text as Linguistic Biosocial Markers
abstract
Research in psychopathology has shown that, at an aggregate level, the patterns of emotional change over time-emotion dynamics-are indicators of one's mental health.One's patterns of emotion change have traditionally been determined through self-reports of emotions; however, there are known issues with accuracy, bias, and ease of data collection.Recent approaches to determining emotion dynamics from one's everyday utterances addresses many of these concerns, but it is not yet known whether these measures of utterance emotion dynamics (UED) correlate with mental health diagnoses.Here, for the first time, we study the relationship between tweet emotion dynamics and mental health disorders.We find that each of the UED metrics studied varied by the user's self-disclosed diagnosis.For example: average valence was significantly higher (i.e., more positive text) in the control group compared to users with ADHD, MDD, and PTSD.Valence variability was significantly lower in the control group compared to ADHD, depression, bipolar disorder, MDD, PTSD, and OCD but not PPD.Rise and recovery rates of valence also exhibited significant differences from the control.This work provides important early evidence for how linguistic cues pertaining to emotion dynamics can play a crucial role as biosocial markers for mental illnesses and aid in the understanding, diagnosis, and management of mental health disorders.
Daniela Teodorescu, Tiffany Cheng, Alona Fyshe, Saif M. Mohammad
EMNLP4
2023 We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields
abstract
Natural Language Processing (NLP) is poised to substantially influence the world.However, significant progress comes hand-in-hand with substantial risks.Addressing them requires broad engagement with various fields of study.Yet, little empirical work examines the state of such engagement (past or current).In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other).We analyzed ∼77k NLP papers, ∼3.1m citations from NLP papers to other papers, and ∼1.8m citations from other papers to NLP papers.We show that, unlike most fields, the cross-field engagement of NLP, measured by our proposed Citation Field Diversity Index (CFDI), has declined from 0.58 in 1980 to 0.31 in 2022 (an all-time low).In addition, we find that NLP has grown more insular-citing increasingly more NLP papers and having fewer papers that act as bridges between fields.NLP citations are dominated by computer science; Less than 8% of NLP citations are to linguistics, and less than 3% are to math and psychology.These findings underscore NLP's urgent need to reflect on its engagement with various fields.
Jan Philip Wahle, Terry Ruas, Mohamed Abdalla 0001, Bela Gipp, Saif M. Mohammad
EMNLP5
2022 Ethics Sheets for AI Tasks
abstract
Several high-profile events, such as the mass testing of emotion recognition systems on vulnerable sub-populations and using question answering systems to make moral judgments, have highlighted how technology will often lead to more adverse outcomes for those that are already marginalized.At issue here are not just individual systems and datasets, but also the AI tasks themselves.In this position paper, I make a case for thinking about ethical considerations not just at the level of individual models and datasets, but also at the level of AI tasks.I will present a new form of such an effort, Ethics Sheets for AI Tasks, dedicated to fleshing out the assumptions and ethical considerations hidden in how a task is commonly framed and in the choices we make regarding the data, method, and evaluation.I will also present a template for ethics sheets with 50 ethical considerations, using the task of emotion recognition as a running example.Ethics sheets are a mechanism to engage with and document ethical considerations before building datasets and systems.Similar to survey articles, a small number of carefully created ethics sheets can serve numerous researchers and developers.
Saif M. Mohammad
ACL (1)1
2022 Geographic Citation Gaps in NLP Research
abstract
In a fair world, people have equitable opportunities to education, to conduct scientific research, to publish, and to get credit for their work, regardless of where they live.However, it is common knowledge among researchers that a vast number of papers accepted at top NLP venues come from a handful of western countries and (lately) China; whereas, very few papers from Africa and South America get published.Similar disparities are also believed to exist for paper citation counts.In the spirit of "what we do not measure, we cannot improve", this work asks a series of questions on the relationship between geographical location and publication success (acceptance in top NLP venues and citation impact).We first created a dataset of 70,000 papers from the ACL Anthology, extracted their meta-information, and generated their citation network.We then show that not only are there substantial geographical disparities in paper acceptance and citation but also that these disparities persist even when controlling for a number of variables such as venue of publication and sub-field of NLP.Further, despite some steps taken by the NLP community to improve geographical diversity, we show that the disparity in publication metrics across locations is still on an increasing trend since the early 2000s.We release our code and dataset here
Mukund Rungta, Janvijay Singh, Saif M. Mohammad, Diyi Yang
EMNLP3
2022 TUSC: Emotion Word Usage in Tweets from US and Canada
Krishnapriya Vishnubhotla, Saif M. Mohammad
LREC2
2022 D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research
abstract
DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research activity, productivity, focus, bias, accessibility, and impact of computer science research. We present an initial analysis focused on the volume of computer science research (e.g., number of papers, authors, research activity), trends in topics of interest, and citation patterns. Our findings show that computer science is a growing research field (15% annually), with an active and collaborative researcher community. While papers in recent years present more bibliographical entries in comparison to previous decades, the average number of citations has been declining. Investigating papers’ abstracts reveals that recent topic trends are clearly reflected in D3. Finally, we list further applications of D3 and pose supplemental research questions. The D3 dataset, our findings, and source code are publicly available for research purposes.
Jan Philip Wahle, Terry Ruas, Saif M. Mohammad, Bela Gipp
LREC3
2022 Ethics Sheet for Automatic Emotion Recognition and Sentiment Analysis
abstract
Abstract The importance and pervasiveness of emotions in our lives makes affective computing a tremendously important and vibrant line of work. Systems for automatic emotion recognition (AER) and sentiment analysis can be facilitators of enormous progress (e.g., in improving public health and commerce) but also enablers of great harm (e.g., for suppressing dissidents and manipulating voters). Thus, it is imperative that the affective computing community actively engage with the ethical ramifications of their creations. In this article, I have synthesized and organized information from AI Ethics and Emotion Recognition literature to present fifty ethical considerations relevant to AER. Notably, this ethics sheet fleshes out assumptions hidden in how AER is commonly framed, and in the choices often made regarding the data, method, and evaluation. Special attention is paid to the implications of AER on privacy and social groups. Along the way, key recommendations are made for responsible AER. The objective of the ethics sheet is to facilitate and encourage more thoughtfulness on why to automate, how to automate, and how to judge success well before the building of AER systems. Additionally, the ethics sheet acts as a useful introductory document on emotion recognition (complementing survey articles).
Saif M. Mohammad
Comput. Linguistics1
2021 Ruddit: Norms of Offensiveness for English Reddit Comments
abstract
Rishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis, Saif M. Mohammad, Ekaterina Shutova. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Rishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis, Saif M. Mohammad, Ekaterina Shutova
ACL/IJCNLP (1)5
2020 Lexichrome: Text Construction and Lexical Discovery with Word-Color Associations Using Interactive Visualization
abstract
Based on word-color associations from a comprehensive, crowdsourced lexicon, we present Lexichrome: a web application that explores the popular perception of relationships between English words and eleven basic color terms using interactive visualization. Lexichrome provides three complementary visualizations: "Palette" presents the diversity of word-color associations across the color palette; "Words" reveals the color associations of individual words using a dictionary-like interface; "Roget's Thesaurus" uncovers color association patterns in different semantic categories found in the thesaurus. Finally, our text editor allows users to compose their own texts and examine the resultant chromatic fingerprints throughout the process. We studied the utility of Lexichrome in a two-part qualitative user study with nine participants from various writing-intensive professions. We find that the presence of word-color associations promotes awareness surrounding word choice, editorial decision, and audience reception, and introduce a variety of use cases, features, and opportunities applicable to creative writing, corporate communication, and journalism.
Chris Kim, Uta Hinrichs, Saif M. Mohammad, Christopher Collins 0001
Conference on Designing Interactive Systems3
2020 Examining Citations of Natural Language Processing Literature
abstract
We extracted information from the ACL Anthology (AA) and Google Scholar (GS) to examine trends in citations of NLP papers.We explore questions such as: how well cited are papers of different types (journal articles, conference papers, demo papers, etc.)? how well cited are papers from different areas of within NLP? etc. Notably, we show that only about 56% of the papers in AA are cited ten or more times.CL Journal has the most cited papers, but its citation dominance has lessened in recent years.On average, long papers get almost three times as many citations as short papers; and papers on sentiment classification, anaphora resolution, and entity recognition have the highest median citations.The analyses presented here, and the associated dataset of NLP papers mapped to citations, have a number of uses including: understanding how the field is growing and quantifying the impact of different types of papers.
Saif M. Mohammad
ACL1
2020 Gender Gap in Natural Language Processing Research: Disparities in Authorship and Citations
abstract
Disparities in authorship and citations across gender can have substantial adverse consequences not just on the disadvantaged genders, but also on the field of study as a whole.Measuring gender gaps is a crucial step towards addressing them.In this work, we examine female first author percentages and the citations to their papers in Natural Language Processing (1965 to 2019).We determine aggregatelevel statistics using an existing manually curated author-gender list as well as first names strongly associated with a gender.We find that only about 29% of first authors are female and only about 25% of last authors are female.Notably, this percentage has not improved since the mid 2000s.We also show that, on average, female first authors are cited less than male first authors, even when controlling for experience and area of research.Finally, we discuss the ethical considerations involved in automatic demographic analysis.
Saif M. Mohammad
ACL1
2020 PoKi: A Large Dataset of Poems by Children
abstract
Child language studies are crucial in improving our understanding of child well-being; especially in determining the factors that impact happiness, the sources of anxiety, techniques of emotion regulation, and the mechanisms to cope with stress. However, much of this research is stymied by the lack of availability of large child-written texts. We present a new corpus of child-written text, PoKi, which includes about 62 thousand poems written by children from grades 1 to 12. PoKi is especially useful in studying child language because it comes with information about the age of the child authors (their grade). We analyze the words in PoKi along several emotion dimensions (valence, arousal, dominance) and discrete emotions (anger, fear, sadness, joy). We use non-parametric regressions to model developmental differences from early childhood to late-adolescence. Results show decreases in valence that are especially pronounced during mid-adolescence, while arousal and dominance peaked during adolescence. Gender differences in the developmental trajectory of emotions are also observed. Our results support and extend the current state of emotion development research.
Will E. Hipson, Saif M. Mohammad
LREC2
2020 SOLO: A Corpus of Tweets for Examining the State of Being Alone
abstract
The state of being alone can have a substantial impact on our lives, though experiences with time alone diverge significantly among individuals. Psychologists distinguish between the concept of solitude, a positive state of voluntary aloneness, and the concept of loneliness, a negative state of dissatisfaction with the quality of one’s social interactions. Here, for the first time, we conduct a large-scale computational analysis to explore how the terms associated with the state of being alone are used in online language. We present SOLO (State of Being Alone), a corpus of over 4 million tweets collected with query terms solitude, lonely, and loneliness. We use SOLO to analyze the language and emotions associated with the state of being alone. We show that the term solitude tends to co-occur with more positive, high-dominance words (e.g., enjoy, bliss) while the terms lonely and loneliness frequently co-occur with negative, low-dominance words (e.g., scared, depressed), which confirms the conceptual distinctions made in psychology. We also show that women are more likely to report on negative feelings of being lonely as compared to men, and there are more teenagers among the tweeters that use the word lonely than among the tweeters that use the word solitude.
Svetlana Kiritchenko, Will E. Hipson, Robert J. Coplan, Saif M. Mohammad
LREC4
2020 NLP Scholar: A Dataset for Examining the State of NLP Research
abstract
Google Scholar is the largest web search engine for academic literature that also provides access to rich metadata associated with the papers. The ACL Anthology (AA) is the largest repository of articles on Natural Language Processing (NLP). We extracted information from AA for about 44 thousand NLP papers and identified authors who published at least three papers there. We then extracted citation information from Google Scholar for all their papers (not just their AA papers). This resulted in a dataset of 1.1 million papers and associated Google Scholar information. We aligned the information in the AA and Google Scholar datasets to create the NLP Scholar Dataset – a single unified source of information (from both AA and Google Scholar) for tens of thousands of NLP papers. It can be used to identify broad trends in productivity, focus, and impact of NLP research. We present here initial work on analyzing the volume of research in NLP over the years and identifying the most cited papers in NLP. We also list a number of additional potential applications.
Saif M. Mohammad
LREC1
2020 WordWars: A Dataset to Examine the Natural Selection of Words
abstract
There is a growing body of work on how word meaning changes over time: mutation. In contrast, there is very little work on how different words compete to represent the same meaning, and how the degree of success of words in that competition changes over time: natural selection. We present a new dataset, WordWars, with historical frequency data from the early 1800s to the early 2000s for monosemous English words in over 5000 synsets. We explore three broad questions with the dataset: (1) what is the degree to which predominant words in these synsets have changed, (2) how do prominent word features such as frequency, length, and concreteness impact natural selection, and (3) what are the differences between the predominant words of the 2000s and the predominant words of early 1800s. We show that close to one third of the synsets undergo a change in the predominant word in this time period. Manual annotation of these pairs shows that about 15% of these are orthographic variations, 25% involve affix changes, and 60% have completely different roots. We find that frequency, length, and concreteness all impact natural selection, albeit in different ways.
Saif M. Mohammad
LREC1
2019 AffectiveTweets: a Weka Package for Analyzing Affect in Tweets
abstract
AffectiveTweets is a set of programs for analyzing emotion and sentiment of social media messages such as tweets. It is implemented as a package for the Weka machine learning workbench and provides methods for calculating state-of-the-art affect analysis features from tweets that can be fed into machine learning algorithms implemented in Weka. It also implements methods for building affective lexicons and distant supervision methods for training affective models from unlabeled tweets. The package was used by several teams in the shared tasks: EmoInt 2017 and Affect in Tweets SemEval 2018 Task 1.
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer, Saif M. Mohammad
J. Mach. Learn. Res.4
2018 Obtaining Reliable Human Ratings of Valence, Arousal, and Dominance for 20, 000 English Words
abstract
Words play a central role in language and thought.Factor analysis studies have shown that the primary dimensions of meaning are valence, arousal, and dominance (VAD).We present the NRC VAD Lexicon, which has human ratings of valence, arousal, and dominance for more than 20,000 English words.We use Best-Worst Scaling to obtain fine-grained scores and address issues of annotation consistency that plague traditional rating scale methods of annotation.We show that the ratings obtained are vastly more reliable than those in existing lexicons.We also show that there exist statistically significant differences in the shared understanding of valence, arousal, and dominance across demographic variables such as age, gender, and personality.
Saif M. Mohammad
ACL (1)1
2018 Word Affect Intensities
Saif M. Mohammad
LREC1
2018 Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories
Saif M. Mohammad, Svetlana Kiritchenko
LREC1
2018 WikiArt Emotions: An Annotated Dataset of Emotions Evoked by Art
Saif M. Mohammad, Svetlana Kiritchenko
LREC1
2018 Quantifying Qualitative Data for Understanding Controversial Issues
Michael Wojatzki, Saif M. Mohammad, Torsten Zesch, Svetlana Kiritchenko
LREC2
2018 Data and systems for medication-related text classification and concept normalization from Twitter: insights from the Social Media Mining for Health (SMM4H)-2017 shared task
abstract
Objective: We executed the Social Media Mining for Health (SMM4H) 2017 shared tasks to enable the community-driven development and large-scale evaluation of automatic text processing methods for the classification and normalization of health-related text from social media. An additional objective was to publicly release manually annotated data. Materials and Methods: We organized 3 independent subtasks: automatic classification of self-reports of 1) adverse drug reactions (ADRs) and 2) medication consumption, from medication-mentioning tweets, and 3) normalization of ADR expressions. Training data consisted of 15 717 annotated tweets for (1), 10 260 for (2), and 6650 ADR phrases and identifiers for (3); and exhibited typical properties of social-media-based health-related texts. Systems were evaluated using 9961, 7513, and 2500 instances for the 3 subtasks, respectively. We evaluated performances of classes of methods and ensembles of system combinations following the shared tasks. Results: Among 55 system runs, the best system scores for the 3 subtasks were 0.435 (ADR class F1-score) for subtask-1, 0.693 (micro-averaged F1-score over two classes) for subtask-2, and 88.5% (accuracy) for subtask-3. Ensembles of system combinations obtained best scores of 0.476, 0.702, and 88.7%, outperforming individual systems. Discussion: Among individual systems, support vector machines and convolutional neural networks showed high performance. Performance gains achieved by ensembles of system combinations suggest that such strategies may be suitable for operational systems relying on difficult text classification tasks (eg, subtask-1). Conclusions: Data imbalance and lack of context remain challenges for natural language processing of social media text. Annotated data from the shared task have been made available as reference standards for future studies (http://dx.doi.org/10.17632/rxwfb3tysd.1).
Abeed Sarker, Maksim Belousov, Jasper Friedrichs, Kai Hakala, Svetlana Kiritchenko, Farrokh Mehryary, Sifei Han, Tung Tran 0001, Anthony Rios, Ramakanth Kavuluru, Berry de Bruijn, Filip Ginter, Debanjan Mahata, Saif M. Mohammad, Goran Nenadic, Graciela Gonzalez-Hernandez
J. Am. Medical Informatics Assoc.14
2017 Stance and Sentiment in Tweets
abstract
We can often detect from a person’s utterances whether he or she is in favor of or against a given target entity—one’s stance toward the target. However, a person may express the same stance toward a target by using negative or positive language. Here for the first time we present a dataset of tweet–target pairs annotated for both stance and sentiment. The targets may or may not be referred to in the tweets, and they may or may not be the target of opinion in the tweets. Partitions of this dataset were used as training and test sets in a SemEval-2016 shared task competition. We propose a simple stance detection system that outperforms submissions from all 19 teams that participated in the shared task. Additionally, access to both stance and sentiment annotations allows us to explore several research questions. We show that although knowing the sentiment expressed by a tweet is beneficial for stance classification, it alone is not sufficient. Finally, we use additional unlabeled data through distant supervision techniques and word embeddings to further improve stance classification.
Saif M. Mohammad, Parinaz Sobhani, Svetlana Kiritchenko
ACM Trans. Internet Techn.1
2016 Happy Accident: A Sentiment Composition Lexicon for Opposing Polarity Phrases
Svetlana Kiritchenko, Saif M. Mohammad
LREC2
2016 A Dataset for Detecting Stance in Tweets
Saif M. Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu 0001, Colin Cherry
LREC1
2016 Sentiment Lexicons for Arabic Social Media
Saif M. Mohammad, Mohammad Salameh, Svetlana Kiritchenko
LREC1
2016 Capturing Reliable Fine-Grained Sentiment Associations by Crowdsourcing and Best-Worst Scaling
abstract
Access to word-sentiment associations is useful for many applications, including sentiment analysis, stance detection, and linguistic analysis. However, manually assigning fine-grained sentiment association scores to words has many challenges with respect to keeping annotations consistent. We apply the annotation technique of Best-Worst Scaling to obtain real-valued sentiment association scores for words and phrases in three different domains: general English, English Twitter, and Arabic Twitter. We show that on all three domains the ranking of words by sentiment remains remarkably consistent even when the annotation process is repeated with a different set of annotators. We also, for the first time, determine the minimum difference in sentiment association that is perceptible to native speakers of a language.
Svetlana Kiritchenko, Saif M. Mohammad
HLT-NAACL2
2016 Sentiment Composition of Words with Opposing Polarities
abstract
In this paper, we explore sentiment composition in phrases that have at least one positive and at least one negative word-phrases like happy accident and best winter break.We compiled a dataset of such opposing polarity phrases and manually annotated them with real-valued scores of sentiment association.Using this dataset, we analyze the linguistic patterns present in opposing polarity phrases.Finally, we apply several unsupervised and supervised techniques of sentiment composition to determine their efficacy on this dataset.Our best system, which incorporates information from the phrase's constituents, their parts of speech, their sentiment association scores, and their embedding vectors, obtains an accuracy of over 80% on the opposing polarity phrases.
Svetlana Kiritchenko, Saif M. Mohammad
HLT-NAACL2
2016 Determining Word-Emotion Associations from Tweets by Multi-label Classification
abstract
The automatic detection of emotions in Twitter posts is a challenging task due to the informal nature of the language used in this platform. In this paper, we propose a methodology for expanding the NRC word-emotion association lexicon for the language used in Twitter. We perform this expansion using multi-label classification of words and compare different word-level features extracted from unlabelled tweets such as unigrams, Brown clusters, POS tags, and word2vec embeddings. The results show that the expanded lexicon achieves major improvements over the original lexicon when classifying tweets into emotional categories. In contrast to previous work, our methodology does not depend on tweets annotated with emotional hashtags, thus enabling the identification of emotional words from any domain-specific collection using unlabelled tweets.
Felipe Bravo-Marquez, Eibe Frank, Saif M. Mohammad, Bernhard Pfahringer
WI3
2016 How Translation Alters Sentiment
abstract
Sentiment analysis research has predominantly been on English texts. Thus there exist many sentiment resources for English, but less so for other languages. Approaches to improve sentiment analysis in a resource-poor focus language include: (a) translate the focus language text into a resource-rich language such as English, and apply a powerful English sentiment analysis system on the text, and (b) translate resources such as sentiment labeled corpora and sentiment lexicons from English into the focus language, and use them as additional resources in the focus-language sentiment analysis system. In this paper we systematically examine both options. We use Arabic social media posts as stand-in for the focus language text. We show that sentiment analysis of English translations of Arabic texts produces competitive results, w.r.t. Arabic sentiment analysis. We show that Arabic sentiment analysis systems benefit from the use of automatically translated English sentiment lexicons. We also conduct manual annotation studies to examine why the sentiment of a translation is different from the sentiment of the source word or text. This is especially relevant for building better automatic translation systems. In the process, we create a state-of-the-art Arabic sentiment analysis system, a new dialectal Arabic sentiment lexicon, and the first Arabic-English parallel corpus that is independently annotated for sentiment by Arabic and English speakers.
Saif M. Mohammad, Mohammad Salameh, Svetlana Kiritchenko
J. Artif. Intell. Res.1
2015 Sentiment after Translation: A Case-Study on Arabic Social Media Posts
abstract
Mohammad Salameh, Saif Mohammad, Svetlana Kiritchenko. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Mohammad Salameh, Saif M. Mohammad, Svetlana Kiritchenko
HLT-NAACL2
2015 Using Hashtags to Capture Fine Emotion Categories from Tweets
abstract
Detecting emotions in microblogs and social media posts has applications for industry, health, and security. Statistical, supervised automatic methods for emotion detection rely on text that is labeled for emotions, but such data are rare and available for only a handful of basic emotions. In this article, we show that emotion‐word hashtags are good manual labels of emotions in tweets. We also propose a method to generate a large lexicon of word–emotion associations from this emotion‐labeled tweet corpus. This is the first lexicon with real‐valued word–emotion association scores. We begin with experiments for six basic emotions and show that the hashtag annotations are consistent and match with the annotations of trained judges. We also show how the extracted tweet corpus and word–emotion associations can be used to improve emotion classification accuracy in a different nontweet domain. Eminent psychologist Robert Plutchik had proposed that emotions have a relationship with personality traits. However, empirical experiments to establish this relationship have been stymied by the lack of comprehensive emotion resources. Because personality may be associated with any of the hundreds of emotions and because our hashtag approach scales easily to a large number of emotions, we extend our corpus by collecting tweets with hashtags pertaining to 585 fine emotions. Then, for the first time, we present experiments to show that fine emotion categories such as those of excitement, guilt, yearning, and admiration are useful in automatically detecting personality from text. Stream‐of‐consciousness essays and collections of Facebook posts marked with personality traits of the author are used as test sets.
Saif M. Mohammad, Svetlana Kiritchenko
Comput. Intell.1
2015 Sentiment, emotion, purpose, and style in electoral tweets
Saif M. Mohammad, Xiaodan Zhu 0001, Svetlana Kiritchenko, Joel D. Martin
Inf. Process. Manag.1
2015 Experiments with three approaches to recognizing lexical entailment
abstract
Abstract Inference in natural language often involves recognizing lexical entailment (RLE), that is, identifying whether one word entails another. For example,buyentailsown. Two general strategies for RLE have been proposed: One strategy is to manually construct an asymmetric similarity measure for context vectors (directional similarity) and another is to treat RLE as a problem of learning to recognize semantic relations using supervised machine-learning techniques (relation classification). In this paper, we experiment with two recent state-of-the-art representatives of the two general strategies. The first approach is an asymmetric similarity measure (an instance of thedirectional similaritystrategy), designed to capture the degree to which the contexts of a word,a, form a subset of the contexts of another word,b. The second approach (an instance of therelation classificationstrategy) represents a word pair,a:b, with a feature vector that is the concatenation of the context vectors ofaandb, and then applies supervised learning to a training set of labeled feature vectors. In addition, we introduce a third approach that is a new instance of therelation classificationstrategy. The third approach represents a word pair,a:b, with a feature vector in which the features are the differences in the similarities ofaandbto a set of reference words. All three approaches use vector space models of semantics, based on word–context matrices. We perform an extensive evaluation of the three approaches using three different datasets. The proposed new approach (similarity differences) performs significantly better than the other two approaches on some datasets and there is no dataset for which it is significantly worse. Along the way, we address some of the concerns raised in past research, regarding the treatment of RLE as a problem of semantic relation classification, and we suggest, it is beneficial to make connections between the research in lexical entailment and the research in semantic relation classification.
Peter D. Turney, Saif M. Mohammad
Nat. Lang. Eng.2
2014 An Empirical Study on the Effect of Negation Words on Sentiment
abstract
Negation words, such as no and not, play a fundamental role in modifying sentiment of textual expressions. We will refer to a negation word as the negator and the text span within the scope of the negator as the argument. Commonly used heuristics to estimate the sentiment of negated expressions rely simply on the sentiment of argument (and not on the negator or the argument itself). We use a sentiment treebank to show that these existing heuristics are poor estimators of sentiment. We then modify these heuristics to be dependent on the negators and show that this improves prediction. Next, we evaluate a recently proposed composition model (Socher et al., 2013) that relies on both the negator and the argument. This model learns the syntax and semantics of the negator's argument with a recursive neural network. We show that this approach performs better than those mentioned above. In addition, we explicitly incorporate the prior sentiment of the argument and observe that this information can help reduce fitting errors.
Xiaodan Zhu 0001, Saif M. Mohammad, Svetlana Kiritchenko
ACL (1)3
2014 Sentiment Analysis of Short Informal Texts
abstract
We describe a state-of-the-art sentiment analysis system that detects (a) the sentiment of short informal textual messages such as tweets and SMS (message-level task) and (b) the sentiment of a word or a phrase within a message (term-level task). The system is based on a supervised statistical text classification approach leveraging a variety of surface-form, semantic, and sentiment features. The sentiment features are primarily derived from novel high-coverage tweet-specific sentiment lexicons. These lexicons are automatically generated from tweets with sentiment-word hashtags and from tweets with emoticons. To adequately capture the sentiment of words in negated contexts, a separate sentiment lexicon is generated for negated words. The system ranked first in the SemEval-2013 shared task `Sentiment Analysis in Twitter' (Task 2), obtaining an F-score of 69.02 in the message-level task and 88.93 in the term-level task. Post-competition improvements boost the performance to an F-score of 70.45 (message-level task) and 89.50 (term-level task). The system also obtains state-of-the-art performance on two additional datasets: the SemEval-2013 SMS test set and a corpus of movie review excerpts. The ablation experiments demonstrate that the use of the automatically generated lexicons results in performance gains of up to 6.5 absolute percentage points.
Svetlana Kiritchenko, Xiaodan Zhu 0001, Saif M. Mohammad
J. Artif. Intell. Res.3
2013 Crowdsourcing a Word-Emotion Association Lexicon
abstract
Even though considerable attention has been given to the polarity of words (positive and negative) and the creation of large polarity lexicons, research in emotion analysis has had to rely on limited and small emotion lexicons. In this paper, we show how the combined strength and wisdom of the crowds can be used to generate a large, high‐quality, word–emotion and word–polarity association lexicon quickly and inexpensively. We enumerate the challenges in emotion annotation in a crowdsourcing scenario and propose solutions to address them. Most notably, in addition to questions about emotions associated with terms, we show how the inclusion of a word choice question can discourage malicious data entry, help to identify instances where the annotator may not be familiar with the target term (allowing us to reject such annotations), and help to obtain annotations at sense level (rather than at word level). We conducted experiments on how to formulate the emotion‐annotation questions, and show that asking if a term is associated with an emotion leads to markedly higher interannotator agreement than that obtained by asking if a term evokes an emotion.
Saif M. Mohammad, Peter D. Turney
Comput. Intell.1
2013 Computing Lexical Contrast
abstract
Knowing the degree of semantic contrast between words has widespread application in natural language processing, including machine translation, information retrieval, and dialogue systems. Manually created lexicons focus on opposites, such as hot and cold. Opposites are of many kinds such as antipodals, complementaries, and gradable. Existing lexicons often do not classify opposites into the different kinds, however. They also do not explicitly list word pairs that are not opposites but yet have some degree of contrast in meaning, such as warm and cold or tropical and freezing. We propose an automatic method to identify contrasting word pairs that is based on the hypothesis that if a pair of words, A and B, are contrasting, then there is a pair of opposites, C and D, such that A and C are strongly related and B and D are strongly related. (For example, there exists the pair of opposites hot and cold such that tropical is related to hot, and freezing is related to cold.) We will call this the contrast hypothesis. We begin with a large crowdsourcing experiment to determine the amount of human agreement on the concept of oppositeness and its different kinds. In the process, we flesh out key features of different kinds of opposites. We then present an automatic and empirical measure of lexical contrast that relies on the contrast hypothesis, corpus statistics, and the structure of a Roget-like thesaurus. We show how, using four different data sets, we evaluated our approach on two different tasks, solving “most contrasting word” questions and distinguishing synonyms from opposites. The results are analyzed across four parts of speech and across five different kinds of opposites. We show that the proposed measure of lexical contrast obtains high precision and large coverage, outperforming existing methods.
Saif M. Mohammad, Bonnie J. Dorr, Graeme Hirst, Peter D. Turney
Comput. Linguistics1
2013 Generating Extractive Summaries of Scientific Paradigms
abstract
Researchers and scientists increasingly find themselves in the position of having to quickly understand large amounts of technical material. Our goal is to effectively serve this need by using bibliometric text mining and summarization techniques to generate summaries of scientific literature. We show how we can use citations to produce automatically generated, readily consumable, technical extractive summaries. We first propose C-LexRank, a model for summarizing single scientific articles based on citations, which employs community detection and extracts salient information-rich sentences. Next, we further extend our experiments to summarize a set of papers, which cover the same scientific topic. We generate extractive summaries of a set of Question Answering (QA) and Dependency Parsing (DP) papers, their abstracts, and their citation sentences and show that citations have unique information amenable to creating a summary.
Vahed Qazvinian, Dragomir R. Radev, Saif M. Mohammad, Bonnie J. Dorr, David M. Zajic, Michael Whidby, Taesun Moon
J. Artif. Intell. Res.3
2012 Portable Features for Classifying Emotional Text
Saif M. Mohammad
HLT-NAACL1
2012 From once upon a time to happily ever after: Tracking emotions in mail and books
Saif M. Mohammad
Decis. Support Syst.1
2009 Estimating Semantic Distance Using Soft Semantic Constraints in Knowledge-Source - Corpus Hybrid Models
Yuval Marton, Saif M. Mohammad, Philip Resnik
EMNLP2
2009 Generating High-Coverage Semantic Orientation Lexicons From Overtly Marked Words and a Thesaurus
Saif M. Mohammad, Cody Dunne, Bonnie J. Dorr
EMNLP1
2009 Using Citations to Generate surveys of Scientific Paradigms
Saif M. Mohammad, Bonnie J. Dorr, Melissa Egan, Ahmed Awadallah 0001, Pradeep Muthukrishnan, Vahed Qazvinian, Dragomir R. Radev, David M. Zajic
HLT-NAACL1
2008 Computing Word-Pair Antonymy
Saif M. Mohammad, Bonnie J. Dorr, Graeme Hirst
EMNLP1
2007 Cross-Lingual Distributional Profiles of Concepts for Measuring Semantic Distance
Saif M. Mohammad, Iryna Gurevych, Graeme Hirst, Torsten Zesch
EMNLP-CoNLL1
2006 Determining Word Sense Dominance Using a Thesaurus
Saif M. Mohammad, Graeme Hirst
EACL1
2006 Distributional measures of concept-distance: A task-oriented evaluation
Saif M. Mohammad, Graeme Hirst
EMNLP1
2004 Combining Lexical and Syntactic Features for Supervised Word Sense Disambiguation
Saif M. Mohammad, Ted Pedersen
CoNLL1
2003 Guaranteed Pre-tagging for the Brill Tagger
Saif M. Mohammad, Ted Pedersen
CICLing1