Mohammad Taher Pilehvar

dblp:136/8704 · DBLP profile ↗
← Back
49ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0003-3694-4006ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 10 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length
abstract
Users often rely on Large Language Models (LLMs) for processing multiple documents or performing analysis over a number of instances.For example, analysing the overall sentiment of a number of movie reviews requires an LLM to process the sentiment of each review individually in order to provide a final aggregated answer.While LLM performance on such individual tasks is generally high, there has been little research on how LLMs perform when dealing with multi-instance inputs.In this paper, we perform a comprehensive evaluation of the multi-instance processing (MIP) ability of LLMs for tasks in which they excel individually.The results show that all LLMs follow a pattern of slight performance degradation for small numbers of instances (≈20-100), followed by a performance collapse on larger instance counts.Crucially, our analysis shows that while context length is associated with this degradation, the number of instances has a stronger effect on the final results.This finding suggests that when optimising LLM performance for MIP, attention should be paid to both context length and, in particular, instance count. 1
Jingxuan Chen, Mohammad Taher Pilehvar, José Camacho-Collados
ACL (1)2
2026 Synthia: Scalable Grounded Persona Generation from Social Media Data
abstract
Vahid Rahimzadeh, Erfan Moosavi Monazzah, Mohammad Taher Pilehvar, Yadollah Yaghoobzadeh. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Vahid Rahimzadeh, Erfan Moosavi Monazzah, Mohammad Taher Pilehvar, Yadollah Yaghoobzadeh
ACL (1)3
2025 FarExStance: Explainable Stance Detection for Farsi
abstract
We introduce FarExStance, a new dataset for explainable stance detection in Farsi. Each instance in this dataset contains a claim, the stance of an article or social media post towards that claim, and an extractive explanation which provides evidence for the stance label. We compare the performance of a fine-tuned multilingual RoBERTa model to several large language models in zero-shot, few-shot, and parameter-efficient fine-tuned settings on our new dataset. On stance detection, the most accurate models are the fine-tuned RoBERTa model, the LLM Aya-23-8B which has been fine-tuned using parameter-efficient fine-tuning, and few-shot Claude-3.5-Sonnet. Regarding the quality of the explanations, our automatic evaluation metrics indicate that few-shot GPT-4o generates the most coherent explanations, while our human evaluation reveals that the best Overall Explanation Score (OES) belongs to few-shot Claude-3.5-Sonnet. The fine-tuned Aya-32-8B model produced explanations most closely aligned with the reference explanations.
Majid Zarharan, Maryam Hashemi, Malika Behroozrazegh, Sauleh Eetemadi, Mohammad Taher Pilehvar, Jennifer Foster
COLING5
2025 LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
abstract
Why do gradient-based explanations struggle with Transformers, and how can we improve them? We identify gradientflow imbalances in Transformers that violate FullGradcompleteness, a critical property for attribution faithfulness that CNNs naturally possess. To address this issue, we introduce LibraGrad—a theoretically grounded post-hoc approach that corrects gradient imbalances through pruning and scaling of backward paths, without changing the forward pass or adding computational overhead. We evaluate LibraGrad using three metric families: Faithfulness, which quantifies prediction changes under perturbations of the most and least relevant features; Completeness Error, which measures attribution conservation relative to model outputs; and Segmentation AP, which assesses alignment with human perception. Extensive experiments across 8 architectures, 4 model sizes, and 5 datasets show that LibraGrad universally enhances gradient-based methods, outperforming existing white-box methods—including Transformer-specific approaches—across all metrics. We demonstrate superior qualitative results through two complementary evaluations: precise text-prompted region highlighting on CLIP models and accurate class discrimination between co-occurring animals on ImageNet-finetuned models—two settings on which existing methods often struggle. Libra-Grad is effective even on the attention-free MLP-Mixer architecture, indicating potential for extension to other modern architectures. Our code is freely available at https://nightmachinery.github.io/LibraGrad/.
Faridoun Mehri, Mahdieh Soleymani Baghshah, Mohammad Taher Pilehvar
CVPR3
2025 NormXLogit: The Head-on-Top Never Lies
abstract
With new large language models (LLMs) emerging frequently, it is important to consider the potential value of model-agnostic approaches that can provide interpretability across a variety of architectures.While recent advances in LLM interpretability show promise, many rely on complex, model-specific methods with high computational costs.To address these limitations, we propose Nor-mXLogit, a novel technique for assessing the significance of individual input tokens.This method operates based on the input and output representations associated with each token.First, we demonstrate that the norm of word embeddings can be utilized as a measure of token importance.Second, we reveal a significant relationship between a token's importance and how predictive its representation is of the model's final output.Extensive analyses indicate that our approach outperforms existing gradient-based methods in terms of faithfulness and offers competitive performance compared to leading architecture-specific techniques.
Sina Abbasi, Mohammad Reza Modarres, Mohammad Taher Pilehvar
EMNLP3
2025 Morables: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
abstract
As LLMs excel on standard reading comprehension benchmarks, attention is shifting toward evaluating their capacity for complex abstract reasoning and inference.Literaturebased benchmarks, with their rich narrative and moral depth, provide a compelling framework for evaluating such deeper comprehension skills.Here, we present MORABLES, a humanverified benchmark built from fables and short stories drawn from historical literature.The main task is structured as multiple-choice questions targeting moral inference, with carefully crafted distractors that challenge models to go beyond shallow, extractive question answering.To further stress-test model robustness, we introduce adversarial variants designed to surface LLM vulnerabilities and shortcuts due to issues such as data contamination.Our findings show that, while larger models outperform smaller ones, they remain susceptible to adversarial manipulation and often rely on superficial patterns rather than true moral reasoning.This brittleness results in significant self-contradiction, with the best models refuting their own answers in roughly 20% of cases depending on the framing of the moral choice.Interestingly, reasoning-enhanced models fail to bridge this gap, suggesting that scale -not reasoning ability -is the primary driver of performance.
Matteo Marcuzzo, Alessandro Zangari, Andrea Albarelli, José Camacho-Collados, Mohammad Taher Pilehvar
EMNLP5
2025 Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
abstract
Accurately measuring gender stereotypical bias in language models is a complex task with many hidden aspects.Current benchmarks have underestimated this multifaceted challenge and failed to capture the full extent of the problem.This paper examines the inconsistencies between intrinsic stereotype benchmarks.We propose that currently available benchmarks each capture only partial facets of gender stereotypes, and when considered in isolation, they provide just a fragmented view of the broader landscape of bias in language models.Using StereoSet and CrowS-Pairs as case studies, we investigated how data distribution affects benchmark results.By applying a framework from social psychology to balance the data of these benchmarks across various components of gender stereotypes, we demonstrated that even simple balancing techniques can significantly improve the correlation between different measurement approaches.Our findings underscore the complexity of gender stereotyping in language models and point to new directions for developing more refined techniques to detect and reduce bias.
Mahdi Zakizadeh, Mohammad Taher Pilehvar
EMNLP2
2025 Pun Unintended: LLMs and the Illusion of Humor Understanding
abstract
Puns are a form of humorous wordplay that exploits polysemy and phonetic similarity.While LLMs have shown promise in detecting puns, we show in this paper that their understanding often remains shallow, lacking the nuanced grasp typical of human interpretation.By systematically analyzing and reformulating existing pun benchmarks, we demonstrate how subtle changes in puns are sufficient to mislead LLMs.Our contributions include comprehensive and nuanced pun detection benchmarks, human evaluation across recent LLMs, and an analysis of the robustness challenges these models face in processing puns.
Alessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar, José Camacho-Collados
EMNLP4
2025 PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
abstract
Erfan Moosavi Monazzah, Vahid Rahimzadeh, Yadollah Yaghoobzadeh, Azadeh Shakery, Mohammad Taher Pilehvar. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Erfan Moosavi Monazzah, Vahid Rahimzadeh, Yadollah Yaghoobzadeh, Azadeh Shakery, Mohammad Taher Pilehvar
NAACL (Long Papers)5
2024 Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator Rationales
abstract
With the alarming rise of hate speech in online communities, the demand for effective NLP models to identify instances of offensive language has reached a critical point. However, the development of such models heavily relies on the availability of annotated datasets, which are scarce, particularly for less-studied languages. To bridge this gap for the Persian language, we present a novel dataset specifically tailored to multi-label hate speech detection. Our dataset, called Phate, consists of an extensive collection of over seven thousand manually-annotated Persian tweets, offering a rich resource for training and evaluating hate speech detection models on this language. Notably, each annotation in our dataset specifies the targeted group of hate speech and includes a span of the tweet which elucidates the rationale behind the assigned label. The incorporation of these information expands the potential applications of our dataset, facilitating the detection of targeted online harm or allowing the benchmark to serve research on interpretability of hate speech detection models. The dataset, annotation guideline, and all associated codes are accessible at https://github.com/Zahra-D/Phate.
Zahra Delbari, Nafise Sadat Moosavi, Mohammad Taher Pilehvar
AAAI3
2024 TweetTER: A Benchmark for Target Entity Retrieval on Twitter without Knowledge Bases
abstract
Entity linking is a well-established task in NLP consisting of associating entity mentions with entries in a knowledge base. Current models have demonstrated competitive performance in standard text settings. However, when it comes to noisy domains such as social media, certain challenges still persist. Typically, to evaluate entity linking on existing benchmarks, a comprehensive knowledge base is necessary and models are expected to possess an understanding of all the entities contained within the knowledge base. However, in practical scenarios where the objective is to retrieve sentences specifically related to a particular entity, strict adherence to a complete understanding of all entities in the knowledge base may not be necessary. To address this gap, we introduce TweetTER (Tweet Target Entity Retrieval), a novel benchmark that aims to bridge the challenges in entity linking. The distinguishing feature of this benchmark is its approach of re-framing entity linking as a binary entity retrieval task. This enables the evaluation of language models’ performance without relying on a conventional knowledge base, providing a more practical and versatile evaluation framework for assessing the effectiveness of language models in entity retrieval tasks.
Kiamehr Rezaee, José Camacho-Collados, Mohammad Taher Pilehvar
LREC/COLING3
2024 RepMatch: Quantifying Cross-Instance Similarities in Representation Space
abstract
Advances in dataset analysis techniques have enabled more sophisticated approaches to analyzing and characterizing training data instances, often categorizing data based on attributes such as "difficulty".In this work, we introduce RepMatch, a novel method that characterizes data through the lens of similarity.RepMatch quantifies the similarity between subsets of training instances by comparing the knowledge encoded in models trained on them, overcoming the limitations of existing analysis methods that focus solely on individual instances and are restricted to within-dataset analysis.Our framework allows for a broader evaluation, enabling similarity comparisons across arbitrary subsets of instances, supporting both dataset-to-dataset and instance-to-dataset analyses.We validate the effectiveness of Rep-Match across multiple NLP tasks, datasets, and models.Through extensive experimentation, we demonstrate that RepMatch can effectively compare datasets, identify more representative subsets of a dataset (that lead to better performance than randomly selected subsets of equivalent size), and uncover heuristics underlying the construction of some challenge datasets.
Mohammad Modarres, Sina Abbasi, Mohammad Taher Pilehvar
EMNLP3
2024 BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages
abstract
Large language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect the daily habits, customs, and lifestyles of different regions. That is, information about the food people eat for their birthday celebrations, spices they typically use, musical instruments youngsters play or the sports they practice in school is not always explicitly written online. To address this issue, we introduce BLEnD, a hand-crafted benchmark designed to evaluate LLMs' everyday knowledge across diverse cultures and languages. The benchmark comprises 52.6k question-answer pairs from 16 countries/regions, in 13 different languages, including low-resource ones such as Amharic, Assamese, Azerbaijani, Hausa, and Sundanese. We evaluate LLMs in two formats: short-answer questions, and multiple-choice questions. We show that LLMs perform better in cultures that are more present online, with a maximum 57.34% difference in GPT-4, the best-performing model, in the short-answer format.Furthermore, we find that LLMs perform better in their local languages for mid-to-high-resource languages. Interestingly, for languages deemed to be low-resource, LLMs provide better answers in English. We make our dataset publicly available at: https://github.com/nlee0212/BLEnD.
Junho Myung, Nayeon Lee, Yi Zhou 0019, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Pérez-Almendros, Abinew Ali Ayele, Víctor Gutiérrez-Basulto, Yazmín Ibáñez-García, Hwaran Lee, Shamsuddeen Hassan Muhammad, Ki-Woong Park, Anar Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, José Camacho-Collados, Alice Oh
NeurIPS19
2023 DecompX: Explaining Transformers Decisions by Propagating Token Decomposition
abstract
Ali Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ali Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
ACL (1)5
2023 Guide the Learner: Controlling Product of Experts Debiasing Method Based on Token Attribution Similarities
abstract
Several proposals have been put forward in recent years for improving out-of-distribution (OOD) performance through mitigating dataset biases.A popular workaround is to train a robust model by re-weighting training examples based on a secondary biased model.Here, the underlying assumption is that the biased model resorts to shortcut features.Hence, those training examples that are correctly predicted by the biased model are flagged as being biased and are down-weighted during the training of the main model.However, assessing the importance of an instance merely based on the predictions of the biased model may be too naive.It is possible that the prediction of the main model can be derived from another decisionmaking process that is distinct from the behavior of the biased model.To circumvent this, we introduce a fine-tuning strategy that incorporates the similarity between the main and biased model attribution scores in a Product of Experts (PoE) loss function to further improve OOD performance.With experiments conducted on natural language inference and fact verification benchmarks, we show that our method improves OOD results while maintaining in-distribution (ID) performance.1 ⋆ Work done as a Master's student at Iran University of Science and
Ali Modarressi, Hossein Amirkhani, Mohammad Taher Pilehvar
EACL3
2023 Pars-OFF: A Benchmark for Offensive Language Detection on Farsi Social Media
abstract
With the increasing use of social media with its ability for users to share comments immediately, the extent of a system to identify offensive content has become a necessity in all languages. Due to the lack of publicly available resources on offensive language identification for Farsi, which has more than 110 million speakers, we present Pars-OFF, a three-layered annotated corpus for offensive language detection in Farsi to fill the existing gap. The introduced corpus contains 10,563 data samples. The tweets have been collected with a combination of similarity-based and keyword-based data selection techniques to avoid severe unbalancedness. Additionally, as a baseline, this article reports the performance of the traditional machine learning approaches and Transformer based models over the Pars-OFF dataset. The best performance was obtained by the BERT+fastText model, yielding the F1-Macro score of 89.57.
Taha Shangipour Ataei, Kamyar Darvishi, Soroush Javdan, Amin Pourdabiri, Behrouz Minaei-Bidgoli, Mohammad Taher Pilehvar
IEEE Trans. Affect. Comput.6
2022 Incorporating Stock Market Signals for Twitter Stance Detection
abstract
Costanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Costanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier
ACL (1)3
2022 AdapLeR: Speeding up Inference by Adaptive Length Reduction
abstract
Pre-trained language models have shown stellar performance in various downstream tasks.But, this usually comes at the cost of high latency and computation, hindering their usage in resource-limited settings.In this work, we propose a novel approach for reducing the computational cost of BERT with minimal loss in downstream performance.Our method dynamically eliminates less contributing tokens through layers, resulting in shorter lengths and consequently lower computational cost.To determine the importance of each token representation, we train a Contribution Predictor for each layer using a gradient-based saliency method.Our experiments on several diverse classification tasks show speedups up to 22x during inference time without much sacrifice in performance.We also validate the quality of the selected tokens in our method using human annotations in the ERASER benchmark.In comparison to other widely used strategies for selecting important tokens, such as saliency and attention, our proposed method has a significantly lower false positive rate in generating rationales.Our code is freely available at https://github.com/amodaresi/ AdapLeR.
Ali Modarressi, Hosein Mohebbi, Mohammad Taher Pilehvar
ACL (1)3
2022 An Empirical Study on the Transferability of Transformer Modules in Parameter-efficient Fine-tuning
abstract
Parameter-efficient fine-tuning approaches have recently garnered a lot of attention.Having considerably lower number of trainable weights, these methods can bring about scalability and computational effectiveness.In this paper, we look for optimal subnetworks and investigate the capability of different transformer modules in transferring knowledge from a pre-trained model to a downstream task.Our empirical results suggest that every transformer module in BERT can act as a winning ticket: fine-tuning each specific module while keeping the rest of the network frozen can lead to comparable performance to the full fine-tuning.Among different modules, LayerNorms exhibit the best capacity for knowledge transfer with limited trainable weights, to the extent that, with only 0.003% of all parameters in the layer-wise analysis, they show acceptable performance on various target tasks.On the reasons behind their effectiveness, we argue that their notable performance could be attributed to their high-magnitude weights compared to that of the other modules in the pre-trained BERT.The code for this paper is freely available at https://github.com/ m-tajari/transformer-transferability.
Mohammad Akbar-Tajari, Sara Rajaee, Mohammad Taher Pilehvar
EMNLP3
2022 Looking at the Overlooked: An Analysis on the Word-Overlap Bias in Natural Language Inference
abstract
It has been shown that NLI models are usually biased with respect to the word-overlap between premise and hypothesis; they take this feature as a primary cue for predicting the entailment label.In this paper, we focus on an overlooked aspect of the overlap bias in NLI models: the reverse word-overlap bias.Our experimental results demonstrate that current NLI models are highly biased towards the non-entailment label on instances with low overlap, and the existing debiasing methods, which are reportedly successful on existing challenge datasets, are generally ineffective in addressing this category of bias.We investigate the reasons for the emergence of the overlap bias and the role of minority examples in its mitigation.For the former, we find that the word-overlap bias does not stem from pre-training, and for the latter, we observe that in contrast to the accepted assumption, eliminating minority examples does not affect the generalizability of debiasing methods with respect to the overlap bias.All the code and relevant data are available at: https: //github.com/sara-rajaee/reverse_bias Overlap Sample LabelFull (1.0) P: A little kid in blue is sledding down a snowy hill.H: A little kid in blue sledding.Entailment P: The young lady is giving the old man a hug.H: The young man is giving the old man a hug.Non-Entailment 12 13 = 0.923 P: A woman in a blue shirt and green hat looks up at the camera.H: A woman wearing a blue shirt and green hat looks at the camera Entailment 11 12 = 0.917 P: Two men in wheelchairs are reaching in the air for a basketball.H: Two women in wheelchairs are reaching in the air for a basketball.Non-Entailment 1 14 = 0.071 P: Several young people sit at a table playing poker.H: Youthful Human beings are gathered around a flat surface to play a card game.
Sara Rajaee, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
EMNLP3
2022 GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers
abstract
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
NAACL-HLT4
2022 PheneBank: a literature-based database of phenotypes
abstract
MOTIVATION: Significant effort has been spent by curators to create coding systems for phenotypes such as the Human Phenotype Ontology, as well as disease-phenotype annotations. We aim to support the discovery of literature-based phenotypes and integrate them into the knowledge discovery process. RESULTS: PheneBank is a Web-portal for retrieving human phenotype-disease associations that have been text-mined from the whole of Medline. Our approach exploits state-of-the-art machine learning for concept identification by utilizing an expert annotated rare disease corpus from the PMC Text Mining subset. Evaluation of the system for entities is conducted on a gold-standard corpus of rare disease sentences and for associations against the Monarch initiative data. AVAILABILITY AND IMPLEMENTATION: The PheneBank Web-portal freely available at http://www.phenebank.org. Annotated Medline data is available from Zenodo at DOI: 10.5281/zenodo.1408800. Semantic annotation software is freely available for non-commercial use at GitHub: https://github.com/pilehvar/phenebank. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mohammad Taher Pilehvar, Adam Bernard, Damian Smedley, Nigel Collier
Bioinform.1
2021 WiC-TSV: An Evaluation Benchmark for Target Sense Verification of Words in Context
abstract
Anna Breit, Artem Revenko, Kiamehr Rezaee, Mohammad Taher Pilehvar, Jose Camacho-Collados. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Anna Breit, Artem Revenko, Kiamehr Rezaee, Mohammad Taher Pilehvar, José Camacho-Collados
EACL4
2021 Exploring the Role of BERT Token Representations to Explain Sentence Probing Results
abstract
Several studies have been carried out on revealing linguistic features captured by BERT.This is usually achieved by training a diagnostic classifier on the representations obtained from different layers of BERT.The subsequent classification accuracy is then interpreted as the ability of the model in encoding the corresponding linguistic property.Despite providing insights, these studies have left out the potential role of token representations.In this paper, we provide a more in-depth analysis on the representation space of BERT in search for distinct and meaningful subspaces that can explain the reasons behind these probing results.Based on a set of probing tasks and with the help of attribution methods we show that BERT tends to encode meaningful knowledge in specific token representations (which are often ignored in standard classification setups), allowing the model to detect syntactic and semantic abnormalities, and to distinctively separate grammatical number and tense subspaces.1
Hosein Mohebbi, Ali Modarressi, Mohammad Taher Pilehvar
EMNLP (1)3
2021 Analysis and Evaluation of Language Models for Word Sense Disambiguation
abstract
Abstract Transformer-based language models have taken many fields in NLP by storm. BERT and its derivatives dominate most of the existing evaluation benchmarks, including those for Word Sense Disambiguation (WSD), thanks to their ability in capturing context-sensitive semantic nuances. However, there is still little knowledge about their capabilities and potential limitations in encoding and recovering word senses. In this article, we provide an in-depth quantitative and qualitative analysis of the celebrated BERT model with respect to lexical ambiguity. One of the main conclusions of our analysis is that BERT can accurately capture high-level sense distinctions, even when a limited number of examples is available for each word sense. Our analysis also reveals that in some cases language models come close to solving coarse-grained noun disambiguation under ideal conditions in terms of availability of training data and computing resources. However, this scenario rarely occurs in real-world settings and, hence, many practical challenges remain even in the coarse-grained setting.We also perform an in-depth comparison of the two main language model based WSD strategies, namely, fine-tuning and feature extraction, finding that the latter approach is more robust with respect to sense bias and it can better exploit limited available training data. In fact, the simple feature extraction strategy of averaging contextualized embeddings proves robust even using only three training sentences per word sense, with minimal improvements obtained by increasing the size of this training data.
Daniel Loureiro, Kiamehr Rezaee, Mohammad Taher Pilehvar, José Camacho-Collados
Comput. Linguistics3
2020 Will-They-Won't-They: A Very Large Dataset for Stance Detection on Twitter
abstract
Costanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Costanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier
ACL3
2020 XL-WiC: A Multilingual Benchmark for Evaluating Semantic Contextualization
abstract
The ability to correctly model distinct meanings of a word is crucial for the effectiveness of semantic representation techniques.However, most existing evaluation benchmarks for assessing this criterion are tied to sense inventories (usually WordNet), restricting their usage to a small subset of knowledge-based representation techniques.The Word-in-Context dataset (WiC) addresses the dependence on sense inventories by reformulating the standard disambiguation task as a binary classification problem; but, it is limited to the English language.We put forward a large multilingual benchmark, XL-WiC, featuring gold standards in 12 new languages from varied language families and with different degrees of resource availability, opening room for evaluation scenarios such as zero-shot cross-lingual transfer.We perform a series of experiments to determine the reliability of the datasets and to set performance baselines for several recent contextualized multilingual models.Experimental results show that even when no tagged instances are available for a target language, models trained solely on the English data can attain competitive performance in the task of distinguishing different meanings of a word, even for distant languages.XL-WiC is available at https://pilehvar.github.io/xlwic/.
Alessandro Raganato, Tommaso Pasini, José Camacho-Collados, Mohammad Taher Pilehvar
EMNLP (1)4
2019 Unseen Word Representation by Aligning Heterogeneous Lexical Semantic Spaces
abstract
Word embedding techniques heavily rely on the abundance of training data for individual words. Given the Zipfian distribution of words in natural language texts, a large number of words do not usually appear frequently or at all in the training data. In this paper we put forward a technique that exploits the knowledge encoded in lexical resources, such as WordNet, to induce embeddings for unseen words. Our approach adapts graph embedding and cross-lingual vector space transformation techniques in order to merge lexical knowledge encoded in ontologies with that derived from corpus statistics. We show that the approach can provide consistent performance improvements across multiple evaluation benchmarks: in-vitro, on multiple rare word similarity datasets, and invivo, in two downstream text classification tasks.
Victor Prokhorov, Mohammad Taher Pilehvar, Dimitri Kartsaklis, Pietro Liò, Nigel Collier
AAAI2
2018 Which Melbourne? Augmenting Geocoding with Maps
abstract
The purpose of text geolocation is to associate geographic information contained in a document with a set (or sets) of coordinates, either implicitly by using linguistic features and/or explicitly by using geographic metadata combined with heuristics.We introduce a geocoder (location mention disambiguator) that achieves state-of-the-art (SOTA) results on three diverse datasets by exploiting the implicit lexical clues.Moreover, we propose a new method for systematic encoding of geographic metadata to generate two distinct views of the same text.To that end, we introduce the Map Vector (MapVec), a sparse representation obtained by plotting prior geographic probabilities, derived from population figures, on a World Map.We then integrate the implicit (language) and explicit (map) features to significantly improve a range of metrics.We also introduce an open-source dataset for geoparsing of news events covering global disease outbreaks and epidemics to help future evaluation in geoparsing.
Milan Gritta, Mohammad Taher Pilehvar, Nigel Collier
ACL (1)2
2018 Mapping Text to Knowledge Graph Entities using Multi-Sense LSTMs
abstract
This paper addresses the problem of mapping natural language text to knowledge base entities.The mapping process is approached as a composition of a phrase or a sentence into a point in a multi-dimensional entity space obtained from a knowledge graph.The compositional model is an LSTM equipped with a dynamic disambiguation mechanism on the input word embeddings (a Multi-Sense LSTM), addressing polysemy issues.Further, the knowledge base space is prepared by collecting random walks from a graph enhanced with textual features, which act as a set of semantic bridges between text and knowledge base entities.The ideas of this work are demonstrated on largescale text-to-entity mapping and entity classification tasks, with state of the art results.* This paper is dedicated to the memory of Euripides Kartsaklis, a man who loved technology.
Dimitri Kartsaklis, Mohammad Taher Pilehvar, Nigel Collier
EMNLP2
2018 Large-scale Exploration of Neural Relation Classification Architectures
abstract
Experimental performance on the task of relation classification has generally improved using deep neural network architectures.One major drawback of reported studies is that individual models have been evaluated on a very narrow range of datasets, raising questions about the adaptability of the architectures, while making comparisons between approaches difficult.In this work, we present a systematic large-scale analysis of neural relation classification architectures on six benchmark datasets with widely varying characteristics.We propose a novel multi-channel LSTM model combined with a CNN that takes advantage of all currently popular linguistic and architectural features.Our 'Man for All Seasons' approach achieves state-of-the-art performance on two datasets.More importantly, in our view, the model allowed us to obtain direct insights into the continued challenges faced by neural language models on this task.
Hoang-Quynh Le, Duy-Cat Can, Sinh T. Vu, Mohammad Taher Pilehvar, Nigel Collier
EMNLP5
2018 Card-660: A Reliable Evaluation Framework for Rare Word Representation Models
abstract
Rare word representation has recently enjoyed a surge of interest, owing to the crucial role that effective handling of infrequent words can play in accurate semantic understanding.However, there is a paucity of reliable benchmarks for evaluation and comparison of these techniques.We show in this paper that the only existing benchmark (the Stanford Rare Word dataset) suffers from low-confidence annotations and limited vocabulary; hence, it does not constitute a solid comparison framework.In order to fill this evaluation gap, we propose CAmbridge Rare word Dataset (CARD-660), an expert-annotated word similarity dataset which provides a highly reliable, yet challenging, benchmark for rare word representation techniques.Through a set of experiments we show that even the best mainstream word embeddings, with millions of words in their vocabularies, are unable to achieve performances higher than 0.43 (Pearson correlation) on the dataset, compared to a human-level upperbound of 0.90.We release the dataset and the annotation materials at https:// pilehvar.github.io/card-660/.
Mohammad Taher Pilehvar, Dimitri Kartsaklis, Victor Prokhorov, Nigel Collier
EMNLP1
2018 From Word To Sense Embeddings: A Survey on Vector Representations of Meaning
abstract
Over the past years, distributed semantic representations have proved to be effective and flexible keepers of prior knowledge to be integrated into downstream applications. This survey focuses on the representation of meaning. We start from the theoretical background behind word vector space models and highlight one of their major limitations: the meaning conflation deficiency, which arises from representing a word with all its possible meanings as a single vector. Then, we explain how this deficiency can be addressed through a transition from the word level to the more fine-grained level of word senses (in its broader acceptation) as a method for modelling unambiguous lexical meaning. We present a comprehensive overview of the wide range of techniques in the two main branches of sense representation, i.e., unsupervised and knowledge-based. Finally, this survey covers the main evaluation procedures and applications for this type of representation, and provides an analysis of four of its important aspects: interpretability, sense granularity, adaptability to different domains and compositionality.
José Camacho-Collados, Mohammad Taher Pilehvar
J. Artif. Intell. Res.2
2017 Vancouver Welcomes You! Minimalist Location Metonymy Resolution
abstract
Named entities are frequently used in a metonymic manner.They serve as references to related entities such as people and organisations.Accurate identification and interpretation of metonymy can be directly beneficial to various NLP applications, such as Named Entity Recognition and Geographical Parsing.Until now, metonymy resolution (MR) methods mainly relied on parsers, taggers, dictionaries, external word lists and other handcrafted lexical resources.We show how a minimalist neural approach combined with a novel predicate window method can achieve competitive results on the Se-mEval 2007 task on Metonymy Resolution.Additionally, we contribute with a new Wikipedia-based MR dataset called RelocaR, which is tailored towards locations as well as improving previous deficiencies in annotation guidelines.
Milan Gritta, Mohammad Taher Pilehvar, Nut Limsopatham, Nigel Collier
ACL (1)2
2017 Towards a Seamless Integration of Word Senses into Downstream NLP Applications
abstract
Lexical ambiguity can impede NLP systems from accurate understanding of semantics.Despite its potential benefits, the integration of sense-level information into NLP systems has remained understudied.By incorporating a novel disambiguation algorithm into a state-of-the-art classification model, we create a pipeline to integrate sense-level information into downstream NLP applications.We show that a simple disambiguation of the input text can lead to consistent performance improvement on multiple topic categorization and polarity detection datasets, particularly when the fine granularity of the underlying sense inventory is reduced and the document is sufficiently large.Our results also point to the need for sense representation research to focus more on in vivo evaluations which target the performance in downstream NLP applications rather than artificial benchmarks.
Mohammad Taher Pilehvar, José Camacho-Collados, Roberto Navigli, Nigel Collier
ACL (1)1
2016 Embeddings for Word Sense Disambiguation: An Evaluation Study
abstract
Recent years have seen a dramatic growth in the popularity of word embeddings mainly owing to their ability to capture semantic information from massive amounts of textual content.As a result, many tasks in Natural Language Processing have tried to take advantage of the potential of these distributional models.In this work, we study how word embeddings can be used in Word Sense Disambiguation, one of the oldest tasks in Natural Language Processing and Artificial Intelligence.We propose different methods through which word embeddings can be leveraged in a state-of-the-art supervised WSD system architecture, and perform a deep analysis of how different parameters affect performance.We show how a WSD system that makes use of word embeddings alone, if designed properly, can provide significant performance improvement over a state-ofthe-art WSD system that incorporates several standard WSD features.
Ignacio Iacobacci, Mohammad Taher Pilehvar, Roberto Navigli
ACL (1)2
2016 De-Conflated Semantic Representations
abstract
One major deficiency of most semantic representation techniques is that they usually model a word type as a single point in the semantic space, hence conflating all the meanings that the word can have.Addressing this issue by learning distinct representations for individual meanings of words has been the subject of several research studies in the past few years.However, the generated sense representations are either not linked to any sense inventory or are unreliable for infrequent word senses.We propose a technique that tackles these problems by de-conflating the representations of words based on the deep knowledge that can be derived from a semantic network.Our approach provides multiple advantages in comparison to the previous approaches, including its high coverage and the ability to generate accurate representations even for infrequent word senses.We carry out evaluations on six datasets across two semantic similarity tasks and report state-of-the-art results on most of them.
Mohammad Taher Pilehvar, Nigel Collier
EMNLP1
2016 Nasari: Integrating explicit knowledge and corpus statistics for a multilingual representation of concepts and entities
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli
Artif. Intell.2
2015 A Unified Multilingual Semantic Representation of Concepts
abstract
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli
ACL (1)2
2015 SensEmbed: Learning Sense Embeddings for Word and Relational Similarity
abstract
Ignacio Iacobacci, Mohammad Taher Pilehvar, Roberto Navigli. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Ignacio Iacobacci, Mohammad Taher Pilehvar, Roberto Navigli
ACL (1)2
2015 NASARI: a Novel Approach to a Semantically-Aware Representation of Items
abstract
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli
HLT-NAACL2
2015 Reserating the awesometastic: An automatic extension of the WordNet taxonomy for novel terms
abstract
This paper presents CROWN, an automatically con-structed extension of WordNet that augments its taxonomy with novel lemmas from Wiktionary. CROWN fills the important gap in WordNet’s lexi-con for slang, technical, and rare lemmas, and more than doubles its current size. In two evaluations, we demonstrate that the construction procedure is accu-rate and has a significant impact on a WordNet-based algorithm encountering novel lemmas. 1
David Jurgens, Mohammad Taher Pilehvar
HLT-NAACL2
2015 An Open-source Framework for Multi-level Semantic Similarity Measurement
abstract
We present an open source, freely available Java implementation of Align, Disambiguate, and Walk (ADW), a state-of-the-art approach for measuring semantic similarity based on the Personalized PageRank algorithm.A pair of linguistic items, such as phrases or sentences, are first disambiguated using an alignment-based disambiguation technique and then modeled using random walks on the WordNet graph.ADW provides three main advantages: (1) it is applicable to all types of linguistic items, from word senses to texts;(2) it is all-in-one, i.e., it does not need any additional resource, training or tuning; and(3) it has proven to be highly reliable at different lexical levels and multiple evaluation benchmarks.We are releasing the source code at https://github.com/pilehvar/adw/.We also provide at http://lcl.uniroma1.it/adw/a Web interface and a Java API that can be seamlessly integrated into other NLP systems requiring semantic similarity measurement.
Mohammad Taher Pilehvar, Roberto Navigli
HLT-NAACL1
2015 From senses to texts: An all-in-one graph-based approach for measuring semantic similarity
Mohammad Taher Pilehvar, Roberto Navigli
Artif. Intell.1
2014 A Robust Approach to Aligning Heterogeneous Lexical Resources
abstract
Lexical resource alignment has been an active field of research over the last decade. However, prior methods for align-ing lexical resources have been either spe-cific to a particular pair of resources, or heavily dependent on the availability of hand-crafted alignment data for the pair of resources to be aligned. Here we present a unified approach that can be applied to an arbitrary pair of lexical resources, includ-ing machine-readable dictionaries with no network structure. Our approach leverages a similarity measure that enables the struc-tural comparison of senses across lexical resources, achieving state-of-the-art per-formance on the task of aligning WordNet to three different collaborative resources: Wikipedia, Wiktionary and OmegaWiki. 1
Mohammad Taher Pilehvar, Roberto Navigli
ACL (1)1
2014 A Large-Scale Pseudoword-Based Evaluation Framework for State-of-the-Art Word Sense Disambiguation
abstract
The evaluation of several tasks in lexical semantics is often limited by the lack of large amounts of manual annotations, not only for training purposes, but also for testing purposes. Word Sense Disambiguation (WSD) is a case in point, as hand-labeled datasets are particularly hard and time-consuming to create. Consequently, evaluations tend to be performed on a small scale, which does not allow for in-depth analysis of the factors that determine a systems' performance. In this paper we address this issue by means of a realistic simulation of large-scale evaluation for the WSD task. We do this by providing two main contributions: First, we put forward two novel approaches to the wide-coverage generation of semantically aware pseudowords (i.e., artificial words capable of modeling real polysemous words); second, we leverage the most suitable type of pseudoword to create large pseudosense-annotated corpora, which enable a large-scale experimental framework for the comparison of state-of-the-art supervised and knowledge-based algorithms. Using this framework, we study the impact of supervision and knowledge on the two major disambiguation paradigms and perform an in-depth analysis of the factors which affect their performance.
Mohammad Taher Pilehvar, Roberto Navigli
Comput. Linguistics1
2013 Align, Disambiguate and Walk: A Unified Approach for Measuring Semantic Similarity
Mohammad Taher Pilehvar, David Jurgens, Roberto Navigli
ACL (1)1
2013 Paving the Way to a Large-scale Pseudosense-annotated Dataset
Mohammad Taher Pilehvar, Roberto Navigli
HLT-NAACL1
2011 TEP: Tehran English-Persian Parallel Corpus
Mohammad Taher Pilehvar, Heshaam Faili, Abdol Hamid Pilevar
CICLing (2)1