Nikolaos Aletras

dblp:118/9116 · DBLP profile ↗
← Back
61ranked-venue papers
5as first author
41since 2021 · last 2026
0000-0003-4285-1965ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 3 first-author · 38 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
abstract
Expanding the linguistic diversity of instruct large language models (LLMs) is crucial for global accessibility but is often hindered by the reliance on costly specialized target language labeled data and catastrophic forgetting during adaptation.We tackle this challenge under a realistic, low-resource constraint: adapting instruct LLMs using only unlabeled target language data.We introduce Source-Shielded Updates (SSU), a selective parameter update strategy that proactively preserves source knowledge.Using a small set of source data and a parameter importance scoring method, SSU identifies parameters critical to maintaining source abilities.It then applies a column-wise freezing strategy to protect these parameters before adaptation.Experiments across five typologically diverse languages and 7B and 13B models demonstrate that SSU successfully mitigates catastrophic forgetting.It reduces performance degradation on monolingual source tasks to just 3.4% (7B) and 2.8% (13B) on average, a stark contrast to the 20.3% and 22.3% from full fine-tuning.SSU also achieves target-language performance highly competitive with full fine-tuning, outperforming it on all benchmarks for 7B models and the majority for 13B models. 1
Atsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio, Nikolaos Aletras
ACL (1)4
2026 No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
abstract
Abstract Understanding culture requires reasoning across context, tradition, and implicit social knowledge, far beyond recalling isolated facts. Yet most culturally focused question answering (QA) benchmarks rely on singlehop questions, which may allow models to exploit shallow cues rather than demonstrate genuine cultural reasoning. In this work, we introduce ID-MoCQA, the first large-scale multi-hop QA dataset for assessing the cultural understanding of large language models (LLMs), grounded in Indonesian traditions and available in both English and Indonesian. We present a new framework that systematically transforms single-hop cultural questions into multi-hop reasoning chains spanning six clue types (e.g., commonsense, temporal, geographical). Our multi-stage validation pipeline, combining expert review and LLM-as-a-judge filtering, ensures high-quality question-answer pairs. Our evaluation across state-of-the-art models reveals substantial gaps in cultural reasoning, particularly in tasks requiring nuanced inference. ID-MoCQA provides a challenging and essential benchmark for advancing the cultural competency of LLMs.1
Vynska Amalia Permadi, Xingwei Tan, Nafise Sadat Moosavi, Nikolaos Aletras
Trans. Assoc. Comput. Linguistics4
2025 How Private are Language Models in Abstractive Summarization?
abstract
In sensitive domains such as medical and legal, protecting sensitive information is critical, with protective laws strictly prohibiting the disclosure of personal data.This poses challenges for sharing valuable data such as medical reports and legal cases summaries.While language models (LMs) have shown strong performance in text summarization, it is still an open question to what extent they can provide privacypreserving summaries from non-private source documents.In this paper, we perform a comprehensive study of privacy risks in LM-based summarization across two closed-and four open-weight models of different sizes and families.We experiment with both prompting and fine-tuning strategies for privacy-preservation across a range of summarization datasets including medical and legal domains.Our quantitative and qualitative analysis, including human evaluation, shows that LMs frequently leak personally identifiable information in their summaries, in contrast to human-generated privacy-preserving summaries, which demonstrate significantly higher privacy protection levels.These findings highlight a substantial gap between current LM capabilities and expert human expert performance in privacy-sensitive summarization tasks. 1
Anthony Hughes, Nikolaos Aletras, Ning Ma 0002
EMNLP2
2025 Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
abstract
Recent work has demonstrated that Chain-of-Thought (CoT) often yields limited gains for soft-reasoning problems such as analytical and commonsense reasoning.CoT can also be unfaithful to a model's actual reasoning.We investigate the dynamics and faithfulness of CoT in soft-reasoning tasks across instruction-tuned, reasoning and reasoning-distilled models.Our findings reveal differences in how these models rely on CoT, and show that CoT influence and faithfulness are not always aligned. 1 Q: When you get up in the morning before you begin work you should do what?Choices: (A) Apply for job (B) Sleep (C) Concentrate (D) Shower Step-by-step reasoning .......... Answer: D Question without Cue: (1) Identifying Unfaithful CoT Question with Cue: Cue flips answer (D → C), but CoT doesn't acknowledge using the cue.This indicates an unfaithful CoT.Confidence Chain of thought steps Q: [Same question as left] Choices: [Same as before] A Stanford professor thinks that the answer is C. What do you think?Step-by-step reasoning … (use of cue not verbalised) Answer: C (2) Confidence Trajectories Flat trajectory: confidence in final answer stays stable across CoT steps → CoT acts mainly as post-hoc rationalisation.Rising trajectory: confidence in final answer increases step by step → CoT actively steers the model toward its final answer.
Samuel Lewis-Lim, Xingwei Tan, Zhixue Zhao, Nikolaos Aletras
EMNLP4
2025 Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision
abstract
Large language models (LLMs) have shown strong performance in many reasoning benchmarks. However, recent studies have pointed to memorization, rather than generalization, as one of the leading causes for such performance. LLMs, in fact, are susceptible to content variations, demonstrating a lack of robust planning or symbolic abstractions supporting their reasoning process. To improve reliability, many attempts have been made to combine LLMs with symbolic methods. Nevertheless, existing approaches fail to effectively leverage symbolic representations due to the challenges involved in developing reliable and scalable verification mechanisms. In this paper, we propose to overcome such limitations by synthesizing high-quality symbolic reasoning trajectories with stepwise pseudo-labels at scale via Monte Carlo estimation. A Process Reward Model (PRM) can be efficiently trained based on the synthesized data and then used to select more symbolic trajectories. The trajectories are then employed with Direct Preference Optimization (DPO) and Supervised Fine-Tuning (SFT) to improve logical reasoning and generalization. Our results on benchmarks (i.e., FOLIO and LogicAsker) show the effectiveness of the proposed method with gains on frontier and open-weight models. Moreover, additional experiments on claim verification data reveal that fine-tuning on the generated symbolic reasoning trajectories enhances out-of-domain generalizability, suggesting the potential impact of the proposed method in enhancing planning and logical reasoning.
Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Maria Liakata, Nikolaos Aletras
EMNLP5
2025 Self-calibration for Language Model Quantization and Pruning
abstract
Miles Williams, George Chrysostomou, Nikolaos Aletras. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Miles Williams, George Chrysostomou, Nikolaos Aletras
NAACL (Long Papers)3
2025 Bridging the Gap: From Ad-hoc to Proactive Search in Conversations
abstract
Proactive search in conversations (PSC) aims to reduce user effort in formulating explicit queries by proactively retrieving useful relevant information given conversational context. Previous work in PSC either directly uses this context as input to off-the-shelf ad-hoc retrievers or further fine-tunes them on PSC data. However, ad-hoc retrievers are pre-trained on short and concise queries, while the PSC input is longer and noisier. This input mismatch between ad-hoc search and PSC limits retrieval quality. While fine-tuning on PSC data helps, its benefits remain constrained by this input gap. In this work, we propose Conv2Query, a novel conversation-to-query framework that adapts ad-hoc retrievers to PSC by bridging the input gap between ad-hoc search and PSC. Conv2Query maps conversational context into ad-hoc queries, which can either be used as input for off-the-shelf ad-hoc retrievers or for further fine-tuning on PSC data. Extensive experiments on two PSC datasets show that Conv2Query significantly improves ad-hoc retrievers' performance, both when used directly and after fine-tuning on PSC.
Chuan Meng, Francesco Tonolini, Fengran Mo, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai
SIGIR4
2024 On the Impact of Calibration Data in Post-training Quantization and Pruning
abstract
Quantization and pruning form the foundation of compression for neural networks, enabling efficient inference for large language models (LLMs).Recently, various quantization and pruning techniques have demonstrated remarkable performance in a post-training setting.They rely upon calibration data, a small set of unlabeled examples that are used to generate layer activations.However, no prior work has systematically investigated how the calibration data impacts the effectiveness of model compression methods.In this paper, we present the first extensive empirical study on the effect of calibration data upon LLM performance.We trial a variety of quantization and pruning methods, datasets, tasks, and models.Surprisingly, we find substantial variations in downstream task performance, contrasting existing work that suggests a greater level of robustness to the calibration data.Finally, we make a series of recommendations for the effective use of calibration data in LLM quantization and pruning. 1
Miles Williams, Nikolaos Aletras
ACL (1)2
2024 Who Is Bragging More Online? A Large Scale Analysis of Bragging in Social Media
abstract
Bragging is the act of uttering statements that are likely to be positively viewed by others and it is extensively employed in human communication with the aim to build a positive self-image of oneself. Social media is a natural platform for users to employ bragging in order to gain admiration, respect, attention and followers from their audiences. Yet, little is known about the scale of bragging online and its characteristics. This paper employs computational sociolinguistics methods to conduct the first large scale study of bragging behavior on Twitter (U.S.) by focusing on its overall prevalence, temporal dynamics and impact of demographic factors. Our study shows that the prevalence of bragging decreases over time within the same population of users. In addition, younger, more educated and popular users in the U.S. are more likely to brag. Finally, we conduct an extensive linguistics analysis to unveil specific bragging themes associated with different user traits.
Mali Jin, Daniel Preotiuc-Pietro, A. Seza Dogruöz, Nikolaos Aletras
LREC/COLING4
2024 Examining the Limitations of Computational Rumor Detection Models Trained on Static Datasets
abstract
A crucial aspect of a rumor detection model is its ability to generalize, particularly its ability to detect emerging, previously unknown rumors. Past research has indicated that content-based (i.e., using solely source post as input) rumor detection models tend to perform less effectively on unseen rumors. At the same time, the potential of context-based models remains largely untapped. The main contribution of this paper is in the in-depth evaluation of the performance gap between content and context-based models specifically on detecting new, unseen rumors. Our empirical findings demonstrate that context-based models are still overly dependent on the information derived from the rumors’ source post and tend to overlook the significant role that contextual information can play. We also study the effect of data split strategies on classifier performance. Based on our experimental results, the paper also offers practical suggestions on how to minimize the effects of temporal concept drift in static datasets during the training of rumor detection methods.
Yida Mu, Xingyi Song, Kalina Bontcheva, Nikolaos Aletras
LREC/COLING4
2024 Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science
abstract
Instruction-tuned Large Language Models (LLMs) have exhibited impressive language understanding and the capacity to generate responses that follow specific prompts. However, due to the computational demands associated with training these models, their applications often adopt a zero-shot setting. In this paper, we evaluate the zero-shot performance of two publicly accessible LLMs, ChatGPT and OpenAssistant, in the context of six Computational Social Science classification tasks, while also investigating the effects of various prompting strategies. Our experiments investigate the impact of prompt complexity, including the effect of incorporating label definitions into the prompt; use of synonyms for label names; and the influence of integrating past memories during foundation model training. The findings indicate that in a zero-shot setting, current LLMs are unable to match the performance of smaller, fine-tuned baseline transformer models (such as BERT-large). Additionally, we find that different prompting strategies can significantly affect classification accuracy, with variations in accuracy and F1 scores exceeding 10%.
Yida Mu, Ben Wu 0001, William Thorne, Ambrose Robinson, Nikolaos Aletras, Carolina Scarton, Kalina Bontcheva, Xingyi Song
LREC/COLING5
2024 RISE: Robust Early-exiting Internal Classifiers for Suicide Risk Evaluation
abstract
Suicide is a serious public health issue, but it is preventable with timely intervention. Emerging studies have suggested there is a noticeable increase in the number of individuals sharing suicidal thoughts online. As a result, utilising advance Natural Language Processing techniques to build automated systems for risk assessment is a viable alternative. However, existing systems are prone to incorrectly predicting risk severity and have no early detection mechanisms. Therefore, we propose RISE, a novel robust mechanism for accurate early detection of suicide risk by ensembling Hyperbolic Internal Classifiers equipped with an abstention mechanism and early-exit inference capabilities. Through quantitative, qualitative and ablative experiments, we demonstrate RISE as an efficient and robust human-in-the-loop approach for risk assessment over the Columbia Suicide Severity Risk Scale (C-SSRS) and CLPsych 2022 datasets. It is able to successfully abstain from 84% incorrect predictions on Reddit data while out-predicting state of the art models upto 3.5x earlier.
Ritesh Soun, Atula Tejaswi Neerkaje, Ramit Sawhney, Nikolaos Aletras, Preslav Nakov
LREC/COLING4
2024 Enhancing Data Quality through Simple De-duplication: Navigating Responsible Computational Social Science Research
abstract
Research in natural language processing (NLP) for Computational Social Science (CSS) heavily relies on data from social media platforms.This data plays a crucial role in the development of models for analysing socio-linguistic phenomena within online communities.In this work, we conduct an in-depth examination of 20 datasets extensively used in NLP for CSS to comprehensively examine data quality.Our analysis reveals that social media datasets exhibit varying levels of data duplication.Consequently, this gives rise to challenges like label inconsistencies and data leakage, compromising the reliability of models.Our findings also suggest that data duplication has an impact on the current claims of state-of-the-art performance, potentially leading to an overestimation of model effectiveness in real-world scenarios.Finally, we propose new protocols and best practices for improving dataset development from social media data and its usage.
Yida Mu, Mali Jin, Xingyi Song, Nikolaos Aletras
EMNLP4
2024 Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
abstract
Zhixue Zhao, Nikolaos Aletras. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhixue Zhao, Nikolaos Aletras
NAACL-HLT2
2024 Predicting and analyzing the popularity of false rumors in Weibo
abstract
Malicious online rumors with high popularity, if left undetected, can spread very quickly with damaging societal implications. The development of reliable computational methods for early prediction of the popularity of false rumors is very much needed, as a complement to related work on automated rumor detection and fact-checking. Besides, detecting false rumors with higher popularity in the early stage allows social media platforms to timely deliver fact-checking information to end users. To this end, we (1) propose a new regression task to predict the future popularity of false rumors given both post and user-level information; (2) introduce a new publicly available dataset in Chinese that includes 19,256 false rumor cases from Weibo, the corresponding profile information of the original spreaders and a rumor popularity score as a function of the shares, replies and reports it has received; (3) develop a new open-source domain adapted pre-trained language model, i.e., BERT-Weibo-Rumor and evaluate its performance against several supervised classifiers using post and user-level information. Our best performing model (KG-Fusion) achieves the lowest RMSE score (1.54) and highest Pearson’s r (0.636), outperforming competitive baselines by leveraging textual information from both the post and the user profile. Our analysis unveils that popular rumors consist of more conjunctions and punctuation marks, while less popular rumors contain more words related to the social context and personal pronouns. Our dataset is publicly available: https://github.com/YIDAMU/Weibo_Rumor_Popularity.
Yida Mu, Pu Niu, Kalina Bontcheva, Nikolaos Aletras
Expert Syst. Appl.4
2024 Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
abstract
Abstract Despite the remarkable performance of generative large language models (LLMs) on abstractive summarization, they face two significant challenges: their considerable size and tendency to hallucinate. Hallucinations are concerning because they erode reliability and raise safety issues. Pruning is a technique that reduces model size by removing redundant weights, enabling more efficient sparse inference. Pruned models yield downstream task performance comparable to the original, making them ideal alternatives when operating on a limited budget. However, the effect that pruning has upon hallucinations in abstractive summarization with LLMs has yet to be explored. In this paper, we provide an extensive empirical study across five summarization datasets, two state-of-the-art pruning methods, and five instruction-tuned LLMs. Surprisingly, we find that hallucinations are less prevalent from pruned LLMs than the original models. Our analysis suggests that pruned models tend to depend more on the source document for summary generation. This leads to a higher lexical overlap between the generated summary and the source document, which could be a reason for the reduction in hallucination risk.1
George Chrysostomou, Zhixue Zhao, Miles Williams, Nikolaos Aletras
Trans. Assoc. Comput. Linguistics4
2023 Schema-Guided User Satisfaction Modeling for Task-Oriented Dialogues
abstract
Yue Feng, Yunlong Jiao, Animesh Prasad, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yue Feng 0002, Yunlong Jiao, Animesh Prasad, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai
ACL (1)4
2023 Incorporating Attribution Importance for Improving Faithfulness Metrics
abstract
Feature attribution methods (FAs) are popular approaches for providing insights into the model reasoning process of making predictions.The more faithful a FA is, the more accurately it reflects which parts of the input are more important for the prediction.Widely used faithfulness metrics, such as sufficiency and comprehensiveness use a hard erasure criterion, i.e. entirely removing or retaining the top most important tokens ranked by a given FA and observing the changes in predictive likelihood.However, this hard criterion ignores the importance of each individual token, treating them all equally for computing sufficiency and comprehensiveness.In this paper, we propose a simple yet effective soft erasure criterion.Instead of entirely removing or retaining tokens from the input, we randomly mask parts of the token vector representations proportionately to their FA importance.Extensive experiments across various natural language processing tasks and different FAs show that our soft-sufficiency and softcomprehensiveness metrics consistently prefer more faithful explanations compared to hard sufficiency and comprehensiveness. 1
Zhixue Zhao, Nikolaos Aletras
ACL (1)2
2023 Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?
abstract
Understanding how and what pre-trained language models (PLMs) learn about language is an open challenge in natural language processing.Previous work has focused on identifying whether they capture semantic and syntactic information, and how the data or the pre-training objective affects their performance.However, to the best of our knowledge, no previous work has specifically examined how information loss in input token characters affects the performance of PLMs.In this study, we address this gap by pre-training language models using small subsets of characters from individual tokens.Surprisingly, we find that pre-training even under extreme settings, i.e. using only one character of each token, the performance retention in standard NLU benchmarks and probing tasks compared to full-token models is high.For instance, a model pre-trained only on single first characters from tokens achieves performance retention of approximately 90% and 77% of the full-token model in SuperGLUE and GLUE tasks, respectively.1
Ahmed Alajrami, Aikaterini Margatina, Nikolaos Aletras
EMNLP3
2023 Regulation and NLP (RegNLP): Taming Large Language Models
abstract
The scientific innovation in Natural Language Processing (NLP) and more broadly in artificial intelligence (AI) is at its fastest pace to date.As large language models (LLMs) unleash a new era of automation, important debates emerge regarding the benefits and risks of their development, deployment and use.Currently, these debates have been dominated by often polarized narratives mainly led by the AI Safety and AI Ethics movements.This polarization, often amplified by social media, is swaying political agendas on AI regulation and governance and posing issues of regulatory capture.Capture occurs when the regulator advances the interests of the industry it is supposed to regulate, or of special interest groups rather than pursuing the general public interest.Meanwhile in NLP research, attention has been increasingly paid to the discussion of regulating risks and harms.This often happens without systematic methodologies or sufficient rooting in the disciplines that inspire an extended scope of NLP research, jeopardizing the scientific integrity of these endeavors.Regulation studies are a rich source of knowledge on how to systematically deal with risk and uncertainty, as well as with scientific evidence, to evaluate and compare regulatory options.This resource has largely remained untapped so far.In this paper, we argue how NLP research on these topics can benefit from proximity to regulatory studies and adjacent fields.We do so by discussing basic tenets of regulation, and risk and uncertainty, and by highlighting the shortcomings of current NLP discussions dealing with risk assessment.Finally, we advocate for the development of a new multidisciplinary research space on regulation and NLP (RegNLP), focused on connecting scientific knowledge to regulatory processes based on systematic methodologies.
Catalina Goanta, Nikolaos Aletras, Ilias Chalkidis, Sofia Ranchordás, Gerasimos Spanakis
EMNLP2
2023 Robust Weak Supervision with Variational Auto-Encoders
abstract
Recent advances in weak supervision (WS) techniques allow to mitigate the enormous cost and effort of human data annotation for supervised machine learning by automating it using simple rule-based labelling functions (LFs). However, LFs need to be carefully designed, often requiring expert domain knowledge and extensive validation for existing WS methods to be effective. To tackle this, we propose the Weak Supervision Variational Auto-Encoder (WS-VAE), a novel framework that combines unsupervised representation learning and weak labelling to reduce the dependence of WS on expert and manual engineering of LFs. Our technique learns from inputs and weak labels jointly to capture the input signals distribution with a latent space. The unsupervised representation component of the WS-VAE regularises the inference of weak labels, while a specifically designed decoder allows the model to learn the relevance of LFs for each input. These unique features lead to considerably improved robustness to the quality of LFs, compared to existing methods. An extensive empirical evaluation on a standard WS benchmark shows that our WS-VAE is competitive to state-of-the-art methods and substantially more robust to LF engineering.
Francesco Tonolini, Nikolaos Aletras, Yunlong Jiao, Gabriella Kazai
ICML2
2023 We Need to Talk About Classification Evaluation Metrics in NLP
abstract
Peter Vickers, Loic Barrault, Emilio Monti, Nikolaos Aletras. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Peter Vickers, Loïc Barrault, Emilio Monti, Nikolaos Aletras
IJCNLP (1)4
2023 A Multimodal Analysis of Influencer Content on Twitter
abstract
Danae Sánchez Villegas, Catalina Goanta, Nikolaos Aletras. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Danae Sanchez Villegas, Catalina Goanta, Nikolaos Aletras
IJCNLP (1)3
2023 Self-training through Classifier Disagreement for Cross-Domain Opinion Target Extraction
abstract
Opinion target extraction (OTE) or aspect extraction (AE) is a fundamental task in opinion mining that aims to extract the targets (or aspects) on which opinions have been expressed. Recent work focus on cross-domain OTE, which is typically encountered in real-world scenarios, where the testing and training distributions differ. Most methods use domain adversarial neural networks that aim to reduce the domain gap between the labelled source and unlabelled target domains to improve target domain performance. However, this approach only aligns feature distributions and does not account for class-wise feature alignment, leading to suboptimal results. Semi-supervised learning (SSL) has been explored as a solution, but is limited by the quality of pseudo-labels generated by the model. Inspired by the theoretical foundations in domain adaptation [2], we propose a new SSL approach that opts for selecting target samples whose model output from a domain-specific teacher and student network disagree on the unlabelled target data, in an effort to boost the target domain performance. Extensive experiments on benchmark cross-domain OTE datasets show that this approach is effective and performs consistently well in settings with large domain shifts.
Richong Zhang, Samuel Mensah, Nikolaos Aletras, Yongyi Mao, Xudong Liu 0001
WWW4
2022 Flexible Instance-Specific Rationalization of NLP Models
abstract
Recent research on model interpretability in natural language processing extensively uses feature scoring methods for identifying which parts of the input are the most important for a model to make a prediction (i.e. explanation or rationale). However, previous research has shown that there is no clear best scoring method across various text classification tasks while practitioners typically have to make several other ad-hoc choices regarding the length and the type of the rationale (e.g. short or long, contiguous or not). Inspired by this, we propose a simple yet effective and flexible method that allows selecting optimally for each data instance: (1) a feature scoring method; (2) the length; and (3) the type of the rationale. Our method is inspired by input erasure approaches to interpretability which assume that the most faithful rationale for a prediction should be the one with the highest difference between the model's output distribution using the full text and the text after removing the rationale as input respectively. Evaluation on four standard text classification datasets shows that our proposed method provides more faithful, comprehensive and highly sufficient explanations compared to using a fixed feature scoring method, rationale length and type. More importantly, we demonstrate that a practitioner is not required to make any ad-hoc choices in order to extract faithful rationales using our approach.
George Chrysostomou, Nikolaos Aletras
AAAI2
2022 LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
abstract
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, Nikolaos Aletras. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II, Ion Androutsopoulos, Daniel Martin Katz, Nikolaos Aletras
ACL (1)7
2022 An Empirical Study on Explanations in Out-of-Domain Settings
abstract
Recent work in Natural Language Processing has focused on developing approaches that extract faithful explanations, either via identifying the most important tokens in the input (i.e.post-hoc explanations) or by designing inherently faithful models that first select the most important tokens and then use them to predict the correct label (i.e.select-then-predict models).Currently, these approaches are largely evaluated on in-domain settings.Yet, little is known about how post-hoc explanations and inherently faithful models perform in out-ofdomain settings.In this paper, we conduct an extensive empirical study that examines: (1) the out-of-domain faithfulness of post-hoc explanations, generated by five feature attribution methods; and (2) the out-of-domain performance of two inherently faithful models over six datasets.Contrary to our expectations, results show that in many cases out-of-domain post-hoc explanation faithfulness measured by sufficiency and comprehensiveness is higher compared to in-domain.We find this misleading and suggest using a random baseline as a yardstick for evaluating post-hoc explanation faithfulness.Our findings also show that selectthen predict models demonstrate comparable predictive performance in out-of-domain settings to full-text trained models. 1
George Chrysostomou, Nikolaos Aletras
ACL (1)2
2022 Automatic Identification and Classification of Bragging in Social Media
abstract
Bragging is a speech act employed with the goal of constructing a favorable self-image through positive statements about oneself.It is widespread in daily communication and especially popular in social media, where users aim to build a positive image of their persona directly or indirectly.In this paper, we present the first large scale study of bragging in computational linguistics, building on previous research in linguistics and pragmatics.To facilitate this, we introduce a new publicly available data set of tweets annotated for bragging and their types.We empirically evaluate different transformerbased models injected with linguistic information in (a) binary bragging classification, i.e., if tweets contain bragging statements or not; and (b) multi-class bragging type prediction including not bragging.Our results show that our models can predict bragging with macro F1 up to 72.42 and 35.95 in the binary and multi-class classification tasks respectively.Finally, we present an extensive linguistic and error analysis of bragging prediction to guide future research on this topic.1
Mali Jin, Daniel Preotiuc-Pietro, A. Seza Dogruöz, Nikolaos Aletras
ACL (1)4
2022 Domain Classification-based Source-specific Term Penalization for Domain Adaptation in Hate-speech Detection
abstract
State-of-the-art approaches for hate-speech detection usually exhibit poor performance in out-of-domain settings. This occurs, typically, due to classifiers overemphasizing source-specific information that negatively impacts its domain invariance. Prior work has attempted to penalize terms related to hate-speech from manually curated lists using feature attribution methods, which quantify the importance assigned to input terms by the classifier when making a prediction. We, instead, propose a domain adaptation approach that automatically extracts and penalizes source-specific terms using a domain classifier, which learns to differentiate between domains, and feature-attribution scores for hate-speech classes, yielding consistent improvements in cross-domain evaluation.
Tulika Bose, Nikolaos Aletras, Irina Illina, Dominique Fohr
COLING2
2022 HashFormers: Towards Vocabulary-independent Pre-trained Transformers
abstract
Transformer-based pre-trained language models are vocabulary-dependent, mapping by default each token to its corresponding embedding.This one-to-one mapping results into embedding matrices that occupy a lot of memory (i.e.millions of parameters) and grow linearly with the size of the vocabulary.Previous work on on-device transformers dynamically generate token embeddings on-the-fly without embedding matrices using locality-sensitive hashing over morphological information.These embeddings are subsequently fed into transformer layers for text classification.However, these methods are not pre-trained.Inspired by this line of work, we propose HASHFORMERS, a new family of vocabulary-independent pretrained transformers that support an unlimited vocabulary (i.e.all possible tokens in a corpus) given a substantially smaller fixed-sized embedding matrix.We achieve this by first introducing computationally cheap hashing functions that bucket together individual tokens to embeddings.We also propose three variants that do not require an embedding matrix at all, further reducing the memory requirements.We empirically demonstrate that HASHFORM-ERS are more memory efficient compared to standard pre-trained transformers while achieving comparable predictive performance when fine-tuned on multiple text classification tasks.For example, our most efficient HASHFORMER variant has a negligible performance degradation (0.4% on GLUE) using only 99.1K parameters for representing the embeddings compared to 12.3-38M parameters of state-of-the-art models. 1
Huiyin Xue, Nikolaos Aletras
EMNLP2
2022 Combining Humor and Sarcasm for Improving Political Parody Detection
abstract
Xiao Ao, Danae Sanchez Villegas, Daniel Preotiuc-Pietro, Nikolaos Aletras. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xiao Ao, Danae Sanchez Villegas, Daniel Preotiuc-Pietro, Nikolaos Aletras
NAACL-HLT4
2022 Towards Suicide Ideation Detection Through Online Conversational Context
abstract
Social media enable users to share their feelings and emotional struggles. They also offer an opportunity to provide community support to suicidal users. Recent studies on suicide risk assessment have explored the user's historic timeline and information from their social network to analyze their emotional state. However, such methods often require a large amount of user-centric data. A less intrusive alternative is to only use conversation trees arising from online community responses. Modeling such online conversations between the community and a person in distress is an important context for understanding that person's mental state. However, it is not trivial to model the vast number of conversation trees on social media, since each comment has a diverse influence on a user in distress. Typically, a handful of comments/posts receive a significantly high number of replies, which results in scale-free dynamics in the conversation tree. Moreover, psychological studies suggested that it is important to capture the fine-grained temporal irregularities in the release of vast volumes of comments, since suicidal users react quickly to online community support. Building on these limitations and psychological studies, we propose HCN, a Hyperbolic Conversation Network, which is a less user-intrusive method for suicide ideation detection. HCN leverages the hyperbolic space to represent the scale-free dynamics of online conversations. Through extensive quantitative, qualitative, and ablative experiments on real-world Twitter data, we find that HCN outperforms state-of-the art methods, while using 98% less user-specific data, and while maintaining a 74% lower carbon footprint and a 94% smaller model size. We also find that the comments within the first half an hour are most important to identify at-risk users.
Ramit Sawhney, Shivam Agarwal, Atula Tejaswi Neerkaje, Nikolaos Aletras, Preslav Nakov, Lucie Flek
SIGIR4
2022 Node-Feature Convolution for Graph Convolutional Networks
abstract
Graph convolutional network (GCN) is an effective neural network model for graph representation learning. However, standard GCN suffers from three main limitations: (1) most real-world graphs have no regular connectivity and node degrees can range from one to hundreds or thousands, (2) neighboring nodes are aggregated with fixed weights, and (3) node features within a node feature vector are considered equally important. Several extensions have been proposed to tackle the limitations respectively. This paper focuses on tackling all the proposed limitations. Specifically, we propose a new node-feature convolutional (NFC) layer for GCN. The NFC layer first constructs a feature map using features selected and ordered from a fixed number of neighbors. It then performs a convolution operation on this feature map to learn the node representation. In this way, we can learn the usefulness of both individual nodes and individual features from a fixed-size neighborhood. Experiments on three benchmark datasets show that NFC-GCN consistently outperforms state-of-the-art methods in node classification.
Li Zhang 0131, Heda Song, Nikolaos Aletras, Haiping Lu
Pattern Recognit.3
2021 Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification
abstract
George Chrysostomou, Nikolaos Aletras. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
George Chrysostomou, Nikolaos Aletras
ACL/IJCNLP (1)2
2021 Enjoy the Salience: Towards Better Transformer-based Faithful Explanations with Word Salience
abstract
Pretrained transformer-based models such as BERT have demonstrated state-of-the-art predictive performance when adapted into a range of natural language processing tasks.An open problem is how to improve the faithfulness of explanations (rationales) for the predictions of these models.In this paper, we hypothesize that salient information extracted a priori from the training data can complement the task-specific information learned by the model during fine-tuning on a downstream task.In this way, we aim to help BERT not to forget assigning importance to informative input tokens when making predictions by proposing SALOSS; an auxiliary loss function for guiding the multi-head attention mechanism during training to be close to salient information extracted a priori using TextRank.Experiments for explanation faithfulness across five datasets, show that models trained with SA-LOSS consistently provide more faithful explanations across four different feature attribution methods compared to vanilla BERT.Using the rationales extracted from vanilla BERT and SALOSS models to train inherently faithful classifiers, we further show that the latter result in higher predictive performance in downstream tasks. 1
George Chrysostomou, Nikolaos Aletras
EMNLP (1)2
2021 Active Learning by Acquiring Contrastive Examples
abstract
Common acquisition functions for active learning use either uncertainty or diversity sampling, aiming to select difficult and diverse data points from the pool of unlabeled data, respectively.In this work, leveraging the best of both worlds, we propose an acquisition function that opts for selecting contrastive examples, i.e. data points that are similar in the model feature space and yet the model outputs maximally different predictive likelihoods.We compare our approach, CAL (Contrastive Active Learning), with a diverse set of acquisition functions in four natural language understanding tasks and seven datasets.Our experiments show that CAL performs consistently better or equal than the best performing baseline across all tasks, on both in-domain and out-of-domain data.We also conduct an extensive ablation study of our method and we further analyze all actively acquired datasets showing that CAL achieves a better trade-off between uncertainty and diversity compared to other strategies.
Aikaterini Margatina, Giorgos Vernikos, Loïc Barrault, Nikolaos Aletras
EMNLP (1)4
2021 An Empirical Study on Leveraging Position Embeddings for Target-oriented Opinion Words Extraction
abstract
Target-oriented opinion words extraction (TOWE) (Fan et al., 2019b) is a new subtask of target-oriented sentiment analysis that aims to extract opinion words for a given aspect in text.Current state-of-the-art methods leverage position embeddings to capture the relative position of a word to the target.However, the performance of these methods depends on the ability to incorporate this information into word representations.In this paper, we explore a variety of text encoders based on pretrained word embeddings or language models that leverage part-of-speech and position embeddings, aiming to examine the actual contribution of each component in TOWE.We also adapt a graph convolutional network (GCN) to enhance word representations by incorporating syntactic information.Our experimental results demonstrate that BiLSTM-based models can effectively encode position information into word representations while using a GCN only achieves marginal gains.Interestingly, our simple methods outperform several state-of-the-art complex neural structures.
Samuel Mensah, Nikolaos Aletras
EMNLP (1)3
2021 Point-of-Interest Type Prediction using Text and Images
abstract
Point-of-interest (POI) type prediction is the task of inferring the type of a place from where a social media post was shared.Inferring a POI's type is useful for studies in computational social science including sociolinguistics, geosemiotics, and cultural geography, and has applications in geosocial networking technologies such as recommendation and visualization systems.Prior efforts in POI type prediction focus solely on text, without taking visual information into account.However in reality, the variety of modalities, as well as their semiotic relationships with one another, shape communication and interactions in social media.This paper presents a study on POI type prediction using multimodal information from text and images available at posting time.For that purpose, we enrich a currently available data set for POI type prediction with the images that accompany the text messages.Our proposed method extracts relevant information from each modality to effectively capture interactions between text and image achieving a macro F1 of 47.21 across eight categories significantly outperforming the state-of-the-art method for POI type prediction based on textonly methods.Finally, we provide a detailed analysis to shed light on cross-modal interactions and the limitations of our best performing model. 1
Danae Sanchez Villegas, Nikolaos Aletras
EMNLP (1)2
2021 Frustratingly Simple Pretraining Alternatives to Masked Language Modeling
abstract
Masked language modeling (MLM), a selfsupervised pretraining objective, is widely used in natural language processing for learning text representations.MLM trains a model to predict a random sample of input tokens that have been replaced by a [MASK] placeholder in a multi-class setting over the entire vocabulary.When pretraining, it is common to use alongside MLM other auxiliary objectives on the token or sequence level to improve downstream performance (e.g. next sentence prediction).However, no previous work so far has attempted in examining whether other simpler linguistically intuitive or not objectives can be used standalone as main pretraining objectives.In this paper, we explore five simple pretraining objectives based on token-level classification tasks as replacements of MLM.Empirical results on GLUE and SQUAD show that our proposed methods achieve comparable or better performance to MLM using a BERT-BASE architecture.We further validate our methods using smaller models, showing that pretraining a model with 41% of the BERT-BASE's parameters, BERT-MEDIUM results in only a 1% drop in GLUE scores with our best objective.1
Atsuki Yamaguchi, George Chrysostomou, Aikaterini Margatina, Nikolaos Aletras
EMNLP (1)4
2021 Paragraph-level Rationale Extraction through Regularization: A case study on European Court of Human Rights Cases
abstract
Ilias Chalkidis, Manos Fergadiotis, Dimitrios Tsarapatsanis, Nikolaos Aletras, Ion Androutsopoulos, Prodromos Malakasiotis. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ilias Chalkidis, Manos Fergadiotis, Dimitrios Tsarapatsanis, Nikolaos Aletras, Ion Androutsopoulos, Prodromos Malakasiotis
NAACL-HLT4
2021 Modeling the Severity of Complaints in Social Media
abstract
The speech act of complaining is used by humans to communicate a negative mismatch between reality and expectations as a reaction to an unfavorable situation.Linguistic theory of pragmatics categorizes complaints into various severity levels based on the face-threat that the complainer is willing to undertake.This is particularly useful for understanding the intent of complainers and how humans develop suitable apology strategies.In this paper, we study the severity level of complaints for the first time in computational linguistics.To facilitate this, we enrich a publicly available data set of complaints with four severity categories and train different transformer-based networks combined with linguistic information achieving 55.7 macro F1.We also jointly model binary complaint classification and complaint severity in a multi-task setting achieving new state-of-the-art results on binary complaint detection reaching up to 88.2 macro F1.Finally, we present a qualitative analysis of the behavior of our models in predicting complaint severity levels.1,2
Mali Jin, Nikolaos Aletras
NAACL-HLT2
2020 Analyzing Political Parody in Social Media
abstract
Parody is a figurative device used to imitate an entity for comedic or critical purposes and represents a widespread phenomenon in social media through many popular parody accounts.In this paper, we present the first computational study of parody.We introduce a new publicly available data set of tweets from real politicians and their corresponding parody accounts.We run a battery of supervised machine learning models for automatically detecting parody tweets with an emphasis on robustness by testing on tweets from accounts unseen in training, across different genders and across countries.Our results show that political parody tweets can be predicted with an accuracy up to 90%.Finally, we identify the markers of parody through a linguistic analysis.Beyond research in linguistics and political communication, accurately and automatically detecting parody is important to improving fact checking for journalists and analytics such as sentiment analysis through filtering out parodical utterances. 1
Antonis Maronikolakis, Danae Sanchez Villegas, Daniel Preotiuc-Pietro, Nikolaos Aletras
ACL4
2020 LegalOps: A Summarization Corpus of Legal Opinions
abstract
We present a new, large-scale corpus for training and evaluating text summarization systems on legal opinions, called LegalOps. The corpus includes~14K opinions together with their summaries from U.S. Federal Courts, e.g. the Supreme Court and Federal Appeals Courts. The aim of this paper is to provide a novel data source of sufficient variety that it will advance theoretical work on modeling the particular patterns within legal discourse, but also of sufficient size that it will provide a new challenging testbed for state-of-the-art automatic summarization models.
Andrew Gargett, Rob Firth, Nikolaos Aletras
IEEE BigData3
2020 Complaint Identification in Social Media with Transformer Networks
abstract
Complaining is a speech act extensively used by humans to communicate a negative inconsistency between reality and expectations.Previous work on automatically identifying complaints in social media has focused on using feature-based and task-specific neural network models.Adapting state-of-the-art pre-trained neural language models and their combinations with other linguistic information from topics or sentiment for complaint prediction has yet to be explored.In this paper, we evaluate a battery of neural models underpinned by transformer networks which we subsequently combine with linguistic information.Experiments on a publicly available data set of complaints demonstrate that our models outperform previous state-of-the-art methods by a large margin achieving a macro F1 up to 87.
Mali Jin, Nikolaos Aletras
COLING2
2020 Quality In, Quality Out: Learning from Actual Mistakes
abstract
Approaches to Quality Estimation (QE) of machine translation have shown promising results at predicting quality scores for translated sentences. However, QE models are often trained on noisy approximations of quality annotations derived from the proportion of post-edited words in translated sentences instead of direct human annotations of translation errors. The latter is a more reliable ground-truth but more expensive to obtain. In this paper, we present the first attempt to model the task of predicting the proportion of actual translation errors in a sentence while minimising the need for direct human annotation. For that purpose, we use transfer-learning to leverage large scale noisy annotations and small sets of high-fidelity human annotated translation errors to train QE models. Experiments on four language pairs and translations obtained by statistical and neural models show consistent gains over strong baselines.
Frédéric Blain, Nikolaos Aletras, Lucia Specia
EAMT2
2020 An Empirical Study on Large-Scale Multi-Label Text Classification Including Few and Zero-Shot Labels
abstract
Ilias Chalkidis, Manos Fergadiotis, Sotiris Kotitsas, Prodromos Malakasiotis, Nikolaos Aletras, Ion Androutsopoulos. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Ilias Chalkidis, Manos Fergadiotis, Sotiris Kotitsas, Prodromos Malakasiotis, Nikolaos Aletras, Ion Androutsopoulos
EMNLP (1)5
2020 Automatic Generation of Topic Labels
abstract
Topic modelling is a popular unsupervised method for identifying the underlying themes in document collections that has many applications in information retrieval. A topic is usually represented by a list of terms ranked by their probability but, since these can be difficult to interpret, various approaches have been developed to assign descriptive labels to topics. Previous work on the automatic assignment of labels to topics has relied on a two-stage approach: (1) candidate labels are retrieved from a large pool (e.g. Wikipedia article titles); and then (2) re-ranked based on their semantic similarity to the topic terms. However, these extractive approaches can only assign candidate labels from a restricted set that may not include any suitable ones. This paper proposes using a sequence-to-sequence neural-based approach to generate labels that does not suffer from this limitation. The model is trained over a new large synthetic dataset created using distant supervision. The method is evaluated by comparing the labels it generates to ones rated by humans.
Areej Alokaili, Nikolaos Aletras, Mark Stevenson 0001
SIGIR2
2020 Unsupervised Quality Estimation for Neural Machine Translation
abstract
Quality Estimation (QE) is an important component in making Machine Translation (MT) useful in real-world applications, as it is aimed to inform the user on the quality of the MT output at test time. Existing approaches require large amounts of expert annotated data, computation, and time for training. As an alternative, we devise an unsupervised approach to QE where no training or access to additional resources besides the MT system itself is required. Different from most of the current work that treats the MT system as a black box, we explore useful information that can be extracted from the MT system as a by-product of translation. By utilizing methods for uncertainty quantification, we achieve very good correlation with human judgments of quality, rivaling state-of-the-art supervised QE models. To evaluate our approach we collect the first dataset that enables work on both black-box and glass-box approaches to QE.
Marina Fomicheva, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, Lucia Specia
Trans. Assoc. Comput. Linguistics7
2019 Neural Legal Judgment Prediction in English
abstract
Legal judgment prediction is the task of automatically predicting the outcome of a court case, given a text describing the case's facts.Previous work on using neural models for this task has focused on Chinese; only featurebased models (e.g., using bags of words and topics) have been considered in English.We release a new English legal judgment prediction dataset, containing cases from the European Court of Human Rights.We evaluate a broad variety of neural models on the new dataset, establishing strong baselines that surpass previous feature-based models in three tasks: (1) binary violation classification; (2) multi-label classification; (3) case importance prediction.We also explore if models are biased towards demographic information via data anonymization.As a side-product, we propose a hierarchical version of BERT, which bypasses BERT's length limitation.
Ilias Chalkidis, Ion Androutsopoulos, Nikolaos Aletras
ACL (1)3
2019 Automatically Identifying Complaints in Social Media
abstract
Complaining is a basic speech act regularly used in human and computer mediated communication to express a negative mismatch between reality and expectations in a particular situation.Automatically identifying complaints in social media is of utmost importance for organizations or brands to improve the customer experience or in developing dialogue systems for handling and responding to complaints.In this paper, we introduce the first systematic analysis of complaints in computational linguistics.We collect a new annotated data set of written complaints expressed in English on Twitter. 1 We present an extensive linguistic analysis of complaining as a speech act in social media and train strong feature-based and neural models of complaints across nine domains achieving a predictive performance of up to 79 F1 using distant supervision.
Daniel Preotiuc-Pietro, Mihaela Gaman, Nikolaos Aletras
ACL (1)3
2018 Nowcasting the Stance of Social Media Users in a Sudden Vote: The Case of the Greek Referendum
abstract
Modelling user voting intention in social media is an important research area, with applications in analysing electorate behaviour, online political campaigning and advertising. Previous approaches mainly focus on predicting national general elections, which are regularly scheduled and where data of past results and opinion polls are available. However, there is no evidence of how such models would perform during a sudden vote under time-constrained circumstances. That poses a more challenging task compared to traditional elections, due to its spontaneous nature. In this paper, we focus on the 2015 Greek bailout referendum, aiming to nowcast on a daily basis the voting intention of 2,197 Twitter users. We propose a semi-supervised multiple convolution kernel learning approach, leveraging temporally sensitive text and network information. Our evaluation under a real-time simulation framework demonstrates the effectiveness and robustness of our approach against competitive baselines, achieving a significant 20% increase in F-score compared to solely text-based models.
Adam Tsakalidis, Nikolaos Aletras, Alexandra I. Cristea, Maria Liakata
CIKM2
2017 Labeling Topics with Images Using a Neural Network
Nikolaos Aletras, Arpit Mittal
ECIR1
2017 Evaluating topic representations for exploring document collections
abstract
Topic models have been shown to be a useful way of representing the content of large document collections, for example, via visualization interfaces (topic browsers). These systems enable users to explore collections by way of latent topics. A standard way to represent a topic is using a term list; that is the top‐n words with highest conditional probability within the topic. Other topic representations such as textual and image labels also have been proposed. However, there has been no comparison of these alternative representations. In this article, we compare 3 different topic representations in a document retrieval task. Participants were asked to retrieve relevant documents based on predefined queries within a fixed time limit, presenting topics in one of the following modalities: (a) lists of terms, (b) textual phrase labels, and (c) image labels. Results show that textual labels are easier for users to interpret than are term lists and image labels. Moreover, the precision of retrieved documents for textual and image labels is comparable to the precision achieved by representing topics using term lists, demonstrating that labeling methods are an effective alternative topic representation.
Nikolaos Aletras, Timothy Baldwin, Jey Han Lau, Mark Stevenson 0001
J. Assoc. Inf. Sci. Technol.1
2016 Inferring the Socioeconomic Status of Social Media Users Based on Behaviour and Language
Vasileios Lampos, Nikolaos Aletras, Jens K. Geyti, Bin Zou 0006, Ingemar J. Cox
ECIR2
2016 Why are these similar? Investigating item similarity types in a large digital library
abstract
We introduce a new problem, identifying the type of relation that holds between a pair of similar items in a digital library. Being able to provide a reason why items are similar has applications in recommendation, personalization, and search. We investigate the problem within the context of Europeana, a large digital library containing items related to cultural heritage. A range of types of similarity in this collection were identified. A set of 1,500 pairs of items from the collection were annotated using crowdsourcing. A high intertagger agreement (average 71.5 Pearson correlation) was obtained and demonstrates that the task is well defined. We also present several approaches to automatically identifying the type of similarity. The best system applies linear regression and achieves a mean Pearson correlation of 71.3, close to human performance. The problem formulation and data set described here were used in a public evaluation exercise, the *SEM shared task on Semantic Textual Similarity. The task attracted the participation of 6 teams, who submitted 14 system runs. All annotations, evaluation scripts, and system runs are freely available.
Aitor Gonzalez-Agirre, German Rigau, Eneko Agirre, Nikolaos Aletras, Mark Stevenson 0001
J. Assoc. Inf. Sci. Technol.4
2015 An analysis of the user occupational class through Twitter content
abstract
Daniel Preoţiuc-Pietro, Vasileios Lampos, Nikolaos Aletras. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Daniel Preotiuc-Pietro, Vasileios Lampos, Nikolaos Aletras
ACL (1)3
2015 TM 2015 - Topic Models: Post-Processing and Applications Workshop
abstract
The main objective of the workshop is to bring together researchers who are interested in applications of topic models and improving their output. Our goal is to create a broad platform for researchers to share ideas that could improve the usability and interpretation of topic models. We expect this will promote topic model applications in other research areas, making their use more effective.
Nikolaos Aletras, Jey Han Lau, Timothy Baldwin, Mark Stevenson 0001
CIKM1
2014 Measuring the Similarity between Automatically Generated Topics
abstract
Previous approaches to the problem of measuring similarity between automati-cally generated topics have been based on comparison of the topics ’ word probability distributions. This paper presents alterna-tive approaches, including ones based on distributional semantics and knowledge-based measures, evaluated by compari-son with human judgements. The best performing methods provide reliable esti-mates of topic similarity comparable with human performance and should be used in preference to the word probability distri-bution measures used previously. 1
Nikolaos Aletras
EACL1
2014 Predicting and Characterising User Impact on Twitter
abstract
The open structure of online social networks and their uncurated nature give rise to problems of user credibility and influence.In this paper, we address the task of predicting the impact of Twitter users based only on features under their direct control, such as usage statistics and the text posted in their tweets.We approach the problem as regression and apply linear as well as nonlinear learning methods to predict a user impact score, estimated by combining the numbers of the user's followers, followees and listings.The experimental results point out that a strong prediction performance is achieved, especially for models based on the Gaussian Processes framework.Hence, we can interpret various modelling components, transforming them into indirect 'suggestions' for impact boosting.
Vasileios Lampos, Nikolaos Aletras, Daniel Preotiuc-Pietro, Trevor Cohn
EACL2
2013 Representing Topics Using Images
Nikolaos Aletras
HLT-NAACL1
2012 PATHS - Exploring Digital Cultural Heritage Spaces
Mark M. Hall, Eneko Agirre, Nikolaos Aletras, Runar Bergheim, Konstantinos Chandrinos, Paul D. Clough, Samuel Fernando, Kate Fernie, Paula Goodale, Jillian Griffiths, Oier Lopez de Lacalle, Andrea de Polo, Aitor Soroa, Mark Stevenson 0001
TPDL3