EDBT 2026 Demo / reviewers in the wild / expert
John Pavlopoulos
dblp:09/269
· DBLP profile ↗
36ranked-venue papers
17as first author
30since 2021 · last 2026
0000-0001-9188-7425ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 15 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Diachronic Representations of Ancient Greek Letterforms
John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler |
ICDAR (3) | 1 |
| 2025 | DiaShift: An Explainable System for Temporal Diagnostic Shift Detection in Clinical NotesabstractClinical diagnoses often evolve (shift) over the course of a patient's hospital stay. These shifts are crucial for the physician yet they are hidden in free-text notes. Unsurprisingly, identifying them is challenging due to language variation, physician perspectives, and documentation practices. This work proposes a three-stage pipeline to detect and categorize diagnostic shifts (DSs) between pairs of physician-authored notes. We introduce a DS schema that captures changes in diagnostic reasoning across sequential notes. Our method first aligns semantically similar text segments, then filters potential DSs by identifying inconsistent segment pairs using Natural Language Inference, and finally classifies the nature of the shift using a large language model (LLM). Using real-world data from the MIMIC-III dataset, we applied our system to 400 randomly sampled sequential physician-authored note pairs — each from the same Intensive Care Unit (ICU) stay of a patient — spanning 137 ICU stays and 115 physicians. Our system detected shift in 70% of note pairs, and in 78% of those cases, the shift was a diagnostic refinement. Manual evaluation showed a Precision of 0.7 for the most critical shift type (Overturn). Notably, the justifications provided by our system were found to be accurate, supporting explainability. Ablation studies highlight that semantic alignment and NLIbased shift filtering are critical for enabling accurate and explainable LLM-based shift classification. This study demonstrates the feasibility of LLM-based DS analysis and motivates future work on broader validation and deployment. Juli Bakagianni, Kalliopi Dalakleidi, Konstantinos Stamatis, John Pavlopoulos |
BIBE | 4 |
| 2025 | From Pen to Prediction: Handwriting-Based Alzheimer's DetectionabstractEarly diagnosis of Alzheimer's disease (AD) is essential for timely intervention and effective care. This paper examines handwriting analysis as an accessible and non-invasive way of early detection by focusing on the comparison of raw handwriting images and tabular features. Two data sets were considered, the DARWIN handwriting dataset, which contains raw images and tabular data from pen movement, and the Alzheimer's disease dataset (ADD), which presents the patient's history and cognitive assessments in tabular format. A range of classification methods including machine learning (Random Forest, SVM, and XGBoost) was tested on tabular data from the two datasets, a deep learning Swin Transformer for image classification, and a multimodal approach that integrated both. Random Forest outperformed other models on DARWIN tabular data$(83.03 \% \pm 1.18)$, while XGBoost was the best on ADD$(83.53 \% \pm 3.44)$. The Swin Transformer also performed consistently on handwriting images$(80.02 \% \pm 0.87)$, capturing features associated with stroke tremors and fluency, as well as other visual aspects of the dataset. A late fusion model incorporating both modalities achieved the highest overall accuracy of$89.15 \% \pm 1.73$, showing that the image and the tabular features produce a complementary diagnostic value. These results indicate that handwriting includes fine neuromotor features related to early AD that can surpass clinical conventional data. We also present ablation studies on task order with respect to image training and end-to-end multimodal learning. These findings provide further evidence of the benefits of modular fusion in situations where data is restricted. Handwriting samples have the potential to become a useable and scalable resource in AD screening because of their low cost, ease of collection, and acquisition logistics, which even accommodate home-based settings. Maria Boumpi, Kalliopi Dalakleidi, John Pavlopoulos |
BIBE | 3 |
| 2025 | Evaluation and Facilitation of Online Discussions in the LLM Era: A SurveyabstractKaterina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé, Danai Myrtzani, Theodoros Evgeniou, Ion Androutsopoulos, John Pavlopoulos. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Katerina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé, Danai Myrtzani, Theodoros Evgeniou, Ion Androutsopoulos, John Pavlopoulos |
EMNLP | 8 |
| 2025 | Mind the gap: from plausible to valid self-explanations in large language modelsabstractAbstract This paper investigates the reliability of explanations generated by large language models (LLMs) when prompted to explain their previous output. We evaluate two kinds of such self-explanations ( SE )—extractive and counterfactual—using state-of-the-art LLMs (1B to 70B parameters) on three different classification tasks (both objective and subjective). In line with Agarwal et al. (Faithfulness versus plausibility: On the (Un)reliability of explanations from large language models. 2024. https://doi.org/10.48550/arXiv.2402.04614 ), our findings indicate a gap between perceived and actual model reasoning: while SE largely correlate with human judgment (i.e. are plausible ), they do not fully and accurately follow the model’s decision process (i.e. are not faithful ). Additionally, we show that counterfactual SE are not even necessarily valid in the sense of actually changing the LLM’s prediction. Our results suggest that extractive SE provide the LLM’s “guess” at an explanation based on training data. Conversely, counterfactual SE can help understand the LLM’s reasoning: We show that the issue of validity can be resolved by sampling counterfactual candidates at high temperature—followed by a validity check—and introducing a formula to estimate the number of tries needed to generate valid explanations. This simple method produces plausible and valid explanations that offer a 16 times faster alternative to SHAP on average in our experiments. Korbinian Randl, John Pavlopoulos, Aron Henriksson, Tony Lindgren |
Mach. Learn. | 2 |
| 2024 | HoLM: Analyzing the Linguistic Unexpectedness in Homeric PoetryabstractThe authorship of the Homeric poems has been a matter of debate for centuries. Computational approaches such as language modeling exist that can aid experts in making crucial headway. We observe, however, that such work has, thus far, only been carried out at the level of lengthier excerpts, but not individual verses, the level at which most suspected interpolations occur. We address this weakness by presenting a corpus of Homeric verses, each complemented with a score quantifying linguistic unexpectedness based on Perplexity. We assess the nature of these scores by exploring their correlation with named entities, the frequency of character n-grams, and (inverse) word frequency, revealing robust correlations with the latter two. This apparent bias can be partly overcome by simply dividing scores for unexpectedness by the maximum term frequency per verse. John Pavlopoulos, Ryan Sandell, Maria Konstantinidou, Chiara Bozzone |
LREC/COLING | 1 |
| 2024 | Deciphering Emotional Landscapes in the Iliad: A Novel French-Annotated Dataset for Emotion RecognitionabstractOne of the most significant pieces of ancient Greek literature, the Iliad, is part of humanity’s collective cultural heritage. This work aims to provide the scientific community with an emotion-labeled dataset for classical literature and Western mythology in particular. To model the emotions of the poem, we use a multi-variate time series. We also evaluated the dataset by means of two methods. We compare the manual classification against a dictionary-based benchmark as well as employ a state-of-the-art deep learning masked language model that has been tuned using our data. Both evaluations return encouraging results (MSE and MAE Macro Avg 0.101 and 0.188 respectively) and highlight some interesting phenomena. Davide Picca, John Pavlopoulos |
LREC/COLING | 2 |
| 2024 | Still All Greeklish to Me: Greeklish to Greek TransliterationabstractModern Greek is normally written in the Greek alphabet. In informal online messages, however, Greek is often written using characters available on Latin-character keyboards, a form known as Greeklish. Originally used to bypass the lack of support for the Greek alphabet in older computers, Greeklish is now also used to avoid switching languages on multilingual keyboards, hide spelling mistakes, or as a form of slang. There is no consensus mapping, hence the same Greek word can be written in numerous different ways in Greeklish. Even native Greek speakers may struggle to understand (or be annoyed by) Greeklish, which requires paying careful attention to context to decipher. Greeklish may also be a problem for NLP models trained on Greek datasets written in the Greek alphabet. Experimenting with a range of statistical and deep learning models on both artificial and real-life Greeklish data, we find that: (i) prompting large language models (e.g., GPT-4) performs impressively well with few- or even zero-shot training, outperforming several fine-tuned encoder-decoder models; however (ii) a twenty years old statistical Greeklish transliteration model is still very competitive; and (iii) the problem is still far from having been solved; (iv) nevertheless, downstream Greek NLP systems that need to cope with Greeklish, such as moderation classifiers, can benefit significantly even with the current non-perfect transliteration systems. We make all our code, models, and data available and suggest future improvements, based on an analysis of our experimental results. Anastasios Toumazatos, John Pavlopoulos, Ion Androutsopoulos, Stavros Vassos |
LREC/COLING | 2 |
| 2024 | Revisiting Silhouette Aggregation
John Pavlopoulos, Georgios Vardakas, Aristidis Likas |
DS (1) | 1 |
| 2024 | Evaluating the Reliability of Self-explanations in Large Language Models
Korbinian Randl, John Pavlopoulos, Aron Henriksson, Tony Lindgren |
DS (1) | 2 |
| 2024 | Polarized Opinion Detection Improves the Detection of Toxic LanguageabstractDistance from unimodality (DFU) has been found to correlate well with human judgment for the assessment of polarized opinions.However, its un-normalized nature makes it less intuitive and somewhat difficult to exploit in machine learning (e.g., as a supervised signal).In this work a normalized version of this measure, called nDFU, is proposed that leads to better assessment of the degree of polarization.Then, we propose a methodology for K-class text classification, based on nDFU, that exploits polarized texts in the dataset.Such polarized instances are assigned to a separate K+1 class, so that a K+1-class classifier is trained.An empirical analysis on three datasets for abusive language detection, shows that nDFU can be used to model polarized annotations and prevent them from harming the classification performance.Finally, we further exploit nDFU to specify conditions that could explain polarization given a dimension and present text examples that polarized the annotators when the dimension was gender and race.Our code is available at https://github.com/ipavlopoulos/ndfu. John Pavlopoulos, Aristidis Likas |
EACL (1) | 1 |
| 2024 | Towards a Greek Proverb Atlas: Computational Spatial Exploration and Attribution of Greek ProverbsabstractProverbs carry wisdom transferred orally from generation to generation.Based on the place they were recorded, this study introduces a publicly-available and machine-actionable dataset of more than one hundred thousand Greek proverb variants.By quantifying the spatial distribution of proverbs, we show that the most widespread proverbs come from the mainland while the least widespread proverbs come primarily from the islands.By focusing on the least dispersed proverbs, we present the most frequent tokens per location and undertake a benchmark in geographical attribution, using text classification and regression (text geocoding).Our results show that this is a challenging task for which specific locations can be attributed more successfully compared to others.The potential of our resource and benchmark is showcased by two novel applications.First, we extracted terms moving the regression prediction toward the four cardinal directions.Second, we leveraged conformal prediction to attribute 3,676 unregistered proverbs with statistically rigorous predictions of locations each of these proverbs was possibly registered in. John Pavlopoulos, Panagiotis Louridas, Panagiotis Filos |
EMNLP | 1 |
| 2024 | Fraud detection with natural language processingabstractAbstract Automated fraud detection can assist organisations to safeguard user accounts, a task that is very challenging due to the great sparsity of known fraud transactions. Many approaches in the literature focus on credit card fraud and ignore the growing field of online banking. However, there is a lack of publicly available data for both. The lack of publicly available data hinders the progress of the field and limits the investigation of potential solutions. With this work, we: (a) introduce FraudNLP , the first anonymised, publicly available dataset for online fraud detection, (b) benchmark machine and deep learning methods with multiple evaluation measures, (c) argue that online actions do follow rules similar to natural language and hence can be approached successfully by natural language processing methods. Petros Boulieris, John Pavlopoulos, Alexandros Xenos, Vasilis Vassalos |
Mach. Learn. | 2 |
| 2024 | Explainable dating of greek papyri imagesabstractAbstract Greek literary papyri, which are unique witnesses of antique literature, do not usually bear a date. They are thus currently dated based on palaeographical methods, with broad approximations which often span more than a century. We created a dataset of 242 images of papyri written in “bookhand” scripts whose date can be securely assigned, and we used it to train algorithms for the task of dating, showing its challenging nature. To address data scarcity, we extended our dataset by segmenting each image into its respective text lines. By using the line-based version of our dataset, we trained a Convolutional Neural Network, equipped with a fragmentation-based augmentation strategy, and we achieved a mean absolute error of 54 years. The results improve further when the task is cast as a multi-class classification problem, predicting the century. Using our network, we computed precise date estimations for papyri whose date is disputed or vaguely defined, employing explainability to understand dating-driving features. John Pavlopoulos, Maria Konstantinidou, Elpida Perdiki, Isabelle Marthot-Santaniello, Holger Essler, Georgios Vardakas, Aristidis Likas |
Mach. Learn. | 1 |
| 2024 | Automotive fault nowcasting with machine learning and natural language processingabstractAbstract Automated fault diagnosis can facilitate diagnostics assistance, speedier troubleshooting, and better-organised logistics. Currently, most AI-based prognostics and health management in the automotive industry ignore textual descriptions of the experienced problems or symptoms. With this study, however, we propose an ML-assisted workflow for automotive fault nowcasting that improves on current industry standards. We show that a multilingual pre-trained Transformer model can effectively classify the textual symptom claims from a large company with vehicle fleets, despite the task’s challenging nature due to the 38 languages and 1357 classes involved. Overall, we report an accuracy of more than 80% for high-frequency classes and above 60% for classes with reasonable minimum support, bringing novel evidence that automotive troubleshooting management can benefit from multilingual symptom text classification. John Pavlopoulos, Alv Romell, Jacob Curman, Olof Steinert, Tony Lindgren, Markus Borg, Korbinian Randl |
Mach. Learn. | 1 |
| 2023 | Dating Greek Papyri with Text RegressionabstractJohn Pavlopoulos, Maria Konstantinidou, Isabelle Marthot-Santaniello, Holger Essler, Asimina Paparigopoulou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. John Pavlopoulos, Maria Konstantinidou, Isabelle Marthot-Santaniello, Holger Essler, Asimina Paparigopoulou |
ACL (1) | 1 |
| 2023 | Leveraging the Spatiotemporal Analysis of Meisho-e Landscapes
Konstantina Liagkou, John Pavlopoulos, Ewa Machotka |
DS | 2 |
| 2023 | Explaining the Chronological Attribution of Greek Papyri Images
John Pavlopoulos, Maria Konstantinidou, Georgios Vardakas, Isabelle Marthot-Santaniello, Elpida Perdiki, Dimitris Koutsianos, Aristidis Likas, Holger Essler |
DS | 1 |
| 2023 | Machine Learning for Ancient Languages: A SurveyabstractAbstract Ancient languages preserve the cultures and histories of the past. However, their study is fraught with difficulties, and experts must tackle a range of challenging text-based tasks, from deciphering lost languages to restoring damaged inscriptions, to determining the authorship of works of literature. Technological aids have long supported the study of ancient texts, but in recent years advances in artificial intelligence and machine learning have enabled analyses on a scale and in a detail that are reshaping the field of humanities, similarly to how microscopes and telescopes have contributed to the realm of science. This article aims to provide a comprehensive survey of published research using machine learning for the study of ancient texts written in any language, script, and medium, spanning over three and a half millennia of civilizations around the ancient world. To analyze the relevant literature, we introduce a taxonomy of tasks inspired by the steps involved in the study of ancient documents: digitization, restoration, attribution, linguistic analysis, textual criticism, translation, and decipherment. This work offers three major contributions: first, mapping the interdisciplinary field carved out by the synergy between the humanities and machine learning; second, highlighting how active collaboration between specialists from both fields is key to producing impactful and compelling scholarship; third, highlighting promising directions for future work in this field. Thus, this work promotes and supports the continued collaborative impetus between the humanities and machine learning. Thea Sommerschield, Yannis M. Assael, John Pavlopoulos, Vanessa Stefanak, Andrew W. Senior, Chris Dyer, John Bodel, Jonathan Prag, Ion Androutsopoulos, Nando de Freitas |
Comput. Linguistics | 3 |
| 2022 | From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil TransferabstractJohn Pavlopoulos, Leo Laugier, Alexandros Xenos, Jeffrey Sorensen, Ion Androutsopoulos. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. John Pavlopoulos, Léo Laugier, Alexandros Xenos, Jeffrey S. Sorensen, Ion Androutsopoulos |
ACL (1) | 1 |
| 2022 | Enriching Grammatical Error Correction Resources for Modern GreekabstractGrammatical Error Correction (GEC), a task of Natural Language Processing (NLP), is challenging for underepresented languages. This issue is most prominent in languages other than English. This paper addresses the issue of data and system sparsity for GEC purposes in the modern Greek Language. Following the most popular current approaches in GEC, we develop and test an MT5 multilingual text-to-text transformer for Greek. To our knowledge this the first attempt to create a fully-fledged GEC model for Greek. Our evaluation shows that our system reaches up to 52.63% F0.5 score on part of the Greek Native Corpus (GNC), which is 16% below the winning system of the BEA-19 shared task on English GEC. In addition, we provide an extended version of the Greek Learner Corpus (GLC), on which our model reaches up to 22.76% F0.5. Previous versions did not include corrections with the annotations which hindered the potential development of efficient GEC systems. For that reason we provide a new set of corrections. This new dataset facilitates an exploration of the generalisation abilities and robustness of our system, given that the assessment is conducted on learner data while the training on native data. Katerina Korre, John Pavlopoulos |
LREC | 2 |
| 2022 | A Study of Distant Viewing of ukiyo-e printsabstractThis paper contributes to studying relationships between Japanese topography and places featured in early modern landscape prints, so-called ukiyo-e or ‘pictures of the floating world’. The printed inscriptions on these images feature diverse place-names, both man-made and natural formations. However, due to the corpus’s richness and diversity, the precise nature of artistic mediation of the depicted places remains little understood. In this paper, we explored a new analytical approach based on the macroanalysis of images facilitated by Natural Language Processing technologies. This paper presents a small dataset with inscriptions on prints that have been annotated by an art historian for included place-name entities. Our dataset is released for public use. By fine-tuning and applying a Japanese BERT-based Name Entity Recogniser, we provide a use-case of a macroanalysis of a visual dataset that is hosted by the digital database of the Art Research Center at the Ritsumeikan University, Kyoto. Our work studies the relationship between topography and its visual renderings in early modern Japanese ukiyo-e landscape prints, demonstrating how an art historian’s work can be improved with Natural Language Processing toward distant viewing of visual datasets. We release our dataset and code for public use: https://github.com/connalia/ukiyo-e_meisho_nlp Konstantina Liagkou, John Pavlopoulos, Ewa Machotka |
LREC | 2 |
| 2022 | Sentiment Analysis of Homeric Text: The 1st Book of IliadabstractSentiment analysis studies are focused more on online customer reviews or social media, and less on literary studies. The problem is greater for ancient languages, where the linguistic expression of sentiments may diverge from modern linguistic forms. This work presents the outcome of a sentiment annotation task of the first Book of Iliad, an ancient Greek poem. The annotators were provided with verses translated into modern Greek and they annotated the perceived emotions and sentiments verse by verse. By estimating the fraction of annotators that found a verse as belonging to a specific sentiment class, we model the poem’s perceived sentiment as a multi-variate time series. By experimenting with a state of the art deep learning masked language model, pre-trained on modern Greek and fine-tuned to estimate the sentiment of our data, we registered a mean squared error of 0.063. This low error indicates that sentiment estimators built on our dataset can potentially be used as mechanical annotators, hence facilitating the distant reading of Homeric text. Our dataset is released for public use. John Pavlopoulos, Alexandros Xenos, Davide Picca |
LREC | 1 |
| 2022 | Handwritten Paleographic Greek Text Recognition: A Century-Based ApproachabstractToday classicists are provided with a great number of digital tools which, in turn, offer possibilities for further study and new research goals. In this paper we explore the idea that old Greek handwriting can be machine-readable and consequently, researchers can study the target material fast and efficiently. Previous studies have shown that Handwritten Text Recognition (HTR) models are capable of attaining high accuracy rates. However, achieving high accuracy HTR results for Greek manuscripts is still considered to be a major challenge. The overall aim of this paper is to assess HTR for old Greek manuscripts. To address this statement, we study and use digitized images of the Oxford University Bodleian Library Greek manuscripts. By manually transcribing 77 images, we created and present here a new dataset for Handwritten Paleographic Greek Text Recognition. The dataset instances were organized by establishing as a leading factor the century to which the manuscript and hence the image belongs. Experimenting then with an HTR model we show that the error rate depends on the century of the image. Paraskevi Platanou, John Pavlopoulos |
LREC | 2 |
| 2022 | Diagnostic captioning: a surveyabstractAbstract Diagnostic captioning (DC) concerns the automatic generation of a diagnostic text from a set of medical images of a patient collected during an examination. DC can assist inexperienced physicians, reducing clinical errors. It can also help experienced physicians produce diagnostic reports faster. Following the advances of deep learning, especially in generic image captioning, DC has recently attracted more attention, leading to several systems and datasets. This article is an extensive overview of DC. It presents relevant datasets, evaluation measures, and up-to-date systems. It also highlights shortcomings that hinder DC’s progress and proposes future directions. John Pavlopoulos, Vasiliki Kougia, Ion Androutsopoulos, Dimitris Papamichail |
Knowl. Inf. Syst. | 1 |
| 2021 | Customized Neural Predictive Medical Text: A Use-Case on Caregivers
John Pavlopoulos, Panagiotis Papapetrou |
AIME | 1 |
| 2021 | Automated Grading of Exam Responses: An Extensive Classification Benchmark
Jimmy Ljungman, Vanessa Lislevand, John Pavlopoulos, Alexandra Farazouli, Zed Lee, Panagiotis Papapetrou, Uno Fors |
DS | 3 |
| 2021 | Sentiment Nowcasting During the COVID-19 Pandemic
Ioanna Miliou, John Pavlopoulos, Panagiotis Papapetrou |
DS | 2 |
| 2021 | Civil Rephrases Of Toxic Texts With Self-Supervised TransformersabstractPlatforms that support online commentary, from social networks to news sites, are increasingly leveraging machine learning to assist their moderation efforts.But this process does not typically provide feedback to the author that would help them contribute according to the community guidelines.This is prohibitively time-consuming for human moderators to do, and computational approaches are still nascent.This work focuses on models that can help suggest rephrasings of toxic comments in a more civil manner.Inspired by recent progress in unpaired sequence-tosequence tasks, a self-supervised learning model is introduced, called CAE-T5 1 .CAE-T5 employs a pre-trained text-to-text transformer, which is fine tuned with a denoising and cyclic auto-encoder loss.Experimenting with the largest toxicity detection dataset to date (Civil Comments) our model generates sentences that are more fluent and better at preserving the initial content compared to earlier text style transfer systems which we compare with using several scoring systems and human evaluation. Léo Laugier, John Pavlopoulos, Jeffrey S. Sorensen, Lucas Dixon |
EACL | 2 |
| 2021 | RTEX: A novel framework for ranking, tagging, and explanatory diagnostic captioning of radiography examsabstractOBJECTIVE: The study sought to assist practitioners in identifying and prioritizing radiography exams that are more likely to contain abnormalities, and provide them with a diagnosis in order to manage heavy workload more efficiently (eg, during a pandemic) or avoid mistakes due to tiredness. MATERIALS AND METHODS: This article introduces RTEx, a novel framework for (1) ranking radiography exams based on their probability to be abnormal, (2) generating abnormality tags for abnormal exams, and (3) providing a diagnostic explanation in natural language for each abnormal exam. Our framework consists of deep learning and retrieval methods and is assessed on 2 publicly available datasets. RESULTS: For ranking, RTEx outperforms its competitors in terms of nDCG@k. The tagging component outperforms 2 strong competitor methods in terms of F1. Moreover, the diagnostic captioning component, which exploits the predicted tags to constrain the captioning process, outperforms 4 captioning competitors with respect to clinical precision and recall. DISCUSSION: RTEx prioritizes abnormal exams toward the improvement of the healthcare workflow by introducing a ranking method. Also, for each abnormal radiography exam RTEx generates a set of abnormality tags alongside a diagnostic text to explain the tags and guide the medical expert. Human evaluation of the produced text shows that employing the generated tags offers consistency to the clinical correctness and that the sentences of each text have high clinical accuracy. CONCLUSIONS: This is the first framework that successfully combines 3 tasks: ranking, tagging, and diagnostic captioning with focus on radiography exams that contain abnormalities. Vasiliki Kougia, John Pavlopoulos, Panagiotis Papapetrou, Max Gordon |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Toxicity Detection: Does Context Really Matter?abstractModeration is crucial to promoting healthy online discussions.Although several 'toxicity' detection datasets and models have been published, most of them ignore the context of the posts, implicitly assuming that comments may be judged independently.We investigate this assumption by focusing on two questions: (a) does context affect the human judgement, and (b) does conditioning on context improve performance of toxicity detection systems?We experiment with Wikipedia conversations, limiting the notion of context to the previous post in the thread and the discussion title.We find that context can both amplify or mitigate the perceived toxicity of posts.Moreover, a small but significant subset of manually labeled posts (5% in one of our experiments) end up having the opposite toxicity labels if the annotators are not provided with context.Surprisingly, we also find no evidence that context actually improves the performance of toxicity classifiers, having tried a range of classifiers and mechanisms to make them context aware.This points to the need for larger datasets of comments annotated in context.We make our code and data publicly available. John Pavlopoulos, Jeffrey S. Sorensen, Lucas Dixon, Nithum Thain, Ion Androutsopoulos |
ACL | 1 |
| 2020 | Clinical Predictive Keyboard using Statistical and Neural Language ModelingabstractA language model can be used to predict the next word during authoring, to correct spelling or to accelerate writing (e.g., in sms or emails). Language models, however, have only been applied in a very small scale to assist physicians during authoring (e.g., discharge summaries or radiology reports). But along with the assistance to the physician, computer-based systems which expedite the patient's exit also assist in decreasing the hospital infections. We employed statistical and neural language modeling to predict the next word of a clinical text and assess all the models in terms of accuracy and keystroke discount in two datasets with radiology reports. We show that a neural language model can achieve as high as 51.3% accuracy in radiology reports (one out of two words predicted correctly). We also show that even when the models are employed only for frequent words, the physician can save valuable time. John Pavlopoulos, Panagiotis Papapetrou |
CBMS | 1 |
| 2017 | Deeper Attention to Abusive User Content ModerationabstractExperimenting with a new dataset of 1.6M user comments from a news portal and an existing dataset of 115K Wikipedia talk page comments, we show that an RNN operating on word embeddings outpeforms the previous state of the art in moderation, which used logistic regression or an MLP classifier with character or word n-grams.We also compare against a CNN operating on word embeddings, and a word-list baseline.A novel, deep, classificationspecific attention mechanism improves the performance of the RNN further, and can also highlight suspicious words for free, without including highlighted words in the training data.We consider both fully automatic and semi-automatic moderation. John Pavlopoulos, Prodromos Malakasiotis, Ion Androutsopoulos |
EMNLP | 1 |
| 2015 | An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competitionabstractBACKGROUND: This article provides an overview of the first BIOASQ challenge, a competition on large-scale biomedical semantic indexing and question answering (QA), which took place between March and September 2013. BIOASQ assesses the ability of systems to semantically index very large numbers of biomedical scientific articles, and to return concise and user-understandable answers to given natural language questions by combining information from biomedical articles and ontologies. RESULTS: The 2013 BIOASQ competition comprised two tasks, Task 1a and Task 1b. In Task 1a participants were asked to automatically annotate new PUBMED documents with MESH headings. Twelve teams participated in Task 1a, with a total of 46 system runs submitted, and one of the teams performing consistently better than the MTI indexer used by NLM to suggest MESH headings to curators. Task 1b used benchmark datasets containing 29 development and 282 test English questions, along with gold standard (reference) answers, prepared by a team of biomedical experts from around Europe and participants had to automatically produce answers. Three teams participated in Task 1b, with 11 system runs. The BIOASQ infrastructure, including benchmark datasets, evaluation mechanisms, and the results of the participants and baseline methods, is publicly available. CONCLUSIONS: A publicly available evaluation infrastructure for biomedical semantic indexing and QA has been developed, which includes benchmark datasets, and can be used to evaluate systems that: assign MESH headings to published articles or to English questions; retrieve relevant RDF triples from ontologies, relevant articles and snippets from PUBMED Central; produce "exact" and paragraph-sized "ideal" answers (summaries). The results of the systems that participated in the 2013 BIOASQ competition are promising. In Task 1a one of the systems performed consistently better from the NLM's MTI indexer. In Task 1b the systems received high scores in the manual evaluation of the "ideal" answers; hence, they produced high quality summaries as answers. Overall, BIOASQ helped obtain a unified view of how techniques from text classification, semantic indexing, document and passage retrieval, question answering, and text summarization can be combined to allow biomedical experts to obtain concise, user-understandable answers to questions reflecting their real information needs. George Tsatsaronis 0001, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry Artières, Axel-Cyrille Ngonga Ngomo, Norman Heino, Éric Gaussier, Liliana Barrio-Alvers, Michael Schroeder 0001, Ion Androutsopoulos, Georgios Paliouras |
BMC Bioinform. | 12 |
| 2014 | A Vague Sense Classifier for Detecting Vague Definitions in OntologiesabstractVagueness is a common human knowledge and linguistic phenomenon, typically manifested by predicates that lack clear applicability conditions and boundaries such as High, Expert or Bad.In the context of ontologies and semantic data, the usage of such predicates within ontology element definitions (classes, relations etc.) can hamper the latter's quality, primarily in terms of shareability and meaning explicitness.With that in mind, we present in this paper a vague word sense classifier that may help both ontology creators and consumers to automatically detect vague ontology definitions and, thus, assess their quality better. Panos Alexopoulos, John Pavlopoulos |
EACL | 2 |
| 2014 | Multi-Granular Aspect Aggregation in Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis estimates the sentiment expressed for each particular aspect (e.g., battery, screen) of an entity (e.g., smartphone).Different words or phrases, however, may be used to refer to the same aspect, and similar aspects may need to be aggregated at coarser or finer granularities to fit the available space or satisfy user preferences.We introduce the problem of aspect aggregation at multiple granularities.We decompose it in two processing phases, to allow previous work on term similarity and hierarchical clustering to be reused.We show that the second phase, where aspects are clustered, is almost a solved problem, whereas further research is needed in the first phase, where semantic similarity measures are employed.We also introduce a novel sense pruning mechanism for WordNet-based similarity measures, which improves their performance in the first phase.Finally, we provide publicly available benchmark datasets. John Pavlopoulos, Ion Androutsopoulos |
EACL | 1 |