EDBT 2026 Demo / reviewers in the wild / expert
Ion Androutsopoulos
dblp:87/6723
· DBLP profile ↗
46ranked-venue papers
7as first author
14since 2021 · last 2025
0009-0000-2969-0509ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News ArticlesabstractNikolaos Nikolaidis, Nicolas Stefanovitch, Purificação Silvano, Dimitar Iliyanov Dimitrov, Roman Yangarber, Nuno Guimarães, Elisa Sartori, Ion Androutsopoulos, Preslav Nakov, Giovanni Da San Martino, Jakub Piskorski. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nikolaos Nikolaidis 0004, Nicolas Stefanovitch, Purificação Silvano, Dimitar Iliyanov Dimitrov, Roman Yangarber, Nuno Guimarães, Elisa Sartori, Ion Androutsopoulos, Preslav Nakov, Giovanni Da San Martino, Jakub Piskorski |
ACL (1) | 8 |
| 2025 | Evaluation and Facilitation of Online Discussions in the LLM Era: A SurveyabstractKaterina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé, Danai Myrtzani, Theodoros Evgeniou, Ion Androutsopoulos, John Pavlopoulos. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Katerina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé, Danai Myrtzani, Theodoros Evgeniou, Ion Androutsopoulos, John Pavlopoulos |
EMNLP | 7 |
| 2025 | Building Open-Retrieval Conversational Question Answering Systems by Generating Synthetic Data and Decontextualizing User QuestionsabstractWe consider open-retrieval conversational question answering (OR-CONVQA), an extension of question answering where system responses need to be (i) aware of dialog history and (ii) grounded in documents (or document fragments) retrieved per question. Domain-specific OR-CONVQA training datasets are crucial for real-world applications, but hard to obtain. We propose a pipeline that capitalizes on the abundance of plain text documents in organizations (e.g., product documentation) to automatically produce realistic OR-CONVQA dialogs with annotations. Similarly to real-world humanannotated OR-CONVQA datasets, we generate in-dialog question-answer pairs, self-contained (decontextualized, e.g., no referring expressions) versions of user questions, and propositions (sentences expressing prominent information from the documents) the system responses are grounded in. We show how the synthetic dialogs can be used to train efficient question rewriters that decontextualize user questions, allowing existing dialog-unaware retrievers to be utilized. The retrieved information and the decontextualized question are then passed on to an LLM that generates the system’s response. Christos Vlachos, Nikolaos Stylianou, Alexandra Fiotaki, Spiros Methenitis, Elisavet Palogiannidi, Themos Stafylakis, Ion Androutsopoulos |
SIGDIAL | 7 |
| 2024 | Still All Greeklish to Me: Greeklish to Greek TransliterationabstractModern Greek is normally written in the Greek alphabet. In informal online messages, however, Greek is often written using characters available on Latin-character keyboards, a form known as Greeklish. Originally used to bypass the lack of support for the Greek alphabet in older computers, Greeklish is now also used to avoid switching languages on multilingual keyboards, hide spelling mistakes, or as a form of slang. There is no consensus mapping, hence the same Greek word can be written in numerous different ways in Greeklish. Even native Greek speakers may struggle to understand (or be annoyed by) Greeklish, which requires paying careful attention to context to decipher. Greeklish may also be a problem for NLP models trained on Greek datasets written in the Greek alphabet. Experimenting with a range of statistical and deep learning models on both artificial and real-life Greeklish data, we find that: (i) prompting large language models (e.g., GPT-4) performs impressively well with few- or even zero-shot training, outperforming several fine-tuned encoder-decoder models; however (ii) a twenty years old statistical Greeklish transliteration model is still very competitive; and (iii) the problem is still far from having been solved; (iv) nevertheless, downstream Greek NLP systems that need to cope with Greeklish, such as moderation classifiers, can benefit significantly even with the current non-perfect transliteration systems. We make all our code, models, and data available and suggest future improvements, based on an analysis of our experimental results. Anastasios Toumazatos, John Pavlopoulos, Ion Androutsopoulos, Stavros Vassos |
LREC/COLING | 3 |
| 2024 | Should I try multiple optimizers when fine-tuning a pre-trained Transformer for NLP tasks? Should I tune their hyperparameters?abstractNefeli Gkouti, Prodromos Malakasiotis, Stavros Toumpis, Ion Androutsopoulos. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Nefeli Gkouti, Prodromos Malakasiotis, Stavros Toumpis, Ion Androutsopoulos |
EACL (1) | 4 |
| 2023 | Machine Learning for Ancient Languages: A SurveyabstractAbstract Ancient languages preserve the cultures and histories of the past. However, their study is fraught with difficulties, and experts must tackle a range of challenging text-based tasks, from deciphering lost languages to restoring damaged inscriptions, to determining the authorship of works of literature. Technological aids have long supported the study of ancient texts, but in recent years advances in artificial intelligence and machine learning have enabled analyses on a scale and in a detail that are reshaping the field of humanities, similarly to how microscopes and telescopes have contributed to the realm of science. This article aims to provide a comprehensive survey of published research using machine learning for the study of ancient texts written in any language, script, and medium, spanning over three and a half millennia of civilizations around the ancient world. To analyze the relevant literature, we introduce a taxonomy of tasks inspired by the steps involved in the study of ancient documents: digitization, restoration, attribution, linguistic analysis, textual criticism, translation, and decipherment. This work offers three major contributions: first, mapping the interdisciplinary field carved out by the synergy between the humanities and machine learning; second, highlighting how active collaboration between specialists from both fields is key to producing impactful and compelling scholarship; third, highlighting promising directions for future work in this field. Thus, this work promotes and supports the continued collaborative impetus between the humanities and machine learning. Thea Sommerschield, Yannis M. Assael, John Pavlopoulos, Vanessa Stefanak, Andrew W. Senior, Chris Dyer, John Bodel, Jonathan Prag, Ion Androutsopoulos, Nando de Freitas |
Comput. Linguistics | 9 |
| 2022 | LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishabstractIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, Nikolaos Aletras. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II, Ion Androutsopoulos, Daniel Martin Katz, Nikolaos Aletras |
ACL (1) | 5 |
| 2022 | FiNER: Financial Numeric Entity Recognition for XBRL TaggingabstractLefteris Loukas, Manos Fergadiotis, Ilias Chalkidis, Eirini Spyropoulou, Prodromos Malakasiotis, Ion Androutsopoulos, Georgios Paliouras. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Lefteris Loukas, Manos Fergadiotis, Ilias Chalkidis, Eirini Spyropoulou, Prodromos Malakasiotis, Ion Androutsopoulos, Georgios Paliouras |
ACL (1) | 6 |
| 2022 | From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil TransferabstractJohn Pavlopoulos, Leo Laugier, Alexandros Xenos, Jeffrey Sorensen, Ion Androutsopoulos. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. John Pavlopoulos, Léo Laugier, Alexandros Xenos, Jeffrey S. Sorensen, Ion Androutsopoulos |
ACL (1) | 5 |
| 2022 | Diagnostic captioning: a surveyabstractAbstract Diagnostic captioning (DC) concerns the automatic generation of a diagnostic text from a set of medical images of a patient collected during an examination. DC can assist inexperienced physicians, reducing clinical errors. It can also help experienced physicians produce diagnostic reports faster. Following the advances of deep learning, especially in generic image captioning, DC has recently attracted more attention, leading to several systems and datasets. This article is an extensive overview of DC. It presents relevant datasets, evaluation measures, and up-to-date systems. It also highlights shortcomings that hinder DC’s progress and proposes future directions. John Pavlopoulos, Vasiliki Kougia, Ion Androutsopoulos, Dimitris Papamichail |
Knowl. Inf. Syst. | 3 |
| 2022 | Deception detection in text and its relation to the cultural dimension of individualism/collectivismabstractAbstract Automatic deception detection is a crucial task that has many applications both in direct physical and in computer-mediated human communication. Our focus is on automatic deception detection in text across cultures. In this context, we view culture through the prism of the individualism/collectivism dimension, and we approximate culture by using country as a proxy. Having as a starting point recent conclusions drawn from the social psychology discipline, we explore if differences in the usage of specific linguistic features of deception across cultures can be confirmed and attributed to cultural norms in respect to the individualism/collectivism divide. In addition, we investigate if a universal feature set for cross-cultural text deception detection tasks exists. We evaluate the predictive power of different feature sets and approaches. We create culture/language-aware classifiers by experimenting with a wide range of n-gram features from several levels of linguistic analysis, namely phonology, morphology and syntax, other linguistic cues like word and phoneme counts, pronouns use, etc., and token embeddings. We conducted our experiments over eleven data sets from five languages (English, Dutch, Russian, Spanish, and Romanian), from six countries (United States of America, Belgium, India, Russia, Mexico, and Romania), and we applied two classification methods, namely logistic regression and fine-tuned BERT models. The results showed that the undertaken task is fairly complex and demanding. Furthermore, there are indications that some linguistic cues of deception have cultural origins and are consistent in the context of diverse domains and data set settings for the same language. This is more evident for the usage of pronouns and the expression of sentiment in deceptive language. The results of this work show that the automatic deception detection across cultures and languages cannot be handled in unified manners and that such approaches should be augmented with knowledge about cultural differences and the domains of interest. Katerina Papantoniou, Panagiotis Papadakos, Theodore Patkos, Giorgos Flouris, Ion Androutsopoulos, Dimitris Plexousakis |
Nat. Lang. Eng. | 5 |
| 2021 | A Neural Model for Joint Document and Snippet Ranking in Question Answering for Large Document CollectionsabstractDimitris Pappas, Ion Androutsopoulos. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Dimitris Pappas, Ion Androutsopoulos |
ACL/IJCNLP (1) | 2 |
| 2021 | MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transferabstractWe introduce MULTI-EURLEX, a new multilingual dataset for topic classification of legal documents.The dataset comprises 65k European Union (EU) laws, officially translated in 23 languages, annotated with multiple labels from the EUROVOC taxonomy.We highlight the effect of temporal concept drift and the importance of chronological, instead of random splits.We use the dataset as a testbed for zeroshot cross-lingual transfer, where we exploit annotated training documents in one language (source) to classify documents in another language (target).We find that fine-tuning a multilingually pretrained model (XLM-ROBERTA, MT5) in a single source language leads to catastrophic forgetting of multilingual knowledge and, consequently, poor zero-shot transfer to other languages.Adaptation strategies, namely partial fine-tuning, adapters, BITFIT, LNFIT, originally proposed to accelerate finetuning for new end-tasks, help retain multilingual knowledge from pretraining, substantially improving zero-shot cross-lingual transfer, but their impact also depends on the pretrained model used and the size of the label set.Language ISO Member Countries where official EU Speakers (%) Number of Documents Words per document code Native Total Train Dev. Ilias Chalkidis, Manos Fergadiotis, Ion Androutsopoulos |
EMNLP (1) | 3 |
| 2021 | Paragraph-level Rationale Extraction through Regularization: A case study on European Court of Human Rights CasesabstractIlias Chalkidis, Manos Fergadiotis, Dimitrios Tsarapatsanis, Nikolaos Aletras, Ion Androutsopoulos, Prodromos Malakasiotis. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ilias Chalkidis, Manos Fergadiotis, Dimitrios Tsarapatsanis, Nikolaos Aletras, Ion Androutsopoulos, Prodromos Malakasiotis |
NAACL-HLT | 5 |
| 2020 | Toxicity Detection: Does Context Really Matter?abstractModeration is crucial to promoting healthy online discussions.Although several 'toxicity' detection datasets and models have been published, most of them ignore the context of the posts, implicitly assuming that comments may be judged independently.We investigate this assumption by focusing on two questions: (a) does context affect the human judgement, and (b) does conditioning on context improve performance of toxicity detection systems?We experiment with Wikipedia conversations, limiting the notion of context to the previous post in the thread and the discussion title.We find that context can both amplify or mitigate the perceived toxicity of posts.Moreover, a small but significant subset of manually labeled posts (5% in one of our experiments) end up having the opposite toxicity labels if the annotators are not provided with context.Surprisingly, we also find no evidence that context actually improves the performance of toxicity classifiers, having tried a range of classifiers and mechanisms to make them context aware.This points to the need for larger datasets of comments annotated in context.We make our code and data publicly available. John Pavlopoulos, Jeffrey S. Sorensen, Lucas Dixon, Nithum Thain, Ion Androutsopoulos |
ACL | 5 |
| 2020 | An Empirical Study on Large-Scale Multi-Label Text Classification Including Few and Zero-Shot LabelsabstractIlias Chalkidis, Manos Fergadiotis, Sotiris Kotitsas, Prodromos Malakasiotis, Nikolaos Aletras, Ion Androutsopoulos. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Ilias Chalkidis, Manos Fergadiotis, Sotiris Kotitsas, Prodromos Malakasiotis, Nikolaos Aletras, Ion Androutsopoulos |
EMNLP (1) | 6 |
| 2019 | Neural Legal Judgment Prediction in EnglishabstractLegal judgment prediction is the task of automatically predicting the outcome of a court case, given a text describing the case's facts.Previous work on using neural models for this task has focused on Chinese; only featurebased models (e.g., using bags of words and topics) have been considered in English.We release a new English legal judgment prediction dataset, containing cases from the European Court of Human Rights.We evaluate a broad variety of neural models on the new dataset, establishing strong baselines that surpass previous feature-based models in three tasks: (1) binary violation classification; (2) multi-label classification; (3) case importance prediction.We also explore if models are biased towards demographic information via data anonymization.As a side-product, we propose a hierarchical version of BERT, which bypasses BERT's length limitation. Ilias Chalkidis, Ion Androutsopoulos, Nikolaos Aletras |
ACL (1) | 2 |
| 2019 | Large-Scale Multi-Label Text Classification on EU LegislationabstractWe consider Large-Scale Multi-Label Text Classification (LMTC) in the legal domain.We release a new dataset of 57k legislative documents from EUR-LEX, annotated with ∼4.3k EUROVOC labels, which is suitable for LMTC, few-and zero-shot learning.Experimenting with several neural classifiers, we show that BIGRUs with label-wise attention perform better than other current state of the art methods.Domain-specific WORD2VEC and context-sensitive ELMO embeddings further improve performance.We also find that considering only particular zones of the documents is sufficient.This allows us to bypass BERT's maximum text length limit and finetune BERT, obtaining the best results in all but zero-shot learning cases. Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Ion Androutsopoulos |
ACL (1) | 4 |
| 2019 | SUM-QE: a BERT-based Summary Quality Estimation ModelabstractStratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, Ion Androutsopoulos. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, Ion Androutsopoulos |
EMNLP/IJCNLP (1) | 4 |
| 2018 | Deep Relevance Ranking using Enhanced Document-Query InteractionsabstractWe explore several new models for document relevance ranking, building upon the Deep Relevance Matching Model (DRMM) of Guo et al. (2016).Unlike DRMM, which uses context-insensitive encodings of terms and query-document term interactions, we inject rich context-sensitive encodings throughout our models, inspired by PACRR's (Hui et al., 2017) convolutional n-gram matching features, but extended in several ways including multiple views of query and document inputs.We test our models on datasets from the BIOASQ question answering challenge (Tsatsaronis et al., 2015) and TREC ROBUST 2004(Voorhees, 2005), showing they outperform BM25-based baselines, DRMM, and PACRR. Ryan T. McDonald, Georgios-Ioannis Brokos, Ion Androutsopoulos |
EMNLP | 3 |
| 2018 | BioRead: A New Dataset for Biomedical Reading Comprehension
Dimitris Pappas, Ion Androutsopoulos, Harris Papageorgiou |
LREC | 2 |
| 2018 | Ontology Driven Extraction of Research Processes
Vayianos Pertsas, Panos Constantopoulos, Ion Androutsopoulos |
ISWC (1) | 3 |
| 2017 | Deeper Attention to Abusive User Content ModerationabstractExperimenting with a new dataset of 1.6M user comments from a news portal and an existing dataset of 115K Wikipedia talk page comments, we show that an RNN operating on word embeddings outpeforms the previous state of the art in moderation, which used logistic regression or an MLP classifier with character or word n-grams.We also compare against a CNN operating on word embeddings, and a word-list baseline.A novel, deep, classificationspecific attention mechanism improves the performance of the RNN further, and can also highlight suspicious words for free, without including highlighted words in the training data.We consider both fully automatic and semi-automatic moderation. John Pavlopoulos, Prodromos Malakasiotis, Ion Androutsopoulos |
EMNLP | 3 |
| 2017 | Extracting contract elementsabstractWe study how contract element extraction can be automated. We provide a labeled dataset with gold contract element annotations, along with an unlabeled dataset of contracts that can be used to pre-train word embeddings. Both datasets are provided in an encoded form to bypass privacy issues. We describe and experimentally compare several contract element extraction methods that use manually written rules and linear classifiers (logistic regression, SVMs) with hand-crafted features, word embeddings, and part-of-speech tag embeddings. The best results are obtained by a hybrid method that combines machine learning (with hand-crafted features and embeddings) and manually written post-processing rules. Ilias Chalkidis, Ion Androutsopoulos, Achilleas Michos |
ICAIL | 2 |
| 2017 | A Deep Learning Approach to Contract Element ExtractionabstractWe explore how deep learning methods can be used for contract element extraction. We show that a BILSTM operating on word, POS tag, and token-shape embeddings outperforms the linear sliding-window classifiers of our previous work, without any manually written rules. Further improvements are observed by stacking an additional LSTM on top of the BILSTM, or by adding a CRF layer on top of the BILSTM. The stacked BILSTM-LSTM misclassifies fewer tokens, but the BILSTM-CRF combination performs better when methods are evaluated for their ability to extract entire, possibly multi-token contract elements. Ilias Chalkidis, Ion Androutsopoulos |
JURIX | 2 |
| 2017 | A Personalized Global Filter To Predict RetweetsabstractInformation shared on Twitter is ever increasing and users-recipients are overwhelmed by the number of tweets they receive, many of which of no interest. Filters that estimate the interest of each incoming post can alleviate this problem, for example by allowing users to sort incoming posts by predicted interest (e.g., "top stories" vs. "most recent" in Facebook). Global and personal filters have been used to detect interesting posts in social networks. Global filters are trained on large collections of posts and reactions to posts (e.g., retweets), aiming to predict how interesting a post is for a broad audience. In contrast, personal filters are trained on posts received by a particular user and the reactions of the particular user. Personal filters can provide recommendations tailored to a particular user's interests, which may not coincide with the interests of the majority of users that global filters are trained to predict. On the other hand, global filters are typically trained on much larger datasets compared to personal filters. Hence, global filters may work better in practice, especially with new users, for which personal filters may have very few training instances ("cold start" problem). Michail Vougioukas, Ion Androutsopoulos, Georgios Paliouras |
UMAP | 2 |
| 2015 | An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competitionabstractBACKGROUND: This article provides an overview of the first BIOASQ challenge, a competition on large-scale biomedical semantic indexing and question answering (QA), which took place between March and September 2013. BIOASQ assesses the ability of systems to semantically index very large numbers of biomedical scientific articles, and to return concise and user-understandable answers to given natural language questions by combining information from biomedical articles and ontologies. RESULTS: The 2013 BIOASQ competition comprised two tasks, Task 1a and Task 1b. In Task 1a participants were asked to automatically annotate new PUBMED documents with MESH headings. Twelve teams participated in Task 1a, with a total of 46 system runs submitted, and one of the teams performing consistently better than the MTI indexer used by NLM to suggest MESH headings to curators. Task 1b used benchmark datasets containing 29 development and 282 test English questions, along with gold standard (reference) answers, prepared by a team of biomedical experts from around Europe and participants had to automatically produce answers. Three teams participated in Task 1b, with 11 system runs. The BIOASQ infrastructure, including benchmark datasets, evaluation mechanisms, and the results of the participants and baseline methods, is publicly available. CONCLUSIONS: A publicly available evaluation infrastructure for biomedical semantic indexing and QA has been developed, which includes benchmark datasets, and can be used to evaluate systems that: assign MESH headings to published articles or to English questions; retrieve relevant RDF triples from ontologies, relevant articles and snippets from PUBMED Central; produce "exact" and paragraph-sized "ideal" answers (summaries). The results of the systems that participated in the 2013 BIOASQ competition are promising. In Task 1a one of the systems performed consistently better from the NLM's MTI indexer. In Task 1b the systems received high scores in the manual evaluation of the "ideal" answers; hence, they produced high quality summaries as answers. Overall, BIOASQ helped obtain a unified view of how techniques from text classification, semantic indexing, document and passage retrieval, question answering, and text summarization can be combined to allow biomedical experts to obtain concise, user-understandable answers to questions reflecting their real information needs. George Tsatsaronis 0001, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry Artières, Axel-Cyrille Ngonga Ngomo, Norman Heino, Éric Gaussier, Liliana Barrio-Alvers, Michael Schroeder 0001, Ion Androutsopoulos, Georgios Paliouras |
BMC Bioinform. | 21 |
| 2015 | Evaluation measures for hierarchical classification: a unified view and novel approaches
Aris Kosmopoulos, Ioannis Partalas, Éric Gaussier, Georgios Paliouras, Ion Androutsopoulos |
Data Min. Knowl. Discov. | 5 |
| 2014 | Multi-Granular Aspect Aggregation in Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis estimates the sentiment expressed for each particular aspect (e.g., battery, screen) of an entity (e.g., smartphone).Different words or phrases, however, may be used to refer to the same aspect, and similar aspects may need to be aggregated at coarser or finer granularities to fit the available space or satisfy user preferences.We introduce the problem of aspect aggregation at multiple granularities.We decompose it in two processing phases, to allow previous work on term similarity and hierarchical clustering to be reused.We show that the second phase, where aspects are clustered, is almost a solved problem, whereas further research is needed in the first phase, where semantic similarity measures are employed.We also introduce a novel sense pruning mechanism for WordNet-based similarity measures, which improves their performance in the first phase.Finally, we provide publicly available benchmark datasets. John Pavlopoulos, Ion Androutsopoulos |
EACL | 2 |
| 2014 | Web-scale classification: web classification in the big data eraabstractThis paper provides an overview of the workshop Web-Scale Classification: Web Classification in the Big Data Era which was held in New York City, on February 28th as a workshop of the seventh International Conference on Web Search and Data Mining. The goal of the workshop was to discuss and assess recent research focusing on classification and mining in Web-scale category systems. The workshop brought together members of several communities such web mining, machine learning, text classification and social media mining. Ioannis Partalas, Massih-Reza Amini, Ion Androutsopoulos, Thierry Artières, Patrick Gallinari, Éric Gaussier, Georgios Paliouras |
WSDM | 3 |
| 2013 | Generating Natural Language Descriptions from OWL Ontologies: the NaturalOWL SystemabstractWe present NaturalOWL, a natural language generation system that produces texts describing individuals or classes of OWL ontologies. Unlike simpler OWL verbalizers, which typically express a single axiom at a time in controlled, often not entirely fluent natural language primarily for the benefit of domain experts, we aim to generate fluent and coherent multi-sentence texts for end-users. With a system like NaturalOWL, one can publish information in OWL on the Web, along with automatically produced corresponding texts in multiple languages, making the information accessible not only to computer programs and domain experts, but also end-users. We discuss the processing stages of NaturalOWL, the optional domain-dependent linguistic resources that the system can use at each stage, and why they are useful. We also present trials showing that when the domain-dependent llinguistic resources are available, NaturalOWL produces significantly better texts compared to a simpler verbalizer, and that the resources can be created with relatively light effort. Ion Androutsopoulos, Gerasimos Lampouras, Dimitrios Galanis |
J. Artif. Intell. Res. | 1 |
| 2012 | Extractive Multi-Document Summarization with Integer Linear Programming and Support Vector Regression
Dimitrios Galanis, Gerasimos Lampouras, Ion Androutsopoulos |
COLING | 3 |
| 2011 | A Generate and Rank Approach to Sentence Paraphrasing
Prodromos Malakasiotis, Ion Androutsopoulos |
EMNLP | 2 |
| 2010 | An extractive supervised two-stage method for sentence compression
Dimitrios Galanis, Ion Androutsopoulos |
HLT-NAACL | 2 |
| 2010 | A Survey of Paraphrasing and Textual Entailment MethodsabstractParaphrasing methods recognize, generate, or extract phrases, sentences, or longer natural language expressions that convey almost the same information. Textual entailment methods, on the other hand, recognize, generate, or extract pairs of natural language expressions, such that a human who reads (and trusts) the first element of a pair would most likely infer that the other element is also true. Paraphrasing can be seen as bidirectional textual entailment and methods from the two areas are often similar. Both kinds of methods are useful, at least in principle, in a wide range of natural language processing applications, including question answering, summarization, text generation, and machine translation. We summarize key ideas from the two areas by considering in turn recognition, generation, and extraction methods, also pointing to prominent articles and resources. Ion Androutsopoulos, Prodromos Malakasiotis |
J. Artif. Intell. Res. | 1 |
| 2009 | Finding Short Definitions of Terms on Web Pages
Gerasimos Lampouras, Ion Androutsopoulos |
EMNLP | 2 |
| 2007 | Word Sense Disambiguation with Spreading Activation Networks Generated from Thesauri
George Tsatsaronis 0001, Michalis Vazirgiannis, Ion Androutsopoulos |
IJCAI | 3 |
| 2007 | Source authoring for multilingual generation of personalised object descriptionsabstractWe present the source authoring facilities of a natural language generation system that produces personalised descriptions of objects in multiple natural languages starting from language-independent symbolic information in ontologies and databases as well as pieces of canned text. The system has been tested in applications ranging from museum exhibitions to presentations of computer equipment for sale. We discuss the architecture of the overall system, the resources that the authors manipulate, the functionality of the authoring facilities, the system's personalisation mechanisms, and how they relate to source authoring. A usability evaluation of the authoring facilities is also presented, followed by more recent work on reusing information extracted from existing databases and documents, and supporting the OWL ontology specification language. Ion Androutsopoulos, Jon Oberlander, Vangelis Karkaletsis |
Nat. Lang. Eng. | 1 |
| 2004 | Learning to Identify Single-Snippet Answers to Definition Questions
Spyridoula Miliaraki, Ion Androutsopoulos |
COLING | 2 |
| 2003 | A Memory-Based Approach to Anti-Spam Filtering for Mailing Lists
Georgios Sakkis, Ion Androutsopoulos, Georgios Paliouras, Vangelis Karkaletsis, Constantine D. Spyropoulos, Panagiotis Stamatopoulos |
Inf. Retr. | 2 |
| 2002 | Ellogon: A New Text Engineering Platform
Georgios Petasis, Vangelis Karkaletsis, Georgios Paliouras, Ion Androutsopoulos, Constantine D. Spyropoulos |
LREC | 4 |
| 2001 | Stacking Classifiers for Anti-Spam Filtering of E-Mail
Georgios Sakkis, Ion Androutsopoulos, Georgios Paliouras, Vangelis Karkaletsis, Constantine D. Spyropoulos, Panagiotis Stamatopoulos |
EMNLP | 2 |
| 2000 | Selectional Restrictions in HPS
Ion Androutsopoulos, Robert Dale |
COLING | 1 |
| 2000 | An experimental comparison of naive bayesian and keyword-based anti-spam filtering with personal e-mail messagesabstractThe growing problem of unsolicited bulk e-mail, also known as “spam”, has generated a need for reliable anti-spam e-mail filters. Filters of this type have so far been based mostly on manually constructed keyword patterns. An alternative approach has recently been proposed, whereby a Naive Bayesian classifier is trained automatically to detect spam messages. We test this approach on a large collection of personal e-mail messages, which we make publicly available in “encrypted” form contributing towards standard benchmarks. We introduce appropriate cost-sensitive measures, investigating at the same time the effect of attribute-set size, training-corpus size, lemmatization, and stop lists, issues that have not been explored in previous experiments. Finally, the Naive Bayesian filter is compared, in terms of performance, to a filter that uses keyword patterns, and which is part of a widely used e-mail reader. Ion Androutsopoulos, John Koutsias, Konstantinos Chandrinos, Constantine D. Spyropoulos |
SIGIR | 1 |
| 1998 | Time, tense and aspect in natural language database interfaces
Ion Androutsopoulos, Graeme D. Ritchie, Peter Thanisch |
Nat. Lang. Eng. | 1 |
| 1995 | Natural language interfaces to databases - an introductionabstractAbstract This paper is an introduction to natural language interfaces to databases (NLIDBS). A brief overview of the history of NLIDBS is first given. Some advantages and disadvantages of NLIDBS are then discussed, comparing NLIDBS to formal query languages, form-based interfaces, and graphical interfaces. An introduction to some of the linguistic problems NLIDBS have to confront follows, for the benefit of readers less familiar with computational linguistics. The discussion then moves on to NLIDB architectures, portability issues, restricted natural language input systems (including menu-based NLIDBS), and NLIDBS with reasoning capabilities. Some less explored areas of NLIDB research are then presented, namely database updates, meta-knowledge questions, temporal questions, and multi-modal NLIDBS. The paper ends with reflections on the current state of the art. Ion Androutsopoulos, Graeme D. Ritchie, Peter Thanisch |
Nat. Lang. Eng. | 1 |