VLDB 2026 Research / reviewers in the wild / expert
Efstathios Stamatatos
dblp:16/2424
· DBLP profile ↗
19ranked-venue papers in the field
5as first author
5since 2021 · last 2025
0000-0002-1336-9128ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 18 (5 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Overview of PAN 2025: Generative AI Detection, Multilingual Text Detoxification, Multi-author Writing Style Analysis, and Generative Plagiarism Detection - Extended Abstract
Janek Bevendorff, Daryna Dementieva, Maik Fröbe, Bela Gipp, André Greiner-Petter, Jussi Karlgren, Maximilian Mayerl, Preslav Nakov, Alexander Panchenko, Martin Potthast, Artem Shelmanov, Efstathios Stamatatos, Benno Stein 0001, Yuxia Wang 0003, Matti Wiegmann, Eva Zangerle |
ECIR (5) | 12 |
| 2024 | Overview of PAN 2024: Multi-author Writing Style Analysis, Multilingual Text Detoxification, Oppositional Thinking Analysis, and Generative AI Authorship Verification - Extended Abstract
Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashraf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korencic, Maximilian Mayerl, Animesh Mukherjee 0001, Alexander Panchenko, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Alisa Smirnova, Efstathios Stamatatos, Benno Stein 0001, Mariona Taulé, Dmitry Ustalov, Matti Wiegmann, Eva Zangerle |
ECIR (6) | 16 |
| 2023 | Overview of PAN 2023: Authorship Verification, Multi-author Writing Style Analysis, Profiling Cryptocurrency Influencers, and Trigger Detection - Extended Abstract
Janek Bevendorff, Mara Chinea-Rios, Marc Franco-Salvador, Annina Heini, Erik Körner, Krzysztof Kredens, Maximilian Mayerl, Piotr Pezik, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (3) | 12 |
| 2022 | Overview of PAN 2022: Authorship Verification, Profiling Irony and Stereotype Spreaders, Style Change Detection, and Trigger Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Elisabetta Fersini, Annina Heini, Mike Kestemont, Krzysztof Kredens, Maximilian Mayerl, Reyner Ortega-Bueno, Piotr Pezik, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (2) | 13 |
| 2021 | Overview of PAN 2021: Authorship Verification, Profiling Hate Speech Spreaders on Twitter, and Style Change Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Gretel Liz De la Peña Sarracén, Mike Kestemont, Enrique Manjavacas, Ilia Markov, Maximilian Mayerl, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (2) | 11 |
| 2020 | Shared Tasks on Authorship Analysis at PAN 2020
Janek Bevendorff, Bilal Ghanem, Anastasia Giahanou, Mike Kestemont, Enrique Manjavacas, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Günther Specht, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Eva Zangerle |
ECIR (2) | 10 |
| 2020 | Improved algorithms for extrinsic author verification
Nektaria Potha, Efstathios Stamatatos |
Knowl. Inf. Syst. | 2 |
| 2019 | Dynamic Ensemble Selection for Author Verification
Nektaria Potha, Efstathios Stamatatos |
ECIR (1) | 2 |
| 2019 | A Decade of Shared Tasks in Digital Text Forensics at PAN
Martin Potthast, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001 |
ECIR (2) | 3 |
| 2019 | Open-Set Web Genre Identification Using Distributional Features and Nearest Neighbors Distance Ratio
Dimitrios A. Pritsos, Anderson Rocha 0001, Efstathios Stamatatos |
ECIR (2) | 3 |
| 2019 | Improving author verification based on topic modelingabstractAuthorship analysis attempts to reveal information about authors of digital documents enabling applications in digital humanities, text forensics, and cyber‐security. Author verification is a fundamental task where, given a set of texts written by a certain author, we should decide whether another text is also by that author. In this article we systematically study the usefulness of topic modeling in author verification. We examine several author verification methods that cover the main paradigms, namely, intrinsic (attempt to solve a one‐class classification task) and extrinsic (attempt to solve a binary classification task) methods as well as profile‐based (all documents of known authorship are treated cumulatively) and instance‐based (each document of known authorship is treated separately) approaches combined with well‐known topic modeling methods such as Latent Semantic Indexing (LSI) and Latent Dirichlet Allocation (LDA). We use benchmark data sets and demonstrate that LDA is better combined with extrinsic methods, while the most effective intrinsic method is based on LSI. Moreover, topic modeling seems to be particularly effective for profile‐based approaches and the performance is enhanced when latent topics are extracted by an enriched set of documents. The comparison to state‐of‐the‐art methods demonstrates the great potential of the approaches presented in this study. It is also demonstrates that even when genre‐agnostic external documents are used, the proposed extrinsic models are very competitive. Nektaria Potha, Efstathios Stamatatos |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2018 | Masking topic-related information to enhance authorship attributionabstractAuthorship attribution attempts to reveal the authors of documents. In recent years, research in this field has grown rapidly. However, the performance of state‐of‐the‐art methods is heavily affected when text of known authorship and texts under investigation differ in topic and/or genre. So far, it is not clear how to quantify the personal style of authors in a way that is not affected by topic shifts or genre variations. In this paper, a set of text distortion methods are used attempting to mask topic‐related information. These methods transform the input texts into a more topic‐neutral form while maintaining the structure of documents associated with the personal style of the author. Using a controlled corpus that includes a fine‐grained range of topics and genres it is demonstrated how the proposed approach can be combined with existing authorship attribution methods to enhance their performance in very challenging tasks, especially in cross‐topic attribution. We also examine cross‐genre attribution and the most challenging, yet realistic, cross‐topic‐and‐genre attribution scenarios and show how the proposed techniques should be tuned to enhance performance in these tasks. Finally, we demonstrate that there are important differences in attribution effectiveness when either conversational genres, nonconversational genres, or a mix of them are considered. Efstathios Stamatatos |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Who Wrote the Web? Revisiting Influential Author Identification Research Applicable to Information Retrieval
Martin Potthast, Sarah Braun, Tolga Buz, Fabian Duffhauss, Florian Friedrich, Jörg Marvin Gülzow, Jakob Köhler, Winfried Lötzsch, Maike Elisa Müller, Robert Paßmann, Bernhard Reinke, Lucas Rettenmeier, Thomas Rometsch, Timo Sommer, Michael Träger, Sebastian Wilhelm, Benno Stein 0001, Efstathios Stamatatos, Matthias Hagen |
ECIR | 19 |
| 2013 | Open-Set Classification for Automated Genre Identification
Dimitrios A. Pritsos, Efstathios Stamatatos |
ECIR | 2 |
| 2011 | Plagiarism detection based on structural informationabstractIn this paper a novel method for detecting plagiarized passages in document collections is presented. In contrast to previous work in this field that uses mainly content terms to represent documents, the proposed method is based on structural information provided by occurrences of a small list of stopwords (i.e., very frequent words). We show that stopword n-grams are able to capture local syntactic similarities between suspicious and original documents. Moreover, an algorithm for detecting the exact boundaries of plagiarized and source passages is proposed. Experimental results on a publicly-available corpus demonstrate that the performance of the proposed approach is competitive when compared with the best reported results. More importantly, it achieves significantly better results when dealing with difficult plagiarism cases where the plagiarized passages are highly modified by replacing most of the words or phrases with synonyms to hide the similarity with the source documents. Efstathios Stamatatos |
CIKM | 1 |
| 2011 | Plagiarism detection using stopword n-gramsabstractIn this paper a novel method for detecting plagiarized passages in document collections is presented. In contrast to previous work in this field that uses content terms to represent documents, the proposed method is based on a small list of stopwords (i.e., very frequent words). We show that stopword n-grams reveal important information for plagiarism detection since they are able to capture syntactic similarities between suspicious and original documents and they can be used to detect the exact plagiarized passage boundaries. Experimental results on a publicly available corpus demonstrate that the performance of the proposed approach is competitive when compared with the best reported results. More importantly, it achieves significantly better results when dealing with difficult plagiarism cases where the plagiarized passages are highly modified and most of the words or phrases have been replaced with synonyms. Efstathios Stamatatos |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2009 | Learning to recognize webpage genres
Ioannis Kanaris, Efstathios Stamatatos |
Inf. Process. Manag. | 2 |
| 2009 | A survey of modern authorship attribution methodsabstractAbstract Authorship attribution supported by statistical or computational methods has a long history starting from the 19th century and is marked by the seminal study of Mosteller and Wallace (1964) on the authorship of the disputed “Federalist Papers.” During the last decade, this scientific field has been developed substantially, taking advantage of research advances in areas such as machine learning, information retrieval, and natural language processing. The plethora of available electronic texts (e.g., e‐mail messages, online forum messages, blogs, source code, etc.) indicates a wide variety of applications of this technology, provided it is able to handle short and noisy text from multiple candidate authors. In this article, a survey of recent advances of the automated approaches to attributing authorship is presented, examining their characteristics for both text representation and text classification. The focus of this survey is on computational requirements and settings rather than on linguistic or literary issues. We also discuss evaluation methodologies and criteria for authorship attribution studies and list open questions that will attract future work in this area. Efstathios Stamatatos |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Author identification: Using text sampling to handle the class imbalance problem
Efstathios Stamatatos |
Inf. Process. Manag. | 1 |