VLDB 2026 Research / reviewers in the wild / expert
Thomas Papastergiou
dblp:173/8953
· DBLP profile ↗
9ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Triangulation-Based Graph Rewiring for Graph Neural NetworksabstractGraph Neural Networks (GNNs) have emerged as the leading paradigm for learning over graph-structured data. However, their performance is limited by issues inherent to graph topology, most notably oversquashing and oversmoothing. Recent advances in graph rewiring aim to mitigate these limitations by modifying the graph topology to promote more effective information propagation. In this work, we introduce TRIGON, a novel framework that constructs enriched, non-planar triangulations by learning to select relevant triangles from multiple graph views. By jointly optimizing triangle selection and downstream classification performance, our method produces a rewired graph with markedly improved structural properties such as reduced diameter, increased spectral gap, and lower effective resistance compared to existing rewiring methods. Empirical results demonstrate that TRIGON outperforms state-of-the-art approaches on node classification tasks across a range of homophilic and heterophilic benchmarks. Hugo Attali, Thomas Papastergiou, Nathalie Pernelle, Fragkiskos D. Malliaros |
CIKM | 2 |
| 2025 | An In-depth Analysis of the Linguistic Characteristics of Science Claims on the Web and their Impact on Fact-checkingabstractWeb claims, seen as assertions shared on the web and eligible for fact-checking, are at the heart of online discourse. They have been studied extensively on a variety of downstream tasks such as fact-checking, claim retrieval, bias detection, argument mining, or viewpoint discovery. On the other hand, claims originating from scientific publications have also been the subject of several downstream NLP tasks. However, research carried out so far has yet to focus on scientific web claims, which are scientific claims made on the web (e.g., on social media and news articles). The process of detecting and fact-checking a claim from the web can be very different depending on whether the claim is scientific or not, thus making it crucial for the developed datasets, methods, and models to make a distinction between the two. With this work, we aim at understanding what makes this distinction necessary, by understanding the linguistic differences between scientific and non-scientific claims on the web, and the impact those differences have on existing downstream tasks. To do so, we manually annotate 1,524 web claims from established benchmarks for fact-checking-related tasks, and we run statistical tests to analyze and compare the linguistic features of each group. We find that scientific claims on the web use more analytical speech, but also use more sentiment-related speech, more expressions of physical motion, and have distinct parts of speech (PoS) and punctuation styles. We also conduct experiments showing that BERT-based language models perform worse on scientific web claims by up to 17 F1 points for several downstream tasks. To understand why, we develop a novel methodology to map predictive tokens of language models to explainable linguistic features and find that language models fail to detect a specific subset of predictive features of scientific web claims. We conclude by stating that language models aimed at studying scientific web claims ought to be trained on scientific web discourse, as opposed to being trained only on generic web discourse or only on scientific text from scientific publications. Salim Hafid, Sebastian Schellhammer, Yavuz Selim Kartal, Thomas Papastergiou, Stefan Dietze, Sandra Bringay, Konstantin Todorov |
ACM Trans. Web | 4 |
| 2024 | A SHAP-based controversy analysis through communities on Twitter
Samy Benslimane, Thomas Papastergiou, Jérôme Azé, Sandra Bringay, Maximilien Servajean, Caroline Mollevi |
World Wide Web (WWW) | 2 |
| 2023 | Explaining controversy through community analysis on TwitterabstractControversy refers to content attracting different point-of-views, as well as positive and negative feedback on a specific event, gathering users into different communities. Research on controversy led to two main categories of works: controversy detection/quantification and controversy explainability. When the former aims to quantify controversy on a topic, the latter aims to understand why a topic is controversial or not. This paper mainly contributes to the controversy explainability. We analyze topic discussions on Twitter from the community perspective to investigate the power of text in classifying tweets into the right community. We propose a SHAP-based pipeline to quantify impactful text features on predictions of three tweet classifiers. We also rely on the use of different text features namely BERT, TF − IDF, and LIWC. The results we obtain from both SHAP plots and statistical analysis show clearly significant impacts of some text features in classifying tweets.It also highlights the relevance of the study as well as the potential benefits of combining text and user interactions to quantify controversy. Samy Benslimane, Thomas Papastergiou, Jérôme Azé, Sandra Bringay, Caroline Mollevi, Maximilien Servajean |
IDEAS | 2 |
| 2023 | Stale TLS Certificates: Investigating Precarious Third-Party Access to Valid TLS KeysabstractCertificate authorities enable TLS server authentication by generating certificates that attest to the mapping between a domain name and a cryptographic keypair, for up to 398 days. This static, name-to-key caching mechanism belies a complex reality: a tangle of dynamic infrastructure involving domains, servers, cryptographic keys, etc. When any of these operations changes, the authentication information in a certificate becomes stale and no longer accurately reflects reality. In this work, we examine the broader phenomenon of certificate invalidation events and discover three classes of security-relevant events that enable a third-party to impersonate a domain outside of their control. Longitudinal measurement of these precarious scenarios reveals that they affect over 15K new domains per day, on average. Unfortunately, modern certificate revocation provides little recourse, so we examine the potential impact of reducing certificate lifetimes (cache duration): shortening the current 398-day limit to 90 days yields a 75% decrease in precarious access to valid TLS keys. Zane Ma, Aaron Faulkenberry, Thomas Papastergiou, Zakir Durumeric, Michael D. Bailey, Angelos D. Keromytis, Fabian Monrose, Manos Antonakakis |
IMC | 3 |
| 2021 | Understanding the Growth and Security Considerations of ECS
Athanasios Kountouras, Panagiotis Kintis, Athanasios Avgetidis, Thomas Papastergiou, Charles Lever, Michalis Polychronakis, Manos Antonakakis |
NDSS | 4 |
| 2020 | IoTFinder: Efficient Large-Scale Identification of IoT Devices via Passive DNS Traffic AnalysisabstractBeing able to enumerate potentially vulnerable IoT devices across the Internet is important, because it allows for assessing global Internet risks and enables network operators to check the hygiene of their own networks. To this end, in this paper we propose IoTFinder, a system for efficient, large-scalepassiveidentification of IoT devices. Specifically, we leverage distributed passive DNS data collection, and develop a machine learning-based system that aims to accurately identify a large variety of IoT devices based solely on theirDNS fingerprints. Our system is independent of whether the devices reside behind a NAT or other middleboxes, or whether they are assigned an IPv4 or IPv6 address. We design IoTFinder as a multi-label classifier, and evaluate its accuracy in several different settings, including computing detection results over a third-party IoT traffic dataset and DNS traffic collected at a US-based ISP hosting more than 40 million clients. The experimental results show that our approach allows for accurately detecting many diverse IoT devices, even when they are hosted behind a NAT and their traffic is “mixed” with traffic generated by other IoT and non-IoT devices hosted in the same local network. Roberto Perdisci, Thomas Papastergiou, Omar Alrawi, Manos Antonakakis |
EuroS&P | 2 |
| 2017 | A distributed proximal gradient descent method for tensor completionabstractIn this paper, we propose a novel proximal method for CANDECOMP/PARAFAC (CP) decomposition to deal with the tensor completion problem. This approach is based on solving local optimization problems, rather than confronting the entire optimization problem at once. In addition, we propose two distributed algorithms for dealing with data of high dimensionality that scale up to tensors of dimension 108×108×108. We show that our proximal method outperforms the Stochastic Gradient Descent (SGD) for CP decomposition in terms of convergence accuracy by a factor of 2.8 and it can efficiently and effectively reconstruct a color image by observing only 10% of its entries. Experimental results show that our distributed methods are scalable in terms of dimensionality, factorization rank, number of machines and perform efficiently in both dense and sparse settings. The proposed distributed proximal approach outperforms existing distributed methods in terms of speed of convergence by a factor of two. Moreover, it can successfully recover a hyperspectral image, by observing 10% of its values. Thomas Papastergiou, Vasileios Megalooikonomou |
IEEE BigData | 1 |
| 2008 | On clustering tree structured data with categorical nature
Basilis Boutsinas, Thomas Papastergiou |
Pattern Recognit. | 2 |