VLDB 2026 Research / reviewers in the wild / expert
Stergios Chatzikyriakidis
dblp:133/2452
· DBLP profile ↗
5ranked-venue papers
2as first author
2since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRDD+: An Extended Greek Dialectal Dataset with Cross-Architecture Fine-tuning EvaluationabstractWe present an extended Greek Dialectal Dataset (GRDD+) 1that complements the existing GRDD dataset with more data from Cretan, Cypriot, Pontic and Northern Greek, while we add six new varieties: Greco-Corsican, Griko (Southern Italian Greek), Maniot, Heptanesian, Tsakonian, and Katharevusa Greek. The result is a dataset with total size 6,374,939 words and 10 varieties. This is the first dataset with such variation and size to date. We conduct a number of fine-tuning experiments to see the effect of good quality dialectal data on a number of LLMs. We fine-tune three model architectures (Llama-3-8B, Llama-3.1-8B, Krikri-8B) and compare the results to frontier models (Claude-3.7-Sonnet, Gemini-2.5, ChatGPT-5). Stergios Chatzikyriakidis, Dimitrios Papadakis, Sevasti-Ioanna Papaioannou, Erofili Psaltaki |
LREC | 1 |
| 2026 | Mathematical structures in natural language semantics: Zawadowski's contribution to linguistics
Stergios Chatzikyriakidis, Robin Cooper |
Math. Struct. Comput. Sci. | 1 |
| 2020 | Improving the Precision of Natural Textual Entailment Problem DatasetsabstractIn this paper, we propose a method to modify natural textual entailment problem datasets so that they better reflect a more precise notion of entailment. We apply this method to a subset of the Recognizing Textual Entailment datasets. We thus obtain a new corpus of entailment problems, which has the following three characteristics: 1. it is precise (does not leave out implicit hypotheses) 2. it is based on “real-world” texts (i.e. most of the premises were written for purposes other than testing textual entailment). 3. its size is 150. Broadly, the method that we employ is to make any missing hypotheses explicit using a crowd of experts. We discuss the relevance of our method in improving existing NLI datasets to be more fit for precise reasoning and we argue that this corpus can be the basis a first step towards wide-coverage testing of precise natural-language inference systems. Jean-Philippe Bernardy, Stergios Chatzikyriakidis |
LREC | 2 |
| 2019 | What Kind of Natural Language Inference are NLP Systems Learning: Is this Enough?abstractIn this paper, we look at Natural Language Inference, arguing that the notion of inference the current NLP systems are learning is much narrower compared to the range of inference patterns found in human reasoning. We take a look at the history and the nature of creating datasets for NLI. We discuss the datasets that are mainly used today for the relevant tasks and show why those are not enough to generalize to other reasoning tasks, e.g. logical and legal reasoning, or reasoning in dialogue settings. We then proceed to propose ways in which this can be remedied, effectively producing more realistic datasets for NLI. Lastly, we argue that the NLP community could have been too hasty to altogether dismiss symbolic approaches in the study of NLI, given that these might still be relevant for more fine-grained cases of reasoning. As such, we argue for a more pluralistic take on tackling NLI, favoring hybrid rather than non-hybrid approaches. Jean-Philippe Bernardy, Stergios Chatzikyriakidis |
ICAART (2) | 2 |
| 2018 | Shami: A Corpus of Levantine Arabic Dialects
Chatrine Qwaider, Motaz Saad, Stergios Chatzikyriakidis, Simon Dobnik |
LREC | 3 |