VLDB 2026 Research / reviewers in the wild / expert
Adam Tsakalidis
dblp:150/6942
· DBLP profile ↗
13ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0003-1831-0683ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Information extraction and text analysis · 30% Representation and self-supervised learning · 21% Deep learning architectures and training · 21% | |
| Databases, data mining, and information retrieval
2 papers |
Web and social media mining · 75% Data mining · 25% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 79% Computational social science and digital humanities · 21% |
Topics — the 13 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
change detection |
0.8 | 1 | 2024 | TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024 |
Machine learning › Representation and self-supervised learning › representation learning
dynamic representation learning |
0.8 | 1 | 2024 | TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024 |
Machine learning › Deep learning architectures and training › positional encoding
rotary position embedding |
0.8 | 1 | 2024 | TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024 |
Machine learning › Representation and self-supervised learning › representation learning › sequence representation
temporal representation learning |
0.8 | 1 | 2024 | TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024 |
Medical and health informatics › clinical monitoring › patient monitoring
dementia monitoring |
0.7 | 1 | 2023 | A Digital Language Coherence Marker for Monitoring Dementia · EMNLP 2023 |
Natural language and speech › Language models and text generation › text summarization
opinion summarization |
0.6 | 1 | 2022 | Unsupervised Opinion Summarisation in the Wasserstein Space · EMNLP 2022 |
Natural language and speech › Language models and text generation
text summarization |
0.6 | 1 | 2022 | Unsupervised Opinion Summarisation in the Wasserstein Space · EMNLP 2022 |
Web and social media mining › social media analysis
microblog analysis |
0.5 | 1 | 2021 | Evaluation of Thematic Coherence in Microblogs · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic change detection |
0.4 | 1 | 2020 | Sequential Modelling of the Evolution of Word Representations for Semantic Change Detection · EMNLP (1) 2020 |
Data mining › text mining
text classification |
0.3 | 1 | 2017 | Towards Real-Time, Country-Level Location Classification of Worldwide Tweets · IEEE Trans. Knowl. Data Eng. 2017 |
Web and social media mining › location-based social network analysis
tweet geolocation |
0.3 | 1 | 2017 | Towards Real-Time, Country-Level Location Classification of Worldwide Tweets · IEEE Trans. Knowl. Data Eng. 2017 |
Web and social media mining
social media analysis |
0.1 | 1 | 2017 | Towards Real-Time, Country-Level Location Classification of Worldwide Tweets · IEEE Trans. Knowl. Data Eng. 2017 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.3temporal logical consistency learning · 1.3longitudinal evaluation · 1.3self-supervised learning · 0.8wasserstein distance · 0.6variational autoencoder · 0.6GRU decoder · 0.6word embeddings · 0.4sequential modeling · 0.4temporal generalization · 0.3feature combination · 0.3classifier training · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TempoFormer: A Transformer for Temporally-aware Representations in Change DetectionabstractDynamic representation learning plays a pivotal role in understanding the evolution of linguistic content over time.On this front both context and time dynamics as well as their interplay are of prime importance.Current approaches model context via pre-trained representations, which are typically temporally agnostic.Previous work on modelling context and temporal dynamics has used recurrent methods, which are slow and prone to overfitting.Here we introduce TempoFormer, the first task-agnostic transformer-based and temporally-aware model for dynamic representation learning.Our approach is jointly trained on inter and intra context dynamics and introduces a novel temporal variation of rotary positional embeddings.The architecture is flexible and can be used as the temporal representation foundation of other models or applied to different transformer-based architectures.We show new SOTA performance on three different realtime change detection tasks. Talia Tseriotou, Adam Tsakalidis, Maria Liakata |
EMNLP | 2 |
| 2023 | Creation and evaluation of timelines for longitudinal user postsabstractThere is increasing interest to work with user generated content in social media, especially textual posts over time.Currently there is no consistent way of segmenting user posts into timelines in a meaningful way that improves the quality and cost of manual annotation.Here we propose a set of methods for segmenting longitudinal user posts into timelines likely to contain interesting moments of change in a user's behaviour, based on their online posting activity.We also propose a novel framework for evaluating timelines and show its applicability in the context of two different social media datasets.Finally, we present a discussion of the linguistic content of highly ranked timelines.1 Anthony Hills, Adam Tsakalidis, Federico Nanni, Ioannis Zachos, Maria Liakata |
EACL | 2 |
| 2023 | A Digital Language Coherence Marker for Monitoring DementiaabstractThe use of spontaneous language to derive appropriate digital markers has become an emergent, promising and non-intrusive method to diagnose and monitor dementia.Here we propose methods to capture language coherence as a cost-effective, human-interpretable digital marker for monitoring cognitive changes in people with dementia.We introduce a novel task to learn the temporal logical consistency of utterances in short transcribed narratives and investigate a range of neural approaches.We compare such language coherence patterns between people with dementia and healthy controls and conduct a longitudinal evaluation against three clinical bio-markers to investigate the reliability of our proposed digital coherence marker.The coherence marker shows a significant difference between people with mild cognitive impairment, those with Alzheimer's Disease and healthy controls.Moreover our analysis shows high association between the coherence marker and the clinical bio-markers as well as generalisability potential to other related conditions. Dimitris Gkoumas, Adam Tsakalidis, Maria Liakata |
EMNLP | 2 |
| 2022 | Identifying Moments of Change from Longitudinal User TextabstractAdam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim, Jiayu Song, Maria Liakata. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Adam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim, Jiayu Song, Maria Liakata |
ACL (1) | 1 |
| 2022 | Unsupervised Opinion Summarisation in the Wasserstein SpaceabstractOpinion summarisation synthesises opinions expressed in a group of documents discussing the same topic to produce a single summary.Recent work has looked at opinion summarisation of clusters of social media posts.Such posts are noisy and have unpredictable structure, posing additional challenges for the construction of the summary distribution and the preservation of meaning compared to online reviews, which has been so far the focus of opinion summarisation.To address these challenges we present WassOS, an unsupervised abstractive summarization model which makes use of the Wasserstein distance.A Variational Autoencoder is used to get the distribution of documents/posts, and the distributions are disentangled into separate semantic and syntactic spaces.The summary distribution is obtained using the Wasserstein barycenter of the semantic and syntactic distributions.A latent variable sampled from the summary distribution is fed into a GRU decoder with a transformer layer to produce the final summary.Our experiments on multiple datasets including Twitter clusters, Reddit threads, and reviews show that WassOS almost always outperforms the state-of-the-art on ROUGE metrics and consistently produces the best summaries with respect to meaning preservation according to human evaluations. Jiayu Song, Iman Munire Bilal, Adam Tsakalidis, Rob Procter, Maria Liakata |
EMNLP | 3 |
| 2022 | Template-based Abstractive Microblog Opinion SummarisationabstractAbstract We introduce the task of microblog opinion summarization (MOS) and share a dataset of 3100 gold-standard opinion summaries to facilitate research in this domain. The dataset contains summaries of tweets spanning a 2-year period and covers more topics than any other public Twitter summarization dataset. Summaries are abstractive in nature and have been created by journalists skilled in summarizing news articles following a template separating factual information (main story) from author opinions. Our method differs from previous work on generating gold-standard summaries from social media, which usually involves selecting representative posts and thus favors extractive summarization models. To showcase the dataset’s utility and challenges, we benchmark a range of abstractive and extractive state-of-the-art summarization models and achieve good performance, with the former outperforming the latter. We also show that fine-tuning is necessary to improve performance and investigate the benefits of using different sample sizes. Iman Munire Bilal, Bo Wang 0034, Adam Tsakalidis, Dong Nguyen 0002, Rob Procter, Maria Liakata |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Evaluation of Thematic Coherence in MicroblogsabstractIman Munire Bilal, Bo Wang, Maria Liakata, Rob Procter, Adam Tsakalidis. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Iman Munire Bilal, Bo Wang 0034, Maria Liakata, Rob Procter, Adam Tsakalidis |
ACL/IJCNLP (1) | 5 |
| 2020 | Sequential Modelling of the Evolution of Word Representations for Semantic Change DetectionabstractSemantic change detection concerns the task of identifying words whose meaning has changed over time.Current state-of-the-art approaches operating on neural embeddings detect the level of semantic change in a word by comparing its vector representation in two distinct time periods, without considering its evolution through time.In this work, we propose three variants of sequential models for detecting semantically shifted words, effectively accounting for the changes in the word representations over time.Through extensive experimentation under various settings with synthetic and real data we showcase the importance of sequential modelling of word vectors through time for semantic change detection.Finally, we compare different approaches in a quantitative manner, demonstrating that temporal modelling of word representations yields a clear-cut advantage in performance. Adam Tsakalidis, Maria Liakata |
EMNLP (1) | 1 |
| 2018 | Nowcasting the Stance of Social Media Users in a Sudden Vote: The Case of the Greek ReferendumabstractModelling user voting intention in social media is an important research area, with applications in analysing electorate behaviour, online political campaigning and advertising. Previous approaches mainly focus on predicting national general elections, which are regularly scheduled and where data of past results and opinion polls are available. However, there is no evidence of how such models would perform during a sudden vote under time-constrained circumstances. That poses a more challenging task compared to traditional elections, due to its spontaneous nature. In this paper, we focus on the 2015 Greek bailout referendum, aiming to nowcast on a daily basis the voting intention of 2,197 Twitter users. We propose a semi-supervised multiple convolution kernel learning approach, leveraging temporally sensitive text and network information. Our evaluation under a real-time simulation framework demonstrates the effectiveness and robustness of our approach against competitive baselines, achieving a significant 20% increase in F-score compared to solely text-based models. Adam Tsakalidis, Nikolaos Aletras, Alexandra I. Cristea, Maria Liakata |
CIKM | 1 |
| 2018 | Can We Assess Mental Health Through Social Media and Smart Devices? Addressing Bias in Methodology and Evaluation
Adam Tsakalidis, Maria Liakata, Theodoros Damoulas, Alexandra I. Cristea |
ECML/PKDD (3) | 1 |
| 2017 | Towards Real-Time, Country-Level Location Classification of Worldwide TweetsabstractThe increase of interest in using social media as a source for research has motivated tackling the challenge of automatically geolocating tweets, given the lack of explicit location information in the majority of tweets. In contrast to much previous work that has focused on location classification of tweets restricted to a specific country, here we undertake the task in a broader context by classifying global tweets at the country level, which is so far unexplored in a real-time scenario. We analyze the extent to which a tweet's country of origin can be determined by making use of eight tweet-inherent features for classification. Furthermore, we use two datasets, collected a year apart from each other, to analyze the extent to which a model trained from historical tweets can still be leveraged for classification of new tweets. With classification experiments on all 217 countries in our datasets, as well as on the top 25 countries, we offer some insights into the best use of tweet-inherent features for an accurate country-level classification of tweets. We find that the use of a single feature, such as the use of tweet content alone-the most widely used feature in previous work-leaves much to be desired. Choosing an appropriate combination of both tweet content and metadata can actually lead to substantial improvements of between 20 and 50 percent. We observe that tweet content, the user's self-reported location and the user's real name, all of which are inherent in a tweet and available in a real-time scenario, are particularly useful to determine the country of origin. We also experiment on the applicability of a model trained on historical tweets to classify new tweets, finding that the choice of a particular combination of features whose utility does not fade over time can actually lead to comparable performance, avoiding the need to retrain. However, the difficulty of achieving accurate classification increases slightly for countries with multiple commonalities, especially for English and Spanish speaking countries. Arkaitz Zubiaga, Alexander Voß, Rob Procter, Maria Liakata, Bo Wang 0034, Adam Tsakalidis |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2016 | Combining Heterogeneous User Generated Data to Sense Well-beingabstractIn this paper we address a new problem of predicting affect and well-being scales in a real-world setting of heterogeneous, longitudinal and non-synchronous textual as well as non-linguistic data that can be harvested from on-line media and mobile phones. We describe the method for collecting the heterogeneous longitudinal data, how features are extracted to address missing information and differences in temporal alignment, and how the latter are combined to yield promising predictions of affect and well-being on the basis of widely used psychological scales. We achieve a coefficient of determination (R^2) of 0.71-0.76 and a correlation coefficient of 0.68-0.87 which is higher than the state-of-the art in equivalent multi-modal tasks for affect. Adam Tsakalidis, Maria Liakata, Theodoros Damoulas, Brigitte Jellinek, Weisi Guo, Alexandra I. Cristea |
COLING | 1 |
| 2014 | An Ensemble Model for Cross-Domain Polarity Classification on Twitter
Adam Tsakalidis, Symeon Papadopoulos, Ioannis Kompatsiaris |
WISE (2) | 1 |