Adam Tsakalidis

dblp:150/6942 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0003-1831-0683ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Information extraction and text analysis · 30% Representation and self-supervised learning · 21% Deep learning architectures and training · 21%
Databases, data mining, and information retrieval
2 papers
Web and social media mining · 75% Data mining · 25%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 79% Computational social science and digital humanities · 21%

Topics — the 13 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
change detection
0.812024
TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024
Machine learning › Representation and self-supervised learning › representation learning
dynamic representation learning
0.812024
TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024
Machine learning › Deep learning architectures and training › positional encoding
rotary position embedding
0.812024
TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024
Machine learning › Representation and self-supervised learning › representation learning › sequence representation
temporal representation learning
0.812024
TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
TempoFormer: A Transformer for Temporally-aware Representations in Change Detection · EMNLP 2024
Medical and health informatics › clinical monitoring › patient monitoring
dementia monitoring
0.712023
A Digital Language Coherence Marker for Monitoring Dementia · EMNLP 2023
Natural language and speech › Language models and text generation › text summarization
opinion summarization
0.612022
Unsupervised Opinion Summarisation in the Wasserstein Space · EMNLP 2022
Natural language and speech › Language models and text generation
text summarization
0.612022
Unsupervised Opinion Summarisation in the Wasserstein Space · EMNLP 2022
Web and social media mining › social media analysis
microblog analysis
0.512021
Evaluation of Thematic Coherence in Microblogs · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic change detection
0.412020
Sequential Modelling of the Evolution of Word Representations for Semantic Change Detection · EMNLP (1) 2020
Data mining › text mining
text classification
0.312017
Towards Real-Time, Country-Level Location Classification of Worldwide Tweets · IEEE Trans. Knowl. Data Eng. 2017
Web and social media mining › location-based social network analysis
tweet geolocation
0.312017
Towards Real-Time, Country-Level Location Classification of Worldwide Tweets · IEEE Trans. Knowl. Data Eng. 2017
Web and social media mining
social media analysis
0.112017
Towards Real-Time, Country-Level Location Classification of Worldwide Tweets · IEEE Trans. Knowl. Data Eng. 2017

Methods — techniques the papers use, named apart from their topics

transformer · 1.3temporal logical consistency learning · 1.3longitudinal evaluation · 1.3self-supervised learning · 0.8wasserstein distance · 0.6variational autoencoder · 0.6GRU decoder · 0.6word embeddings · 0.4sequential modeling · 0.4temporal generalization · 0.3feature combination · 0.3classifier training · 0.3
YearPublicationVenuePosition
2024 TempoFormer: A Transformer for Temporally-aware Representations in Change Detection
abstract
Dynamic representation learning plays a pivotal role in understanding the evolution of linguistic content over time.On this front both context and time dynamics as well as their interplay are of prime importance.Current approaches model context via pre-trained representations, which are typically temporally agnostic.Previous work on modelling context and temporal dynamics has used recurrent methods, which are slow and prone to overfitting.Here we introduce TempoFormer, the first task-agnostic transformer-based and temporally-aware model for dynamic representation learning.Our approach is jointly trained on inter and intra context dynamics and introduces a novel temporal variation of rotary positional embeddings.The architecture is flexible and can be used as the temporal representation foundation of other models or applied to different transformer-based architectures.We show new SOTA performance on three different realtime change detection tasks.
Talia Tseriotou, Adam Tsakalidis, Maria Liakata
EMNLP2
2023 Creation and evaluation of timelines for longitudinal user posts
abstract
There is increasing interest to work with user generated content in social media, especially textual posts over time.Currently there is no consistent way of segmenting user posts into timelines in a meaningful way that improves the quality and cost of manual annotation.Here we propose a set of methods for segmenting longitudinal user posts into timelines likely to contain interesting moments of change in a user's behaviour, based on their online posting activity.We also propose a novel framework for evaluating timelines and show its applicability in the context of two different social media datasets.Finally, we present a discussion of the linguistic content of highly ranked timelines.1
Anthony Hills, Adam Tsakalidis, Federico Nanni, Ioannis Zachos, Maria Liakata
EACL2
2023 A Digital Language Coherence Marker for Monitoring Dementia
abstract
The use of spontaneous language to derive appropriate digital markers has become an emergent, promising and non-intrusive method to diagnose and monitor dementia.Here we propose methods to capture language coherence as a cost-effective, human-interpretable digital marker for monitoring cognitive changes in people with dementia.We introduce a novel task to learn the temporal logical consistency of utterances in short transcribed narratives and investigate a range of neural approaches.We compare such language coherence patterns between people with dementia and healthy controls and conduct a longitudinal evaluation against three clinical bio-markers to investigate the reliability of our proposed digital coherence marker.The coherence marker shows a significant difference between people with mild cognitive impairment, those with Alzheimer's Disease and healthy controls.Moreover our analysis shows high association between the coherence marker and the clinical bio-markers as well as generalisability potential to other related conditions.
Dimitris Gkoumas, Adam Tsakalidis, Maria Liakata
EMNLP2
2022 Identifying Moments of Change from Longitudinal User Text
abstract
Adam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim, Jiayu Song, Maria Liakata. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Adam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim, Jiayu Song, Maria Liakata
ACL (1)1
2022 Unsupervised Opinion Summarisation in the Wasserstein Space
abstract
Opinion summarisation synthesises opinions expressed in a group of documents discussing the same topic to produce a single summary.Recent work has looked at opinion summarisation of clusters of social media posts.Such posts are noisy and have unpredictable structure, posing additional challenges for the construction of the summary distribution and the preservation of meaning compared to online reviews, which has been so far the focus of opinion summarisation.To address these challenges we present WassOS, an unsupervised abstractive summarization model which makes use of the Wasserstein distance.A Variational Autoencoder is used to get the distribution of documents/posts, and the distributions are disentangled into separate semantic and syntactic spaces.The summary distribution is obtained using the Wasserstein barycenter of the semantic and syntactic distributions.A latent variable sampled from the summary distribution is fed into a GRU decoder with a transformer layer to produce the final summary.Our experiments on multiple datasets including Twitter clusters, Reddit threads, and reviews show that WassOS almost always outperforms the state-of-the-art on ROUGE metrics and consistently produces the best summaries with respect to meaning preservation according to human evaluations.
Jiayu Song, Iman Munire Bilal, Adam Tsakalidis, Rob Procter, Maria Liakata
EMNLP3
2022 Template-based Abstractive Microblog Opinion Summarisation
abstract
Abstract We introduce the task of microblog opinion summarization (MOS) and share a dataset of 3100 gold-standard opinion summaries to facilitate research in this domain. The dataset contains summaries of tweets spanning a 2-year period and covers more topics than any other public Twitter summarization dataset. Summaries are abstractive in nature and have been created by journalists skilled in summarizing news articles following a template separating factual information (main story) from author opinions. Our method differs from previous work on generating gold-standard summaries from social media, which usually involves selecting representative posts and thus favors extractive summarization models. To showcase the dataset’s utility and challenges, we benchmark a range of abstractive and extractive state-of-the-art summarization models and achieve good performance, with the former outperforming the latter. We also show that fine-tuning is necessary to improve performance and investigate the benefits of using different sample sizes.
Iman Munire Bilal, Bo Wang 0034, Adam Tsakalidis, Dong Nguyen 0002, Rob Procter, Maria Liakata
Trans. Assoc. Comput. Linguistics3
2021 Evaluation of Thematic Coherence in Microblogs
abstract
Iman Munire Bilal, Bo Wang, Maria Liakata, Rob Procter, Adam Tsakalidis. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Iman Munire Bilal, Bo Wang 0034, Maria Liakata, Rob Procter, Adam Tsakalidis
ACL/IJCNLP (1)5
2020 Sequential Modelling of the Evolution of Word Representations for Semantic Change Detection
abstract
Semantic change detection concerns the task of identifying words whose meaning has changed over time.Current state-of-the-art approaches operating on neural embeddings detect the level of semantic change in a word by comparing its vector representation in two distinct time periods, without considering its evolution through time.In this work, we propose three variants of sequential models for detecting semantically shifted words, effectively accounting for the changes in the word representations over time.Through extensive experimentation under various settings with synthetic and real data we showcase the importance of sequential modelling of word vectors through time for semantic change detection.Finally, we compare different approaches in a quantitative manner, demonstrating that temporal modelling of word representations yields a clear-cut advantage in performance.
Adam Tsakalidis, Maria Liakata
EMNLP (1)1
2018 Nowcasting the Stance of Social Media Users in a Sudden Vote: The Case of the Greek Referendum
abstract
Modelling user voting intention in social media is an important research area, with applications in analysing electorate behaviour, online political campaigning and advertising. Previous approaches mainly focus on predicting national general elections, which are regularly scheduled and where data of past results and opinion polls are available. However, there is no evidence of how such models would perform during a sudden vote under time-constrained circumstances. That poses a more challenging task compared to traditional elections, due to its spontaneous nature. In this paper, we focus on the 2015 Greek bailout referendum, aiming to nowcast on a daily basis the voting intention of 2,197 Twitter users. We propose a semi-supervised multiple convolution kernel learning approach, leveraging temporally sensitive text and network information. Our evaluation under a real-time simulation framework demonstrates the effectiveness and robustness of our approach against competitive baselines, achieving a significant 20% increase in F-score compared to solely text-based models.
Adam Tsakalidis, Nikolaos Aletras, Alexandra I. Cristea, Maria Liakata
CIKM1
2018 Can We Assess Mental Health Through Social Media and Smart Devices? Addressing Bias in Methodology and Evaluation
Adam Tsakalidis, Maria Liakata, Theodoros Damoulas, Alexandra I. Cristea
ECML/PKDD (3)1
2017 Towards Real-Time, Country-Level Location Classification of Worldwide Tweets
abstract
The increase of interest in using social media as a source for research has motivated tackling the challenge of automatically geolocating tweets, given the lack of explicit location information in the majority of tweets. In contrast to much previous work that has focused on location classification of tweets restricted to a specific country, here we undertake the task in a broader context by classifying global tweets at the country level, which is so far unexplored in a real-time scenario. We analyze the extent to which a tweet's country of origin can be determined by making use of eight tweet-inherent features for classification. Furthermore, we use two datasets, collected a year apart from each other, to analyze the extent to which a model trained from historical tweets can still be leveraged for classification of new tweets. With classification experiments on all 217 countries in our datasets, as well as on the top 25 countries, we offer some insights into the best use of tweet-inherent features for an accurate country-level classification of tweets. We find that the use of a single feature, such as the use of tweet content alone-the most widely used feature in previous work-leaves much to be desired. Choosing an appropriate combination of both tweet content and metadata can actually lead to substantial improvements of between 20 and 50 percent. We observe that tweet content, the user's self-reported location and the user's real name, all of which are inherent in a tweet and available in a real-time scenario, are particularly useful to determine the country of origin. We also experiment on the applicability of a model trained on historical tweets to classify new tweets, finding that the choice of a particular combination of features whose utility does not fade over time can actually lead to comparable performance, avoiding the need to retrain. However, the difficulty of achieving accurate classification increases slightly for countries with multiple commonalities, especially for English and Spanish speaking countries.
Arkaitz Zubiaga, Alexander Voß, Rob Procter, Maria Liakata, Bo Wang 0034, Adam Tsakalidis
IEEE Trans. Knowl. Data Eng.6
2016 Combining Heterogeneous User Generated Data to Sense Well-being
abstract
In this paper we address a new problem of predicting affect and well-being scales in a real-world setting of heterogeneous, longitudinal and non-synchronous textual as well as non-linguistic data that can be harvested from on-line media and mobile phones. We describe the method for collecting the heterogeneous longitudinal data, how features are extracted to address missing information and differences in temporal alignment, and how the latter are combined to yield promising predictions of affect and well-being on the basis of widely used psychological scales. We achieve a coefficient of determination (R^2) of 0.71-0.76 and a correlation coefficient of 0.68-0.87 which is higher than the state-of-the art in equivalent multi-modal tasks for affect.
Adam Tsakalidis, Maria Liakata, Theodoros Damoulas, Brigitte Jellinek, Weisi Guo, Alexandra I. Cristea
COLING1
2014 An Ensemble Model for Cross-Domain Polarity Classification on Twitter
Adam Tsakalidis, Symeon Papadopoulos, Ioannis Kompatsiaris
WISE (2)1