George Tsatsaronis 0001

dblp:14/2922 · also Georgios Tsatsaronis, Giorgios Tsatsaronis · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
7since 2021 · last 2025
0000-0003-2116-2933ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 13 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Extracting, Detecting, and Generating Research Questions for Scientific Articles
abstract
The volume of academic articles is increasing rapidly, reflecting the growing emphasis on research and scholarship across different science disciplines. This rapid growth necessitates the development of tools for more efficient and rapid understanding of these articles. Clear and well-defined Research Questions (RQs) in research articles can help guide scholarly inquiries. However, many academic studies lack a proper definition of RQs in their articles. This research addresses this gap by presenting a comprehensive framework for the systematic extraction, detection, and generation of RQs from scientific articles. The extraction component uses a set of regular expressions to identify articles containing well-defined RQs. The detection component aims to identify more complex RQs in articles, beyond those captured by the rule-based extraction method. The RQ generation focuses on creating RQs for articles that lack them. We integrate all these components to build a pipeline to extract RQs or generate them based on the articles’ full text. We evaluate the performance of the designed pipeline on a set of metrics designed to assess the quality of RQs. Our results indicate that the proposed pipeline can reliably detect RQs and generate high-quality ones.
Sina Taslimi, Artemis Çapari, Hosein Azarbonyad, Zi Long Zhu, Zubair Afzal, Evangelos Kanoulas, George Tsatsaronis 0001
COLING7
2025 Question-Answer Extraction from Scientific Articles Using Knowledge Graphs and Large Language Models
abstract
When deciding to read an article or incorporate it into their research, scholars often seek to quickly identify and understand its main ideas.In this paper, we aim to extract these key concepts and contributions from scientific articles in the form of Question and Answer (QA) pairs.We propose two distinct approaches for generating QAs.The first approach involves selecting salient paragraphs, using a Large Language Model (LLM) to generate questions, ranking these questions by the likelihood of obtaining meaningful answers, and subsequently generating answers.This method relies exclusively on the content of the articles.However, assessing an article's novelty typically requires comparison with the existing literature.Therefore, our second approach leverages a Knowledge Graph (KG) for QA generation.We construct a KG by fine-tuning an Entity Relationship (ER) extraction model on scientific articles and using it to build the graph.We then employ a salient triplet extraction method to select the most pertinent ERs per article, utilizing metrics such as the centrality of entities based on a triplet TF-IDF-like measure.This measure assesses the saliency of a triplet based on its importance within the article compared to its prevalence in the literature.For evaluation, we generate QAs using both approaches and have them assessed by Subject Matter Experts (SMEs) through a set of predefined metrics to evaluate the quality of both questions and answers.Our evaluations demonstrate that the KG-based approach effectively captures the main ideas discussed in the articles.Furthermore, our findings indicate that fine-tuning the ER extraction model on our scientific corpus is crucial for extracting high-quality triplets from such documents.
Hosein Azarbonyad, Zi Long Zhu, Georgios Cheirmpos, Zubair Afzal, Vikrant Yadav, George Tsatsaronis 0001
SIGIR6
2024 Scalable Patent Classification with Aggregated Multi-View Ranking
abstract
Automated patent classification typically involves assigning labels to a patent from a taxonomy, using multi-class multi-label classification models. However, classification-based models face challenges in scaling to large numbers of labels, struggle with generalizing to new labels, and fail to effectively utilize the rich information and multiple views of patents and labels. In this work, we propose a multi-view ranking-based method to address these limitations. Our method consists of four ranking-based models that incorporate different views of patents and a meta-model that aggregates and re-ranks the candidate labels given by the four ranking models. We compared our approach against the state-of-the-art baselines on two publicly available patent classification datasets, USPTO-2M and CLEF-IP-2011. We demonstrate that our approach can alleviate the aforementioned limitations and achieve a new state-of-the-art performance by a significant margin.
Dan Li 0015, Vikrant Yadav, Zi Long Zhu, Maziar Moradi Fard, Zubair Afzal, George Tsatsaronis 0001
LREC/COLING6
2024 ScienceDirect Topic Pages: A Knowledge Base of Scientific Concepts Across Various Science Domains
Artemis Çapari, Hosein Azarbonyad, George Tsatsaronis 0001, Zubair Afzal, Judson Dunham
SIGIR3
2023 Generating Topic Pages for Scientific Concepts Using Scientific Publications
Hosein Azarbonyad, Zubair Afzal, George Tsatsaronis 0001
ECIR (2)3
2022 Find the Funding: Entity Linking with Incomplete Funding Knowledge Bases
abstract
Automatic extraction of funding information from academic articles adds significant value to industry and research communities, including tracking research outcomes by funding organizations, profiling researchers and universities based on the received funding, and supporting open access policies. Two major challenges of identifying and linking funding entities are: (i) sparse graph structure of the Knowledge Base (KB), which makes the commonly used graph-based entity linking approaches suboptimal for the funding domain, (ii) missing entities in KB, which (unlike recent zero-shot approaches) requires marking entity mentions without KB entries as NIL. We propose an entity linking model that can perform NIL prediction and overcome data scarcity issues in a time and data-efficient manner. Our model builds on a transformer-based mention detection and a bi-encoder model to perform entity linking. We show that our model outperforms strong existing baselines.
Gizem Aydin, Seyed Amin Tabatabaei, George Tsatsaronis 0001, Faegheh Hasibi
COLING3
2021 Is your document novel? Let attention guide you. An attention-based model for document-level novelty detection
abstract
Abstract Detecting, whether a document contains sufficient new information to be deemed as novel , is of immense significance in this age of data duplication. Existing techniques for document-level novelty detection mostly perform at the lexical level and are unable to address the semantic-level redundancy. These techniques usually rely on handcrafted features extracted from the documents in a rule-based or traditional feature-based machine learning setup. Here, we present an effective approach based on neural attention mechanism to detect document-level novelty without any manual feature engineering. We contend that the simple alignment of texts between the source and target document(s) could identify the state of novelty of a target document. Our deep neural architecture elicits inference knowledge from a large-scale natural language inference dataset, which proves crucial to the novelty detection task. Our approach is effective and outperforms the standard baselines and recent work on document-level novelty detection by a margin of $\sim$ 3% in terms of accuracy.
Tirthankar Ghosal, Vignesh Edithal, Asif Ekbal, Pushpak Bhattacharyya, Srinivasa Satya Sameer Kumar Chivukula, George Tsatsaronis 0001
Nat. Lang. Eng.6
2019 EigenSent: Spectral sentence embeddings using higher-order Dynamic Mode Decomposition
abstract
Distributed representation of words, or word embeddings, have motivated methods for calculating semantic representations of word sequences such as phrases, sentences and paragraphs.Most of the existing methods to do so either use algorithms to learn such representations, or improve on calculating weighted averages of the word vectors.In this work, we experiment with spectral methods of signal representation and summarization as mechanisms for constructing such word-sequence embeddings in an unsupervised fashion.In particular, we explore an algorithm rooted in fluid-dynamics, known as higher-order Dynamic Mode Decomposition, which is designed to capture the eigenfrequencies, and hence the fundamental transition dynamics, of periodic and quasi-periodic systems.It is empirically observed that this approach, which we call EigenSent, can summarize transitions in a sequence of words and generate an embedding that can represent well the sequence itself.To the best of the authors' knowledge, this is the first application of a spectral decomposition and signal summarization technique on text, to create sentence embeddings.We test the efficacy of this algorithm in creating sentence embeddings on three public datasets, where it performs appreciably well.Moreover it is also shown that, due to the positive combination of their complementary properties, concatenating the embeddings generated by EigenSent with simple word vector averaging achieves state-of-the-art results.
Subhradeep Kayal, George Tsatsaronis 0001
ACL (1)2
2018 Novelty Goes Deep. A Deep Neural Solution To Document Level Novelty Detection
abstract
The rapid growth of documents across the web has necessitated finding means of discarding redundant documents and retaining novel ones. Capturing redundancy is challenging as it may involve investigating at a deep semantic level. Techniques for detecting such semantic redundancy at the document level are scarce. In this work we propose a deep Convolutional Neural Networks (CNN) based model to classify a document as novel or redundant with respect to a set of relevant documents already seen by the system. The system is simple and do not require any manual feature engineering. Our novel scheme encodes relevant and relative information from both source and target texts to generate an intermediate representation which we coin as the Relative Document Vector (RDV). The proposed method outperforms the existing state-of-the-art on a document-level novelty detection dataset by a margin of ∼5% in terms of accuracy. We further demonstrate the effectiveness of our approach on a standard paraphrase detection dataset where paraphrased passages closely resemble to semantically redundant documents.
Tirthankar Ghosal, Vignesh Edithal, Asif Ekbal, Pushpak Bhattacharyya, George Tsatsaronis 0001, Srinivasa Satya Sameer Kumar Chivukula
COLING5
2017 Detecting Aggressive Behavior in Discussion Threads Using Text Mining
Filippos Ventirozos, Iraklis Varlamis, George Tsatsaronis 0001
CICLing (2)3
2016 Learning Domain Labels Using Conceptual Fingerprints: An In-Use Case Study in the Neurology Domain
Zubair Afzal, George Tsatsaronis 0001, Marius A. Doornenbal, Pascal Coupet, Michelle Gregory
EKAW2
2015 An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
abstract
BACKGROUND: This article provides an overview of the first BIOASQ challenge, a competition on large-scale biomedical semantic indexing and question answering (QA), which took place between March and September 2013. BIOASQ assesses the ability of systems to semantically index very large numbers of biomedical scientific articles, and to return concise and user-understandable answers to given natural language questions by combining information from biomedical articles and ontologies. RESULTS: The 2013 BIOASQ competition comprised two tasks, Task 1a and Task 1b. In Task 1a participants were asked to automatically annotate new PUBMED documents with MESH headings. Twelve teams participated in Task 1a, with a total of 46 system runs submitted, and one of the teams performing consistently better than the MTI indexer used by NLM to suggest MESH headings to curators. Task 1b used benchmark datasets containing 29 development and 282 test English questions, along with gold standard (reference) answers, prepared by a team of biomedical experts from around Europe and participants had to automatically produce answers. Three teams participated in Task 1b, with 11 system runs. The BIOASQ infrastructure, including benchmark datasets, evaluation mechanisms, and the results of the participants and baseline methods, is publicly available. CONCLUSIONS: A publicly available evaluation infrastructure for biomedical semantic indexing and QA has been developed, which includes benchmark datasets, and can be used to evaluate systems that: assign MESH headings to published articles or to English questions; retrieve relevant RDF triples from ontologies, relevant articles and snippets from PUBMED Central; produce "exact" and paragraph-sized "ideal" answers (summaries). The results of the systems that participated in the 2013 BIOASQ competition are promising. In Task 1a one of the systems performed consistently better from the NLM's MTI indexer. In Task 1b the systems received high scores in the manual evaluation of the "ideal" answers; hence, they produced high quality summaries as answers. Overall, BIOASQ helped obtain a unified view of how techniques from text classification, semantic indexing, document and passage retrieval, question answering, and text summarization can be combined to allow biomedical experts to obtain concise, user-understandable answers to questions reflecting their real information needs.
George Tsatsaronis 0001, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry Artières, Axel-Cyrille Ngonga Ngomo, Norman Heino, Éric Gaussier, Liliana Barrio-Alvers, Michael Schroeder 0001, Ion Androutsopoulos, Georgios Paliouras
BMC Bioinform.1
2014 PYTHIA: Employing Lexical and Semantic Features for Sentiment Analysis
Ioannis Manousos Katakis, Iraklis Varlamis, George Tsatsaronis 0001
ECML/PKDD (3)3
2013 Temporal Classifiers for Predicting the Expansion of Medical Subject Headings
George Tsatsaronis 0001, Iraklis Varlamis, Nattiya Kanhabua, Kjetil Nørvåg
CICLing (1)1
2013 Semantic smoothing for text clustering
Jamal Abdul Nasir, Iraklis Varlamis, Asim Karim, George Tsatsaronis 0001
Knowl. Based Syst.4
2012 SemaFor: semantic document indexing using semantic forests
abstract
Traditional document indexing techniques store documents using easily accessible representations, such as inverted indices, which can efficiently scale for large document sets. These structures offer scalable and efficient solutions in text document management tasks, though, they omit the cornerstone of the documents' purpose: meaning. They also neglect semantic relations that bind terms into coherent fragments of text that convey messages. When semantic representations are employed, the documents are mapped to the space of concepts and the similarity measures are adapted appropriately to better fit the retrieval tasks. However, these methods can be slow both at indexing and retrieval time. In this paper we propose SemaFor, an indexing algorithm for text documents, which uses semantic spanning forests constructed from lexical resources, like Wikipedia, and WordNet, and spectral graph theory in order to represent documents for further processing.
George Tsatsaronis 0001, Iraklis Varlamis, Kjetil Nørvåg
CIKM1
2011 Visualizing Bibliographic Databases as Graphs and Mining Potential Research Synergies
abstract
Bibliographic databases are a prosperous field for data mining research and social network analysis. They contain rich information, which can be analyzed across different dimensions(e.g., author, year, venue, topic) and can be exploited in multiple ways. The representation and visualization of bibliographic databases as graphs and the application of data mining techniques can help us uncover interesting knowledge concerning potential synergies between researchers, possible matchings between researchers and venues, or even the ideal venue for presenting a research work. In this paper, we propose a novel representation model for bibliographic data, which combines co-authorship and content similarity information, and allows for the formation of scientific networks. Using a graph visualization tool from the biological domain, we are able to provide comprehensive visualizations that help us uncover hidden relations between authors and suggest potential synergies between researchers or groups.
Iraklis Varlamis, George Tsatsaronis 0001
ASONAM2
2011 How to Become a Group Leader? or Modeling Author Types Based on Graph Mining
George Tsatsaronis 0001, Iraklis Varlamis, Sunna Torge 0001, Matthias Reimann, Kjetil Nørvåg, Michael Schroeder 0001, Matthias Zschunke
TPDL1
2011 TRUMIT: A Tool to Support Large-Scale Mining of Text Association Rules
Robert Neumayer, George Tsatsaronis 0001, Kjetil Nørvåg
ECML/PKDD (3)2
2011 A Knowledge-Based Semantic Kernel for Text Classification
Jamal Abdul Nasir, Asim Karim, George Tsatsaronis 0001, Iraklis Varlamis
SPIRE3
2011 Data mining in software engineering
abstract
The increased availability of data created as part of the software development process allows us to apply novel analysis techniques on the data and use the results to guide the process's optimization. In this paper we describe various data sources an
Maria Halkidi, Diomidis Spinellis, George Tsatsaronis 0001, Michalis Vazirgiannis
Intell. Data Anal.3
2010 An Experimental Study on Unsupervised Graph-based Word Sense Disambiguation
George Tsatsaronis 0001, Iraklis Varlamis, Kjetil Nørvåg
CICLing1
2010 SemanticRank: Ranking Keywords and Sentences Using Semantic Graphs
George Tsatsaronis 0001, Iraklis Varlamis, Kjetil Nørvåg
COLING1
2010 KDTA: Automated Knowledge-Driven Text Annotation
Katerina Papantoniou, George Tsatsaronis 0001, Georgios Paliouras
ECML/PKDD (3)2
2010 Text Relatedness Based on a Word Thesaurus
abstract
The computation of relatedness between two fragments of text in an automated manner requires taking into account a wide range of factors pertaining to the meaning the two fragments convey, and the pairwise relations between their words. Without doubt, a measure of relatedness between text segments must take into account both the lexical and the semantic relatedness between words. Such a measure that captures well both aspects of text relatedness may help in many tasks, such as text retrieval, classification and clustering. In this paper we present a new approach for measuring the semantic relatedness between words based on their implicit semantic links. The approach exploits only a word thesaurus in order to devise implicit semantic links between words. Based on this approach, we introduce Omiotis, a new measure of semantic relatedness between texts which capitalizes on the word-to-word semantic relatedness measure (SR) and extends it to measure the relatedness between texts. We gradually validate our method: we first evaluate the performance of the semantic relatedness measure between individual words, covering word-to-word similarity and relatedness, synonym identification and word analogy; then, we proceed with evaluating the performance of our method in measuring text-to-text semantic relatedness in two tasks, namely sentence-to-sentence similarity and paraphrase recognition. Experimental evaluation shows that the proposed method outperforms every lexicon-based method of semantic relatedness in the selected tasks and the used data sets, and competes well against corpus-based and hybrid approaches.
George Tsatsaronis 0001, Iraklis Varlamis, Michalis Vazirgiannis
J. Artif. Intell. Res.1
2009 Omiotis: A Thesaurus-Based Measure of Text Relatedness
George Tsatsaronis 0001, Iraklis Varlamis, Michalis Vazirgiannis, Kjetil Nørvåg
ECML/PKDD (2)1
2007 Word Sense Disambiguation with Spreading Activation Networks Generated from Thesauri
George Tsatsaronis 0001, Michalis Vazirgiannis, Ion Androutsopoulos
IJCAI1
2005 Word Sense Disambiguation for Exploiting Hierarchical Thesauri in Text Classification
Dimitrios Mavroeidis, George Tsatsaronis 0001, Michalis Vazirgiannis, Martin Theobald, Gerhard Weikum
PKDD2