Konstantina Nikolaidou

dblp:292/3661 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2024
0000-0002-9332-3188ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 DiffusionPen: Towards Controlling the Style of Handwritten Text Generation
Konstantina Nikolaidou, George Retsinas, Giorgos Sfikas, Marcus Liwicki
ECCV (85)1
2024 Enhancing CRNN HTR Architectures with Transformer Blocks
George Retsinas, Konstantina Nikolaidou, Giorgos Sfikas
ICDAR (4)2
2024 How GANs assist in Covid-19 pandemic era: a review
Yahya Sherif Solayman Mohamed Saleh, Hamam Mokayed, Konstantina Nikolaidou, Lama Alkhaled, Yan Chai Hum
Multim. Tools Appl.3
2023 WordStylist: Styled Verbatim Handwritten Text Generation with Latent Diffusion Models
Konstantina Nikolaidou, George Retsinas, Vincent Christlein, Mathias Seuret, Giorgos Sfikas, Elisa H. Barney Smith, Hamam Mokayed, Marcus Liwicki
ICDAR (2)1
2022 Investigating the Effect of Using Synthetic and Semi-synthetic Images for Historical Document Font Classification
Konstantina Nikolaidou, Richa Upadhyay, Mathias Seuret, Marcus Liwicki
DAS1
2022 Potential Idiomatic Expression (PIE)-English: Corpus for Classes of Idioms
abstract
We present a fairly large, Potential Idiomatic Expression (PIE) dataset for Natural Language Processing (NLP) in English. The challenges with NLP systems with regards to tasks such as Machine Translation (MT), word sense disambiguation (WSD) and information retrieval make it imperative to have a labelled idioms dataset with classes such as it is in this work. To the best of the authors’ knowledge, this is the first idioms corpus with classes of idioms beyond the literal and the general idioms classification. In particular, the following classes are labelled in the dataset: metaphor, simile, euphemism, parallelism, personification, oxymoron, paradox, hyperbole, irony and literal. We obtain an overall inter-annotator agreement (IAA) score, between two independent annotators, of 88.89%. Many past efforts have been limited in the corpus size and classes of samples but this dataset contains over 20,100 samples with almost 1,200 cases of idioms (with their meanings) from 10 classes (or senses). The corpus may also be extended by researchers to meet specific needs. The corpus has part of speech (PoS) tagging from the NLTK library. Classification experiments performed on the corpus to obtain a baseline and comparison among three common models, including the BERT model, give good results. We also make publicly available the corpus and the relevant codes for working with it for NLP tasks.
Tosin P. Adewumi, Roshanak Vadoodi, Aparajita Tripathy, Konstantina Nikolaidou, Foteini Liwicki, Marcus Liwicki
LREC4
2022 A survey of historical document image datasets
abstract
Abstract This paper presents a systematic literature review of image datasets for document image analysis, focusing on historical documents, such as handwritten manuscripts and early prints. Finding appropriate datasets for historical document analysis is a crucial prerequisite to facilitate research using different machine learning algorithms. However, because of the very large variety of the actual data (e.g., scripts, tasks, dates, support systems, and amount of deterioration), the different formats for data and label representation, and the different evaluation processes and benchmarks, finding appropriate datasets is a difficult task. This work fills this gap, presenting a meta-study on existing datasets. After a systematic selection process (according to PRISMA guidelines), we select 65 studies that are chosen based on different factors, such as the year of publication, number of methods implemented in the article, reliability of the chosen algorithms, dataset size, and journal outlet. We summarize each study by assigning it to one of three pre-defined tasks: document classification, layout structure, or content analysis. We present the statistics, document type, language, tasks, input visual aspects, and ground truth information for every dataset. In addition, we provide the benchmark tasks and results from these papers or recent competitions. We further discuss gaps and challenges in this domain. We advocate for providing conversion tools to common formats (e.g., COCO format for computer vision tasks) and always providing a set of evaluation metrics, instead of just one, to make results comparable across studies.
Konstantina Nikolaidou, Mathias Seuret, Hamam Mokayed, Marcus Liwicki
Int. J. Document Anal. Recognit.1
2021 Detecting COVID-19 from Audio Recording of Coughs Using Random Forests and Support Vector Machines
abstract
The detection of COVID-19 is and will remain in the foreseeable future a crucial challenge, making the development of tools for the task important.One possible approach, on the confines of speech and audio processing, is detecting potential COVID-19 cases based on cough sounds.We propose a simple, yet robust method based on the well-known ComParE 2016 feature set, and two classical machine learning models, namely Random Forests, and Support Vector Machines (SVMs).Furthermore, we combine the two methods, by calculating the weighted average of their predictions.Our results in the DiCOVA challenge show that this simple approach leads to a robust solution while producing competitive results.Based on the Area Under the Receiver Operating Characteristic Curve (AUC ROC) score, both classical machine learning methods we applied markedly outperform the baseline provided by the challenge organisers.Moreover, their combination attains an AUC ROC score of 85.21, positioning us at fourth place on the leaderboard (where the second team attained a similar, 85.43 score).Here, we would describe this system in more detail, and analyse the resulting models, drawing conclusions, and determining future work directions.
Isabella Södergren, Maryam Pahlavan Nodeh, Prakash Chandra Chhipa, Konstantina Nikolaidou, György Kovács 0001
Interspeech4