EDBT 2026 Demo / reviewers in the wild / expert
Vassilis Katsouros
dblp:77/5060 · also Vassilios Katsouros
· DBLP profile ↗
35ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0002-4185-2344ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Disentangling Local and Global Semantics in Diffusion Models for Image EditingabstractAbstract Diffusion models have achieved state-of-the-art image synthesis, yet unlike GANs, they lack a well-structured latent space for intuitive image editing. Existing diffusion-based editing methods often rely on supervised fine-tuning or text-based guidance, while recent unsupervised techniques leveraging the model’s bottleneck layer suffer from one or more key limitations: (i) they focus only on global attributes, (ii) fail to disentangle local and global semantics, or (iii) require extensive human intervention. To fill this gap, we first propose an unsupervised method for localized image editing in pre-trained unconditional diffusion models that disentangles local and global semantics in the model’s latent space. Given an input image and a user-specified region of interest, our approach uses the denoising network’s Jacobian to map that region to a corresponding latent subspace. We then separate this subspace into shared (global) and region-specific components to uncover latent directions that control local attributes. These directions generalize across images, enabling semantically consistent edits without retraining. We go one step further by extending our method to minimize manual supervision by automatically inferring edit directions from a single reference image and generating region masks without human input. Experiments on multiple datasets show that our method yields more localized, high-fidelity edits than state-of-the-art approaches. Manos Plitsis, Theodoros Kouzelis, Panagiotis Koromilas, Vassilis Katsouros, Mihalis A. Nicolaou, Yannis Panagakis |
Int. J. Comput. Vis. | 4 |
| 2025 | Pay (Cross) Attention to the Melody: Curriculum Masking for Single-Encoder Melodic Harmonization
Maximos Kaliakatsos-Papakostas, Dimos Makris, Konstantinos Soiledis, Konstantinos-Theodoros Tsamis, Vassilis Katsouros, Emilios Cambouropoulos |
IEEE Big Data | 5 |
| 2025 | Old Greek OCR Result Correction Using LLMsabstractRecognition of historical documents is still an active research field due to the relatively low recognition accuracy achieved when processing old fonts or low-quality images. In this work, we investigate the use of Large Language Models (LLMs) for the correction of the OCR for old Greek documents. We examine two different old Greek datasets, one machine printed and one typewritten, using a Deep Network based OCR together with several known and easy-to-use LLMs for the correction of the result. Additionally, we synthetically produce erroneous texts and change the LLM prompts in order to further study the behavior of LLMs for correcting old Greek noisy text. Experimental results highlight the potential of LLMs for OCR correction of old Greek documents especially for the cases that the recognition results are relatively poor. Andreas Evaggelatos, Konstantinos Palaiologos, Basilios Gatos, Panagiotis Kaddas, Aikaterini Christopoulou, Vassilis Katsouros |
DocEng | 6 |
| 2025 | Medusa: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
INTERSPEECH | 7 |
| 2024 | AI-Enabled Art Education: Unleashing Creative Potential and Exploring Co-Creation Frontiers
Vassilis Evangelidis, Helena G. Theodoropoulou, Vassilis Katsouros, Chairi Kiourt |
CSEDU (2) | 3 |
| 2024 | Investigating Personalization Methods in Text to Music GenerationabstractIn this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio diffusers with two established personalization methods. We experiment with the effect of audio-specific data augmentation on the overall system performance and assess different training strategies. For evaluation, we construct a novel dataset with prompts and music clips. We consider both embedding-based and music-specific metrics for quantitative evaluation, as well as a user study for qualitative evaluation. Our analysis shows that similarity metrics are in accordance with user preferences and that current personalization approaches tend to learn rhythmic music concepts more easily than melody. The code, dataset, and example material of this study are open to the research community1. Manos Plitsis, Theodoros Kouzelis, Georgios Paraskevopoulos, Vassilis Katsouros, Yannis Panagakis |
ICASSP | 4 |
| 2024 | The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
Georgios Paraskevopoulos, Chara Tsoukala, Athanasios Katsamanis, Vassilis Katsouros |
INTERSPEECH | 4 |
| 2024 | Sample-Efficient Unsupervised Domain Adaptation of Speech Recognition Systems: A Case Study for Modern GreekabstractModern speech recognition systems exhibit rapid performance degradation under domain shift. This issue is especially prevalent in data-scarce settings, such as low-resource languages, where the diversity of training data is limited. In this work, we propose M2DS2, a simple and sample-efficient fine-tuning strategy for large pre-trained speech models, based on mixed source and target domain self-supervision. We find that including source domain self-supervision stabilizes training and avoids mode collapse of the latent representations. For evaluation, we collect HParl, a 120-hour speech corpus for Greek, consisting of plenary sessions in the Greek Parliament. We merge HParl with two popular Greek corpora to create GREC-MD, a test-bed for multi-domain evaluation of Greek ASR systems. In our experiments, we find that, while other Unsupervised Domain Adaptation baselines fail in this resource-constrained environment, M2DS2 yields significant improvements for cross-domain adaptation, even when only a few hours of in-domain audio are available. When we relax the problem in a weakly supervised setting, we find that independent adaptation for audio using M2DS2 and language using simple LM augmentation techniques is particularly effective, yielding word error rates comparable to the fully supervised baselines. Georgios Paraskevopoulos, Theodoros Kouzelis, Georgios Rouvalis, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents
Panagiotis Kaddas, Konstantinos Palaiologos, Basilios Gatos, Vassilis Katsouros, Katerina Christopoulou |
ICDAR (4) | 4 |
| 2023 | Weakly-supervised forced alignment of disfluent speech using phoneme-level modeling
Theodoros Kouzelis, Georgios Paraskevopoulos, Athanasios Katsamanis, Vassilis Katsouros |
INTERSPEECH | 4 |
| 2022 | A Few-Sample Strategy for Guitar Tablature Transcription Based on Inharmonicity Analysis and Playability ConstraintsabstractThe prominent strategical approaches regarding the problem of guitar tablature transcription rely either on fingering patterns encoding or on the extraction of string-related audio features. The current work combines the two aforementioned strategies in an explicit manner by employing two discrete components for string-fret classification. It extends older few-sample modeling strategies by introducing various adaptation schemes for the first stage of audio processing, taking advantage of the inharmonic characteristics of guitar sound. Physical limitations and common standards of human performers are incorporated in a genetic algorithm which constitutes a second contextual-based module that further processes the initial audio-based predictions. The proposed methods are evaluated on both annotated guitar performances and isolated note recordings. Grigoris Bastas, Stefanos Koutoupis, Maximos Kaliakatsos-Papakostas, Vassilis Katsouros, Petros Maragos |
ICASSP | 4 |
| 2022 | Zero-Shot Cross-lingual Aphasia Detection using Automatic Speech Recognition
Gerasimos Chatzoudis, Manos Plitsis, Spyridoula Stamouli, Athanasia-Lida Dimou, Athanasios Katsamanis, Vassilis Katsouros |
INTERSPEECH | 6 |
| 2022 | SciPar: A Collection of Parallel Corpora from Scientific AbstractsabstractThis paper presents SciPar, a new collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries of universities and national archives. We describe first how we harvested and processed metadata from 86, mainly European, repositories to extract bilingual titles and abstracts, and then how we mined high quality sentence pairs in a wide range of scientific areas and sub-disciplines. In total, the resource includes 9.17 million segment alignments in 31 language pairs and is publicly available via the ELRC-SHARE repository. The bilingual corpora in this collection could prove valuable in various applications, such as cross-lingual plagiarism detection or adapting Machine Translation systems for the translation of scientific texts and academic writing in general, especially for language pairs which include English. Dimitrios Roussis, Vassilis Papavassiliou, Prokopis Prokopidis, Stelios Piperidis, Vassilis Katsouros |
LREC | 5 |
| 2022 | Regotron: Regularizing the Tacotron2 Architecture Via Monotonic Alignment LossabstractDeep learning Text-to-Speech (TTS) systems have achieved impressive generated speech quality, close to human parity. However, they suffer from training stability issues and in-correct alignment between the intermediate acoustic representation and the text input. In this work, we propose Regotron, a regularized Tacotron2 version which alleviates the training issues by augmenting the objective function with an additional term, which penalizes non-monotonic alignments in the location-sensitive attention mechanism. By introducing this regularization term we demonstrate its effectiveness to stabilize the training process, produce a monotonic attention quicker (13% of the total number of epochs compared to Tacotron2) and reduce the alignment errors during inference. Moreover, Regotron has minimal additional computational overhead, reduces common TTS mistakes and at the same time achieves improved speech naturalness according to subjective mean opinion scores (MOS) collected from 50 evaluators. Efthymios Georgiou, Kosmas Kritsis, Georgios Paraskevopoulos, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
SLT | 5 |
| 2021 | Attention-based Multimodal Feature Fusion for Dance Motion GenerationabstractRecent advances in deep learning have enabled the extraction of high-level skeletal features from raw images and video sequences, paving the way for new possibilities in a variety of artificial intelligence tasks, including automatically synthesized human motion sequences. In this paper we present a system that combines 2D skeletal data and musical information to generate skeletal dancing sequences. The architecture is implemented solely with convolutional operations and trained by following a teacher-force supervised learning approach, while the synthesis of novel motion sequences follows an autoregressive process. Additionally, by employing an attention mechanism we fuse the latent representations of past music and motion information in order to condition the generation process. For assessing the system performance, we generated 900 sequences and evaluated the perceived realism, motion diversity and multimodality of the generated sequences based on various diversity metrics. Kosmas Kritsis, Aggelos Gkiokas, Aggelos Pikrakis, Vassilis Katsouros |
ICMI | 4 |
| 2020 | Air-Writing Recognition using Deep Convolutional and Recurrent Neural Network ArchitecturesabstractIn this paper, we explore deep learning architectures applied to the air-writing recognition problem where a person writes text freely in the three dimensional space. We focus on handwritten digits, namely from 0 to 9, which are structured as multidimensional time-series acquired from a Leap Motion Controller (LMC) sensor. We examine both dynamic and static approaches to model the motion trajectory. We train and compare several state-of-the-art convolutional and recurrent architectures. Specifically, we employed a Long Short-Term Memory (LSTM) network and also its bidirectional counterpart (BLSTM) in order to map the input sequence to a vector of fixed dimensionality, which is subsequently passed to a dense layer for classification among the targeted air-handwritten classes. In the second architecture we adopt 1D Convolutional Neural Networks (CNNs) to encode the input features before feeding them to an LSTM neural network (CNN-LSTM). The third architecture is a Temporal Convolutional Network (TCN) that uses dilated causal convolutions. Finally, a deep CNN architecture for automating the feature learning and classification from raw input data is presented. The performance evaluation has been carried out on a dataset of 10 participants, who wrote each digit at least 10 times, resulting in almost 1200 examples. Grigoris Bastas, Kosmas Kritsis, Vassilis Katsouros |
ICFHR | 3 |
| 2018 | Musical track popularity mining dataset: Extension & experimentation
Ioannis Karydis, Aggelos Gkiokas, Vassilis Katsouros, Lazaros S. Iliadis |
Neurocomputing | 3 |
| 2016 | Recognition of Greek Polytonic on Historical Degraded Texts Using HMMsabstractOptical Character Recognition (OCR) of ancient Greek polytonic scripts is a challenging task due to the large number of character classes, resulting from variations of diacritical marks on the vowel letters. Classical OCR systems require a character segmentation phase, which in the case of Greek polytonic scripts is the main source of errors that finally affects the overall OCR performance. This paper suggests a character segmentation free HMM-based recognition system and compares its performance with other commercial, open source, and state-of-the art OCR systems. The evaluation has been carried out on a challenging novel dataset of Greek polytonic degraded texts and has shown that HMM-based OCR yields character and word level error rates of 8.61% and 25.30% respectively, which outperforms most of the available OCR systems and it is comparable with the performance of the state-of-the-art system based on LSTM Networks proposed recently. Vassilis Katsouros, Vassilis Papavassiliou, Foteini Liwicki, Basilios Gatos |
DAS | 1 |
| 2016 | Towards Multi-Purpose Spectral Rhythm Features: An Application to Dance Style, Meter and Tempo EstimationabstractThis paper addresses the extraction of multipurpose spectral rhythm features that simultaneously tackle a variety of rhythm analysis tasks, namely, dance style classification, meter estimation, and tempo estimation. The term spectral rhythm features emanates from the origin of the extracted features, which is the periodicity function (PF), a spectral representation that encapsulates the salience of the rhythm frequencies. Two dimensionality reduction techniques applied on the PF to extract expressive and compact features are compared, namely, a linear transformation resulting from Principal Component Analysis and a nonlinear mapping derived from a Restricted Boltzmann Machine. Subsequently, the derived features were used as input to an SVM classifier for each task. Moreover, an additional method is proposed that reformulates the well-studied tempo estimation task as a combination of multiple binary classification sub-problems. Evaluation was performed on a large number of datasets demonstrating that the same set of features learned from the PF provide a robust rhythmic representation that achieved comparable results to the current state-of-the-art methods for the aforementioned tasks. Aggelos Gkiokas, Vassilis Katsouros, George Carayannis |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | GRPOLY-DB: An old Greek polytonic document image databaseabstractRecognition of old Greek document images containing polytonic (multi accent) characters is a challenging task due to the large number of existing character classes (more than 270) which cannot be handled sufficiently by current OCR technologies. Taking into account that the Greek polytonic system was used from the late antiquity until recently, a large amount of scanned Greek documents still remains without full test search capabilities. In order to assist the progress of relevant research, this paper introduces the first publicly available old Greek polytonic database GRPOLY-DB for the evaluation of several document image processing tasks. It contains both machine-printed and handwritten documents as well as annotation with ground-truth information that can be used for training and evaluation of the most commou document image processing tasks, i.e.. text line and word segmentation, test recognition, isolated character recognition and word spotting. Results using several representative baseline technologies are also presented in order to help researchers evaluate their methods and advance the frontiers of old Greek document image recognition and word spotting. Basilios Gatos, Nikolaos Stamatopoulos, Georgios Louloudis, Giorgos Sfikas, George Retsinas, Vassilis Papavassiliou, Fotini Sunistira, Vassilis Katsouros |
ICDAR | 8 |
| 2015 | Recognition of historical Greek polytonic scripts using LSTM networksabstractThis paper reports on high-performance Optical Character Recognition (OCR) experiments using Long Short-Term Memory (LSTM) Networks for Greek polytonic script. Even though there are many Greek polytonic manuscripts, the digitization of such documents has not been widely applied, and very limited work has been done on the recognition of such scripts. We have collected a large number of diverse document pages of Greek polytonic scripts in a novel database, called Polyton-DB, containing 15; 689 textlines of synthetic and authentic printed scripts and performed baseline experiments using LSTM Networks. Evaluation results show that the character error rate obtained with LSTM varies from 5.51% to 14.68% (depending on the document) and is better than two well-known OCR engines, namely, Tesseract and ABBYY FineReader. Foteini Liwicki, Adnan Ul-Hasan, Vassilis Papavassiliou, Basilios Gatos, Vassilis Katsouros, Marcus Liwicki |
ICDAR | 5 |
| 2015 | Recognition of online handwritten mathematical formulas using probabilistic SVMs and stochastic context free grammars
Foteini Liwicki, Vassilis Katsouros, George Carayannis |
Pattern Recognit. Lett. | 2 |
| 2014 | Recognition of Spatial Relations in Mathematical FormulasabstractA critical issue in recognition of mathematical expressions is the identification of the spatial relations of the symbols or/and sub-expressions that comprise the entire mathematical formula. This paper addresses the problem of structural analysis of mathematical expressions by constructing appropriate feature vectors to represent the spatial affinity of the objects (mathematical symbols or sub-expressions) under examination and employing two popular machine learning techniques: (i) Support Vector Machines (SVM) and (ii) Artificial Neural Networks (ANN) to recognize the spatial relation between these objects. In order to evaluate the proposed techniques, we use Math Brush, a large publicly available dataset of mathematical expressions with annotated spatial relations, and a subset of spatial relations derived from the mathematical expressions the CROHME 2012 dataset. The experimental results give an overall mean error rate of 2.8% for the SVM and 3.4% for the ANN classifiers respectively, which are at par with other approaches evaluated on the same datasets. Foteini Liwicki, Vassilis Papavassiliou, Vassilis Katsouros, George Carayannis |
ICFHR | 3 |
| 2012 | Music tempo estimation and beat tracking by applying source separation and metrical relationsabstractIn this paper, we present tempo estimation and beat tracking algorithms by utilizing percussive/harmonic separation of the audio signal, in order to extract filterbank energies and chroma features from the respective components. Periodicity analysis is carried out by the convolution of feature sequences with a bank of resonators. Target tempo is estimated from the resulting periodicity vector by incorporating metrical relations knowledge. Tempo estimation is followed by a local tempo refinement method to enhance the beat-tracking algorithm. Beat tracking involves the computation of the beat saliencies derived from the resonators responses and proposes a distance measure between candidate beats locations. A dynamic programming algorithm is adopted to find the optimal “path” of beats. Both tempo estimation and beat tracking methods were submitted on MIREX 2011, while the tempo estimation algorithm was also evaluated on ISMIR 2004 Tempo Induction Evaluation Exchange Dataset. Aggelos Gkiokas, Vassilis Katsouros, George Carayannis, Themos Stafylakis |
ICASSP | 2 |
| 2012 | A Morphology Based Approach for Binarization of Handwritten DocumentsabstractDocument image binarization is an initial though critical stage towards the recognition of the text components of a document. This paper describes an efficient method based on mathematical morphology for extracting text regions from degraded handwritten document images. The basic stages of our approach are: (a) top-hat-by-reconstruction to produce a filtered image with reasonable even background, (b) region growing starting from a set of seed points and attaching to each seed similar intensity neighboring pixels and (c) conditional extension of the initially detected text regions based on the values of the second derivative of the filtered image. The method was evaluated on the benchmarking dataset of the International Document Image Binarization Contest (DIBCO 2011) and show promising results. Vassilis Papavassiliou, Foteini Liwicki, Vassilis Katsouros, George Carayannis |
ICFHR | 3 |
| 2012 | A System for Recognition of On-Line Handwritten Mathematical ExpressionsabstractWe present a system for recognizing online mathematical expressions (ME). Symbol recognition is based on a template elastic matching distance between pen direction features. The structural analysis of the ME is based on extracting the baseline of the ME and then classifying symbols into levels above and below the baseline. The symbols are then sequentially analyzed using six spatial relations and a respective 2d structure is processed to give the resulting MathML representation of the ME. The system was evaluated on the Competition on Recognition of Online Handwritten Mathematical Expressions (CROHME) 2011 datasets and demonstrates promising results. Foteini Liwicki, Vassilis Papavassiliou, Vassilis Katsouros, George Carayannis |
ICFHR | 3 |
| 2011 | Closed-form expressions vs. BIC: A comparison for speaker clusteringabstractIn this paper, the use of closed-form expressions is compared to the BIC approximation, with respect to speaker clustering. We first show that the particular BIC setting which is commonly used in this task, namely the approximation of the marginal with respect to the model parameters and conditional with respect to the latent variables likelihood, belongs to an exponential family, and hence admits a closed-form expression by attaching conjugate priors. We then formalize the role of the tuning parameter as a hyperparameter of the prior and finally we explain the several proposed setting global, local and segmental based on the strength of the prior. Experiments are carried out for the speaker clustering task and improvement over the BIC approximation is reported. Themos Stafylakis, Xavier Anguera Miró, Vassilis Katsouros, George Carayannis |
ICASSP | 3 |
| 2011 | Enhancing Handwritten Word Segmentation by Employing Local Spatial FeaturesabstractThis paper proposes an enhancement of our previously presented word segmentation method (ILSPLWseg) [1] by exploiting local spatial features. ILSP-LWseg is based on a gap metric that exploits the objective function of a soft-margin linear SVM that separates successive connected components (CCs). Then a global threshold for the gap metrics is estimated and used to classify the candidate gaps in "within" or "between" words classes. In the proposed enhancement the initial categorization is examined against the local features (i.e. margin and slope of the linear classifier for every pair of CCs in each text line) and a refined classification is applied for each text line. The method was tested on the benchmarking datasets of ICDAR07, ICDAR09 and ICFHR10 handwriting segmentation contests and performs better than the winning algorithm. Foteini Liwicki, Vassilis Papavassiliou, Themos Stafylakis, Vassilis Katsouros |
ICDAR | 4 |
| 2010 | A new penalty term for the BIC with respect to speaker diarizationabstractIn this paper we examine a new penalty term for the Bayesian Information Criterion (BIC) that is suited to the problem of speaker diarization. Based on our previous approach of penalizing each cluster only with its effective sample size - an approach we called segmental - we propose a stricter penalty term. The criterion we derive retains the main property of the Segmental-BIC, i.e. it approximates the evidence of overall partitions of the data and simultaneously leads to a pairwise dissimilarity measure that is completely defined by the pair of clusters in question. The experimental results show significant improvement in diarization accuracy on the ESTER benchmark. Themos Stafylakis, Georgios Tzimiropoulos, Vassilis Katsouros, George Carayannis |
ICASSP | 3 |
| 2010 | A Morphological Approach for Text-Line Segmentation in Handwritten DocumentsabstractDocument image segmentation to text lines is a critical stage towards unconstrained handwritten document recognition. Although morphological operations proved to be effective in processing machine-printed documents for several issues, similar methods for unconstraint-handwritten documents lack accuracy. We propose an efficient method based on binary morphology for text-line segmentation in such documents. The basic steps of our approach are: a) sub sampling and binary rank order filtering to enhance the text-line structures and b) applying dilations and (p,q)-th generalized foreground rank openings successively to join close and horizontally overlapping regions while preventing a merge in the vertical direction. The method tested on the benchmarking dataset of the ICDAR07 handwriting segmentation contest and show remarkable results. Vassilis Papavassiliou, Vassilis Katsouros, George Carayannis |
ICFHR | 2 |
| 2010 | Handwritten document image segmentation into text lines and words
Vassilis Papavassiliou, Themos Stafylakis, Vassilis Katsouros, George Carayannis |
Pattern Recognit. | 3 |
| 2009 | Redefining the Bayesian information criterion for speaker diarisationabstractA novel approach to the Bayesian Information Criterion (BIC) is introduced. The new criterion redefines the penalty terms of the BIC, such that each parameter is penalized with the effective sample size is trained with. Contrary to Local-BIC, the proposed criterion scores overall clustering hypotheses and therefore is not restricted to hierarchical clustering algorithms. Contrary to Global-BIC, it provides a local dissimilarity measure that depends only the statistics of the examined clusters and not on the overall sample size. We tested our criterion with two benchmark tests and found significant improvement in performance in the speaker diarisation task. Copyright © 2009 ISCA. Themos Stafylakis, Vassilis Katsouros, George Carayannis |
INTERSPEECH | 2 |
| 2008 | Robust text-line and word segmentation for handwritten documents imagesabstractThis paper addresses the problem of automatic text-line and word segmentation in handwritten document images. Two novel approaches are presented, one for each task. In text-line segmentation a Viterbi algorithm is proposed while an SVM-based metric is adopted to locate words in each text-line. The overall algorithm was tested in the ICDAR2007 handwriting segmentation contest and showed highly promising results. Themos Stafylakis, Vassilis Papavassiliou, Vassilis Katsouros, George Carayannis |
ICASSP | 3 |
| 2007 | Efficient combination of parametric spaces, models and metrics for speaker diarization1abstractIn this paper we present a method of combining several acoustic parametric spaces, statistical models and distance metrics in speaker diarization task. Focusing our interest on the post-segmentation part of the problem, we adopt an incremental feature selection and fusion algorithm based on the Maximum Entropy Principle and Iterative Scaling Algorithm that combines several statistical distance measures on speech-chunk pairs. By this approach, we place the merging-of-chunks clustering process into a probabilistic framework. We also propose a decomposition of the input space according to gender, recording conditions and chunk lengths. The algorithm produced highly competitive results compared to GMM-UBM state-of-the-art methods. Themos Stafylakis, Vassilis Katsouros, George Carayannis |
ASRU | 2 |
| 2007 | A Parametric Spectral-Based Method for Verification of Text in VideosabstractA new method for verifying text areas detected in video streams is proposed. The algorithm explores the spectral properties of the horizontal projection of candidate text regions in order to reduce the high amount of false alarms that most text detection algorithms suffer from. The full algorithm (text localization followed by verification and temporal redundancy module) has been tested on newscast video sequences (MPEG-1-720 x 576 resolution-184 minutes). The detection module produced 94.82% recall rate but only 51.84% precision rate. The addition of the verification module increased the precision rate to 78.93% keeping the recall rate almost unaffected. Vassilis Papavassiliou, Themos Stafylakis, Vassilis Katsouros, George Carayannis |
ICDAR | 3 |