EDBT 2026 Demo / reviewers in the wild / expert
Constantine Kotropoulos
dblp:k/CKotropoulos · also Costas Kotropoulos
· DBLP profile ↗
122ranked-venue papers
21as first author
14since 2021 · last 2026
0000-0001-9939-7930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 17 first-author · 7 since 2021Artificial intelligence and machine learning · 41 · 5 first-author · 6 since 2021Systems, architecture and hardware · 7 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parameter-Efficient and Adaptive Fine-Tuning for Long-Tailed Ancient Characters Recognition
Aouaidjia Kamel, Constantine Kotropoulos, Chongsheng Zhang |
ICDAR (3) | 3 |
| 2025 | DeepENF: A data-driven Electric Network Frequency estimation framework
Ioannis Tsingalis, Constantine Kotropoulos |
Pattern Recognit. Lett. | 2 |
| 2024 | Interpretable Face Aging: Enhancing Conditional Adversarial Autoencoders with Lime ExplanationsabstractAn innovative approach is proposed that leverages a perturbation explainable system within the Conditional Adversarial Autoencoder (CAAE) framework. The incorporation of the perturbation-based explainable system in the CAAE model harnesses the explanatory power of Local Interpretable Model-Agnostic Explanations (LIME). LIME generates perturbations in the latent space of the CAAE and provides insightful explanations for the discrepancies between fake and real face images. By indicating the areas that contribute most significantly to the aging process, LIME guides the adversarial training process to focus on those aspects, resulting in corrective feedback to the discriminator. The performance of the proposed framework, against state-of-the-art methods, is assessed by objective figures of merit demonstrating superior results in face aging. Christos Korgialas, Evangelia Pantraki, Constantine Kotropoulos |
ICASSP | 3 |
| 2024 | Applying the Neural Bellman-Ford Model to the Single Source Shortest Path Problem
Spyridon Drakakis, Constantine Kotropoulos |
ICPRAM | 2 |
| 2024 | Mobile Phone Identification from Recorded Speech Signals Using Non-Speech Segments and Universal Background Model Adaptation
Dimitrios Kritsiolis, Constantine Kotropoulos |
ICPRAM | 2 |
| 2024 | On Spectrogram Analysis in a Multiple Classifier Fusion Framework for Power Grid Classification Using Electric Network FrequencyabstractThe Electric Network Frequency (ENF) serves as a unique signature inherent to power distribution systems. Here, a novel approach for power grid classification is developed, leveraging ENF. Spectrograms are generated from audio and power recordings across different grids, revealing distinctive ENF patterns that aid in grid classification through a fusion of classifiers. Four traditional machine learning classifiers plus a Convolutional Neural Network (CNN), optimized using Neural Architecture Search, are developed for One-vs-All classification. This process generates numerous predictions per sample, which are then compiled and used to train a shallow multi-label neural network specifically designed to model the fusion process, ultimately leading to the conclusive class prediction for each sample. Experimental findings reveal that both validation and testing accuracy outperform those of current state-of-the-art classifiers, underlining the effectiveness and robustness of the proposed methodology. Georgios Tzolopoulos, Christos Korgialas, Constantine Kotropoulos |
ICPRAM | 3 |
| 2023 | Electric Network Frequency Detection Using Least Absolute DeviationsabstractElectric Network Frequency (ENF) is a fingerprint in multi-media forensics applications. ENF is a weak signal that is difficult to be detected. This difficulty stems from the existence of colored wide-sense stationary Gaussian noise in ENF as well as due to many unknown random parameters. However, several ENF detectors have been proposed, motivating the related research. In this paper, a novel Least Absolute Deviations-based ENF detector is proposed that is coined as LAD-Likelihood Ratio Test (LAD-LRT). The performance of the LAD-LRT detector is thoroughly analyzed concerning test statistic distribution and threshold selection. The aim is to develop a detector that detects ENF more accurately in short-length recordings than the state-of-the-art Least-Squares (LS)-LRT and naive-LRT detectors. Thorough evaluation using benchmark audio recordings demonstrate the effectiveness of the proposed detector. Christos Korgialas, Constantine Kotropoulos |
ICASSP | 2 |
| 2023 | Dual Hypergraph Features for Path Inference in Wikipedia LinksabstractHere, path extrapolation is assessed in a graph by extracting the most informative features from nodes and edges. Wikipedia articles are the graph nodes, and the links between articles are the graph edges. A graph neural network called GRETEL is used as baseline. New features are extracted by exploiting the information from the data. By employing the dual hypergraph transformation the structural role of nodes and edges is interchanged, enabling more complex relationships to be captured. Experimental evidence is disclosed demonstrating that the features extracted from the dual hypergraph are more de-scriptive than those extracted from the original graph, improving GRETEL performance. Anastasia-Sotiria Toufa, Constantine Kotropoulos, Ioannis Tsingalis |
IJCNN | 2 |
| 2023 | An Autoregressive Graph Convolutional Long Short-Term Memory Hybrid Neural Network for Accurate Prediction of COVID-19 CasesabstractEfficient prediction of COVID-19 cases could prepare the healthcare system to accommodate the COVID-19 cases in the forthcoming days and improve the overall resource management. A hybrid model comprised of an autoregressive filter, a graph convolutional neural network (GCN), and a long short-term memory neural network is proposed for COVID-19 cases prediction in USA. It captures accurately both linearities and nonlinearities present in the time series. An adjacency matrix is exploited in GCN that relies on Granger causality tests applied to historical COVID-19 cases for each state in USA. By doing so, the latent information about the spread of the virus is captured efficiently and the prediction performance of the hybrid model is improved, revealing which state truly affects the other ones. The proposed method outperforms the state-of-the-art techniques. Myrsini Ntemi, Ioannis Sarridis, Constantine Kotropoulos |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2022 | Cross-lingual transfer learning: A PARAFAC2 approach
Evangelia Pantraki, Ioannis Tsingalis, Constantine Kotropoulos |
Pattern Recognit. Lett. | 3 |
| 2021 | Blackman-Tukey spectral estimation and electric network frequency matching from power mains and speech recordingsabstractAbstract Forensic applications exploit electric network frequency (ENF) as a fingerprint to determine multimedia content authenticity, as well as the time and region of multimedia recording. ENF is present at a nominal frequency of 50/60 Hz and its harmonics. Strong interference due to speech content deteriorates ENF estimation accuracy. Herein, the authors propose a non‐parametric approach for ENF estimation, which incorporates a customised lag window design into the Blackman–Tukey spectral estimation method. Leakage reduction is formulated as a problem of energy maximisation within the main lobe of the spectral window. The proposed approach is compared to state‐of‐the‐art methods for ENF estimation. Maximum correlation coefficient and minimum standard deviation of errors are employed to measure ENF estimation accuracy. Hypothesis testing is performed to determine whether the improvements in ENF estimation accuracy of the proposed approach over the state‐of‐the‐art methods are statistically significant. Experimental results and statistical tests indicate that the proposed approach improves ENF estimation against many state‐of‐the‐art methods. Georgios Karantaidis, Constantine Kotropoulos |
IET Signal Process. | 2 |
| 2021 | Face aging using global and pyramid generative adversarial networks
Evangelia Pantraki, Constantine Kotropoulos |
Mach. Vis. Appl. | 2 |
| 2021 | A jump-diffusion particle filter for price prediction
Myrsini Ntemi, Constantine Kotropoulos |
Signal Process. | 2 |
| 2021 | Adaptive hypergraph learning with multi-stage optimizations for image and tag recommendation
Georgios Karantaidis, Ioannis Sarridis, Constantine Kotropoulos |
Signal Process. Image Commun. | 3 |
| 2020 | Digit Recognition Applied to Reconstructed Audio Signals Using Deep LearningabstractCompressed sensing allows signal reconstruction from a few measurements. This work proposes a complete pipeline for digit recognition applied to audio reconstructed signals. The reconstruction procedure exploits the assumption that the original signal lies in the range of a generator. A pretrained generator of a Generative Adversarial Network generates audio digits. A new method for reconstruction is proposed, using only the most active segment of the signal, i.e., the segment with the highest energy. The underlying assumption is that such segment offers a more compact representation, preserving the meaningful content of signal. Cases when the reconstruction produces noise, instead of digit, are treated as outliers. In order to detect and reject them, three unsupervised indicators are used, namely, the total energy of reconstructed signal, the predictions of an one-class Support Vector Machine, and the confidence of a pretrained classifier used for recognition. This classifier is based on neural networks architectures and is pretrained on original audio recordings, employing three input representations, i.e., raw audio, spectrogram, and gammatonegram. Experiments are conducted, analyzing both the quality of reconstruction and the performance of classifiers in digit recognition, demonstrating that the proposed method yields higher performance in both the quality of reconstruction and digit recognition accuracy. Anastasia-Sotiria Toufa, Constantine Kotropoulos |
ICPR | 2 |
| 2020 | A dynamic dyadic particle filter for price prediction
Myrsini Ntemi, Constantine Kotropoulos |
Signal Process. | 2 |
| 2019 | Leveraging Image-to-image Translation Generative Adversarial Networks for Face AgingabstractHere, face images of a specific age class are translated to images of different age classes in an unsupervised manner that enables training on independent sets of images for each age class. In order to learn pairwise translations between age classes, we adopt the UNsupervised Image-to-image Translation framework that employs Variational AutoEncoders and Generative Adversarial Networks. By mapping face images of different age classes to shared latent representations, the most personalized and abstract facial characteristics are preserved. To effectively diffuse age class information, a pyramid of local, neighbour, and global encoders is employed so that the latent representations progressively cover an increased age range. The proposed framework is applied to the FGNET aging database and compared to state-of-the-art techniques and the ground truth. Appealing experimental results demonstrate the ability of the proposed method to efficiently capture both intense and subtle aging effects. Evangelia Pantraki, Constantine Kotropoulos, Andreas Lanitis |
ICASSP | 2 |
| 2019 | Block Randomized Optimization for Adaptive Hypergraph LearningabstractThe high-order relations between the content in social media sharing platforms are frequently modeled by a hypergraph. Either hypergraph Laplacian matrix or the adjacency matrix is a big matrix. Randomized algorithms are used for low-rank factorizations in order to approximately decompose and eventually invert such big matrices fast. Here, block randomized Singular Value Decomposition (SVD) via subspace iteration is integrated within adaptive hypergraph weight estimation for image tagging, as a first approach. Specifically, creating low-rank submatrices along the main diagonal by tessellation permits fast matrix inversions via randomized SVD. Moreover, a second approach is proposed for solving the linear system in the optimization problem of hypergraph learning by employing the conjugate gradient method. Both proposed approaches achieve high accuracy in image tagging measured by F1score and succeed to reduce the computational requirements of adaptive hypergraph weight estimation. Georgios Karantaidis, Ioannis Sarridis, Constantine Kotropoulos |
ICIP | 3 |
| 2019 | A Simple Algorithm for Non-Negative Sparse Principal Component AnalysisabstractA novel adaptive method that computes sparse and non-negative eigenvectors is proposed. The proposed method achieves the sparsity in the derived eigenvectors implicitly through non-negative constraints and does not include any additional parameters except the learning rate. Although adding constraints to the standard Principal Component Analysis (PCA) leads to a reduction in the explained variance, the proposed method performs competitively with the standard PCA and other PCA variants. The assessment of the proposed method is conducted by performing a quantitative and qualitative evaluation. Ioannis Tsingalis, Constantine Kotropoulos |
ICIP | 2 |
| 2018 | M-estimators for robust multidimensional scaling employing ℓ2, 1 norm regularization
Fotios D. Mandanas, Constantine Kotropoulos |
Pattern Recognit. | 2 |
| 2016 | Adaptive algorithms for hypergraph learningabstractSocial media sharing platforms enable image content as well as context information (e.g., user friendships, geo-tags assigned to images) to be jointly analyzed in order to achieve accurate image annotation or successful image recommendation. The context information is expressed frequently in terms of high-order relations, such as the relations among users, tags, and images. Hypergraphs can model the aforementioned high-order relations between their vertices (i.e., users, user social groups, tags, geo-tags, and images) by hyper-edges, whose influence can be assessed by properly estimating their weights. Here, an efficient adaptive hypergraph weight estimation is proposed for image tagging. In particular, both equality and inequality constraints enforced during hypergraph learning are taken into account and an efficient adaptation step selection using the Armijo rule is proposed. Experiments conducted on a dataset demonstrate the superior performance of the proposed approach compared to the state-of-the-art. Aikaterini Chasapi, Constantine Kotropoulos, Konstantinos Pliakos |
ICASSP | 2 |
| 2015 | A maximum correntropy criterion for robust multidimensional scalingabstractMultidimensional Scaling (MDS) refers to a class of dimensionality reduction techniques applied to pairwise dissimilarities between objects, so that the interpoint distances in the space of reduced dimensions approximate the initial pairwise dissimilarities as closely as possible. Here, a unified framework is proposed, where the MDS is treated as maximization of a correntropy criterion, which is solved by half-quadratic optimization in a multiplicative formulation. The proposed algorithm is coined as Multiplicative Half-Quadratic MDS (MHQMDS). Its performance is assessed for potential functions associated to various M-estimators, because the correntropy criterion is closely related to the Welsch M-estimator. Three state-of-the-art MDS techniques, namely the Scaling by Majorizing a Complicated Function (SMACOF), the Robust Euclidean Embedding (REE), and the Robust MDS (RMDS), are implemented under the same conditions. The experimental results indicate that the MHQMDS, relying on the M-estimators, performs better than the aforementioned state-of-the-art competing techniques. Fotios D. Mandanas, Constantine Kotropoulos |
ICASSP | 2 |
| 2015 | Weight estimation in hypergraph learningabstractThe unremitting rising popularity of social media has led to an exponential increase in web activity as manifested by the vast volume of uploaded images. This boundless volume of image data has triggered the interest in image tagging. Here, an efficient hypergraph weight estimation scheme is proposed that improves the accuracy of image tagging, using hypergraph learning. The proposed method models high-order relations between hypergraph vertices (i.e., users, user social groups, tags, geo-tags, and images) by hyperedges. The information captured by the hyperedges is efficiently distilled by estimating the hyperedge weights. Experiments conducted on a dataset crawled from Flickr demonstrate the effectiveness of the proposed approach. Specifically, an average precision of 91% at 26% recall has been achieved for image tagging. Konstantinos Pliakos, Constantine Kotropoulos |
ICASSP | 2 |
| 2015 | Image tag recommendation based on novel tensor structures and their decompositionsabstractIn this paper, we address the problem of image tagging and we propose automatic methods for image tagging, using tensor decompositions. Tensors are a suitable way of mathematically representing multilink relations. Another, complementary structure that captures the aforementioned high-order relations is the hypergraph. More specifically, three different matrices are derived from the hypergraph, namely, the incidence, adjacency, and affinity matrices. The just mentioned matrices are used to create slices of a novel tensor structure, which combines users' and images' relations. Four methods are exploited to decompose the tensor, i.e., the Higher Order Singular Value Decomposition (HOSVD), the Canonical Decomposition/Parallel Factor Analysis (CANDECOMP/ PARAFAC, CP), the Non-negative Tensor Factor Analysis (NTF), and Tucker Decomposition (TD). Experiments conducted on a dataset retrieved from Flickr demonstrate the potential of the proposed approach. Panagiotis Barmpoutis, Constantine Kotropoulos, Konstantinos Pliakos |
ISPA | 2 |
| 2015 | Greek folk music classification into two genres using lyrics and audio via canonical correlation analysisabstractWe are interested in Greek folk music genre classification by resorting to canonical correlation analysis (CCA). Here, the genre is related to the place of origin of the song. The CCA learns a linear transformation of the song lyrics descriptors that is highly correlated with their genre labels as well as another linear transformation of the audio features extracted from music recordings, which is maximally correlated with their genre labels. In the latter task, thanks to the deep CCA (DCCA), deep nonlinear transformations of the audio features are learnt, which are maximally correlated with the genre labels. Experimental findings are disclosed for a two-class genre recognition problem, employing folk songs originated from Pontus and Asia Minor. It is demonstrated that the CCA achieves an average accuracy of 97.02% across the 5 folds, when the term frequency-inverse document frequency features model the song lyrics. By modeling the music signal of each song with 28 mel-frequency cepstral coefficients (MFCCs) extracted from each frame and averaged over all frames, the average accuracy of the CCA drops to 72.9% across the 5 folds. The DCCA yields an accuracy of 69% for audio-based genre recognition. Nikoletta Bassiou, Constantine Kotropoulos, Anastasios Papazoglou-Chalikias |
ISPA | 2 |
| 2015 | Video summarization based on shot boundary detection with penalized contrastsabstractIn this paper, we propose a novel technique for shot boundary detection and video summarization that is based on change point detection with penalized contrasts. The proposed method extracts a proper time series from the video, calculating the mean level of the hue component of consecutive video frames in the Hue Saturation Value color space. Change point detection is applied to the time series, estimating the number of shot transitions and their location. The posterior distribution of the change point sequence is defined, that splits the video into homogenous temporal segments. A representative frame is selected in each segment and similar or meaningless keyframes are deleted from the summary. The resulting summary is of comparable quality to that of the state of the art techniques. Paschalina Medentzidou, Constantine Kotropoulos |
ISPA | 2 |
| 2014 | Simultaneous image tagging and geo-location prediction within hypergraph ranking frameworkabstractThe development of social media has led to a burst of interest in image-related metadata information, such as tags and geo-tags. Tags are semantic keywords that are assigned to an image. Image tagging enables the users of social media sharing platforms to annotate images, facilitating image search and content description. Despite the volume of related research, issues such as accuracy or efficiency still remain open problems. Here, a novel method for simultaneous image tagging and geo-location prediction is proposed that is based on hypergraph learning. The method is further improved by enforcing group sparsity constraints. It fully exploits various types of information, such as social, image-related metadata, or similarities based on visual attributes. Experiments on a dataset crawled from Flickr demonstrate F1at 10 top ranked tags equal to 0.558 for image tagging and cumulative geotagging prediction rate at 3 top ranks equal to 83%. Konstantinos Pliakos, Constantine Kotropoulos |
ICASSP | 2 |
| 2014 | PLSA driven image annotation, classification, and tourism recommendationabstractA burst of interest in image annotation and recommendation has been witnessed. Despite the huge effort made by the scientific community in the aforementioned research areas, accuracy or efficiency still remain open problems. Here, efficient methods for image annotation, visual image content classification as well as touristic place of interest (POI) recommendation are developed within the same framework. In particular, semantic image annotation and touristic POI recommendation harness the geo-information associated to images. Both semantic image annotation and visual image content classification resort to Probabilistic Latent Semantic Analysis (PLSA). Several tourist destinations, strongly related to the query image, are recommended, using hypergraph ranking. Experimental results were conducted on a large image dataset of Greek sites, demonstrating the potential of the proposed methods. Semantic image annotation by means of PLSA has achieved an average precision of 90% at 10% recall. The average accuracy of content-based image classification is 80%. An average precision of 90% is measured at 1% recall for tourism recommendation. Konstantinos Pliakos, Constantine Kotropoulos |
ICIP | 2 |
| 2014 | Social image search exploiting joint visual-textual information within a fuzzy hypergraph frameworkabstractThe unremitting growth of social media popularity is manifested by the vast volume of images uploaded to the web. Despite the extensive research efforts, there are still open problems in accurate or efficient image search methods. The majority of existing methods, dedicated to image search, treat the image visual content and the semantic information captured by the social image tags, separately or in a sequential manner. Here, a novel and efficient method is proposed, exploiting visual and textual information simultaneously. The joint visual-textual information is captured by a fuzzy hypergraph powered by the term-frequency and inverse-document-frequency (tf-idf) weighting scheme. Experimental results conducted on two datasets substantiate the merits of the proposed method. Indicatively, an average precision of 77% is measured at 1% recall for image-based queries. Konstantinos Pliakos, Constantine Kotropoulos |
MMSP | 2 |
| 2014 | Greek folk music denoising under a symmetric α-stable noise assumptionabstractThe noise in musical audio recordings is assumed to obey an α-stable distribution. A sparse linear regression framework with structured priors is elaborated. Markov Chain Monte Carlo is used to infer the clean music signal model and the α-stable noise distribution parameters. The musical audio recordings are processed both as a whole and in segments by using a sine-bell window for analysis and overlap-and-add reconstruction. Experiments on noisy Greek folk music excerpts demonstrate better denoising under the α-stable noise assumption than the Gaussian white noise one, and when processing is performed in segments rather than in full recordings. Nikoletta Bassiou, Constantine Kotropoulos, Ioannis Pitas |
QSHINE | 2 |
| 2014 | Elastic Net subspace clustering applied to pop/rock music structure analysis
Yannis Panagakis, Constantine Kotropoulos |
Pattern Recognit. Lett. | 2 |
| 2014 | Music genre classification via joint sparse low-rank representation of audio featuresabstractA novel framework for music genre classification, namely the joint sparse low-rank representation (JSLRR) is proposed in order to: 1) smooth the noise in the test samples, and 2) identify the subspaces that the test samples lie onto. An efficient algorithm is proposed for obtaining the JSLRR and a novel classifier is developed, which is referred to as the JSLRR-based classifier. Special cases of the JSLRR-based classifier are the joint sparse representation-based classifier and the low-rank representation-based one. The performance of the three aforementioned classifiers is compared against that of the sparse representation-based classifier, the nearest subspace classifier, the support vector machines, and the nearest neighbor classifier for music genre classification on six manually annotated benchmark datasets. The best classification results reported here are comparable with or slightly superior than those obtained by the state-of-the-art music genre classification methods. Yannis Panagakis, Constantine Kotropoulos, Gonzalo R. Arce |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Online PLSA: Batch Updating Techniques Including Out-of-Vocabulary WordsabstractA novel method is proposed for updating an already trained asymmetric and symmetric probabilistic latent semantic analysis (PLSA) model within the context of a varying document stream. The proposed method is coined online PLSA (oPLSA). The oPLSA employs a fixed-size moving window over a document stream to incorporate new documents and at the same time to discard old ones (i.e., documents that fall outside the scope of the window). In addition, the oPLSA assimilates new words that had not been previously seen (out-of-vocabulary words), and discards the words that exclusively appear in the documents to be thrown away. To handle the new words, Good-Turing estimates for the probabilities of unseen words are exploited. The experimental results demonstrate the superiority in terms of accuracy of the oPLSA over well known PLSA updating methods, such as the PLSA folding-in (PLSA fold.), the PLSA rerun from the breakpoint, the quasi-Bayes PLSA, and the Incremental PLSA. A comparison with respect to the CPU run time reveals that the oPLSA is the second fastest method after the PLSA fold. However, the better accuracy of the oPLSA than that of the PLSA fold. pays off the longer computation time. The oPLSA and the other PLSA updating methods together with online LDA are tested for document clustering and F1 scores are also reported. Nikoletta Bassiou, Constantine Kotropoulos |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Music recommendation using hypergraphs and group sparsityabstractA challenging problem in multimedia recommendation is to model a variety of relations, such as social, friend, listening, or tagging ones in a unified framework and to exploit all these sources of information. In this paper, music recommendation problem is expressed as a hypergraph ranking problem, introducing group sparsity constraints. By doing so, one can control how the different data groups (i.e., sets of hypergraph vertices) affect the recommendation process. Experiments on a dataset collected from Last.fm demonstrate that the accuracy is significantly increased by exploiting the group structure of the data. Preliminary results are also presented for Greek folk music recommendation. Antonis Theodoridis, Constantine Kotropoulos, Yannis Panagakis |
ICASSP | 2 |
| 2012 | Automatic music tagging by low-rank representationabstractA novel multi-label annotation method is proposed and applied to music tagging. Each music recording is represented by its auditory temporal modulations (ATMs). Given a set of training music recordings represented by the tag-music recording matrix having zero-one (indicator) vectors of the tags associated with each recording in its columns along with the matrix of the ATM representations in its columns, a low-rank weight matrix is sought, such that the tag-music recording matrix is expressed as the product of the weight matrix and the matrix of the ATM representations plus an error matrix. Clearly, such a weight matrix captures the relationships between the labels (i.e., tags) and the audio features. It can be derived by solving a convex nuclear norm minimization problem, if the tag-music recording matrix and the matrix of the ATM representations are assumed to be jointly low-rank. Having found the weight matrix, the annotation vector for labeling any test music recording can be obtained by multiplying the weight matrix with its ATM representation. The just outlined method is referred to as low-rank representation based multi-label annotation (LRRMA). The LRRMA outperforms the state-of-the-art auto-tagging systems, when applied to the CAL500 dataset in a 5-fold cross-validation experimental protocol. Yannis Panagakis, Constantine Kotropoulos |
ICASSP | 2 |
| 2011 | Automatic music tagging via PARAFAC2abstractAutomatic music tagging is addressed by resorting to auditory temporal modulations and Parallel Factor Analysis 2 (PARAFAC2). The starting point is to represent each music recording by its auditory temporal modulations. Then, an irregular third order tensor is formed. The first slice contains the vectorized training temporal modulations, while the second slice contains the corresponding multi-label vectors. The PARAFAC2 is employed to effectively harness the multi-label information for dimensionality reduction. Any vectorized test auditory representation of temporal modulations is first projected onto the semantic space derived via the PARAFAC2 and the coefficient vector is obtained. Then, the annotation vector is obtained by multiplying this coefficient vector by the left singular vectors of the second slice (i.e., the slice associated to the label vector). The proposed framework, outperforms the state-of-the-art auto-tagging systems, when applied to the CAL500 dataset in a 10-fold cross-validation experimental protocol. Yannis Panagakis, Constantine Kotropoulos |
ICASSP | 2 |
| 2011 | RPLSA: A novel updating scheme for Probabilistic Latent Semantic Analysis
Nikoletta Bassiou, Constantine Kotropoulos |
Comput. Speech Lang. | 2 |
| 2011 | Long distance bigram models applied to word clustering
Nikoletta Bassiou, Constantine Kotropoulos |
Pattern Recognit. | 2 |
| 2010 | Music genre classification via Topology Preserving Non-Negative Tensor Factorization and sparse representationsabstractMotivated by the rich, psycho-physiologically grounded properties of auditory cortical representations and the power of sparse representation-based classifiers, we propose a robust music genre classification framework. Its first pilar is a novel multilinear subspace analysis method that reduces the dimensionality of cortical representations of music signals, while preserving the topology of the cortical representations. Its second pilar is the sparse representation based classification, that models any test cortical representation as a sparse weighted sum of dictionary atoms, which stem from training cortical representations of known genre, by assuming that the representations of music recordings of the same genre are close enough in the tensor space they lie. Accordingly, the dimensionality reduction is made in a compatible manner to the working principle of the sparse-representation based classification. Music genre classification accuracy of 93.7% and 94.93% is reported on the GTZAN and the ISMIR2004 Genre datasets, respectively. Both accuracies outperform any accuracy ever reported for state of the art music genre classification algorithms applied to the aforementioned datasets. Yannis Panagakis, Constantine Kotropoulos |
ICASSP | 2 |
| 2010 | Word Clustering Using PLSA Enhanced with Long Distance BigramsabstractProbabilistic latent semantic analysis is enhanced with long distance bigram models in order to improve word clustering. The long distance bigram probabilities and the interpolated long distance bigram probabilities at varying distances within a context capture different aspects of contextual information. In addition, the baseline bigram, which incorporates trigger-pairs for various histories, is tested in the same framework. The experimental results collected on publicly available corpora (CISI, Cran field, Medline, and NPL) demonstrate the superiority of the long distance bigrams over the baseline bigrams as well as the superiority of the interpolated long distance bigrams against the long distance bigrams and the baseline bigram with trigger-pairs in yielding more compact clusters containing less outliers. Nikoletta Bassiou, Constantine Kotropoulos |
ICPR | 2 |
| 2010 | Ensemble Discriminant Sparse Projections Applied to Music Genre ClassificationabstractResorting to the rich, psycho-physiologically grounded, properties of the slow temporal modulations of music recordings, a novel classifier ensemble is built, which applies discriminant sparse projections. More specifically, over complete dictionaries are learned and sparse coefficient vectors are extracted to optimally approximate the slow temporal modulations of the training music recordings. The sparse coefficient vectors are then projected to the principal subspaces of their within-class and between-class covariance matrices. Decisions are taken with respect to the minimum Euclidean distance from the class mean sparse coefficient vectors, which undergo the aforementioned projections. The application of majority voting to the decisions taken by 10 individual classifiers, which are trained on the 10 training folds defined by stratified 10-fold cross-validation on the GTZAN dataset, yields a music genre classification accuracy of 84.96% on average. The latter exceeds by 2.46% the highest accuracy previously reported without employing any sparse representations. Constantine Kotropoulos, Gonzalo R. Arce, Yannis Panagakis |
ICPR | 1 |
| 2010 | Speaker Diarization Exploiting the Eigengap Criterion and Cluster EnsemblesabstractA novel system for speaker diarization is proposed that combines the eigengap criterion and cluster ensembles. No explicit assumptions on the number of speakers are made. Two variants of the system are developed. The first variant does not cluster the speech segments that are detected as outliers, while the second one does. The aforementioned system variants are assessed with respect to various metrics, such as the overall classification error, the average cluster purity, and the average speaker purity. Experiments are conducted on two-person dialogue scenes in movies as well as on news broadcasts from MDE RT-03 Training Data Speech Corpus released by the U.S. National Institute of Standards and Technology. In the latter case, the diarization error rate is also reported. It is demonstrated that the clustering performance does not degrade when outliers are present. Moreover, thanks to the eigengap criterion, the evaluation metrics are improved. Nikoletta Bassiou, Vassiliki Moschou, Constantine Kotropoulos |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Non-Negative Tensor Factorization Applied to Music Genre ClassificationabstractMusic genre classification techniques are typically applied to the data matrix whose columns are the feature vectors extracted from music recordings. In this paper, a feature vector is extracted using a texture window of one sec, which enables the representation of any 30 sec long music recording as a time sequence of feature vectors, thus yielding a feature matrix. Consequently, by stacking the feature matrices associated to any dataset recordings, a tensor is created, a fact which necessitates studying music genre classification using tensors. First, a novel algorithm for non-negative tensor factorization (NTF) is derived that extends the non-negative matrix factorization. Several variants of the NTF algorithm emerge by employing different cost functions from the class of Bregman divergences. Second, a novel supervised NTF classifier is proposed, which trains a basis for each class separately and employs basis orthogonalization. A variety of spectral, temporal, perceptual, energy, and pitch descriptors is extracted from 1000 recordings of the GTZAN dataset, which are distributed across 10 genre classes. The NTF classifier performance is compared against that of the multilayer perceptron and the support vector machines by applying a stratified 10-fold cross validation. A genre classification accuracy of 78.9% is reported for the NTF classifier demonstrating the superiority of the aforementioned multilinear classifier over several data matrix-based state-of-the-art classifiers. Emmanouil Benetos, Constantine Kotropoulos |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Non-Negative Multilinear Principal Component Analysis of Auditory Temporal Modulations for Music Genre ClassificationabstractMotivated by psychophysiological investigations on the human auditory system, a bio-inspired two-dimensional auditory representation of music signals is exploited, that captures the slow temporal modulations. Although each recording is represented by a second-order tensor (i.e., a matrix), a third-order tensor is needed to represent a music corpus. Non-negative multilinear principal component analysis (NMPCA) is proposed for the unsupervised dimensionality reduction of the third-order tensors. The NMPCA maximizes the total tensor scatter while preserving the non-negativity of auditory representations. An algorithm for NMPCA is derived by exploiting the structure of the Grassmann manifold. The NMPCA is compared against three multilinear subspace analysis techniques, namely the non-negative tensor factorization, the high-order singular value decomposition, and the multilinear principal component analysis as well as their linear counterparts, i.e., the non-negative matrix factorization, the singular value decomposition, and the principal components analysis in extracting features that are subsequently classified by either support vector machine or nearest neighbor classifiers. Three different sets of experiments conducted on the GTZAN and the ISMIR2004 Genre datasets demonstrate the superiority of NMPCA against the aforementioned subspace analysis techniques in extracting more discriminating features, especially when the training set has small cardinality. The best classification accuracies reported in the paper exceed those obtained by the state-of-the-art music genre classification algorithms applied to both datasets. Yannis Panagakis, Constantine Kotropoulos, Gonzalo R. Arce |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Information Loss of the Mahalanobis Distance in High Dimensions: Application to Feature SelectionabstractWhen an infinite training set is used, the Mahalanobis distance between a pattern measurement vector of dimensionality D and the center of the class it belongs to is distributed as a chi(2) with D degrees of freedom. However, the distribution of Mahalanobis distance becomes either Fisher or Beta depending on whether cross validation or resubstitution is used for parameter estimation in finite training sets. The total variation between chi(2) and Fisher, as well as between chi(2) and Beta, allows us to measure the information loss in high dimensions. The information loss is exploited then to set a lower limit for the correct classification rate achieved by the Bayes classifier that is used in subset feature selection. Dimitrios Ververidis, Constantine Kotropoulos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Robust Detection of Phone Boundaries Using Model Selection Criteria With Few ObservationsabstractAutomatic phone segmentation techniques based on model selection criteria are studied. We investigate the phone boundary detection efficiency of entropy- and Bayesian- based model selection criteria in continuous speech based on the DISTBIC hybrid segmentation algorithm. DISTBIC is a text-independent bottom-up approach that identifies sequential model changes by combining metric distances with statistical hypothesis testing. Using robust statistics and small sample corrections in the baseline DISTBIC algorithm, phone boundary detection accuracy is significantly improved, while false alarms are reduced. We also demonstrate further improvement in phonemic segmentation by taking into account how the model parameters are related in the probability density functions of the underlying hypotheses as well as in the model selection via the information complexity criterion and by employing M-estimators of the model parameters. The proposed DISTBIC variants are tested on the NTIMIT database and the achievedF1measure is 74.7% using a 20-ms tolerance in phonemic segmentation. George Almpanidis, Margarita Kotti, Constantine Kotropoulos |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Gender classification in two Emotional Speech databasesabstractGender classification is a challenging problem, which finds applications in speaker indexing, speaker recognition, speaker diarization, annotation and retrieval of multimedia databases, voice synthesis, smart human-computer interaction, biometrics, social robots etc. Although it has been studied for more than thirty years, by no means it is a solved problem. Processing emotional speech in order to identify speakerpsilas gender makes the problem even more interesting. A large pool of 1379 features is created including 605 novel features. A branch and bound feature selection algorithm is applied to select a subset of 15 features among the 1379 originally extracted. Support vector machines with various kernels are tested as gender classifiers, when applied to two databases, namely: the Berlin database of Emotional Speech and the Danish Emotional Speech database. The reported classification results out perform those obtained by state-of-the-art techniques, since a perfect classification accuracy is obtained. Margarita Kotti, Constantine Kotropoulos |
ICPR | 2 |
| 2008 | Speaker segmentation and clustering
Margarita Kotti, Vassiliki Moschou, Constantine Kotropoulos |
Signal Process. | 3 |
| 2008 | Fast and accurate sequential floating forward feature selection with the Bayes classifier applied to speech emotion recognition
Dimitrios Ververidis, Constantine Kotropoulos |
Signal Process. | 2 |
| 2008 | Phonemic segmentation using the generalised Gamma distribution and small sample Bayesian information criterion
George Almpanidis, Constantine Kotropoulos |
Speech Commun. | 2 |
| 2008 | Computationally Efficient and Robust BIC-Based Speaker SegmentationabstractAn algorithm for automatic speaker segmentation based on the Bayesian information criterion (BIC) is presented. BIC tests are not performed for every window shift, as previously, but when a speaker change is most probable to occur. This is done by estimating the next probable change point thanks to a model of utterance durations. It is found that the inverse Gaussian fits best the distribution of utterance durations. As a result, less BIC tests are needed, making the proposed system less computationally demanding in time and memory, and considerably more efficient with respect to missed speaker change points. A feature selection algorithm based on branch and bound search strategy is applied in order to identify the most efficient features for speaker segmentation. Furthermore, a new theoretical formulation of BIC is derived by applying centering and simultaneous diagonalization. This formulation is considerably more computationally efficient than the standard BIC, when the covariance matrices are estimated by other estimators than the usual maximum-likelihood ones. Two commonly used pairs of figures of merit are employed and their relationship is established. Computational efficiency is achieved through the speaker utterance modeling, whereas robustness is achieved by feature selection and application of BIC tests at appropriately selected time instants. Experimental results indicate that the proposed modifications yield a superior performance compared to existing approaches. Margarita Kotti, Emmanouil Benetos, Constantine Kotropoulos |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Audio-Assisted Movie Dialogue DetectionabstractAbstract—An audio-assisted system is investigated that detects if a movie scene is a dialogue or not. The system is based on actor indicator functions. That is, functions which define if an actor speaks at a certain time instant. In particular, the cross-correlation and the magnitude of the corresponding the cross-power spectral density of a pair of indicator functions are input to various classifiers, such as voted perceptrons, radial basis function networks, random trees, and support vector machines for dialogue/non-dialogue detection. To boost classifier efficiency AdaBoost is also exploited. The aforementioned classifiers are trained using ground truth indicator functions determined by human annotators for 41 dialogue and another 20 non-dialogue audio instances. For testing, actual indicator functions are derived by applying audio activity detection and actor clustering to audio recordings. 23 instances are randomly chosen among the aforementioned 41 dialogue instances, 17 of which correspond to dialogue scenes and 6 to non-dialogue ones. Accuracy ranging between 0.739 and 0.826 is reported. Index Terms—Audio activity detection, cross-correlation, crosspower spectral density, dialogue detection, indicator functions, speaker clustering. I. Margarita Kotti, Dimitrios Ververidis, Georgios Evangelopoulos, Yannis Panagakis, Constantine Kotropoulos, Petros Maragos, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2007 | Systematic comparison of BIC-based speaker segmentation systemsabstractUnsupervised speaker change detection is addressed in this paper. Three speaker segmentation systems are examined. The first system investigates the AudioSpectrumCentroid and the AudioWaveformEnvelope features, implements a dynamic fusion scheme, and applies the Bayesian Information Criterion (BIC). The second system consists of three modules. In the first module, a second-order statistic-measure is extracted; the Euclidean distance and the T2Hotelling statistic are applied sequentially in the second module; and BIC is utilized in the third module. The third system, first uses a metric-based approach, in order to detect potential speaker change points, and then the BIC criterion is applied to validate the previously detected change points. Experiments are carried out on a dataset, which is created by concatenating speakers from the TIMIT database. A systematic performance comparison among the three systems is carried out by means of one-way ANOVA method and post hoc Tukey's method. Vassiliki Moschou, Margarita Kotti, Emmanouil Benetos, Constantine Kotropoulos |
MMSP | 4 |
| 2007 | Using Adaptive Genetic Algorithms to Improve Speech Emotion RecognitionabstractIn this paper, adaptive genetic algorithms are employed to search for the worst performing features with respect to the probability of correct classification achieved by the Bayes classifier in a first stage. These features are subsequently excluded from sequential floating feature selection that employs the probability of correct classification of the Bayes classifier as criterion. In a second stage, adaptive genetic algorithms search for the worst performing utterances with respect to the same criterion. The sequential application of both stages is demonstrated to improve speech emotion recognition on the Danish Emotional Speech database. Mohammad Hosein Sedaaghi, Constantine Kotropoulos, Dimitrios Ververidis |
MMSP | 2 |
| 2007 | Color image histogram equalization by absolute discounting back-off
Nikoletta Bassiou, Constantine Kotropoulos |
Comput. Vis. Image Underst. | 2 |
| 2007 | A neural network approach to audio-assisted movie dialogue detection
Margarita Kotti, Emmanouil Benetos, Constantine Kotropoulos, Ioannis Pitas |
Neurocomputing | 3 |
| 2007 | Assessment of self-organizing map variants for clustering with application to redistribution of emotional speech patterns
Vassiliki Moschou, Dimitrios Ververidis, Constantine Kotropoulos |
Neurocomputing | 3 |
| 2007 | Combining text and link analysis for focused crawling - An application for vertical search engines
George Almpanidis, Constantine Kotropoulos, Ioannis Pitas |
Inf. Syst. | 2 |
| 2007 | Enhanced Eigen-Audioframes for Audiovisual Scene Change DetectionabstractIn this paper, a novel audio-visual scene change detection algorithm is presented and evaluated experimentally. An enhanced set of eigen-audioframes is created that is related to an audio signal subspace, where audio background changes are easily discovered. An analysis is presented that justifies why this subspace favors scene change detection. Additionally, a novel process is developed in order to detect audio scene change candidates in this subspace. Visual information is used to align audio scene change indications with neighboring video shot changes and, accordingly, to reduce the false alarm rate of the audio-only scene change detection. Moreover, video fade effects are identified and used independently in order to track scene changes. The false alarm rate is reduced further by extracting acoustic features in order to verify that the scene change indications are valid. The detection methodology was tested on newscast videos provided by the TRECVID2003 video test set. The experimental results demonstrate that the proposed method achieves an F-measure exceeding 0.85. Accordingly, it effectively tackles the scene change detection problem Marios Kyperountas, Constantine Kotropoulos, Ioannis Pitas |
IEEE Trans. Multim. | 2 |
| 2006 | Feature Selection Based on Mutual Correlation
Michal Haindl, Petr Somol, Dimitrios Ververidis, Constantine Kotropoulos |
CIARP | 4 |
| 2006 | On the Variants of the Self-Organizing Map That Are Based on Order Statistics
Vassiliki Moschou, Dimitrios Ververidis, Constantine Kotropoulos |
ICANN (1) | 3 |
| 2006 | Musical Instrument Classification using Non-Negative Matrix Factorization Algorithms and Subset Feature SelectionabstractIn this paper, a class, of algorithms for automatic classification of individual musical instrument sounds is presented. Several perceptual features used in sound classification applications as well as MPEG-7 descriptors were measured for 300 sound recordings consisting of 6 different musical instrument classes. Subsets of the feature set are selected using branch-and-bound search, obtaining the most suitable features for classification, A class of classifiers is developed based on the non-negative matrix factorization (NMF). The standard NMF method is examined as well as its modifications: the local, the sparse, and the discriminant NMF. The experimental results compare feature subsets of varying sizes alongside the various NMF algorithms. It has been found that a subset containing the mean and die variance of the first mel-frequency cepstral coefficient and the audiospectrumflatness descriptor along with the means of the audiospectrumenvelope and the audiospectrumspread descriptors when is fed to a standard NMF classifier yields an accuracy exceeding 95% Emmanouil Benetos, Margarita Kotti, Constantine Kotropoulos |
ICASSP (5) | 3 |
| 2006 | Self Organizing Maps for Reducing the Number of Clusters by One on Simplex SubspacesabstractThis paper deals with N-dimensional patterns that are represented as points on the (N - 1)-dimensional simplex. The elements of such patterns could be the posterior class probabilities for N classes, given a feature vector derived by the Bayes classifier for example. Such patterns form N clusters on the (N - 1)-dimensional simplex. We are interested in reducing the number of clusters to N - 1 in order to redistribute the features assigned to a particular class in the N - 1 simplex over the remaining N - 1 classes in an optimal manner by using a self-organizing map. An application of the proposed solution to the re-assignment of emotional speech features classified as neutral into the emotional states of anger, happiness, surprise, and sadness on the Danish emotional speech database is presented Constantine Kotropoulos, Vassiliki Moschou |
ICASSP (5) | 1 |
| 2006 | Voice Activity Detection with Generalized Gamma DistributionabstractIn this work, we model speech samples with the generalized Gamma distribution and evaluate the efficiency of such modelling for voice activity detection. Using a computationally inexpensive maximum likelihood approach, we employ the Bayesian information criterion for identifying the phoneme boundaries in noisy speech George Almpanidis, Constantine Kotropoulos |
ICME | 2 |
| 2006 | Applying Supervised Classifiers Based on Non-negative Matrix Factorization to Musical Instrument ClassificationabstractIn this paper, a new approach for automatic audio classification using non-negative matrix factorization (NMF) is presented. Training is performed onto each audio class individually, whilst during the test phase each test recording is projected onto the several training matrices. Experiments demonstrating the efficiency of the proposed approach were performed for musical instrument classification. Several perceptual features as well as MPEG-7 descriptors were measured for 300 sound recordings consisting of 6 different musical instrument classes. Subsets of the feature set were selected using branch-and-bound search, in order to obtain the most discriminating features for classification. Several NMF techniques were utilized, namely the standard NMF method, the local NMF, and the sparse NMF. The experiments demonstrate an almost perfect classification (classification error 1.0%), outperforming the state-of-the-art techniques tested for the aforementioned experiment Emmanouil Benetos, Margarita Kotti, Constantine Kotropoulos |
ICME | 3 |
| 2006 | Automatic Speaker Segmentation using Multiple Features and Distance Measures: A Comparison of Three ApproachesabstractThis paper addresses the problem of unsupervised speaker change detection. Three systems based on the Bayesian Information Criterion (BIC) are tested. The first system investigates the AudioSpectrumCentroid and the AudioWaveformEnvelope features, implements a dynamic thresholding followed by a fusion scheme, and finally applies BIC. The second method is a real-time one that uses a metric-based approach employing the line spectral pairs and the BIC to validate a potential speaker change point. The third method consists of three modules. In the first module, a measure based on second-order statistics is used; in the second module, the Euclidean distance and T2 Hotelling statistic are applied; and in the third module, the BIC is utilized. The experiments are carried out on a dataset created by concatenating speakers from the TIMIT database, that is referred to as the TIMIT data set. A comparison between the performance of the three systems is made based on t-statistics. Margarita Kotti, Luis P. M. Martins, Emmanouil Benetos, Jaime S. Cardoso 0001, Constantine Kotropoulos |
ICME | 5 |
| 2006 | Musical instrument classification using non-negative matrix factorization algorithmsabstractIn this paper, a class of algorithms for automatic classification of individual musical instrument sounds is presented. Several perceptual features used in general sound classification applications were measured for 300 sound recordings consisting of 6 different musical instrument classes (piano, violin, cello, flute, bassoon and soprano saxophone). In addition, MPEG-7 basic spectral and spectral basis descriptors were considered, providing an effective combination for accurately describing the spectral and timbral audio characteristics. The audio files were split using 70% of the available data for training and the remaining 30% for testing. A classifier was developed based on non-negative matrix factorization (NMF) techniques, thus introducing a novel application of NMF. The standard NMF method was examined, as well as its modifications: the local, the sparse, and the discriminant NMF. Experimental results are presented to compare MPEG-7 spectral basis representations with MPEG-7 basic spectral features alongside the various NMF algorithms. The results indicate that the use of the spectrum projection coefficients for feature extraction and the standard NMF classifier yields an accuracy exceeding 95% Emmanouil Benetos, Margarita Kotti, Constantine Kotropoulos |
ISCAS | 3 |
| 2006 | Automatic speaker change detection with the Bayesian information criterion using MPEG-7 features and a fusion schemeabstractThis paper addresses unsupervised speaker change detection, a necessary step for several indexing tasks. We assume that there is no prior knowledge either on the number of speakers or their identities. Features included in the MPEG-7 audio prototype are investigated such as the AudioWaveformEnvelope and the AudioSpectrumCentroid. The model selection criterion is the Bayesian information criterion (BIC). A multiple pass algorithm is proposed. It uses a dynamic thresholding for scalar features and a fusion scheme so as to refine the segmentation results. It also models every speaker by a multivariate Gaussian probability density function and whenever new information is available, the respective model is updated. The experiments are carried out on a dataset created by concatenating speakers from the TIMIT database, that is referred to as the TIMIT data set. It is and demonstrated that the performance of the proposed multiple pass algorithm is better than that of other approaches Margarita Kotti, Emmanouil Benetos, Constantine Kotropoulos |
ISCAS | 3 |
| 2006 | Demonstrating the stability of support vector machines for classification
Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas |
Signal Process. | 2 |
| 2006 | Emotional speech recognition: Resources, features, and methods
Dimitrios Ververidis, Constantine Kotropoulos |
Speech Commun. | 2 |
| 2005 | Modified Gauss-Seidel affine projection algorithm for acoustic echo cancellationabstractThe paper proposes a robust and stable fast affine projection algorithm based on the Gauss-Seidel method, the so called modified Gauss-Seidel fast affine projection algorithm. The proposed algorithm is generalized for simplified Volterra filters as well. The computational complexity of both the modified Gauss-Seidel fast affine projection algorithm and its generalization for simplified Volterra filters is derived and their performance for acoustic echo cancellation is assessed. Felix Albu, Constantine Kotropoulos |
ICASSP (3) | 2 |
| 2005 | Emotional Speech Classification Using Gaussian Mixture Models and the Sequential Floating Forward Selection AlgorithmabstractEmotional speech classification can be treated as a supervised learning task where the statistical properties of emotional speech segments are the features and the emotional styles form the labels. The Akaike criterion is used for estimating automatically the number of Gaussian densities that model the probability density function of the emotional speech features. A procedure for reducing the computational burden of crossvalidation in sequential floating forward selection algorithm is proposed that applies the t-test on the probability of correct classification for the Bayes classifier designed for various feature sets. For the Bayes classifier, the sequential floating forward selection algorithm is found to yield a higher probability of correct classification by 3% than that of the sequential forward selection algorithm either taking into account the gender information or ignoring it. The experimental results indicate that the utterances from isolated words and sentences are more colored emotional than those from paragraphs. Without taking into account the gender information, the probability of correct classification for the Bayes classifier admits a maximum when the probability density function of emotional speech features extracted from the aforementioned utterances is modeled as a mixture of 2 Gaussian densities Dimitrios Ververidis, Constantine Kotropoulos |
ICME | 2 |
| 2004 | Audio PCA in a novel multimedia scheme for scene change detectionabstractA novel scene change detection algorithm is proposed in this paper that exploits both audio and video information. Audio frames are projected to the eigenspace and their distance from a reference noise eigenframe is calculated. An analysis is presented that explains why this subspace favors scene cut detection. Video information is used to align audio scene change indications with neighboring shot changes in the visual data, and accordingly to reduce the false alarm rate. Moreover, video fade effects are identified and used independently in order to track scene changes. The detection technique was tested on newscast videos provided by the TRECVID 2003 video test set. The experimental results show that the aforementioned methods used to process the audio and video information complement each other well in tackling the scene change detection problem. Marios Kyperountas, Zuzana Cernekova, Constantine Kotropoulos, Marios A. Gavrielides, Ioannis Pitas |
ICASSP (4) | 3 |
| 2004 | Automatic emotional speech classificationabstractOur purpose is to design a useful tool which can be used in psychology to automatically classify utterances into five emotional states such as anger, happiness, neutral, sadness, and surprise. The major contribution of the paper is to rate the discriminating capability of a set of features for emotional speech recognition. A total of 87 features has been calculated over 500 utterances from the Danish Emotional Speech database. The sequential forward selection method (SFS) has been used in order to discover a set of 5 to 10 features which are able to classify the utterances in the best way. The criterion used in SFS is the cross-validated correct classification score of one of the following classifiers: nearest mean and Bayes classifier where class pdf are approximated via Parzen windows or modelled as Gaussians. After selecting the 5 best features, we reduce the dimensionality to two by applying principal component analysis. The result is a 51.6% /spl plusmn/ 3% correct classification rate at 95% confidence interval for the five aforementioned emotions, whereas a random classification would give a correct classification rate of 20%. Furthermore, we find out those two-class emotion recognition problems whose error rates contribute heavily to the average error and we indicate that a possible reduction of the error rates reported in this paper would be achieved by employing two-class classifiers and combining them. Dimitrios Ververidis, Constantine Kotropoulos, Ioannis Pitas |
ICASSP (1) | 2 |
| 2004 | MPEG-4 compliant reproduction of face animation created in MayaabstractThis work presents a method for extracting facial information encoded in MPEG-4 format from an animated face model in Maya. The method generates the appropriate data, as specified in the MPEG-4 standard, such as the face definition parameters (FDPs), the face animation parameters (FAPs) and the facial animation table (FAT). The described procedure was implemented as a plug-in for Maya, which requires as inputs only the animated face model and the correspondence between its vertices and the FDPs. The extracted data were tested on a publicly available FAP-player in order to demonstrate the faithful reproduction of the face animation originally created with the high-end 3D graphics modeller. The capabilities of the publicly available FAP player have also been augmented in order to enable loading any proprietary face model and its corresponding FAT. Charalampos Laftsidis, Constantine Kotropoulos, Ioannis Pitas |
ICME | 2 |
| 2004 | Searching relevant syllable context by clustering for alignment in modern Greek speechabstractWe present some results on the optimal design of syllable databases for modern Greek. We show that speaker gender or stress are not crucial criteria for voiced syllables, whereas distance to silence has to be taken into account. Such syllable databases are exploited in an alignment system based on dynamic time warping for modern Greek speech. The overall architecture of the alignment system is also presented. R. Rispoli, Constantine Kotropoulos, Ioannis Pitas |
ICME | 2 |
| 2004 | Automatic detection of vocal fold paralysis and edemaabstractIn this paper we propose a combined scheme of linear prediction analysis for feature extraction along with linear projection methods for feature reduction followed by known pattern recognition methods on the purpose of discriminating between normal and pathological voice samples. Two different cases of speech under vocal fold pathology are examined: vocal fold paralysis and vocal fold edema. Three known classifiers are tested and compared in both cases, namely the Fisher linear discriminant, the -nearest neighbor classifier, and the nearest mean classifier. The performance of each classifier is evaluated in terms of the probabilities of false alarm and detection or the receiver operating characteristic. The datasets used are part of a database of disordered speech developed by Massachusetts Eye and Ear Infirmary. The experimental results indicate that vocal fold paralysis and edema can easily be detected by any of the aforementioned classifiers. Maria Marinaki, Constantine Kotropoulos, Ioannis Pitas, Nicos Maglaveras |
INTERSPEECH | 2 |
| 2004 | Marginal median SOM for document organization and retrieval
Apostolos Georgakis, Constantine Kotropoulos, Alexandros Xafopoulos, Ioannis Pitas |
Neural Networks | 2 |
| 2004 | Language identification in web documents using discrete HMMs
Alexandros Xafopoulos, Constantine Kotropoulos, George Almpanidis, Ioannis Pitas |
Pattern Recognit. | 2 |
| 2003 | CAML - A Universal Configuration Language for Dialogue Systems
Gergely Kovásznai, Constantine Kotropoulos, Ioannis Pitas |
DEXA | 2 |
| 2003 | Video shot segmentation using singular value decompositionabstractA new method for detecting shot boundaries in video sequences using singular value decomposition (SVD) is proposed. The method relies on performing singular value decomposition on the matrix A created from 3D histograms of single frames. We have used SVD for its capabilities to derive a low dimensional refined feature space from a high dimensional raw feature space, where pattern similarity can easily be detected. The method can detect cuts and gradual transitions, such as dissolves and fades, which cannot be detected easily by entropy measures. Zuzana Cernekova, Constantine Kotropoulos, Ioannis Pitas |
ICASSP (3) | 2 |
| 2003 | ICA and Gabor representation for facial expression recognitionabstractTwo hybrid systems for classifying seven categories of human facial expression are proposed. The first system combines independent component analysis (ICA) and support vector machines (SVMs). The original face image database is decomposed into linear combinations of several basis images, where the corresponding coefficients of these combinations are fed up into SVMs instead of an original feature vector comprised of grayscale image pixel values. The classification accuracy of this system is compared against that of baseline techniques that combine ICA with either two-class cosine similarity classifiers or two-class maximum correlation classifiers, when we classify facial expressions into these seven classes. We found that, ICA decomposition combined with SVMs outperforms the aforementioned baseline classifiers. The second system proposed operates in two steps: first, a set of Gabor wavelets (GWs) is applied to the original face image database and, second, the new features obtained are classified by using either SVMs or cosine similarity classifiers or maximum correlation classifier. The best facial expression recognition rate is achieved when Gabor wavelets are combined with SVMs. Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas |
ICIP (2) | 2 |
| 2003 | Video shot segmentation using singular value decompositionabstractA new method for detecting shot boundaries in video sequences using singular value decomposition (SVD) is proposed. The method relies on performing singular value decomposition on the matrix A created from 3D histograms of single frames. We have used SVD for its capabilities to derive a low dimensional refined feature space from a high dimensional raw feature space, where pattern similarity can easily be detected. The method can detect cuts and gradual transitions, such as dissolves and fades, which cannot be detected easily by entropy measures. Zuzana Cernekova, Constantine Kotropoulos, Ioannis Pitas |
ICME | 2 |
| 2003 | Segmentation of ultrasonic images using Support Vector Machines
Constantine Kotropoulos, Ioannis Pitas |
Pattern Recognit. Lett. | 1 |
| 2003 | Preface
Constantine Kotropoulos, Ioannis Pitas, Athina P. Petropulu |
Pattern Recognit. Lett. | 1 |
| 2002 | A SOM Variant Based on the Wilcoxon Test for Document Organization and Retrieval
Apostolos Georgakis, Constantine Kotropoulos, Ioannis Pitas |
ICANN | 2 |
| 2002 | Recent advances in biometric person authenticationabstractBiometrics is an emerging topic in the field of signal processing. While technologies (e.g. audio, video) for biometrics have mostly been studied separately, ultimately, biometric technologies could find their strongest role as interwined and complementary pieces of a multi-modal authentication system. In this paper, a short overview of voice, fingerprint, and face authentication algorithms is provided. Jean-Luc Dugelay, Jean-Claude Junqua, Constantine Kotropoulos, Roland Kuhn 0001, Florent Perronnin, Ioannis Pitas |
ICASSP | 3 |
| 2002 | On the stability of support vector machines for face detectionabstractIn this paper we study the stability of support vector machines in face detection by decomposing their average prediction error into the bias, variance, and aggregation effect terms. Such an analysis indicates whether bagging, a method for generating multiple versions of a classifier from bootstrap samples of a training set, and combining their outcomes by majority voting, is expected to improve the accuracy of the classifier. We estimate the bias, variance, and aggregation effect by using bootstrap smoothing techniques when support vector machines are applied to face detection in the AT & T face database and we demonstrate that support vector machines are stable classifiers. Accordingly, bagging is not expected to improve their face detection accuracy. Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas |
ICIP (3) | 2 |
| 2002 | Application of support vector machines classifiers to visual speech recognitionabstractIn this paper we propose a visual speech recognition network based on support vector machines. Each word of the dictionary is modeled by a set of temporal sequences of visemes. Each viseme is described by a support vector machine, and the temporal character of speech is modeled by integrating the support vector machines as nodes into a Viterbi decoding lattice. Experiments conducted on a small visual speech recognition task using very simple features demonstrate a word recognition rate on the level of the best rates previously reported even without training the state transition probabilities in the Viterbi lattices. This proves the suitability of support vector machines for visual speech recognition. Mihaela Gordan, Constantine Kotropoulos, Apostolos Georgakis, Ioannis Pitas |
ICIP (3) | 2 |
| 2002 | Face verification using elastic graph matching based on morphological signal decomposition
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas |
Signal Process. | 2 |
| 2001 | Frontal face detection using support vector machines and back-propagation neural networksabstractFace detection is a key problem in building systems that perform face recognition/verification and model-based image coding. Two algorithms for face detection that employ either support vector machines or backpropagation feedforward neural networks are described, and their performance is tested on the same frontal face database using the false acceptance and false rejection rates as quantitative figures of merit. The aforementioned algorithms can replace the explicitly-defined knowledge for facial regions and facial features in mosaic-based face detection algorithms. Nikoletta Bassiou, Constantine Kotropoulos, T. Kosmidis, Ioannis Pitas |
ICIP (1) | 2 |
| 2001 | Combining support vector machines for accurate face detectionabstractThe paper proposes the application of majority voting on the output of several support vector machines in order to select the most suitable learning machine for frontal face detection. The first experimental results indicate a significant reduction of the rate of false positive patterns. Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas |
ICIP (1) | 2 |
| 2001 | Using Support Vector Machines to Enhance the Performance of Elastic Graph Matching for Frontal Face AuthenticationabstractA novel method for enhancing the performance of elastic graph matching in frontal face authentication is proposed. The starting point is to weigh the local similarity values at the nodes of an elastic graph according to their discriminatory power. Powerful and well-established optimization techniques are used to derive the weights of the linear combination. More specifically, we propose a novel approach that reformulates Fisher's discriminant ratio to a quadratic optimization problem subject to a set of inequality constraints by combining statistical pattern recognition and support vector machines (SVM). Both linear and nonlinear SVM are then constructed to yield the optimal separating hyperplanes and the optimal polynomial decision surfaces, respectively. The method has been applied to frontal face authentication on the M2VTS database. Experimental results indicate that the performance of morphological elastic graph matching is highly improved by using the proposed weighting technique. Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | Face authentication by using elastic graph matching and support vector machinesabstractA novel method for enhancing the performance of elastic graph matching in face authentication is proposed. The starting point is to weigh the local matching errors at the nodes of an elastic graph according to their discriminatory power. We propose a novel approach to discriminant analysis that re-formulates Fisher's linear discriminant ratio to a quadratic optimization problem subject to inequality constraints by combining statistical pattern recognition and support vector machines. The method is applied to frontal face authentication on the M2VTS database. Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas |
ICASSP | 2 |
| 2000 | Using Support Vector Machines for Face Authentication Based on Elastic Graph MatchingabstractA novel method for enhancing the performance of elastic graph matching in face authentication is proposed. Our objective is to weigh the local matching errors at the nodes of an elastic graph according to their discriminatory power. We propose a novel approach to discriminant analysis that re-formulates Fisher's linear discriminant ratio to a quadratic optimization problem subject to inequality constraints by combining statistical pattern recognition and support vector machines. The method is applied to frontal face authentication on the M2VTS database. Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas |
ICIP | 2 |
| 2000 | Comparison of Face Verification Results on the XM2VTS DatabaseabstractPresents results of the face verification contest that was organized in conjunction with International Conference on Pattern Recognition 2000. Participants had to use identical data sets from a large, publicly available multimodal database XM2VTSDB. Training and evaluation was carried out according to an a priori known protocol. Verification results of all tested algorithms have been collected and made public on the XM2VTSDB website, facilitating large scale experiments on classifier combination and fusion. Tested methods included, among others, representatives of the most common approaches to face verification -elastic graph matching, Fisher's linear discriminant and support vector machines. Jiri Matas, Miroslav Hamouz, Kenneth Jonsson, Josef Kittler, Yongping Li, Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas, Teewoon Tan, Hong Yan 0001, Fabrizio Smeraldi, N. Capdevielle, Wulfram Gerstner, Yousri Abdeljaoued, Josef Bigün, Souheil Ben Yacoub, Eddy Mayoraz |
ICPR | 6 |
| 2000 | Morphological elastic graph matching applied to frontal face authentication under well-controlled and real conditions
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. | 1 |
| 2000 | Frontal face authentication using morphological elastic graph matchingabstractA novel dynamic link architecture based on multiscale morphological dilation-erosion is proposed for frontal face authentication. Instead of a set of Gabor filters tuned to different orientations and scales, multiscale morphological operations are employed to yield a feature vector at each node of the reference grid. Linear projection algorithms for feature selection and automatic weighting of the nodes according to their discriminatory power succeed to increase the authentication capability of the method. The performance of the morphological dynamic link architecture is evaluated in terms of the receiver operating characteristic in the M2VTS face image database. The comparison with other frontal face authentication algorithms indicates that the morphological dynamic link architecture with discriminatory power coefficients is the best algorithm with respect to the equal error rate achieved. Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 2000 | Frontal Face Authentication Using Discriminating Grids with Morphological Feature VectorsabstractA novel elastic graph matching procedure based on multiscale morphological operations, the so called morphological dynamic link architecture, is developed for frontal face authentication. Fast algorithms for implementing mathematical morphology operations are presented. Feature selection by employing linear projection algorithms is proposed. Discriminatory power coefficients that weigh the matching error at each grid node are derived. The performance of morphological dynamic link architecture in frontal face authentication is evaluated in terms of the receiver operating characteristic on the M2VTS face image database. Preliminary results for face recognition using the proposed technique are also presented. Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Multim. | 1 |
| 1999 | Compensating for variable recording conditions in frontal face authentication algorithmsabstractThis paper addresses the problem of compensating for variable recording conditions such as changes in illumination, scale differences, and varying face position. It is well known that the performance of any face authentication/recognition algorithm deteriorates significantly in the presence of the aforementioned conditions as well as the expression variations. The use of simple and powerful pre-processing techniques aiming at compensating for variable recording conditions prior to the application of any authentication algorithm is proposed. It is shown that such an approach overcomes indeed the image variations and guarantees an almost stable performance for the Morphological Dynamic Link Architecture developed within the European research project M2VTS. Anastasios Tefas, Yann Menguy, Constantine Kotropoulos, Gaël Richard, Ioannis Pitas, Philip Lockwood |
ICASSP | 3 |
| 1998 | Variants of Dynamic Link Architecture Based on Mathematical Morphology for Frontal Face AuthenticationabstractTwo novel variants of dynamic link architecture that are based on mathematical morphology and incorporate coefficients which weigh the contribution of each node in elastic graph matching according to its discriminatory power are developed. They are the so called Morphological Dynamic Link Architecture and the Morphological Signal Decomposition-Dynamic Lint Architecture. The proposed variants are tested for face authentication in a cooperative scenario where the candidates claim an identity to be checked. Their performance is evaluated in terms of their receiver operating characteristic and the equal error rate achieved in M2VTS database. An equal error rate in the range 3.7-6.8% is reported. Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas |
CVPR | 2 |
| 1998 | Face Verification based on Morphological Shape Decomposition
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas |
FG | 2 |
| 1998 | Face authentication using variants of elastic graph matching based on mathematical morphology that incorporate local discriminant coefficientsabstractTwo novel variants of dynamic link architecture that are based on mathematical morphology and incorporate local coefficients which weigh the contribution of each node according to its discriminatory power in elastic graph matching are proposed, namely, the morphological dynamic link architecture and the morphological signal decomposition-dynamic link architecture. They are tested for face authentication in a cooperative scenario where the candidates claim an identity to be checked. Their performance is evaluated in terms of their receiver operating characteristics and the equal error rate achieved in the M2VTS database. An equal error rate of 6.6%-6.8% is reported. Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas |
ICASSP | 1 |
| 1998 | Frontal Face Authentication using Variants of Dynamic Link Matching Based on Mathematical MorphologyabstractTwo variants of dynamic link matching based on mathematical morphology are developed and tested for frontal face authentication, namely, the morphological dynamic link architecture and the morphological signal decomposition-dynamic link architecture. Local coefficients which weigh the contribution of each node in elastic graph matching according to its discriminatory power are derived. The performance of the proposed algorithms is evaluated in terms of their receiver operating characteristic and the equal error rate (EER) achieved in the M2VTS database. The comparison with other frontal face authentication algorithms developed within M2VTS project indicates that morphological dynamic link architecture with discriminatory power coefficients is ranked as the best algorithm in terms of the EER. Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas |
ICIP (1) | 1 |
| 1997 | Rule-based face detection in frontal viewsabstractFace detection is a key problem in building automated systems that perform face recognition. A very attractive approach for face detection is based on multiresolution images (also known as mosaic images). Motivated by the simplicity of this approach, a rule-based face detection algorithm in frontal views is developed that extends the work of G. Yang and T.S. Huang (see Pattern Recognition, vol.27, no.1, p.53-63, 1994). The proposed algorithm has been applied to frontal views extracted from the European ACTS M2VTS database that contains the videosequences of 37 different persons. It has been found that the algorithm provides a correct facial candidate in all cases. However, the success rate of the detected facial features (e.g. eyebrows/eyes, nostrils/nose, and mouth) that validate the choice of a facial candidate is found to be 86.5% under the most strict evaluation conditions. Constantine Kotropoulos, Ioannis Pitas |
ICASSP | 1 |
| 1997 | Face Authentication Based on Morphological Grid MatchingabstractA novel dynamic link architecture based on multiscale morphological dilation-erosion is proposed for face verification in a cooperative scenario where the candidates claim an identity that is to be checked. The performance of the morphological dynamic link architecture (MDLA) is evaluated in terms of the receiver operating characteristic (ROC) for several threshold selections on the matching error in the M2VTS database. The experimental results indicate that the proposed method outperforms the dynamic link matching with Gabor based feature vectors. Constantine Kotropoulos, Ioannis Pitas |
ICIP (1) | 1 |
| 1996 | Multichannel adaptive L-filters in color image filteringabstractThree novel adaptive multichannel L-filters based on marginal data ordering are proposed. They rely on well-known algorithms for the unconstrained minimization of the mean squared error (MSE), namely, the least mean squares (LMS), the normalized LMS (NLMS) and the LMS-Newton (LMSN) algorithm. Performance comparisons in color image filtering have been made both in RGB and U/sup */V/sup */W/sup */ color spaces. The proposed adaptive multichannel L-filters outperform the other candidates in noise suppression for color images corrupted by mixed impulsive and additive white contaminated Gaussian noise. Constantine Kotropoulos, Ioannis Pitas, Maria Gabrani |
ICIP (1) | 1 |
| 1996 | Nonlinear adaptive filters for speckle suppression in ultrasonic images
Eleftherios Kofidis, Sergios Theodoridis, Constantine Kotropoulos, Ioannis Pitas |
Signal Process. | 3 |
| 1996 | Adaptive LMS L-filters for noise suppression in imagesabstractSeveral adaptive least mean squares (LMS) L-filters, both constrained and unconstrained ones, are developed for noise suppression in images and compared in this paper. First, the location-invariant LMS L-filter for a nonconstant signal corrupted by zero-mean additive white noise is derived. It is demonstrated that the location-invariant LMS L-filter can be described in terms of the generalized linearly constrained adaptive processing structure proposed by Griffiths and Jim (1982). Subsequently, the normalized and the signed error LMS L-filters are studied. A modified LMS L-filter with nonhomogeneous step-sizes is also proposed in order to accelerate the rate of convergence of the adaptive L-filter. Finally, a signal-dependent adaptive filter structure is developed to allow a separate treatment of the pixels that are close to the edges from the pixels that belong to homogeneous image regions. Constantine Kotropoulos, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 1996 | Order statistics learning vector quantizerabstractWe propose a novel class of learning vector quantizers (LVQs) based on multivariate data ordering principles. A special case of the novel LVQ class is the median LVQ, which uses either the marginal median or the vector median as a multivariate estimator of location. The performance of the proposed marginal median LVQ in color image quantization is demonstrated by experiments. Ioannis Pitas, Constantine Kotropoulos, Nikos Nikolaidis 0001, Ruikang Yang, Moncef Gabbouj |
IEEE Trans. Image Process. | 2 |
| 1994 | Multichannel L-filter design based on marginal data orderingabstractWe address the design of multichannel L-filters based on marginal data ordering using the mean-squared-error as fidelity criterion. Design procedures subject to the constraints of unbiased or location-invariant estimation or without imposing any constraint are discussed. It is shown by simulations that the proposed multichannel L-filters perform better than other multichannel nonlinear filters such as the vector median, the marginal alpha-trimmed mean, the marginal median, the multichannel modified trimmed mean and the multichannel double-window trimmed mean, the multivariate ranked-order estimators as well as their single-channel counterparts.> Constantine Kotropoulos, Ioannis Pitas |
ICASSP (3) | 1 |
| 1994 | Cellular LMS L-filters for Noise Suppression in Still Images and Image SequencesabstractA novel class of nonlinear adaptive L-filters based on cellular neural networks topology is presented. Like cellular neural systems and cellular automata as well, processing nodes, called cells, communicate with each other directly only through its nearest neighbors exchanging information. Each cell is an adaptive LMS L-filter. The proposed filters share the best features of both adaptive filters and cellular neural network topologies; their adaptive structure tracks image nonstationarities and their local interconnection feature makes it suitable for VLSI implementation. Cellular adaptive LMS L-filters are suited for high-speed parallel adaptive image filtering. Some interesting applications to image and image sequence filtering are demonstrated.> Maria Gabrani, Constantine Kotropoulos, Ioannis Pitas |
ICIP (1) | 2 |
| 1994 | Adaptive LMS L-filters for smoothing noisy imagesabstractSeveral adaptive LMS L-filters, both constrained and unconstrained ones, are developed for noise suppression in images and being compared in this paper. First, the location-invariant LMS L-filter for a nonconstant signal corrupted by zero-mean additive white noise is derived. Subsequently, the normalized and the sign LMS L-filters are studied. It is shown that both these filters turn to be identical for a certain choice of the adaptation step-size. A modified LMS L-filter with nonhomogeneous step-sizes is also proposed in order to accelerate the rate of convergence of the adaptive L-filter. Finally, a signal-dependent adaptive filter structure is developed to allow a separate treatment of the pixels that are close to the edges from the pixels that belong to homogeneous image regions. Constantine Kotropoulos, Ioannis Pitas |
ICPR (3) | 1 |
| 1994 | A Class of Order Statistics Learning Vector QuantizersabstractA novel class of Learning Vector Quantizers (LVQs) based on multivariate order statistics is proposed in order to overcome the drawback that the estimators for obtaining the reference vectors in LVQ do not have robustness either against erroneous choices for the winner vector or against the outliers that may exist in vector-valued observations. The performance of the proposed variants of LVQ is demonstrated by experiments. In the case of marginal median LVQ, its asymptotic properties are derived as well.> Ioannis Pitas, Constantine Kotropoulos, Nikos Nikolaidis 0001, Ruikang Yang, Moncef Gabbouj |
ISCAS | 2 |
| 1994 | Nonlinear ultrasonic image processing based on signal-adaptive filters and self-organizing neural networksabstractTwo approaches for ultrasonic image processing are examined. First, signal-adaptive maximum likelihood (SAML) filters are proposed for ultrasonic speckle removal. It is shown that in the case of displayed ultrasound (US) image data the maximum likelihood (ML) estimator of the original (noiseless) signal closely resembles the L(2) mean which has been proven earlier to be the ML estimator of the original signal in US B-mode data. Thus, the design of signal-adaptive L(2) mean filters is treated for US B-mode data and displayed US image data as well. Secondly, the segmentation of ultrasonic images using self-organizing neural networks (NN) is investigated. A modification of the learning vector quantizer (L(2 ) LVQ) is proposed in such a way that the weight vectors of the output neurons correspond to the L(2) mean instead of the sample arithmetic mean of the input observations. The convergence in the mean and in the mean square of the proposed L(2) LVQ NN are studied. L(2) LVQ is combined with signal-adaptive filtering in order to allow preservation of image edges and details as well as maximum speckle reduction in homogeneous regions. Constantine Kotropoulos, Xanthippos C. Magnisalis, Ioannis Pitas, Michael G. Strintzis |
IEEE Trans. Image Process. | 1 |
| 1993 | A Variant of Learning Vector Quantizer Based on Split-Merge Statistical Tests
Constantine Kotropoulos, Ioannis Pitas |
CAIP | 1 |
| 1993 | Voronoi tessellation and Delauney triangulation using Euclidean disk growing in Z2
Constantine Kotropoulos, Ioannis Pitas, A. Maglara |
ICASSP (5) | 1 |
| 1993 | A variant of learning vector quantizer based on the L2 mean for segmentation of ultrasonic images
Constantine Kotropoulos, Ioannis Pitas, Xanthippos C. Magnisalis, Michael G. Strintzis |
ISCAS | 1 |
| 1992 | A texture-based approach to the segmentation of seismic images
Ioannis Pitas, Constantine Kotropoulos |
Pattern Recognit. | 2 |
| 1992 | Constrained adaptive LMS L-filters
Constantine Kotropoulos, Ioannis Pitas |
Signal Process. | 1 |
| 1991 | Constrained adaptive LMS L-filtersabstractTwo novel adaptive nonlinear filter structures are proposed which are based on linear combinations of order statistics. These adaptive schemes are modifications of the standard LMS (least mean square) algorithm and have the ability to incorporate constraints imposed on coefficients in order to permit location invariant and unbiased estimation of a constant signal in the presence of additive white noise. The convergence properties of the proposed filters are considered. Both of them can adapt well to a variety of noise probability distributions ranging from short-tailed to long-tailed ones. Simulation examples are given.> Constantine Kotropoulos, Ioannis Pitas |
ICASSP | 1 |
| 1989 | Texture analysis and segmentation of seismic imagesabstractA method is proposed for the texture analysis and segmentation of geophysical images. It is based on the detection of the seismic horizons and on the calculation of their features (e.g. length, average reflection strength, signature). These features represent the texture of the seismic image. The horizons are clustered into classes according to one or several of their features. Each cluster represents a distinct texture characteristic of the seismic image. After this initial clustering, the points of each horizon are used as seeds for geophysical image segmentation. All pixels in the seismic image are clustered in those classes, according to their geometric proximity to points lying on classified horizons. Thus the entire seismic image is classified on the basis of seismic texture patterns. Two methods are proposed for clustering pixels according to their geometric proximity to reference points. The first is based on Voronoi tesselation and mathematical morphology. The second is based on a so-called radiation model for region growing. Simulation examples are presented.> Ioannis Pitas, Constantine Kotropoulos |
ICASSP | 2 |