Stefanos D. Kollias

dblp:k/StefanosDKollias · DBLP profile ↗
← Back
158ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0003-2899-0598ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 71 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 10Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 EdgeScan for IoT Contextual Understanding With Edge Computing and Image Captioning
abstract
The emergence of Edge Computing has shifted the processing capabilities in proximity to the Internet of Things (IoT) data sources, offering solutions to latency and bandwidth constraints applications. This shift complements Cloud Computing, especially in handling real-time data processing and enhancing processing. Image processing, particularly image captioning for smart monitoring systems, benefits greatly from this synergy. Image captioning plays a crucial role in understanding visual data. While early methods excelled in encoder-decoder frameworks and attention mechanisms, they often overlooked semantic representations which are essential for comprehensive image understanding. To address this gap, we introduce the EdgeScan framework, leveraging Edge Computing for image analysis and semantic feature extractions closer to data sources. EdgeScan integrates visual and semantic features to create more informative and enriched image captions that enhance image captioning accuracy. The EdgeScan image captioning model architecture is capable of 1) learning the salient image region’s specific feature representation and 2) co-embedding visual attention and semantic attributes in one space for feature fusion. This improves models’ ability to interpret and respond to data in a meaningful way, which is particularly valuable for IoT applications that require a deep understanding of the semantics of diverse and constantly changing data for efficient operation. Extensive experiments were conducted on the MS-COCO dataset to demonstrate the superiority of EdgeScan in both quantitative and qualitative performance, achieving highest consensus-based image description evaluation score of 120.9, as well as notable scores of 78.6 for BLEU@1 and 57.7 for recall-oriented understudy for gisting evaluation metrics, promising advancements in IoT-driven image understanding and competitiveness against the state of the art.
Deema Abdal Hafeth, Mohammed Al-Khafajiy, Stefanos D. Kollias
IEEE Internet Things J.3
2024 Uncertainty-Guided Contrastive Learning For Single Source Domain Generalisation
abstract
In the context of single domain generalisation, the objective is for models that have been exclusively trained on data from a single domain to demonstrate strong performance when confronted with various unfamiliar domains. In this paper, we introduce a novel model referred to as Contrastive Uncertainty Domain Generalisation Network (CUDGNet). The key idea is to augment the source capacity in both input and label spaces through the fictitious domain generator and jointly learn the domain invariant representation of each class through contrastive learning. Extensive experiments on two Single Source Domain Generalisation (SSDG) datasets demonstrate the effectiveness of our approach, which surpasses the state-of-the-art single-DG methods by up to 7.08%. Our method also provides efficient uncertainty estimation at inference time from a single forward pass through the generator subnetwork.
Anastasios Arsenos, Dimitris Kollias, Evangelos Petrongonas, Christos Skliros, Stefanos D. Kollias
ICASSP5
2023 Cloud-IoT Application for Scene Understanding in Assisted Living: Unleashing the Potential of Image Captioning and Large Language Model (ChatGPT)
abstract
Vision is a vital sense that plays a pivotal role in our understanding of the world. The majority of our external information is acquired through our visual system, which significantly impacts various aspects of our lives, including mobility, cognitive abilities, access to information, and how we interact with both our surroundings and other individuals. Hence, individuals who need assisted living due to visual challenges are left behind and rely on human-driven image captioning services to make sense of their surroundings. In response to this challenge, we have developed a proof-of-concept system that integrates a large language model like ChatGPT to provide assistance to individuals with visual impairments in their daily lives through the utilisation of image captioning techniques. Our proposed model leverages the image captioning technique to describe the user’s environment. It is a fusion of concepts from Deep Learning and the Internet of Things, enabling it to provide more informative and enriched image captions. In this process, ChatGPT is stimulated to generate increasingly detailed and informative descriptions of images, allowing users to gain a deeper understanding of their surroundings. Our findings show that the proposed system generates captions that are contextually relevant to the visual content. These captions can assist individuals in various day-today activities, contributing to an improved quality of life.
Deema Abdal Hafeth, Gokul Lal, Mohammed Al-Khafajiy, Thar Baker, Stefanos D. Kollias
DeSE5
2023 A deep neural architecture for harmonizing 3-D input data analysis and decision making in medical imaging
abstract
Harmonizing the analysis of data , especially of 3-D image volumes, consisting of different number of slices and annotated per volume, is a significant problem in training and using deep neural networks in various applications, including medical imaging . Moreover, unifying the decision making of the networks over different input datasets is crucial for the generation of rich data-driven knowledge and for trusted usage in the applications. This paper presents a new deep neural architecture, named RACNet, which includes routing and feature alignment steps and effectively handles different input lengths and single annotations of the 3-D image inputs, whilst providing highly accurate decisions. In addition, through latent variable extraction from the trained RACNet, a set of anchors are generated providing further insight on the network’s decision making. These can be used to enrich and unify data-driven knowledge extracted from different datasets. An extensive experimental study illustrates the above developments, focusing on COVID-19 diagnosis through analysis of 3-D chest CT scans from databases generated in different countries and medical centers.
Dimitris Kollias, Anastasios Arsenos, Stefanos D. Kollias
Neurocomputing3
2021 Effects of image quality and quantity on building a competitive COVID-19 diagnosis model
abstract
The outbreak of a health crisis, such as Covid-19, leads to decisions that must combine efficiency and speed. Often there is a trade-off between these two values, as the faster a decision is made, the less information is considered. This paper presents a deep learning model pipeline that balances these two values with the primary goal of classifying human lung X-rays into three categories: pneumonia, covid-19 and normal. Through this process, we tried to explore whether the quality of an image can enhance the learning process to a greater extent as opposed to having larger number of images. For this purpose, we follow two approaches by viewing quality and quantity as competing objectives to increasing the level of information obtained. The first is through increasing the number of X-ray images in the dataset, and the second is through improving the quality of the X-ray images. In the first approach, our goal is achieved using a Generative Adversarial Network (GAN) to generate plasmatic covid-19 class X-rays, while in the second approach, we improve the resolution of the X-ray images. To find the hyperparameters in both approaches that lead to better system performance, we exploit the Particle Swarm Optimization (PSO) algorithm. Rapid training and hyperparameter tuning better perform through this algorithm. Our experiments depict the performance that our models, based on the two approaches, achieved. Accuracy reaches 93% while sensitivity reaches 90% over Covid-19 cases. Finally, we conclude which characteristic, quality or quantity, is most useful in our case.
Philippos Skovelef Orfanoudakis, Paraskevi K. Tzouveli, Stefanos D. Kollias
IEEE BigData3
2021 Extracting geographical characteristics about COVID-19 evolution worldwide using machine learning
abstract
Since the beginning of 2020, the whole world has been plagued by the coronavirus pandemic. During the last sixteen months, almost every country in the world has faced several epidemic waves. An intriguing question that arises is whether neighboring countries, similar in regard to their socioeconomic status and the restrictions employed to counter the spread of the virus, showcase similarities in their respective number of cases and deaths. To that end, in this paper we form three clusters of similar countries (European and USA, African-Asian and Latin American) and we use their cumulative data as training data for machine learning models (RNN family, TCN and Attention) that predict the respective cases and deaths of 4 fixed neighboring countries, namely Cyprus, Greece, Italy and Spain. The results of the experiments conducted show that these 4 countries accent bigger similarity with the European cluster, as expected. Thus, evidence is provided bolstering the claim that similar neighboring countries exhibit alike behavior regarding the repercussions of the COVID-19.
Admitos-Rafael Passadakis, Anastasios Vlachos, Paraskevi K. Tzouveli, Stefanos D. Kollias
IEEE BigData4
2021 DeMoCap: Low-Cost Marker-Based Motion Capture
Anargyros Chatzitofis, Dimitrios Zarpalas, Petros Daras, Stefanos D. Kollias
Int. J. Comput. Vis.4
2021 An autoencoder wavelet based deep neural network with attention mechanism for multi-step prediction of plant growth
Bashar Alhnaity, Stefanos D. Kollias, Georgios Leontidis, Shouyong Jiang, Bert Schamp, Simon Pearson
Inf. Sci.2
2020 Capsule Routing via Variational Bayes
abstract
Capsule networks are a recently proposed type of neural network shown to outperform alternatives in challenging shape recognition tasks. In capsule networks, scalar neurons are replaced with capsule vectors or matrices, whose entries represent different properties of objects. The relationships between objects and their parts are learned via trainable viewpoint-invariant transformation matrices, and the presence of a given object is decided by the level of agreement among votes from its parts. This interaction occurs between capsule layers and is a process called routing-by-agreement. In this paper, we propose a new capsule routing algorithm derived from Variational Bayes for fitting a mixture of transforming gaussians, and show it is possible transform our capsule network into a Capsule-VAE. Our Bayesian approach addresses some of the inherent weaknesses of MLE based models such as the variance-collapse by modelling uncertainty over capsule pose parameters. We outperform the state-of-the-art on smallNORB using ≃50% fewer capsules than previously reported, achieve competitive performances on CIFAR-10, Fashion-MNIST, SVHN, and demonstrate significant improvement in MNIST to affNIST generalisation over previous works.1
Fabio De Sousa Ribeiro, Georgios Leontidis, Stefanos D. Kollias
AAAI3
2020 A Compact Sequence Encoding Scheme for Online Human Activity Recognition in HRI Applications
Georgios Tsatiris, Kostas Karpouzis, Stefanos D. Kollias
EANN3
2020 Introducing Routing Uncertainty in Capsule Networks
abstract
Rather than performing inefficient local iterative routing between adjacent capsule layers, we propose an alternative global view based on representing the inherent uncertainty in part-object assignment. In our formulation, the local routing iterations are replaced with variational inference of part-object connections in a probabilistic capsule network, leading to a significant speedup without sacrificing performance. In this way, global context is also considered when routing capsules by introducing global latent variables that have direct influence on the objective function, and are updated discriminatively in accordance with the minimum description length (MDL) principle. We focus on enhancing capsule network properties, and perform a thorough evaluation on pose-aware tasks, observing improvements in performance over previous approaches whilst being more computationally efficient.
Fabio De Sousa Ribeiro, Georgios Leontidis, Stefanos D. Kollias
NeurIPS3
2020 Unified deep learning approach for prediction of Parkinson's disease
abstract
The study presents a novel approach, based on deep learning, for diagnosis of Parkinson's disease through medical imaging. The approach includes analysis and use of the knowledge extracted by deep convolutional and recurrent neural networks when trained with medical images, such as magnetic resonance images and dopamine transporters scans. Internal representations of the trained DNNs constitute the extracted knowledge which is used in a transfer learning and domain adaptation manner, so as to create a unified framework for prediction of Parkinson's across different medical environments. A large experimental study is presented illustrating the ability of the proposed approach to effectively predict Parkinson's, using different medical image sets from real environments.
James Wingate, Ilianna Kollia, Luc M. Bidaut, Stefanos D. Kollias
IET Image Process.4
2020 Deep Bayesian Self-Training
abstract
Acknowledgements The authors would like to thank Mr. George Marandianos, Mrs. Mamatha Thota and Mr. Samuel Bond-Taylor for manually annotating datasets used in this study and of course the reviewers for their constructive feedback that helped to improve the manuscript. We would also like to thank Professor Luc Bidaut for enabling this collaboration. Funding The research presented in this paper was funded by Engineering and Physical Sciences Research Council (Reference Number EP/R005524/1) and Innovate UK (Reference Number 102908), in collaboration with the Olympus Automation Limited Company, for the project Automated Robotic Food Manufacturing System.
Fabio De Sousa Ribeiro, Francesco Calivá, Mark Swainson, Kjartan Gudmundsson, Georgios Leontidis, Stefanos D. Kollias
Neural Comput. Appl.6
2020 A Scalable Test Suite for Continuous Dynamic Multiobjective Optimization
abstract
Dynamic multiobjective optimization (DMO) has gained increasing attention in recent years. Test problems are of great importance in order to facilitate the development of advanced algorithms that can handle dynamic environments well. However, many of the existing dynamic multiobjective test problems have not been rigorously constructed and analyzed, which may induce some unexpected bias when they are used for algorithmic analysis. In this paper, some of these biases are identified after a review of widely used test problems. These include poor scalability of objectives and, more important, problematic overemphasis of static properties rather than dynamics making it difficult to draw accurate conclusion about the strengths and weaknesses of the algorithms studied. A diverse set of dynamics and features is then highlighted that a good test suite should have. We further develop a scalable continuous test suite, which includes a number of dynamics or features that have been rarely considered in literature but frequently occur in real life. It is demonstrated with empirical studies that the proposed test suite is more challenging to the DMO algorithms found in the literature. The test suite can also test algorithms in ways that existing test suites cannot.
Shouyong Jiang, Marcus Kaiser, Shengxiang Yang, Stefanos D. Kollias, Natalio Krasnogor
IEEE Trans. Cybern.4
2019 Predicting Parkinson's Disease using Latent Information extracted from Deep Neural Networks
abstract
This paper presents a new method for medical diagnosis of neurodegenerative diseases, such as Parkinson's, by extracting and using latent information from trained Deep convolutional, or convolutional-recurrent Neural Networks (DNNs). In particular, our approach adopts a combination of transfer learning, k-means clustering and k-Nearest Neighbour classification of deep neural network learned representations to provide enriched prediction of the disease based on MRI and/or DaT Scan data. A new loss function is introduced and used in the training of the DNNs, so as to perform adaptation of the generated learned representations between data from different medical environments. Results are presented using a recently published database of Parkinson's related information, which was generated and evaluated in a hospital environment.
Ilianna Kollia, Andreas Stafylopatis, Stefanos D. Kollias
IJCNN3
2019 Revisiting the medial axis for planar shape decomposition
Nikos Papanelopoulos, Yannis Avrithis, Stefanos D. Kollias
Comput. Vis. Image Underst.3
2018 An End-to-End Deep Neural Architecture for Optical Character Verification and Recognition in Retail Food Packaging
abstract
There exist various types of information in retail food packages, including food product name, ingredients list and use by date. The correct recognition and coding of use by dates is especially critical in ensuring proper distribution of the product to the market and eliminating potential health risks caused by erroneous mislabelling. The latter can have a major negative effect on the health of consumers and consequently raise legal issues for suppliers. In this work, an end-to-end architecture, composed of a dual deep neural network based system is proposed for automatic recognition of use by dates in food package photos. The system includes: a Global level convolutional neural network (CNN) for high-level food package image quality evaluation (blurry/clear/missing use by date statistics); a Local level fully convolutional network (FCN) for use by date ROI localisation. Post ROI extraction, the date characters are then segmented and recognised. The proposed framework is the first to employ deep neural networks for end-to-end automatic use by date recognition in retail packaging photos. It is capable of achieving very good levels of performance on all the aforementioned tasks, despite the varied textual/pictorial content complexity found in food packaging design.
Fabio De Sousa Ribeiro, Liyun Gong, Francesco Calivá, Mark Swainson, Kjartan Gudmundsson, Miao Yu 0001, Georgios Leontidis, Xujiong Ye, Stefanos D. Kollias
ICIP9
2018 Investigating the Best Performing Task Conditions of a Multi-Tasking Learning Model in Healthcare Using Convolutional Neural Networks: Evidence from a Parkinson'S Disease Database
abstract
This paper presents three conditions of Multi-Task Learning (MTL) model architectures based on Deep Neural Networks (DNNs) to predict the Parkinson's Disease (PD) from brain images. It also demonstrates the usefulness of incorporating additional patients' contextual epidemiological information (i.e. their age and sex). Our aim is to investigate which are the patients' clinical data, that when joined with the primary task could perform best and thus provide an improved computational PD prediction model. Our proposed model architectures are evaluated on a new medical dataset, which is presently under development. Our preliminary results suggest the robustness of our proposed systems to analyze and provide an accurate estimate of the status of the disease. Finally, we discuss the lessons learned from our experimental settings with respect to addressing several research questions such as the importance of selecting the auxiliary task(s) respectively, which is the best performing task combination to achieve such improved Parkinson's disease prediction, as well as whether all auxiliary tasks are equally effective.
Aggeliki Vlachostergiou, Athanasios Tagaris, Andreas Stafylopatis, Stefanos D. Kollias
ICIP4
2018 Multi-Task Learning for Predicting Parkinson's Disease Based on Medical Imaging Information
abstract
Parkinson's disease (PD) is a long-term degenerative disorder of the central nervous system, with symptoms generally appearing slowly over time. Predicting the PD disease is critical as motor and non-motor manifestations occur many years after the onset of neurodegeneration, hence its early management of disease is a significant challenge in the field of PD therapeutics. While part of previous studies with respect to the prediction of Parkinson's Disease has been based mainly on brain images, dependencies between additional patients' information have not been taken into account. This observation suggests that prediction of Parkinson's Disease along with additional patients' data with a unified framework should outperform Machine Learning (ML) algorithms that treat different sources of patients' information separately. Our presented framework relies on Multi-Task Learning (MTL) implemented with Deep Neural Networks (DNNs) with shared hidden layers. Our preliminary experimental results confirm the benefits of MTL over Single-Task Learning (STL), underlying the capability of our proposed system to achieve an increased Area Under the Curve (AUC) as high as 92% and helping at the same time to reduce human error.
Aggeliki Vlachostergiou, Athanasios Tagaris, Andreas Stafylopatis, Stefanos D. Kollias
ICIP4
2018 A Deep Learning Approach to Anomaly Detection in Nuclear Reactors
abstract
In this work, a novel deep learning approach to unfold nuclear power reactor signals is proposed. It includes a combination of convolutional neural networks (CNN), denoising autoencoders (DAE) and $k$-means clustering of representations. Monitoring nuclear reactors while running at nominal conditions is critical. Based on analysis of the core reactor neutron flux, it is possible to derive useful information for building fault/anomaly detection systems. By leveraging signal and image pre-processing techniques, the high and low energy spectra of the signals were appropriated into a compatible format for CNN training. Firstly, a CNN was employed to unfold the signal into either twelve or forty-eight perturbation location sources, followed by a $k$-means clustering and $k$-Nearest Neighbour coarse-to-fine procedure, which significantly increases the unfolding resolution. Secondly, a DAE was utilised to denoise and reconstruct power reactor signals at varying levels of noise and/or corruption. The reconstructed signals were evaluated w.r.t. their original counter parts, by way of normalised cross correlation and unfolding metrics. The results illustrate that the origin of perturbations can be localised with high accuracy, despite limited training data and obscured$/$noisy signals, across various levels of granularity.
Francesco Calivá, Fabio De Sousa Ribeiro, Antonios Mylonakis, Christophe Demazière, Paolo Vinai, Georgios Leontidis, Stefanos D. Kollias
IJCNN7
2018 Applications of human action analysis and recognition on wireless network infrastructures: State of the art and real world challenges
abstract
Human action recognition and analysis has given life to a wide variety of real-world applications, ranging from surveillance and human-computer interaction to patient monitoring and rehabilitation. Most action recognition systems, especially smart-home or assistive living applications, depend on network infrastructures for easy data fusion and integration of different sensing modalities. However, despite the fact that action recognition methods have extensively been evaluated for their accuracy and there is a consensus on the ways to provide quality of service in various network infrastructures, there is poor coverage of the inherent challenges of performing human action in real world network-based applications. In this work, we attempt to document these challenges based on representative, state of the art techniques and venture to report on the open issues that need to be resolved by new techniques aiming to provide viable real world applications.
Georgios Tsatiris, Kostas Karpouzis, Stefanos D. Kollias
INISTA3
2018 A novel rule based machine translation scheme from Greek to Greek Sign Language: Production of different types of large corpora and Language Models evaluation
Dimitris Kouremenos, Klimis S. Ntalianis, Stefanos D. Kollias
Comput. Speech Lang.3
2018 On the Beneficial Effect of Noise in Vertex Localization
Konstantinos Raftopoulos, Stefanos D. Kollias, Dionyssios D. Sourlas, Marin Ferecatu
Int. J. Comput. Vis.2
2017 HEAR?INFO: A Modern Mobile-Web Platform Addressed to Hard-of-Hearing Elderly Individuals
abstract
In the concept of the hearing loss awareness, a modern mobile-web platform is hereby presented, aiming to offer constant online access to individuals who are hard of hearing, while presenting them regularly updated information concerning their condition. This information is presented via a specific modified interface, taking into account the special needs of the specific community. After a thorough research in GUI, the software requirements substitute or supplement the lack of integrated sound systems, with visual modifications, caption text and even specially chosen colours. Different applications, including auditory tests and exercises are considered, aiming to the self-awareness and broadening of an individual's fund of knowledge.
Penelope Ioannidou, Panagiotis Katrakazas, Stefanos D. Kollias, Michail Sarafidis, Dimitris Koutsouris
CBMS3
2017 Computer vision based fall detection by a convolutional neural network
abstract
In this work, we propose a novel computer vision based fall detection system, which could be applied for the health-care of the elderly people community. For a recorded video stream, background subtraction is firstly applied to extract the human body silhouette. Extracted silhouettes corresponding to daily activities are applied to construct a convolutional neural network, which is applied for classification of different classes of human postures (e.g., bend, stand, lie and sit) and detection of a fall event (i.e., lying posture is detected in the floor region). As far as we know, this work is the first attempt for the application of the convolutional neural network for the fall detection application. From a dataset of daily activities recorded from multiple people, we show that the proposed method both achieves higher postures classification results than the state-of-the-art classifiers and can successfully detect the fall event with a low false alarm rate.
Miao Yu 0001, Liyun Gong, Stefanos D. Kollias
ICMI3
2016 α-shapes for local feature detection
Christos Varytimidis, Konstantinos Rapantzikos, Yannis Avrithis, Stefanos D. Kollias
Pattern Recognit.4
2014 Improving Local Features by Dithering-Based Image Sampling
Christos Varytimidis, Konstantinos Rapantzikos, Yannis Avrithis, Stefanos D. Kollias
ACCV (2)4
2014 Towards large-scale geometry indexing by feature selection
Giorgos Tolias, Yannis Kalantidis, Yannis Avrithis, Stefanos D. Kollias
Comput. Vis. Image Underst.4
2014 Visual Focus of Attention in Non-calibrated Environments using Gaze Estimation
Stylianos Asteriadis, Kostas Karpouzis, Stefanos D. Kollias
Int. J. Comput. Vis.3
2013 Intelligent and Adaptive Pervasive Future Internet: Smart Cities for the Citizens
George Caridakis, Georgios Siolas, Phivos Mylonas, Stefanos D. Kollias, Andreas Stafylopatis
EANN (2)4
2013 Social and Smart: Towards an Instance of Subconscious Social Intelligence
Manuel Graña, Bruno Apolloni, Maurizio Fiasché, Gian Luca Galliani, C. Zizzo, George Caridakis, Georgios Siolas, Stefanos D. Kollias, F. Barriento, S. San Jose
EANN (2)8
2013 A Particle Swarm Optimization (PSO) Model for Scheduling Nonlinear Multimedia Services in Multicommodity Fat-Tree Cloud Networks
Ioannis M. Stephanakis, Ioannis P. Chochliouros, George Caridakis, Stefanos D. Kollias
EANN (2)4
2013 Social things - The SandS instantiation
abstract
At a time when socialism as an economic option is variously questioned, very few people are against social instances of our life such as entertainment, customer assistance, and so on. This happens with the management of many things accompanying our life as well. We can find both the reason and the evidence for the viability of this trend in one very basic fact: things are social because they work better. However, also in this sphere social politics are highly questionable. Here we introduce the perspective adopted in the European project SandS within a framework of Internet of Things. In this case things are agents interacting on the network within a service centric approach where a sound hierarchy dispatches instructions. It is a complete ecosystem where the social network develops a collective intelligence subtending new concrete functionalities that are centered on the user willing and fostered by his/her feedbacks. The central role of the user reflects on all aspects of the ecosystem, from the family of things which are socially governed: the household appliances (the white goods) that affect our everyday life, up to the employed hardware and software: strictly open source.
Bruno Apolloni, Maurizio Fiasché, Gian Luca Galliani, C. Zizzo, George Caridakis, Georgios Siolas, Stefanos D. Kollias, Manuel Graña, F. Barriento, S. San Jose
WOWMOM7
2013 Exploring trace transform for robust human action recognition
Georgios Goudelis, Kostas Karpouzis, Stefanos D. Kollias
Pattern Recognit.3
2013 Mining User Queries with Markov Chains: Application to Online Image Retrieval
abstract
We propose a novel method for automatic annotation, indexing and annotation-based retrieval of images. The new method, that we call Markovian Semantic Indexing (MSI), is presented in the context of an online image retrieval system. Assuming such a system, the users' queries are used to construct an Aggregate Markov Chain (AMC) through which the relevance between the keywords seen by the system is defined. The users' queries are also used to automatically annotate the images. A stochastic distance between images, based on their annotation and the keyword relevance captured in the AMC, is then introduced. Geometric interpretations of the proposed distance are provided and its relation to a clustering in the keyword space is investigated. By means of a new measure of Markovian state similarity, the mean first cross passage time (CPT), optimality properties of the proposed distance are proved. Images are modeled as points in a vector space and their similarity is measured with MSI. The new method is shown to possess certain theoretical advantages and also to achieve better Precision versus Recall results when compared to Latent Semantic Indexing (LSI) and probabilistic Latent Semantic Indexing (pLSI) methods in Annotation-Based Image Retrieval (ABIR) tasks.
Konstantinos Raftopoulos, Klimis S. Ntalianis, Dionyssios D. Sourlas, Stefanos D. Kollias
IEEE Trans. Knowl. Data Eng.4
2012 Tractable reasoning with vague knowledge using fuzzy EL++
Theofilos P. Mailis, Giorgos Stoilos, Nikos Simou, Giorgos B. Stamou, Stefanos D. Kollias
J. Intell. Inf. Syst.5
2012 Non parametric, self organizing, scalable modeling of spatiotemporal inputs: The sign language paradigm
George Caridakis, Kostas Karpouzis, Athanasios I. Drosopoulos, Stefanos D. Kollias
Neural Networks4
2011 Visual Pathways for Shape Abstraction
Konstantinos Raftopoulos, Stefanos D. Kollias
ICANN (1)2
2011 Affine morphological shape stable boundary regions (SSBR) for image representation
abstract
This paper presents a new structure-based interest region detector called Shape-Stable Region Boundaries (SSRB) which we use for object class recognition. The SSRB interest operator detects stable boundary regions within the multi-scale morphological image representation [1]. To detect robust boundary regions, we perform multi-scale analysis via anisotropic diffusion operators to preserve boundaries and guarantee invariance to affine transformations. We extract the transition boundaries of the diffusivity velocity map and track their evolution at each level of the scale-space. The stability of the boundary shape is subsequently estimated through a minimization process over different scales. Unlike most state of the art detectors which use the Gaussian scale space for multi-scale image representation, our approach is intrinsically affine invariant [1]. Experiments on different benchmark datasets show that SSRB is comparable or superior to state-of the art detectors for both feature matching and object recognition.
Petros Kapsalas, Stefanos D. Kollias
ICIP2
2011 The Global-Local transformation for noise resistant shape representation
Konstantinos Raftopoulos, Stefanos D. Kollias
Comput. Vis. Image Underst.2
2011 VIRaL: Visual Image Retrieval and Localization
Yannis Kalantidis, Giorgos Tolias, Yannis Avrithis, Marios Phinikettos, Evaggelos Spyrou, Phivos Mylonas, Stefanos D. Kollias
Multim. Tools Appl.7
2010 Intelligent content retrieval using a visual vocabulary and geometric constraints
abstract
During the last decades multimedia processing has emerged as an important technology to retrieve content based on similar data. Moreover, recent developments in the fields of high definition (HD) multimedia content and personal content collections (personal camcorders and digital still image cameras) tend to generate a huge volume of multimedia data everyday. Thus, the need for a meaningful, quick organization and access to generated content is now more than necessary; however, it still remains a rather difficult problem to be tackled both by humans and computers. In this paper we propose an intelligent extension of traditional image analysis methodologies towards more efficient digital content retrieval. The main idea is to extend local feature extraction methodologies by introducing additional geometrical constraints in the process. The proposed approach is tested and evaluated on a number of publicly available image datasets and results are very promising.
Evaggelos Spyrou, Yannis Kalantidis, Giorgos Tolias, Phivos Mylonas, Stefanos D. Kollias
FUZZ-IEEE5
2010 Shape-stable region boundary extraction via affine morphological scale space (AMSS)
abstract
In this paper we present a new approach towards the extraction of affine image regions based on detecting shape-stable boundaries from a multi-scale image representation. We construct an affine morphological scale space (AMSS) representation [1], which performs anisotropic diffusion while preserving boundaries and being invariant to affine transformations. We extract the transition boundaries of the diffusivity velocity map and track their evolution at each level of the scale-space. We then determine the stability of the boundary shape through a minimization process over different scales. Unlike most state of the art detectors which use the Gaussian scale space for multi-scale image representation, our approach is intrinsically affine invariant. We evaluate our detector by measuring repeatability of regions in transformed images of the same scene and comparing it to the state-of-the-art region detectors [2].
Petros Kapsalas, Stefanos D. Kollias
ACM Multimedia2
2010 SOMM: Self organizing Markov map for gesture recognition
George Caridakis, Kostas Karpouzis, Athanasios I. Drosopoulos, Stefanos D. Kollias
Pattern Recognit. Lett.4
2009 Dense saliency-based spatiotemporal feature points for action recognition
abstract
Several spatiotemporal feature point detectors have been used in video analysis for action recognition. Feature points are detected using a number of measures, namely saliency, cornerness, periodicity, motion activity etc. Each of these measures is usually intensity-based and provides a different trade-off between density and informativeness. In this paper, we use saliency for feature point detection in videos and incorporate color and motion apart from intensity. Our method uses a multi-scale volumetric representation of the video and involves spatiotemporal operations at the voxel level. Saliency is computed by a global minimization process constrained by pure volumetric constraints, each of them being related to an informative visual aspect, namely spatial proximity, scale and feature similarity (intensity, color, motion). Points are selected as the extrema of the saliency response and prove to balance well between density and informativeness. We provide an intuitive view of the detected points and visual comparisons against state-of-the-art space-time detectors. Our detector outperforms them on the KTH dataset using nearest-neighbor classifiers and ranks among the top using different classification frameworks. Statistics and comparisons are also performed on the more difficult Hollywood human actions (HOHA) dataset increasing the performance compared to current published results.
Konstantinos Rapantzikos, Yannis Avrithis, Stefanos D. Kollias
CVPR3
2009 Towards modeling embodied conversational agent character profiles using appraisal theory predictions in expression synthesis
Lori Malatesta, Amaryllis Raouzaiou, Kostas Karpouzis, Stefanos D. Kollias
Appl. Intell.4
2009 Estimation of behavioral user state based on eye gaze and head pose - application in an e-learning environment
Stylianos Asteriadis, Paraskevi K. Tzouveli, Kostas Karpouzis, Stefanos D. Kollias
Multim. Tools Appl.4
2009 MPEG-4 facial expression synthesis
Lori Malatesta, Amaryllis Raouzaiou, Kostas Karpouzis, Stefanos D. Kollias
Pers. Ubiquitous Comput.4
2009 Spatiotemporal saliency for video classification
Konstantinos Rapantzikos, Nicolas Tsapatsoulis, Yannis Avrithis, Stefanos D. Kollias
Signal Process. Image Commun.4
2009 Using Visual Context and Region Semantics for High-Level Concept Detection
abstract
In this paper we investigate detection of high-level concepts in multimedia content through an integrated approach of visual thesaurus analysis and visual context. In the former, detection is based on model vectors that represent image composition in terms of region types, obtained through clustering over a large data set. The latter deals with two aspects, namely high-level concepts and region types of the thesaurus, employing a model of a priori specified semantic relations among concepts and automatically extracted topological relations among region types; thus it combines both conceptual and topological context. A set of algorithms is presented, which modify either the confidence values of detected concepts, or the model vectors based on which detection is performed. Visual context exploitation is evaluated on TRECVID and Corel data sets and compared to a number of related visual thesaurus approaches.
Phivos Mylonas, Evaggelos Spyrou, Yannis Avrithis, Stefanos D. Kollias
IEEE Trans. Multim.4
2008 A non-intrusive method for user focus of attention estimation in front of a computer monitor
abstract
In this work, we present a system that estimates a user's focus of attention in front of a computer screen, using a Web camera, based on detection and tracking of the user's head position and eye movements. Utilizing machine learning concepts, the system gives real time feedback on the user's attention, by combining information coming from eye gaze, head pose, and distance from the screen. The system is completely un-intrusive and no special hardware (such as infrared cameras or wearable devices) is needed. Furthermore, it adjusts to every user, not necessitating initial calibration, and can work under real and unconstrained conditions in terms of lighting.
Stylianos Asteriadis, Paraskevi K. Tzouveli, Kostas Karpouzis, Stefanos D. Kollias
FG4
2008 Reasoning with qualified cardinality restrictions in fuzzy Description Logics
abstract
Description logics (DLs) are modern knowledge representation formalisms which are used today in many applications for reasoning with structured knowledge. Moreover, they are used in the semantic web (an extension of the current web) through the ontology language OWL. On the other hand fuzzy description logics (fuzzy-DLs) have been proposed as expressive logical formalisms capable of capturing and reasoning with vague and imprecise knowledge in the semantic web. In the current paper we investigate on the problem of reasoning with qualified cardinality restrictions (QCRs) in fuzzy DLs, extending previous results on simple number restrictions, thus we present a tableaux algorithm for the the fuzzy-DL fKD-ALCIQ.
Giorgos Stoilos, Giorgos B. Stamou, Stefanos D. Kollias
FUZZ-IEEE3
2008 Adaptive Reading Assistance for the Inclusion of Students with Dyslexia: The AGENT-DYSL Approach
abstract
Dyslexia is a major barrier to success in education and later on the job as reading skills are fundamental for personal competence development. Children with dyslexia have special learning needs (e.g., more teacher support), which currently only specialized institutions can provide. However, this takes children out of their peer group and causes social problems. On the other side, there is general-purpose reading support software, which are not geared towards children with dyslexia as they lack personalization. AGENT-DYSL brings together speech and image recognition as well as semantic technologies to build a truly adaptive reading support system for children with dyslexia.
Paraskevi K. Tzouveli, Andreas Schmidt 0007, Michael Schneider 0001, Antonios Symvonis, Stefanos D. Kollias
ICALT5
2008 A Neuro-fuzzy Approach to User Attention Recognition
Stylianos Asteriadis, Kostas Karpouzis, Stefanos D. Kollias
ICANN (1)3
2008 Adaptation of Connectionist Weighted Fuzzy Logic Programs with Kripke-Kleene Semantics
Alexandros Chortaras, Giorgos B. Stamou, Andreas Stafylopatis, Stefanos D. Kollias
ICANN (1)4
2008 Semantic Adaptation of Neural Network Classifiers in Image Segmentation
Nikos Simou, Thanos Athanasiadis, Stefanos D. Kollias, Giorgos B. Stamou, Andreas Stafylopatis
ICANN (1)3
2008 Hand trajectory based gesture recognition using self-organizing feature maps and markov models
abstract
This work presents the design and experimental verification of an original system architecture aiming at recognizing gestures based solely on the hand trajectory. Self organizing feature maps are used to model spatial information while Markov models encode the temporal aspect of hand position within a trajectory. A validated classification mechanism is produced through a set of models and a committee machine setup ensures robustness as indicated by the experimental results performed.
George Caridakis, Kostas Karpouzis, Christos Pateritsas, Athanasios I. Drosopoulos, Andreas Stafylopatis, Stefanos D. Kollias
ICME6
2008 User and context adaptive neural networks for emotion recognition
George Caridakis, Kostas Karpouzis, Stefanos D. Kollias
Neurocomputing3
2008 International Conference on Artificial Neural Networks (ICANN 2006)
Stefanos D. Kollias, Andreas Stafylopatis, Wlodzislaw Duch
Neurocomputing1
2008 Semantic representation of multimedia content: Knowledge representation and semantic indexing
Phivos Mylonas, Thanos Athanasiadis, Manolis Wallace, Yannis Avrithis, Stefanos D. Kollias
Multim. Tools Appl.5
2007 f-DLPs: Extending Description Logic Programs with Fuzzy Sets and Fuzzy Logic
abstract
The Semantic Web can be viewed as largely about "Knowledge meets the Web". Thus its vision includes ontologies and rules. A key requirement for the architecture of the Semantic Web is to be able to layer "rules on top of ontologies" and "ontologies on top of rules". This has as a counterpart the definition of a mapping between Description Logics and Logic Programming, which is known as Description Logic Programs. In this paper we extend the Description Logic Programs with fuzzy sets and fuzzy logic in order to be able to represent the imprecision and vagueness of real-life applications. We provide the common semantics of the mapping, and the conditions that must be met for this semantic equivalence, based on the model-theoretic semantics.
Tassos Venetis, Giorgos Stoilos, Giorgos B. Stamou, Stefanos D. Kollias
FUZZ-IEEE4
2007 Probabilistic Video-Based Gesture Recognition Using Self-organizing Feature Maps
George Caridakis, Christos Pateritsas, Athanasios I. Drosopoulos, Andreas Stafylopatis, Stefanos D. Kollias
ICANN (2)5
2007 salienShrink: Saliency-Based Wavelet Shrinkage
abstract
This paper describessalienShrink, a method to denoise images based on computing a map of salient coefficients in the wavelet domain and use it to improve common denoising algorithms. By salient, we refer to those coefficients that correspond mostly to pure signal and should therefore be preserved throughout the denoising procedure. We use a computationally efficient model to detect salient regions in the bands of the multiresolution wavelet transform. These regions are used to obtain a more accurate estimate of the noise level, improving the performance of existing well known shrinkage methods. Extensive experimental results on theBiShrinkmethod show that the proposed method effectively enhances PSNR and improves the visual quality of the denoised images.
Konstantinos Rapantzikos, Yannis Avrithis, Stefanos D. Kollias
ICIP (1)3
2007 Semantic Image Segmentation and Object Labeling
abstract
In this paper, we present a framework for simultaneous image segmentation and object labeling leading to automatic image annotation. Focusing on semantic analysis of images, it contributes to knowledge-assisted multimedia analysis and bridging the gap between semantics and low level visual features. The proposed framework operates at semantic level using possible semantic labels, formally represented as fuzzy sets, to make decisions on handling image regions instead of visual features used traditionally. In order to stress its independence of a specific image segmentation approach we have modified two well known region growing algorithms, i.e., watershed and recursive shortest spanning tree, and compared them to their traditional counterparts. Additionally, a visual context representation and analysis approach is presented, blending global knowledge in interpreting each object locally. Contextual information is based on a novel semantic processing methodology, employing fuzzy algebra and ontological taxonomic knowledge representation. In this process, utilization of contextual knowledge re-adjusts labeling results of semantic region growing, by means of fine-tuning membership degrees of detected concepts. The performance of the overall methodology is evaluated on a real-life still image dataset from two popular domains
Thanos Athanasiadis, Phivos Mylonas, Yannis Avrithis, Stefanos D. Kollias
IEEE Trans. Circuits Syst. Video Technol.4
2006 Adaptive On-Line Neural Network Retraining for Real Life Multimodal Emotion Recognition
Spiros Ioannou, Loïc Kessous, George Caridakis, Kostas Karpouzis, Vered Aharonson, Stefanos D. Kollias
ICANN (1)6
2006 Confronting the Synchronization Problem of Semantic Region Under Geometric Attacks
abstract
In this paper, an affine invariant watermarking scheme, robust to geometric attacks, is proposed and applied to face regions. Initially, face regions are unsupervisedly extracted from an initial image and a normalization procedure, invariant to geometric attacks is applied on each of these regions using a set of specific moment criteria. A spread spectrum-based DS-CDMA watermarking scheme is then used in order to provide a multi bits watermark. The multi bits watermark embedding and detection procedures are then applicable to each normalized face region. Finally, performance of the proposed face regions watermarking scheme is tested under various distortions, providing efficient and robust watermark retrieval
Paraskevi K. Tzouveli, Klimis S. Ntalianis, Stefanos D. Kollias
ICME3
2006 A Connectionist Model for Weighted Fuzzy Programs
abstract
The usefulness of the results of logic programming in real-life applications is sometimes limited due to the inability of this theory to model the uncertain and dynamic character of real environments. Fuzzy logic programming has been lately considered as an important framework for handling uncertainty in logic programming systems. Still, there is a need for modelling adaptation of logic programs and the progress in this area is rather slow. In the present paper, we first extend fuzzy logic programs in a direction that brings them closer to the connectionist approach: we introduce weighted fuzzy programs, which allow the association of significance weights with the atoms that make up the body of a logic rule. The weights add expressiveness to the programs and allow the determination of the degree with which an antecedent affects the value of the rule consequent. Then, we propose a neural network implementation of weighted fuzzy programs that is capable of computing the minimal Herbrand model of a weighted fuzzy program.
Alexandros Chortaras, Giorgos B. Stamou, Andreas Stafylopatis, Stefanos D. Kollias
IJCNN4
2006 Intelligent Facial Analysis and Expression Recognition
abstract
Since facial expressions are a key modality in human communication, the automated analysis of facial images and video for the estimation of the displayed expression is central in the design of intuitive and human friendly computer interaction systems. In this paper we present an intelligent feature extraction system which combines analysis from multiple channels based on their confidence, to result in better, error resilient facial feature boundary detection. Neural networks are a key component of the system. Issues such as uncertainty and lack of confidence in the process of feature extraction are considered during the expression analysis and recognition. Various results are presented which illustrate the performance of the method.
Spiros Ioannou, Manolis Wallace, Stefanos D. Kollias
IJCNN3
2006 Computationally efficient sup-t transitive closure for sparse fuzzy binary relations
Manolis Wallace, Yannis Avrithis, Stefanos D. Kollias
Fuzzy Sets Syst.3
2006 Emotional face expression profiles supported by virtual human ontology
abstract
Abstract Expressive facial animation synthesis of human like characters has had many approaches with good results. MPEG‐4 standard has functioned as the basis of many of those approaches. In this paper we would like to lay out the knowledge of some of those approaches inside an ontology in order to support the modeling of emotional facial animation in virtual humans (VH). Inside this ontology we will present MPEG‐4 facial animation concepts and its relationship with emotion through expression profiles that utilize psychological models of emotions. The ontology allows storing, indexing and retrieving prerecorded synthetic facial animations that can express a given emotion. Also this ontology can be used a refined knowledge base in regards to the emotional facial animation creation. This ontology is made using Web Ontology Language and the results are presented as answered queries. Copyright © 2006 John Wiley & Sons, Ltd.
Alejandra García-Rojas, Frédéric Vexo, Daniel Thalmann, Amaryllis Raouzaiou, Kostas Karpouzis, Stefanos D. Kollias, Laurent Moccozet, Nadia Magnenat-Thalmann
Comput. Animat. Virtual Worlds6
2006 Integrating Multimedia Archives: The Architecture and the Content Layer
abstract
In the last few years, numerous multimedia archives have made extensive use of digitized storage and annotation technologies. Still, the development of single points of access, providing common and uniform access to their data, despite the efforts and accomplishments of standardization organizations, has remained an open issue as it involves the integration of various large-scale heterogeneous and heterolingual systems. This paper describes a mediator system that achieves architectural integration through an extended three-tier architecture and content integration through semantic modeling. The described system has successfully integrated five multimedia archives, quite different in nature and content from each other, while also providing easy and scalable inclusion of more archives in the future.
Manolis Wallace, Thanos Athanasiadis, Yannis Avrithis, Anastasios Delopoulos, Stefanos D. Kollias
IEEE Trans. Syst. Man Cybern. Part A5
2005 Confidence-Based Fusion of Multiple Feature Cues for Facial Expression Recognition
abstract
Since facial expressions are a key modality in human communication, the automated analysis of facial images for the estimation of the displayed expression is essential in the design of intuitive and accessible human computer interaction systems. In most existing rule-based expression recognition approaches, analysis is semiautomatic or requires high quality video. In this paper we propose a feature extraction system which combines analysis from multiple channels based on their confidence, to result in better facial feature boundary detection. The facial features are then used for expression estimation. The proposed approach has been implemented as an extension to an existing expression analysis system in the framework of the IST ERMIS project
Spiros Ioannou, Manolis Wallace, Kostas Karpouzis, Amaryllis Raouzaiou, Stefanos D. Kollias
FUZZ-IEEE5
2005 Handling Uncertainty in Video Analysis with Spatiotemporal Visual Attention
abstract
In natural vision, we center our fixation on the most informative points in a scene in order to reduce our overall uncertainty about the scene and help interpret it. Even if we are looking for a specific stimulus around us, we face a great amount of uncertainty since that stimulus could be in any spatial location. Visual attention (VA) schemes have been proposed by researchers to account for the ability of the human eye to quickly fixate on informative regions. Recently, VA in images, and especially saliency-based VA, became an active research topic of the computer vision community. The proposed work provides an extension towards VA in video sequences by integrating spatiotemporal information. The potential applications include video classification, scene understanding, surveillance and segmentation
Konstantinos Rapantzikos, Yannis Avrithis, Stefanos D. Kollias
FUZZ-IEEE3
2005 Possibilistic Evaluation of Extended Fuzzy Rules in the Presence of Uncertainty
abstract
Characterization fuzzy in term "fuzzy rule base" is currently referred to the ability to define rule antecedents using fuzzy numbers. On the other hand, when it comes to the knowledge described by the rules and to the information contained in rule antecedents, absolute accuracy is assumed. With the emergence of a vast variety of applications of rule based systems, where antecedents are not provided by sensors but rather by complicated processing modules, more efficient rules and rule evaluation structures are needed, that are able to describe knowledge in more intuitive manner and cope with uncertainty in the assumed input. In this paper we propose extended fuzzy rules that allow for optional antecedents and provide a methodology for the possibilistic evaluation of both conventional and extended fuzzy rules in the presence of uncertainty. The work has been successfully applied in a real life problem, for which conventional fuzzy rules and fuzzy rule evaluation were inadequate
Manolis Wallace, Stefanos D. Kollias
FUZZ-IEEE2
2005 Combination of multiple extraction algorithms in the detection of facial features
abstract
Automated analysis of facial images for the estimation of the displayed expression is essential in the design of intuitive and accessible human computer interaction systems. In existing rule-based expression recognition approaches, different feature extraction techniques have been tested that allow for the automatic detection of feature points, providing the required input for a rule based expression analysis; each one of these techniques outperforms others under specific constraints. In this paper we propose a feature extraction system which combines analysis from multiple channels based on their confidence, to result in better, error resilient facial feature boundary detection. The proposed approach has been implemented as an extension to an existing expression analysis system in the framework of the 1ST ERMIS project.
Spiros Ioannou, Manolis Wallace, Kostas Karpouzis, Amaryllis Raouzaiou, Stefanos D. Kollias
ICIP (2)5
2005 Chaotic video objects encryption based on mixed feedback, multiresolution decomposition and time-variant S-boxes
abstract
The increasing popularity of multimedia applications creates an increasing need for secure storage and transmission techniques. Especially wireless communications, which can be easily intercepted, should be protected by eavesdroppers. Towards this direction, several encryption schemes have been proposed in literature, however most of them do not consider regions of interest (video objects). These regions may need better protection, may be the only regions that need protection, or should be accessed separately according to different privileges. For these reasons in this paper a video objects chaotic encryption system is proposed. Initially stereoscopic pairs are analyzed and video objects are automatically extracted using a color segments fusion approach. Next for each video object multiresolution decomposition is performed and the pixels of the lowest resolution level are encrypted using a chaotic cipher module. Finally the encrypted regions are propagated to the higher resolution levels and the encryption process is repeated until the highest level is reached. The system presents robustness against known cryptanalytic attacks, enables layered access of multimedia content and the overall security can be enhanced due to region topology.
Klimis S. Ntalianis, Stefanos D. Kollias
ICIP (2)2
2005 An intelligent system for facial emotion recognition
abstract
An intelligent emotion recognition system, interweaving psychological findings about emotion representation with analysis and evaluation of facial expressions has been generated and its performance has been investigated with experimental real data. Additionally, a fuzzy rule based system has been created for classifying facial expressions to the six archetypal emotion categories. The continuous 2-D emotion space was then examined and a pool of known and novel classification and clustering techniques have been applied to our data obtaining high rates in classification and clustering into quadrants of the emotion representation space.
Roddy Cowie, Ellen Douglas-Cowie, John G. Taylor, Spiros Ioannou, Manolis Wallace, Stefanos D. Kollias
ICME6
2005 Multimodal Emotion Recognition and Expressivity Analysis
abstract
The paper presents the framework of a special session that aims at investigating the best possible techniques for multimodal emotion recognition and expressivity analysis in human computer interaction, based on a common psychological background. The session mainly deals with audio and visual emotion analysis, with physiological signal analysis serving as supplementary to these modalities. Specific topics that are examined include extraction of emotional features and signs from each modality in separate, integration of the outputs of single-mode emotion analysis systems and recognition of the user's emotional state, taking into account emotion models and existing knowledge or demands from both the analysis and synthesis perspective. Various labeling schemes, supply of accordingly labeled test databases, as well as synthesis of expressive avatars and affective interactions, are issues brought up and examined in the proposed framework
Stefanos D. Kollias, Kostas Karpouzis
ICME1
2005 An Optimized Key-Frames Extraction Scheme Based on SVD and Correlation Minimization
abstract
In this paper an optimized and efficient technique for keyframes extraction of video sequences is proposed, which leads to selection of a meaningful set of video frames for each given shot. Initially for each frame, the singular value decomposition method is applied and a diagonal matrix is produced, containing the singular values of the frame. Afterwards, a feature vector is created for each frame, by gathering the respective singular values. Next, all feature vectors of the shot are collected to form the feature vectors basin of this shot. Finally, a genetic algorithm approach is proposed and applied to the vectors basin, for locating frames of minimally correlated feature vectors, which are selected as keyframes. Experimental results indicate the promising performance of the proposed scheme on real life video shots
Klimis S. Ntalianis, Stefanos D. Kollias
ICME2
2005 A String Metric for Ontology Alignment
Giorgos Stoilos, Giorgos B. Stamou, Stefanos D. Kollias
ISWC3
2005 Emotion recognition through facial expression analysis based on a neurofuzzy network
Spiros Ioannou, Amaryllis Raouzaiou, Vassilis Tzouvaras, Theofilos P. Mailis, Kostas Karpouzis, Stefanos D. Kollias
Neural Networks6
2005 Intelligent initialization of resource allocating RBF networks
Manolis Wallace, Nicolas Tsapatsoulis, Stefanos D. Kollias
Neural Networks3
2004 Computationally efficient incremental transitive closure of sparse fuzzy binary relations
abstract
Existing literature in the field of transitive relations focuses mainly on dense, Boolean, undirected relations. With the emergence of a new area of intelligent retrieval, where sparse transitive fuzzy ordering relations are utilized, existing theory and methodologies need to be extended, as to cover the new needs. This work discusses the incremental update of such fuzzy binary relations, while focusing on both storage and computational complexity issues. Moreover, it proposes a novel transitive closure algorithm that has a remarkably low computational complexity (below O(n/sup 2/)) for the average sparse relation; such are the relations encountered in intelligent retrieval.
Manolis Wallace, Stefanos D. Kollias
FUZZ-IEEE2
2004 Facial expression classification based on MPEG-4 FAPs: the use of evidence and prior knowledge for uncertainty removal
abstract
As low resolution shots, rotations of the head with respect to the camera, face deformation due to speech and so on inflict a great deal of uncertainty in FAP measurements, uncertainty is also inherent in the process of expression analysis. We tackle such uncertainty via the observation that user emotions do not typically alter rapidly very often. Thus, possibilistic evidence may be gathered from each frame about the user expression; evidence from the current and recent frames can be combined using evidence theory.
Manolis Wallace, Amaryllis Raouzaiou, Nicolas Tsapatsoulis, Stefanos D. Kollias
FUZZ-IEEE4
2004 Towards a Personalized e-Learning Scheme for Teachers
abstract
One of the most important topics in modern e-learning schemes for teachers is the treatment of information and users at a personalized level. In this framework, the automated extraction of user profiles from an e-learning system is an interesting and important problem. In this paper we present the designing and materializing of such a profile-based system for teachers, which extends on previous work on profile extraction and re-evaluation in the direction of automated extraction of user preferences. Our approach relies on suitably adapted fundamental e-learning models, such as the IEEE e-learning model, as well as on a novel mechanism that creates, updates and uses users' profiles.
Phivos Mylonas, Paraskevi K. Tzouveli, Stefanos D. Kollias
ICALT3
2004 Adaptation of facial feature extraction and rule generation in emotion-analysis systems
abstract
The paper addresses the problem of emotion recognition in faces through an intelligent neuro-fuzzy system, where the extraction of facial features follows the MPEG-4 standard and is adapted to particular environmental conditions and specific persons. These features are associated to symbolic fuzzy predicates providing the classification of facial images according to the underlying emotional states. For this classification we use rules extracted from psychological studies and expression databases including extreme expressions such as those illustrated in Ekman's database. The rules are then refined in realistic conditions, taking into account the extracted features. The experimental results, based both in extreme and naturalistic databases developed in the frameworks of IST ERMIS and NoE HUMAINE, illustrate the capability of the developed system to analyse and recognise facial expressions in human computer interaction applications.
Sivann Ioannou, Amaryllis Raouzaiou, Kostas Karpouzis, Stefanos D. Kollias
IJCNN4
2004 Security of human video objects by incorporating a chaos-based feedback cryptographic scheme
abstract
Security of multimedia files attracts more and more attention and many encryption methods have been proposed in literature. However most cryptographic systems deal with multimedia files as binary large objects, without taking into consideration regions of semantic information. These regions may need better protection or can be the only regions that need protection, depending on the specific application. Towards this direction, in this paper we propose a human video object encryption system based on the chaotic logistic map. Initially face regions are efficiently detected and afterwards body regions are extracted, using geometric information of the location of face regions. Next the pixels of extracted human video objects are encrypted using an iterative cipher module, which is based on a feedback mechanism responsible for mixing the current encryption parameters with encrypted information of the previous step. The system presents robustness against known cryptanalytic attacks, and can save us a great amount of computational resources and time devoted for encrypting the whole contents of a multimedia file.
Paraskevi K. Tzouveli, Klimis S. Ntalianis, Stefanos D. Kollias
ACM Multimedia3
2004 Emotion synthesis in the MPEG-4 framework
abstract
Man-machine interaction (MMI) systems that utilize multimodal information about users' current emotional state are presently at the forefront of interest of the computer vision and artificial intelligence communities. A lifelike human face can enhance interactive applications by providing straightforward feedback to and from the users and stimulating emotional responses from them. In this paper, we present an abstract means of description of facial expressions, by utilizing concepts included in the MPEG-4 standard to synthesize expressions using a reduced representation, suitable for networked and lightweight applications.
Amaryllis Raouzaiou, Kostas Karpouzis, Stefanos D. Kollias
MMSP3
2004 Web Access to Large Audiovisual Assets Based on User Preferences
Kostas Karpouzis, George Moschovitis, Klimis S. Ntalianis, Spiros Ioannou, Stefanos D. Kollias
Multim. Tools Appl.5
2004 A snake model for object tracking in natural sequences
Gavriil Tsechpenakis, Konstantinos Rapantzikos, Nicolas Tsapatsoulis, Stefanos D. Kollias
Signal Process. Image Commun.4
2004 Semantic association of multimedia document descriptions through fuzzy relational algebra and fuzzy reasoning
abstract
According to the emerging MPEG-7 standard, the semantic description of multimedia documents is expressed in terms of semantic entities such as objects, events, concepts, and relations among them. The semantic entities can be used as index terms, in order to support the semantic search process. In this paper, we propose a method that a) applies fuzzy relational operations (closure, composition) and fuzzy rules to expand a semantic encyclopedia and b) uses the encyclopedia to associate the semantic entities with the aid of a fuzzy thesaurus. This method is shown to reduce the need for human intervention in creating semantic descriptions of multimedia documents, as well as correct for incompleteness and inconsistency.
Giorgos Akrivas, Giorgos B. Stamou, Stefanos D. Kollias
IEEE Trans. Syst. Man Cybern. Part A3
2003 Towards Supporting the Teaching of History Using an Intelligent Information System that Relies on the Electronic Road Metaphor
abstract
The new educational system in Greece, as far as teaching of history is concerned, is characterized by a shift away from sterile memorization and towards a critical approach of historical facts and phenomena that would contribute in both the development of pupils' historical concept and conscience and the promotion of critical thought. This reflects the globally accepted goals of teaching history courses. Such teaching goals can be greatly supported by computer-based applications, which can offer access to vast amounts of historical texts and data to be used next to the main scholar textbook and be analyzed by pupils. Still, existing applications seem to be quite inadequate for this purpose, as they require that the pupil be already informed on a matter, before the initiation of a quest for data. We describe an intelligent information system that is designed to facilitate browsing of educational material and historical sources, thus allowing pupils to efficiently retrieve information on topics that are not yet known to them and expand in this way their historical knowledge. This can help in fulfilling the teaching goals of the new educational system.
Manolis Wallace, Mary Stefanou, Kostas Karpouzis, Stefanos D. Kollias
ICALT4
2003 An Intelligent Scheme for Facial Expression Recognition
Amaryllis Raouzaiou, Spiros Ioannou, Kostas Karpouzis, Nicolas Tsapatsoulis, Stefanos D. Kollias, Roddy Cowie
ICANN5
2003 An Emotional Recognition Architecture Based on Human Brain Structure
John G. Taylor, Nickolaos F. Fragopanagos, Roddy Cowie, Ellen Douglas-Cowie, Stavroula-Evita Fotinea, Stefanos D. Kollias
ICANN6
2003 Knowledge Refinement Using Fuzzy Compositional Neural Networks
Vassilis Tzouvaras, Giorgos B. Stamou, Stefanos D. Kollias
ICANN3
2003 Adaptive rule-based recognition of events in video sequences
abstract
Knowledge-based fuzzy inference and neural learning are used in this paper in order to model the event recognition task in semantic video analysis. The advantage of their use is the symbolic nature of the representation of the knowledge concerning the events to be recognized. Moreover, this knowledge can be adapted with the aid of data taken from video sequences. The proposed system has been tested in soccer video sequences for detecting some complex predetermined (and represented in the form of rules) events.
Vassilis Tzouvaras, Gavriil Tsechpenakis, Giorgos B. Stamou, Stefanos D. Kollias
ICIP (2)4
2003 MPEG-4: one multimedia standard to unite all
abstract
A wide variety of multimedia coding and representation standards have emerged during the past years, in the quest to establish a common denominator for both the research community and the related industry to incorporate results and algorithms in commercial products. MPEG-4 aims to improve where past standards proved ineffective or isolated and integrate different modalities to supply different forms of material, suitable for diverse applications and devices. In this paper, we describe the general concepts that MPEG-4 builds upon and how they are related to multi-user interactive environments and gaming.
Kostas Karpouzis, Amaryllis Raouzaiou, Paraskevi K. Tzouveli, Spiros Ioannou, Stefanos D. Kollias
ICME5
2003 Emotion representation for online gaming
abstract
The ability to simulate lifelike interactive characters has many applications in the gaming industry. Human faces may act as visual interfaces that help users feel at home when interacting with a computer because they are accepted as the most expressive means for communicating and recognizing emotions. Thus, a lifelike human face can enhance interactive applications by providing straightforward feedback to and from the users and stimulating emotional responses from them. Thus, the gaming and entertainment industries can benefit from employing believable, expressive characters since such features significantly enhance the atmosphere of a virtual world and communicate messages far more vividly than any textual or speech information. In this paper, we present an abstract means of description of facial expressions, by utilizing concepts included in the MPEG-4 standard. Furthermore, we exploit these concepts to synthesize a wide variety of expressions using a reduced representation, suitable for networked and lightweight applications.
Amaryllis Raouzaiou, Kostas Karpouzis, Stefanos D. Kollias
ICME3
2003 Object tracking in clutter and partial occlusion through rule-driven utilization of Snakes
abstract
Efficient moving object tracking in cases of partial occlusion is a challenging task for the researchers in the fields of computer vision and video processing. In modern coding standards, like MPEG-4 and MPEG-7, the term of video objects is used to define moving objects in a video sequence. Automatic extraction of such objects is by no means trivial, and occlusion is one of the most important problems. This paper reformulates one of the most popular deformable templates for shape modeling and object tracking, the Snakes, in a probabilistic manner, in order to include providence for partial occlusion of the moving objects. Experiments of object tracking in partial occlusion, in complex natural sequences, where temporal clutter, abrupt motion and external lightning changes have been carried out, showing the efficiency of the proposed approach.
Gavriil Tsechpenakis, Konstantinos Rapantzikos, Nicolas Tsapatsoulis, Stefanos D. Kollias
ICME4
2003 A modular approach to facial feature segmentation on real sequences
George N. Votsis, Athanasios I. Drosopoulos, Stefanos D. Kollias
Signal Process. Image Commun.3
2003 An adaptable neural-network model for recursive nonlinear traffic prediction and modeling of MPEG video sources
abstract
Multimedia services and especially digital video is expected to be the major traffic component transmitted over communication networks [such as internet protocol (IP)-based networks]. For this reason, traffic characterization and modeling of such services are required for an efficient network operation. The generated models can be used as traffic rate predictors, during the network operation phase (online traffic modeling), or as video generators for estimating the network resources, during the network design phase (offline traffic modeling). In this paper, an adaptable neural-network architecture is proposed covering both cases. The scheme is based on an efficient recursive weight estimation algorithm, which adapts the network response to current conditions. In particular, the algorithm updates the network weights so that 1) the network output, after the adaptation, is approximately equal to current bit rates (current traffic statistics) and 2) a minimal degradation over the obtained network knowledge is provided. It can be shown that the proposed adaptable neural-network architecture simulates a recursive nonlinear autoregressive model (RNAR) similar to the notation used in the linear case. The algorithm presents low computational complexity and high efficiency in tracking traffic rates in contrast to conventional retraining schemes. Furthermore, for the problem of offline traffic modeling, a novel correlation mechanism is proposed for capturing the burstness of the actual MPEG video traffic. The performance of the model is evaluated using several real-life MPEG coded video sources of long duration and compared with other linear/nonlinear techniques used for both cases. The results indicate that the proposed adaptable neural-network architecture presents better performance than other examined techniques.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
IEEE Trans. Neural Networks3
2003 An efficient fully unsupervised video object segmentation scheme using an adaptive neural-network classifier architecture
abstract
In this paper, an unsupervised video object (VO) segmentation and tracking algorithm is proposed based on an adaptable neural-network architecture. The proposed scheme comprises: 1) a VO tracking module and 2) an initial VO estimation module. Object tracking is handled as a classification problem and implemented through an adaptive network classifier, which provides better results compared to conventional motion-based tracking algorithms. Network adaptation is accomplished through an efficient and cost effective weight updating algorithm, providing a minimum degradation of the previous network knowledge and taking into account the current content conditions. A retraining set is constructed and used for this purpose based on initial VO estimation results. Two different scenarios are investigated. The first concerns extraction of human entities in video conferencing applications, while the second exploits depth information to identify generic VOs in stereoscopic video sequences. Human face/ body detection based on Gaussian distributions is accomplished in the first scenario, while segmentation fusion is obtained using color and depth information in the second scenario. A decision mechanism is also incorporated to detect time instances for weight updating. Experimental results and comparisons indicate the good performance of the proposed scheme even in sequences with complicated content (object bending, occlusion).
Anastasios Doulamis, Nikolaos D. Doulamis, Klimis S. Ntalianis, Stefanos D. Kollias
IEEE Trans. Neural Networks4
2002 Context - Sensitive Query Expansion Based on Fuzzy Clustering of Index Terms
Giorgos Akrivas, Manolis Wallace, Giorgos B. Stamou, Stefanos D. Kollias
FQAS4
2002 Neural Networks Retraining for Unsupervised Video Object Segmentation of Videoconference Sequences
Klimis S. Ntalianis, Nikolaos D. Doulamis, Anastasios Doulamis, Stefanos D. Kollias
ICANN4
2002 Video object articulation using depth-based content segmentation approaches
abstract
Two efficient unsupervised video object segmentation approaches are proposed and then extensively compared in terms of computational cost and quality of segmentation results. Both methods exploit depth information. In particular a depth segments map is initially estimated by analyzing a stereoscopic pair of frames and applying a segmentation algorithm. In the first, a "constrained fusion of color segments" (CFCS), video object segmentation is performed by fusion of color segments according to a depth similarity criterion. In the second approach, first, a dilated version of the boundary of each depth segment is produced and several feature points are estimated on this dilated boundary. Then, for each initial point, a normalized motion geometric space (MGS) is created which determines the only allowed path on to which the point can move. In the last step, each initial point moves on to its MGS and stops according to a weighted stop-function. Experiments on real life stereoscopic sequences are presented to exhibit the speed and accuracy of the proposed schemes.
Nikolaos D. Doulamis, Anastasios Doulamis, Stefanos D. Kollias, Klimis S. Ntalianis
ICIP (2)3
2002 An automatic scheme for stereoscopic video object-based watermarking using qualified significant wavelet trees
abstract
A fully automatic system for embedding visually recognizable watermark patterns to video objects is proposed. The architecture consists of 3 main modules. In the first module, unsupervised video object extraction is performed, by analyzing stereoscopic pairs of frames. Then each video object is decomposed into three levels with ten subbands, using the discrete wavelet transform (DWT) and three pairs of subbands are formed (HL/sub 3/, HL/sub 2/), (LH/sub 3/, LH/sub 2/) and (HH/sub 3/, HH/sub 2/). Next qualified significant wavelet trees (QSWTs) are estimated for the specific pair of subbands that contains the highest energy content compared to the other two pairs. QSWTs are derived from the embedded zerotree wavelet (EZW) algorithm and they are high-energy coefficient paths within the selected pair of subbands. Finally, in the third module, visually recognizable watermark patterns are redundantly embedded to the coefficients of the highest energy QSWTs and the inverse DWT is applied to provide the watermarked video object. The performance of the proposed video object watermarking system is tested under various signal distortions such as JPEG lossy compression, sharpening, blurring and adding different types of noise. Experimental results on real life stereoscopic images are presented to indicate the efficiency and robustness of the proposed scheme.
Nikolaos D. Doulamis, Anastasios Doulamis, Stefanos D. Kollias, Klimis S. Ntalianis
ICIP (3)3
2002 A robust steganographic wavelet-based system for resistant message hiding under error prone networks
abstract
A wavelet-based steganographic method is proposed for robust message hiding. The message is embedded into the most significant wavelet coefficients of a cover image to provide invisibility and resistance against lossy transmission, compression or other distortion. The architecture consists of three modules. In the first module, the initial message is enciphered by an encryption algorithm. The enciphered message is imprinted onto a white-background image to construct the message-image to be hidden. In the second module, the cover image is decomposed into two levels with seven subbands, using the DWT. Next, qualified significant wavelet trees (QSWTs), which are paths of significant wavelet coefficients, are estimated for the highest energy pair of subbands. In the third module, the message-image is redundantly embedded to the coefficients of the best QSWTs and the IDWT is applied to provide the stego-image. The robustness and efficiency of the proposed steganographic system is evaluated under various loss rates, combined with different JPEG compression ratios.
Klimis S. Ntalianis, Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
ICME (2)4
2002 Erratum to: "A fuzzy video content representation for video summarization and content-based retrieval" [Signal Processing 80(6) (2000) 1049-1067]
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
Signal Process.3
2001 Adaptable Neural Networks for Unsupervised Video Object Segmentation of Stereoscopic Sequences
Anastasios Doulamis, Klimis S. Ntalianis, Nikolaos D. Doulamis, Stefanos D. Kollias
ICANN4
2001 A Neural-Network-Based Approach to Adaptive Human Computer Interaction
George N. Votsis, Nikolaos D. Doulamis, Anastasios Doulamis, Nicolas Tsapatsoulis, Stefanos D. Kollias
ICANN5
2001 Tube-embodied gradient vector flow fields for unsupervised video object plane (VOP) segmentation
abstract
In this paper constrained gradient vector flow (GVF) field generation is performed, for fast and accurate unsupervised stereoscopic semantic segmentation. The scheme utilizes the information provided by a depth segments map, produced by stereo analysis methods and incorporation of a segmentation algorithm. Then a Canny edge detector is applied to the depth region and produces an edge map. The edge map is used for tube estimation inside which the GVF field evolves. After generation of the GVF field an active contour is unsupervisedly initialized onto the outer bound of the tube. Finally a greedy approach is adopted and the active contour, guided by the GVF field, extracts the VOP. Experimental results on real life stereoscopic video sequences indicate the efficiency of the proposed scheme.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias, Klimis S. Ntalianis
ICIP (2)3
2001 Automatic Content-Based Organization Of Video Sequences For Multimedia Applications
abstract
In this paper an efficient scheme for automatic organization of stereo-captured video sequences is presented, which exploits foreground VOP information of frames. More specifically after shot cut detection, for each frame of a shot, a fast, unsupervised foreground VOP extraction algorithm is applied, based on depth information and normalized Motion Geometric Spaces. Then for each frame, a feature vector is constructed, containing color and size characteristics of the foreground VOPs within the frame. Afterwards, for a given shot, key frames are extracted using an optimization method for locating minimally correlated feature vectors. Finally correlation links are generated between the extracted key frames to produce a graph-like structure of the video sequence. Experimental results on real life stereoscopic sequences indicate the promising performance of the proposed scheme.
Klimis S. Ntalianis, Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
ICME4
2001 Multiresolution Gradient Vector Flow Field: A Fast Implementation Towards Video Object Plane Segmentation
abstract
In this paper, an efficient scheme for video object segmentation is proposed. The scheme is based on a multiresolution Gradient Vector Flow field (M-GVF) and a Motion Geometric Space (MGS) formulation. In particular the proposed scheme is initialized from an object approximation which can be provided either (a) automatically (unsupervised case) based on a depth map estimation method or (b) semi-automatically by user interaction. In the following, several feature points are estimated on the initial object contour (i.e. depth object) and an M-GVF adapted MGS is created to determine the direction that a feature point is allowed to move to. In this framework, each feature point moves onto its MGS in order to locate the contour of the physical video object. Experimental results are presented to indicate the reliable performance of the proposed scheme on real life stereoscopic and monocular video sequences.
Klimis S. Ntalianis, Nikolaos D. Doulamis, Stefanos D. Kollias, Anastasios Doulamis
ICME3
2001 Visual Information Retrieval From Annotated Large Audiovisual Assets Based On User Profiling And Collaborative Recommendations
abstract
Current multimedia databases contain a wealth of information in the form of audiovisual and text data. Even though efficient search algorithms have been developed for either media, there still exists the need for abstract data presentation and summarization. Moreover, retrieval systems should be capable of providing the user with additional information related to the specific subject of the query, as
Klimis S. Ntalianis, Spiros Ioannou, Kostas Karpouzis, George Moschovitis, Stefanos D. Kollias
ICME5
2001 A multiscale tree-structure for fast browsing and effective transmission of video files
abstract
A multiscale video content organization scheme is proposed for fast browsing and efficient transmission of video sequences. The scheme leads to construction of a five-layer tree structure. At layer 0 the root-node is located, connected to all nodes of layer 1, each corresponding to a class of shots. Then every class of shot is expanded at layer 2. The nodes of this layer represent shots. At the next resolution level (layer 3) nodes represent key-frames of shots. Finally at layer 4 the full resolution level is reached, where nodes correspond to frames of the sequence. Each node contains a viewing element and we focus on the extraction of these elements for layers 1, 2 and 3. Viewing elements of layers 1 and 3 are optimally extracted by minimizing a cross correlation criterion. Additionally viewing elements of layer 2 are selected according to a correlation measure between the mean vector of a shot and each of the frames within this shot. The resulting tree-structure enables a user to quickly and easily detect content of interest, by selecting the viewing element of his/her liking. Experimental results on real-life video sequences indicate the promising performance of the proposed scheme.
Klimis S. Ntalianis, Nikolaos D. Doulamis, Anastasios Doulamis, Ioannis Z. Koukoutsidis, Stefanos D. Kollias
MMSP5
2001 Affine-invariant curve normalization for object shape representation, classification, and retrieval
Yannis Avrithis, Yiannis Xirouhakis, Stefanos D. Kollias
Mach. Vis. Appl.3
2001 Facial Image Indexing in Multimedia Databases
Nicolas Tsapatsoulis, Yannis Avrithis, Stefanos D. Kollias
Pattern Anal. Appl.3
2001 A novel iron loss reduction technique for distribution transformers based on a combined genetic algorithm - neural network approach
abstract
The paper presents an effective method to reduce the iron losses of wound core distribution transformers based on a combined neural network/genetic algorithm approach. The originality of the work presented is that it tackles the iron loss reduction problem during the transformer production phase, while previous works concentrated on the design phase. More specifically, neural networks effectively use measurements taken at the first stages of core construction in order to predict the iron losses of the assembled transformers, while genetic algorithms are used to improve the grouping process of the individual cores by reducing iron losses of assembled transformers. The proposed method has been tested on a transformer manufacturing industry. The results demonstrate the feasibility and practicality of this approach. Significant reduction of transformer iron losses is observed in comparison to the current practice leading to important economic savings for the transformer manufacturer.
Pavlos S. Georgilakis, Nikolaos D. Doulamis, Anastasios Doulamis, Nikos D. Hatziargyriou, Stefanos D. Kollias
IEEE Trans. Syst. Man Cybern. Syst.5
2000 Efficient Face Detection for Multimedia Applications
abstract
Face detection is becoming an important tool in the framework of many multimedia applications. Several face detection algorithms based on skin color characteristics have appeared in the literature. Most of them have generalization problems due to the skin color model they use. We present a study which attempts to minimize the generalization problem by combining the M-RSST color segmentation algorithm with a Gaussian model of the skin color distribution and global shape features. Moreover by associating the resultant segments with a face probability we can index and retrieve facial images from multimedia databases
Nicolas Tsapatsoulis, Yannis Avrithis, Stefanos D. Kollias
ICIP3
2000 Affine-Invariant Curve Normalization for Shape-Based Retrieval
abstract
A novel method for two-dimensional curve normalization with respect to affine transformations is presented in this paper, allowing an affine-invariant curve representation to be obtained without any actual loss of information on the original curve. It can be applied as a pre-processing step to any shape representation, classification, recognition or retrieval technique, since it effectively decouples the problem of affine-invariant description from feature extraction and pattern matching. Curves estimated from object contours are first modeled by cubic B-splines and then normalized in several steps in order to eliminate translation, scaling, skew, starting point, rotation and reflection transformations, based on a combination of curve features including moments and Fourier descriptors.
Yannis Avrithis, Yiannis Xirouhakis, Stefanos D. Kollias
ICPR3
2000 Recursive Non Linear Models for On Line Traffic Prediction of VBR MPEG Coded Video Sources
abstract
Any performance evaluation of broadband networks requires modeling of the actual network traffic. Since multimedia services and especially MPEG coded video streams are expected to be a major traffic component over these networks, modeling of such services and traffic prediction are useful for the reliable operation of the broadband based an Asynchronous Transfer Mode (ATM) networks. In this paper, a recursive implementation of a Non linear AutoRegressive model (RNAR) is presented for on line traffic prediction of Variable Bit Rate (VBR) MPEG-2 video sources. This is accomplished by using an efficient weight adaptation algorithm so that the network provide good performance even in case of highly fluctuated traffic rates. In particular, the network weights are adapted so that the output is approximately equal to the current data while preserving the former knowledge of the network. Experimental results are presented to show the good performance of the proposed scheme. Furthermore, comparison with other linear or non linear techniques is presented to show that the adopted method yields better results than the other ones.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
IJCNN (6)3
2000 On Emotion Recognition of Faces and of Speech Using Neural Networks, Fuzzy Logic and the ASSESS System
abstract
We propose a framework for the processing of face image sequences and speech, using different dynamic techniques to extract appropriate features for emotion recognition. The features will be used by a hybrid classification procedure, employing neural network techniques and fuzzy logic, to accumulate the evidence for the presence of an emotional expression of the face and the speaker's voice.
Winfried A. Fellenz, John G. Taylor, Roddy Cowie, Ellen Douglas-Cowie, Frédéric Piat, Stefanos D. Kollias, Christos Orovas, Bruno Apolloni
IJCNN (2)6
2000 Efficient video summarization based on a fuzzy video content representation
abstract
A fuzzy representation of visual content is proposed, which is useful for video summarization. In particular, a multidimensional fuzzy histogram is constructed for each video frame based on a collection of appropriate features, extracted using video sequence analysis techniques. Then, key frames are selected optimally by minimizing a cross correlation criterion. Experimental results and comparison with other known methods are presented to indicate the good performance of the proposed scheme on real life video recordings.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
ISCAS3
2000 A fuzzy video content representation for video summarization and content-based retrieval
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
Signal Process.3
2000 Efficient summarization of stereoscopic video sequences
abstract
An efficient technique for summarization of stereoscopic video sequences is presented, which extracts a small but meaningful set of video frames using a content-based sampling algorithm. The proposed video-content representation provides the capability of browsing digital stereoscopic video sequences and performing more efficient content-based queries and indexing. Each stereoscopic video sequence is first partitioned into shots by applying a shot-cut detection algorithm so that frames (or stereo pairs) of similar visual characteristics are gathered together. Each shot is then analyzed using stereo-imaging techniques, and the disparity field, occluded areas, and depth map are estimated. A multiresolution implementation of the recursive shortest spanning tree (RSST) algorithm is applied for color and depth segmentation, while fusion of color and depth segments is employed for reliable video object extraction. In particular, color segments are projected onto depth segments so that video objects on the same depth plane are retained, while at the same time accurate object boundaries are extracted. Feature vectors are then constructed using multidimensional fuzzy classification of segment features including size, location, color, and depth. Shot selection is accomplished by clustering similar shots based on the generalized Lloyd-Max algorithm, while for a given shot, key frames are extracted using an optimization method for locating frames of minimally correlated feature vectors. For efficient implementation of the latter method, a genetic algorithm is used. Experimental results are presented, which indicate the reliable performance of the proposed scheme on real-life stereoscopic video sequences.
Nikolaos D. Doulamis, Anastasios Doulamis, Yannis Avrithis, Klimis S. Ntalianis, Stefanos D. Kollias
IEEE Trans. Circuits Syst. Video Technol.5
2000 On-line retrainable neural networks: improving the performance of neural networks in image analysis problems
abstract
A novel approach is presented in this paper for improving the performance of neural-network classifiers in image recognition, segmentation, or coding applications, based on a retraining procedure at the user level. The procedure includes: 1) a training algorithm for adapting the network weights to the current condition; 2) a maximum a posteriori (MAP) estimation procedure for optimally selecting the most representative data of the current environment as retraining data; and 3) a decision mechanism for determining when network retraining should be activated. The training algorithm takes into consideration both the former and the current network knowledge in order to achieve good generalization. The MAP estimation procedure models the network output as a Markov random field (MRF) and optimally selects the set of training inputs and corresponding desired outputs. Results are presented which illustrate the theoretical developments as well as the performance of the proposed approach in real-life experiments.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
IEEE Trans. Neural Networks Learn. Syst.3
1999 Statistical Multiplexing and Quality of Service Control of VBR Mpeg Video Sources
abstract
In this paper efficient modeling of VBR MPEG coded video sources is proposed by appropriately combining properties of frame and GOP layer signals. In particular a Markov chain is presented for modeling the video activity, the states of which correspond to correlated AR models responsible for generating the I, P and B frames. Furthermore, an adaptive implementation of the AR coefficients is accomplished in cases that we are interested in video traffic prediction. Experimental results using long duration sequences care provided to indicate the good performance of the proposed modeling scheme.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
ICIP (1)3
1999 A Neural Network Approach to Interactive Content-Based Retrieval of Video Databases
abstract
A neural network scheme is presented in this paper for adaptive video indexing and retrieval. First, a limited but characteristic amount of frames are extracted from each video scene, by minimizing a cross-correlation criterion. Low level features are extracted to indicate the frame characteristics, such as color and motion segments. This is due to the fact that extraction of high-level, semantic, features from any kind of images is too hard to be implemented. After the key frame extraction, the video queries are implemented directly on this small number of frames. To reduce, however, the limitation of low-level features the human is considered as a part of the process, meaning that he/she is able to assign a degree of appropriateness for each retrieved image of the system and then restart the searching. A feedforward neural network structure is proposed as a parametric distance for the retrieval, mainly due to the highly non linear capabilities. An adaptation mechanism is also proposed for updating the network weights, each time a new image selection is performed by the user. This mechanism can modify the network weights so that the output of the network, after the adaptation, is as much as close to the user's selection while simultaneously performing a minimal degradation of the previous learned data.
Nikolaos D. Doulamis, Anastasios Doulamis, Stefanos D. Kollias
ICIP (2)3
1999 Adaptive Wavelet-Packet Decomposition for Rate Control of Object Oriented Coding of Video Sequences
abstract
A novel object dependent coding scheme based on optimization upon wavelet packet trees is proposed. The method may be employed in the context of object oriented video standard MPEG-4 for constraint distortion minimization for each video object (VO). Optimal transmission rates are evaluated for each VO which guarantee that important VO (like the foreground) are transmitted with distortions which are equal or less than preselected distortion values. The proposed algorithm is applied to object dependent coding of video frames of the Akiyo sequence with peak-signal-to-noise (PSNR) values, which are higher for the foreground than the background. Experimental results, which favor the proposed method, are provided as well.
Ioannis M. Stephanakis, Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
ICIP (2)4
1999 Piecewise Wiener filter model based on fuzzy partition of local wavelet features for image restoration
abstract
Autoregressive Wiener filters are used for prediction and restoration of still frame and video images. Filters of this kind solve a linear optimization problem for the global statistics of an image. They fail when image statistics vary in space (non-stationarity) and when the corrupting noise is nonlinear. A piecewise Wiener filter defined upon a fuzzy partition of the space of local wavelet features is presented and successfully applied to image restoration in the aforementioned cases. Unsupervised clustering of the features using the Bezdek fuzzy c-means algorithm is performed for region estimation and subsequent application of the proper filter h/sub Rk/(n, m) according to a degree of belief /spl mu//sub Rk/. Experimental results indicate increased improvements in signal-to-noise ratios of corrupted images using the proposed method.
Ioannis M. Stephanakis, George Stamou, Stefanos D. Kollias
IJCNN3
1999 A stochastic framework for optimal key frame extraction from MPEG video databases
abstract
A framework for video content representation is proposed in this paper for extracting limited, but meaningful, information of video data directly from MPEG compressed domain. First, the traditional frame-based representation is transformed to a feature-based one. Then, all features are gathered together using a fuzzy formulation and extraction of several key frames is performed for each shot in a content-based rate sampling framework. In particular, our approach is based on minimization of a cross-correlation criterion among video frames of a given shot so as to be located a set of minimally correlated feature vectors. Experimental results indicating the good performance of the proposed scheme are also presented.
Nikolaos D. Doulamis, Anastasios Doulamis, Yannis Avrithis, Stefanos D. Kollias
MMSP4
1999 Modeling and adaptive prediction of VBR MPEG video sources
abstract
In this paper, efficient modeling of variable bit-rate (VBR) MPEG-coded video sources is proposed by appropriately combining properties of frame and GOP (group of pictures) layer signals. In particular, a Markov chain is presented for modeling the video activity and the states of which correspond to correlated autoregressive (AR) models that are responsible for generating the I (intra-frame), P (predictive) and B (bidirectionally predictive) frames. Furthermore, an adaptive implementation of the AR coefficients is accomplished in cases where we are interested in video traffic prediction.
Nikolaos D. Doulamis, Anastasios Doulamis, Stefanos D. Kollias
MMSP3
1999 A pyramidal graph representation for efficient image content description
abstract
In this paper, an efficient image content representation is proposed, which is appropriate for content-based indexing and retrieval. In particular, the image content is represented by a hierarchical graph, constructed in a tree or pyramid structure framework. The basic concept is to represent the spatial relations of the objects in the image with an image-graph and then, the object details with an object-graph. Further decomposition of the objects regions is permitted in a hierarchical framework.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
MMSP3
1999 Efficient use of content-based retrieval in quick browsing: a realization
abstract
Due to recent growth of interest in multimedia applications, an increasing demand has emerged for efficient storage, management and browsing in multimedia databases. This fact has made the content-based retrieval (CBR) concept very popular during the past decade. However, it has been practically shown that CBR tools respond successfully to the inexperienced user queries very rarely. It is for this reason that in commercial systems, queries are supported by the use of annotations or manual indexing. In this work, content-based query, retrieval and indexing capabilities have been combined with an intelligent agent framework over a simple database architecture. The proposed system is based on the ideas presented in (Xirouhakis et al., 1998) which are implemented and further extended in this work.
George N. Votsis, Yiannis Xirouhakis, Kostas Karpouzis, Stefanos D. Kollias
MMSP4
1999 A Stochastic Framework for Optimal Key Frame Extraction from MPEG Video Databases
Yannis Avrithis, Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
Comput. Vis. Image Underst.4
1998 Face extraction from non-uniform background and recognition in compressed domain
abstract
A complete face recognition system is proposed in this paper by introducing the concepts of foreground objects, which are currently used in the MPEG-4 standardization phase, to human identification. The system automatically detects and extracts the human face from the background, even if is not uniform, based on a combination of a retrainable neural network structure and the morphological size distribution technique. In order to combine face images of high quality and low computational complexity, the recognition stage is performed in the compressed domain. Thus, in contrast to existing recognition schemes, the face images are available in their original quality and not only in their transformed representation.
Nicolas Tsapatsoulis, Nikolaos D. Doulamis, Anastasios Doulamis, Stefanos D. Kollias
ICASSP4
1998 Video Content Representation using Optimal Extraction of Frames and Scenes
abstract
An efficient video content representation is proposed using optimal extraction of characteristic frames and scenes. This representation, apart from providing browsing capabilities to digital video databases, also allows more efficient content-based queries and indexing. For performing the frame/scene extraction, a feature vector formulation of the images is proposed based on color and motion segmentation. Then, the scene selection is accomplished by clustering similar scenes based on a distortion criterion. Frame selection is performed using an optimization method for locating a set of minimally correlated feature vectors.
Nikolaos D. Doulamis, Anastasios Doulamis, Yannis Avrithis, Stefanos D. Kollias
ICIP (1)4
1998 A Neural Network based Scheme for Unsupervised Video Object Segmentation
abstract
We propose a neural network based scheme for performing unsupervised video object segmentation, especially for videophone or videoconferencing applications. The procedure includes (a) a training algorithm for adapting the network weights to the current condition, (b) a maximum a posteriori (MAP) estimation procedure for optimally selecting the most representative data of the current environment as retraining data and (c) a decision mechanism for determining when network retraining should be activated. The training algorithm takes into consideration both the former and the current network knowledge in order to achieve good generalization. The MAP estimation procedure models the network output as a Markov random field (MRF) and optimally selects the set of training inputs and corresponding desired outputs, using initial estimates of the human face and body. Finally, a verification mechanism is introduced which augments the training data, exploiting information of the previous and current environment.
Anastasios Doulamis, Nikolaos D. Doulamis, Stefanos D. Kollias
ICIP (2)3
1998 The rendering pipeline in the classroom: a diversified approach
abstract
In this paper we describe an integrated method of teaching an introductory computer graphics course. Most such courses are simply "art-oriented", that is they focus on getting students to use modern commercial software, so as to prepare them for a corresponding career, or concentrate on the basic concepts of graphics theory and merely provide a theoretical foundation, such as simple translations and projections; in this case, they usually fail to motivate the class by producing practical interesting examples. The curriculum that we propose combines theoretical knowledge of introductory computer graphics concepts and techniques with laboratory work in programming or modelling and animation exercises. This set of applied laboratory exercises is relevant to the material taught in class, but also extends to familiarising students with the modern uses of computer generated imagery, such as films, virtual worlds or medical imaging. The feedback from the students, combined with their success in the course, shows that this coupled teaching and immersion material is by far more interesting and challenging, while still providing them with the essential academic background.
Kostas Karpouzis, Stefanos D. Kollias
ITiCSE2
1998 Facial expression recognition using HMM with observation dependent transition matrix
abstract
An expression recognition technique is proposed based on the hidden Markov models (HMM) ability to deal with time sequential data and to provide time scale invariability as well as a learning capability. A feature vector sequence is used for this purpose, which relies on optical flow extraction, as well as directional filtering of the motion field. Segmentation and identification of important facial parts are preceding feature extraction. The HMM is enhanced with an observation dependent transition matrix, being able to cope with the dynamics of emotions and the severe complexity of expressions timing. Experimental results are included illustrating the effectiveness of this method.
Nicolas Tsapatsoulis, Miltiades Leonidou, Stefanos D. Kollias
MMSP3
1998 Efficient browsing in multimedia databases using intelligent agents and content-based retrieval schemes
abstract
An intelligent multimedia system is developed for efficient retrieval and processing of information stored in multimedia databases. Intelligent agents are employed in order to limit the search into minor subsets of the database; a training strategy is introduced for quick browsing in the database contents. Content-based query and retrieval schemes are also employed to increase the effectiveness of the browsing procedure.
Yiannis Xirouhakis, George N. Votsis, Kostas Karpouzis, Stefanos D. Kollias
MMSP4
1998 Prediction of iron losses of wound core distribution transformers based on artificial neural networks
Pavlos S. Georgilakis, Nikos D. Hatziargyriou, Nikolaos D. Doulamis, Anastasios Doulamis, Stefanos D. Kollias
Neurocomputing5
1998 Low bit-rate coding of image sequences using adaptive regions of interest
abstract
An adaptive algorithm for extracting foreground objects from background in videophone or videoconference applications is presented. The algorithm uses a neural network architecture that classifies the video frames in regions of interest (ROI) and non-ROI areas, also being able to automatically adapt its performance to scene changes. The algorithm is incorporated in motion-compensated discrete cosine transform (MC-DCT)-based coding schemes, allocating more bits to ROI than to non-ROI areas. Simulation results are presented, using the Claire and Trevor sequences, which show reconstructed images of better quality, as well as signal-to-noise ratio improvements of about 1.4 dB, compared to those achieved by standard MC-DCT encoders.
Nikolaos D. Doulamis, Anastasios Doulamis, Dimitrios Kalogeras, Stefanos D. Kollias
IEEE Trans. Circuits Syst. Video Technol.4
1997 Improving the performance of MPEG compatible encoders using on line retrainable neural network
abstract
On line retraining of neural network is introduced for extracting foreground/background objects in video sequences. The scheme is applied together with a modification of the rate control of MPEG-1 algorithm. The proposed method is compatible to MPEG-1/2 standard but also can be used as a pre-coding stage for the forthcoming MPEG-4 algorithm. Simulation studies have shown an improvement of about 1.5 dB on average as far the PSNR is concerned compared with the conventional MPEG-1 encoder.
Stefanos D. Kollias, Nikolaos D. Doulamis, Anastasios Doulamis
ICIP (3)1
1997 An image analysis system for automated detection of breast cancer nuclei
abstract
A study for breast cancer nuclei detection is presented. The proposed algorithm determines the centers of nuclei in biopsy images using block-based processing of the images followed by singular value decomposition of each block. The normalized singular value vector, which consists of the normalized singular values of the block in decreasing order, is fed as the input to an appropriate neural network which classifies the block into nuclei and background ones. Examples are presented which illustrate the ability of the proposed technique to model the knowledge provided by experts in the nuclei detection task.
Nicolas Tsapatsoulis, Frank Schnorrenberg, Constantinos S. Pattichis, Stefanos D. Kollias
ICIP (3)4
1996 A multiresolution neural network approach to invariant image recognition
Stefanos D. Kollias
Neurocomputing1
1996 Neural network-assisted effective lossy compression of medical images
abstract
A neural network architecture is proposed and shown to be very effective in performing lossy compression of medical images. A novel ROI-JPEG technique is introduced as the coding platform, in which the neural architecture adoptively selects regions of interest (ROI's) in the images. By letting the selected ROI's be coded with high quality, in contrast to the rest of image areas, high compression ratios are achieved, while retaining the significant (from medical point of view) image content. The performance of the method is illustrated by means of experimental results in real life problems taken from pathology and telemedicine applications. © 1996 IEEE.
Nikos G. Panagiotidis, Dimitrios Kalogeras, Stefanos D. Kollias, Andreas Stafylopatis
Proc. IEEE3
1995 Two-dimensional filter bank design for optimal reconstruction using limited subband information
abstract
In this correspondence, we propose design techniques for analysis and synthesis filters of 2-D perfect reconstruction filter banks (PRFB's) that perform optimal reconstruction when a reduced number of subband signals is used. Based on the minimization of the squared error between the original signal and some low-resolution representation of it, the 2-D filters are optimally adjusted to the statistics of the input images so that most of the signal's energy is concentrated in the first few subband components. This property makes the optimal PRFB's efficient for image compression and pattern representations at lower resolutions for classification purposes. By extending recently introduced ideas from frequency domain principal component analysis to two dimensions, we present results for general 2-D discrete nonstationary and stationary second-order processes, showing that the optimal filters are nonseparable. Particular attention is paid to separable random fields, proving that only the first and last filters of the optimal PRFB are separable in this case. Simulation results that illustrate the theoretical achievements are presented.
Andreas Tirakis, Anastasios Delopoulos, Stefanos D. Kollias
IEEE Trans. Image Process.3
1994 A multiresolution probabilistic neural network for image segmentation
abstract
A multiresolution network for segmenting textures and magnetic resonance images is proposed, based on maximum likelihood estimation. The network incorporates a probabilistic neural architecture to facilitate the generation of likelihood estimates. Further on, an iterative segmentation process is used, which refines the likelihood estimates based upon both the neighbouring estimated likelihoods and the confidence on these estimates. A multiresolution neural network structure which permits a significant reduction of the time needed to solve the segmentation problem is proposed. This is performed by an initial segmentation at lower resolution and subsequent refinement at higher resolutions.>
Stefanos D. Kollias, Dimitrios Kalogeras
ICASSP (2)1
1994 Adaptive classification of textured images using linear prediction and neural networks
Levon Sukissian, Stefanos D. Kollias, Yiannis S. Boutalis
Signal Process.2
1994 Invariant image classification using triple-correlation-based neural networks
abstract
Triple-correlation-based neural networks are introduced and used in this paper for invariant classification of 2D gray scale images. Third-order correlations of an image are appropriately clustered, in spatial or spectral domain, to generate an equivalent image representation that is invariant with respect to translation, rotation, and dilation. An efficient implementation scheme is also proposed, which is robust to distortions, insensitive to additive noise, and classifies the original image using adequate neural network architectures applied directly to 2D image representations. Third-order neural networks are shown to be a specific category of triple-correlation-based networks, applied either to binary or gray-scale images. A simulation study is given, which illustrates the theoretical developments, using synthetic and real image data.
Anastasios Delopoulos, Andreas Tirakis, Stefanos D. Kollias
IEEE Trans. Neural Networks3
1993 Efficient image classification using neural networks and multiresolution analysis
Andreas Tirakis, Stefanos D. Kollias
ICASSP (1)2
1992 A Progressive Scheme for Digital Image Halftoning, Coding of Halftones, and Reconstruction
abstract
A digital halftoning technique for the efficient transformation of gray-scale images into bilevel ones, based on the progressive generation of the bilevel image pixels in a parallel way, is presented. An image distortion criterion, in which the gray-tone image is approximated by a filtered version of the halftoned image, is used for this purpose. A combined scheme is also derived in which continuous-tone images are progressively coded and transmitted in bilevel form and can be reconstructed in gray-scale form.>
Stefanos D. Kollias, Dimitris Anastassiou
IEEE J. Sel. Areas Commun.1
1990 Adaptive classification of textured images using moments and autoregressive models
abstract
An adaptive approach to the classification of textured images is presented, based on the extraction of appropriate features from images. Autoregressive linear prediction models, as well as moments of images, are features which are examined and compared in the paper. Classification is achieved in an adaptive way, using an artificial feedforward neural network, which is trained by examples, using an efficient variant of the backpropagation learning algorithm. It is also shown that an adaptive least squares estimation algorithm can be appropriately interweaved with the network, resulting in an on-line adaptive classification scheme. Simulation results are given, which illustrate the performance of the presented method.
Levon Sukissian, Andreas Tirakis, Stefanos D. Kollias
VCIP3
1989 Image halftoning and reconstruction using a neural network
abstract
Digital image halftoning is treated as an optimization problem to which neural networks provide an efficient parallel solution. An image distortion measure is introduced in which the gray-tone image is approximated by a filtered version of the halftoned image. This distortion measure is minimized by using a near-neighborhood-connected symmetric neural network. The filter used in the distortion measure is estimated on the basis of a training set of gray-tone and bilevel images. This filter is then used for the reconstruction of a gray-tone image from its halftoned version. The above procedure, combined with a postprocessing of the reconstructed image by a nonlinear edge-preserving noise-smoothing filter, provides images of good quality.>
Stefanos D. Kollias, Tu-Chih Tsai, Dimitris Anastassiou
ICASSP1
1988 A fast technique for automatic segmentation and classification of textured images
abstract
A fast computationally efficient method for automatic segmentation and classification of textured images is presented. The method does not necessarily need a-priori information about the textures present in the image, thus avoiding the necessity of a training set of textures. A fast adaptive multichannel technique for autoregressive image model parameter estimation with fast tracking capabilities and a powerful statistical distance measure are appropriately interweaved to form the proposed technique. Specific properties of the estimation part of the algorithm are exploited to reduce greatly the computational complexity of the distance measure. Some interesting extensions of the method are discussed and examples are given which illustrate the performance of the algorithm.>
Yiannis S. Boutalis, Stefanos D. Kollias, George Carayannis, Levon Sukissian
ICASSP2
1987 A fast multichannel approach to adaptive estimation and filtering of two dimensional images
abstract
A fast, computationally efficient method for adaptive image estimation is presented in this paper. The method is based on the multichannel form of the Fast-A-Posteriori-Error Sequential Technique (FAEST) for the estimation of the parameters of 2-D autoregressive (AR) models. The above models are appropriately designed to have the shift invariance property necessary for fast recursive least squares techniques. Moreover inclusion of a forgetting factor, or of a sliding window in the algorithm, permits adaptive image model parameter estimation. The specific properties of the algorithm are examined and it is shown that the proposed estimation scheme is much faster than other existing approaches. Various interesting applications of the method are discussed and examples are given which illustrate the above theoretical results.
Yiannis S. Boutalis, Stefanos D. Kollias, George Carayannis
ICASSP2
1983 A model reduction algorithm by spline approximation in the deconvolution of seismic signals
abstract
This paper presents a new method for obtaining ap - proximate low-order recursive (ARMA) realizations of convolution models. It is based upon optimum piecewise linear or cubic spline approximation of the convolution kernel. The method may be efficiently used in the deconvolution of seismic signal. The basic seismic wavelet is approximated by splines and a fixed-lag smoothing state-space representation, equivalent to the resulting ARMA model, is derived. Use of Kalman filtering gives fixed-lag smoothed estimates of the so-called reflection coefficient sequence. Adaptive estimation is used in the case when the seismic wavelet is not known apriori. Simulation results using a Ricker wavelet are presented, which illustrate the performance af the proposed method.
Stefanos D. Kollias, Cristos C. Halkias
ICASSP1