Petros Maragos

dblp:22/4003 · DBLP profile ↗
← Back
228ranked-venue papers
36as first author
38since 2021 · last 2025
0000-0003-0534-2707ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 173 · 28 first-author · 25 since 2021Artificial intelligence and machine learning · 84 · 8 first-author · 19 since 2021Systems, architecture and hardware · 15 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Power in Unity: Combining in-Domain and out-of-Domain Pre-Training Strategies for EEG-Based Person Identification
abstract
We present the NTUA-IRAL team’s solution for the Person Identification track of the Signal Processing EEG-Music Emotion Recognition Grand Challenge, hosted at ICASSP. Our approach employs an ensemble of three CNNs, each pretrained using a distinct strategy: contrastive pre-training, traditional ImageNet pre-training, and task-specific pre-training on a publicly available EEG dataset. This diverse pre-training regimen enabled our models to achieve a test set accuracy of 100%, earning third place in the challenge subtrack.
Christos Garoufis, Marios Glytsos, Ioanna Chourdaki, Panagiotis Paraskevas Filntisis, Petros Maragos
ICASSP5
2025 Towards Open-Ended Robotic Exploration Using Vision-Inspired Similarity and Foundation Models
abstract
In the domain of robotics, achieving Lifelong Open-ended Learning Autonomy (LOLA) represents a significant milestone, especially in contexts where autonomous agents must adapt to unforeseen environmental variations and evolving objectives. This paper introduces VISOR (VisionSimilarity for Open-ended Robotic exploration), a vision-based framework designed to assist robotic agents in autonomously exploring and learning from new environments and objects, whether through guided or random exploration, without reliance on predefined design considerations. In that direction, VISOR acts as a perception mediator, classifying everything a robot encounters in a scene as either known or unknown. It further identifies potential distractors (e.g., background elements), known categories, or objects specified through text seeds. By leveraging recent advancements in vision foundation models, VISOR operates in a training-free manner. It begins by segmenting a scene into its constituent entities, regardless of familiarity, and then extracts robust visual representations for each one. These representations are compared against an adaptive memory system that evolves over time; unknown objects are assigned unique IDs and added to this memory as new classes, enriching the robot's understanding of its environment. We argue that this evolving memory can facilitate guided exploration through prior knowledge, enhancing the efficiency of robotic exploration, and validate this by designing two exploration scenarios and running both simulated and real-world experiments.
Panagiotis Paraskevas Filntisis, Efthymios Tsaprazlis, Paraskevas Oikonomou, Francesco Mattioli 0003, Vieri G. Santucci, George Retsinas, Petros Maragos
ICRA7
2025 Proactive Tactile Exploration for Object-Agnostic Shape Reconstruction from Minimal Visual Priors
abstract
The perception of an object's surface is important for robotic applications enabling robust object manipulation. The level of accuracy in such a representation affects the outcome of the action planning, especially during tasks that require physical contact, e.g. grasping. In this paper, we propose a novel iterative method for 3D shape reconstruction consisting of two steps. At first, a mesh is fitted on data points acquired from the object's surface, based on a single primitive template. Subsequently, the mesh is properly adjusted to adequately represent local deformities. Moreover, a novel proactive tactile exploration strategy aims at minimizing the total uncertainty with the least number of contacts, while reducing the risk of contact failure in case the estimated surface differs significantly from the real one. The performance of the methodology is evaluated both in 3D simulation and on a real setup.
Paris Oikonomou, George Retsinas, Petros Maragos, Costas S. Tzafestas
ICRA3
2025 Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic Data
abstract
Accurate 6D object pose estimation is essential for robotic grasping and manipulation, particularly in agriculture, where fruits and vegetables exhibit high intra-class variability in shape, size, and texture. The vast majority of existing methods rely on instance-specific CAD models or require depth sensors to resolve geometric ambiguities, making them impractical for real-world agricultural applications. In this work, we introduce PLANTPose, a novel framework for category-level 6D pose estimation that operates purely on RGB input. PLANT-Pose predicts both the 6D pose and deformation parameters relative to a base mesh, allowing a single category-level CAD model to adapt to unseen instances. This enables accurate pose estimation across varying shapes without relying on instance-specific data. To enhance realism and improve generalization, we also leverage Stable Diffusion to refine synthetic training images with realistic texturing, mimicking variations due to ripeness and environmental factors and bridging the domain gap between synthetic data and the real world. Our evaluations on a challenging benchmark that includes bananas of various shapes, sizes, and ripeness status demonstrate the effectiveness of our framework in handling large intraclass variations while maintaining accurate 6D pose predictions, significantly outperforming the state-of-the-art RGB-based approach MegaPose. Our code, data, and models are publicly available at https://github.com/mariosgly/PLANTPose.
Marios Glytsos, Panagiotis Paraskevas Filntisis, George Retsinas, Petros Maragos
IROS4
2025 Instance-Level Composed Image Retrieval
abstract
The progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge—comparable to retrieval among more than 40M random distractors—through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/.
Bill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis, Yannis Avrithis, Petros Maragos, Ondrej Chum, Giorgos Tolias
NeurIPS6
2025 Pre-training for Action Recognition with Automatically Generated Fractal Datasets
Davyd Svyezhentsev, George Retsinas, Petros Maragos
Int. J. Comput. Vis.3
2025 Revisiting Tropical Polynomial Division: Theory, Algorithms, and Application to Neural Networks
abstract
Tropical geometry has recently found several applications in the analysis of neural networks with piecewise linear activation functions. This article presents a new look at the problem of tropical polynomial division and its application to the simplification of neural networks. We analyze tropical polynomials with real coefficients, extending earlier ideas and methods developed for polynomials with integer coefficients. We first prove the existence of a unique quotient-remainder pair and characterize the quotient in terms of the convex bi-conjugate of a related function. Interestingly, the quotient of tropical polynomials with integer coefficients does not necessarily have integer coefficients. Furthermore, we develop a relationship of tropical polynomial division with the computation of the convex hull of unions of convex polyhedra and use it to derive an exact algorithm for tropical polynomial division. An approximate algorithm is also presented, based on an alternation between data partition and linear programming. We also develop special techniques to divide composite polynomials, described as sums or maxima of simpler ones. Finally, we provide numerical results to demonstrate the efficiency of the proposed algorithms, using the MNIST handwritten digits, SVHN, CIFAR-10, and CIFAR-100 datasets, along with an application example in learning model predictive control (LMPC).
Ioannis Kordonis, Petros Maragos
IEEE Trans. Neural Networks Learn. Syst.2
2024 3D Facial Expressions through Analysis-by-Neural-Synthesis
abstract
While existing methods for 3D face reconstruction from in-the-wild images excel at recovering the overall face shape, they commonly miss subtle, extreme, asymmetric, or rarely observed expressions. We improve upon these meth-ods with SMIRK (Spatial Modeling for Image-based Reconstruction of Kinesics), which faithfully reconstructs expres-sive 3D faces from images. We identify two key limitations in existing methods: shortcomings in their self-supervised training formulation, and a lack of expression diversity in the training images. For training, most methods employ differentiable rendering to compare a predicted face mesh with the input image, along with a plethora of additional loss functions. This differentiable rendering loss not only has to provide supervision to optimize for 3D face geom-etry, camera, albedo, and lighting, which is an ill-posed optimization problem, but the domain gap between ren-dering and input image further hinders the learning pro-cess. Instead, SMIRK replaces the differentiable rendering with a neural rendering module that, given the ren-dered predicted mesh geometry, and sparsely sampled pix-els of the input image, generates a face image. As the neural rendering gets color information from sampled im-age pixels, supervising with neural rendering-based reconstruction loss can focus solely on the geometry. Further it enables us to generate images of the input identity with varying expressions while training. These are then utilized as input to the reconstruction model and used as supervision with ground truth geometry. This effectively augments the training data and enhances the generalization for di-verse expressions. Our qualitative, quantitative and partic-ularly our perceptual evaluations demonstrate that SMIRK achieves the new state-of-the art performance on accurate expression reconstruction. For our method's source code, demo video and more, please visit our project webpage: https://georgeretsi.github.io/smirk/.
George Retsinas, Panagiotis Paraskevas Filntisis, Radek Danecek, Victoria Fernández Abrevaya, Anastasios Roussos, Timo Bolkart, Petros Maragos
CVPR7
2024 Augmenting Transformer Autoencoders with Phenotype Classification for Robust Detection of Psychotic Relapses
abstract
Recently, deep autoencoder architectures have received attention for the problem of unsupervised anomaly detection. Detecting psychotic relapses in mental health patients is a crucial challenge, often framed as anomaly detection, given the limited availability of data during relapsing states. In this paper, motivated by the fact that during relapses patients tend to undergo behavioral changes, we augment the classical autoencoder architecture with extra patient identification components. We show that formulating the problem as one of both signal reconstruction and patient identification largely improves the overall precision and robustness of relapse detection and significantly outperforms previous methods with a relative improvement of 15%. In addition, we also explore multiple ways to fuse the identification and reconstruction errors into a unified anomaly score that outperforms the results achieved by each error in isolation.
Niki Efthymiou, George Retsinas, Panagiotis Paraskevas Filntisis, Petros Maragos
ICASSP4
2024 Matrix Factorization in Tropical and Mixed Tropical-Linear Algebras
abstract
Matrix Factorization (MF) has found numerous applications in Machine Learning and Data Mining, including collaborative filtering recommendation systems, dimensionality reduction, data visualization, and community detection. Motivated by the recent successes of tropical algebra and geometry in machine learning, we investigate two problems involving matrix factorization over the tropical algebra. For the first problem, Tropical Matrix Factorization (TMF), which has been studied already in the literature, we propose an improved algorithm that avoids many of the local optima. The second formulation considers the approximate decomposition of a given matrix into the product of three matrices where a usual matrix product is followed by a tropical product. This formulation has a very interesting interpretation in terms of the learning of the utility functions of multiple users. We also present numerical results illustrating the effectiveness of the proposed algorithms, as well as an application to recommendation systems with promising results.
Ioannis Kordonis, Emmanouil Theodosis, George Retsinas, Petros Maragos
ICASSP4
2024 SDPL-SLAM: Introducing Lines in Dynamic Visual SLAM and Multi-Object Tracking
abstract
The need for a robust visual SLAM system operating in real human environments has led to the gradual abandonment of the static world assumption and to the creation of many dynamic SLAM algorithms. Even though there have been many dynamic SLAM proposals, the vast majority of them relied on point features. However, research in static SLAM systems has demonstrated that the use of more complex geometric shapes such as lines can improve performance. Motivated by this we have created a new dynamic SLAM system that estimates the camera poses and the motion of rigid objects, by exploiting both static and dynamic points and lines. Line segments have been incorporated in a novel way in every aspect of our algorithm, by improving their correspondences through optical flow refinement, and by introducing line error terms in both camera and object motion, and in batch optimization. Our proposal has been tested extensively in indoor and outdoor datasets and has achieved significant improvement compared to other state-of-the-art dynamic SLAM systems. Our results demonstrated that line segments enhanced the robustness, thus contributing towards a fully operational SLAM system.Code is publicly available*.
Argyris Manetas, Panagiotis Mermigkas, Petros Maragos
IROS3
2024 OVeNet: Offset Vector Network for Semantic Segmentation
abstract
Semantic segmentation is a fundamental task in visual scene understanding. We focus on the supervised setting, where ground-truth semantic annotations are available. Based on knowledge about the high regularity of real-world scenes, we propose a method for improving class predictions by learning to selectively exploit information from neighboring pixels. In particular, our method is based on the prior that for each pixel, there is a seed pixel in its close neighborhood sharing the same prediction with the former. Motivated by this prior, we design a novel two-head network, named Offset Vector Network (OVeNet), which generates both standard semantic predictions and a dense 2D offset vector field indicating the offset from each pixel to the respective seed pixel, which is used to compute an alternative, seed-based semantic prediction. The two predictions are adaptively fused at each pixel using a learnt dense confidence map for the predicted offset vector field. We supervise offset vectors indirectly via optimizing the seed-based prediction and via a novel loss on the confidence map. Compared to the baseline state-of-the-art architectures HRNet and HRNet+OCR on which OVeNet is built, the latter achieves significant performance gains on three prominent benchmarks for semantic segmentation, namely Cityscapes, ACDC and ADE20K. Code is available at https://github.com/stamatisalex/OVeNet.
Stamatis Alexandropoulos, Christos Sakaridis, Petros Maragos
WACV3
2023 Feather: An Elegant Solution to Effective DNN Sparsification
Athanasios Glentis Georgoulakis, George Retsinas, Petros Maragos
BMVC3
2023 Multimodal Recognition of Valence, Arousal and Dominance via Late-Fusion of Text, Audio and Facial Expressions
abstract
We present an approach for the prediction of valence, arousal, and dominance of people communicating via text/audio/video streams for a translation from and to sign languages.The approach consists of the fusion of the output of three CNN-based models dedicated to the analysis of text, audio, and facial expressions.Our experiments show that any combination of two or three modalities increases prediction performance for valence and arousal.
Fabrizio Nunnari, Annette Rios, Uwe D. Reichel, Chirag Bhuvaneshwara, Panagiotis Paraskevas Filntisis, Petros Maragos, Felix Burkhardt, Florian Eyben, Björn W. Schuller, Sarah Ebling
ESANN6
2023 Relapse Prediction from Long-Term Wearable Data Using Self-Supervised Learning and Survival Analysis
abstract
The introduction of biometric signal analysis in psychiatry could potentially reshape the field by making it more accurate, proactive and personalized. Such biosignals usually acquired from wearables encompass the quantification of human behavior and traits. In this study, we use long-term data acquired from commercial smartwatches, including kinetic and physiological signals, to extract information-thick descriptors that are used for the prediction of subsequent relapses in patients in the psychotic spectrum. Specifically, we propose a novel combination of methods based on Self-Supervised Learning and Survival Analysis that operates on unlabeled and censored data. When combined with other static features that describe the past course of the patient’s health, the proposed methodology yields promising predictive results in terms of two standard survival analysis metrics.
E. Fekas, Athanasia Zlatintsi, Panagiotis Paraskevas Filntisis, Christos Garoufis, Niki Efthymiou, Petros Maragos
ICASSP6
2023 Convolutional Recurrent Neural Networks for the Classification of Cetacean Bioacoustic Patterns
abstract
In this paper we focus on the development of a convolutional recurrent neural network (CRNN) to categorize biosignals collected in the Hellenic Trench, generated by two cetacean species, sperm whales (Physeter macrocephalus) and striped dolphins (Stenella coeruleoalba). We convert audio signals into mel-spectrograms and forward the input into a deep residual network (ResNet), designed to capture spectral patterns. Next, ResNet’s output is reshaped into a time-distributed layer and fed into recurrent network variants, Long Short-Term Memory (LSTMs) or Gated Recurrent Units (GRUs), able to recognize long-term time dependencies on extracted features. The hybrid network perfectly classifies audio signals into three categories (dolphins, sperm whales, ambient noise) while it also exhibits high learning ability on recognising intraclass representations of overlapping acoustic patterns (clicks vs whistles and clicks, both emitted by dolphins). The proposed scheme outperforms traditional Machine Learning (ML) techniques, baseline ResNet and LSTM architectures or their deep parallel combinations.
Dimitris N. Makropoulos, Antigoni Tsiami, Aristides Prospathopoulos, Dimitris Kassis, Alexandros Frantzis, Emmanuel K. Skarsoulis, George Piperakis, Petros Maragos
ICASSP8
2023 Newton-Based Trainable Learning Rate
abstract
Selecting an appropriate learning rate for efficiently training deep neural networks is a difficult process that can be affected by numerous parameters, such as the dataset, the model architecture or even the batch size. In this work, we propose an algorithm for automatically adjusting the learning rate during the training process, assuming a gradient descent formulation. The rationale behind our approach is to train the learning rate along with the model weights. Specifically, we formulate first and second-order gradients w.r.t. the learning rate as functions of consecutive weight gradients, leading to a cost-effective implementation. Our extensive experimental evaluation validates the effectiveness of the proposed method for a plethora of different settings. The proposed method has proven to be robust to both the initial learning rate and the batch size, making it ideal for an off-the-shelf optimizing scheme.
George Retsinas, Giorgos Sfikas, Panagiotis Paraskevas Filntisis, Petros Maragos
ICASSP4
2023 E-Prevention: The ICASSP-2023 Challenge on Person Identification and Relapse Detection from Continuous Recordings of Biosignals
abstract
The e-Prevention challenge concerns the analysis and processing of long-term continuous recordings of biosignals recorded from wearable sensors, i.e., accelerometers, gyroscopes and heart rate monitors embedded in smartwatches, as well as sleep information and daily step count, in order to extract high-level representations of the wearer’s activity and behavior, termed as digital phenotypes. The ability of these digital phenotypes to quantify behavioral patterns and traits will be evaluated in two different tasks: 1) Person Identification, and 2) Relapse Detection in patients in the psychotic spectrum. The long-term data that will be used in this challenge have been acquired during the course of the e-Prevention project, an innovative integrated system for medical support that facilitates effective monitoring and relapse prevention in patients with mental disorders (i.e, schizophrenia and bipolar disorder). Specifically, the data were continuously collected from patients for a monitoring period of up to 2.5 years, while from the control subgroup for a period of 3 months, constituting one of the largest of its kind ever recorded.
Athanasia Zlatintsi, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Christos Garoufis, George Retsinas, Thomas Sounapoglou, Ilias Maglogiannis, Panayiotis Tsanakas, Nikolaos Smyrnis, Petros Maragos
ICASSP10
2023 3D Neural Sculpting (3DNS): Editing Neural Signed Distance Functions
abstract
In recent years, implicit surface representations through neural networks that encode the signed distance have gained popularity and have achieved state-of-the-art results in various tasks (e.g. shape representation, shape reconstruction, and learning shape priors). However, in contrast to conventional shape representations such as polygon meshes, the implicit representations cannot be easily edited and existing works that attempt to address this problem are extremely limited. In this work, we propose the first method for efficient interactive editing of signed distance functions expressed through neural networks, allowing free-form editing. Inspired by 3D sculpting software for meshes, we use a brush-based framework that is intuitive and can in the future be used by sculptors and digital artists. In order to localize the desired surface deformations, we regulate the network by using a copy of it to sample the previously expressed surface. We introduce a novel framework for simulating sculpting-style surface edits, in conjunction with interactive surface sampling and efficient adaptation of network weights. We qualitatively and quantitatively evaluate our method in various different 3D objects and under many different edits. The reported results clearly show that our method yields high accuracy, in terms of achieving the desired edits, while at the same time preserving the geometry outside the interaction areas.
Petros Tzathas, Petros Maragos, Anastasios Roussos
WACV2
2022 Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videos
abstract
In this paper, we introduce a novel deep learning method for photo-realistic manipulation of the emotional state of actors in “in-the-wild” videos. The proposed method is based on a parametric 3D face representation of the actor in the input scene that offers a reliable disentanglement of the facial identity from the head pose and facial expressions. It then uses a novel deep domain translation framework that alters the facial expressions in a consistent and plausible manner, taking into account their dynamics. Finally, the altered facial expressions are used to photo-realistically manipulate the facial region in the input scene based on an especially-designed neural face renderer. To the best of our knowledge, our method is the first to be capable of controlling the actor's facial expressions by even using as a sole input the semantic labels of the manipulated emotions, while at the same time preserving the speech-related lip movements. We conduct extensive qualitative and quantitative evaluations and comparisons, which demonstrate the effectiveness of our approach and the especially promising results that we obtain. Our method opens a plethora of new possibilities for useful applications of neural rendering technologies, ranging from movie post-production and video games to photo-realistic affective avatars.
Foivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, Anastasios Roussos
CVPR3
2022 Enhancing Affective Representations Of Music-Induced Eeg Through Multimodal Supervision And Latent Domain Adaptation
abstract
The study of Music Cognition and neural responses to music has been invaluable in understanding human emotions. Brain signals, though, manifest a highly complex structure that makes processing and retrieving meaningful features challenging, particularly of abstract constructs like affect. Moreover, the performance of learning models is undermined by the limited amount of available neuronal data and their severe inter-subject variability. In this paper we extract efficient, personalized affective representations from EEG signals during music listening. To this end, we employ music signals as a supervisory modality to EEG, aiming to project their semantic correspondence onto a common representation space. We utilize a bi-modal framework by combining an LSTM-based attention model to process EEG and a pre-trained model for music tagging, along with a reverse domain discriminator to align the distributions of the two modalities, further constraining the learning process with emotion tags. The resulting framework can be utilized for emotion recognition both directly, by performing supervised predictions from either modality, and indirectly, by providing relevant music samples to EEG input queries. The experimental findings show the potential of enhancing neuronal data through stimulus information for recognition purposes and yield insights into the distribution and temporal variance of music-induced affective features.
Kleanthis Avramidis, Christos Garoufis, Athanasia Zlatintsi, Petros Maragos
ICASSP4
2022 A Few-Sample Strategy for Guitar Tablature Transcription Based on Inharmonicity Analysis and Playability Constraints
abstract
The prominent strategical approaches regarding the problem of guitar tablature transcription rely either on fingering patterns encoding or on the extraction of string-related audio features. The current work combines the two aforementioned strategies in an explicit manner by employing two discrete components for string-fret classification. It extends older few-sample modeling strategies by introducing various adaptation schemes for the first stage of audio processing, taking advantage of the inharmonic characteristics of guitar sound. Physical limitations and common standards of human performers are incorporated in a genetic algorithm which constitutes a second contextual-based module that further processes the initial audio-based predictions. The proposed methods are evaluated on both annotated guitar performances and isolated note recordings.
Grigoris Bastas, Stefanos Koutoupis, Maximos Kaliakatsos-Papakostas, Vassilis Katsouros, Petros Maragos
ICASSP5
2022 Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language Recognition
abstract
We address the challenging problem of continuous sign language recognition (CSLR) from RGB videos, proposing a novel deep-learning framework that employs spatio-temporal graph convolutional networks (ST-GCNs), which operate on multiple, appropriately fused feature streams, capturing the signer’s pose, shape, appearance, and motion information. In addition to introducing such networks to the continuous recognition problem, our model’s novelty lies on: (i) the feature streams considered and their blending into three ST-GCN modules; (ii) the combination of such modules with bi-directional long short-term memory networks, thus capturing both short-term embedded signing dynamics and long-range feature dependencies; and (iii) the fusion scheme, where the resulting modules operate in parallel, their posteriors aligned via a guiding connectionist temporal classification method, and fused for sign gloss prediction. Notably, concerning (i), in addition to traditional CSLR features, we investigate the utility of 3D human pose and shape parameterization via the "ExPose" approach, as well as 3D skeletal joint information that is regressed from detected 2D joints. We evaluate the proposed system on two well-known CSLR benchmarks, conducting extensive ablations on its modules. We achieve the new state-of-the-art on one of the two datasets, while reaching very competitive performance on the other.
Maria Parelli, Katerina Papadimitriou, Gerasimos Potamianos, Georgios Pavlakos, Petros Maragos
ICASSP5
2022 Neural Network Approximation based on Hausdorff distance of Tropical Zonotopes
Panagiotis Misiakos, Georgios Smyrnis, George Retsinas, Petros Maragos
ICLR4
2022 Child Engagement Estimation in Heterogeneous Child-Robot Interactions Using Spatiotemporal Visual Cues
abstract
Robots are increasingly introduced in various Child-Robot Interactions with educational, entertainment or even therapeutic goals. In order to achieve qualitative inter-actions, robots need to adjust their behavior according to children's response. A robot's ability to successfully estimate partner's engagement is of great importance towards this direction. In this research we propose a method to estimate the engagement level of children during heterogeneous and challenging child-robot interactions. Our method uses the spatiotemporal residual$\mathrm{R}(2+1)\mathrm{D}$blocks to simultaneously leverage the rich RGB and temporal information, which is crucial for the engagement estimation. We present results on three different groups of data, including the PInSoRo open dataset, proving our method's robustness and improvement over previous works.
Dafni Anagnostopoulou, Niki Efthymiou, Christina Papailiou, Petros Maragos
IROS4
2021 From Seq2Seq Recognition to Handwritten Word Embeddings
George Retsinas, Giorgos Sfikas, Christophoros Nikou, Petros Maragos
BMVC4
2021 Exploiting Emotional Dependencies with Graph Convolutional Networks for Facial Expression Recognition
abstract
Over the past few years, deep learning methods have shown remarkable results in many face-related tasks including automatic facial expression recognition (FER) in-the-wild. Meanwhile, numerous models describing the human emotional states have been proposed by the psychology community. However, we have no clear evidence as to which representation is more appropriate and the majority of FER systems use either the categorical or the dimensional model of affect. Inspired by recent work in multi-label classification, this paper proposes a novel multi-task learning (MTL) framework that exploits the dependencies between these two models using a Graph Convolutional Network (GCN) to recognize facial expressions in-the-wild. Specifically, a shared feature representation is learned for both discrete and continuous recognition in a MTL setting. Moreover, the facial expression classifiers and the valence-arousal regressors are learned through a GCN that explicitly captures the dependencies between them. To evaluate the performance of our method under real-world conditions we perform extensive experiments on the AffectNet and Aff-Wild2 datasets. The results of our experiments show that our method is capable of improving the performance across different datasets and backbone architectures. Finally, we also surpass the previous state-of-the-art methods on the categorical model of AffectNet.
Panagiotis Antoniadis, Panagiotis Paraskevas Filntisis, Petros Maragos
FG3
2021 Leveraging Semantic Scene Characteristics and Multi-Stream Convolutional Architectures in a Contextual Approach for Video-Based Visual Emotion Recognition in the Wild
abstract
In this work we tackle the task of video-based visual emotion recognition in the wild. Standard methodologies that rely solely on the extraction of bodily and facial features often fall short of accurate emotion prediction in cases where the aforementioned sources of affective information are inaccessible due to head/body orientation, low resolution and poor illumination. We aspire to alleviate this problem by leveraging visual context in the form of scene characteristics and attributes, as part of a broader emotion recognition framework. Temporal Segment Networks (TSN) constitute the backbone of our proposed model. Apart from the RGB input modality, we make use of dense Optical Flow, following an intuitive multi-stream approach for a more effective encoding of motion. Furthermore, we shift our attention towards skeleton-based learning and leverage action-centric data as means of pretraining a Spatial-Temporal Graph Convolutional Network (ST-GCN) for the task of emotion recognition. Our extensive experiments on the challenging Body Language Dataset (BoLD) verify the superiority of our methods over existing approaches, while by properly incorporating all of the aforementioned modules in a network ensemble, we manage to surpass the previous best published recognition scores, by a large margin.
Ioannis Pikoulis, Panagiotis Paraskevas Filntisis, Petros Maragos
FG3
2021 Deep Convolutional and Recurrent Networks for Polyphonic Instrument Classification from Monophonic Raw Audio Waveforms
abstract
Sound Event Detection and Audio Classification tasks are traditionally addressed through time-frequency representations of audio signals such as spectrograms. However, the emergence of deep neural networks as efficient feature extractors has enabled the direct use of audio signals for classification purposes. In this paper, we attempt to recognize musical instruments in polyphonic audio by only feeding their raw waveforms into deep learning models. Various recurrent and convolutional architectures incorporating residual connections are examined and parameterized in order to build end-to-end classifiers with low computational cost and only minimal preprocessing. We obtain competitive classification scores and useful instrument-wise insight through the IRMAS test set, utilizing a parallel CNN-BiGRU model with multiple residual connections, while maintaining a significantly reduced number of trainable parameters.
Kleanthis Avramidis, Agelos Kratimenos, Christos Garoufis, Athanasia Zlatintsi, Petros Maragos
ICASSP5
2021 Advances in Morphological Neural Networks: Training, Pruning and Enforcing Shape Constraints
abstract
In this paper, we study an emerging class of neural networks, the Morphological Neural networks, from some modern perspectives. Our approach utilizes ideas from tropical geometry and mathematical morphology. First, we state the training of a binary morphological classifier as a Difference-of-Convex optimization problem and extend this method to multiclass tasks. We then focus on general morphological networks trained with gradient descent variants and show, quantitatively via pruning schemes as well as qualitatively, the sparsity of the resulted representations compared to FeedForward networks with ReLU activations as well as the effect the training optimizer has on such compression techniques. Finally, we show how morphological networks can be employed to guarantee monotonicity and present a softened version of a known architecture, based on Maslov Dequantization, which alleviates issues of gradient propagation associated with its "hard" counterparts and moderately improves performance.
Nikolaos Dimitriadis, Petros Maragos
ICASSP2
2021 Independent Sign Language Recognition with 3d Body, Hands, and Face Reconstruction
abstract
Independent Sign Language Recognition is a complex visual recognition problem that combines several challenging tasks of Computer Vision due to the necessity to exploit and fuse information from hand gestures, body features and facial expressions. While many state-of-the-art works have managed to deeply elaborate on these features independently, to the best of our knowledge, no work has adequately combined all three information channels to efficiently recognize Sign Language. In this work, we employ SMPL-X, a contemporary parametric model that enables joint extraction of 3D body shape, face and hands information from a single image. We use this holistic 3D reconstruction for SLR, demonstrating that it leads to higher accuracy than recognition from raw RGB images and their optical flow fed into the state-of-the-art I3D-type network for 3D action recognition and from 2D Openpose skeletons fed into a Recurrent Neural Network. Finally, a set of experiments on the body, face and hand features showed that neglecting any of these, significantly reduces the classification accuracy, proving the importance of jointly modeling body shape, facial expression and hand pose for Sign Language Recognition.
Agelos Kratimenos, Georgios Pavlakos, Petros Maragos
ICASSP3
2021 Sparsity in Max-Plus Algebra and Applications in Multivariate Convex Regression
abstract
In this paper, we study concepts of sparsity in the max-plus algebra and apply them to the problem of multivariate convex regression. We show how to efficiently find sparse (containing many −∞ elements) approximate solutions to max-plus equations by leveraging notions from submodular optimization. Subsequently, we propose a novel method for piecewise-linear surface fitting of convex multivariate functions, with optimality guarantees for the model parameters and an approximately minimum number of affine regions.
Nikos Tsilivis 0001, Anastasios Tsiamis, Petros Maragos
ICASSP3
2021 Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection
abstract
Scene Graph Generators (SGGs) are models that, given an image, build a directed graph where each edge represents a predicted subject predicate object triplet. Most SGGs silently exploit datasets' bias on relationships' context, i.e. its subject and object, to improve recall and neglect spatial and visual evidence, e.g. having seen a glut of data for person wearing shirt, they are overconfident that every person is wearing every shirt. Such imprecise predictions are mainly ascribed to the lack of negative examples for most relationships, which obstructs models from meaningfully learning predicates, even those that have ample positive examples. We first present an indepth investigation of the context bias issue to showcase that all examined state-of-the-art SGGs share the above vulnerabilities. In response, we propose a semi-supervised scheme that forces predicted triplets to be grounded consistently back to the image, in a closed-loop manner. The developed spatial common sense can be then distilled to a student SGG and substantially enhance its spatial reasoning ability. This Grounding Consistency Distillation (GCD) approach is model-agnostic and benefits from the superfluous unlabeled samples to retain the valuable context information and avert memorization of annotations. Furthermore, we demonstrate that current metrics disregard unlabeled samples, rendering themselves incapable of reflecting context bias, then we mine and incorporate during evaluation hard-negatives to reformulate precision as a reliable metric. Extensive experimental comparisons exhibit large quantitative - up to 70% relative precision boost on VG200 dataset - and qualitative improvements to prove the significance of our GCD method and our metrics towards refocusing graph generation as a core aspect of scene understanding. Code available at https://github.com/deeplab-ai/grounding-consistent-vrd.
Markos Diomataris, Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Maragos
ICCV4
2021 Online Weight Pruning Via Adaptive Sparsity Loss
abstract
Pruning neural networks has regained interest in recent years as a means to compress state-of-the-art deep neural networks and enable their deployment on resource-constrained devices. In this paper, we propose a robust sparsity controlling framework that efficiently prunes network parameters during training with minimal computational overhead. We incorporate fast mechanisms to prune individual layers and build upon these to automatically prune the entire network under a user-defined budget constraint. Key to our end-to-end network pruning approach is the formulation of an intuitive and easy-to-implement adaptive sparsity loss used to explicitly control sparsity during training, enabling efficient budget-aware optimization.
George Retsinas, Athena Elafrou, Georgios I. Goumas, Petros Maragos
ICIP4
2021 Deformation-Invariant Networks For Handwritten Text Recognition
abstract
Image deformations under simple geometric restrictions are crucial for Handwriting Text Recognition (HTR), since different writing styles can be viewed as simple geometrical deformations of the same textual elements. In this respect, the usefulness of including deformation invariance to an HTR system is indisputable. We explore different existing strategies for ensuring deformation invariance, including spatial transformers and deformable convolutions, under the context of text recognition, as well as introduce a new deformation-based algorithm, inspired by adversarial learning, which aims to reduce character output uncertainty during evaluation time. The resulting HTR system is shown to achieve state-of-the-art performance on the IAM and RIMES datasets.
George Retsinas, Giorgos Sfikas, Christophoros Nikou, Petros Maragos
ICIP4
2021 Engagement Estimation During Child Robot Interaction Using Deep Convolutional Networks Focusing on ASD Children
abstract
Estimating the engagement of children is an essential prerequisite for constructing natural Child-Robot Interaction. Especially in the case of children with Autism Spectrum Disorder, monitoring the engagement of the other party allows robots to adjust their actions according to the educational and therapeutic goals in hand. In this work we delve into engagement estimation with a focus on children with autism spectrum disorder. We propose deep convolutional architectures for engagement estimation that outperform previous methods, and explore their performance under variable conditions, in four databases depicting ASD and TD children interacting with robots or humans.
Dafni Anagnostopoulou, Niki Efthymiou, Christina Papailiou, Petros Maragos
ICRA4
2021 A linear method for camera pair self-calibration
Nikos Melanitis, Petros Maragos
Comput. Vis. Image Underst.2
2021 Tropical Geometry and Machine Learning
abstract
Tropical geometry is a relatively recent field in mathematics and computer science, combining elements of algebraic geometry and polyhedral geometry. The scalar arithmetic of its analytic part preexisted in the form of max-plus and min-plus semiring arithmetic used in finite automata, nonlinear image processing, convex analysis, nonlinear control, optimization, and idempotent mathematics. Tropical geometry recently emerged in the analysis and extension of several classes of problems and systems in both classical machine learning and deep learning. Three such areas include: 1) deep neural networks with piecewise linear (PWL) activation functions; 2) probabilistic graphical models; and 3) nonlinear regression with PWL functions. In this article, we first summarize introductory ideas and objects of tropical geometry, providing a theoretical framework for both the max-plus algebra that underlies tropical geometry and its extensions to general max algebras. This unifies scalar and vector/signal operations over a class of nonlinear spaces, called weighted lattices, and allows us to provide optimal solutions for algebraic equations used in tropical geometry and generalize tropical geometric objects. Then, we survey the state of the art and recent progress in the aforementioned areas. First, we illustrate a purely geometric approach for studying the representation power of neural networks with PWL activations. Then, we review the tropical geometric analysis of parametric statistical models, such as HMMs; later, we focus on the Viterbi algorithm and related methods for weighted finite-state transducers and provide compact and elegant representations via their formal tropical modeling. Finally, we provide optimal solutions and an efficient algorithm for the convex regression problem, using concepts and tools from tropical geometry and max-plus algebra. Throughout this article, we also outline problems and future directions in machine learning that can benefit from the tropical-geometric point of view.
Petros Maragos, Vasileios Charisopoulos, Emmanouil Theodosis
Proc. IEEE1
2020 From Saturation to Zero-Shot Visual Relationship Detection Using Local Context
Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Maragos
BMVC3
2020 STAViS: Spatio-Temporal AudioVisual Saliency Network
abstract
We introduce STAViS, a spatio-temporal audiovisual saliency network that combines spatio-temporal visual and auditory information in order to efficiently address the problem of saliency estimation in videos. Our approach employs a single network that combines visual saliency and auditory features and learns to appropriately localize sound sources and to fuse the two saliencies in order to obtain a final saliency map. The network has been designed, trained end-to-end, and evaluated on six different databases that contain audiovisual eye-tracking data of a large variety of videos. We compare our method against 8 different state-of-the-art visual saliency models. Evaluation results across databases indicate that our STAViS model outperforms our visual only variant as well as the other state-of-the-art models in the majority of cases. Also, the consistently good performance it achieves for all databases indicates that it is appropriate for estimating saliency "in-the-wild". The code is available at https://github.com/atsiami/STAViS.
Antigoni Tsiami, Petros Koutras, Petros Maragos
CVPR3
2020 An LSTM-Based Dynamic Chord Progression Generation System for Interactive Music Performance
abstract
In this paper, we describe an interactive generative music system, designed to handle polyphonic guitar music. We formulate the problem of chord progression generation as a prediction problem. Thus, we propose utilization of an LSTM-based network architecture incorporating neural attention that is able to learn a mapping between symbolic representations of polyphonic chord progressions and future chord candidates. Furthermore, we have developed a virtual air-guitar controller, utilizing a Kinect device, that uses the above architecture in order to change in real time the guitar chord mapping, depending on the performer's previous performance. The whole system was evaluated both objectively and subjectively. The goal of the objective evaluation was to measure the ability of the system to correctly generate chord candidates for existing chord progressions, as well as identify the type of errors. The subjective evaluation mainly focused on the longer-term behavior of the system, regarding the musical coherence and the variety of the generated progressions. The results were encouraging regarding the ability of our system to generate sound chord progressions, while highlighting a number of issues that require to be resolved.
Christos Garoufis, Athanasia Zlatintsi, Petros Maragos
ICASSP3
2020 Multivariate Tropical Regression and Piecewise-Linear Surface Fitting
abstract
In this paper we propose a novel approach for multivariate convex regression by using as approximation model a maximum of hyperplanes, which we represent as a multivariate max-plus tropical polynomial. Our approach uses concepts from tropical geometry and finds an optimal solution for the model parameters (that minimizes a data fitting error norm) by solving systems of max-plus equations using max-plus algebra and projections on weighted lattices. Our method has lower complexity than most other methods for fitting piecewise-linear (PWL) functions and we apply it to optimal PWL regression for fitting max-plus tropical surfaces to arbitrary data that constitute polyhedral shape approximations.
Petros Maragos, Emmanouil Theodosis
ICASSP1
2020 Person Identification Using Deep Convolutional Neural Networks on Short-Term Signals from Wearable Sensors
abstract
In this work, we explore the discriminating ability of short-term signal patterns (e.g. few minutes long) with respect to the person identification task. We focus on signals recorded by simple wearable devices, such as smart watches, which can measure movements (accelerometer and gyroscope sensors) and biosignals (heart rate monitor). To address the person identification problem, we develop a deep neural network, based on one-dimensional convolutions, which receives raw signals from three different smartwatch sensors and predicts the person wearing the smartwatch. Experimental results indicate that even with signals from wearable sensors collected at intervals of only 10 minutes, different users can be identified with notably high accuracy, revealing the existence of distinct short-term patterns of movement and heart rate between different persons.
George Retsinas, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Emmanouil Theodosis, Athanasia Zlatintsi, Petros Maragos
ICASSP6
2020 Maxpolynomial Division with Application To Neural Network Simplification
abstract
In this work, we further the link between neural networks with piecewise linear activations and tropical algebra. To that end, we introduce the process of Maxpolynomial Division, a geometric method which simulates division of polynomials in the max-plus semiring, while highlighting its key properties and noting its connection to neural networks. Afterwards, we generalize this method and apply it in the context of neural network minimization, for two-layer networks used for binary classification problems, attempting to reduce the size of the hidden layer before the output. A tractable method to find an appropriate divisor and perform the division is introduced and evaluated in the IMDB Movie Review and MNIST datasets, with preliminary experiments demonstrating a capacity of this method to reduce the size of the network, without major loss of performance.
Georgios Smyrnis, Petros Maragos, George Retsinas
ICASSP2
2020 Multiclass Neural Network Minimization via Tropical Newton Polytope Approximation
abstract
The field of tropical algebra is closely linked with the domain of neural networks with piecewise linear activations, since their output can be described via tropical polynomials in the max-plus semiring. In this work, we attempt to make use of methods stemming from a form of approximate division of such polynomials, which relies on the approximation of their Newton Polytopes, in order to minimize networks trained for multiclass classification problems. We make theoretical contributions in this domain, by proposing and analyzing methods which seek to reduce the size of such networks. In addition, we make experimental evaluations on the MNIST and Fashion-MNIST datasets, with our results demonstrating a significant reduction in network size, while retaining adequate performance.
Georgios Smyrnis, Petros Maragos
ICML2
2020 Enhancing Handwritten Text Recognition with N-gram sequence decomposition and Multitask Learning
abstract
Current state-of-the-art approaches in the field of Handwritten Text Recognition are predominately single task with unigram, character level target units. In our work, we utilize a Multi-task Learning scheme, training the model to perform decompositions of the target sequence with target units of different granularity, from fine to coarse. We consider this method as a way to utilize n-gram information, implicitly, in the training process, while the final recognition is performed using only the unigram output. Unigram decoding of such a multi-task approach highlights the capability of the learned internal representations, imposed by the different n-grams at the training step. We select n-grams as our target units and we experiment from unigrams till fourgrams, namely subword level granularities. These multiple decompositions are learned from the network with task-specific CTC losses. Concerning network architectures, we propose two alternatives, namely the Hierarchical and the Block Multi-task. Overall, our proposed model, even though evaluated only on the unigram task, outperforms its counterpart single-task by absolute 2.52% WER and 1.02% CER, in the greedy decoding, without any computational overhead during inference, hinting towards successfully imposing an implicit language model.
Vasiliki Tassopoulou, George Retsinas, Petros Maragos
ICPR3
2019 RecNets: Channel-wise Recurrent Convolutional Neural Networks
George Retsinas, Athena Elafrou, Georgios I. Goumas, Petros Maragos
BMVC4
2019 Tropical Modeling of Weighted Transducer Algorithms on Graphs
abstract
Weighted Finite State Transducers (WFSTs) are versatile graphical automata that can model a great number of problems, ranging from automatic speech recognition to DNA sequencing. Traditional computer science algorithms are employed when working with these automata in order to optimize their size, but also the run time of decoding algorithms. However, these algorithms are not unified under a common framework that would allow for their treatment as a whole. Moreover, the inherent geometrical representation of WFSTs, coupled with the topology-preserving algorithms that operate on them make the structures ideal for tropical analysis. The benefits of such analysis have a twofold nature; first, matrix operations offer a connection to nonlinear vector space and spectral theory, and, second, tropical algebra offers a connection to tropical geometry. In this work we model some of the most frequently used algorithms in WFSTs by using tropical algebra; this provides a theoretical unification and allows us to also analyze aspects of their tropical geometry.
Emmanouil Theodosis, Petros Maragos
ICASSP2
2019 Deeply Supervised Multimodal Attentional Translation Embeddings for Visual Relationship Detection
abstract
Detecting visual relationships, i.e.triplets, has been a challenging Scene Understanding task approached in the past via linguistic priors or spatial information in a single feature branch. We introduce a new deeply supervised two-branch architecture, the Multimodal Attentional Translation Embeddings, where the visual features of each branch are driven by a multimodal attentional mechanism that exploits spatio-linguistic similarities in a low-dimensional space. We present a variety of experiments comparing against all related approaches in the literature, as well as by re-implementing and fine-tuning several of them. Results on the commonly employed VRD dataset [1] show that the proposed method clearly outperforms all others, while we also justify our claims both quantitatively and qualitatively.
Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras, Athanasia Zlatintsi, Petros Maragos
ICIP5
2019 Showcasing Deeply Supervised Multimodal Attentional Translation Embeddings: a Demo for Visual Relationship Detection
abstract
We address the task of Visual Relationship Detection, i.e. the detection oftriplets in an image, introducing Multimodal Attentional Translation Embeddings (ICIP 2019 paper, id 3642). Motivated by the need of visualization and interpretation of the results, as well as the lack of other tools for online predictions on this task, we design and implement the first architecture for live inference of visual relationships on video streams and wild images, including research and engineering extensions, ablation models and a lightweight CPU-version. The code is available at https://bitbucket.org/deeplabai/vrd.
Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras, Athanasia Zlatintsi, Petros Maragos
ICIP5
2019 Video Processing and Learning in Assistive Robotic Applications
abstract
The integration of visual perception to robotic systems is a key research area in recent years. Advances in modern computer vision techniques along with the development of faster and more accurate visual sensors led to the emergence of new methods for robotic visual perception [1]. One area of research is the development of robotic assistive vision for Human-Robot Interaction (HRI) systems [2], [3]. The rapid increase of people with special needs, such as the elderly population, and the simultaneous reduction of personal care staff, reinforce the need for robotic assistants [4], [5]. There are many challenges in this area including the familiarity of these users with new technologies and the domain specific datasets, which are required for training user oriented models. Nowadays, modern assistive and social human-robot interaction requires the multimodal communication with speech, gestures and human movements so as to enhance the classic interaction with only spoken commands.
Petros Koutras, Georgia Chalvatzaki, Antigoni Tsiami, Alexandros Nikolakakis, Costas S. Tzafestas, Petros Maragos
ICIP6
2019 LSTM-based Network for Human Gait Stability Prediction in an Intelligent Robotic Rollator
abstract
In this work, we present a novel framework for on-line human gait stability prediction of the elderly users of an intelligent robotic rollator using Long Short Term Memory (LSTM) networks, fusing multimodal RGB-D and Laser Range Finder (LRF) data from non-wearable sensors. A Deep Learning (DL) based approach is used for the upper body pose estimation. The detected pose is used for estimating the body Center of Mass (CoM) using Unscented Kalman Filter (UKF). An Augmented Gait State Estimation framework exploits the LRF data to estimate the legs' positions and the respective gait phase. These estimates are the inputs of an encoder-decoder sequence to sequence model which predicts the gait stability state as Safe or Fall Risk walking. It is validated with data from real patients, by exploring different network architectures, hyperparameter settings and by comparing the proposed method with other baselines. The presented LSTM-based human gait stability predictor is shown to provide robust predictions of the human stability state, and thus has the potential to be integrated into a general user-adaptive control architecture as a fall-risk alarm.
Georgia Chalvatzaki, Petros Koutras, Jack Hadfield, Xanthi S. Papageorgiou, Costas S. Tzafestas, Petros Maragos
ICRA6
2019 A Deep Learning Approach for Multi-View Engagement Estimation of Children in a Child-Robot Joint Attention Task
abstract
In this work, we tackle the problem of child engagement estimation while children freely interact with a robot in a friendly, room-like environment. We propose a deep learning-based multi-view solution that takes advantage of recent developments in human pose detection. We extract the child's pose from different RGB-D cameras placed regularly in the room, fuse the results and feed them to a deep Neural Network (NN) trained for classifying engagement levels. The deep network contains a recurrent layer, in order to exploit the rich temporal information contained in the pose data. The resulting method outperforms a number of baseline classifiers and provides a promising tool for better automatic understanding of a child's attitude, interest and attention while cooperating with a robot. The goal is to integrate this model in next-generation social robots as an attention monitoring tool during various Child Robot Interaction (CRI) tasks both for Typically Developed (TD) children and children affected by autism (ASD).
Jack Hadfield, Georgia Chalvatzaki, Petros Koutras, Mehdi Khamassi, Costas S. Tzafestas, Petros Maragos
IROS6
2019 A behaviorally inspired fusion approach for computational audiovisual saliency modeling
Antigoni Tsiami, Petros Koutras, Athanasios Katsamanis, Argiro Vatakis, Petros Maragos
Signal Process. Image Commun.5
2018 Multimodal Visual Concept Learning With Weakly Supervised Techniques
abstract
Despite the availability of a huge amount of video data accompanied by descriptive texts, it is not always easy to exploit the information contained in natural language in order to automatically recognize video concepts. Towards this goal, in this paper we use textual cues as means of supervision, introducing two weakly supervised techniques that extend the Multiple Instance Learning (MIL) framework: the Fuzzy Sets Multiple Instance Learning (FSMIL) and the Probabilistic Labels Multiple Instance Learning (PLMIL). The former encodes the spatio-temporal imprecision of the linguistic descriptions with Fuzzy Sets, while the latter models different interpretations of each description's semantics with Probabilistic Labels, both formulated through a convex optimization algorithm. In addition, we provide a novel technique to extract weak labels in the presence of complex semantics, that consists of semantic similarity computations. We evaluate our methods on two distinct problems, namely face and action recognition, in the challenging and realistic setting of movies accompanied by their screenplays, contained in the COGNIMUSE database. We show that, on both tasks, our method considerably outperforms a state-of-the-art weakly supervised approach, as well as other baselines.
Giorgos Bouritsas, Petros Koutras, Athanasia Zlatintsi, Petros Maragos
CVPR4
2018 Multi-View Audio-Articulatory Features for Phonetic Recognition on RTMRI-TIMIT Database
abstract
In this paper, we investigate the use of articulatory information, and more specifically real time Magnetic Resonance Imaging (rtMRI) data of the vocal tract, to improve speech recognition performance. For the purpose of our experiments, we use data from the rtMRI-TIMIT database. Firstly, Scale Invariant Feature Transform (SIFT) features are extracted for each video frame. Afterwards, the SIFT descriptors of each frame are transformed to a single histogram per picture, by using the Bag of Visual Words methodology. Since this kind of articulatory information is difficult to acquire in typical speech recognition setups we only consider it to be available in the training phase. Thus, we use a multi-view setup approach by applying Canonical Correlation Analysis (CCA) to visual and audio data. By using the transformation matrix, acquired during the training stage, we transform both train and test audio data to produce MFCC-articulatory features, which form the input for the recognition system. Experimental results demonstrate improvements in phone recognition in comparison with the audio-based baseline.
Ioannis K. Douros, Athanasios Katsamanis, Petros Maragos
ICASSP3
2018 Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and Adults
abstract
Human-robot interaction (HRI) is a research area of growing interest with a multitude of applications for both children and adult user groups, as, for example, in edutainment and social robotics. Crucial, however, to its wider adoption remains the robust perception of HRI scenes in natural, untethered, and multi-party interaction scenarios, across user groups. Towards this goal, we investigate three focal HRI perception modules operating on data from multiple audio-visual sensors that observe the HRI scene from the far-field, thus bypassing limitations and platform-dependency of contemporary robotic sensing. In particular, the developed modules fuse intra- and/or inter-modality data streams to perform: (i) audio-visual speaker localization; (ii) distant speech recognition; and (iii) visual recognition of hand-gestures. Emphasis is also placed on ensuring high speech and gesture recognition rates for both children and adults. Development and objective evaluation of the three modules is conducted on a corpus of both user groups, collected by our far-field multisensory setup, for an interaction scenario of a question-answering “guess-the-object” collaborative HRI game with a “Furhat” robot. In addition, evaluation of the game incorporating the three developed modules is reported. Our results demonstrate robust far-field audio-visual perception of the multi-party HRI scene.
Antigoni Tsiami, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Petros Koutras, Gerasimos Potamianos, Petros Maragos
ICASSP6
2018 Multimodal Signal Processing and Learning Aspects of Human-Robot Interaction for an Assistive Bathing Robot
abstract
We explore new aspects of assistive living on smart human-robot interaction (HRI) that involve automatic recognition and online validation of speech and gestures in a natural interface, providing social features for HRI. We introduce a whole framework and resources of a real-life scenario for elderly subjects supported by an assistive bathing robot, addressing health and hygiene care issues. We contribute a new dataset and a suite of tools used for data acquisition and a state-of-the-art pipeline for multimodal learning within the framework of the I -Support bathing robot, with emphasis on audio and RGB- D visual streams. We consider privacy issues by evaluating the depth visual stream along with the RGB, using Kinect sensors. The audio-gestural recognition task on this new dataset yields up to 84.5%, while the online validation of the I-Support system on elderly users accomplishes up to 84% when the two modalities are fused together. The results are promising enough to support further research in the area of multimodal recognition for assistive social HRI, considering the difficulties of the specific task.
Athanasia Zlatintsi, Isidoros Rodomagoulakis, Petros Koutras, Athanasios Dometios, Vassilis Pitsikalis, Costas S. Tzafestas, Petros Maragos
ICASSP7
2018 Multi- View Fusion for Action Recognition in Child-Robot Interaction
abstract
Answering the challenge of leveraging computer vision methods in order to enhance Human Robot Interaction (HRI) experience, this work explores methods that can expand the capabilities of an action recognition system in such tasks. A multi-view action recognition system is proposed for integration in HRI scenarios with special users, such as children, in which there is limited data for training and many state-of-the-art techniques face difficulties. Different feature extraction approaches, encoding methods and fusion techniques are combined and tested in order to create an efficient system that recognizes children pantomime actions. This effort culminates in the integration of a robotic platform and is evaluated under an alluring Children Robot Interaction scenario.
Niki Efthymiou, Petros Koutras, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos
ICIP5
2018 Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots
abstract
Child-robot interaction is an interdisciplinary research area that has been attracting growing interest, primarily focusing on edutainment applications. A crucial factor to the successful deployment and wide adoption of such applications remains the robust perception of the child's multi-modal actions, when interacting with the robot in a natural and untethered fashion. Since robotic sensory and perception capabilities are platform-dependent and most often rather limited, we propose a multiple Kinect-based system to perceive the child-robot interaction scene that is robot-independent and suitable for indoors interaction scenarios. The audio-visual input from the Kinect sensors is fed into speech, gesture, and action recognition modules, appropriately developed in this paper to address the challenging nature of child-robot interaction. For this purpose, data from multiple children are collected and used for module training or adaptation. Further, information from the multiple sensors is fused to enhance module performance. The perception system is integrated in a modular multi-robot architecture demonstrating its flexibility and scalability with different robotic platforms. The whole system, called Multi3, is evaluated, both objectively at the module level and subjectively in its entirety, under appropriate child-robot interaction scenarios containing several carefully designed games between children and robots.
Antigoni Tsiami, Petros Koutras, Niki Efthymiou, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos
ICRA6
2018 User-Adaptive Human-Robot Formation Control for an Intelligent Robotic Walker Using Augmented Human State Estimation and Pathological Gait Characterization
abstract
In this paper we describe a control strategy for a user-adaptive human-robot system for an intelligent robotic Mobility Assistive Device (MAD)using raw data from a single laser-range-finder (LRF)mounted on the MAD and scanning the walking area. The proposed control architecture consists of three modules. In the first module, a previously proposed methodology (termed IMM-PDA-PF)delivers the augmented human state estimation of the user by providing robust leg tracking and on-line estimation of the human gait phases. This information is processed at the next module for providing the pathological gait parametrization and characterization, by computing specific gait parameters for each gait cycle. These gait parameters form the feature vector that classifies the user in a certain class related to risk of fall. Those are of particular significance to the system, since the gait parameters and the respective class are used in the third module, i.e. the human-robot formation controller, in order to adapt the desired formation of the human-robot system, by selecting the appropriate control variables. The experimental evaluation comprises gait data from real patients, and demonstrates the stability of the human-robot formation control, indicating the importance of incorporating an on-line gait characterization of the user, using non-wearable and non-invasive methods, in the context of a robotic MAD.
Georgia Chalvatzaki, Xanthi S. Papageorgiou, Petros Maragos, Costas S. Tzafestas
IROS3
2018 Object Assembly Guidance in Child-Robot Interaction using RGB-D based 3D Tracking
abstract
This work examines how and to what benefit an autonomous humanoid robot can supervise a child in an object assembly task. In order to understand the child's actions, a novel 3D object tracking algorithm for RGB-D data is employed. The tracker consists of two stages: the first performs a tracking-by-detection scheme on the color stream, to locate the objects on the image plane, while the second uses a particle filter that operates on the depth data stream to refine the first stage output and infer the objects' rotations. Given the six degrees-of-freedom of the assembly part poses, the system is able to recognize which connections have been completed at any given time. This information is then used to select an appropriate verbal or gestural response for the robot. Experimental results show that (a) the tracking algorithm is accurate, fast and robust to severe occlusions and fast movements, (b) the proposed method of assembly state estimation is indeed effective, and (c) the resulting Child-Robot Interaction scenario is educational and enjoyable for the children involved.
Jack Hadfield, Petros Koutras, Niki Efthymiou, Gerasimos Potamianos, Costas S. Tzafestas, Petros Maragos
IROS6
2017 Photorealistic adaptation and interpolation of facial expressions using HMMS and AAMS for audio-visual speech synthesis
abstract
In this paper, motivated by the continuously increasing presence of intelligent agents in everyday life, we address the problem of expressive photorealistic audio-visual speech synthesis, with a strong focus on the visual modality. Emotion constitutes one of the main driving factors of social life and it is expressed mainly through facial expressions. Synthesis of a talking head capable of expressive audio-visual speech is challenging due to the data overhead that arises when considering the vast number of emotions we would like the talking head to express. In order to tackle this challenge, we propose the usage of two methods, namely Hidden Markov Model (HMM) adaptation and interpolation, with HMMs modeling visual parameters via an Active Appearance Model (AAM) of the face. We show that through HMM adaptation we can successfully adapt a “neutral” talking head to a target emotion with a small amount of adaptation data, as well as that through HMM interpolation we can robustly achieve different levels of intensity for an emotion.
Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Petros Maragos
ICIP3
2017 Demonstration of an HMM-based photorealistic expressive audio-visual speech synthesis system
abstract
Summary form only given. The usage of conversational agents is rapidly increasing in everyday life (cortana, siri, etc.). It has been shown that the inclusion of a talking face, increases the intelligibility of speech and the naturalness of human-computer interaction. Furthermore, an agent capable of expressing emotions has a stronger appeal to the human party and affects the interlocutor's emotional state. The proposed demonstration is a Hidden Markov Model (HMM) based photorealistic audio-visual speech synthesis system, capable of expressing emotions [1, 2]. The system is capable of generating a talking head speaking in three emotions: happiness, anger, and sadness, plus in neutral speaking style. Further capabilities of the system include 1) the usage of HMM interpolation [3] in order to generate speech with mixtures of the original emotions (e.g., both anger and happiness), and speech with different levels of expressiveness (by mixing with the neutral emotion), 2) the usage of HMM adaptation [4], in order to adapt to a target emotion using only a few number of sentences. Equipment In order to showcase our system we will use a laptop and speakers. The system will run fully on the laptop. Demonstration Experience During the demonstration, viewers will have the opportunity to: 1. Watch videos of the talking head speaking in 3 different emotions (plus neutral) and see how the expressive talking head feels more natural compared to the talking head speaking in neutral style. 2. Watch the talking head speaking in two or more emotions at the same time, and see how the weights assigned to each emotion affects the outcome. It will also be of great interest to see which emotion each viewer perceives. In addition, through interpolation with the neutral emotion, viewers will be able to watch the talking head speak in different expressiveness levels for each emotion. 3. See how the neutral talking head can be adapted to speak in another emotion using only a few sentences, and how the number of sentences used affects the expressiveness of the resulting talking head.
Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Petros Maragos
ICIP3
2017 Comparative experimental validation of human gait tracking algorithms for an intelligent robotic rollator
abstract
Tracking human gait accurately and robustly constitutes a key factor for a smart robotic walker, aiming to provide assistance to patients with different mobility impairment. A context-aware assistive robot needs constant knowledge of the user's kinematic state to assess the gait status and adjust its movement properly to provide optimal assistance. In this work, we experimentally validate the performance of two gait tracking algorithms using data from elderly patients; the first algorithm employs a Kalman Filter (KF), while the second one tracks the user legs separately using two probabilistically associated Particle Filters (PFs). The algorithms are compared according to their accuracy and robustness, using data captured from real experiments, where elderly subjects performed specific walking scenarios with physical assistance from a prototype Robotic Rollator. Sensorial data were provided by a laser rangefinder mounted on the robotic platform recording the movement of the user's legs. The accuracy of the proposed algorithms is analysed and validated with respect to ground truth data provided by a Motion Capture system tracking a set of visual markers worn by the patients. The robustness of the two tracking algorithms is also analysed comparatively in a complex maneuvering scenario. Current experimental findings demonstrate the superior performance of the PFs in difficult cases of occlusions and clutter, where KF tracking often fails.
Georgia Chalvatzaki, Xanthi S. Papageorgiou, Costas S. Tzafestas, Petros Maragos
ICRA4
2017 Real-time end-effector motion behavior planning approach using on-line point-cloud data towards a user adaptive assistive bath robot
abstract
Elderly people have particular needs in performing bathing activities, since these tasks require body flexibility. Our aim is to build an assistive robotic bath system, in order to increase the independence and safety of this procedure. Towards this end, the expertise of professional carers for bathing sequences and appropriate motions has to be adopted, in order to achieve natural, physical human - robot interaction. In this paper, a real-time end-effector motion planning method for an assistive bath robot, using on-line Point-Cloud information, is proposed. The visual feedback obtained from Kinect depth sensor is employed to adapt suitable washing paths to the user's body part motion and deformable surface. We make use of a navigation function-based controller, with guarantied globally uniformly asymptotic stability, and bijective transformations for the adaptation of the paths. Experiments were conducted with a rigid rectangular object for validation purposes, while a female subject took part to the experiment in order to evaluate and demonstrate the basic concepts of the proposed methodology.
Athanasios Dometios, Xanthi S. Papageorgiou, Antonis Arvanitakis, Costas S. Tzafestas, Petros Maragos
IROS5
2017 Estimating double support in pathological gaits using an HMM-based analyzer for an intelligent robotic walker
abstract
For a robotic walker designed to assist mobility constrained people, it is important to take into account the different spectrum of pathological walking patterns, which result into completely different needs to be covered for each specific user. For a deployable intelligent assistant robot it is necessary to have a precise gait analysis system, providing real-time monitoring of the user and extracting specific gait parameters, which are associated with the rehabilitation progress and the risk of fall. In this paper, we present a completely non-invasive framework for the on-line analysis of pathological human gait and the recognition of specific gait phases and events. The performance of this gait analysis system is assessed, in particular, as related to the estimation of double support phases, which are typically difficult to extract reliably, especially when applying non-wearable and non-intrusive technologies. Furthermore, the duration of double support phases constitutes an important gait parameter and a critical indicator in pathological gait patterns. The performance of this framework is assessed using real data collected from an ensemble of elderly persons with different pathologies. The estimated gait parameters are experimentally validated using ground truth data provided by a Motion Capture system. The results obtained and presented in this paper demonstrate that the proposed human data analysis (modeling, learning and inference) framework has the potential to support efficient detection and classification of specific walking pathologies, as needed to empower a cognitive robotic mobility-assistance device with user-adaptive and context-aware functionalities.
Georgia Chalvatzaki, Xanthi S. Papageorgiou, Costas S. Tzafestas, Petros Maragos
RO-MAN4
2017 Room-localized spoken command recognition in multi-room, multi-microphone environments
Isidoros Rodomagoulakis, Athanasios Katsamanis, Gerasimos Potamianos, Panagiotis Giannoulis, Antigoni Tsiami, Petros Maragos
Comput. Speech Lang.6
2017 Theoretical Analysis of Active Contours on Graphs
abstract
Active contour models based on partial differential equations have proved successful in image segmentation, yet the study of their formulation on arbitrary geometric graphs, which place no restrictions in the spatial configuration of samples, is still at an early stage. In this paper, we introduce geometric approximations of gradient and curvature on arbitrary graphs, which enable a straightforward extension of active contour models that are formulated through level sets to such general inputs. We prove convergence in probability of our gradient approximation to the true gradient value and derive an asymptotic upper bound for the error of this approximation for the class of random geometric graphs. Two different approaches for the approximation of curvature are presented, and both are also proved to converge in probability in the case of random geometric graphs. We propose neighborhood-based filtering on graphs to improve the accuracy of the aforementioned approximations and define two variants of Gaussian smoothing on graphs which include normalization in order to adapt to graph nonuniformities. The performance of our active contour framework on graphs is demonstrated in the segmentation of regular images and geographical data defined on arbitrary graphs, using geodesic active contours and active contours without edges as representative models in our experiments.
Christos Sakaridis, Kimon Drakopoulos, Petros Maragos
SIAM J. Imaging Sci.3
2017 Video-realistic expressive audio-visual speech synthesis for the Greek language
Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Pirros Tsiakoulis, Petros Maragos
Speech Commun.4
2017 Graph-Driven Diffusion and Random Walk Schemes for Image Segmentation
abstract
We propose graph-driven approaches to image segmentation by developing diffusion processes defined on arbitrary graphs. We formulate a solution to the image segmentation problem modeled as the result of infectious wavefronts propagating on an image-driven graph where pixels correspond to nodes of an arbitrary graph. By relating the popular Susceptible - Infected - Recovered epidemic propagation model to the Random Walker algorithm, we develop the Normalized Random Walker and a lazy random walker variant. The underlying iterative solutions of these methods are derived as the result of infections transmitted on this arbitrary graph. The main idea is to incorporate a degree-aware term into the original Random Walker algorithm in order to account for the node centrality of every neighboring node and to weigh the contribution of every neighbor to the underlying diffusion process. Our lazy random walk variant models the tendency of patients or nodes to resist changes in their infection status. We also show how previous work can be naturally extended to take advantage of this degreeaware term which enables the design of other novel methods. Through an extensive experimental analysis, we demonstrate the reliability of our approach, its small computational burden and the dimensionality reduction capabilities of graph-driven approaches. Without applying any regular grid constraint, the proposed graph clustering scheme allows us to consider pixellevel, node-level approaches and multidimensional input data by naturally integrating the importance of each node to the final clustering or segmentation solution. A software release containing implementations of this work and supplementary material can be found at: http://cvsp.cs.ntua.gr/research/GraphClustering/.
Christos G. Bampis, Petros Maragos, Alan C. Bovik
IEEE Trans. Image Process.2
2016 Multimodal human action recognition in assistive human-robot interaction
abstract
Within the context of assistive robotics we develop an intelligent interface that provides multimodal sensory processing capabilities for human action recognition. Human action is considered in multimodal terms, containing inputs such as audio from microphone arrays, and visual inputs from high definition and depth cameras. Exploring state-of-the-art approaches from automatic speech recognition, and visual action recognition, we multimodally recognize actions and commands. By fusing the unimodal information streams, we obtain the optimum multimodal hypothesis which is to be further exploited by the active mobility assistance robot in the framework of the MOBOT EU research project. Evidence from recognition experiments shows that by integrating multiple sensors and modalities, we increase multimodal recognition performance in the newly acquired challenging dataset involving elderly people while interacting with the assistive robot.
Isidoros Rodomagoulakis, Nikolaos Kardaris, Vassilis Pitsikalis, E. Mavroudi, Athanasios Katsamanis, Antigoni Tsiami, Petros Maragos
ICASSP7
2016 Towards a behaviorally-validated computational audiovisual saliency model
abstract
Computational saliency models aim at predicting, in a bottom-up fashion, where human attention is drawn in the presented (visual, auditory or audiovisual) scene and have been proven useful in applications like robotic navigation, image compression and movie summarization. Despite the fact that well-established auditory and visual saliency models have been validated in behavioral experiments, e.g., by means of eye-tracking, there is no established computational audiovisual saliency model validated in the same way. In this work, building on biologically-inspired models of visual and auditory saliency, we present a joint audiovisual saliency model and introduce the validation approach we follow to show that it is compatible with recent findings of psychology and neuroscience regarding multimodal integration and attention. In this direction, we initially focus on the "pip and pop" effect which has been observed in behavioral experiments and indicates that visual search in sequences of cluttered images can be significantly aided by properly timed non-spatial auditory signals presented alongside the target visual stimuli.
Antigoni Tsiami, Athanasios Katsamanis, Petros Maragos, Argiro Vatakis
ICASSP3
2016 Projective non-negative matrix factorization for unsupervised graph clustering
abstract
We develop an unsupervised graph clustering and image segmentation algorithm based on non-negative matrix factorization. We consider arbitrarily represented visual signals (in 2D or 3D) and use a graph embedding approach for image or point cloud segmentation. We extend a Projective Non-negative Matrix Factorization variant to include local spatial relationships over the image graph. By using properly defined region features, one can apply our method of unsupervised graph clustering for object and image segmentation. To demonstrate this, we apply our ideas on many graph based segmentation tasks such as 2D pixel and super-pixel segmentation and 3D point cloud segmentation. Finally, we show results comparable to those achieved by the only existing work in pixel based texture segmentation using Nonnegative Matrix Factorization, deploying a simple yet effective extension that is parameter free. We provide a detailed convergence proof of our spatially regularized method and various demonstrations as supplementary material. This novel work brings together graph clustering with image segmentation.
Christos G. Bampis, Petros Maragos, Alan C. Bovik
ICIP2
2016 Introducing temporal order of dominant visual word sub-sequences for human action recognition
abstract
We present a novel video representation for human action recognition by considering temporal sequences of visual words. Based on state-of-the-art dense trajectories, we introduce temporal bundles of dominant, that is most frequent, visual words. These are employed to construct a complementary action representation of ordered dominant visual word sequences, that additionally incorporates fine grained temporal information. We exploit the introduced temporal information by applying local sub-sequence alignment that quantifies the similarity between sequences. This facilitates the fusion of our representation with the bag-of-visual-words (BoVW) representation. Our approach incorporates sequential temporal structure and results in a low-dimensional representation compared to the BoVW, while still yielding a descent result when combined with it. Experiments on the KTH, Hollywood2 and the challenging HMDB51 datasets show that the proposed framework is complementary to the BoVW representation, which discards temporal order.
Nikolaos Kardaris, Vassilis Pitsikalis, E. Mavroudi, Petros Maragos
ICIP4
2016 FMRI-based perceptual validation of a computational model for visual and auditory saliency in videos
abstract
In this study, we make use of brain activation data to investigate the perceptual plausibility of a visual and an auditory model for visual and auditory saliency in video processing. These models have already been successfully employed in a number of applications. In addition, we experiment with parameters, modifications and suitable fusion schemes. As part of this work, fMRI data from complex video stimuli were collected, on which we base our analysis and results. The core part of the analysis involves the use of well-established methods for the manipulation of fMRI data and the examination of variability across brain responses of different individuals. Our results indicate a success in confirming the value of these saliency models in terms of perceptual plausibility.
Georgia Panagiotaropoulou, Petros Koutras, Athanasios Katsamanis, Petros Maragos, Athanasia Zlatintsi, Athanassios Protopapas, Eustratios Karavasilis, Nikolaos Smyrnis
ICIP4
2016 A multimedia gesture dataset for human robot communication: Acquisition, tools and recognition results
abstract
Motivated by the recent advances in human-robot interaction we present a new dataset, a suite of tools to handle it and state-of-the-art work on visual gestures and audio commands recognition. The dataset has been collected with an integrated annotation and acquisition web-interface that facilitates on-the-way temporal ground-truths for fast acquisition. The dataset includes gesture instances in which the subjects are not in strict setup positions, and contains multiple scenarios, not restricted to a single static configuration. We accompany it by a valuable suite of tools as the practical interface to acquire audio-visual data in the robotic operating system, a state-of-the-art learning pipeline to train visual gesture and audio command models, and an online gesture recognition system. Finally, we include a rich evaluation of the dataset providing rich and insightfull experimental recognition results.
Isidoros Rodomagoulakis, Nikolaos Kardaris, Vassilis Pitsikalis, Antonis Arvanitakis, Petros Maragos
ICIP5
2016 A Phase-Based Time-Frequency Masking for Multi-Channel Speech Enhancement in Domestic Environments
Alessio Brutti, Antigoni Tsiami, Athanasios Katsamanis, Petros Maragos
INTERSPEECH4
2016 A Platform for Building New Human-Computer Interface Systems that Support Online Automatic Recognition of Audio-Gestural Commands
abstract
We introduce a new framework to build human-computer interfaces that provide online automatic audio-gestural command recognition. The overall system allows the construction of a multimodal interface that recognizes user input expressed naturally as audio commands and manual gestures, captured by sensors such as Kinect. It includes a component for acquiring multimodal user data which is used as input to a module responsible for training audio-gestural models. These models are employed by the automatic recognition component, which supports online recognition of audio-visual modalities. The overall framework is exemplified by a working system use case. This demonstrates the potential of the overall software platform, which can be employed to build other new human-computer interaction systems. Moreover, users may populate libraries of models and/or data, that can be shared in the network. In this way users may reuse or extend existing systems.
Nikolaos Kardaris, Isidoros Rodomagoulakis, Vassilis Pitsikalis, Antonis Arvanitakis, Petros Maragos
ACM Multimedia5
2015 Max-product dynamical systems and applications to audio-visual salient event detection in videos
abstract
This paper introduces a theory for max-product systems by analyzing them as discrete-time nonlinear dynamical systems that obey a superposition of a weighted maximum type and evolve on nonlinear spaces which we call complete weighted lattices. Special cases of such systems have found applications in speech recognition as weighted finite-state transducers and in belief propagation on graphical models. Our theoretical approach establishes their representation in state and input-output spaces using monotone lattice operators, finds analytically their state and output responses using nonlinear convolutions, studies their stability, and provides optimal solutions to solving max-product matrix equations. Further, we apply these systems to extend the Viterbi algorithm in HMMs by adding control inputs and model cognitive processes such as detecting audio and visual salient events in multimodal video streams, which shows good performance as compared to human attention.
Petros Maragos, Petros Koutras
ICASSP1
2015 Multichannel speech enhancement using MEMS microphones
abstract
In this work, we investigate the efficacy of Micro Electro-Mechanical System (MEMS) microphones, a newly developed technology of very compact sensors, for multichannel speech enhancement. Experiments are conducted on real speech data collected using a MEMS microphone array. First, the effectiveness of the array geometry for noise suppression is explored, using a new corpus containing speech recorded in diffuse and localized noise fields with a MEMS microphone array configured in linear and hexagonal array geometries. Our results indicate superior performance of the hexagonal geometry. Then, MEMS microphones are compared to Electret Condenser Microphones (ECMs), using the ATHENA database, which contains speech recorded in realistic smart home noise conditions with hexagonal-type arrays of both microphone types. MEMS microphones exhibit performance similar to ECMs. Good performance, versatility in placement, small size, and low cost, make MEMS microphones attractive for multichannel speech processing.
Z.-I. Skordilis, Antigoni Tsiami, Petros Maragos, Gerasimos Potamianos, Luca Spelgatti, Roberto Sannino
ICASSP3
2015 Unifying the random walker algorithm and the SIR model for graph clustering and image segmentation
abstract
In this paper, we explore the image segmentation task using a graph clustering approach. We formulate this clustering as a diffusion scheme whose steady state is determined by the Random Walker (RW) method. Then, we discover the equivalence of this diffusion with the Susceptible - Infected - Recovered (SIR) model, a well-studied epidemic propagation model. We further argue that using a Region Adjacency Graph (RAG) exploits the clustering properties and leads to a dimensionality reduction. Finally, we propose a novel method called Normalized Random Walker (NRW) algorithm which extends the RW method. Qualitative and quantitative experiments validate the efficiency and robustness of our method, with respect to parameter tuning, seed quality and location.
Christos G. Bampis, Petros Maragos
ICIP2
2015 Estimation of eye gaze direction angles based on active appearance models
abstract
In this paper we demonstrate efficient methods for continuous estimation of eye gaze angles with application to sign language videos. The difficulty of the task lies on the fact that those videos contain images with low face resolution since they are recorded from distance. First, we proceed to the modeling of face and eyes region by training and fitting Global and Local Active Appearance Models (LAAM). Next, we propose a system for eye gaze estimation based on a machine learning approach. In the first stage of our method, we classify gaze into discrete classes using GMMs that are based either on the parameters of the LAAM, or on HOG descriptors for the eyes region. We also propose a method for computing gaze direction angles from GMM log-likelihoods. We qualitatively and quantitatively evaluate our methods on two sign language databases and compare with a state of the art geometric model of the eye based on LAAM landmarks, which provides an estimate in direction angles. Finally, we further evaluate our framework by getting ground truth data from an eye tracking system Our proposed methods, and especially the GMMs using LAAM parameters, demonstrate high accuracy and robustness even in challenging tasks.
Petros Koutras, Petros Maragos
ICIP2
2015 Predicting audio-visual salient events based on visual, audio and text modalities for movie summarization
abstract
In this paper, we present a new and improved synergistic approach to the problem of audio-visual salient event detection and movie summarization based on visual, audio and text modalities. Spatio-temporal visual saliency is estimated through a perceptually inspired frontend based on 3D (space, time) Gabor filters and frame-wise features are extracted from the saliency volumes. For the auditory salient event detection we extract features based on Teager-Kaiser Energy Operator, while text analysis incorporates part-of-speech tagging and affective modeling of single words on the movie subtitles. For the evaluation of the proposed system, we employ an elementary and non-parametric classification technique like KNN. Detection results are reported on the MovSum database, using objective evaluations against ground-truth denoting the perceptually salient events, and human evaluations of the movie summaries. Our evaluation verifies the appropriateness of the proposed methods compared to our baseline system. Finally, our newly proposed summarization algorithm produces summaries that consist of salient and meaningful events, also improving the comprehension of the semantics.
Petros Koutras, Athanasia Zlatintsi, Elias Iosif, Athanasios Katsamanis, Petros Maragos, Alexandros Potamianos
ICIP5
2015 Hidden markov modeling of human pathological gait using laser range finder for an assisted living intelligent robotic walker
abstract
The precise analysis of a patient's or an elderly person's walking pattern is very important for an effective intelligent active mobility assistance robot. This walking pattern can be described by a cyclic motion, which can be modeled using the consecutive gait phases. In this paper, we present a completely non-invasive framework for analyzing and recognizing a pathological human walking gait pattern. Our framework utilizes a laser range finder sensor to detect and track the human legs, and an appropriately synthesized Hidden Markov Model (HMM) for state estimation, and recognition of the gait patterns. We demonstrate the applicability of this setup using real data, collected from an ensemble of different elderly persons with a number of pathologies. The results presented in this paper demonstrate that the proposed human data analysis scheme has the potential to provide the necessary methodological (modeling, inference, and learning) framework for a cognitive behavior-based robot control system. More specifically, the proposed framework has the potential to be used for the classification of specific walking pathologies, which is needed for the development of a context-aware robot mobility assistant.
Xanthi S. Papageorgiou, Georgia Chalvatzaki, Costas S. Tzafestas, Petros Maragos
IROS4
2015 Multimodal gesture recognition via multiple hypotheses rescoring
Vassilis Pitsikalis, Athanasios Katsamanis, Stavros Theodorakis, Petros Maragos
J. Mach. Learn. Res.4
2015 Structure Tensor Total Variation
abstract
We introduce a novel generic energy functional that we employ to solve inverse imaging problems within a variational framework. The proposed regularization family, termed as structure tensor total variation (STV), penalizes the eigenvalues of the structure tensor and is suitable for both grayscale and vector-valued images. It generalizes several existing variational penalties, including the total variation seminorm and vectorial extensions of it. Meanwhile, thanks to the structure tensor's ability to capture first-order information around a local neighborhood, the STV functionals can provide more robust measures of image variation. Further, we prove that the STV regularizers are convex while they also satisfy several invariance properties w.r.t. image transformations. These properties qualify them as ideal candidates for imaging applications. In addition, for the discrete version of the STV functionals we derive an equivalent definition that is based on the patch-based Jacobian operator, a novel linear operator which extends the Jacobian matrix. This alternative definition allow us to derive a dual problem formulation. The duality of the problem paves the way for employing robust tools from convex optimization and enables us to design an efficient and parallelizable optimization algorithm. Finally, we present extensive experiments on various inverse imaging problems, where we compare our regularizers with other competing regularization approaches. Our results are shown to be systematically superior, both quantitatively and visually.
Stamatios Lefkimmiatis, Anastasios Roussos, Petros Maragos, Michael Unser
SIAM J. Imaging Sci.3
2015 A perceptually based spatio-temporal computational framework for visual saliency estimation
Petros Koutras, Petros Maragos
Signal Process. Image Commun.2
2014 Robust far-field spoken command recognition for home automation combining adaptation and multichannel processing
abstract
The paper presents our approach to speech-controlled home automation. We are focusing on the detection and recognition of spoken commands preceded by a key-phrase as recorded in a voice-enabled apartment by a set of multiple microphones installed in the rooms. For both problems we investigate robust modeling, environmental adaptation and multichannel processing to cope with a) insufficient training data and b) the far-field effects and noise in the apartment. The proposed integrated scheme is evaluated in a challenging and highly realistic corpus of simulated audio recordings and achieves F-measure close to 0.70 for key-phrase spotting and word accuracy close to 98% for the command recognition task.
Athanasios Katsamanis, Isidoros Rodomagoulakis, Gerasimos Potamianos, Petros Maragos, Antigoni Tsiami
ICASSP4
2014 Advances on action recognition in videos using an interest point detector based on multiband spatio-temporal energies
abstract
This paper proposes a new visual framework for action recognition in videos, that consists of an energy detector coupled with a carefully designed multiband energy based filterbank. The tracking of video energy is performed using perceptually inspired 3D Gabor filters combined with ideas from Dominant Energy Analysis. Within this framework, we utilize different alternatives such as non-linear energy operators where actions are implicitly considered as manifestations of spatio-temporal oscillations in the dynamic visual stream. Texture and motion decomposition of actions through multiband filtering is the basis of our approach. This new energy-based saliency measure of action videos leads to the extraction of local spatio-temporal interest points that give promising results for the task of action recognition. Such interest points are processed further in order to formulate a robust representation of an action in a video. Theoretical formulation is supported by evaluation in two popular action databases, in which our method seems to outperform the state of the art.
Kevis-Kokitsi Maninis, Petros Koutras, Petros Maragos
ICIP3
2014 Kinect-based multimodal gesture recognition using a two-pass fusion scheme
abstract
We present a new framework for multimodal gesture recognition that is based on a two-pass fusion scheme. In this, we deal with a demanding Kinect-based multimodal dataset, which was introduced in a recent gesture recognition challenge. We employ multiple modalities, i.e., visual cues, such as colour and depth images, as well as audio, and we specifically extract feature descriptors of the hands' movement, handshape, and audio spectral properties. Based on these features, we statistically train separate unimodal gesture-word models, namely hidden Markov models, explicitly accounting for the dynamics of each modality. Multimodal recognition of unknown gesture sequences is achieved by combining these models in a late, two-pass fusion scheme that exploits a set of unimodally generated n-best recognition hypotheses. The proposed scheme achieves 88.2% gesture recognition accuracy in the Kinect-based multimodal dataset, outperforming all recently published approaches on the same challenging multimodal gesture recognition task.
Georgios Pavlakos, Stavros Theodorakis, Vassilis Pitsikalis, Athanasios Katsamanis, Petros Maragos
ICIP5
2014 Hidden Markov modeling of human normal gait using laser range finder for a mobility assistance robot
abstract
For an effective intelligent active mobility assistance robot, the walking pattern of a patient or an elderly person has to be analyzed precisely. A well-known fact is that the walking patterns are gaits, that is, cyclic patterns with several consecutive phases. These cyclic motions can be modeled using the consecutive gait phases. In this paper, we present a completely non-invasive framework for analyzing a normal human walking gait pattern. Our framework utilizes a laser range finder sensor to collect the data, a combination of filters to preprocess these data, and an appropriately synthesized Hidden Markov Model (HMM) for state estimation, and recognition of the gait data. We demonstrate the applicability of this setup using real data, collected from an ensemble of different persons. The results presented in this paper demonstrate that the proposed human data analysis scheme has the potential to provide the necessary methodological (modeling, inference, and learning) framework for a cognitive behavior-based robot control system. More specifically, the proposed framework has the potential to be used for the recognition of abnormal gait patterns and the subsequent classification of specific walking pathologies, which is needed for the development of a context-aware robot mobility assistant.
Xanthi S. Papageorgiou, Georgia Chalvatzaki, Costas S. Tzafestas, Petros Maragos
ICRA4
2014 ATHENA: a Greek multi-sensory database for home automation control uthor: isidoros rodomagoulakis (NTUA, Greece)
Antigoni Tsiami, Isidoros Rodomagoulakis, Panagiotis Giannoulis, Athanasios Katsamanis, Gerasimos Potamianos, Petros Maragos
INTERSPEECH6
2014 The DIRHA simulated corpus
Luca Cristoforetti, Mirco Ravanelli, Maurizio Omologo, Alessandro Sosi, Alberto Abad, Martin Hagmüller, Petros Maragos
LREC7
2014 Dynamic-static unsupervised sequentiality, statistical subunits and lexicon for sign language recognition
Stavros Theodorakis, Vassilis Pitsikalis, Petros Maragos
Image Vis. Comput.3
2013 Novel Representations, Methods, and Algorithms in Computer Vision
Kostas Daniilidis, Petros Maragos, Nikos Paragios
Int. J. Comput. Vis.2
2013 Dynamic affine-invariant shape-appearance handshape features and classification in sign language videos
Anastasios Roussos, Stavros Theodorakis, Vassilis Pitsikalis, Petros Maragos
J. Mach. Learn. Res.4
2013 Multiscale Fractal Analysis of Musical Instrument Signals With Application to Recognition
abstract
In this paper, we explore nonlinear methods, inspired by the fractal theory for the analysis of the structure of music signals at multiple time scales, which is of importance both for their modeling and for their automatic computer-based recognition. We propose the multiscale fractal dimension (MFD) profile as a short-time descriptor, useful to quantify the multiscale complexity and fragmentation of the different states of the music waveform. We have experimentally found that this descriptor can discriminate several aspects among different music instruments, which is verified by further analysis on synthesized sinusoidal signals. We compare the descriptiveness of our features against that of Mel frequency cepstral coefficients (MFCCs), using both static and dynamic classifiers such as Gaussian mixture models (GMMs) and hidden Markov models (HMMs). The method and features proposed in this paper appear to be promising for music signal analysis, due to their capability for multiscale analysis of the signals and their applicability in recognition, as they accomplish an error reduction of up to 32%. These results are quite interesting and render the descriptor of direct applicability in large-scale music classification tasks.
Athanasia Zlatintsi, Petros Maragos
IEEE Trans. Speech Audio Process.2
2013 Multimodal Saliency and Fusion for Movie Summarization Based on Aural, Visual, and Textual Attention
abstract
Multimodal streams of sensory information are naturally parsed and integrated by humans using signal-level feature extraction and higher level cognitive processes. Detection of attention-invoking audiovisual segments is formulated in this work on the basis of saliency models for the audio, visual, and textual information conveyed in a video stream. Aural or auditory saliency is assessed by cues that quantify multifrequency waveform modulations, extracted through nonlinear operators and energy tracking. Visual saliency is measured through a spatiotemporal attention model driven by intensity, color, and orientation. Textual or linguistic saliency is extracted from part-of-speech tagging on the subtitles information available with most movie distributions. The individual saliency streams, obtained from modality-depended cues, are integrated in a multimodal saliency curve, modeling the time-varying perceptual importance of the composite video stream and signifying prevailing sensory events. The multimodal saliency representation forms the basis of a generic, bottom-up video summarization algorithm. Different fusion schemes are evaluated on a movie database of multimodal saliency annotations with comparative results provided across modalities. The produced summaries, based on low-level features and content-independent fusion and selection, are of subjectively high aesthetic and informative quality.
Georgios Evangelopoulos, Athanasia Zlatintsi, Alexandros Potamianos, Petros Maragos, Konstantinos Rapantzikos, Georgios Skoumas, Yannis Avrithis
IEEE Trans. Multim.4
2012 The Dicta-Sign Wiki: Enabling Web Communication for the Deaf
Eleni Efthimiou, Stavroula-Evita Fotinea, Thomas Hanke 0001, John R. W. Glauert, Richard Bowden, Annelies Braffort, Christophe Collet 0002, Petros Maragos, François Lefebvre-Albaret
ICCHP (2)8
2012 Unsupervised classification of extreme facial events using active appearance models tracking for sign language videos
abstract
We propose an Unsupervised method for Extreme States Classification (UnESC) on feature spaces of facial cues of interest. The method is built upon Active Appearance Models (AAM) face tracking and on feature extraction of Global and Local AAMs. UnESC is applied primarily on facial pose, but is shown to be extendable for the case of local models on the eyes and mouth. Given the importance of facial events in Sign Languages we apply the UnESC on videos from two sign language corpora, both American (ASL) and Greek (GSL) yielding promising qualitative and quantitative results. Apart from the detection of extreme facial states, the proposed Un-ESC also has impact for SL corpora lacking any facial annotations.
Epameinondas Antonakos, Vassilis Pitsikalis, Isidoros Rodomagoulakis, Petros Maragos
ICIP4
2012 Dominant spatio-temporal modulations and energy tracking in videos: Application to interest point detection for action recognition
abstract
The presence of multiband amplitude and frequency modulations (AM-FM) in wideband signals, such as textured images or speech, has led to the development of efficient multicomponent modulation models for low-level image and sound analysis. Moreover, compact yet descriptive representations have emerged by tracking, through non-linear energy operators, the dominant model components across time, space or frequency. In this paper, we propose a generalization of such approaches in the 3D spatio-temporal domain and explore the potential of incorporating the Dominant Component Analysis scheme for interest point detection and human action recognition in videos. Within this framework, actions are implicitly considered as manifestations of spatio-temporal oscillations in the dynamic visual stream. Multiband filtering and energy operators are applied to track the source energy in both spatial and temporal frequency bands. A new measure for extracting keypoint locations is formulated as the temporal dominant energy computed over the spatial dominant components, in terms of their modulation energy, of input video frames. Theoretical formulation is supported by evaluation and comparisons in human action classification, which demonstrate the potential of the proposed spatio-temporal detector.
Christos Georgakis 0001, Petros Maragos, Georgios Evangelopoulos, Dimitrios Dimitriadis
ICIP2
2012 Human action recognition using Histographic methods and hidden Markov models for visual martial arts applications
abstract
Human Action Recognition is being used with an increasing rate in applications designed to describe human activity in everyday life. However, some areas still remain far from the epicenter of scientific research, like dynamic problems of detection and classification of movements from visual Martial Arts. With this paper, we are proposing a novel recognition system focused on these types of action, based on the use of local spatio-temporal features from Histographic methods, as those extracted by Histograms of Oriented Gradient (HOG) and Histograms of Optical Flow (HOF), while we also pursue the reduction of the problem's dimensionality by applying on them Principal Components Analysis (PCA). In continuation, we combine these features with the use of Hidden Markov Models (HMM) in order to train models for each different movement. Our system is tested with very encouraging results upon a database comprising sequences of shotokan karate movements (katas), created by us for the needs of this research. In parallel, we additionally attach an educational character to our application, with the extraction of a score for the accuracy of execution of each movement based on the prototypes we have built.
Sotirios Stasinopoulos, Petros Maragos
ICIP2
2012 Recognitionwith raw canonical phonetic movement and handshape subunits on videos of continuous Sign Language
abstract
The visual processing of Sign Language (SL) videos offers multiple interdisciplinary challenges for image processing and recognition. Based on tracking and visual feature extraction, we investigate SL visual phonetic modeling by exploiting statistical subunit (SU) models of movement-position and handshape. We further propose a new framework to construct a data-driven lexicon that retains phonetics' movement information and to perform automatic recognition of continuous SL videos. We construct phonetically meaningful transition SU, named as raw canonical phonetic subunits (SU-CanRaw). Then, we integrate via a Hidden Markov Model multistream scheme the SU-CanRaw extended for both hands, with handshape SU, based on our previous work on Affine-invariant Shape-Appearance Models. By applying the all-inclusive framework on continuous SL videos, we automatically generate a data-driven lexicon that can be further exploited, for automatic analysis of SL corpora, and continuous SL recognition. The recognition experiments, conducted on a newly acquired continuous SL corpus, lead to promising results.
Stavros Theodorakis, Vassilis Pitsikalis, Isidoros Rodomagoulakis, Petros Maragos
ICIP4
2011 On the Effects of Filterbank Design and Energy Computation on Robust Speech Recognition
abstract
In this paper, we examine how energy computation and filterbank design contribute to the overall front-end robustness, especially when the investigated features are applied to noisy speech signals, in mismatched training-testing conditions. In prior work (“Auditory Teager energy cepstrum coefficients for robust speech recognition,” D. Dimitriadis, P. Maragos, and A. Potamianos, in Proc. Eurospeech'05, Sep. 2005), a novel feature set called “Teager energy cepstrum coefficients” (TECCs) has been proposed, employing a dense, smooth filterbank and alternative energy computation schemes. TECCs were shown to be more robust to noise and exhibit improved performance compared to the widely used Mel frequency cepstral coefficients (MFCCs). In this paper, we attempt to interpret these results using a combined theoretical and experimental analysis framework. Specifically, we investigate in detail the connection between the filterbank design, i.e., the filter shape and bandwidth, the energy estimation scheme and the automatic speech recognition (ASR) performance under a variety of additive and/or convolutional noise conditions. For this purpose: 1) the performance of filterbanks using triangular, Gabor, and Gammatone filters with various bandwidths and filter positions are examined under different noisy speech recognition tasks, and 2) the squared amplitude and Teager-Kaiser energy operators are compared as two alternative approaches of computing the signal energy. Our end-goal is to understand how to select the most efficient filterbank and energy computation scheme that are maximally robust under both clean and noisy recording conditions. Theoretical and experimental results show that: 1) the filter bandwidth is one of the most important factors affecting speech recognition performance in noise, while the shape of the filter is of secondary importance, and 2) the Teager-Kaiser operator outperforms (on the average and for most noise types) the squared amplitude energy computation scheme for speech recognition in noisy conditions, especially, for large filter bandwidths. Experimental results show that selecting the appropriate filterbank and energy computation scheme can lead to significant error rate reduction over both MFCC and perceptual linear predicion (PLP) features for a variety of speech recognition tasks. A relative error rate reduction of up to ~ 30% for MFCCs and ~ 39% for PLPs is shown for the Aurora-3 Spanish Task.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
IEEE Trans. Speech Audio Process.2
2010 Model-level data-driven sub-units for signs in videos of continuous Sign Language
abstract
We investigate the issue of sign language automatic phonetic subunit modeling, that is completely data driven and without any prior phonetic information. A first step of visual processing leads to simple and effective region-based visual features. Prior to the sub-unit modeling we propose to employ a pronunciation clustering step with respect to each sign. Afterwards, for each sign and pronunciation group we find the time segmentation at the hidden Markov model (HMM) level. The models employed refer to movements as a sequence of dominant hand positions. The constructed segments are exploited explicitly at the model level via hierarchical clustering of HMMs and lead to the data-driven movement sub-unit construction. The constructed movement sub-units are evaluated in qualitative analysis experiments on data from the Boston University (BU)- 400 American Sign Language corpus showing promising results.
Stavros Theodorakis, Vassilis Pitsikalis, Petros Maragos
ICASSP3
2010 Spatial bayesian surprise for image saliency and quality assessment
abstract
We propose an alternative interpretation of Bayesian surprise in the spatial domain, to account for saliency arising from contrast in image context. Our saliency formulation is integrated in three different application scenaria, with considerable improvements in performance: 1) visual attention prediction, validated using eye- and mouse-tracking data, 2) region of interest detection, to improve scale selection and localization, 3) image quality assessment to achieve better agreement with subjective human evaluations.
Ioannis Gkioulekas, Georgios Evangelopoulos, Petros Maragos
ICIP3
2010 Tensor-based image diffusions derived from generalizations of the Total Variation and Beltrami Functionals
abstract
We introduce a novel functional for vector-valued images that generalizes several variational methods, such as the Total Variation and Beltrami Functionals. This functional is based on the structure tensor that describes the geometry of image structures within the neighborhood of each point. We first generalize the Beltrami functional based on the image patches and using embeddings in high dimensional spaces. Proceeding to the most general form of the proposed functional, we prove that its minimization leads to a nonlinear anisotropic diffusion that is regularized, in the sense that its diffusion tensor contains convolutions with a kernel. Using this result we propose two novel diffusion methods, the Generalized Beltrami Flow and the Tensor Total Variation. These methods combine the advantages of the variational approaches with those of the tensor-based diffusion approaches.
Anastasios Roussos, Petros Maragos
ICIP2
2010 Affine-invariant modeling of shape-appearance images applied on sign language handshape classification
abstract
We propose a novel affine-invariant modeling of handshape-appearance images, which offers a compact and descriptive representation of the hand configurations. Our approach combines: (1) A hybrid representation of both shape and appearance of the hand that models the handshapes without any landmark points. (2) Modeling of the shape-appearance images with a linear combination of variation images that is followed by an affine transformation, which accounts for modest pose variation. (3) Finally, an optimization based fitting process that results on the estimated variation image coefficients that are further employed as features. The proposed modeling is applied on handshapes from Sign Language video data after segmentation and tracking. It is evaluated on extensive experiments of handshape classification, which investigate the effect of the involved parameters and moreover provide a variety of comparisons to base-line approaches found in the literature. The results of at least 10.5% absolute improvement indicate the effectiveness of our approach in the handshape classification problem.
Anastasios Roussos, Stavros Theodorakis, Vassilis Pitsikalis, Petros Maragos
ICIP4
2009 GridNews: A distributed automatic Greek broadcast transcription system
abstract
In this paper, a distributed system storing and retrieving broadcast news data recorded from the Greek television is presented. These multimodal data are processed in a grid computational environment interconnecting distributed data storage and processing subsystems. The innovative element of this system is the implementation of the signal processing algorithms in this grid environment, offering additional flexibility and computational power. Among the developed signal processing modules are: the Segmentor, cutting up the original videos into shorter ones, the classifier, recognizing whether these short videos contain speech or not, the Greek large-vocabulary speech recognizer, transcribing speech into written text, and finally the text search engine and the video retriever. All the processed data are stored and retrieved in geographically distributed storage elements. A user-friendly, Web-based interface is developed, facilitating the transparent import and storage of new multimodal data, their off-line processing and finally, their search and retrieval.
Dimitrios Dimitriadis, A. Metallinou, Ioannis Konstantinou, Georgios I. Goumas, Petros Maragos, Nectarios Koziris
ICASSP5
2009 Video event detection and summarization using audio, visual and text saliency
abstract
Detection of perceptually important video events is formulated here on the basis of saliency models for the audio, visual and textual information conveyed in a video stream. Audio saliency is assessed by cues that quantify multifrequency waveform modulations, extracted through nonlinear operators and energy tracking. Visual saliency is measured through a spatiotemporal attention model driven by intensity, color and motion. Text saliency is extracted from part-of-speech tagging on the subtitles information available with most movie distributions. The various modality curves are integrated in a single attention curve, where the presence of an event may be signified in one or multiple domains. This multimodal saliency curve is the basis of a bottom-up video summarization algorithm, that refines results from unimodal or audiovisual-based skimming. The algorithm performs favorably for video summarization in terms of informativeness and enjoyability.
Georgios Evangelopoulos, Athanasia Zlatintsi, Georgios Skoumas, Konstantinos Rapantzikos, Alexandros Potamianos, Petros Maragos, Yannis Avrithis
ICASSP6
2009 Product-HMMs for automatic sign language recognition
abstract
We address multistream sign language recognition and focus on efficient multistream integration schemes. Alternative approaches are investigated and the application of Product-HMMs (PHMM) is proposed. The PHMM is a variant of the general multistream HMM that also allows for partial asynchrony between the streams. Experiments in classification and isolated sign recognition for the Greek sign language using different fusion methods, show that the PHMMs perform the best. Fusing movement and shape information with the PHMMs has increased sign classification performance by 1,2% in comparison to the Parallel HMM fusion model. Isolated sign recognition rate increased by 8,3% over movement only models and by 1,5% over movement-shape models using multistream HMMs.
Stavros Theodorakis, Athanasios Katsamanis, Petros Maragos
ICASSP3
2009 Poisson-Haar Transform: A nonlinear multiscale representation for photon-limited image denoising
abstract
We present a novel multiscale image representation belonging to the class of multiscale multiplicative decompositions, which we term Poisson-Haar transform. The proposed representation is well-suited for analyzing images degraded by signal-dependent Poisson noise, allowing efficient estimation of their underlying intensity by means of multiscale Bayesian schemes. The Poisson-Haar decomposition has a direct link to the standard 2D Haar wavelet transform, thus retaining many of the properties that have made wavelets successful in signal processing and analysis. The practical relevance and effectiveness of the proposed approach is verified through denoising experiments on simulated and real-world photon-limited images.
Stamatios Lefkimmiatis, George Papandreou, Petros Maragos
ICIP3
2009 Overview of adaptive morphology: Trends and perspectives
abstract
In this paper we briefly overview emerging trends in `Adaptive Morphology', i.e. work related to the theory and/or applications of image analysis filters, systems, or algorithms based on mathematical morphology, that are adaptive w.r.t. to space or intensity or use any other adaptive scheme. We present a new classification of work in this area structured along several major theoretical perspectives. We then sample specific approaches that develop spatially-variant structuring elements or intensity level-adaptive operators, modeled and implemented either via conventional nonlinear digital filtering or via geometric PDEs. Finally, we discuss some applications.
Petros Maragos, Corinne Vachier
ICIP1
2009 Tongue tracking in Ultrasound images with Active Appearance Models
abstract
Tongue Ultrasound imaging is widely used for human speech production analysis and modeling. In this paper, we propose a novel method to automatically detect and track the tongue contour in Ultrasound (US) videos. Our method is built on a variant of Active Appearance Modeling. It incorporates shape prior information and can estimate the entire tongue contour robustly and accurately in a sequence of US frames. Experimental evaluation demonstrates the effectiveness of our approach and its improved performance compared to previously proposed tongue tracking techniques.
Anastasios Roussos, Athanasios Katsamanis, Petros Maragos
ICIP3
2009 Gestural teleoperation of a mobile robot based on visual recognition of sign language static handshapes
abstract
This paper presents results achieved in the frames of a national research project (titled ldquoDIANOEMArdquo), where visual analysis and sign recognition techniques have been explored on Greek Sign Language (GSL) data. Besides GSL modelling, the aim was to develop a pilot application for teleoperating a mobile robot using natural hand signs. A small vocabulary of hand signs has been designed to enable desktopbased teleoperation at a high-level of supervisory telerobotic control. Real-time visual recognition of the hand images is performed by training a multi-layer perceptron (MLP) neural network. Various shape descriptors of the segmented hand posture images have been explored as inputs to the MLP network. These include Fourier shape descriptors on the contour of the segmented hand sign images, moments, compactness, eccentricity, and histogram of the curvature. We have examined which of these shape descriptors are best suited for real-time recognition of hand signs, in relation to the number and choice of hand postures, in order to achieve maximum recognition performance. The hand-sign recognizer has been integrated in a graphical user interface, and has been implemented with success on a pilot application for real-time desktop-based gestural teleoperation of a mobile robot vehicle.
Costas S. Tzafestas, Nikos Mitsou, Nikos Georgakarakos, Olga Diamanti, Petros Maragos, Stavroula-Evita Fotinea, Eleni Efthimiou
RO-MAN5
2009 Reversible Interpolation of Vectorial Images by an Anisotropic Diffusion-Projection PDE
Anastasios Roussos, Petros Maragos
Int. J. Comput. Vis.2
2009 Texture Analysis and Segmentation Using Modulation Features, Generative Models, and Weighted Curve Evolution
abstract
In this work we approach the analysis and segmentation of natural textured images by combining ideas from image analysis and probabilistic modeling. We rely on AM-FM texture models and specifically on the Dominant Component Analysis (DCA) paradigm for feature extraction. This method provides a low-dimensional, dense and smooth descriptor, capturing essential aspects of texture, namely scale, orientation, and contrast. Our contributions are at three levels of the texture analysis and segmentation problems: First, at the feature extraction stage we propose a Regularized Demodulation Algorithm that provides more robust texture features and explore the merits of modifying the channel selection criterion of DCA. Second, we propose a probabilistic interpretation of DCA and Gabor filtering in general, in terms of Local Generative Models. Extending this point of view to edge detection facilitates the estimation of posterior probabilities for the edge and texture classes. Third, we propose the Weighted Curve Evolution scheme that enhances the Region Competition/ Geodesic Active Regions methods by allowing for the locally adaptive fusion of heterogeneous cues. Our segmentation results are evaluated on the Berkeley Segmentation Benchmark, and compare favorably to current state-of-the-art methods.
Iasonas Kokkinos, Georgios Evangelopoulos, Petros Maragos
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Synergy between Object Recognition and Image Segmentation Using the Expectation-Maximization Algorithm
abstract
In this work, we formulate the interaction between image segmentation and object recognition in the framework of the Expectation-Maximization (EM) algorithm. We consider segmentation as the assignment of image observations to object hypotheses and phrase it as the E-step, while the M-step amounts to fitting the object models to the observations. These two tasks are performed iteratively, thereby simultaneously segmenting an image and reconstructing it in terms of objects. We model objects using Active Appearance Models (AAMs) as they capture both shape and appearance variation. During the E-step, the fidelity of the AAM predictions to the image is used to decide about assigning observations to the object. For this, we propose two top-down segmentation algorithms. The first starts with an oversegmentation of the image and then softly assigns image segments to objects, as in the common setting of EM. The second uses curve evolution to minimize a criterion derived from the variational interpretation of EM and introduces AAMs as shape priors. For the M-step, we derive AAM fitting equations that accommodate segmentation information, thereby allowing for the automated treatment of occlusions. Apart from top-down segmentation results, we provide systematic experiments on object detection that validate the merits of our joint segmentation and recognition approach.
Iasonas Kokkinos, Petros Maragos
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Analysis and classification of speech signals by generalized fractal dimension features
Vassilis Pitsikalis, Petros Maragos
Speech Commun.2
2009 Face Active Appearance Modeling and Speech Acoustic Information to Recover Articulation
abstract
We are interested in recovering aspects of vocal tract's geometry and dynamics from speech, a problem referred to as speech inversion. Traditional audio-only speech inversion techniques are inherently ill-posed since the same speech acoustics can be produced by multiple articulatory configurations. To alleviate the ill-posedness of the audio-only inversion process, we propose an inversion scheme which also exploits visual information from the speaker's face. The complex audiovisual-to-articulatory mapping is approximated by an adaptive piecewise linear model. Model switching is governed by a Markovian discrete process which captures articulatory dynamic information. Each constituent linear mapping is effectively estimated via canonical correlation analysis. In the described multimodal context, we investigate alternative fusion schemes which allow interaction between the audio and visual modalities at various synchronization levels. For facial analysis, we employ active appearance models (AAMs) and demonstrate fully automatic face tracking and visual feature extraction. Using the AAM features in conjunction with audio features such as Mel frequency cepstral coefficients (MFCCs) or line spectral frequencies (LSFs) leads to effective estimation of the trajectories followed by certain points of interest in the speech production system. We report experiments on the QSMT and MOCHA databases which contain audio, video, and electromagnetic articulography data recorded in parallel. The results show that exploiting both audio and visual modalities in a multistream hidden Markov model based scheme clearly improves performance relative to either audio or visual-only estimation.
Athanasios Katsamanis, George Papandreou, Petros Maragos
IEEE Trans. Speech Audio Process.3
2009 Adaptive Multimodal Fusion by Uncertainty Compensation With Application to Audiovisual Speech Recognition
abstract
While the accuracy of feature measurements heavily depends on changing environmental conditions, studying the consequences of this fact in pattern recognition tasks has received relatively little attention to date. In this paper, we explicitly take feature measurement uncertainty into account and show how multimodal classification and learning rules should be adjusted to compensate for its effects. Our approach is particularly fruitful in multimodal fusion scenarios, such as audiovisual speech recognition, where multiple streams of complementary time-evolving features are integrated. For such applications, provided that the measurement noise uncertainty for each feature stream can be estimated, the proposed framework leads to highly adaptive multimodal fusion rules which are easy and efficient to implement. Our technique is widely applicable and can be transparently integrated with either synchronous or asynchronous multimodal sequence integration architectures. We further show that multimodal fusion methods relying on stream weights can naturally emerge from our scheme under certain assumptions; this connection provides valuable insights into the adaptivity properties of our multimodal uncertainty compensation approach. We show how these ideas can be practically applied for audiovisual speech recognition. In this context, we propose improved techniques for person-independent visual feature extraction and uncertainty estimation with active appearance models, and also discuss how enhanced audio features along with their uncertainty estimates can be effectively computed. We demonstrate the efficacy of our approach in audiovisual speech recognition experiments on the CUAVE database using either synchronous or asynchronous multimodal integration models.
George Papandreou, Athanasios Katsamanis, Vassilis Pitsikalis, Petros Maragos
IEEE Trans. Speech Audio Process.4
2009 Bayesian Inference on Multiscale Models for Poisson Intensity Estimation: Applications to Photon-Limited Image Denoising
abstract
We present an improved statistical model for analyzing Poisson processes, with applications to photon-limited imaging. We build on previous work, adopting a multiscale representation of the Poisson process in which the ratios of the underlying Poisson intensities (rates) in adjacent scales are modeled as mixtures of conjugate parametric distributions. Our main contributions include: 1) a rigorous and robust regularized expectation-maximization (EM) algorithm for maximum-likelihood estimation of the rate-ratio density parameters directly from the noisy observed Poisson data (counts); 2) extension of the method to work under a multiscale hidden Markov tree model (HMT) which couples the mixture label assignments in consecutive scales, thus modeling interscale coefficient dependencies in the vicinity of image edges; 3) exploration of a 2-D recursive quad-tree image representation, involving Dirichlet-mixture rate-ratio densities, instead of the conventional separable binary-tree image representation involving beta-mixture rate-ratio densities; and 4) a novel multiscale image representation, which we term Poisson-Haar decomposition, that better models the image edge structure, thus yielding improved performance. Experimental results on standard images with artificially simulated Poisson noise and on real photon-limited images demonstrate the effectiveness of the proposed techniques.
Stamatios Lefkimmiatis, Petros Maragos, George Papandreou
IEEE Trans. Image Process.2
2008 Image decomposition into structure and texture subcomponents with multifrequency modulation constraints
abstract
Texture information in images is coupled with geometric macrostructures and piecewise-smooth intensity variations. Decomposing an image f into a geometric structure component u and a texture component v is an inverse estimation problem, essential for understanding and analyzing images depending on their content. In this paper, we present a novel combined approach for simultaneous texture from structure separation and multiband texture modeling. First, we formulate a new, variational decomposition scheme, involving an explicit texture reconstruction constraint (prior) formed by the responses of selected frequency-tuned linear filters. This forms a dasiau + Kvpsila image model of K + 1 components. Subsequent texture modeling is applied to the estimated v component and its consistency is compared to using the complete, initial image f. The decomposition step, functioning as an advanced texture-front end, improves clustering and classification performance, for various multiband features. The proposed method can be generalized to other texture models or applications.
Georgios Evangelopoulos, Petros Maragos
CVPR2
2008 Adaptive and constrained algorithms for inverse compositional Active Appearance Model fitting
abstract
Parametric models of shape and texture such as active appearance models (AAMs) are diverse tools for deformable object appearance modeling and have found important applications in both image synthesis and analysis problems. Among the numerous algorithms that have been proposed for AAM fitting, those based on the inverse-compositional image alignment technique have recently received considerable attention due to their potential for high efficiency. However, existing fitting algorithms perform poorly when used in conjunction with models exhibiting significant appearance variation, such as AAMs trained on multiple-subject human face images. We introduce two enhancements to inverse-compositional AAM matching algorithms in order to overcome this limitation. First, we propose fitting algorithm adaptation, by means of (a) fitting matrix adjustment and (b) AAM mean template update. Second, we show how prior information can be incorporated and constrain the AAM fitting process. The inverse-compositional nature of the algorithm allows efficient implementation of these enhancements. Both techniques substantially improve AAM fitting performance, as demonstrated with experiments on publicly available multi-face datasets.
George Papandreou, Petros Maragos
CVPR2
2008 Audiovisual-to-articulatory speech inversion using Active Appearance Models for the face and Hidden Markov Models for the dynamics
abstract
We are interested in recovering aspects of vocal tract's geometry and dynamics from auditory and visual speech cues. We approach the problem in a statistical framework based on Hidden Markov Models and demonstrate effective estimation of the trajectories followed by certain points of interest in the speech production system. Alternative fusion schemes are investigated to account for asynchrony between the modalities and allow independent modeling of the dynamics of the involved streams. Visual cues are extracted from the speaker's face by means of active appearance modeling. We report experiments on the QSMT database which contains audio, video, and electromagnetic articulography data recorded in parallel. The results show that exploiting both audio and visual modalities in a multistream HMM based scheme clearly improves performance relative to either audio or visual-only estimation.
Athanasios Katsamanis, George Papandreou, Petros Maragos
ICASSP3
2008 Multisensor multiband cross-energy tracking for feature extraction and recognition
abstract
In this paper, we present a multisensor multiband energy tracking scheme for robust feature extraction in noisy environments. We introduce a multisensor feature extraction algorithm which combines both the spatial and frequency information incorporated in the speech signals captured by a microphone array. This is based on the estimation of cross-energies over multiple sensors and minimization of an error term due to noise. The relevant noise-analysis is given. Automatic speech recognition (ASR) experiments at various SNR levels demonstrate that the newly proposed frontend performs better than alternative schemes, especially in noisy conditions.
Stamatios Lefkimmiatis, Petros Maragos, Athanasios Katsamanis
ICASSP2
2008 Image inpainting with a wavelet domain Hidden Markov tree model
abstract
We present a novel technique for image inpainting, the problem of filling-in missing image parts. Image inpainting is ill-posed and we adopt a probabilistic model-based approach to regularize it. The main elements of our image model are, first, an over-complete complex-wavelet image representation, which ensures good shift invariance and directional selectivity and, second, a discrete-state/continuous-observation hidden Markov tree model for the wavelet coefficients, which captures key statistical properties of natural image wavelet responses, such as heavy-tailed histograms and persistence of large wavelet coefficients across scales. We show how these ideas can be integrated into a multi-scale generative process for natural images and present alternative deterministic and Markov chain Monte Carlo algorithms for image inpainting under this model. We demonstrate the effectiveness of the method in digitally restoring images of ancient wall-paintings.
George Papandreou, Petros Maragos, Anil C. Kokaram
ICASSP2
2008 Geodesic active regions for segmentation and tracking of human gestures in sign language videos
abstract
Reliable segmentation and motion tracking algorithms are required to achieve gesture detection and tracking for human-machine interaction. In this paper we present an efficient method for detecting and tracking moving hands in sign language video frames. We make use of the geodesic active region framework in conjunction with new color and motion forces; color information is provided by a skin color model, while motion information is derived from the optical flow field. Extensive experimentation indicates that the proposed algorithm behaves sufficiently well for gesture detection and tracking.
Olga Diamanti, Petros Maragos
ICIP2
2008 Texture modulation-constrained image decomposition
abstract
Texture modeling and separation of structure in images are treated in synergy. A variational image decomposition scheme is formulated using explicit texture reconstruction constraints from the outputs of linear filters tuned to different spatial frequencies and orientations. Relevant to the texture image part information is reconstructed using modulation modeling and component selection. The general formulation leads to a u + Kv model of K + 1 image components, with multiple texture subcomponents.
Georgios Evangelopoulos, Petros Maragos
ICIP2
2008 Movie summarization based on audiovisual saliency detection
abstract
Based on perceptual and computational attention modeling studies, we formulate measures of saliency for an audiovisual stream. Audio saliency is captured by signal modulations and related multi-frequency band features, extracted through nonlinear operators and energy tracking. Visual saliency is measured by means of a spatiotemporal attention model driven by various feature cues (intensity, color, motion). Audio and video curves are integrated in a single attention curve, where events may be enhanced, suppressed or vanished. The presence of salient events is signified on this audiovisual curve by geometrical features such as local extrema, sharp transition points and level sets. An audiovisual saliency-based movie summarization algorithm is proposed and evaluated. The algorithm is shown to perform very well in terms of summary informativeness and enjoyability for movie clips of various genres.
Georgios Evangelopoulos, Konstantinos Rapantzikos, Alexandros Potamianos, Petros Maragos, Nancy Zlatintsi, Yannis Avrithis
ICIP4
2008 Photon-limited image denoising by inference on multiscale models
abstract
We present an improved statistical model of Poisson processes, with applications to photon-limited imaging. We build on previous work, adopting a multiscale representation of the Poisson process in which the ratios of the underlying Poisson intensities (rates) in adjacent scales are modeled as mixtures of conjugate parametric distributions. Our main novel contributions are (1) a rigorous and robust regularized expectation-maximization (EM) algorithm for maximum-likelihood estimation of the rate-ratio density parameters directly from the observed Poisson data (counts); (2) extension of the method to work under a scale-recursive hidden Markov tree model (HMT) which couples the mixture label assignments in consecutive scales, thus modeling inter-scale coefficient dependencies in the vicinity of edges; and (3) exploration of a fully 2-D quad-tree image partitioning, involving Dirichlet-mixture rate-ratio densities, instead of the conventional separable binary image partitioning involving Beta-mixture rate-ratio densities. Experimental intensity estimation results on standard images with artificially simulated Poisson noise and photon-limited images with real shot noise demonstrate the effectiveness of the proposed approach.
Stamatios Lefkimmiatis, George Papandreou, Petros Maragos
ICIP3
2008 A PDE formulation for viscous morphological operators with extensions to intensity-adaptiveoperators
abstract
Viscous morphological operators have shown very good performance in regularizing various image analysis tasks such as detection of intensity-varying boundaries and segmentation. This paper presents a novel formulation of viscous morphological operators as solutions of nonlinear partial differential equations (PDEs) of the hyperbolic type with level-varying speed. Efficient numerical algorithms are also developed to solve these PDEs and generate the viscous operations. It also generalizes the viscous operators by studying the class of intensity level-varying operators, of which special cases are intensity adaptive connected operators such as volume openings and viscous reconstruction filters. We present both theoretical aspects and applications of the above ideas. Index Terms — Morphological operations, Adaptive filters, Partial differential equations. 1.
Petros Maragos, Corinne Vachier
ICIP1
2008 An inpainting system for automatic image structure - texture restoration with text removal
abstract
In this paper we deal with the inpainting problem and with the problem of finding text in images. We first review many of the methods used for structure and texture inpaintings. The novel contribution of the paper is the combination of the inpainting techniques with the techniques of finding text in images and a simple morphological algorithm that links them. This combination results in an automatic system for text removal and image restoration that requires no user interface at all. Examples on real images show very good performance of the proposed system and the importance of the new linking algorithm.
Eftychios A. Pnevmatikakis, Petros Maragos
ICIP2
2008 Computational analysis and learning for a biologically motivated model of boundary detection
Iasonas Kokkinos, Rachid Deriche, Olivier D. Faugeras, Petros Maragos
Neurocomputing4
2008 Audio-Assisted Movie Dialogue Detection
abstract
Abstract—An audio-assisted system is investigated that detects if a movie scene is a dialogue or not. The system is based on actor indicator functions. That is, functions which define if an actor speaks at a certain time instant. In particular, the cross-correlation and the magnitude of the corresponding the cross-power spectral density of a pair of indicator functions are input to various classifiers, such as voted perceptrons, radial basis function networks, random trees, and support vector machines for dialogue/non-dialogue detection. To boost classifier efficiency AdaBoost is also exploited. The aforementioned classifiers are trained using ground truth indicator functions determined by human annotators for 41 dialogue and another 20 non-dialogue audio instances. For testing, actual indicator functions are derived by applying audio activity detection and actor clustering to audio recordings. 23 instances are randomly chosen among the aforementioned 41 dialogue instances, 17 of which correspond to dialogue scenes and 6 to non-dialogue ones. Accuracy ranging between 0.739 and 0.826 is reported. Index Terms—Audio activity detection, cross-correlation, crosspower spectral density, dialogue detection, indicator functions, speaker clustering. I.
Margarita Kotti, Dimitrios Ververidis, Georgios Evangelopoulos, Yannis Panagakis, Constantine Kotropoulos, Petros Maragos, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.6
2008 Generalized Flooding and Multicue PDE-Based Image Segmentation
abstract
Image segmentation remains an important, but hard-to-solve, problem since it appears to be application dependent with usually no a priori information available regarding the image structure. Moreover, the increasing demands of image analysis tasks in terms of segmentation results' quality introduce the necessity of employing multiple cues for improving image segmentation results. In this paper, we attempt to incorporate cues such as intensity contrast, region size, and texture in the segmentation procedure and derive improved results compared to using individual cues separately. We emphasize on the overall segmentation procedure, and we propose efficient simplification operators and feature extraction schemes, capable of quantifying important characteristics, like geometrical complexity, rate of change in local contrast variations, and orientation, that eventually favor the final segmentation result. Based on the well-known morphological paradigm of watershed transform segmentation, which exploits intensity contrast and region size criteria, we investigate its partial differential equation (PDE) formulation, and we extend it in order to satisfy various flooding criteria, thus making it applicable to a wider range of images. Going a step further, we introduce a segmentation scheme that couples contrast criteria in flooding with texture information. The modeling of the proposed scheme is done via PDEs and the efficient incorporation of the available contrast and texture information, is done by selecting an appropriate cartoon-texture image decomposition scheme. The proposed coupled segmentation scheme is driven by two separate image components: cartoon U (for contrast information) and texture component V. The performance of the proposed segmentation scheme is demonstrated through a complete set of experimental results and substantiated using quantitative and qualitative criteria.
Anastasia Sofou, Petros Maragos
IEEE Trans. Image Process.2
2007 Multiband, multisensor robust features for noisy speech recognition
abstract
This paper presents a novel feature extraction scheme tak-ing advantage of both the nonlinear modulation speech model and the spatial diversity of speech and noise signals in a mul-tisensor environment. Herein, we propose applying robust fea-tures to speech signals captured by a multisensor array mini-mizing a noise energy criterion over multiple frequency bands. We show that we can achieve improved recognition perfor-mance by minimizing the Teager-Kaiser energy of the noise-corrupted signals in different frequency bands. These Multi-band, Multisensor Cepstral (MBSC) features are inspired by similar ones already been applied to single-microphone noisy Speech Recognition tasks with significantly improved results. The recognition results show that the proposed features can per-form better than the widely-used MFCC features.
Dimitrios Dimitriadis, Petros Maragos, Stamatios Lefkimmiatis
INTERSPEECH2
2007 Advanced front-end for robust speech recognition in extremely adverse environments
abstract
In this paper, a unified approach to speech enhancement, feature extraction and feature normalization for speech recognition in adverse recording conditions is presented. The proposed frontend system consists of several different, independent, processing modules. Each of the algorithms contained in these modules has been independently applied to the problem of speech recognition in noise, significantly improving the recognition rates. In this work, these algorithms are merged in a single front-end and their combined performance is demonstrated. Specifically, the proposed advanced front-end extracts noise-invariant features via the following modules: Wiener filtering, voice-activity detection, robust feature extraction (nonlinear modulation or fractal features), parameter equalization and frame-dropping. The advanced front-end is applied to extremely adverse environments where most feature extraction schemes fail. We show that by combining speech enhancement, robust feature extraction and feature normalization up to a fivefold error rate reduction can be achieved for certain tasks.
Dimitrios Dimitriadis, José C. Segura, Luz García 0001, Alexandros Potamianos, Petros Maragos, Vassilis Pitsikalis
INTERSPEECH5
2007 Audiovisual-to-Articulatory Speech Inversion Using HMMs
abstract
We address the problem of audiovisual speech inversion, namely recovering the vocal tract's geometry from auditory and visual speech cues. We approach the problem in a statistical framework, combining ideas from multistream Hidden Markov Models and canonical correlation analysis, and demonstrate effective estimation of the trajectories followed by certain points of interest in the speech production system. Our experiments show that exploiting both audio and visual modalities clearly improves performance relative to either audio-only or visual-only estimation. We report experiments on the QSMT database which contains audio, video, and electromagnetic articulography data recorded in parallel.
Athanasios Katsamanis, George Papandreou, Petros Maragos
MMSP3
2007 Multimodal Fusion and Learning with Uncertain Features Applied to Audiovisual Speech Recognition
abstract
We study the effect of uncertain feature measurements and show how classification and learning rules should be adjusted to compensate for it. Our approach is particularly fruitful in multimodal fusion scenarios, such as audio-visual speech recognition, where multiple streams of complementary features whose reliability is time-varying are integrated. For such applications, by taking the measurement noise uncertainty of each feature stream into account, the proposed framework leads to highly adaptive multimodal fusion rules for classification and learning which are widely applicable and easy to implement. We further show that previous multimodal fusion methods relying on stream weights fall under our scheme under certain assumptions; this provides novel insights into their applicability for various tasks and suggests new practical ways for estimating the stream weights adaptively. The potential of our approach is demonstrated in audio-visual speech recognition experiments.
George Papandreou, Athanasios Katsamanis, Vassilis Pitsikalis, Petros Maragos
MMSP4
2007 An Audio-Visual Saliency Model for Movie Summarization
abstract
A saliency-based method for generating video summaries is presented, which exploits coupled audiovisual information from both media streams. Efficient and advanced speech and image processing algorithms to detect key frames that are acoustically and visually salient are used. Promising results are shown from experiments on a movie database.
Konstantinos Rapantzikos, Georgios Evangelopoulos, Petros Maragos, Yannis Avrithis
MMSP3
2007 A generalized estimation approach for linear and nonlinear microphone array post-filters
Stamatios Lefkimmiatis, Petros Maragos
Speech Commun.2
2007 Multigrid Geometric Active Contour Models
abstract
Geometric active contour models are very popular partial differential equation-based tools in image analysis and computer vision. We present a new multigrid algorithm for the fast evolution of level-set-based geometric active contours and compare it with other established numerical schemes. We overcome the main bottleneck associated with most numerical implementations of geometric active contours, namely the need for very small time steps to avoid instability, by employing a very stable fully 2-D implicit-explicit time integration numerical scheme. The proposed scheme is more accurate and has improved rotational invariance properties compared with alternative split schemes, particularly when big time steps are utilized. We then apply properly designed multigrid methods to efficiently solve the occurring sparse linear system. The combined algorithm allows for the rapid evolution of the contour and convergence to its final configuration after very few iterations. Image segmentation experiments demonstrate the efficiency and accuracy of the method.
George Papandreou, Petros Maragos
IEEE Trans. Image Process.2
2006 Bottom-Up & Top-down Object Detection using Primal Sketch Features and Graphical Models
abstract
A combination of techniques that is becoming increasingly popular is the construction of part-based object representations using the outputs of interest-point detectors. Our contributions in this paper are twofold: first, we propose a primal-sketch-based set of image tokens that are used for object representation and detection. Second, top-down information is introduced based on an efficient method for the evaluation of the likelihood of hypothesized part locations. This allows us to use graphical model techniques to complement bottom-up detection, by proposing and finding the parts of the object that were missed by the front-end feature detection stage. Detection results for four object categories validate the merits of this joint top-down and bottom-up approach.
Iasonas Kokkinos, Petros Maragos, Alan L. Yuille
CVPR (2)2
2006 An optimum microphone array post-filter for speech applications
abstract
This paper proposes a post-filtering estimation scheme for mul-tichannel noise reduction. The proposed method extends and im-proves the existing Zelinski’s and, the most general and prominent, McCowan’s post-filtering methods that use the auto- and cross-spectral densities of the multichannel input signals to estimate the transfer function of the Wiener post-filter. A major drawback of these two speech enhancement algorithms is that the noise power spectrum at the beamformer’s output is over-estimated and there-fore the derived filters are sub-optimal in the Wiener sense. The proposed method deals with this problem and can be considered as an optimal post-filter that is appropriate for a wide variety of different noise fields. In experiments over real-noise multichannel recordings, the proposed technique is shown to obtain a significant headstart over the other methods in terms of signal-to-noise ratio and speech degradation measures. In addition it is used for ASR experiments where promising preliminary results are presented.
Stamatios Lefkimmiatis, Dimitrios Dimitriadis, Petros Maragos
INTERSPEECH3
2006 Adaptive multimodal fusion by uncertainty compensation
abstract
While the accuracy of feature measurements heavily depends on changing environmental conditions, studying the consequences of this fact in pattern recognition tasks has received relatively little attention to date. In this work we explicitly take into account feature measurement uncertainty and we show how classification rules should be adjusted to compensate for its effects. Our approach is particularly fruitful in multimodal fusion scenarios, such as audio-visual speech recognition, where multiple streams of complementary time-evolving features are integrated. For such applications, provided that the measurement noise uncertainty for each feature stream can be estimated, the proposed framework leads to highly adaptive multimodal fusion rules which are widely applicable and easy to implement. We further show that previous multimodal fusion methods relying on stream weights fall under our scheme under certain assumptions; this provides novel insights into their applicability for various tasks and suggests new practical ways for estimating the stream weights adaptively. The potential of our approach is demonstrated in audio-visual speech recognition using either synchronous or asynchronous models.
Vassilis Pitsikalis, Athanasios Katsamanis, George Papandreou, Petros Maragos
INTERSPEECH4
2006 Continuous energy demodulation methods and application to speech analysis
Dimitrios Dimitriadis, Petros Maragos
Speech Commun.2
2006 Filtered Dynamics and Fractal Dimensions for Noisy Speech Recognition
abstract
We explore methods from fractals and dynamical systems theory for robust processing and recognition of noisy speech. A speech signal is embedded in a multidimensional phase-space and is subsequently filtered exploiting aspects of its unfolded dynamics. Invariant measures (fractal dimensions) of the filtered signal are used as features in automatic speech recognition (ASR). We evaluate the new proposed features as well as the previously proposed multiscale fractal dimension via ASR experiments on the Aurora 2 database. The conducted experiments demonstrate relative improved word accuracy for the fractal features, especially at lower signal-to-noise ratio, when they are combined with the mel-frequency cepstral coefficients
Vassilis Pitsikalis, Petros Maragos
IEEE Signal Process. Lett.2
2006 Multiband Modulation Energy Tracking for Noisy Speech Detection
abstract
The ability to accurately locate the boundaries of speech activity is an important attribute of any modern speech recognition, processing, or transmission system. The effort in this paper is the development of efficient, sophisticated features for speech detection in noisy environments, using ideas and techniques from recent advances in speech modeling and analysis, like presence of modulations in speech formants, energy separation and multiband filtering. First we present a method, conceptually based on a classic speech-silence discrimination procedure, that uses some newly developed, short-time signal analysis tools and provide for it a detection theoretic motivation. The new energy and spectral content representations are derived through filtering the signal in various frequency bands, estimating the Teager-Kaiser energy for each and demodulating the most active one in order to derive the signal's dominant AM-FM components. This modulation approach demonstrated an improved robustness in noise over the classic algorithm, reaching an average error reduction of 33.5% under 5-30-dB noise. Second, by incorporating alternative modulation energy features in voice activity detection, improvement in overall misclassification error of a high hit rate detector reached 7.5% and 9.5% on different benchmarks
Georgios Evangelopoulos, Petros Maragos
IEEE Trans. Speech Audio Process.2
2005 A Cross-Validatory Statistical Approach to Scale Selection for Image Denoising by Nonlinear Diffusion
abstract
Scale-spaces induced by diffusion processes play an important role in many computer vision tasks. Automatically selecting the most appropriate scale for a particular problem is a central issue for the practical applicability of such scale-space techniques. This paper concentrates on automatic scale selection when nonlinear diffusion scale-spaces are utilized for image denoising. The problem is studied in a statistical model selection framework and cross-validation techniques are utilized to address it in a principled way. The proposed novel algorithms do not require knowledge of the noise variance and have acceptable computational cost. Extensive experiments on natural images show that the proposed methodology leads to robust algorithms, which outperform existing techniques for a wide range of noise types and noise levels.
George Papandreou, Petros Maragos
CVPR (1)2
2005 An Expectation Maximization Approach to the Synergy between Image Segmentation and Object Categorization
abstract
In this work, we deal with the problem of modelling and exploiting the interaction between the processes of image segmentation and object categorization. We propose a novel framework to address this problem that is based on the combination of the expectation maximization (EM) algorithm and generative models for object categories. Using a concise formulation of the interaction between these two processes, segmentation is interpreted as the E step, assigning observations to models, whereas object detection/analysis is modelled as the M-step, fitting models to observations. We present in detail the segmentation and detection processes comprising the E and M steps and demonstrate results on the joint detection and segmentation of the object categories of faces and cars.
Iasonas Kokkinos, Petros Maragos
ICCV2
2005 Image denoising in nonlinear scale-spaces: automatic scale selection via cross-validation
abstract
Multiscale, i.e. scale-space image analysis is a powerful framework for many image processing tasks. A fundamental issue with such scale-space techniques is the automatic selection of the most salient scale for a particular application. This paper considers optimal scale selection when nonlinear diffusion and morphological scale-spaces are utilized for image denoising. The problem is studied from a statistical model selection viewpoint and cross-validation techniques are utilized to address it in a principled way. The proposed novel algorithms do not require knowledge of the noise variance, have acceptable computational cost and are readily integrated with a wide class of scale-space inducing processes which require setting of a scale parameter. Our experiments show that this methodology leads to robust algorithms, which outperform existing scale-selection techniques for a wide range of noise types and noise levels.
George Papandreou, Petros Maragos
ICIP (1)2
2005 Coupled geometric and texture PDE-based segmentation
abstract
In this paper, along with recent trends in segmentation using multiple image cues, we examine the integration of modulation texture features, image contrast and region size for decomposition of an image in homogenous regions. First, we propose the use of a morphological PDE-based segmentation scheme of the watershed type, based on seeded region-growing and level curve evolution with speed depending on contrast and size. Second, we analyze object surface texture by modelling image variations as local spatial modulation components estimated via multi-frequency filtering and instantaneous energy-tracking operators. By separately exploiting contrast and texture information, through multiscale image decomposition, we propose a PDE-based coupled segmentation method. Experimental results on various classes of images such as soilsections, aerial and natural scenes indicate that the combined effect of image decomposition and multi-cue segmentation improves the overall segmentation process.
Anastasia Sofou, Georgios Evangelopoulos, Petros Maragos
ICIP (2)3
2005 Auditory Teager energy cepstrum coefficients for robust speech recognition
abstract
In this paper, a feature extraction algorithm for robust speech recognition is introduced. The feature extraction algorithm is motivated by the human auditory processing and the nonlinear Teager-Kaiser energy operator that estimates the true energy of the source of a resonance. The proposed features are labeled as Teager Energy Cepstrum Coefficients (TECCs). TECCs are computed by first filtering the speech signal through a dense non constant-Q Gammatone filterbank and then by estimating the "true" energy of the signal's source, i.e., the short-time average of the output of the Teager-Kaiser energy operator. Error analysis and speech recognition experiments show that the TECCs and the mel frequency cepstrum coefficients (MFCCs) perform similarly for clean recording conditions; while the TECCs perform significantly better than the MFCCs for noisy recognition tasks. Specifically, relative word error rate improvement of 60% over the MFCC baseline is shown for the Aurora-3 database for the high-mismatch condition. Absolute error rate improvement ranging from 5% to 20% is shown for a phone recognition task in (various types of additive) noise.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
INTERSPEECH2
2005 Speech event detection using multiband modulation energy
abstract
The need for efficient, sophisticated features for speech event detection is inherent in state of the art processing, enhancement and recognition systems. We explore ideas and techniques from non-linear speech modeling and analysis, like modulations and multiband filtering and propose new energy and spectral content features derived through filtering in multiple frequency bands and tracking dominant modulation energy in terms of the Teager-Kaiser Energy of separate AM-FM components. We present a detection-theoretic motivation and incorporate them in two detection schemes namely word boundary and voice activity detection. The modulation approach demonstrated noisy speech endpoint detection accuracy, reaching ∼40% error reduction on NTIMIT. In a voice activity scheme, improvement in overall misclassification error of a high hit-rate detector reached 7.5% on Aurora 2 and 9.5% on Aurora 3 databases.
Georgios Evangelopoulos, Petros Maragos
INTERSPEECH2
2005 Advances in statistical estimation and tracking of AM-FM speech components
abstract
In this paper we present two extensions of a statistical framework to demodulate speech resonances, which are modeled as AM-FM signals. The first approach utilizes bandpass filtering and a standard demodulation algorithm which regularizes instantaneous amplitude and frequency estimates. The second employs particle filtering techniques to allow temporal variations of the parameters that are connected with spectral characteristics of the analyzed signal. Results are presented on both synthetic and real speech signals and improved performance is demonstrated. Both approaches appear to cope quite satisfactorily with the nonstationarity of speech signals. 1.
Athanasios Katsamanis, Petros Maragos
INTERSPEECH2
2005 Soil image segmentation and texture analysis: a computer vision approach
abstract
Automated processing of digitized soilsection images reveals elements of soil structure and draws primary estimates of bioecological importance, like ground fertility and changes in terrestrial ecosystems. We examine a sophisticated integration of some modern methods from computer vision for image feature extraction, texture analysis, and segmentation into homogeneous regions, relevant to soil micromorphology. First, we propose the use of a morphological partial differential equation-based segmentation scheme based on seeded region-growing and level curve evolution with speed depending on image contrast. Second, we analyze surface texture information by modeling image variations as local modulation components and using multifrequency filtering and instantaneous nonlinear energy-tracking operators to estimate spatial modulation energy. By separately exploiting contrast and texture information, through multiscale image smoothing, we propose a joint image segmentation method for further interpretation of soil images and feature measurements. Our experimental results in images digitized under different specifications and scales demonstrate the efficacy of our proposed computational methods for soil structure analysis. We also briefly demonstrate their applicability to remote sensing images.
Anastasia Sofou, Georgios Evangelopoulos, Petros Maragos
IEEE Geosci. Remote. Sens. Lett.3
2005 Robust AM-FM Features for Speech Recognition
abstract
In this letter, a nonlinear AM-FM speech model is used to extract robust features for speech recognition. The proposed features measure the amount of amplitude and frequency modulation that exists in speech resonances and attempt to model aspects of the speech acoustic information that the commonly used linear source-filter model fails to capture. The robustness and discriminability of the AM-FM features is investigated in combination with mel cepstrum coefficients (MFCCs). It is shown that these hybrid features perform well in the presence of noise, both in terms of phoneme-discrimination (J-measure) and in terms of speech recognition performance in several different tasks. Average relative error rate reduction up to 11% for clean and 46% for mismatched noisy conditions is achieved when AM-FM features are combined with MFCCs.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
IEEE Signal Process. Lett.2
2005 Nonlinear speech analysis using models for chaotic systems
abstract
In this paper, we use concepts and methods from chaotic systems to model and analyze nonlinear dynamics in speech signals. The modeling is done not on the scalar speech signal, but on its reconstructed multidimensional attractor by embedding the scalar signal into a phase space. We have analyzed and compared a variety of nonlinear models for approximating the dynamics of complex systems using a small record of their observed output. These models include approximations based on global or local polynomials as well as approximations inspired from machine learning such as radial basis function networks, fuzzy-logic systems and support vector machines. Our focus has been on facilitating the application of the methods of chaotic signal analysis even when only a short time series is available, like phonemes in speech utterances. This introduced an increased degree of difficulty that was dealt with by resorting to sophisticated function approximation models that are appropriate for short data sets. Using these models enabled us to compute for short time series of speech sounds useful features like Lyapunov exponents that are used to assist in the characterization of chaotic systems. Several experimental insights are reported on the possible applications of such nonlinear models and features.
Iasonas Kokkinos, Petros Maragos
IEEE Trans. Speech Audio Process.2
2004 A Fast Multigrid Implicit Algorithm for the Evolution of Geodesic Active Contours
George Papandreou, Petros Maragos
CVPR (2)2
2004 A Biologically Motivated and Computationally Tractable Model of Low and Mid-Level Vision Tasks
Iasonas Kokkinos, Rachid Deriche, Petros Maragos, Olivier D. Faugeras
ECCV (2)3
2004 Modeling resonances with phase modulated self-similar processes [speech processing example]
abstract
In this paper, we propose a nonlinear model for time-varying random resonances where the instantaneous phase (and frequency) of a sinusoidal oscillation is allowed to vary proportionally to a random process that belongs to the class of /spl alpha/-stable self-similar stochastic processes. This is a general model that includes phase modulations by fractional Brownian motion or fractional stable Levy motion as special cases. We explore theoretically this random modulation model and derive analytically its autocorrelation and power spectrum. We also propose an algorithm to fit this model to arbitrary resonances with random phase modulation. Further, we apply the above ideas to some speech data and demonstrate that the model is suitable for fricative sounds.
Alexandros G. Dimakis, Petros Maragos
ICASSP (2)2
2004 Modulation-feature based textured image segmentation using curve evolution
abstract
In this paper we incorporate recent results from AM-FM models for texture analysis into the variational model of image segmentation and examine the potential benefits of using the combination of these two approaches for texture segmentation. Using the dominant components analysis (DCA) technique we obtain a low-dimensional, yet rich texture feature vector that proves to be useful for texture segmentation. We use an unsupervised scheme for texture segmentation, where only the number of regions is known a-priori. Experimental results on both synthetic and challenging real-world images demonstrate the potential of the proposed combination.
Iasonas Kokkinos, Georgios Evangelopoulos, Petros Maragos
ICIP3
2004 Advances in texture analysis-energy dominant component & multiple hypothesis testing
Iasonas Kokkinos, Georgios Evangelopoulos, Petros Maragos
ICIP3
2003 PDE-based modeling of image segmentation using volumic flooding
abstract
The classical case of morphological segmentation is based on the watershed transform, constructed by flooding the gradient image, which is seen as a topographic surface, with constant height speed. Changing the flooding criteria, (e.g. constant-speed height, area or volume) yields different segmentation results. In the field of PDEs and curve evolution the classic watershed transform can be modelled as the solution of an eikonal PDE. In this paper we model the watershed segmentation based on a volume flooding criterion via a different eikonal PDE. Then we solve this PDE using the fast marching method, which is a specific algorithm from the methodology of level sets. In addition, we attempt to exploit the advantages of image segmentation using PDE-based volume flooding over the classic height flooding.
Anastasia Sofou, Petros Maragos
ICIP (2)2
2003 Lattice Fuzzy Signal Operators and Generalized Image Gradients
Petros Maragos, Vassilis Tzouvaras, Giorgos B. Stamou
IFSA1
2003 Robust energy demodulation based on continuous models with application to speech recognition
abstract
In this paper, we develop improved schemes for simultaneous speech interpolation and demodulation based on continuous-time models. This leads to robust algorithms to estimate the instantaneous amplitudes and frequencies of the speech resonances and extract novel acoustic features for ASR. The continous-time models retain the excellent time resolution of the ESAs based on discrete energy operators and perform better in the presence of noise. We also introduce a robust algorithm based on the ESAs for amplitude compensation of the filtered signals. Furthermore, we use robust nonlinear modulation features to enhance the classic cepstrum-based features and use the augmented feature set for ASR applications. ASR experiments show promising evidence that the robust modulation features improve recognition. 1.
Dimitrios Dimitriadis, Petros Maragos
INTERSPEECH2
2003 Nonlinear analysis of speech signals: generalized dimensions and lyapunov exponents
abstract
In this paper, we explore modern methods and algorithms from fractal/chaotic systems theory for modeling speech signals in a multidimensional phase space and extracting characteristic invariant measures like generalized fractal dimensions and Lyapunov exponents. Such measures can capture valuable information for the characterisation of the multidimensional phase space- which is closer to the true dynamics- since they are sensitive to the frequency with which the attractor visits different regions and the rate of exponential divergence of nearby orbits, respectively. Further we examine the classification capability of related nonlinear features over broad phoneme classes. The results of these preliminary experiments indicate that the information carried by these novel nonlinear feature sets is important and useful. 1.
Vassilis Pitsikalis, Iasonas Kokkinos, Petros Maragos
INTERSPEECH3
2003 Algebraic and PDE Approaches for Lattice Scale-Spaces with Global Constraints
Petros Maragos
Int. J. Comput. Vis.1
2002 Modulation features for speech recognition
abstract
Automatic speech recognition (ASR) systems can benefit from including into their acoustic processing part new features that account for various nonlinear and time-varying phenomena during speech production. In this paper, we develop robust methods to extract novel acoustic features from speech signals of the modulation type based on time-varying models for speech analysis. Further, we integrate the new speech features with the standard linear ones (mel-frequency cesptrum) to develop a augmented set of acoustic features and demonstrate its efficacy by showing significant improvements in HMM-based word recognition over the TIMIT database.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
ICASSP2
2002 Speech analysis and feature extraction using chaotic models
abstract
Nonlinear systems based on chaos theory can model various aspects of the nonlinear dynamic phenomena occuring during speech production. In this paper, we explore modem methods and algorithms from chaotic systems theory for modeling speech signals in a multidimensional phase space and for extracting nonlinear acoustic features. Further, we integrate these chaotic-type features with the standard linear ones (based on cepstrum) to develop a generalized hybrid set of short-time acoustic features for speech signals and demonstrate its efficacy by showing significant improvements in HMM-based word recognition.
Vassilis Pitsikalis, Petros Maragos
ICASSP2
2001 An improved energy demodulation algorithm using splines
abstract
A new algorithm is proposed for demodulating discrete-time AM-FM signals, which first interpolates the signals with smooth splines and then uses the continuous-time energy separation algorithm (ESA) based on the Teager-Kaiser energy operator. This spline-based ESA retains the excellent time resolution of the ESA based on discrete energy operators but performs better in the presence of noise. Further, its dependence on smooth splines allows some optimal trade-off between data fitting versus smoothing.
Dimitrios Dimitriadis, Petros Maragos
ICASSP2
2001 Generalized multiscale connected operators with applications to granulometric image analysis
abstract
In this paper, generalized granulometric size distributions and size histograms (a.k.a 'pattern spectra') are developed using generalized multiscale lattice operators of the opening and closing type. The generalized size histograms are applied to granulometric analysis of soil-section images. An interesting structure is obtained when the histogram is based on area openings. Furthermore, a fast implementation of the generalized size histograms is presented using threshold analysis-synthesis. Comparisons with size distributions based on conventional morphological operators indicate that the generalized histograms provide a more direct and informative description of the image content in objects with scale-dependent geometric attributes. Applications are also developed for studying the structure of soilsection images.
Anastasios Doulamis, Nikolaos D. Doulamis, Petros Maragos
ICIP (3)3
2001 Synthesis and applications of lattice image operators based on fuzzy norms
abstract
We use concepts from the lattice-based theory of morphological operators and fuzzy sets to develop generalized lattice image operators that can be expressed as nonlinear convolutions that are suprema or infima of fuzzy intersection or union norms. Our emphasis (different from previous works) is the construction of pairs of fuzzy dilation and erosion operators that form lattice adjunctions. This guarantees that their composition will be a valid algebraic opening or closing. The power, but also the difficulty, in applying these fuzzy operators to image analysis is the large variety of fuzzy norms and the absence of systematic ways in selecting them. Towards this goal, we have performed extensive experiments in applying these fuzzy operators to various nonlinear filtering and image analysis tasks, attempting first to understand the effect that the type of fuzzy norm and the shape/size of the structuring function has on the resulting new image operators. Further, we have developed some new fuzzy edge gradients and optimized their usage for edge detection on test problems via a parametric fuzzy norm.
Petros Maragos, Vassilis Tzouvaras, Giorgos B. Stamou
ICIP (1)1
2001 Segmentation of soil section images using connected operators
abstract
Segmentation of soil section images is an important task for automating the measurement of the grains' properties as well as for detecting and recognizing objects in the soil, important for its bioecological quality. We apply several types of morphological systems to watershed-based segmentation of soil section images. We use efficient connected operators such as reconstruction open-closing and area open-closing as well as some relatively new operators, the levelings, for image denoising, simplification and feature/marker extraction. Further, we introduce an improvement of the reconstruction operators used in segmentation, based on a generalized multiscale connectivity analysis.
Anastasia Sofou, Costas S. Tzafestas, Petros Maragos
ICIP (3)3
2001 Time-frequency distributions for automatic speech recognition
abstract
The use of general time-frequency distributions as features for automatic speech recognition (ASR) is discussed in the context of hidden Markov classifiers. Short-time averages of quadratic operators, e.g., energy spectrum, generalized first spectral moments, and short-time averages of the instantaneous frequency, are compared to the standard front end features, and applied to ASR. Theoretical and experimental results indicate a close relationship among these feature sets.
Alexandros Potamianos, Petros Maragos
IEEE Trans. Speech Audio Process.2
2000 A PDE Approach to Nonlinear Image Simplification via Levelings and Reconstruction Filters
abstract
We present a nonlinear partial differential equation (PDE) that models the generation of a large class of advanced morphological filters, the levelings and the openings/closings by reconstruction. These types of filters are very useful in numerous image analysis and vision tasks ranging from enhancement, feature detection, image simplification, to segmentation. The developed PDE models these nonlinear filters as the limit of a controlled growth starting from an initial seed signal. This growth is of the multiscale dilation or erosion type and the controlling mechanism is a switch that reverses the growth when the difference between the current evolution and a reference signal switches signs. We discuss theoretical aspects of this PDE, propose a discrete algorithm for its numerical solution and corresponding filter implementation, and provide insights via several experiments. Finally, we outline its use for improving the Gaussian scale-space by using the latter as initial seed to generate multiscale levelings that have a superior preservation of image edges and boundaries.
Petros Maragos, Fernand Meyer
ICIP1
2000 Curve Evolution, Differential Morphology, and Distance Transforms Applied to Multiscale and Eikonal Problems
abstract
In differential morphology, multiscale dilations and erosions are modeled via nonlinear partial differential equations (PDEs) in scale-space. Curve evolution employs methods of differential geometry to study the differential equations governing the p
Petros Maragos, Muhammad Akmal Butt
Fundam. Informaticae1
2000 Nonlinear Scale-Space Representation with Morphological Levelings
Fernand Meyer, Petros Maragos
J. Vis. Commun. Image Represent.2
2000 Neural networks with hybrid morphological/rank/linear nodes: a unifying framework with applications to handwritten character recognition
Lúcio F. C. Pessoa, Petros Maragos
Pattern Recognit.2
2000 Multicomponent AM-FM demodulation via periodicity-based algebraic separation and energy-based demodulation
abstract
Previously investigated multicomponent AM-FM demodulation techniques either assume that the individual component signals are spectrally isolated from each other or that the components can be isolated by linear time-invariant filtering techniques and, consequently, break down in the case where the components overlap spectrally or when one of the components is stronger than the other. In this paper, we present a nonlinear algorithm for the separation and demodulation of discrete-time multicomponent AM-FM signals. Our approach divides the demodulation problem into two independent tasks: algebraic separation of the components based on periodicity assumptions and then monocomponent demodulation of each component by instantaneously tracking and separating its source energy into its amplitude and frequency parts. The proposed new algorithm avoids the shortcomings of previous approaches and works well for extremely small spectral separations of the components and for a wide range of relative amplitude/power ratios. We present its theoretical analysis and experimental results and outline its application to demodulation of cochannel FM voice signals.
Balasubramaniam Santhanam, Petros Maragos
IEEE Trans. Commun.2
1999 Speech analysis and synthesis using an AM-FM modulation model
Alexandros Potamianos, Petros Maragos
Speech Commun.2
1998 Harmonic analysis and restoration of separation methods for periodic signal mixtures: Algebraic separation versus comb filtering
Balasubramaniam Santhanam, Petros Maragos
Signal Process.2
1998 Optimum design of chamfer distance transforms
abstract
The distance transform has found many applications in image analysis. Chamfer distance transforms are a class of discrete algorithms that offer a good approximation to the desired Euclidean distance transform at a lower computational cost. They can also give integer-valued distances that are more suitable for several digital image processing tasks. The local distances used to compute a chamfer distance transform are selected to minimize an approximation error. A new geometric approach is developed to find optimal local distances. This new approach is easier to visualize than the approaches found in previous work, and can be easily extended to chamfer metrics that use large neighborhoods. A new concept of critical local distances is presented which reduces the computational complexity of the chamfer distance transform without increasing the maximum approximation error.
Muhammad Akmal Butt, Petros Maragos
IEEE Trans. Image Process.2
1998 MRL-filters: a general class of nonlinear systems and their optimal design for image processing
abstract
A class of morphological/rank/linear (MRL)-filters is presented as a general nonlinear tool for image processing. They consist of a linear combination between a morphological/rank filter and a linear filter. A gradient steepest descent method is proposed to optimally design these filters, using the averaged least mean squares (LMS) algorithm. The filter design is viewed as a learning process, and convergence issues are theoretically and experimentally investigated. A systematic approach is proposed to overcome the problem of nondifferentiability of the nonlinear filter component and to improve the numerical robustness of the training algorithm, which results in simple training equations. Image processing applications in system identification and image restoration are also presented, illustrating the simplicity of training MRL-filters and their effectiveness for image/signal processing.
Lúcio F. C. Pessoa, Petros Maragos
IEEE Trans. Image Process.2
1997 Demodulation of discrete multicomponent AM-FM signals using periodic algebraic separation and energy demodulation
abstract
Existing multicomponent AM-FM demodulation algorithms either assume spectrally distinct components or components separable via linear filtering and break down when the components overlap spectrally or if one of the components is stronger than the other. In this paper, we present a nonlinear algorithm for multicomponent AM-FM demodulation which avoids the above shortcomings and works well even for extremely small spectral separation of the components. The proposed algorithm separates the multicomponent demodulation problem into two tasks: periodicity-based algebraic separation of the components and then monocomponent demodulation via energy-based methods.
Balasubramaniam Santhanam, Petros Maragos
ICASSP2
1997 On using fractal features of speech sounds in automatic speech recognition
Petros Maragos, Alexandros Potamianos
EUROSPEECH1
1997 Speech analysis and synthesis using an AM-FM modulation model
abstract
In this paper, the AM‐FM modulation model is applied to speech analysis, synthesis and coding. The AM‐FM model represents the speech signal as the sum of formant resonance signals each of which contains amplitude and frequency modulation. Multiband filtering and demodulation using the energy separation algorithm are the basic tools used for speech analysis. First, multiband demodulation analysis (MDA) is applied to the problem of fundamental frequency estimation using the average instantaneous frequency as estimates of pitch harmonics. The MDA pitch tracking algorithm is shown to produce smooth and accurate fundamental frequency contours. Next, the AM‐FM modulation vocoder is introduced, which represents speech as the sum of resonance signals. A time-varying filterbank is used to extract the formant bands and then the energy separation algorithm is used to demodulate the resonance signals into the amplitude envelope and instantaneous frequency signals. EAcient modeling and coding (at 4.8‐9.6 kbits/sec) algorithms are proposed for the amplitude envelope and instantaneous frequency of speech resonances. Finally, the perceptual importance of modulations in speech resonances is investigated and it is shown that amplitude modulation patterns are both speaker and phone dependent. ” 1999 Elsevier Science B.V. All rights reserved. Zusammenfassung
Alexandros Potamianos, Petros Maragos
EUROSPEECH2
1997 Lattice calculus of the morphological slope transform
Henk J. A. M. Heijmans, Petros Maragos
Signal Process.2
1996 Morphological/rank neural networks and their adaptive optimal design for image processing
abstract
We formulate a general class of neural network based filters, where each node is a morphological/rank operation. This type of system is computationally efficient since no multiplications are necessary. The introduction of such networks is partially motivated from observations that internal structures of a neuron can generate logic operations. An efficient adaptive optimal design procedure is proposed for these networks, based on the back-propagation algorithm. The procedure is optimal under the LMS criterion. Finally, experimental results are illustrated in problems of noise cancellation, encouraging the use of such class of systems and its training algorithm as important tools for nonlinear signal and image processing.
Lúcio F. C. Pessoa, Petros Maragos
ICASSP2
1996 Energy demodulation of two-component AM-FM signals with application to speaker separation
abstract
In this paper, an efficient low-complexity algorithm is presented for the separation and demodulation of two-component AM-FM signals and applied to the separation of voice-modulated FM signals. The proposed algorithm is based on the generating differential equation of the mixture signal and nonlinear differential energy operators.
Balasubramaniam Santhanam, Petros Maragos
ICASSP2
1996 Partial differential equations in image analysis: continuous modeling, discrete processing
abstract
Most traditional methodologies in digital image processing start with a discrete model which is directly related to the subsequent discrete processing of the image data. We review some examples and focus on some details of an emerging new methodology that starts from some continuous models provided by PDEs and proceeds with discrete processing of the image data via the numerical implementation of these PDEs on some discrete grid. We examine the PDEs modeling multiscale morphological image analysis and the related eikonal PDE of optics. The PDE approach is very promising because it provides many new mathematical models and has connections with the physics of imaging.
Petros Maragos, Muhammad Akmal Butt
ICIP (3)1
1996 Energy demodulation of two-component AM-FM signal mixtures
abstract
An algorithm for the separation and energy-based demodulation of two-component mixtures of AM-FM signals is presented. The proposed algorithm is based on the generating differential or difference equation of the mixture signal and nonlinear differential energy operators.
Balasubramaniam Santhanam, Petros Maragos
IEEE Signal Process. Lett.2
1996 Guest Editorial Introduction to the Special Issue on Nonlinear Image Processing
Gonzalo R. Arce, Petros Maragos, Yrjö Neuvo, Ioannis Pitas
IEEE Trans. Image Process.2
1996 Differential morphology and image processing
abstract
Image processing via mathematical morphology has traditionally used geometry to intuitively understand morphological signal operators and set or lattice algebra to analyze them in the space domain. We provide a unified view and analytic tools for morphological image processing that is based on ideas from differential calculus and dynamical systems. This includes ideas on using partial differential or difference equations (PDEs) to model distance propagation or nonlinear multiscale processes in images. We briefly review some nonlinear difference equations that implement discrete distance transforms and relate them to numerical solutions of the eikonal equation of optics. We also review some nonlinear PDEs that model the evolution of multiscale morphological operators and use morphological derivatives. Among the new ideas presented, we develop some general 2-D max/min-sum difference equations that model the space dynamics of 2-D morphological systems (including the distance computations) and some nonlinear signal transforms, called slope transforms, that can analyze these systems in a transform domain in ways conceptually similar to the application of Fourier transforms to linear systems. Thus, distance transforms are shown to be bandpass slope filters. We view the analysis of the multiscale morphological PDEs and of the eikonal PDE solved via weighted distance transforms as a unified area in nonlinear image processing, which we call differential morphology, and briefly discuss its potential applications to image processing and computer vision.
Petros Maragos
IEEE Trans. Image Process.1
1995 Speech formant frequency and bandwidth tracking using multiband energy demodulation
abstract
The AM-FM modulation model and a multiband analysis/demodulation scheme is applied to speech formant frequency and bandwidth tracking. Filtering is performed by a bank of Gabor bandpass filters. Each band is demodulated to amplitude envelope and instantaneous frequency signals using the energy separation algorithm. Short-time formant frequency and bandwidth estimates are obtained from the instantaneous amplitude and frequency signals and their merits are presented. The estimates are used to determine the formant locations and bandwidths. Performance and computational issues (frequency domain implementation) are discussed. Overall, the multiband demodulation approach to formant tracking is easy to implement, provides accurate formant frequency and realistic bandwidth estimates, and performs well in the presence of nasalization.
Alexandros Potamianos, Petros Maragos
ICASSP2
1995 Min-max classifiers: Learnability, design and application
Ping-Fai Yang, Petros Maragos
Pattern Recognit.2
1995 Higher order differential energy operators
abstract
Instantaneous signal operators /spl Upsi//sub k/(x)=x/spl dot/x/sup (k-1)/-xx/sup (k)/ of integer orders k are proposed to measure the cross energy between a signal x and its derivatives. These higher order differential energy operators contain as a special case, for k=2, the Teager-Kaiser (1990) operator. When applied to (possibly modulated) sinusoids, they yield several new energy measurements useful for parameter estimation or AM-FM demodulation. Applying them to sampled signals involves replacing derivatives with differences that lead to several useful discrete energy operators defined on an extremely short window of samples.>
Petros Maragos, Alexandros Potamianos
IEEE Signal Process. Lett.1
1994 Differential Morphology: Multiscale Image Dynamics, Max-Min Difference Equations, and Slope Transforms
abstract
We develop 2D max-min difference equations that model the space dynamics of 2D morphological systems and some nonlinear signal transforms, called slope transforms, that can analyze these systems in a transform domain in a way conceptually analogous to the application of Fourier transforms to linear systems. Further, we discuss some nonlinear partial differential equations (PDEs) that model the evolution of multiscale morphological filters of the max-min type. These PDEs are related to the eikonal equation. Solutions of the eikonal PDE are proposed based on 2D min difference equations that can compute distance transforms. We view the analysis of the multiscale morphological PDEs and of the eikonal PDE solved via max-min equations as a unified area in nonlinear image processing which we call differential morphology. Its potential applications include distance-path finding, segmentation, gridless halftoning, and shape from shading.>
Petros Maragos
ICIP (2)1
1994 Demodulation of Images Modeled by Amplitude-Frequency Modulations Using Multidimensional Energy Separation
abstract
Locally narrowband images can be modeled as 2D spatial AM-FM signals with several applications in image texture analysis and computer vision. In this paper we formulate such an image demodulation problem, and present a solution based on the multidimensional energy operator /spl Phi/(f)=/spl par//spl nabla/f/spl par//sup 2/-f/spl nabla//sup 2/f. We discuss some interesting properties of this multidimensional operator and develop multidimensional energy separation algorithms to estimate the amplitude envelope and instantaneous frequencies of 2D spatially-varying AM-FM signals. Experiments are also presented on applying this 2D energy demodulation algorithm to estimate the instantaneous amplitude contrast and spatial frequencies of image textures bandpass filtered via Gabor filters. The attractive features of the multidimensional energy operator and the 2D energy separation algorithm are their simplicity, efficiency, and ability to track instantaneously-varying spatial modulation patterns.>
Petros Maragos, Alan C. Bovik
ICIP (3)1
1994 Morphological systems: Slope transforms and max-min difference and differential equations
Petros Maragos
Signal Process.1
1994 A comparison of the energy operator and the Hilbert transform approach to signal and speech demodulation
Alexandros Potamianos, Petros Maragos
Signal Process.2
1994 A system for finding speech formants and modulations via energy separation
abstract
This correspondence presents an experimental system that uses an energy-tracking operator and a related energy separation algorithm to automatically find speech formants and amplitude/frequency modulations in voiced speech segments. Initial estimates of formant center frequencies are provided by either LPC or morphological spectral peak picking. These estimates are then shown to be improved by a combination of bandpass filtering and iterative application of energy separation.>
Helen M. Hanson, Petros Maragos, Alexandros Potamianos
IEEE Trans. Speech Audio Process.2
1993 Finding speech formants and modulations via energy separation: with application to a vocoder
Helen M. Hanson, Petros Maragos, Alexandros Potamianos
ICASSP (2)2
1993 Morphological systems for character image processing and recognition
Ping-Fai Yang, Petros Maragos
ICASSP (5)2
1992 Evolution equations for continuous-scale morphology
abstract
Several nonlinear partial differential equations that model the scale evolution associated with continuous-space multiscale morphological erosions, dilations, openings, and closings are discussed. These systems relate the infinitesimal evolution of the multiscale signal ensemble in scale space to a nonlinear operator acting on the space of signals. The type of this nonlinear operator is determined by the shape and dimensionality of the structuring element used by the morphological operators, generally taking the form of nonlinear algebraic functions of certain differential operators.>
Roger W. Brockett, Petros Maragos
ICASSP2
1992 On separating amplitude from frequency modulations using energy operators
abstract
To estimate the amplitude envelope and instantaneous frequency of an AM-FM signal the authors developed a novel approach that uses nonlinear combinations of instantaneous signal outputs from an energy-tracking operator to separate its output energy product into its amplitude modulation and frequency modulation components. This energy separation algorithm is then applied to search for modulations in speech resonances, which the authors model using AM-FM signals. The theoretical and experimental results demonstrate that the energy separation algorithm, due to its low computational complexity and instantaneously adapting nature, is very useful in detecting modulation patterns in speech and other time-varying signals.>
Petros Maragos, James F. Kaiser, Thomas F. Quatieri
ICASSP1
1991 Affine models for image matching and motion detection
abstract
A model is developed for detecting the displacement field in spatiotemporal image sequences. It allows for affine shape deformations of corresponding spatial regions and for affine transformations of the image intensity range. The model includes the block matching method as a special case. A least-squares algorithm is used to find the model parameters. It is experimentally demonstrated that the affine matching model performs better than other standard approaches. The resulting 2-D motion estimates are then used by a 3-D affine model and a least-squares algorithm that recover 3-D rigid body motion and depth from two perspective views.>
Chiou-Shann Fuh, Petros Maragos
ICASSP2
1991 Fractal aspects of speech signals: dimension and interpolation
abstract
The nonlinear dynamics of air flow during speech production may often result in some small or large degree of turbulence. The author quantifies the geometry of speech turbulence, as reflected in the fragmentation of the time signal, by using fractal models. He describes an efficient algorithm for estimating the short-time fractal dimension of speech segmentation and sound classification. He also develops a method for fractal speech interpolation which can be used to synthesize controlled amounts of turbulence in speech or to increase its sampling rate by preserving not its bandwidth (as is classically done) but rather its fractal dimension.>
Petros Maragos
ICASSP1
1991 Speech nonlinearities, modulations, and energy operators
abstract
An AM-FM model for representing modulations in speech resonances is investigated. Specifically, an FM model is proposed for the time-varying formants whose amplitude varies as the envelope of an AM signal. To detect the modulations the energy operator Psi ( chi ) and its discrete counterpart are applied. It is found that Psi can approximately track the envelope of AM signals, the instantaneous frequency of FM signals, and the product of these two functions in the general case of AM-FM signals. Several experiments on the application of this AM-FM modeling to speech signals, band pass filtered via Gabor filters are reported.>
Petros Maragos, Thomas F. Quatieri, James F. Kaiser
ICASSP1
1990 Fractal excitation signals for CELP speech coders
abstract
A class of random signals is presented as a new excitation codebook for stochastic predictive speech coders. These signals, known as fractional noises, depend on a single parameter, 0<H<1, that controls the low- or high-pass trend of their power spectrum. Preliminary experiments have shown their performance improves when H is limited to a subinterval of
Petros Maragos, Kenneth L. Young
ICASSP1
1990 Threshold Superposition in Morphological Image Analysis Systems
abstract
It is shown that four composite morphological systems, namely morphological edge detection, peak/valley extraction, skeletonization, and shape-size distributions obey a weak linear superposition, called threshold-linear superposition. The output image signal or measurement from each system is shown to be the sum of outputs due to input binary images that result from thresholding the input gray-level image at all levels. These results are generalized to a vector space formulation, e.g. to any finite linear combination of simple morphological systems. Thus many such systems processing gray-level images are reduced to corresponding binary image processing systems, which are easier to analyze and implement.>
Petros Maragos, Robert D. Ziff
IEEE Trans. Pattern Anal. Mach. Intell.1
1990 Morphological systems for multidimensional signal processing
abstract
The basic theory and applications of a set-theoretic approach to image analysis called mathematical morphology are reviewed. The goals are to show how the concepts of mathematical morphology geometrical structure in signals to illuminate the ways that morphological systems can enrich the theory and applications of multidimensional signal processing. The topics covered include: applications to nonlinear filtering (morphological and rank-order filters, multiscale smoothing, morphological sampling, and morphological correlation); applications to image analysis (feature extraction, shape representation and description, size distributions, and fractals); and representation theorems, which shows how a large class of nonlinear and linear signal operators can be realized as a combination of simple morphological operations.
Petros Maragos, Ronald W. Schafer
Proc. IEEE1
1989 Region-based optical flow estimation
abstract
A correspondence method is developed for determining optical flow where the primitive motion tokens to be matched between consecutive time frames are regions. The computation of optical flow consists of three stages: region extraction, region matching, and optical flow smoothing. The computation is completed by smoothing the initial optical flow, where the sparse velocity data are either smoothed with a vector median filter or interpolated to obtain dense velocity estimates by using a motion-coherence regularization. The proposed region-based method for optical flow is simple, computationally efficient, and more robust than iterative gradient methods, especially for medium-range motion.>
Chiou-Shann Fuh, Petros Maragos
CVPR2
1989 Morphological correlation and mean absolute error criteria
abstract
The mean absolute error criterion for signal detection and matching is linked with a morphological signal correlation (a sum of minima). Several properties of this nonlinear correlation are investigated, its performance for signal detection is compared with that of the classical (sum of products) linear correlation, and its statistical form is calculated for speckle patterns.>
Petros Maragos
ICASSP1
1989 A Representation Theory for Morphological Image and Signal Processing
abstract
A unifying theory for many concepts and operations encountered in or related to morphological image and signal analysis is presented. The unification requires a set-theoretic methodology, where signals are modeled as sets, systems (signal transformations) are viewed as set mappings, and translational-invariant systems are uniquely characterized by special collections of input signals. This approach leads to a general representation theory, in which any translation-invariant, increasing, upper semicontinuous system can be presented exactly as a minimal nonlinear superposition of morphological erosions or dilations. The theory is used to analyze some special cases of image/signal analysis systems, such as morphological filters, median and order-statistic filters, linear filters, and shape recognition transforms. Although the developed theory is algebraic, its prototype operations are well suited for shape analysis; hence, the results also apply to systems that extract information about the geometrical structure of signals.>
Petros Maragos
IEEE Trans. Pattern Anal. Mach. Intell.1
1989 Pattern Spectrum and Multiscale Shape Representation
abstract
The results of a study on multiscale shape description, smoothing and representation are reported. Multiscale nonlinear smoothing filters are first developed, using morphological opening and closings. G. Matheron (1975) used openings and closings to obtain probabilistic size distributions of Euclidean-space sets (continuous binary images). These distributions are used to develop a concept of pattern spectrum (a shape-size descriptor). A pattern spectrum is introduced for continuous graytone images and arbitrary multilevel signals, as well as for discrete images, by developing a discrete-size family of patterns. Large jumps in the pattern spectrum at a certain scale indicate the existence of major (protruding or intruding) substructures of the signal at the scale. An entropy-like shape-size complexity measure is also developed based on the pattern spectrum. For shape representation, a reduced morphological skeleton transform is introduced for discrete binary and graytone images. This transform is a sequence of skeleton components (sparse images) which represent the original shape at various scales. It is shown that the partially reconstructed images from the inverse transform on subsequences of skeleton components are the openings of the image at a scale determined by the number of eliminated components; in addition, two-way correspondences are established among the degree of shape smoothing via multiscale openings or closings, the pattern spectrum zero values, and the elimination or nonexistence of skeleton components at certain scales.>
Petros Maragos
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 Morphology-based symbolic image modeling, multi-scale nonlinear smoothing, and pattern spectrum
abstract
The author develops a symbolic modeling of images based on their shape-size information. First, multiscale multishape structural distributions in the image are modeled by morphological openings, and a related shape-size descriptor, the pattern spectrum, is developed that can detect critical scales. Then the image is modeled as a nonlinear superposition of simpler parts (the symbols), which are translated and scaled shape patterns drawn from a finite collection. The model parameters are found by using the information from openings and pattern spectrum, and by local searches at points of generalized skeletons. The results appear promising for multiscale image analysis and shape recognition.>
Petros Maragos
CVPR1
1988 Symbolic signal representation using nonlinear filtering
abstract
A symbolic modeling method is developed for images and imagelike signals based on their shape-size structure. A signal is represented as a minimal nonlinear superposition of simpler parts, namely, the symbols, which are translated and scaled shape patterns drawn from a finite collection. To obtain the model parameters, multiscale morphological openings and the related pattern spectrum are used to help decide which shape patterns exist in the signal and at which scales. Then the pattern locations are found via a local search at points of generalized morphological skeletons. This model is useful for signal abstractions and object recognition.>
Petros Maragos
ICASSP1
1988 Optimal Morphological Approaches To Image Matching And Object Detection
abstract
It is shown that minimizing the C1 match- ing error is equivalent to maximizing a nonlinear signal cor- relation (a sum of minima), which is related to morpholog- ical filtering. This approach is optimum in rnany formula- tions of the image matching and object detection problem. Further, a closely related approach is outlined for object detection using rank order filters.
Petros Maragos
ICCV1
1987 Pattern spectrum of images and morphological shape-size complexity
abstract
By using morphological opening and closing set operations, a pattern spectrum of binary images can be developed, which measures the image content relative to patterns of arbitrary shape and size. In this paper, the pattern spectrum of discrete images is further generalized, is directly related to the skeleton (medial axis) transform, is used to derive a shape-size complexity measure of the image and its immediate background, and all these concepts are extended to graytone images. The results of this study indicate that the pattern spectrum can be used to analyze and enrich the medial axis (or other skeleton-like) transforms, and quantify the roughness of the image boundary or surface.
Petros Maragos
ICASSP1
1986 Applications of morphological filtering to image analysis and processing
abstract
This paper summarizes some applications of morphological erosions, dilations, openings, and closings to image edge-detection, cleaning of impulsive-noise, median-filtering, region-filling, skeletonization, and 2-D shape recognition.
Petros Maragos, Ronald W. Schafer
ICASSP1
1985 A unification of linear, median, order-statistics and morphological filters under mathematical morphology
abstract
This paper presents a summary of a unified theory, based upon mathematical morphology, of all translation-invariant and increasing systems. Examples of such systems are morphological transformations of signals, order-statistics filters and some linear shift-invariant filters. Our theoretical research showed that every such system can be uniquely represented by the minimal elements of its kernel and realized as a minimal combination of morphological erosions.
Petros Maragos, Ronald W. Schafer
ICASSP1
1984 Multichannel linear predictive coding of color images
abstract
This paper reports on a preliminary study of applying single-channel (scalar) and multichannel (vector) 2-D linear prediction to color image modeling and coding. Also, the novel idea of a multi-input single-output 2-D ADPCM coder is introduced. The results of this study indicate that texture information in multispectral images can be represented by linear prediction coefficients or matrices, where as the prediction error conveys edge-information. Moreover, by using a single-channel edge-information we obtained, from original color images of 24 bits/pixel, reconstructed images of good quality at information rates of 1 bit/pixel or less.
Petros Maragos, Russell M. Mersereau, Ronald W. Schafer
ICASSP1
1984 Morphological skeleton representation and coding of binary images
abstract
This paper presents a preliminary study on using Mathematical Morphology to represent and code a binary or a grey-tone image by parts of its skeleton, a thinned version of the image. An image can be uniquely decomposed into skeleton components, and then reconstructed by dilating these components. Since, for a certain category of imagery, the skeleton components possess a lower entropy than the original image, a run-length or entropy coding scheme can be used to achieve representation or transmission of the image at a lower information rate than originally required.
Petros Maragos, Ronald W. Schafer
ICASSP1
1983 Two-dimensional linear predictive analysis of arbitrarily-shaped regions
abstract
This paper is concerned with the use of 2-D linear prediction for image segmentation. It begins with a brief summary of the mathematics involved in 2-D linear predictive analysis of arbitrarily-shaped regions. Then, it introduces a 2-D LPC distance measure based on the error residual of 2-D linear prediction. Finally, it describes how the above results can be applied to image segmentation using a simple cluster seeking algorithm. The results indicate that arbitrarily-shaped image regions can be well identified and clustered using as features their 2-D LPC parameters.
Petros Maragos, Russell M. Mersereau, Ronald W. Schafer
ICASSP1
1982 Some experiments in ADPCM coding of images
abstract
This paper is a preliminary report on a study of the application of two-dimensional linear prediction in image quantization. The study has focused on three major concerns: implementation of an adaptive linear predictor, adaptive quantization of the prediction error signal, and the adaptive predictive coding of density (logarithm of intensity) images. The results of the study indicate that through the use of adaptive prediction and quantization, a high level of image fidelity can be obtained for both intensity and density images at information rates well below one bit/pixel.
Petros Maragos, Russell M. Mersereau, Ronald W. Schafer
ICASSP1