Gerasimos Potamianos

dblp:78/2516 · DBLP profile ↗
← Back
81ranked-venue papers
17as first author
8since 2021 · last 2025
0000-0002-9833-7124ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 71 · 14 first-author · 7 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 4 since 2021Systems, architecture and hardware · 2Theory of computation · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Resource-Efficient and Noise-Robust Modality Fusion for Audio-Visual Speech Recognition
abstract
Resource-efficient audio-visual fusion techniques often struggle to maintain robust performance across varying acoustic noise conditions in speech recognition tasks. This paper introduces a dynamic routing approach for noise-robust audio-visual fusion, which adaptively directs features to noise-specific subnetworks. Unlike attention-based fusion methods that incur quadratic computational complexity, our approach maintains constant complexity, making it suitable for resource-constrained applications. We evaluate our method on the "Lip Reading in the Wild" dataset for isolated (within context) word recognition, augmented with diverse acoustic noise conditions. Results demonstrate significant improvements over feature concatenation and gating baselines of similar computational complexity. In extreme noise (-20 dB), our approach achieves 89.51% mean word accuracy, outperforming the next-best method (bimodal gating) by 0.63% absolute. Notably, it also excels in low-noise scenarios (20 dB), with only a marginal 0.03% absolute accuracy decrease compared to the audio-only system, while the second-best method shows a 0.54% absolute decline. These results demonstrate our method’s effectiveness in mitigating catastrophic fusion and maintaining high performance across diverse acoustic environments, while remaining computationally efficient.
Alexandros Koumparoulis, Gerasimos Potamianos
ICASSP2
2025 A Multi-Stream Framework Utilizing 3D Human Reconstruction for Cued Speech Recognition
Katerina Papadimitriou, Gerasimos Potamianos
INTERSPEECH2
2024 Multimodal Continuous Fingerspelling Recognition via Visual Alignment Learning
Katerina Papadimitriou, Gerasimos Potamianos
INTERSPEECH2
2024 A large corpus for the recognition of Greek Sign Language gestures
Katerina Papadimitriou, Galini Sapountzaki, Kyriaki Vasilaki, Eleni Efthimiou, Stavroula-Evita Fotinea, Gerasimos Potamianos
Comput. Vis. Image Underst.6
2023 Sign Language Recognition via Deformable 3D Convolutions and Modulated Graph Convolutional Networks
abstract
Automatic sign language recognition (SLR) remains challenging, especially when employing RGB video alone (i.e., with no depth or special glove-based input) and under a signer-independent (SI) framework, due to inter-personal signing variation. In this paper, we address SI isolated SLR from RGB video, proposing an innovative deep-learning framework that leverages multi-modal appearanceand skeleton-based information. Specifically, we propose three components for the first time in SLR: (i) a modified version of the ResNet2+1D network to capture signing appearance information, where spatial and temporal convolutions are substituted by their deformable counterparts, accomplishing both prevalent spatial modeling potential and motion-aware modeling adaptability; (ii) a novel spatio-temporal graph convolutional network (ST-GCN) that integrates a GCN variant, involving weight and affinity modulation for modeling diverse correlations between different body joints beyond the physical human skeleton structure, followed by a self-attention layer and a temporal convolution; and (iii) the “PIXIE” 3D human pose and shape regressor to generate 3D joint-rotation parameterization used for ST-GCN graph construction. Both appearance- and skeleton-based streams are ensembled in the proposed system and evaluated on two datasets of isolated signs, one in Turkish and one in Greek. Our system outperforms the state-of-the-art on the second set, yielding 53% relative error rate reduction (2.45% absolute), while it performs on par with the best reported system on the first.
Katerina Papadimitriou, Gerasimos Potamianos
ICASSP2
2023 Multimodal Locally Enhanced Transformer for Continuous Sign Language Recognition
Katerina Papadimitriou, Gerasimos Potamianos
INTERSPEECH2
2022 Accurate and Resource-Efficient Lipreading with Efficientnetv2 and Transformers
abstract
We present a novel resource-efficient end-to-end architecture for lipreading that achieves state-of-the-art results on a popular and challenging benchmark. In particular, we make the following contributions: First, inspired by the recent success of the EfficientNet architecture in image classification and our earlier work on resource-efficient lipreading models (MobiLipNet), we introduce Efficient-Nets to the lipreading task. Second, we show that the currently most popular in the literature 3D front-end contains a max-pool layer that prohibits networks from reaching superior performance and propose its removal. Finally, we improve our system’s back-end robustness by including a Transformer encoder. We evaluate our proposed system on the “Lipreading In-The-Wild” (LRW) corpus, a database containing short video segments from BBC TV broadcasts. The proposed network (T-variant) attains 88.53% word accuracy, a 0.17% absolute improvement over the current state-of-the-art, while being five times less computationally intensive. Further, an up-scaled version of our model (L-variant) achieves 89.52%, a new state-of-the-art result on the LRW corpus.
Alexandros Koumparoulis, Gerasimos Potamianos
ICASSP2
2022 Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language Recognition
abstract
We address the challenging problem of continuous sign language recognition (CSLR) from RGB videos, proposing a novel deep-learning framework that employs spatio-temporal graph convolutional networks (ST-GCNs), which operate on multiple, appropriately fused feature streams, capturing the signer’s pose, shape, appearance, and motion information. In addition to introducing such networks to the continuous recognition problem, our model’s novelty lies on: (i) the feature streams considered and their blending into three ST-GCN modules; (ii) the combination of such modules with bi-directional long short-term memory networks, thus capturing both short-term embedded signing dynamics and long-range feature dependencies; and (iii) the fusion scheme, where the resulting modules operate in parallel, their posteriors aligned via a guiding connectionist temporal classification method, and fused for sign gloss prediction. Notably, concerning (i), in addition to traditional CSLR features, we investigate the utility of 3D human pose and shape parameterization via the "ExPose" approach, as well as 3D skeletal joint information that is regressed from detected 2D joints. We evaluate the proposed system on two well-known CSLR benchmarks, conducting extensive ablations on its modules. We achieve the new state-of-the-art on one of the two datasets, while reaching very competitive performance on the other.
Maria Parelli, Katerina Papadimitriou, Gerasimos Potamianos, Georgios Pavlakos, Petros Maragos
ICASSP3
2020 Audio-Assisted Image Inpainting for Talking Faces
abstract
The goal of our work is to complete missing areas of images of talking faces, exploiting information from both the visual and audio modalities. Existing image inpainting methods rely solely on visual content that doesn't always provide sufficient information for the task. To counter this, we propose a neural network that employs an encoder-decoder architecture with a bimodal fusion mechanism, thus taking into account both visual and audio content. Our proposed method demonstrates consistently superior performance over a baseline visual-only model, reaching for example up to 17% relative improvement in mean absolute error. The presented model is applicable to practical video editing tasks, such as object and overlay-text removal from talking faces, where existing lip and face generation works are not applicable as they require clean input.
Alexandros Koumparoulis, Gerasimos Potamianos, Samuel Thomas 0001, Edmilson da Silva Morais
ICASSP2
2020 A Deep Learning Approach to Object Affordance Segmentation
abstract
Learning to understand and infer object functionalities is an important step towards robust visual intelligence. Significant research efforts have recently focused on segmenting the object parts that enable specific types of human-object interaction, the so-called "object affordances". However, most works treat it as a static semantic segmentation problem, focusing solely on object appearance and relying on strong supervision and object detection. In this paper, we propose a novel approach that exploits the spatio-temporal nature of human-object interaction for affordance segmentation. In particular, we design an autoencoder that is trained using ground-truth labels of only the last frame of the sequence, and is able to infer pixel-wise affordance labels in both videos and static images. Our model surpasses the need for object labels and bounding boxes by using a soft-attention mechanism that enables the implicit localization of the interaction hotspot. For evaluation purposes, we introduce the SOR3D-AFF corpus, which consists of human-object interaction sequences and supports 9 types of affordances in terms of pixel-wise annotation, covering typical manipulations of tool-like objects. We show that our model achieves competitive results compared to strongly supervised methods on SOR3D-AFF, while being able to predict affordances for similar unseen objects in two affordance image-only datasets.
Spyridon Thermos, Petros Daras, Gerasimos Potamianos
ICASSP3
2020 Resource-Adaptive Deep Learning for Visual Speech Recognition
abstract
We focus on the problem of efficient architectures for lipreading that allow trading-off computational resources for visual speech recognition accuracy. In particular, we make two contributions: First, we introduce MobiLipNetV3, an efficient and accurate lipreading model, based on our earlier work on MobiLipNetV2 and incorporating recent advances in convolutional neural network architectures. Second, we propose a novel recognition paradigm, called MultiRate Ensemble (MRE), that combines a “lean” and a “full” MobiLipNetV3 in the lipreading pipeline, with the latter applied at a lower frame rate. This architecture yields a family of systems offering multiple accuracy vs. efficiency operating points depending on the frame-rate decimation of the “full” model, thus allowing adaptation to the available device resources. We evaluate our approach on the TCD-TIMIT corpus, popular in speaker-independent lipreading of continuous speech. The proposed MRE family of systems can be up to 73 times more efficient compared to residual neural network based lipreading, and up to twice as MobiLipNetV2, while in both cases reaching up to 8% absolute WER reduction, depending on the MRE chosen operating point. For example, a temporal decimation of three yields a 7% absolute WER reduction and a 26% relative decrease in computations over MobiLipNetV2. © 2020 ISCA
Alexandros Koumparoulis, Gerasimos Potamianos, Samuel Thomas 0001, Edmilson da Silva Morais
INTERSPEECH2
2020 Multimodal Sign Language Recognition via Temporal Deformable Convolutional Sequence Learning
abstract
In this paper we address the challenging problem of sign language recognition (SLR) from videos, introducing an end-to-end deep learning approach that relies on the fusion of a number of spatio-temporal feature streams, as well as a fully convolutional encoder-decoder for prediction. Specifically, we examine the contribution of optical flow, human skeletal features, as well as appearance features of handshapes and mouthing, in conjunction with a temporal deformable convolutional attention-based encoder-decoder for SLR. To our knowledge, this is the first use in this task of a fully convolutional multi-step attention-based encoder-decoder employing temporal deformable convolutional block structures. We conduct experiments on three sign language datasets and compare our approach to existing state-of-the-art SLR methods, demonstrating its superiority. © 2020 ISCA
Katerina Papadimitriou, Gerasimos Potamianos
INTERSPEECH2
2020 Deep sensorimotor learning for RGB-D object recognition
Spyridon Thermos, Georgios Th. Papadopoulos, Petros Daras, Gerasimos Potamianos
Comput. Vis. Image Underst.4
2019 MobiLipNet: Resource-Efficient Deep Learning Based Lipreading
abstract
Recent works in visual speech recognition utilize deep learning advances to improve accuracy. Focus however has been primarily on recognition performance, while ignoring the computational burden of deep architectures. In this paper we address these issues concurrently, aiming at both high computational efficiency and recognition accuracy in lipreading. For this purpose, we investigate the MobileNet convolutional neural network architectures, recently proposed for image classification. In addition, we extend the 2D convolutions of MobileNets to 3D ones, in order to better model the spatio-temporal nature of the lipreading problem. We investigate two architectures in this extension, introducing the temporal dimension as part of either the depthwise or the pointwise MobileNet convolutions. To further boost computational efficiency, we also consider using pointwise convolutions alone, as well as networks operating on half the mouth region. We evaluate the proposed architectures on speaker-independent visual-only continuous speech recognition on the popular TCD-TIMIT corpus. Our best system outperforms a baseline CNN by 4.27% absolute in word error rate and over 12 times in computational efficiency, whereas, compared to a state-of-the-art ResNet, it is 37 times more efficient at a minor 0.07% absolute error rate degradation. Copyright © 2019 ISCA
Alexandros Koumparoulis, Gerasimos Potamianos
INTERSPEECH2
2019 End-to-End Convolutional Sequence Learning for ASL Fingerspelling Recognition
abstract
Although fingerspelling is an often overlooked component of sign languages, it has great practical value in the communication of important context words that lack dedicated signs. In this paper we consider the problem of fingerspelling recognition in videos, introducing an end-to-end lexicon-free model that consists of a deep auto-encoder image feature learner followed by an attention-based encoder-decoder for prediction. The feature extractor is a vanilla auto-encoder variant, employing a quadratic activation function. The learned features are subsequently fed into the attention-based encoder-decoder. The latter deviates from traditional recurrent neural network architectures, being a fully convolutional attention-based encoder-decoder that is equipped with a multi-step attention mechanism relying on a quadratic alignment function and gated linear units over the convolution output. The introduced model is evaluated on the TTIC/UChicago fingerspelling video dataset, where it outperforms previous approaches in letter accuracy under all three, signer-dependent, -adapted, and -independent, experimental paradigms. Copyright © 2019 ISCA
Katerina Papadimitriou, Gerasimos Potamianos
INTERSPEECH2
2018 Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and Adults
abstract
Human-robot interaction (HRI) is a research area of growing interest with a multitude of applications for both children and adult user groups, as, for example, in edutainment and social robotics. Crucial, however, to its wider adoption remains the robust perception of HRI scenes in natural, untethered, and multi-party interaction scenarios, across user groups. Towards this goal, we investigate three focal HRI perception modules operating on data from multiple audio-visual sensors that observe the HRI scene from the far-field, thus bypassing limitations and platform-dependency of contemporary robotic sensing. In particular, the developed modules fuse intra- and/or inter-modality data streams to perform: (i) audio-visual speaker localization; (ii) distant speech recognition; and (iii) visual recognition of hand-gestures. Emphasis is also placed on ensuring high speech and gesture recognition rates for both children and adults. Development and objective evaluation of the three modules is conducted on a corpus of both user groups, collected by our far-field multisensory setup, for an interaction scenario of a question-answering “guess-the-object” collaborative HRI game with a “Furhat” robot. In addition, evaluation of the game incorporating the three developed modules is reported. Our results demonstrate robust far-field audio-visual perception of the multi-party HRI scene.
Antigoni Tsiami, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Petros Koutras, Gerasimos Potamianos, Petros Maragos
ICASSP5
2018 Multi- View Fusion for Action Recognition in Child-Robot Interaction
abstract
Answering the challenge of leveraging computer vision methods in order to enhance Human Robot Interaction (HRI) experience, this work explores methods that can expand the capabilities of an action recognition system in such tasks. A multi-view action recognition system is proposed for integration in HRI scenarios with special users, such as children, in which there is limited data for training and many state-of-the-art techniques face difficulties. Different feature extraction approaches, encoding methods and fusion techniques are combined and tested in order to create an efficient system that recognizes children pantomime actions. This effort culminates in the integration of a robotic platform and is evaluated under an alluring Children Robot Interaction scenario.
Niki Efthymiou, Petros Koutras, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos
ICIP4
2018 Attention-Enhanced Sensorimotor Object Recognition
abstract
Sensorimotor learning, namely the process of understanding the physical world by combining visual and motor information, has been recently investigated, achieving promising results for the task of 2D/3D object recognition. Following the recent trend in computer vision, powerful deep neural networks (NNs) have been used to model the “sensory” and “motor” information, namely the object appearance and affordance. However, the existing implementations cannot efficiently address the spatio-temporal nature of the human-object interaction. Inspired by recent work on attention-based learning, this paper introduces an attention-enhanced NN-based model that learns to selectively focus on parts of the physical interaction where the object appearance is corrupted by occlusions and deformations. The model's attention mechanism relies on the confidence of classifying an object based solely on its appearance. Three metrics are used to measure the latter, namely the prediction entropy, the average N-best likelihood difference, and the N-best likelihood dispersion. Evaluation of the attention-enhanced model on the SOR3D dataset reports 33% and 26% relative improvement over the appearance-only and the spatio-temporal fusion baseline models, respectively.
Spyridon Thermos, Georgios Th. Papadopoulos, Petros Daras, Gerasimos Potamianos
ICIP4
2018 Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots
abstract
Child-robot interaction is an interdisciplinary research area that has been attracting growing interest, primarily focusing on edutainment applications. A crucial factor to the successful deployment and wide adoption of such applications remains the robust perception of the child's multi-modal actions, when interacting with the robot in a natural and untethered fashion. Since robotic sensory and perception capabilities are platform-dependent and most often rather limited, we propose a multiple Kinect-based system to perceive the child-robot interaction scene that is robot-independent and suitable for indoors interaction scenarios. The audio-visual input from the Kinect sensors is fed into speech, gesture, and action recognition modules, appropriately developed in this paper to address the challenging nature of child-robot interaction. For this purpose, data from multiple children are collected and used for module training or adaptation. Further, information from the multiple sensors is fused to enhance module performance. The perception system is integrated in a modular multi-robot architecture demonstrating its flexibility and scalability with different robotic platforms. The whole system, called Multi3, is evaluated, both objectively at the module level and subjectively in its entirety, under appropriate child-robot interaction scenarios containing several carefully designed games between children and robots.
Antigoni Tsiami, Petros Koutras, Niki Efthymiou, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos
ICRA5
2018 Object Assembly Guidance in Child-Robot Interaction using RGB-D based 3D Tracking
abstract
This work examines how and to what benefit an autonomous humanoid robot can supervise a child in an object assembly task. In order to understand the child's actions, a novel 3D object tracking algorithm for RGB-D data is employed. The tracker consists of two stages: the first performs a tracking-by-detection scheme on the color stream, to locate the objects on the image plane, while the second uses a particle filter that operates on the depth data stream to refine the first stage output and infer the objects' rotations. Given the six degrees-of-freedom of the assembly part poses, the system is able to recognize which connections have been completed at any given time. This information is then used to select an appropriate verbal or gestural response for the robot. Experimental results show that (a) the tracking algorithm is accurate, fast and robust to severe occlusions and fast movements, (b) the proposed method of assembly state estimation is indeed effective, and (c) the resulting Child-Robot Interaction scenario is educational and enjoyable for the children involved.
Jack Hadfield, Petros Koutras, Niki Efthymiou, Gerasimos Potamianos, Costas S. Tzafestas, Petros Maragos
IROS4
2018 Deep View2View Mapping for View-Invariant Lipreading
abstract
Recently, visual-only and audio-visual speech recognition have made significant progress thanks to deep-learning based, trainable visual front-ends (VFEs), with most research focusing on frontal or near-frontal face videos. In this paper, we seek to expand the applicability of VFEs targeted on frontal face views to non-frontal ones, without making assumptions on the VFE type, and allowing systems trained on frontal-view data to be applied on mismatched, non-frontal videos. For this purpose, we adapt the “pix2pix” model, recently proposed for image translation tasks, to transform non-frontal speaker mouth regions to frontal, employing a convolutional neural network architecture, which we call “view2view”. We develop our approach on the OuluVS2 multiview lipreading dataset, allowing training of four such networks that map views at predefined non-frontal angles (up to profile) to frontal ones, which we subsequently feed to a frontal-view VFE. We compare the “view2view” network against a baseline that performs linear cross-view regression at the VFE space. Results on visual-only, as well as audio-visual automatic speech recognition over multiple acoustic noise conditions, demonstrate that the “view2view” significantly outperforms the baseline, narrowing the performance gap from an ideal, matched scenario of view-specific systems. Improvements are retained when the approach is coupled with an automatic view estimator.
Alexandros Koumparoulis, Gerasimos Potamianos
SLT2
2017 Deep Affordance-Grounded Sensorimotor Object Recognition
abstract
It is well-established by cognitive neuroscience that human perception of objects constitutes a complex process, where object appearance information is combined with evidence about the so-called object affordances, namely the types of actions that humans typically perform when interacting with them. This fact has recently motivated the sensorimotor approach to the challenging task of automatic object recognition, where both information sources are fused to improve robustness. In this work, the aforementioned paradigm is adopted, surpassing current limitations of sensorimotor object recognition research. Specifically, the deep learning paradigm is introduced to the problem for the first time, developing a number of novel neuro-biologically and neuro-physiologically inspired architectures that utilize state-of-the-art neural networks for fusing the available information sources in multiple ways. The proposed methods are evaluated using a large RGB-D corpus, which is specifically collected for the task of sensorimotor object recognition and is made publicly available. Experimental results demonstrate the utility of affordance information to object recognition, achieving an up to 29% relative error reduction by its inclusion.
Spyridon Thermos, Georgios Th. Papadopoulos, Petros Daras, Gerasimos Potamianos
CVPR4
2017 Room-localized spoken command recognition in multi-room, multi-microphone environments
Isidoros Rodomagoulakis, Athanasios Katsamanis, Gerasimos Potamianos, Panagiotis Giannoulis, Antigoni Tsiami, Petros Maragos
Comput. Speech Lang.3
2016 Audio-visual speech activity detection in a two-speaker scenario incorporating depth information from a profile or frontal view
abstract
Motivated by increasing popularity of depth visual sensors, such as the Kinect device, we investigate the utility of depth information in audio-visual speech activity detection. A two-subject scenario is assumed, allowing to also consider speech overlap. Two sensory setups are employed, where depth video captures either a frontal or profile view of the subjects, and is subsequently combined with the corresponding planar video and audio streams. Further, multi-view fusion is regarded, using audio and planar video from a sensor at the complementary view setup. Support vector machines provide temporal speech activity classification for each visually detected subject, fusing the available modality streams. Classification results are further combined to yield speaker diarization. Experiments are reported on a suitable audio-visual corpus recorded by two Kinects. Results demonstrate the benefits of depth information, particularly in the frontal depth view setup, reducing speech activity detection and speaker diarization errors over systems that ignore it.
Spyridon Thermos, Gerasimos Potamianos
SLT2
2015 Multichannel speech enhancement using MEMS microphones
abstract
In this work, we investigate the efficacy of Micro Electro-Mechanical System (MEMS) microphones, a newly developed technology of very compact sensors, for multichannel speech enhancement. Experiments are conducted on real speech data collected using a MEMS microphone array. First, the effectiveness of the array geometry for noise suppression is explored, using a new corpus containing speech recorded in diffuse and localized noise fields with a MEMS microphone array configured in linear and hexagonal array geometries. Our results indicate superior performance of the hexagonal geometry. Then, MEMS microphones are compared to Electret Condenser Microphones (ECMs), using the ATHENA database, which contains speech recorded in realistic smart home noise conditions with hexagonal-type arrays of both microphone types. MEMS microphones exhibit performance similar to ECMs. Good performance, versatility in placement, small size, and low cost, make MEMS microphones attractive for multichannel speech processing.
Z.-I. Skordilis, Antigoni Tsiami, Petros Maragos, Gerasimos Potamianos, Luca Spelgatti, Roberto Sannino
ICASSP4
2015 Detecting audio-visual synchrony using deep neural networks
Etienne Marcheret, Gerasimos Potamianos, Josef Vopicka, Vaibhava Goel
INTERSPEECH2
2014 Robust far-field spoken command recognition for home automation combining adaptation and multichannel processing
abstract
The paper presents our approach to speech-controlled home automation. We are focusing on the detection and recognition of spoken commands preceded by a key-phrase as recorded in a voice-enabled apartment by a set of multiple microphones installed in the rooms. For both problems we investigate robust modeling, environmental adaptation and multichannel processing to cope with a) insufficient training data and b) the far-field effects and noise in the apartment. The proposed integrated scheme is evaluated in a challenging and highly realistic corpus of simulated audio recordings and achieves F-measure close to 0.70 for key-phrase spotting and word accuracy close to 98% for the command recognition task.
Athanasios Katsamanis, Isidoros Rodomagoulakis, Gerasimos Potamianos, Petros Maragos, Antigoni Tsiami
ICASSP3
2014 ATHENA: a Greek multi-sensory database for home automation control uthor: isidoros rodomagoulakis (NTUA, Greece)
Antigoni Tsiami, Isidoros Rodomagoulakis, Panagiotis Giannoulis, Athanasios Katsamanis, Gerasimos Potamianos, Petros Maragos
INTERSPEECH5
2012 A hierarchical approach with feature selection for emotion recognition from speech
Panagiotis Giannoulis, Gerasimos Potamianos
LREC2
2011 Special Section on Interactive Multimedia
abstract
The five papers in this special section present effective solutions to tackle interactive multimedia challenges.
Shueng-Han Gary Chan, Pascal Frossard, Gerasimos Potamianos
IEEE Trans. Multim.4
2010 Joint estimation of DOA and speech based on EM beamforming
abstract
In this paper, we propose a multi-microphone joint optimal estimation of the direction of arrival (DOA) and the source speech signal through newly introduced EM beamforming. This produces a posterior PDF for the DOA, based only on the reliable speech spectrum. By maximizing over the posterior PDF of the DOA, we achieve maximum a posteriori DOA estimation. After convergence, the estimated source spectrum through weighted sum in the Bayesian sense is a maximum likelihood estimate (MLE). This is a sufficient statistic for minimum mean square error (MMSE) optimal estimation using a subsequent single channel MMSE filter.
Lae-Hoon Kim, Mark Hasegawa-Johnson, Gerasimos Potamianos, Vit Libal
ICASSP3
2009 Audio-visual automatic speech recognition and related bimodal speech technologies: A review of the state-of-the-art and open problems
abstract
Summary form only given. The presentation will provide an overview of the main research achievements and the state-of-the-art in the area of audiovisual speech processing, mainly focusing in the area of audio-visual automatic speech recognition. The topic has been of interest in the speech research community due to the potential of increased robustness to acoustic noise that the visual modality holds. Nevertheless, significant challenges remain that have hindered practical applications of the technology most notably difficulties with visual speech information extraction and audio-visual fusion algorithms that remain robust to the audio-visual environment variability inherent in practical, unconstrained interaction scenarios and audio-visual data sources, for example multiparty interaction in smart spaces, broadcast news, etc. These challenges are also shared across a number of interesting audio-visual speech technologies beyond the core speech recognition problem, where the visual modality has the potential to resolve ambiguity inherent in the audio signal alone; for example, speech activity detection, speaker diarization, and source separation.
Gerasimos Potamianos
ASRU1
2009 Long-time span acoustic activity analysis from far-field sensors in smart homes
abstract
Smart homes for the aging population have recently started attracting the attention of the research community. One of the problems of interest is this of monitoring the activities of daily living (ADLs) of the elderly, in order to help identify critical problems, aiming to improve their protection and general well-being. In this paper, we report on our initial attempts to recognize such activities, based on input from networks of far-field microphones distributed inside the home. We propose two approaches to the problem: The first models the entire activity, which typically covers long time spans, with a single statistical model, for example a hidden Markov model (HMM), a Gaussian mixture model (GMM), or GMM super-vectors in conjunction with support vector machines (SVMs). The second is a two-step approach: It first performs acoustic event detection (AED) to locate distinctive events, characteristic of the ADLs, and it is subsequently followed by a post-processing stage that employs activity-specific language models (LMs) to classify the output sequences of detected events into ADLs. Experiments are reported on a corpus containing a small number of acted ADLs, collected as part of the Netcarity Integrated Project inside a two-room smart home. Our results show that SVM GMM supervector modeling improves six-class ADL classification accuracy to 76%, compared to 56% achieved by the GMMs, while also outperforming HMMs by 8% absolute. Preliminary results from LM scoring of acoustic event sequences are comparable to those from GMMs on a three-class ADL classification task.
Jing Huang 0019, Xiaodan Zhuang, Vit Libal, Gerasimos Potamianos
ICASSP4
2009 Acoustic fall detection using Gaussian mixture models and GMM supervectors
abstract
We present a system that detects human falls in the home environment, distinguishing them from competing noise, by using only the audio signal from a single far-field microphone. The proposed system models each fall or noise segment by means of a Gaussian mixture model (GMM) supervector, whose Euclidean distance measures the pairwise difference between audio segments. A support vector machine built on a kernel between GMM supervectors is employed to classify audio segments into falls and various types of noise. Experiments on a dataset of human falls, collected as part of the Netcarity project, show that the method improves fall classification F-score to 67% from 59% of a baseline GMM classifier. The approach also effectively addresses the more difficult fall detection problem, where audio segment boundaries are unknown. Specifically, we employ it to reclassify confusable segments produced by a dynamic programming scheme based on traditional GMMs. Such post-processing improves a fall detection accuracy metric by 5% relative.
Xiaodan Zhuang, Jing Huang 0019, Gerasimos Potamianos, Mark Hasegawa-Johnson
ICASSP3
2009 Robust audio-visual speech synchrony detection by generalized bimodal linear prediction
Kshitiz Kumar, Jirí Navrátil 0001, Etienne Marcheret, Vit Libal, Gerasimos Potamianos
INTERSPEECH5
2008 A multi-modal spoken dialog system for interactive TV
abstract
In this demonstration we present a novel prototype system that implements a multi-modal interface for control of the television. This system combines the standard TV remote control with a dialog management based natural language speech interface to allow users to efficiently interact with the TV, and to seamlessly alternate between the two modalities. One of the main objectives of this system is to make the unwieldy Electronic Program Guide information more navigable by the use of voice to filter and locate programs of interest.
Rajesh Balchandran, Mark E. Epstein, Gerasimos Potamianos, Ladislav Serédi
ICMI3
2007 Kernel-Based 3D Tracking
abstract
We present a computer vision system for robust object tracking in 3D by combining evidence from multiple calibrated cameras. This kernel-based 3D tracker is automatically bootstrapped by constructing 3D point clouds. These points clouds are then clustered and used to initialize the trackers and validate their performance. The framework describes a complete tracking system that fuses appearance features from all available camera sensors and is capable of automatic initialization and drift detection. Its elegance resides in its inherent ability to handle problems encountered by various 2D trackers, including scale selection, occlusion, view-dependence, and correspondence across views. Tracking results for an indoor smart room and a multi-camera outdoor surveillance scenario are presented. We demonstrate the effectiveness of this unified approach by comparing its performance to a baseline 3D tracker that fuses results of independent 2D trackers, as well as comparing the re-initialization results to known ground truth.
Ambrish Tyagi, Mark A. Keck, James W. Davis, Gerasimos Potamianos
CVPR4
2007 Dynamic Stream Weight Modeling for Audio-Visual Speech Recognition
abstract
To generate optimal multi-stream audio-visual speech recognition performance, appropriate dynamic weighting of each modality is desired. In this paper, we propose to estimate such weights based on a combination of acoustic signal space observations and single-modality audio and visual speech model likelihoods. Two modeling approaches are investigated for such weight estimation: one based on a sigmoid fitting function, the other employing Gaussian mixture models. Reported experiments demonstrate that the later approach outperforms sigmoid based modeling, and is dramatically superior to the static weighting scheme.
Etienne Marcheret, Vit Libal, Gerasimos Potamianos
ICASSP (4)3
2007 Detection, diarization, and transcription of far-field lecture speech
abstract
Speech processing of lectures recorded inside smart rooms has recently attracted much interest. In particular, the topic has been central to the Rich Transcription (RT) Meeting Recognition Evaluation campaign series, sponsored by NIST, with emphasis placed on benchmarking speech activity detection (SAD), speaker diarization (SPKR), speech-to-text (STT), and speakerattributed STT (SASTT) technologies. In this paper, we present the IBM systems developed to address these tasks in preparation for the RT 2007 evaluation, focusing on the far-field condition of lecture data collected as part of European project CHIL. For their development, the systems are benchmarked on a subset of the RT Spring 2006 (RT06s) evaluation test set, where they yield significant improvements for all SAD, SPKR, and STT tasks over RT06s results; for example, a 16% relative reduction in word error rate is reported in STT, attributed to a number of system advances discussed here. Initial results are also presented on SASTT, a task newly introduced in 2007 in place of the discontinued SAD. Index Terms: speech processing, speech recognition, speaker diarization, speech activity detection, lectures, smart rooms.
Jing Huang 0019, Etienne Marcheret, Karthik Visweswariah, Vit Libal, Gerasimos Potamianos
INTERSPEECH5
2007 A unified approach to multi-pose audio-visual ASR
abstract
The vast majority of studies in the field of audio-visual automatic speech recognition (AVASR) assumes frontal images of a speaker's face, but this cannot always be guaranteed in practice. Hence our recent research efforts have concentrated on extracting visual speech information from non-frontal faces, in particular the profile view. The introduction of additional views to an AVASR system increases the complexity of the system, as it has to deal with the different visual features associated with the various views. In this paper, we propose the use of linear regression to find a transformation matrix based on synchronous frontal and profile visual speech data, which is used to normalize the visual speech in each viewpoint into a single uniform view. In our experiments for the task of multi-speaker lipreading, we show that this "pose-invariant" technique reduces train/test mismatch between visual speech features of different views, and is of particular benefit when there is more training data for one viewpoint over another (e.g. frontal over profile).
Patrick Lucey, Gerasimos Potamianos, Sridha Sridharan
INTERSPEECH2
2007 An Embedded System for In-Vehicle Visual Speech Activity Detection
abstract
We present a system for automatically detecting driver's speech in the automobile domain using visual-only information extracted from the driver's mouth region. The work is motivated by the desire to eliminate manual push-to-talk activation of the speech recognition engine in newly designed voice interfaces in the typically noisy car environment, aiming at reducing driver cognitive load and increasing naturalness of the interaction. The proposed system uses a camera mounted on the rearview mirror to monitor the driver, detect face boundaries and facial features, and finally employ lip motion clues to recognize visual speech activity. In particular, the designed algorithm has very low computational cost, which allows real-time implementation on currently available inexpensive embedded platforms, as described in the paper. Experiments are also reported on a small multi-speaker database collected in moving automobiles, that demonstrate promising accuracy.
Vit Libal, Jonathan H. Connell, Gerasimos Potamianos, Etienne Marcheret
MMSP3
2006 Person Tracking in Smart Rooms using Dynamic Programming and Adaptive Subspace Learning
abstract
We present a robust vision system for single person tracking inside a smart room using multiple synchronized, calibrated, stationary cameras. The system consists of two main components, namely initialization and tracking, assisted by an additional component that detects tracking drift. The main novelty lies in the adaptive tracking mechanism that is based on subspace learning of the tracked person appearance in selected two-dimensional camera views. The sub-space is learned on the fly, during tracking, but in contrast to the traditional literature approach, an additional "forgetting" mechanism is introduced, as a means to reduce drifting. The proposed algorithm replaces mean-shift tracking, previously employed in our work. By combining the proposed technique with a robust initialization component that is based on face detection and spatio-temporal dynamic programming, the resulting vision system significantly outperforms previously reported systems for the task of tracking the seminar presenter in data collected as part of the CHIL project
ZhenQiu Zhang, Gerasimos Potamianos, Stephen M. Chu, Jilin Tu, Thomas S. Huang
ICME2
2006 Lipreading Using Profile Versus Frontal Views
abstract
Visual information from a speaker's mouth region is known to improve automatic speech recognition robustness. However, the vast majority of audio-visual automatic speech recognition (AVASR) studies assume frontal images of the speaker's face. In contrast, this paper investigates extracting visual speech information from the speaker's profile view, and, to our knowledge, constitutes the first real attempt to attack this problem. As with any AVASR system, the overall recognition performance depends heavily on the visual front end. This is especially the case with profile-view data, as the facial features are heavily compacted compared to the frontal scenario. In this paper, we particularly describe our visual front end approach, and report experiments on a multi-subject, small-vocabulary, bimodal, multi-sensory database that contains synchronously captured audio with frontal and profile face video. Our experiments show that AVASR is possible from profile views with moderate performance degradation compared to frontal video data
Patrick Lucey, Gerasimos Potamianos
MMSP2
2005 Improved face finding in visually challenging environments
abstract
Finding faces in visually challenging environments is crucial to many applications, such as audio-visual automatic speech recognition, video indexing, person recognition, and video surveillance. In this study, we investigate several algorithms to improve face detection accuracy in visually challenging environments using the IBM appearance based face detection system. The algorithms considered are trainable skintone pre-screening, Hamming windowing of the face images, DCT coefficient selection, and the AdaBoost technique. When these methods are combined, an up to 68% relative reduction in face detection error is observed on visually challenging datasets.
Jintao Jiang, Gerasimos Potamianos, Giridharan Iyengar
ICME2
2005 Automatic Speech Activity Detection, Source Localization, and Speech Recognition on the Chil Seminar Corpus
abstract
To realize the long-term goal of ubiquitous computing, technological advances in multi-channel acoustic analysis are needed in order to solve several basic problems, including speaker localization and tracking, speech activity detection (SAD) and distant-talking automatic speech recognition (ASR). The European Commission integrated project CHIL, “ Computers in the Human Interaction Loop”, aims to make significant advances in these three technologies. In this work, we report the results of our initial automatic source localization, speech activity detection, and speech recognition experiments on the CHIL seminar corpus, which is comprised of spontaneous speech collected by both near- and far-field microphones. In addition to the audio sensors, the seminars were also recorded by calibrated video cameras. This simultaneous audio-visual data capture enables the realistic evaluation of component technologies as was never possible with earlier data bases.
Dusan Macho, Jaume Padrell, Alberto Abad, Climent Nadeu, Javier Hernando, John W. McDonough, Matthias Wölfel, Ulrich Klee, Maurizio Omologo, Alessio Brutti, Piergiorgio Svaizer, Gerasimos Potamianos, Stephen M. Chu
ICME12
2005 Speech activity detection fusing acoustic phonetic and energy features
Etienne Marcheret, Karthik Visweswariah, Gerasimos Potamianos
INTERSPEECH3
2004 Improved face and feature finding for audio-visual speech recognition in visually challenging environments
abstract
Visual information in a speaker's face is known to improve the robustness of automatic speech recognition (ASR). However, most studies in audio-visual ASR have focused on "visually clean" data to benefit ASR in noise. This paper is a follow up on a previous study that investigated audio-visual ASR in visually challenging environments. It focuses on visual speech front end processing, and it proposes an improved, appearance based face and feature detection algorithm that utilizes Gaussian mixture model classifiers. This method is shown to improve the accuracy of face and feature detection, and thus visual speech recognition, over our previously used baseline system. In turn, this translates to improved audio-visual ASR, resulting in a 10% relative reduction of the word-error-rate in noisy speech.
Jintao Jiang, Gerasimos Potamianos, Harriet J. Nock, Giridharan Iyengar, Chalapathy Neti
ICASSP (5)2
2004 Towards practical deployment of audio-visual speech recognition
abstract
Much progress has been achieved during the past two decades in audio-visual automatic speech recognition (AVASR). However, challenges persist that hinder AVASR deployment in practical situations, most notably, robust and fast extraction of visual speech features. We review our efforts in overcoming this problem, based on an appearance-based visual feature representation of the speaker's mouth region. We cover three topics in particular. Firstly, we discuss AVASR in realistic, visually challenging domains, where lighting, background, and head-pose vary significantly. To enhance visual-front-end robustness in such environments, we employ an improved statistical-based face detection algorithm that significantly outperforms our baseline scheme. However, visual-only recognition remains inferior to visually "clean" (studio-like) data, thus demonstrating the importance of accurate mouth region extraction. We then consider a wearable audio-visual sensor to capture the mouth region directly, thus eliminating face detection. Its use improves visual-only recognition, even over full-face videos recorded in the studio-like environment. Finally, we address the speed issue in visual feature extraction, by discussing our real-time AVASR prototype implementation. The reported progress demonstrates the feasibility of practical AVASR.
Gerasimos Potamianos, Chalapathy Neti, Jing Huang 0019, Jonathan H. Connell, Stephen M. Chu, Vit Libal, Etienne Marcheret, Norman Haas, Jintao Jiang
ICASSP (3)1
2004 Multistage information fusion for audio-visual speech recognition
abstract
The paper looks into the information fusion problem in the context of audio-visual speech recognition. Existing approaches to audio-visual fusion typically address the problem in either the feature domain or the decision domain. We consider a hybrid approach that aims to take advantage of both the feature fusion and the decision fusion methodologies. We introduce a general formulation to facilitate information fusion at multiple stages, followed by an experimental study of a set of fusion schemes allowed by the framework. The proposed method is implemented on a real-time audio-visual speech recognition system, and evaluated on connected digit recognition tasks under varying acoustic conditions. The results show that the multistage fusion system consistently achieves lower word error rates than the reference feature fusion and decision fusion systems. It is further shown that removing the audio only channel from the multistage system leads to only minimal degradations in recognition performance while providing a noticeable reduction in computational load.
Stephen M. Chu, Vit Libal, Etienne Marcheret, Chalapathy Neti, Gerasimos Potamianos
ICME5
2004 Efficient likelihood computation in multi-stream HMM based audio-visual speech recognition
Etienne Marcheret, Stephen M. Chu, Vaibhava Goel, Gerasimos Potamianos
INTERSPEECH4
2004 Mutual information based visual feature selection for lipreading
abstract
Image transforms, such as the discrete cosine, are widely used to extract visual features from the speaker's mouth region to be used in automatic speechreading and audio-visual speech recognition. Typically, the spatial frequency components with the highest energy in the transform space are retained for recognition. This paper proposes an alternative technique for selecting such features, by utilizing the mutual information criterion instead. Mutual information between each individual spatial frequency component and the speech classes of interest is employed as a measure of its appropriateness for speech classification. The highest mutual information components are then selected as visual speech features. Extensions to this scheme by using joint mutual information between candidate feature pairs and classes are also considered. The algorithm is tested on visual-only speech recognition of connected-digit strings, using an appropriate audio-visual database. For low-dimensional visual feature vectors, the proposed method significantly outperforms features selected by means of energy, reducing word error rate by as much as 20% relative. These gains however diminish as higher feature dimensionalities are allowed.
Patricia Scanlon, Gerasimos Potamianos, Vit Libal, Stephen M. Chu
INTERSPEECH2
2004 Audio-visual speech recognition using an infrared headset
Jing Huang 0019, Gerasimos Potamianos, Jonathan H. Connell, Chalapathy Neti
Speech Commun.2
2003 Audio-visual speaker recognition using time-varying stream reliability prediction
abstract
We examine a time-varying, context dependent, information fusion methodology for multi-stream authentication based on audio and video data collected simultaneously during a user's interaction with a system. Scores obtained from the two data streams are combined based on the relative local richness, as compared to the training data or derived model, and on the stability of each stream. The results show that the proposed technique outperforms the use of video or audio data alone as well as the use of fused data streams (via concatenation). Of particular note is that the performance improvements are achieved for clean, high quality speech, whereas previous efforts focused on degraded speech conditions.
Upendra V. Chaudhari, Ganesh N. Ramaswamy, Gerasimos Potamianos, Chalapathy Neti
ICASSP (5)3
2003 Frame-dependent multi-stream reliability indicators for audio-visual speech recognition
abstract
We investigate the use of local, frame-dependent reliability indicators of the audio and visual modalities, as a means of estimating stream exponents of multi-stream hidden Markov models for audio-visual automatic speech recognition. We consider two such indicators at each modality, defined as functions of the speech-class conditional observation probabilities of appropriate audio-or visual-only classifiers. We subsequently map the four reliability indicators into the stream exponents of a state-synchronous, two-stream hidden Markov model, as a sigmoid function of their linear combination. We propose two algorithms to estimate the sigmoid weights, based on the maximum conditional likelihood and minimum classification error criteria. We demonstrate the superiority of the proposed approach on a connected-digit audio-visual speech recognition task, under varying audio channel noise conditions. Indeed, the use of the estimated, frame-dependent stream exponents results in a significantly smaller word error rate than using global stream exponents. In addition, it outperforms utterance-level exponents, even though the latter utilize a-priori knowledge of the utterance noise level.
Ashutosh Garg 0001, Gerasimos Potamianos, Chalapathy Neti, Thomas S. Huang
ICASSP (1)2
2003 Information fusion and decision cascading for audio-visual speaker recognition based on time-varying stream reliability prediction
abstract
We examine the techniques for multi-modal biometric information fusion for verification and identification of speakers, where the reliability of each data stream, either audio of video, is modeled with parameters that are time-varying and depend on the context created by its local behavior. The complementary nature and the time dependent relative reliability of audio and video data is studied in the context of verification and identification, on data collected during a user's interaction with an automated system. Of significance is that this data is not corrupted artificially. Particular focus is directed to verification and its ability to refine identification decisions, by indicating a level of confidence in the system decisions. Results show more striking effects for verification, when using time-dependent fusion, than for identification.
Upendra V. Chaudhari, Ganesh N. Ramaswamy, Gerasimos Potamianos, Chalapathy Neti
ICME3
2003 A real-time prototype for small-vocabulary audio-visual ASR
abstract
We present a prototype for the automatic recognition of audio-visual speech, developed to augment the IBM ViaVoice/spl trade/ speech recognition system. Frontal face, full frame video is captured through a USB 2.0 interface by means of an inexpensive PC camera, and processed to obtain appearance-based visual features. Subsequently, these are combined with audio features, synchronously extracted from the acoustic signal, using a simple discriminant feature fusion technique. On the average, the required computations utilize approximately 67% of a Pentium/spl trade/ 4, 1.8 GHz processor, leaving the remaining resources available to hidden Markov model based speech recognition. Real-time performance is there- fore achieved for small-vocabulary tasks, such as connected-digit recognition. In the paper, we discuss the prototype architecture based on the ViaVoice engine, the basic algorithms employed, and their necessary modifications to ensure real-time performance and causality of the visual front end processing. We benchmark the resulting system performance on stored videos against prior research experiments, and we report a close match between the two.
Jonathan H. Connell, Norman Haas, Etienne Marcheret, Chalapathy Neti, Gerasimos Potamianos, Senem Velipasalar
ICME5
2003 Frame-dependent multi-stream reliability indicators for audio-visual speech recognition
abstract
We investigate the use of local, frame-dependent reliability indicators of the audio and visual modalities, as a means of estimating stream exponents of multi-stream hidden Markov models for audio-visual automatic speech recognition. We consider two such indicators at each modality, defined as functions of the speech-class conditional observation probabilities of appropriate audio or visual-only classifiers. We subsequently map the four reliability indicators into the stream exponents of a state-synchronous, two-stream hidden Markov model, as a sigmoid function of their linear combination. We propose two algorithms to estimate the sigmoid weights, based on the maximum conditional likelihood and minimum classification error criteria. We demonstrate the superiority of the proposed approach on a connected-digit audio-visual speech recognition task, under varying audio channel noise conditions. Indeed, the use of estimated, frame-dependent stream exponents results in a significantly smaller word error rare than using global stream exponents. In addition, it outperforms utterance-level exponents, even though the latter utilize a-priori knowledge of the utterance noise level.
Ashutosh Garg 0001, Gerasimos Potamianos, Chalapathy Neti, Thomas S. Huang
ICME2
2003 Audio-visual speech recognition in challenging environments
abstract
Visual speech information is known to improve accuracy and noise robustness of automatic speech recognizers. However, todate, all audio-visual ASR work has concentrated on “visually clean” data with limited variation in the speaker’s frontal pose, lighting, and background. In this paper, we investigate audiovisual ASR in two practical environments that present significant challenges to robust visual processing: (a) Typical offices, where data are recorded by means of a portable PC equipped with an inexpensive web camera, and (b) automobiles, with data collected at three approximate speeds. The performance of all components of a state-of-the-art audio-visual ASR system is reported on these two sets and benchmarked against “visually clean” data recorded in a studio-like environment. Not surprisingly, both audio- and visual-only ASR degrade, more than doubling their respective word error rates. Nevertheless, visual speech remains beneficial to ASR.
Gerasimos Potamianos, Chalapathy Neti
INTERSPEECH1
2003 Recent advances in the automatic recognition of audiovisual speech
abstract
Visual speech information from the speaker's mouth region has been successfully shown to improve noise robustness of automatic speech recognizers, thus promising to extend their usability in the human computer interface. In this paper, we review the main components of audiovisual automatic speech recognition (ASR) and present novel contributions in two main areas: first, the visual front-end design, based on a cascade of linear image transforms of an appropriate video region of interest, and subsequently, audiovisual speech integration. On the latter topic, we discuss new work on feature and decision fusion combination, the modeling of audiovisual speech asynchrony, and incorporating modality reliability estimates to the bimodal recognition process. We also briefly touch upon the issue of audiovisual adaptation. We apply our algorithms to three multisubject bimodal databases, ranging from small- to large-vocabulary recognition tasks, recorded in both visually controlled and challenging environments. Our experiments demonstrate that the visual modality improves ASR over all conditions and data considered, though less so for visually challenging environments and large vocabulary tasks.
Gerasimos Potamianos, Chalapathy Neti, Guillaume Gravier, Andrew W. Senior
Proc. IEEE1
2002 Noisy audio feature enhancement using audio-visual speech data
abstract
We investigate improving automatic speech recognition (ASR) in noisy conditions by enhancing noisy audio features using visual speech captured from the speaker's face. The enhancement is achieved by applying a linear filter to the concatenated vector of noisy audio and visual features, obtained by mean square error estimation of the clean audio features in a training stage. The performance of the enhanced audio features is evaluated on two ASR tasks: A connected digits task and speaker-independent, large-vocabulary, continuous speech recognition. In both cases and at sufficiently low signal-to-noise ratios (SNRs), ASR trained on the enhanced audio features significantly outperforms ASR trained on the noisy audio, achieving for example a 46% relative reduction in word error rate on the digits task at −3.5 dB SNR. However, the method fails to capture the full visual modality benefit to ASR, as demonstrated by its comparison to discriminant audio-visual feature fusion introduced in previous work.
Roland Göcke, Gerasimos Potamianos, Chalapathy Neti
ICASSP2
2002 Maximum entropy and MCE based HMM stream weight estimation for audio-visual ASR
abstract
In this paper, we propose a new fast and flexible algorithm based on the maximum entropy (MAXENT) criterion to estimate stream weights in a state-synchronous multi-stream HMM. The technique is compared to the minimum classification error (MCE) criterion and to a brute-force, grid-search optimization of the WER on both a small and a large vocabulary audio-visual continuous speech recognition task. When estimating global stream weights, the MAXENT approach gives comparable results to the grid-search and the MCE. Estimation of state dependent weights is also considered: We observe significant improvements in both the MAXENT and MCE criteria, which, however, do not result in significant WER gains.
Guillaume Gravier, Scott Axelrod, Gerasimos Potamianos, Chalapathy Neti
ICASSP3
2002 Audio-visual speech enhancement with AVCDCN (audio-visual codebook dependent cepstral normalization)
abstract
We introduce a non-linear enhancement technique called audio-visual codebook dependent cepstral normalization (AVCDCN) and we consider its use with both audio-only and audio-visual speech recognition. AVCDCN is inspired from CDCN, an audio-only enhancement technique that approximates the nonlinear effect of noise on speech with a piecewise constant function. Our experiments show that the use of visual information in AVCDCN allows significant performance gains over CDCN.
Sabine Deligne, Gerasimos Potamianos, Chalapathy Neti
INTERSPEECH2
2001 Weighting schemes for audio-visual fusion in speech recognition
abstract
We demonstrate an improvement in the state-of-the-art large vocabulary continuous speech recognition (LVCSR) performance, under clean and noisy conditions, by the use of visual information, in addition to the traditional audio one. We take a decision fusion approach for the audio-visual information, where the single-modality (audio- and visual- only) HMM classifiers are combined to recognize audio-visual speech. More specifically, we tackle the problem of estimating the appropriate combination weights for each of the modalities. Two different techniques are described: the first uses an automatically extracted estimate of the audio stream reliability in order to modify the weights for each modality (both clean and noisy audio results are reported), while the second is a discriminative model combination approach where weights on pre-defined model classes are optimized to minimize WER (clean audio only results).
Hervé Glotin, D. Vergyr, Chalapathy Neti, Gerasimos Potamianos, Jürgen Lüttin
ICASSP4
2001 Asynchronous stream modeling for large vocabulary audio-visual speech recognition
abstract
Addresses the problem of audio-visual information fusion to provide highly robust speech recognition. We investigate methods that make different assumptions about asynchrony and conditional dependence across streams and propose a technique based on composite HMMs that can account for stream asynchrony and different levels of information integration. We show how these models can be trained jointly based on maximum likelihood estimation. Experiments, performed for a speaker-independent large vocabulary continuous speech recognition task and different integration methods, show that best performance is obtained by asynchronous stream integration. This system reduces the error rate at a 8.5 dB SNR with additive speech "babble" noise by 27 % relative over audio-only models and by 12 % relative over traditional audio-visual models using concatenative feature fusion.
Jürgen Lüttin, Gerasimos Potamianos, Chalapathy Neti
ICASSP2
2001 Hierarchical discriminant features for audio-visual LVCSR
abstract
We propose the use of a hierarchical, two-stage discriminant transformation for obtaining audio-visual features that improve automatic speech recognition. Linear discriminant analysis (LDA), followed by a maximum likelihood linear transform (MLLT) is first applied to MFCC based audio-only features, as well as on visual only features, obtained by a discrete cosine transform of the video region of interest. Subsequently, a second stage of LDA and MLLT is applied to the concatenation of the resulting single modality features. The obtained audio-visual features are used to train a traditional HMM based speech recognizer. Experiments on the IBM ViaVoice/sup TM/ audio-visual database demonstrate that the proposed feature fusion method improves speaker-independent, large vocabulary, continuous speech recognition (LVCSR) for both clean and noisy audio conditions considered. A 24% relative word error rate reduction over an audio-only system is achieved in the latter case.
Gerasimos Potamianos, Jürgen Lüttin, Chalapathy Neti
ICASSP1
2001 Improved ROI and within frame discriminant features for lipreading
abstract
We study three aspects of designing appearance based visual features for automatic lipreading: (a) the choice of the video region of interest (ROI) on which image transform features are obtained; (b) the extraction of speech discriminant features at each frame; (c) the use of temporal information to improve visual speech modeling. With respect to (a), we propose a ROI that includes the speaker's jaw and cheeks, in addition to the traditionally used mouth/lip region. With respect to (b) and (c), we propose the use of a two-stage linear discriminant analysis, both within a single frame and across a large number of frames. On a large-vocabulary, continuous-speech, audio-visual database, the proposed visual features result in a 13% absolute reduction in visual-only word error rate over a baseline visual front end, and in an additional 28% relative improvement in audio-visual over audio-only phonetic classification accuracy.
Gerasimos Potamianos, Chalapathy Neti
ICIP (3)1
2001 A Comparison Of Model And Transform-Based Visual Features For Audio-Visual LVCSR
abstract
Four different visual speech parameterisation methods are compared on a large vocabulary, continuous, audio-visual speech recognition task using the IBM ViaVoice TM audio-visual speech database. Three are direct mouth image region based transforms; discrete cosine and wavelet transforms, and principal component analysis. The fourth uses a statistical model of shape and appearance called an active appearance model, to track and obtain model parameters describing the entire face. All parameterisations are compared experimentally using hidden Markov models (HMM's) in a speaker independent test. Visualonly HMM's are used to rescore lattices obtained from audio models trained in noisy conditions. 1.
Iain A. Matthews, Gerasimos Potamianos, Chalapathy Neti, Jürgen Lüttin
ICME2
2001 Large-vocabulary audio-visual speech recognition by machines and humans
abstract
We compare automatic recognition with human perception of audio-visual speech, in the large-vocabulary, continuous speech recognition (LVCSR) domain. Specifically, we study the benefit of the visual modality for both machines and humans, when combined with audio degraded by speech-babble noise at various signal-to-noise ratios (SNRs). We first consider an automatic speechreading system with a pixel based visual front end that uses feature fusion for bimodal integration, and we compare its performance with an audio-only LVCSR system. We then describe results of human speech perception experiments, where subjects are asked to transcribe audio-only and audiovisual utterances at various SNRs. For both machines and humans, we observe approximately a 6 dB effective SNR gain compared to the audio-only performance at 10 dB, however such gains significantly diverge at other SNRs. Furthermore, automatic audio-visual recognition outperforms human audioonly speech perception at low SNRs. 1.
Gerasimos Potamianos, Chalapathy Neti, Giridharan Iyengar, Eric Helmuth
INTERSPEECH1
2001 Robust detection of visual ROI for automatic speechreading
abstract
We present our work on visual pruning in an audio-visual (AV) speech recognition scenario. Visual speech information has been successfully used in circumstances where audio-only recognition suffers (e.g. noisy environments). Tracking and extraction of region-of-interest (ROI) (e.g., speaker's mouth region) from video is an essential component of such systems. It is important for the visual front-end to handle tracking errors that result in noisy visual data and hamper performance. We present our robust visual front-end, investigate methods to prune visual noise and its effect on the performance of the AV speech recognition systems. Specifically, we estimate the "goodness of ROI" using Gaussian mixture models and our experiments indicate that significant performance gains are achieved with good quality visual data.
Giridharan Iyengar, Gerasimos Potamianos, Chalapathy Neti, Tanveer A. Faruquie, Ashish Verma 0001
MMSP2
2001 Large-vocabulary audio-visual speech recognition: a summary of the Johns Hopkins Summer 2000 Workshop
abstract
We report a summary of the Johns Hopkins Summer 2000 Workshop on audio-visual automatic speech recognition (ASR) in the large-vocabulary, continuous speech domain. Two problems of audio-visual ASR were mainly addressed: visual feature extraction and audio-visual information fusion. First, image transform and model-based visual features were considered, obtained by means of the discrete cosine transform (DCT) and active appearance models, respectively. The former were demonstrated to yield superior automatic speech reading. Subsequently, a number of feature fusion and decision fusion techniques for combining the DCT visual features with traditional acoustic ones were implemented and compared. Hierarchical discriminant feature fusion and asynchronous decision fusion by means of the multi-stream hidden Markov model consistently improved ASR for both clean and noisy speech. Compared to an equivalent audio-only recognizer, introducing the visual modality reduced ASR word error rate by 7% relative in clean speech, and by 27% relative at an 8.5 dB SNR audio condition.
Chalapathy Neti, Gerasimos Potamianos, Jürgen Lüttin, Iain A. Matthews, Hervé Glotin, Dimitra Vergyri
MMSP2
2000 Perceptual interfaces for information interaction: joint processing of audio and visual information for human-computer interaction
abstract
We are exploiting the human perceptual principle of sensory integration (the joint use of audio and visual information) to improve the recognition of human activity (speech recognition, speech event detection and speaker change), intent (intent to speak) and human identity (speaker recognition), particularly in the presence of acoustic degradation due to noise and channel. In this paper, we present experimental results in a variety of contexts that demonstrate the benefit of joint audio-visual processing.
Chalapathy Neti, Giridharan Iyengar, Gerasimos Potamianos, Andrew W. Senior, Benoît Maison
INTERSPEECH3
2000 Stream confidence estimation for audio-visual speech recognition
abstract
We investigate the use of single modality confidence measures as a means of estimating adaptive, local weights for improved audio -visual automatic speech recognition. We limit our work to the toy problem of audio-visual phonetic classification by means of a two-stream Gaussian mixture model (GMM), where each stream models the class conditional audio- or visual-only observation probability, raised to an appropriate exponent. We consider such stream exponents as two-dimensional piecewise constant functions of the audio and visual stream local confidences, and we estimate them by minimizing the misclassification error on a held-out data set. Three stream confidence measures are investigated, namely the stream entropy, the n-best likelihood ratio average, and an n-best stream likelihood dispersion measure. The later results in superior audio-visual phonetic classification, as indicated by our experiments on a 260-subject, 40-hour long, large vocabulary, continuous speech audio-visual dataset. By using local, dispersion-based stream exponents, we achieve an additional 20% phone classification accuracy improvement over the improvement that global stream exponents add to clean audio -only phonetic classification. The performance of the algorithm however still falls significantly short of an "oracle" (cheating) confidence estimation scheme.
Gerasimos Potamianos, Chalapathy Neti
INTERSPEECH1
1999 Speaker adaptation for audio-visual speech recognition
Gerasimos Potamianos, Alexandros Potamianos
EUROSPEECH1
1998 Discriminative training of HMM stream exponents for audio-visual speech recognition
abstract
We propose the use of discriminative training by means of the generalized probabilistic descent (GPB) algorithm to estimate hidden Markov model (HMM) stream exponents for audio-visual speech recognition. Synchronized audio and visual features are used to respectively train audio-only and visual-only single-stream HMMs of identical topology by maximum likelihood. A two-stream HMM is then obtained by combining the two single-stream HMMs and introducing exponents that weigh the log-likelihood of each stream. We present the GPD algorithm for stream exponent estimation, consider a possible initialization, and apply it to the single speaker connected letters task of the AT&T bimodal database. We demonstrate the superior performance of the resulting multi-stream HMM to the audio-only, visual-only, and audio-visual single-stream HMMs.
Gerasimos Potamianos, Hans Peter Graf
ICASSP1
1998 An Image Transform Approach for HMM based Automatic Lipreading
Gerasimos Potamianos, Hans Peter Graf, Eric Cosatto
ICIP (3)1
1998 Linear discriminant analysis for speechreading
abstract
This paper investigates the use of Fisher-Rao (1965) linear discriminant analysis (LDA) as a means of visual feature extraction for hidden Markov model based automatic speechreading. For every video frame, a three-dimensional region of interest containing the speaker's mouth over a sequence of adjacent frames is lexicographically arranged into a data vector. Such vectors are then projected onto the space of the most discriminant "eigensequences", estimated by means of LDA on a training set of image sequence vectors, labeled from a set of a-priori chosen classes. The resulting projections, as well as their first and second derivatives over time, are used as features for automatic speechreading. The proposed method is applied to single-speaker, multi-speaker, and speaker-independent visual-only recognition tasks, consistently outperforming principal component analysis and discrete wavelet transform based visual features. Specific issues relevant to LDA are also discussed, namely, class selection, automatic data class labelling, and dimensionality reduction prior to LDA.
Gerasimos Potamianos, Hans Peter Graf
MMSP1
1998 A study of n-gram and decision tree letter language modeling methods
abstract
The goal of this paper is to investigate various language model smoothing techniques and decision tree based language model design algorithms. For this purpose, we build language models for printable characters (letters), based on the Brown corpus. We consider two classes of models for the text generation process: the n-gram language model and various decision tree based language models. In the first part of the paper, we compare the most popular smoothing algorithms applied to the former. We conclude that the bottom-up deleted interpolation algorithm performs the best in the task of n-gram letter language model smoothing, significantly outperforming the back-off smoothing technique for large values of n. In the second part of the paper, we consider various decision tree development algorithms. Among them, a K-means clustering type algorithm for the design of the decision tree questions gives the best results. However, the n-gram language model outperforms the decision tree language models for letter language modeling. We believe that this is due to the predictive nature of letter strings, which seems to be naturally modeled by n-grams. Das Ziel dieses Beitrags ist verschiedene Techniken zur Glättung von Sprachmodellen und Algorithmen zum Entwurf von Sprachmodellen auf der Basis von Entscheidungsbäumen zu untersuchen. Zu diesem Zweck verwenden wir den Brown-Korpus um Modelle von Buchstabenfolgen zu erstellen. Wir betrachten zwei Klassen von Modellen zur Textgenerierung: das n-Gramm Sprachmodell sowie verschiedene auf Entscheidungsbäumen basierende Verfahren. Im ersten Teil dieses Beitrags vergleichen wir die am häufigsten benutzten Glättungsalgorithmen angewandt auf n-Gramme. Wir folgern, daß der “bottom-up deleted interpolation”-Algorithmus am besten zur Glättung von n-Gramm Sprachmodellen geeignet ist und für große n dem “back-off”-Verfahren deutlich überlegen ist. Im zweiten Teil dieses Beitrags betrachten wir dann verschiedene Algorithmen zur Bildung von Entscheidungsbäumen. Unter diesen erzielt ein K-means-ähnlicher Algorithmus die besten Ergebnisse beim Entwurf der Fragen die der Entscheidungsbaum stellt. Für die Modellierung von Buchstabenfolgen erzielt das n-Gramm Sprachmodell aber trotzdem noch bessere Ergebnisse als alle Entscheidungsbäume. Wir glauben, daß dies durch die Fähigkeit der n-Gramme Buchstabenfolgen vorherzusagen begründet ist. Le but de cet article est d'étudier différentes techniques de lissage de modèles de langage et différents algorithmes de construction de modèles de langage à base d'arbres de décision. Pour cela, nous construisons des modèles de langage pour des caractères écrits (lettres) à partir du Brown corpus. Nous considérons deux classes de modèles pour le processus de génération du texte: le modèle de langage n-gram, et différents modèles de langage à base d'arbres de décision. Dans la première partie de l'article, nous comparons les algorithmes de lissage les plus couramment appliqués au modèle de langage n-gram. L'algorithme “bottom-up deleted interpolation” donne les meilleurs résultats pour le lissage du modèle de langage n-gram, dépassant de façon significative la technique de lissage “back-off” pour de grandes valeurs de n. Dans la seconde partie de l'article, nous considérons différents algorithmes de développement d'arbres de décision. Parmi eux, un algorithme de type classification K-means aboutit aux meilleurs résultats pour la construction des arbres de décision. Cependant, le modèle de langage n-gram fournit de meilleurs résultats que les modèles de langage à arbres de décision pour la modélisation du langage écrit. Nous croyons que cela est dû à la nature prédictive des chaı̂nes de caractères, qui semble être naturellement modélisée par les n-grams.
Gerasimos Potamianos, Frederick Jelinek
Speech Commun.1
1997 Stochastic approximation algorithms for partition function estimation of Gibbs random fields
abstract
We present an analysis of previously proposed Monte Carlo algorithms for estimating the partition function of a Gibbs random field. We show that this problem reduces to estimating one or more expectations of suitable functionals of the Gibbs states with respect to properly chosen Gibbs distributions. As expected, the resulting estimators are consistent. Certain generalizations are also provided. We study computational complexity with respect to grid size and show that Monte Carlo partition function estimation algorithms can be classified into two categories: E-type algorithms that are of exponential complexity and P-type algorithms that are of polynomial complexity, Turing reducible to the problem of sampling from the Gibbs distribution. E-type algorithms require estimating a single expectation, whereas, P-type algorithms require estimating a number of expectations with respect to Gibbs distributions which are chosen to be sufficiently "close" to each other. In the latter case, the required number of expectations is of polynomial order with respect to grid size. We compare computational complexity by using both theoretical results and simulation experiments. We determine the most efficient E-type and P-type algorithms and conclude that P-type algorithms are more appropriate for partition function estimation. We finally suggest a practical and efficient P-type algorithm for this task.
Gerasimos Potamianos, John K. Goutsias
IEEE Trans. Inf. Theory1
1993 An analysis of Monte Carlo methods for likelihood estimation of Gibbsian images
Gerasimos Potamianos, John K. Goutsias
ICASSP (5)1
1993 Partition function estimation of Gibbs random field images using Monte Carlo simulations
abstract
A Monte Carlo simulation technique for estimating the partition function of a general Gibbs random field image is proposed. By expressing the partition function as an expectation, an importance sampling approach for estimating it using Monte Carlo simulations is developed. As expected, the resulting estimators are unbiased and consistent. Computations can be performed iteratively by using simple Monte Carlo algorithms with remarkable success, as demonstrated by simulations. The work concentrates on binary, second-order Gibbs random fields defined on a rectangular lattice. However, the proposed methods can be easily extended to more general Gibbs random fields. Their potential contribution to optimal parameter estimation and hypothesis testing problems for general Gibbs random field images using a likelihood approach is anticipated.>
Gerasimos Potamianos, John K. Goutsias
IEEE Trans. Inf. Theory1
1991 A novel method for computing the partition function of Markov random field images using Monte Carlo simulations
abstract
The authors present a new Monte Carlo simulation technique for the estimation of the partition function of a general Markov random field (MRF), which results in unbiased, consistent and asymptotically efficient estimates. This technique gives extremely accurate results, as demonstrated by simulations. Use of more efficient algorithms can boost the performance and accuracy of the method, and yield more reliable estimates.>
Gerasimos Potamianos, John K. Goutsias
ICASSP1