VLDB 2026 Research / reviewers in the wild / expert
Aggelos K. Katsaggelos
dblp:k/AggelosKAggelos
· DBLP profile ↗
460ranked-venue papers
20as first author
43since 2021 · last 2026
0000-0003-4554-0070ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 392 · 16 first-author · 25 since 2021Artificial intelligence and machine learning · 33 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 4 first-author · 7 since 2021Computer networks · 15 · 1 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E2Detect: Object Detection from Event Camera via Sparse Feature Pyramid Recovery
Chris Henry, Zhu Li 0001, Aggelos K. Katsaggelos |
ISCAS | 3 |
| 2026 | Probabilistic smooth attention for deep multiple instance learning in medical imagingabstractThe Multiple Instance Learning (MIL) paradigm is attracting plenty of attention in medical imaging classification, where labeled data is scarce. MIL methods cast medical images as bags of instances (e.g. patches in whole slide images, or slices in CT scans), and only bag labels are required for training. Deep MIL approaches have obtained promising results by aggregating instance-level representations via an attention mechanism to compute the bag-level prediction. These methods typically capture both local interactions among adjacent instances and global, long-range dependencies through various mechanisms. However, they treat attention values deterministically, potentially overlooking uncertainty in the contribution of individual instances. In this work we propose a novel probabilistic framework that estimates a probability distribution over the attention values, and accounts for both global and local interactions. In a comprehensive evaluation involving eleven state-of-the-art baselines and three medical datasets, we show that our approach achieves top predictive performance in different metrics. Moreover, the probabilistic treatment of the attention provides uncertainty maps that are interpretable in terms of illness localization. Francisco M. Castro-Macías, Pablo Morales-Alvarez, Yunan Wu, Rafael Molina 0001, Aggelos K. Katsaggelos |
Pattern Recognit. | 5 |
| 2026 | JustRAIGS: Justified Referral in AI Glaucoma Screening ChallengeabstractA major contributor to permanent vision loss is glaucoma. Early diagnosis is crucial for preventing vision loss due to glaucoma, making glaucoma screening essential. A more affordable method of glaucoma screening can be achieved by applying artificial intelligence to evaluate color fundus photographs (CFPs). We present the Justified Referral in AI Glaucoma Screening (JustRAIGS) challenge to further develop these AI algorithms for glaucoma screening and to assess their efficacy. To support this challenge, we have generated a distinctive big dataset containing more than 110,000 meticulously labeled CFPs obtained from approximately 60,000 patients and 500 distinct screening centers in the USA. Our objective is to assess the practicality of creating advanced and dependable AI systems that can take a CFP as input and produce the probability of referable glaucoma, as well as outputs for glaucoma justification by integrating both binary and multi-label classification tasks. This paper presents the evaluation of solutions provided by nine teams, recognizing the team with the highest level of performance. The highest achieved score of sensitivity at a specificity level of 95% was 85%, and the highest achieved score of Hamming losses average was 0.13. Additionally, we test the top three participants' algorithms on an external dataset to validate the performance and generalization of these models. The outcomes of this research can offer valuable insights into the development of intelligent systems for detecting glaucoma. Ultimately, findings can aid in the early detection and treatment of glaucoma patients, hence decreasing preventable vision impairment and blindness caused by glaucoma. Yeganeh Madadi, Hina Raja, Koen A. Vermeer, Hans G. Lemij, Xiaoqin Huang, Gitaek Kwon, Adrian Galdran, Miguel Ángel González Ballester, Dan Presil, Kristhian Aguilar, Victor F. Cavalcante, Celso B. Carvalho, Waldir S. S. Júnior, Mateus Oliveira, Charilaos Apostolidis, Aggelos K. Katsaggelos, Tomasz Kubrak, Ángela Casado, Jónathan Heras, Marcos Ortega 0001, Lucía Ramos, Philippe Zhang, Weili Jiang, Pierre-Henri Conze, Mathieu Lamard, Gwenolé Quellec, Mostafa El Habib Daho, Madukuri Shaurya, Anumeha Varma, Siamak Yousefi |
IEEE Trans. Medical Imaging | 21 |
| 2025 | Multi-Modal Hand-to-Mouth Gesture Recognition in Activity-Oriented RGB-Thermal Footage (Student Abstract)abstractHealth-risk behaviors such as overeating and smoking have a profound impact on public health, making their monitoring and mitigation critical. Wearable RGB-Thermal cameras are being employed to monitor these behaviors by capturing hand-to-mouth (HTM) gestures, which are central to them. However, detection models relying on single modalities—either RGB or thermal—often struggle to accurately distinguish these confounding gestures due to inherent sensor limitations, such as sensitivity to lighting conditions or thermal occlusions. We present a family of fusion models that integrate RGB and thermal video data using early-, decision- , and a novel mid-fusion architecture, RGB-Thermal Fusion Video Network (RTFVNet), designed to enhance the recognition of HTM gestures associated with eating and smoking. Our evaluation shows that while decision fusion achieves the highest F1-score of 88% (0.44 TFLOPs), RTFVNet offers an optimal balance between performance (85%) and complexity (0.37 TFLOPs) for gesture classification of eating, smoking, and non-gesture activities. Glenn Fernandes, Meixi Lu, Farzad Shahabi, Aggelos K. Katsaggelos, Nabil Alshurafa |
AAAI | 5 |
| 2025 | Longitudinal Wrist PPG Analysis for Reliable Hypertension Risk Screening Using Deep LearningabstractHypertension is a leading risk factor for cardiovascular diseases. Traditional blood pressure monitoring methods are cumbersome and inadequate for continuous tracking, prompting the development of PPG-based cuffless blood pressure monitoring wearables. This study leverages deep learning models, including ResNet and Transformer, to analyze wrist PPG data collected with a smartwatch for efficient hypertension risk screening, eliminating the need for handcrafted PPG features. Using the Home Blood Pressure Monitoring (HBPM) longitudinal dataset of 448 subjects and five-fold cross-validation, our model was trained on over 68k spot-check instances from 358 subjects and tested on real-world continuous recordings of 90 subjects. The compact ResNet model with 0.124M parameters performed significantly better than traditional machine learning methods, demonstrating its effectiveness in distinguishing between healthy and abnormal cases in real-world scenarios. Jiyang Li, Ramy Hussein, Xiaoyu Li 0007, Guangpu Zhu, Aggelos K. Katsaggelos, Zijing Zeng, Yelei Li |
ICASSP | 7 |
| 2025 | Advancing Limited-Angle CT Reconstruction through Diffusion-Based Sinogram CompletionabstractLimited Angle Computed Tomography (LACT) often faces significant challenges due to missing angular information. Unlike previous methods that operate in the image domain, we propose a new method that focuses on sinogram inpainting. We leverage MR-SDEs, a variant of diffusion models that characterize the diffusion process with mean-reverting stochastic differential equations, to fill in missing angular data at the projection level. Furthermore, by combining distillation with constraining the output of the model using the pseudo-inverse of the inpainting matrix, the diffusion process is accelerated and done in a step, enabling efficient and accurate sinogram completion. A subsequent post-processing module back-projects the inpainted sinogram into the image domain and further refines the reconstruction, effectively suppressing artifacts while preserving critical structural details. Quantitative experimental results demonstrate that the proposed method achieves state-of-the-art performance in both perceptual and fidelity quality, offering a promising solution for LACT reconstruction in scientific and clinical applications. Santiago Lopez Tapia, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2025 | Enhanced heart failure mortality prediction through model-independent hybrid feature selection and explainable machine learningabstractHeart failure (HF) remains a significant public health challenge with high mortality rates. Machine learning (ML) techniques offer a promising approach to predict HF mortality, potentially improving clinical outcomes. However, the effectiveness of these techniques heavily depends on the quality and relevance of the features used. This study introduces a novel hybrid feature selection methodology that combines Extremely Randomized Trees (Extra-Trees) and non-linear correlation measures to enhance 1-year all-cause mortality prediction in HF patients using echocardiographic and key demographic data. Unlike existing feature selection methods that are often tied to specific ML models and produce inconsistent feature sets across different algorithms, our proposed approach is model-independent, ensuring robustness and generalizability. Moreover, the optimal number of predictive features is identified through loss graph inspection, leading to a compact and highly informative subset of seven features. We trained and evaluated seven widely-used ML models on both the full feature set and the selected subset, finding that most models maintained or improved their predictive performance despite an 80% reduction in features. Model interpretability was enhanced using SHapley Additive exPlanations (SHAP), allowing for a detailed examination of how individual features influence predictions. To further assess its effectiveness, we compared our methodology against widely known feature selection techniques across all seven ML models. The results underscore the superiority of our proposed feature set in accurately predicting HF mortality over conventional methods, offering new opportunities for personalized management strategies based on a streamlined and explainable feature subset. Georgios Petmezas, Vasileios E. Papageorgiou, Vassilios Vassilikos, Efstathios Pagourelias, Dimitrios Tachmatzidis, George Tsaklidis, Aggelos K. Katsaggelos, Nicos Maglaveras |
J. Biomed. Informatics | 7 |
| 2025 | Automatic Camera Movement Generation With Enhanced Immersion for Virtual CinematographyabstractUser-generated cinematic creations are gaining popularity as our daily entertainment, yet it is a challenge to master cinematography for producing immersive contents. Many existing automatic methods focus on roughly controlling predefined shot types or movement patterns, which struggle to engage viewers with the actor's circumstances. Real-world cinematographic rules show that directors can create immersion by comprehensively synchronizing the camera with the actor. Inspired by this strategy, we propose a deep camera control framework that enables actor-camera synchronization in three aspects, considering frame aesthetics, spatial action, and emotional status in the 3D virtual stage. Following rule-of-thirds, our framework first modifies the initial camera placement to position the actor aesthetically. This adjustment is facilitated by a weakly-supervised adjustor that analyzes frame composition via camera projection. We then design a GAN model that can adversarially synthesize fine-grained camera movement based on the actor's action and psychological state, using an encoder-decoder generator to map kinematics and emotional variables into camera trajectories. Moreover, we incorporate a regularizer to align the generated stylistic variances with specific emotional categories and intensities. The experimental results show that our proposed method yields immersive cinematic videos of high quality, both quantitatively and qualitatively. Live examples can be found in the supplementary video. Xinyi Wu 0001, Haohong Wang, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 3 |
| 2025 | LCNet: Lightweight Cycle Network Driven by Physical and Deep Prior for Compressed SensingabstractDeep learning (DL) networks have recently achieved excellent performance on image compressed sensing. However, most existing methods rely on burdened and complex network structures, resulting in significant computational and storage requirements that defeat the purpose of compressed sensing. This severely hinders their applicability in real-world resource-limited devices. In this paper, a lightweight cycle network driven by physical and deep priors for image compressed sensing is proposed which integrates the learning of the sensing matrix and compressive image reconstruction. Specifically, the regularization terms and a likelihood term derived from the physical observation model are learned in an end-to-end cycle network, simultaneously estimating the reconstructed image and sensing matrix in the image and feature domains. Moreover, a dual-domain fusion reconstruction module is proposed. It creates simulated measurement residuals for enhancing reconstruction in the compressed domain, which leads to high reconstruction performance and reduces computational load by bonding together the compressed image domains in the cyclic network. Extensive experiments demonstrate that our model delivers superior performance and alleviates model complexity, which is of great importance in low-budget applications. Shuowen Yang, Fernando Pérez-Bueno, Hanlin Qin, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 5 |
| 2024 | CPDR: Towards Highly-Efficient Salient Object Detection via Crossed Post-decoder Refinement
Yijie Li 0003, Hewei Wang 0001, Aggelos K. Katsaggelos |
BMVC | 3 |
| 2024 | Real-World Atmospheric Turbulence Correction Via Domain AdaptationabstractAtmospheric turbulence, a common phenomenon in daily life, is primarily caused by the uneven heating of the Earth’s surface. This phenomenon results in distorted and blurred acquired images or videos and can significantly impact downstream vision tasks, particularly those that rely on capturing clear, stable images or videos from outdoor environments, such as accurately detecting or recognizing objects. Therefore, people have proposed ways to simulate atmospheric turbulence and designed effective deep learning-based methods to remove the atmospheric turbulence effect. However, these synthesized turbulent images can not cover all the range of real-world turbulence effects. Though the models have achieved great performance for synthetic scenarios, there always exists a performance drop when applied to real-world cases. Moreover, reducing real-world turbulence is a more challenging task as there are no clean ground truth counter-parts provided to the models during training. In this paper, we propose a real-world atmospheric turbulence mitigation model under a domain adaptation framework, which links the supervised simulated atmospheric turbulence correction with the unsupervised real-world atmospheric turbulence correction. We will show our proposed method enhances performance in real-world atmospheric turbulence scenarios, improving both image quality and downstream vision tasks. Xijun Wang 0003, Santiago Lopez Tapia, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2024 | Bayesian Blind Image Deconvolution using an Hyperbolic-Secant priorabstractIn this paper we propose the use of the Hyperbolic Secant (HS) distribution as a prior for the Blind Image Deconvolution (BID) problem. It is well-known that when high-pass filters are applied to natural images, the resulting coefficients are sparse. We leverage this property using the HS distribution, a seldom explored Super Gaussian distribution with suitable properties for this problem. Using the Pólya-Gamma distribution, we derive an explicit Gaussian Scale Mixture representation. This representation is then used to propose a novel variational Bayesian algorithm that outperforms state-of-the-art BID methods. Francisco M. Castro-Macías, Fernando Pérez-Bueno, Miguel Vega, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 6 |
| 2024 | Sm: enhanced localization in Multiple Instance Learning for medical imaging classificationabstractMultiple Instance Learning (MIL) is widely used in medical imaging classification to reduce the labeling effort.
While only bag labels are available for training, one typically seeks predictions at both bag and instance levels (classification and localization tasks, respectively). Early MIL methods treated the instances in a bag independently. Recent methods account for global and local dependencies among instances. Although they have yielded excellent results in classification, their performance in terms of localization is comparatively limited. We argue that these models have been designed to target the classification task, while implications at the instance level have not been deeply investigated. Motivated by a simple observation -- that neighboring instances are likely to have the same label -- we propose a novel, principled, and flexible mechanism to model local dependencies. It can be used alone or combined with any mechanism to model global dependencies (e.g., transformers). A thorough empirical validation shows that our module leads to state-of-the-art performance in localization while being competitive or superior in classification. Our code is at https://github.com/Franblueee/SmMIL. Francisco M. Castro-Macías, Pablo Morales-Alvarez, Yunan Wu, Rafael Molina 0001, Aggelos K. Katsaggelos |
NeurIPS | 5 |
| 2024 | Hyperbolic Secant representation of the logistic function: Application to probabilistic Multiple Instance Learning for CT intracranial hemorrhage detectionabstractMultiple Instance Learning (MIL) is a weakly supervised paradigm that has been successfully applied to many different scientific areas and is particularly well suited to medical imaging. Probabilistic MIL methods, and more specifically Gaussian Processes (GPs), have achieved excellent results due to their high expressiveness and uncertainty quantification capabilities. One of the most successful GP-based MIL methods, VGPMIL, resorts to a variational bound to handle the intractability of the logistic function. Here, we formulate VGPMIL using Pólya-Gamma random variables. This approach yields the same variational posterior approximations as the original VGPMIL, which is a consequence of the two representations that the Hyperbolic Secant distribution admits. This leads us to propose a general GP-based MIL method that takes different forms by simply leveraging distributions other than the Hyperbolic Secant one. Using the Gamma distribution we arrive at a new approach that obtains competitive or superior predictive performance and efficiency. This is validated in a comprehensive experimental study including one synthetic MIL dataset, two well-known MIL benchmarks, and a real-world medical problem. We expect that this work provides useful ideas beyond MIL that can foster further research in the field. Francisco M. Castro-Macías, Pablo Morales-Alvarez, Yunan Wu, Rafael Molina 0001, Aggelos K. Katsaggelos |
Artif. Intell. | 5 |
| 2024 | An end-to-end approach to combine attention feature extraction and Gaussian Process models for deep multiple instance learning in CT hemorrhage detection
Jose Pérez-Cano, Yunan Wu, Arne Schmidt 0005, Miguel López-Pérez, Pablo Morales-Alvarez, Rafael Molina 0001, Aggelos K. Katsaggelos |
Expert Syst. Appl. | 7 |
| 2024 | A deep learning method for predicting the COVID-19 ICU patient outcome fusing X-rays, respiratory sounds, and ICU parameters
Yunan Wu, Bruno Miguel Machado Rocha, Evangelos Kaimakamis, Grigorios-Aris Cheimariotis, Georgios Petmezas, Evangelos Chatzis, Vassilis Kilintzis, Leandros Stefanopoulos, Diogo Pessoa, Alda Marques, Paulo Carvalho 0001, Rui Pedro Paiva, Serafeim-Chrysovalantis Kotoulas, Militsa Bitzani, Aggelos K. Katsaggelos, Nicos Maglaveras |
Expert Syst. Appl. | 15 |
| 2024 | Focused active learning for histopathological image classification
Arne Schmidt 0005, Pablo Morales-Alvarez, Lee A. D. Cooper, Lee A. Newberg, Andinet Enquobahrie, Rafael Molina 0001, Aggelos K. Katsaggelos |
Medical Image Anal. | 7 |
| 2024 | Real-Time Lightweight Video Super-Resolution With RRED-Based Perceptual ConstraintabstractReal-time video services are gaining popularity in our daily life, yet limited network bandwidth can constrain the delivered video quality. Video Super Resolution (VSR) technology emerges as a key solution to enhance user experience by reconstructing high-resolution (HR) videos. The existing real-time VSR frameworks have primarily emphasized spatial quality metrics like PSNR and SSIM, which often lack consideration of temporal coherence, a critical factor for accurately reflecting the overall quality of super-resolved videos. Inspired by Video Quality Assessment (VQA) strategies, we propose a dual-frame training framework and a lightweight multi-branch network to address VSR processing in real time. Such designs thoroughly leverage the spatio-temporal correlations between consecutive frames so as to ensure efficient video restoration. Furthermore, we incorporate ST-RRED, a powerful VQA approach that separately measures spatial and temporal consistency aligning with human perception principles, into our loss functions. This guides us to synthesize quality-aware perceptual features across both space and time for realistic reconstruction. Our model demonstrates remarkable efficiency, achieving near real-time processing of 4K videos. Compared to the state-of-the-art lightweight model MRVSR, ours is more compact and faster, 60% smaller in size (0.483M vs. 1.21M parameters), and 106% quicker (96.44fps vs. 46.7fps on 1080p frames), with significantly improved perceptual quality. Xinyi Wu 0001, Santiago Lopez Tapia, Xijun Wang 0003, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Development of a Miniaturized Mechanoacoustic Sensor for Continuous, Objective Cough Detection, Characterization and Physiologic Monitoring in Children With Cystic FibrosisabstractCough is an important symptom in children with acute and chronic respiratory disease. Daily cough is common in Cystic Fibrosis (CF) and increased cough is a symptom of pulmonary exacerbation. To date, cough assessment is primarily subjective in clinical practice and research. Attempts to develop objective, automatic cough counting tools have faced reliability issues in noisy environments and practical barriers limiting long-term use. This single-center pilot study evaluated usability, acceptability and performance of a mechanoacoustic sensor (MAS), previously used for cough classification in adults, in 36 children with CF over brief and multi-day periods in four cohorts. Children whose health was at baseline and who had symptoms of pulmonary exacerbation were included. We trained, validated, and deployed custom deep learning algorithms for accurate cough detection and classification from other vocalization or artifacts with an overall area under the receiver-operator characteristic curve (AUROC) of 0.96 and average precision (AP) of 0.93. Child and parent feedback led to a redesign of the MAS towards a smaller, more discreet device acceptable for daily use in children. Additional improvements optimized power efficiency and data management. The MAS's ability to objectively measure cough and other physiologic signals across clinic, hospital, and home settings is demonstrated, particularly aided by an AUROC of 0.97 and AP of 0.96 for motion artifact rejection. Examples of cough frequency and physiologic parameter correlations with participant-reported outcomes and clinical measurements for individual patients are presented. The MAS is a promising tool in objective longitudinal evaluation of cough in children with CF. Andreas Tzavelis, John Palla, Radhika Mathur, Brittany Bedford, Yung-Hsuan Wu, Jacob Trueb, Hee-Sup Shin, Hany M. Arafa, Hyoyoung Jeong, Jay Young Kwak, Jennifer Chiang, Sydney Schulz, Tina M. Carter, Vittobai Rangaraj, Aggelos K. Katsaggelos, Susanna A. McColley, John A. Rogers |
IEEE J. Biomed. Health Informatics | 16 |
| 2024 | Automated Adaptive Cinematography for User Interaction in Open WorldabstractAdvancements in wearable technology and their capacity to interpret user movements, transforming them into interactive actions in virtual environments, have sparked an increased demand for user flexibility within these spaces. A direct outcome of this growing trend is the imperative need for automated cinematography in expansive, open-world scenarios. Nevertheless, the task of interpreting these interactive sequences through automated cinematography in unconstrained environments involves significant computational challenges. In response to this, we introduce the Automated Adaptive Cinematography for Open-world Generative Adversarial Network (AACOGAN) -an innovative solution that addresses these issues. Contrary to traditional models, which require comprehensive prior knowledge about scenes, characters, and objects, AACOGAN identifies and models the relationships among user interactions, object positions, and camera movements during the process of user engagement. This novel approach allows the model to function effectively even in open-world scenarios riddled with numerous uncertain factors. In the experimental phase, we developed and employed theMineStory Dataset, designed specifically for automatic cinematography in open-world scenarios. We devised and implemented novel metrics that are more congruent with the distinctive features of open-world scenarios. These innovative metrics provide a more nuanced understanding of the performance and effectiveness of our proposed method. Experimental findings substantiate that AACOGAN significantly enhances automatic cinematography performance within open-world contexts, including an average augmentation of 73% in the correlation between user interactions and camera trajectories, and an increase of up to 32.9% in the quality of multi-focus scenes. Therefore, AACOGAN emerges as an efficient, and innovative solution for creating appropriate camera shots in a myriad of interactive motions in open-world scenarios. An exemplary video footage can be found athttps://youtu.be/pbSHF-uxomw. Zixiao Yu, Xinyi Wu 0001, Haohong Wang, Aggelos K. Katsaggelos, Jian Ren 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | A Joint Intensity-Neuromorphic Event Imaging System With Bandwidth-Limited Communication ChannelabstractWe present a novel adaptive multimodal intensity-event algorithm to optimize an overall objective of object tracking under bit rate constraints for a host-chip architecture. The chip is a computationally resource-constrained device acquiring high-resolution intensity frames and events, while the host is capable of performing computationally expensive tasks. We develop a joint intensity-neuromorphic event rate-distortion compression framework with a quadtree (QT)-based compression of intensity and events scheme. The goal of this compression framework is to optimally allocate bits to the intensity frames and neuromorphic events based on the minimum distortion at a given communication channel capacity. The data acquisition on the chip is driven by the presence of objects of interest in the scene as detected by an object detector. The most informative intensity and event data are communicated to the host under rate constraints so that the best possible tracking performance is obtained. The detection and tracking of objects in the scene are done on the distorted data at the host. Intensity and events are jointly used in a fusion framework to enhance the quality of the distorted images, in order to improve the object detection and tracking performance. The performance assessment of the overall system is done in terms of the multiple object tracking accuracy (MOTA) score. Compared with using intensity modality only, there is an improvement in MOTA using both these modalities in different scenarios. Srutarshi Banerjee, Henry Chopp, Zihao W. Wang, Oliver Cossairt, Aggelos K. Katsaggelos |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Crowdsourcing Segmentation of Histopathological Images Using Annotations Provided by Medical Students
Miguel López-Pérez, Pablo Morales-Alvarez, Lee A. D. Cooper, Rafael Molina 0001, Aggelos K. Katsaggelos |
AIME | 5 |
| 2023 | Thermal Spread Functions (TSF): Physics-Guided Material ClassificationabstractRobust and non-destructive material classification is a challenging but crucial first-step in numerous vision applications. We propose a physics-guided material classification framework that relies on thermal properties of the object. Our key observation is that the rate of heating and cooling of an object depends on the unique intrinsic properties of the material, namely the emissivity and diffusivity. We leverage this observation by gently heating the objects in the scene with a low-power laser for a fixed duration and then turning it off, while a thermal camera captures measurements during the heating and cooling process. We then take this spatial and temporal “thermal spread function” (TSF) to solve an inverse heat equation using the finite-differences approach, resulting in a spatially varying estimate of diffusivity and emissivity. These tuples are then used to train a classifier that produces a fine-grained material label at each spatial pixel. Our approach is extremely simple requiring only a small light source (low power laser) and a thermal camera, and produces robust classification results with 86% accuracy over 16 classes11Code: https://github.com/aniketdashpute/TSF. Aniket Dashpute, Vishwanath Saragadam, Emma Alexander, Florian Willomitzer, Aggelos K. Katsaggelos, Ashok Veeraraghavan, Oliver Cossairt |
CVPR | 5 |
| 2023 | BKinD-3D: Self-Supervised 3D Keypoint Discovery from Multi-View VideosabstractQuantifying motion in 3D is important for studying the behavior of humans and other animals, but manual pose annotations are expensive and time-consuming to obtain. Self-supervised keypoint discovery is a promising strategy for estimating 3D poses without annotations. However, current keypoint discovery approaches commonly process single 2D views and do not operate in the 3D space. We propose a new method to perform self-supervised keypoint discovery in 3D from multi-view videos of behaving agents, without any keypoint or bounding box supervision in 2D or 3D. Our method, BKinD-3D, uses an encoder-decoder architecture with a 3D volumetric heatmap, trained to reconstruct spatiotemporal differences across multiple views, in addition to joint length constraints on a learned 3D skeleton of the subject. In this way, we discover keypoints without requiring manual supervision in videos of humans and rats, demonstrating the potential of 3D keypoint discovery for studying behavior. Jennifer J. Sun, Lili Karashchuk, Amil Dravid, Serim Ryou, Sonia Fereidooni, John C. Tuthill, Aggelos K. Katsaggelos, Bingni W. Brunton, Georgia Gkioxari, Ann Kennedy, Yisong Yue, Pietro Perona |
CVPR | 7 |
| 2023 | Spiking GLOM: Bio-Inspired Architecture for Next-Generation Object RecognitionabstractToday, artificial neural networks (ANNs) have demonstrated extraordinary abilities in many cognition tasks. Nevertheless, the limitations of many ANN-based techniques are evident, such as the low energy efficiency and the lack of interpretability. To alleviate these problems, researchers have directed their attention to bio-inspired models, including energy-efficient Spiking Neural Networks (SNNs) and the GLOM model representing part-whole hierarchies in neural networks. In this paper, we propose a novel bio-inspired solution to next-generation object recognition. Specifically, we propose an energy-efficient and interpretable model – Spiking GLOM by introducing spiking neurons and neuronal dynamics into the GLOM model. Moreover, we evaluate our model and its variants on CIFAR-10. Extensive experiments demonstrate the effectiveness of our proposed models for object recognition and show the superiority of our models in energy efficiency and interpretability. Srutarshi Banerjee, Henry Chopp, Aggelos K. Katsaggelos, Oliver Cossairt |
ICIP | 4 |
| 2023 | Deep Robust Image Restoration Using the Moore-Penrose Blur InverseabstractThis paper proposes a deep learning model for robust image restoration when the degradation is not precisely known. We show how the Moore-Penrose pseudo-inverse of a blur convolution operator can be approximated by a Wiener filter’s impulse response. The image restoration problem is then cast as the learning of a residual on the frequencies where the blurring filter is zero which, when added to the Wiener restoration, will satisfy the image formation model. A Dynamic Filter Network removes artifacts introduced by inaccurate blur estimations and other image formation model inconsistencies. The experiments conducted on synthetic and real image datasets assert the performance and robustness of the proposed method and show its superiority to existing ones. Santiago Lopez Tapia, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2023 | Variational Deep Atmospheric Turbulence Correction for VideoabstractThis paper presents a novel variational deep-learning approach for video atmospheric turbulence correction. We modify and tailor a Nonlinear Activation Free Network to video restoration. By including it in a variational inference framework, we boost the model’s performance and stability. This is achieved through conditioning the model on features extracted by a variational autoencoder (VAE). Furthermore, we enhance these features by making the encoder of the VAE include information pertinent to the image formation via a new loss based on the prediction of parameters of the geometrical distortion and the spatially variant blur responsible for the video sequence degradation. Experiments on a comprehensive synthetic video dataset demonstrate the effectiveness and reliability of the proposed method and validate its superiority compared to existing state-of-the-art approaches. Santiago Lopez Tapia, Xijun Wang 0003, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2023 | Deep Bayesian Blind Color Deconvolution of Histological ImagesabstractHistological images are often tainted with two or more stains to reveal their underlying structures and conditions. Blind Color Deconvolution (BCD) techniques separate colors (stains) and structural information (concentrations), which is useful for the processing, data augmentation, and classification of such images. Classical BCD methods rely on a complicated optimization procedure that has to be carried out on each image independently, i.e., they are not amortized methods. In contrast, once they have been trained, deep neural networks can be used in a fast, amortized manner on unseen inputs. Unfortunately, the lack of large databases of ground truth color and concentrations has limited the development of deep models for BCD. In this work, we propose a deep variational Bayesian BCD neural network (BCD-Net) for stain separation and concentration estimation. BCD-Net is trained by maximizing the evidence lower bound of the observed images, which does not require the use of ground truth examples of stains and concentrations. Results obtained using two multicenter databases (Camelyon-17 and a stain separation benchmark) demonstrate the effectiveness of BCD-Net in the stain separation tasks, while drastically reducing the computation time compared to classical non-amortized methods. Shuowen Yang, Fernando Pérez-Bueno, Francisco M. Castro-Macías, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2023 | Smooth Attention for Deep Multiple Instance Learning: Application to CT Intracranial Hemorrhage Detection
Yunan Wu, Francisco M. Castro-Macías, Pablo Morales-Alvarez, Rafael Molina 0001, Aggelos K. Katsaggelos |
MICCAI (5) | 5 |
| 2023 | A Novel Automatic Content Generation and Optimization FrameworkabstractWith the rapid growth of IoT multimedia devices, more and more content is being delivered in multimedia forms which generally requires more effort and resources from users to create and could be challenging to streamline. In this article, we propose Text2Animation (T2A), a framework that helps to generate complex multimedia content, and animation, from the simple textual script input. By leveraging recent advances in computational cinematography and video understanding, the purposed workflow further reduces associated knowledge requirements dramatically for cinematography. By jointly incorporating the fidelity and aesthetic models, T2A jointly considers the comprehensiveness of the visual presentation of the input script and the compliance of generated video with given cinematography specifications. The virtual camera placement in a 3-D environment is mapped into an optimization problem that can be resolved by using dynamic programming to achieve the lowest computational complexity. Experimental results show that T2A can reduce the manual animation production process by around 74%, and the new optimization framework can improve the perceptual quality of the output video by up to 35%. More video footage can be found athttps://www.youtube.com/watch?v=MMTJbmWL3gs. Zixiao Yu, Haohong Wang, Aggelos K. Katsaggelos, Jian Ren 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Probabilistic fusion of crowds and experts for the search of gravitational waves
Pablo Ruiz 0002, Pablo Morales-Alvarez, Scott Coughlin, Rafael Molina 0001, Aggelos K. Katsaggelos |
Knowl. Based Syst. | 5 |
| 2023 | Reinforcement Learning for Adaptive Video Compressive SensingabstractWe apply reinforcement learning to video compressive sensing to adapt the compression ratio. Specifically, video snapshot compressive imaging (SCI), which captures high-speed video using a low-speed camera is considered in this work, in which multiple ( B ) video frames can be reconstructed from a snapshot measurement. One research gap in previous studies is how to adapt B in the video SCI system for different scenes. In this article, we fill this gap utilizing reinforcement learning (RL). An RL model, as well as various convolutional neural networks for reconstruction, are learned to achieve adaptive sensing of video SCI systems. Furthermore, the performance of an object detection network using directly the video SCI measurements without reconstruction is also used to perform RL-based adaptive video compressive sensing. Our proposed adaptive SCI method can thus be implemented in low cost and real time. Our work takes the technology one step further towards real applications of video SCI. Sidi Lu, Xin Yuan 0002, Aggelos K. Katsaggelos, Weisong Shi |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2022 | Visual Explanations for Convolutional Neural Networks via Latent Traversal of Generative Adversarial Networks (Student Abstract)abstractLack of explainability in artificial intelligence, specifically deep neural networks, remains a bottleneck for implementing models in practice. Popular techniques such as Gradient-weighted Class Activation Mapping (Grad-CAM) provide a coarse map of salient features in an image, which rarely tells the whole story of what a convolutional neural network(CNN) learned. Using COVID-19 chest X-rays, we present a method for interpreting what a CNN has learned by utilizing Generative Adversarial Networks (GANs). Our GAN framework disentangles lung structure from COVID-19 features. Using this GAN, we can visualize the transition of a pair of COVID negative lungs in a chest radiograph to a COVID positive pair by interpolating in the latent space of the GAN, which provides fine-grained visualization of how the CNN responds to varying features within the lungs. Amil Dravid, Aggelos K. Katsaggelos |
AAAI | 2 |
| 2022 | Investigating the Potential of Auxiliary-Classifier Gans for Image Classification in Low Data RegimesabstractGenerative Adversarial Networks (GANs) have shown promise in augmenting datasets and boosting convolutional neural network (CNN) performance on image classification tasks. But they introduce more hyperparameters to tune as well as the need for additional time and computational power to train, supplementary to the CNN. In this work, we examine the potential for Auxiliary-Classifier GANs (AC-GANs) as a ’one-stop-shop’ architecture for image classification, particularly in low data regimes. Additionally, we explore modifications to the typical AC-GAN framework, changing the generator’s latent space sampling scheme and employing a Wasserstein loss with gradient penalty to stabilize the simultaneous training of image synthesis and classification. Through experiments on images of varying resolutions and complexity, we demonstrate that AC-GANs show promise in image classification, achieving competitive performance with standard CNNs. These methods can be employed as an ’all-in-one’ framework with particular utility in the absence of large amounts of training data. Amil Dravid, Florian Schiffers, Yunan Wu, Oliver Cossairt, Aggelos K. Katsaggelos |
ICASSP | 5 |
| 2022 | Event - Driven Tactile Learning with Location Spiking NeuronsabstractThe sense of touch is essential for a variety of daily tasks. New advances in event-based tactile sensors and Spiking Neural Networks (SNNs) spur the research in event-driven tactile learning. However, SNN -enabled event-driven tactile learning is still in its infancy due to the limited representative abilities of existing spiking neurons and high spatio-temporal complexity in the data. In this paper, to improve the representative capabilities of existing spiking neurons, we propose a novel neuron model called “location spiking neuron”, which enables us to extract features of event-based data in a novel way. Moreover, based on the classical Time Spike Response Model (TSRM), we develop a specific location spiking neuron model - Location Spike Response Model (LSRM) that serves as a new building block of SNNs11The TSRM is the classical SRM in the literature. We add the character “T” to highlight its difference with the LSRM.• Furthermore, we propose a hybrid model which combines an SNN with TSRM neurons and an SNN with LSRM neurons to capture the complex spatio-temporal dependencies in the data. Extensive experiments demonstrate the significant improvements of our models over other works on event-driven tactile learning and show the superior energy efficiency of our models and location spiking neurons, which may unlock their potential on neuromorphic hardware. Srutarshi Banerjee, Henry Chopp, Aggelos K. Katsaggelos, Oliver Cossairt |
IJCNN | 4 |
| 2022 | Guided Event Filtering: Synergy Between Intensity Images and Neuromorphic Events for High Performance ImagingabstractMany visual and robotics tasks in real-world scenarios rely on robust handling of high speed motion and high dynamic range (HDR) with effectively high spatial resolution and low noise. Such stringent requirements, however, cannot be directly satisfied by a single imager or imaging modality, rather by multi-modal sensors with complementary advantages. In this paper, we address high performance imaging by exploring the synergy between traditional frame-based sensors with high spatial resolution and low sensor noise, and emerging event-based sensors with high speed and high dynamic range. We introduce a novel computational framework, termed Guided Event Filtering (GEF), to process these two streams of input data and output a stream of super-resolved yet noise-reduced events. To generate high quality events, GEF first registers the captured noisy events onto the guidance image plane according to our flow model. it then performs joint image filtering that inherits the mutual structure from both inputs. Lastly, GEF re-distributes the filtered event frame in the space-time volume while preserving the statistical characteristics of the original events. When the guidance images under-perform, GEF incorporates an event self-guiding mechanism that resorts to neighbor events for guidance. We demonstrate the benefits of GEF by applying the output high quality events to existing event-based algorithms across diverse application categories, including high speed object tracking, depth estimation, high frame-rate video synthesis, and super resolution/HDR/color image restoration. Peiqi Duan 0002, Zihao W. Wang, Boxin Shi, Oliver Cossairt, Tiejun Huang 0001, Aggelos K. Katsaggelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Scalable Variational Gaussian Processes for Crowdsourcing: Glitch Detection in LIGO
Pablo Morales-Alvarez, Pablo Ruiz 0002, Scott Coughlin, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Lossy Event Compression Based On Image-Derived Quad Trees And Poisson Disk SamplingabstractEvent cameras have provided new opportunities for tackling visual tasks under challenging scenarios over conventional RGB cameras. However, not much focus has been given on event compression algorithms. The main challenge for compressing events is its unique asynchronous form. To address this problem, we propose a novel event compression algorithm based on a quad tree (QT) segmentation map derived from the adjacent intensity images. The QT informs 2D spatial priority within the 3D space-time volume. In the event encoding step, events are first aggregated over time to form polarity-based event histograms. The histograms are then variably sampled via Poisson Disk Sampling prioritized by the QT based segmentation map. Next, differential encoding and run length encoding are employed for encoding the spatial and polarity information of the sampled events, respectively, followed by Huffman encoding to produce the final encoded events. Our algorithm achieves greater than $6 \times$ higher compression compared to the state of the art. Srutarshi Banerjee, Zihao W. Wang, Henry Chopp, Oliver Cossairt, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2021 | A Data Fusion Method For The Delayering Of X-Ray Fluorescence Images Of Painted Works Of ArtabstractIn this manuscript, we address the problem of studying layer structure in X-ray Fluorescence (XRF) elemental maps of paintings through the incorporation of reflectance imaging spectral data in the visible or near IR range. We propose a conceptually flexible approach, which involves an initial clustering step for the visible hyperspectral reflectance data (RIS) and the formation of a synthetic surface XRF image. Considering the difference of the full and synthetic surface XRF images, surface and subsurface correlated features are then identified. Results are demonstrated on real and simulated data. Lionel Fiske, Aggelos K. Katsaggelos, Maurice C. G. Aalders, Matthias Alfeld, Marc Walton, Oliver Cossairt |
ICIP | 2 |
| 2021 | Human Vision-Like Robust Object RecognitionabstractPrevious research always solely utilizes Artificial Neural Networks (ANNs) or Spiking Neural Networks (SNNs) for object recognition. However, evidence in neuroscience suggests that the visual processing in human vision is performed hierarchically in the combination of analog and digital processing. To construct a more human vision-like object recognition system, we propose a general hierarchical ANN-SNN model. We evaluate our model and its variants on two popular datasets to show its effectiveness, robustness, efficiency, and generality. Extensive experiments clearly demonstrate the superiority of our proposed models for robust object recognition. Srutarshi Banerjee, Henry Chopp, Aggelos K. Katsaggelos, Oliver Cossairt |
ICIP | 5 |
| 2021 | Skinscan: Low-Cost 3D-Scanning for Dermatologic Diagnosis and DocumentationabstractThe utilization of computational photography becomes increasingly essential in the medical field. Today, imaging techniques for dermatology range from two-dimensional (2D) color imagery with a mobile device to professional clinical imaging systems measuring additional detailed three-dimensional (3D) data. The latter are commonly expensive and not accessible to a broad audience. In this work, we propose a novel system and software framework that relies only on low-cost (and even mobile) commodity devices present in every household to measure detailed 3D information of the human skin with a 3D-gradient-illumination-based method. We believe that our system has great potential for early-stage diagnosis and monitoring of skin diseases, especially in vastly populated or underdeveloped areas. Merlin A. Nau, Florian Schiffers, Andreas K. Maier, Jack Tumblin, Marc Walton, Aggelos K. Katsaggelos, Florian Willomitzer, Oliver Cossairt |
ICIP | 8 |
| 2021 | Improving Acquisition Speed of X-Ray Ptychography Through Spatial Undersampling and RegularizationabstractX-ray ptychography is one of the versatile techniques for nanometer resolution imaging. The magnitude of the diffraction patterns is recorded on a detector, and the phase of the diffraction patterns is estimated using phase retrieval techniques. Most phase retrieval algorithms make the solution well-posed by relying on the constraints imposed by the overlapping region between neighboring diffraction pattern samples. As the overlap between neighboring diffraction patterns reduces, the problem becomes ill-posed, and the object cannot be recovered. To avoid the ill-posedness, we investigate the effect of regularizing the phase retrieval algorithm with image priors for various overlap ratios between the neighboring diffraction patterns. We show that the object can be faithfully reconstructed at low overlap ratios by regularizing the phase retrieval algorithm with image priors such as Total-Variation prior and Structure Tensor Prior. We also show the effectiveness of our proposed algorithm on real data acquired from an IC chip with a coherent X-ray beam. Prasan A. Shedligeri, Florian Schiffers, Semih Barutcu, Pablo Ruiz 0002, Aggelos K. Katsaggelos, Oliver Cossairt |
ICIP | 5 |
| 2021 | Combining Attention-Based Multiple Instance Learning and Gaussian Processes for CT Hemorrhage Detection
Yunan Wu, Arne Schmidt 0005, Enrique Hernández-Sánchez, Rafael Molina 0001, Aggelos K. Katsaggelos |
MICCAI (2) | 5 |
| 2020 | Joint Filtering of Intensity Images and Neuromorphic Events for High-Resolution Noise-Robust ImagingabstractWe present a novel computational imaging system with high resolution and low noise. Our system consists of a traditional video camera which captures high-resolution intensity images, and an event camera which encodes high-speed motion as a stream of asynchronous binary events. To process the hybrid input, we propose a unifying framework that first bridges the two sensing modalities via a noise-robust motion compensation model, and then performs joint image filtering. The filtered output represents the temporal gradient of the captured space-time volume, which can be viewed as motion-compensated event frames with high resolution and low noise. Therefore, the output can be widely applied to many existing event-based algorithms that are highly dependent on spatial resolution and noise robustness. In experimental results performed on both publicly available datasets as well as our contributing RGB-DAVIS dataset, we show systematic performance improvement in applications such as high frame-rate video synthesis, feature/corner detection and tracking, as well as high dynamic range image reconstruction. Zihao W. Wang, Peiqi Duan 0002, Oliver Cossairt, Aggelos K. Katsaggelos, Tiejun Huang 0001, Boxin Shi |
CVPR | 4 |
| 2020 | Super Gaussian Priors for Blind Color Deconvolution of Histological ImagesabstractColor deconvolution aims at separating multi-stained images into single stained ones. In digital histopathological images, true stain color vectors vary between images and need to be estimated to obtain stain concentrations and separate stain bands. These band images can be used for image analysis purposes and, once normalized, utilized with other multi-stained images (from different laboratories and obtained using different scanners) for classification purposes. In this paper we propose the use of Super Gaussian (SG) priors for each stain concentration together with the similarity to a given reference matrix for the color vectors. Variational inference and an evidence lower bound are utilized to automatically estimate all the latent variables. The proposed methodology is tested on real images and compared to classical and state-of-the-art methods for histopathological blind image color deconvolution. Fernando Pérez-Bueno, Miguel Vega, Valery Naranjo, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2020 | Variational Bayesian Blind Color Deconvolution of Histopathological ImagesabstractMost whole-slide histological images are stained with two or more chemical dyes. Slide stain separation or color deconvolution is a crucial step within the digital pathology workflow. In this paper, the blind color deconvolution problem is formulated within the Bayesian framework. Starting from a multi-stained histological image, our model takes into account both spatial relations among the concentration image pixels and similarity between a given reference color-vector matrix and the estimated one. Using Variational Bayes inference, three efficient new blind color deconvolution methods are proposed which provide automated procedures to estimate all the model parameters in the problem. A comparison with classical and current state-of-the-art color deconvolution algorithms using real images has been carried out demonstrating the superiority of the proposed approach. Natalia Hidalgo-Gavira, Javier Mateos, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 5 |
| 2020 | Adaptive Image Sampling Using Deep Learning and Its Application on X-Ray Fluorescence Image ReconstructionabstractThis paper presents an adaptive image sampling algorithm based on Deep Learning (DL). It consists of an adaptive sampling mask generation network which is jointly trained with an image inpainting network. The sampling rate is controlled by the mask generation network, and a binarization strategy is investigated to make the sampling mask binary. In addition to the image sampling and reconstruction process, we show how it can be extended and used to speed up raster scanning such as the X-Ray fluorescence (XRF) image scanning process. Recently XRF laboratory-based systems have evolved into lightweight and portable instruments thanks to technological advancements in both X-Ray generation and detection. However, the scanning time of an XRF image is usually long due to the long exposure requirements (e.g., 100 μs - 1 ms per point). We propose an XRF image in painting approach to address the long scanning times, thus speeding up the scanning process, while being able to reconstruct a high quality XRF image. The proposed adaptive image sampling algorithm is applied to the RGB image of the scanning target to generate the sampling mask. The XRF scanner is then driven according to the sampling mask to scan a subset of the total image pixels. Finally, we inpaint the scanned XRF image by fusing the RGB image to reconstruct the full scan XRF image. The experiments show that the proposed adaptive sampling algorithm is able to effectively sample the image and achieve a better reconstruction accuracy than that of existing methods. Qiqin Dai, Henry Chopp, Emeline Pouyet, Oliver Cossairt, Marc Walton, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 6 |
| 2019 | Direct Estimation of Weights and Efficient Training of Deep Neural Networks without SGDabstractWe argue that learning a hierarchy of features in a hierarchical dataset requires lower layers to approach convergence faster than layers above them. We show that, if this assumption holds, we can analytically approximate the outcome of stochastic gradient descent (SGD) for each layer. We find that the weights should converge to a class-based PCA, with some weights in every layer dedicated to principal components of each label class. The class-based PCA allows us to train layers directly, without SGD, often leading to a dramatic decrease in training complexity. We demonstrate the effectiveness of this by using our results to replace one and two convolutional layers in networks trained on MNIST, CIFAR10 and CIFAR100 datasets, showing that our method achieves performance superior or comparable to similar architectures trained using SGD. Nima Dehmamy, Neda Rohani, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2019 | Multi-frame Super-resolution for Time-of-flight ImagingabstractRecently, time-of-flight (ToF) sensors have emerged as a promising three-dimensional sensing technology that can be manufactured inexpensively in a compact size. However, current state-of-the-art ToF sensors suffer from low spatial resolution due to physical limitations in the fabrication process. In this paper, we analyze the ToF sensor's output as a complex value coupling the depth and intensity information in a phasor representation. Based on this analysis, we introduce a novel multi-frame superresolution technique that can improve both spatial resolution in intensity and depth images simultaneously. We believe our proposed method can benefit numerous applications where high resolution depth sensing is desirable, such as precision automated navigation and collision avoidance. Fengqiang Li, Pablo Ruiz 0002, Oliver Cossairt, Aggelos K. Katsaggelos |
ICASSP | 4 |
| 2019 | Pigment Unmixing of Hyperspectral Images of Paintings Using Deep Neural NetworksabstractIn this paper, the problem of automatic nonlinear unmixing of hyperspectral reflectance data using works of art as test cases is described. We use a deep neural network to decompose a given spectrum quantitatively to the abundance values of pure pigments. We show that adding another step to identify the constituent pigments of a given spectrum leads to more accurate unmixing results. Towards this, we use another deep neural network to identify pigments first and integrate this information to different layers of the network used for pigment unmixing. As a test set, the hyperspectral images of a set of mock-up paintings consisting of a broad palette of pigment mixtures, and pure pigment exemplars, were measured. The results of the algorithm on the mock-up test set are reported and analyzed. Neda Rohani, Emeline Pouyet, Marc Walton, Oliver Cossairt, Aggelos K. Katsaggelos |
ICASSP | 5 |
| 2019 | Spatially Adaptive Losses for Video Super-resolution with GANsabstractDeep Learning techniques and more specifically Generative Adversarial Networks (GANs) have recently been used for solving the video super-resolution (VSR) problem. In some of the published works, feature-based perceptual losses have also been used, resulting in promising results. While there has been work in the literature incorporating temporal information into the loss function, studies which make use of the spatial activity to improve GAN models are still lacking. Towards this end, this paper aims to train a GAN guided by a spatially adaptive loss function. Experimental results demonstrate that the learned model achieves improved results with sharper images, fewer artifacts and less noise. Xijun Wang 0003, Alice Lucas, Santiago Lopez Tapia, Xinyi Wu 0001, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP | 6 |
| 2019 | Gan-Based Video Super-Resolution With Direct Regularized Inversion of the Low-Resolution Formation ModelabstractWhile high and ultra high definition displays are becoming popular, most of the available content has been acquired at much lower resolutions. In this work we propose to pseudo-invert with regularization the image formation model using GANs and perceptual losses. Our model, which does not require the use of motion compensation, utilizes explicitly the low resolution image formation model and additionally introduces two feature losses which are used to obtain perceptually improved high resolution images. The experimental validation shows that our approach outperforms current video super resolution learning based models. Santiago Lopez Tapia, Alice Lucas, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2019 | Efficient Fine-Tuning of Neural Networks for Artifact Removal in Deep Learning for Inverse Imaging ProblemsabstractWhile Deep Neural Networks trained for solving inverse imaging problems (such as super-resolution, denoising, or inpainting tasks) regularly achieve new state-of-the-art restoration performance, this increase in performance is often accompanied with undesired artifacts generated in their solution. These artifacts are usually specific to the type of neural network architecture, training, or test input image used for the inverse imaging problem at hand. In this paper, we propose a fast, efficient post-processing method for reducing these artifacts. Given a test input image and its known image formation model, we fine-tune the parameters of the trained network and iteratively update them using a data consistency loss. We show that in addition to being efficient and applicable to large variety of problems, our post-processing through fine-tuning approach enhances the solution originally provided by the neural network by maintaining its restoration quality while reducing the observed artifacts, as measured qualitatively and quantitatively. Alice Lucas, Santiago Lopez Tapia, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2019 | A 3D Cross-Hemisphere Neighborhood Difference Convnet for Chronic Stroke Lesion SegmentationabstractThe following topics are dealt with: learning (artificial intelligence); feature extraction; convolutional neural nets; image classification; image segmentation; image representation; object detection; video signal processing; image colour analysis; neural nets. Yanran Wang 0002, Hengkang Wang, Sophia Chen, Aggelos K. Katsaggelos, Adam Martersteck, James Higgins, Virginia B. Hill, Todd B. Parrish |
ICIP | 4 |
| 2019 | Sinogram Image Completion for Limited Angle Tomography With Generative Adversarial NetworksabstractIn this paper, we present a novel approach based on deep neural network for solving the limited angle tomography problem. The limited angle views in tomography cause severe artifacts in the tomographic reconstruction. We use deep convolutional generative adversarial networks (DCGAN) to fill in the missing information in the sino-gram domain. By using the continuity loss and the two-ends method, the image completion in the sinogram domain is done effectively, resulting in high quality reconstructions with fewer artifacts. The sinogram completion method can be applied to different problems such as ring artifact removal and truncated tomography problems. Seunghwan Yoo, Mark Wolfman, Doga Gürsoy, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2019 | Learning from crowds with variational Gaussian processes
Pablo Ruiz 0002, Pablo Morales-Alvarez, Rafael Molina 0001, Aggelos K. Katsaggelos |
Pattern Recognit. | 4 |
| 2019 | Generative Adversarial Networks and Perceptual Losses for Video Super-ResolutionabstractVideo super-resolution (VSR) has become one of the most critical problems in video processing. In the deep learning literature, recent works have shown the benefits of using adversarial-based and perceptual losses to improve the performance on various image restoration tasks; however, these have yet to be applied for video super-resolution. In this paper, we propose a generative adversarial network (GAN)-based formulation for VSR. We introduce a new generator network optimized for the VSR problem, named VSRResNet, along with new discriminator architecture to properly guide VSRResNet during the GAN training. We further enhance our VSR GAN formulation with two regularizers, a distance loss in feature-space and pixel-space, to obtain our final VSRResFeatGAN model. We show that pre-training our generator with the mean-squared-error loss only quantitatively surpasses the current state-of-the-art VSR models. Finally, we employ the PercepDist metric to compare the state-of-the-art VSR models. We show that this metric more accurately evaluates the perceptual quality of SR solutions obtained from neural networks, compared with the commonly used PSNR/SSIM metrics. Finally, we show that our proposed model, the VSRResFeatGAN model, outperforms the current state-of-the-art SR models, both quantitatively and qualitatively. Alice Lucas, Santiago Lopez Tapia, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2018 | Efficient Video Object Segmentation via Network ModulationabstractVideo object segmentation targets segmenting a specific object throughout a video sequence when given only an annotated first frame. Recent deep learning based approaches find it effective to fine-tune a general-purpose segmentation model on the annotated frame using hundreds of iterations of gradient descent. Despite the high accuracy that these methods achieve, the fine-tuning process is inefficient and fails to meet the requirements of real world applications. We propose a novel approach that uses a single forward pass to adapt the segmentation model to the appearance of a specific object. Specifically, a second meta neural network named modulator is trained to manipulate the intermediate layers of the segmentation network given limited visual and spatial information of the target object. The experiments show that our approach is 70× faster than fine-tuning approaches and achieves similar accuracy. Our model and code have been released at https://github.com/linjieyangsc/video_seg. Yanran Wang 0002, Xuehan Xiong, Jianchao Yang, Aggelos K. Katsaggelos |
CVPR | 5 |
| 2018 | 3D Image Reconstruction from Multi-Focus Microscope: Axial Super-Resolution and Multiple-Frame ProcessingabstractMulti-focus microscope (MFM) provides a way to obtain 3D information by simultaneously capturing multiple focal planes. The naive method for MFM reconstruction is to stack the sub-images with alignment. However, the resolution in the z-axis in this method is limited by the number of acquired focal planes. In this work we build on a recent reconstruction algorithm for MFM, using information from multiple frames to improve the reconstruction quality. We propose two multiple-frame MFM image reconstruction algorithms: batch and recursive approaches. In the batch approach, we take multiple MFM frames and jointly estimate the 3D image and the motion for each frame. In the recursive approach, we utilize the reconstructed image from the previous frame. Experimental results show that the proposed algorithms produce a sequence of 3D object reconstruction with high quality that enable reconstruction of dynamic extended objects. Seunghwan Yoo, Pablo Ruiz 0002, Xiang Huang 0006, Kuan He, Nicola J. Ferrier, Mark Hereld, Alan Selewa, Matthew Daddysman, Norbert Scherer, Oliver Cossairt, Aggelos K. Katsaggelos |
ICASSP | 11 |
| 2018 | ADP: Automatic differentiation ptychographyabstractPtychography is an imaging technique which aims to recover the complex-valued exit wavefront of an object from a set of its diffraction pattern magnitudes. Ptychography is one of the most popular techniques for sub-30 nanometer imaging as it does not suffer from the limitations of typical lens based imaging techniques. The object can be reconstructed from the captured diffraction patterns using iterative phase retrieval algorithms. Over time many algorithms have been proposed for iterative reconstruction of the object based on manually derived update rules. In this paper, we adapt automatic differentiation framework to solve practical and complex ptychographic phase retrieval problems and demonstrate its advantages in terms of speed, accuracy, adaptability and generalizability across different scanning techniques. Sushobhan Ghosh, Youssef S. G. Nashed, Oliver Cossairt, Aggelos K. Katsaggelos |
ICCP | 4 |
| 2018 | Direct: Deep Discriminative Embedding for Clustering of Ligo DataabstractIn this paper, benefiting from the strong ability of deep neural network in estimating non-linear functions, we propose a discriminative embedding function to be used as a feature extractor for clustering tasks. The trained embedding function transfers knowledge from the domain of a labeled set of morphologically-distinct images, known as classes, to a new domain within which new classes can potentially be isolated and identified. Our target application in this paper is the Gravity Spy Project, which is an effort to characterize transient, non-Gaussian noise present in data from the Advanced Laser Interferometer Gravitational-wave Observatory, or LIGO. Accumulating large, labeled sets of noise features and identifying of new classes of noise lead to a better understanding of their origin, which makes their removal from the data and/or detectors possible. Sara Bahaadini, Neda Rohani, Aggelos K. Katsaggelos, Vahid Noroozi, Scott Coughlin, Michael Zevin |
ICIP | 3 |
| 2018 | Blind Color Deconvolution of Histopathological Images Using a Variational Bayesian ApproachabstractWhole-slide histological images are routinely used by medical doctors in diagnosis. Most of these images are stained with the very common and inexpensive hematoxylin and eosin dyes. Slide stain separation and color normalization are crucial steps within the digital pathology workflow which require a previous color deconvolution step. This image processing task is not easy, especially when working with images taken from different microscopes and slides stained in different laboratories. In this paper, based on Variational Bayes inference, an efficient new blind color deconvolution method is proposed. The new model takes into account both spatial relations among image pixels and similarity to a given reference color-vector matrix. A comparison with classical and current state-of-the-art color deconvolution algorithms, using real images with known ground truth hematoxylin and eosin values, has been carried out. This comparison has demonstrated the superiority of the proposed approach. Natalia Hidalgo-Gavira, Javier Mateos, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2018 | An Interior Point Method for Nonnegative Sparse Signal ReconstructionabstractWe present a primal-dual interior point method (IPM) with a novel preconditioner to solve the ℓ1-norm regularized least square problem for nonnegative sparse signal reconstruction. IPM is a second-order method that uses both gradient and Hessian information to compute effective search directions and achieve super-linear convergence rates. It therefore requires many fewer iterations than first-order methods such as iterative shrinkage/thresholding algorithms (ISTA) that only achieve sub-linear convergence rates. However, each iteration of IPM is more expensive than in ISTA because it needs to evaluate an inverse of a Hessian matrix to compute the Newton direction. We propose to approximate each Hessian matrix by a diagonal matrix plus a rank-one matrix. This approximation matrix is easily invertible using the Sherman-Morrison formula, and is used as a novel preconditioner in a preconditioned conjugate gradient method to compute a truncated Newton direction. We demonstrate the efficiency of our algorithm in compressive 3D volumetric image reconstruction. Numerical experiments show favorable results of our method in comparison with previous interior point based and iterative shrinkage/thresholding based algorithms. Xiang Huang 0006, Kuan He, Seunghwan Yoo, Oliver Cossairt, Aggelos K. Katsaggelos, Nicola J. Ferrier, Mark Hereld |
ICIP | 5 |
| 2018 | Generative Adversarial Networks and Perceptual Losses for Video Super-ResolutionabstractRecent research on image super-resolution (SR) has shown that the use of perceptual losses such as feature-space loss functions and adversarial training can greatly improve the perceptual quality of the resulting SR output. In this paper, we extend the use of these perceptual-focused approaches for image SR to that of video SR. We design a 15-block residual neural network, VSRResNet, which is pre-trained on a the traditional mean -squared -error (MSE) loss and later fine-tuned with a feature-space loss function in an adversarial setting. We show that our proposed system, VSRRes-FeatGAN, produces super-resolved frames of much higher perceptual quality than those provided by the MSE-based model. Alice Lucas, Aggelos K. Katsaggelos, Santiago Lopez Tapia, Rafael Molina 0001 |
ICIP | 2 |
| 2018 | Probabilistic Matrix Completion for Image Phase RetrievalabstractIn this paper we address the Phase Retrieval problem, which aims to recover the phase of the Fourier transform of a signal when only magnitude measurements are available. Following recent developments in Phase Retrieval, the problem can be transformed into a convex semidefinite programming optimization problem, which can be solved using Matrix Completion techniques. In this paper the acquisition process is modeled using a likelihood function, which splits the original problem into two convex optimization problems, and alternates between the solution of each of them. To relate both convex problems we introduce a heuristic, which results in fast convergence of the proposed method. Petros Nyfantis, Pablo Ruiz 0002, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2018 | Video Error Concealment Using Deep Neural NetworksabstractThis paper presents an adaptable decoder-like model for video error concealment through optical flow prediction using deep neural networks. The horizontal and vertical motion fields from previous optical flows are separated and passed through two parallel pipelines with convolutional and long short-term memory layers. The combined output from these two networks, the predicted flow, is then used to reconstruct the degraded portion of the future video frame. Unlike current methods that use pixel or voxel information, we propose an architecture that uses three previous optical flows obtained through a flow generation step. The generator portion of the network can be easily interchanged with other methods, increasing the adaptability of the model. The network is trained in supervised mode and the performance is evaluated using standard video quality metrics by comparing the reconstructed frames from our prediction and the generated ground truth. Arun Sankisa, Arjun Punjabi, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2018 | Bayesian Approach for Automatic Joint Parameter Estimation in 3D Image Reconstruction from Multi-Focus MicroscopeabstractWe present a Bayesian approach for 3D image reconstruction of an extended object imaged with multi-focus microscopy (MFM). MFM simultaneously captures multiple sub-images of different focal planes to provide 3D information of the sample. The naive method to reconstruct the object is to stack the sub-images along the z-axis, but the result suffers from poor resolution in the z-axis. The maximum a posteriori framework provides a way to reconstruct a 3D image according to its observation model and prior knowledge. It jointly estimates the 3D image and the model parameters. Experimental results with synthetic and real experimental data show that it enables the high-quality 3D reconstruction of an extended object from MFM. Seunghwan Yoo, Pablo Ruiz 0002, Xiang Huang 0006, Kuan He, Itay Gdor, Alan Selewa, Matthew Daddysman, Nicola J. Ferrier, Mark Hereld, Norbert Scherer, Oliver Cossairt, Aggelos K. Katsaggelos |
ICIP | 13 |
| 2018 | Fully Automated Blind Color Deconvolution of Histopathological Images
Natalia Hidalgo-Gavira, Javier Mateos, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
MICCAI (2) | 5 |
| 2018 | Machine learning for Gravity Spy: Glitch classification and dataset
Sara Bahaadini, Vahid Noroozi, Neda Rohani, Scott Coughlin, Michael Zevin, Joshua R. Smith 0003, Vicky Kalogera, Aggelos K. Katsaggelos |
Inf. Sci. | 8 |
| 2018 | Variational Gaussian process for multisensor classification problems
Neda Rohani, Pablo Ruiz 0002, Rafael Molina 0001, Aggelos K. Katsaggelos |
Pattern Recognit. Lett. | 4 |
| 2017 | Shape-from-Shifting: Uncalibrated Photometric Stereo with a Mobile DeviceabstractSurface shape scanning techniques, such as laser scanning and photometric stereo, are widespread analytical tools used in the field of cultural heritage. Compared to regular 2D RGB photos, 3D surface scans provide higher fidelity of an object's surface shape which assist conservators, art historians, and archaeologists in understanding how these artworks and artifacts are made and to digitally document them for purposes of conservation. However, current state-of-the-art 3D surface scanning tools used in art conservation are often expensive and bulky-such as light dome structures that are often over 1 m in diameter. In this paper, we introduce mobile shape-from-shifting (SfS): a simple, low-cost and streamlined photometric stereo framework for scanning planar surfaces with a consumer mobile device coupled to a low-cost add-on component. Our free-form mobile SfS framework relaxes the rigorous hardware and other complex requirements inherent to conventional 3D scanning tools. This is achieved by taking a sequence of photos with the on-board camera and flash of a mobile device. The sequence of captures are used to reconstruct high quality normal maps using nearlight photometric stereo algorithms, which are of comparable quality to conventional photometric stereo. We demonstrate 3D surface reconstructions with SfS on different materials and scales. Moreover, the mobile SfS technique can be used "in the wild" so that 3D scans may be performed in their natural environment, eliminating the need for transport to a laboratory setting. With the elegant design and low cost, we believe our Mobile SfS can greatly benefit the conservation community by providing a userfriendly and cost-effective solution for 3D surface scanning. Chia-Kai Yeh, Fengqiang Li, Gianluca Pastorelli, Marc Walton, Aggelos K. Katsaggelos, Oliver Cossairt |
eScience | 5 |
| 2017 | Deep multi-view models for glitch classificationabstractNon-cosmic, non-Gaussian disturbances known as “glitches”, show up in gravitational-wave data of the Advanced Laser Interferometer Gravitational-wave Observatory, or aLIGO. In this paper, we propose a deep multi-view convolutional neural network to classify glitches automatically. The primary purpose of classifying glitches is to understand their characteristics and origin, which facilitates their removal from the data or from the detector entirely. We visualize glitches as spectrograms and leverage the state-of-the-art image classification techniques in our model. The suggested classifier is a multi-view deep neural network that exploits four different views for classification. The experimental results demonstrate that the proposed model improves the overall accuracy of the classification compared to traditional single view algorithms. Sara Bahaadini, Neda Rohani, Scott Coughlin, Michael Zevin, Vicky Kalogera, Aggelos K. Katsaggelos |
ICASSP | 6 |
| 2017 | A bayesian multi-frame image super-resolution algorithm using the Gaussian Information FilterabstractMulti-frame image super-resolution (SR) is an image processing technology applicable to any digital, pixilated camera that is limited, by construction, to a certain number of pixels. The objective of SR is to utilize signal processing to overcome the physical limitation and emulate the “capabilities” of a camera with a higher-density pixel array. SR is well known to be an ill-posed problem and, consequently, state-of-the-art solutions approach it statistically, typically making use of Bayesian inference. Unfortunately, direct marginalization of the posterior distribution resulting from the Bayesian modeling is not analytically tractable. An approximation method, such as Variational Bayesian Inference (VBI), is a powerful tool that retains the advantages of statistical modeling. However, its derivation is tedious and model specific. In this paper, we propose an alternative approximate inference methodology, based upon the well-established, Gaussian Information Filter, which offers a much simpler mathematical derivation while retaining the statistical advantages of VBI. Matthew Woods, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2017 | Ptychnet: CNN based fourier ptychographyabstractFourier ptychography is an imaging technique that overcomes the diffraction limit of conventional cameras with applications in microscopy and long range imaging. Diffraction blur causes resolution loss in both cases. In Fourier ptychography, a coherent light source illuminates an object, which is then imaged from multiple viewpoints. The reconstruction of the object from these set of recordings can be obtained by an iterative phase retrieval algorithm. However, the retrieval process is slow and does not work well under certain conditions. In this paper, we propose a new reconstruction algorithm that is based on convolutional neural networks and demonstrate its advantages in terms of speed and performance. Armin Kappeler, Sushobhan Ghosh, Jason Holloway, Oliver Cossairt, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2017 | Passive millimeter wave image classification with large scale Gaussian processesabstractPassive Millimeter Wave Images (PMMWIs) are being increasingly used to identify and localize objects concealed under clothing. Taking into account the quality of these images and the unknown position, shape, and size of the hidden objects, large data sets are required to build successful classification/detection systems. Kernel methods, in particular Gaussian Processes (GPs), are sound, flexible, and popular techniques to address supervised learning problems. Unfortunately, their computational cost is known to be prohibitive for large scale applications. In this work, we present a novel approach to PMMWI classification based on the use of Gaussian Processes for large data sets. The proposed methodology relies on linear approximations to kernel functions through random Fourier features. Model hyperparameters are learned within a variational Bayes inference scheme. Our proposal is well suited for real-time applications, since its computational cost at training and test times is much lower than the original GP formulation. The proposed approach is tested on a unique, large, and real PMMWI database containing a broad variety of sizes, types, and locations of hidden objects. Pablo Morales-Alvarez, Adrián Pérez-Suay, Rafael Molina 0001, Gustau Camps-Valls, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2017 | Spike and slab variational inference for blind image deconvolutionabstractIn this work, we propose a new variational blind deconvolution method for spike and slab prior models. Soft-sparse or shrinkage priors such as the Laplace and other related Gaussian Scale Mixture priors may not be ideal sparsity promoting priors. They assign zero probability mass to events we may be interested in assigning a probability greater than zero. The truly sparse nature of the spike and slab priors allows us to discard irrelevant information in the blur estimation process, resulting in improved performance. We present an efficient inference algorithm to estimate the unknown blur kernel in the filter space, from which we estimate the final deblurred image. The VB approach we propose in this paper handles the inference in a much more efficient way than MCMC, and is more accurate than the standard mean field variational approximation. We prove the efficacy of our method by means of a series of experiments on both synthetically generated and real images. Juan G. Serra, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2017 | Greedy Bayesian double sparsity dictionary learningabstractThis work presents a greedy Bayesian dictionary learning (DL) algorithm where not only the signals but also the dictionary representation matrix accept a sparse representation. This double-sparsity (DS) model has been shown to be superior to the standard sparse approach in some image processing tasks, where sparsity is only imposed on the signal coefficients. We present a new Bayesian approach which addresses typical shortcomings of regularization-based DS algorithms: the prior knowledge of the true noise level and the need of parameter tuning. Our model estimates the noise and sparsity levels as well as the model parameters from the observations and frequently outperforms state-of-the-art dictionary based techniques by taking into account the uncertainty of the estimates. Additionally, we introduce a versatile notation which generalizes denoising, inpainting and compressive sensing problem formulations. Finally, theoretical results are validated with denoising experiments on a set of images. Juan G. Serra, Salvador Villena, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2017 | Sparse Representation-Based Multiple Frame Video Super-ResolutionabstractIn this paper, we propose two multiple-frame super-resolution (SR) algorithms based on dictionary learning (DL) and motion estimation. First, we adopt the use of video bilevel DL, which has been used for single-frame SR. It is extended to multiple frames by using motion estimation with sub-pixel accuracy. We propose a batch and a temporally recursive multi-frame SR algorithm, which improves over single-frame SR. Finally, we propose a novel DL algorithm utilizing consecutive video frames, rather than still images or individual video frames, which further improves the performance of the video SR algorithms. Extensive experimental comparisons with the state-of-the-art SR algorithms verify the effectiveness of our proposed multiple-frame video SR approach. Qiqin Dai, Seunghwan Yoo, Armin Kappeler, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2017 | Robust and Low-Rank Representation for Fast Face Identification With OcclusionsabstractIn this paper, we propose an iterative method to address the face identification problem with block occlusions. Our approach utilizes a robust representation based on two characteristics in order to model contiguous errors (e.g., block occlusion) effectively. The first fits to the errors a distribution described by a tailored loss function. The second describes the error image as having a specific structure (resulting in low-rank in comparison with image size). We will show that this joint characterization is effective for describing errors with spatial continuity. Our approach is computationally efficient due to the utilization of the alternating direction method of multipliers. A special case of our fast iterative algorithm leads to the robust representation method, which is normally used to handle non-contiguous errors (e.g., pixel corruption). Extensive results on representative face databases (in constrained and unconstrained environments) document the effectiveness of our method over existing robust representation methods with respect to both identification rates and computational time. Michael Iliadis, Haohong Wang, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2017 | Bayesian K-SVD Using Fast Variational InferenceabstractRecent work in signal processing in general and image processing in particular deals with sparse representation related problems. Two such problems are of paramount importance: an overriding need for designing a well-suited overcomplete dictionary containing a redundant set of atoms-i.e., basis signals-and how to find a sparse representation of a given signal with respect to the chosen dictionary. Dictionary learning techniques, among which we find the popular K-singular value decomposition algorithm, tackle these problems by adapting a dictionary to a set of training data. A common drawback of such techniques is the need for parameter-tuning. In order to overcome this limitation, we propose a fully-automated Bayesian method that considers the uncertainty of the estimates and produces a sparse representation of the data without prior information on the number of non-zeros in each representation vector. We follow a Bayesian approach that uses a three-tiered hierarchical prior to enforce sparsity on the representations and develop an efficient variational inference framework that reduces computational complexity. Furthermore, we describe a greedy approach that speeds up the whole process. Finally, we present experimental results that show superior performance on two different applications with real images: denoising and inpainting. Juan G. Serra, Matteo Testa, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2016 | Compressive reconstruction for 3D incoherent holographic microscopyabstractIncoherent holography has recently attracted significant research interest due to its flexibility for a wide variety of light sources. In this paper, we use compressive sensing to reconstruct a three-dimensional volumetric object from its two-dimensional Fresnel incoherent correlation hologram. We show how compressed sensing enables reconstruction without out-of-focus artifacts, when compared to conventional back-propagation recovery. Finally, we analyze the reconstruction guarantees of the proposed approach both numerically and theoretically and compare that with coherent holography. Oliver Cossairt, Kuan He, Ruibo Shang, Nathan Matsuda, Xiang Huang 0006, Aggelos K. Katsaggelos, Leonidas Spinoulas, Seunghwan Yoo |
ICIP | 7 |
| 2016 | Multi-model robust error correction for face recognitionabstractIn this work we present a general framework for robust error estimation in face recognition. The proposed formulation allows the simultaneous use of various loss functions for modeling the residual in face images, which usually follows non-standard distributions, depending on the image capturing conditions. Our method extends the current vast literature offering flexibility in the selection of the residual modeling characteristics but, at the same time, considering many existing algorithms as special cases. As such, it proves robust for a range of error inducing factors, such as, varying illumination, occlusion, pixel corruption, disguise or their combinations. Extensive simulations document the superiority of selecting multiple models for representing the noise term in face recognition problems, allowing the algorithm to achieve near-optimal performance in most of the tested face databases. Finally, the multi-model residual representation offers useful insights into understanding how different noise types affect face recognition rates. Michael Iliadis, Leonidas Spinoulas, Albert S. Berahas, Haohong Wang, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2016 | Block based video alignment with linear time and space complexityabstractVideo retrieval and video copy detection are well studied problems. The goal is to find the matching video in a database from a given query video. Typically, these query videos are short and aligning the query video is of secondary importance. Short sequences can be aligned using dynamic time warping. But, since time and memory usage increases quadratically with the length of the sequences, such process is not suitable for the alignment of two full length movies. A typical feature film is between 70 and 210 minutes long. Our goal is to find an accurate frame-by-frame alignment of a full length original film and a copy that has inserted and deleted sequences (e.g., commercial breaks or censorship), as well as differences in quality, format and framerate. We propose a fast, robust and memory efficient video sequence alignment algorithm which has linear space and time complexity. Armin Kappeler, Michael Iliadis, Haohong Wang, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2016 | Super-resolution of compressed videos using convolutional neural networksabstractConvolutional neural networks (CNN) have been successfully applied to image super-resolution (SR) as well as other image restoration tasks. In this paper, we consider the problem of compressed video super-resolution. Traditional SR algorithms for compressed videos rely on information from the encoder such as frame type or quantizer step, whereas our algorithm only requires the compressed low resolution frames to reconstruct the high resolution video. We propose a CNN that is trained on both the spatial and the temporal dimensions of compressed videos to enhance their spatial resolution. Consecutive frames are motion compensated and used as input to a CNN that provides super-resolved video frames as output. Our network is pretrained with images, which significantly improves the performance over random initialization. In extensive experimental evaluations, we trained the state-of-the-art image and video superresolution algorithms on compressed videos and compared their performance to our proposed method. Armin Kappeler, Seunghwan Yoo, Qiqin Dai, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2016 | Multiframe blind deconvolution of passive millimeter wave images using variational dirichlet blur kernel estimationabstractPassive Millimeter Wave Images currently used to detect hidden threats suffer from low resolution, blur, and a very low signal-to-noise-ratio. These shortcomings render threat detection, both visual and automatic, very challenging. Furthermore, due to the presence of very severe noise, most of the blind image restoration methods fail to recover the system blurring kernel from a single image. In this paper we propose a robust Bayesian multiframe blind image deconvolution method that approximates the posterior distribution of the blur by a Dirichlet distribution. We show that this approach naturally incorporates the non-negativity and normalization constraints for the blur and cope well with the image noise. The performance of the proposed method is tested on both synthetic and real images. Javier Mateos, Antonio López, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2016 | A novel cumulative distortion metric and a no-reference sparse prediction model for packet prioritization in encoded video transmissionabstractIn this paper we propose a new quality metric to estimate the impact of packet loss on the perceptual quality of encoded video sequences transmitted over error-prone networks. The proposed metric, henceforth referred to as Cumulative Distortion using Structural Similarity (CDSSIM), quantifies the overall structural distortion resulting from bidirectional error propagation in predictively coded, motion compensated videos. Furthermore, we present a No-Reference (NR) sparse regression model to predict the proposed CDSSIM metric using pre-defined features associated with slice loss. The Least Absolute Shrinkage and Selection Operator (LASSO) method is applied for two resolution formats with features extracted solely from the encoded bit-stream. Standardized statistical performance measures show that the model can predict the cumulative distortion to a high degree of accuracy. We further evaluate the results using a Quartile-Based Prioritization (QBP) scheme and demonstrate that the predicted data provides an effective way to prioritize packets for video streaming applications. Arun Sankisa, Katerina Pandremmenou, Lisimachos P. Kondi, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2016 | Bayesian logistic regression with sparse general representation prior for multispectral image classificationabstractIn this work we address the multispectral image classification problem from a Bayesian perspective. We develop an algorithm which utilizes the logistic regression function as the observation model in a probabilistic framework, Super-Gaussian (SG) priors which promote sparsity on the adaptive coefficients, and Variational inference to obtain estimates of all the model unknowns. The proposed algorithm is validated on both synthetic and real experiments and compared with other state-of-the-art methods, such as Support Vector Machine and Gaussian Processes, demonstrating its improved performance. Juan G. Serra, Pablo Ruiz 0002, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2016 | A deep symmetry convnet for stroke lesion segmentationabstractStroke is one of the leading causes of death and disability. Clinically, to establish stroke patient prognosis, an accurate delineation of brain lesion is essential, which is time consuming and prone to subjective errors. In this paper, we propose a novel method call Deep Lesion Symmetry ConvNet to automatically segment chronic stroke lesions using MRI. An 8-layer 3D convolutional neural network is constructed to handle the MRI voxels. An additional CNN stream using the corresponding symmetric MRI voxels is combined, leading to a significant improvement in system performance. The high average dice coefficient achieved on our dataset based on data collected from three research labs demonstrates the effectiveness of our method. Yanran Wang 0002, Aggelos K. Katsaggelos, Todd B. Parrish |
ICIP | 2 |
| 2016 | Joint Data Filtering and Labeling Using Gaussian Processes and Alternating Direction Method of MultipliersabstractSequence labeling aims at assigning a label to every sample of a signal (or pixel of an image) while considering the sequentiality (or vicinity) of the samples. To perform this task, many works in the literature first filter and then label the data. Unfortunately, the filtering, which is performed independently from the labeling, is far from optimal and frequently makes the latter task harder. In this paper, a novel approach that trains a Gaussian process classifier and estimates the coefficients of an optimal filter jointly is presented. The new approach, based on Bayesian modeling and alternating direction method of multipliers (ADMMs) optimization, performs both tasks simultaneously. All unknowns are treated as stochastic variables, which are estimated using variational inference and filtering and labeling are linked with the use of ADMM. In the experimental section, synthetic and real experiments are presented to compare the proposed method with other existing approaches. Pablo Ruiz 0002, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2015 | Dictionary-based multiple frame video super-resolutionabstractIn this paper, we propose a multiple-frame super-resolution (SR) algorithm based on dictionary learning and motion estimation. We adopt the use of multiple bilevel dictionaries which have also been used for single-frame SR. Multiple frames compensated through sub-pixel motion are considered. By simultaneously solving for a batch of patches from multiple frames, the proposed multiple-frame SR algorithm improves over single frame SR. We also propose a novel dictionary learning algorithm based on which dictionaries are trained from consecutive video frames, rather than still images or individual video frames, which further improves the performance of the developed video SR algorithm. Extensive experimental comparisons with state-of-the-art SR algorithms verifies the effectiveness of our proposed multiple-frame SR approach. Qiqin Dai, Seunghwan Yoo, Armin Kappeler, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2015 | Image super-resolution from compressed sensing observationsabstractIn this work we propose a novel framework to obtain High Resolution (HR) images from Compressed Sensing (CS) imaging systems capturing multiple Low Resolution (LR) images of the same scene. The proposed CS Super Resolution (SR) approach combines existing CS reconstruction algorithms with an LR to HR approach based on the use of a Super Gaussian (SG) regularization term. The reconstruction is formulated as a constrained optimization problem which is solved using the Alternate Direction Methods of Multipliers (ADMM). The image estimation subproblem is solved using Majorization-Minimization (MM) while the CS reconstruction becomes an l1-minimization subject to a quadratic constraint. The performed experiments show that the proposed method compares favorably to classical SR methods at compression ratio 1, obtaining excellent SR reconstructions at ratios below one. Wael Saafin, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2015 | Sampling optimization for on-chip compressive videoabstractIn this paper, we consider the problem of on-chip temporal compressive sensing for video reconstruction at high frame-rates without the need of any additional optical components. We devise an optimization scheme in order to achieve adequate spatio-temporal sampling of subsequent frames under maximal capturing speed, based on the bandwidth constraints of a sensor. We test this optimization strategy on a commercially available camera and propose a set of reconstruction steps that can achieve reasonable performance but, at the same time, accommodate high-resolution video reconstruction under realistic time requirements. Our analysis constitutes a set of first steps bringing high-speed compressive video capture within the realm of commercial availability. Leonidas Spinoulas, Oliver Cossairt, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2015 | Distortion Estimation Using Structural Similarity for Video Transmission over Wireless NetworksabstractEfficient streaming of video over wireless networks requires real-time assessment of distortion due to packet loss, especially because predictive coding at the encoder can cause inter-frame propagation of errors and impact the overall quality of the transmitted video. This paper presents an algorithm to evaluate the expected receiver distortion on the source side by utilizing encoder information, transmission channel characteristics and error concealment. Specifically, distinct video transmission units, Group of Blocks (GOBs), are iteratively built at the source by taking into account macroblock coding modes and motion-compensated error concealment for three different combinations of packet loss. Distortion of these units is then calculated using the structural similarity (SSIM) metric and they are stochastically combined to derive the overall expected distortion. The proposed model provides a more accurate estimate of the distortion that closely models quality as perceived through the human visual system. When incorporated into a content-aware utility function, preliminary experimental results show improved packet ordering & scheduling efficiency and overall video signal at the receiver. Arun Sankisa, Aggelos K. Katsaggelos, Peshala V. Pahalawatta |
ISM | 2 |
| 2015 | Cardiac magnetic resonance image-based classification of the risk of arrhythmias in post-myocardial infarction patients
Lasya Priya Kotu, Kjersti Engan, Reza Borhani, Aggelos K. Katsaggelos, Stein Ørn, Leik Woie, Trygve Eftestøl |
Artif. Intell. Medicine | 4 |
| 2015 | Audiovisual Fusion: Challenges and New ApproachesabstractIn this paper, we review recent results on audiovisual (AV) fusion. We also discuss some of the challenges and report on approaches to address them. One important issue in AV fusion is how the modalities interact and influence each other. This review will address this question in the context of AV speech processing, and especially speech recognition, where one of the issues is that the modalities both interact but also sometimes appear to desynchronize from each other. An additional issue that sometimes arises is that one of the modalities may be missing at test time, although it is available at training time; for example, it may be possible to collect AV training data while only having access to audio at test time. We will review approaches to address this issue from the area of multiview learning, where the goal is to learn a model or representation for each of the modalities separately while taking advantage of the rich multimodal training data. In addition to multiview learning, we also discuss the recent application of deep learning (DL) toward AV fusion. We finally draw conclusions and offer our assessment of the future in the area of AV fusion. Aggelos K. Katsaggelos, Sara Bahaadini, Rafael Molina 0001 |
Proc. IEEE | 1 |
| 2015 | Preconditioning for Underdetermined Linear Systems with Sparse SolutionsabstractPerformance guarantees for the algorithms deployed to solve underdetermined linear systems with sparse solutions are based on the assumption that the involved system matrix has the form of an incoherent unit norm tight frame. Learned dictionaries, which are popular in sparse representations, often do not meet the necessary conditions for signal recovery. In compressed sensing (CS), recovery rates have been improved substantially with optimized projections; however, these techniques do not produce binary matrices, which are more suitable for hardware implementation. In this paper, we consider an underdetermined linear system with sparse solutions and propose a preconditioning technique that yields a system matrix having the properties of an incoherent unit norm tight frame. While existing work in preconditioning concerns greedy algorithms, the proposed technique is based on recent theoretical results for standard numerical solvers such as BP and OMP. Our simulations show that the proposed preconditioning improves the recovery rates both in sparse representations and CS; the results for CS are comparable to optimized projections. Evaggelia Tsiligianni, Lisimachos P. Kondi, Aggelos K. Katsaggelos |
IEEE Signal Process. Lett. | 3 |
| 2015 | Variational Dirichlet Blur Kernel EstimationabstractBlind image deconvolution involves two key objectives: 1) latent image and 2) blur estimation. For latent image estimation, we propose a fast deconvolution algorithm, which uses an image prior of nondimensional Gaussianity measure to enforce sparsity and an undetermined boundary condition methodology to reduce boundary artifacts. For blur estimation, a linear inverse problem with normalization and nonnegative constraints must be solved. However, the normalization constraint is ignored in many blind image deblurring methods, mainly because it makes the problem less tractable. In this paper, we show that the normalization constraint can be very naturally incorporated into the estimation process by using a Dirichlet distribution to approximate the posterior distribution of the blur. Making use of variational Dirichlet approximation, we provide a blur posterior approximation that considers the uncertainty of the estimate and removes noise in the estimated kernel. Experiments with synthetic and real data demonstrate that the proposed method is very competitive to the state-of-the-art blind image restoration methods. Xu Zhou 0005, Javier Mateos, Fugen Zhou, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 5 |
| 2014 | Learning filters in Gaussian process classification problemsabstractMany real classification tasks are oriented to sequence (neighbor) labeling, that is, assigning a label to every sample of a signal while taking into account the sequentiality (or neighborhood) of the samples. This is normally approached by first filtering the data and then performing classification. In consequence, both processes are optimized separately, with no guarantee of global optimality. In this work we utilize Bayesian modeling and inference to jointly learn a classifier and estimate an optimal filterbank. Variational Bayesian inference is used to approximate the posterior distributions of all unknowns, resulting in an iterative procedure to estimate the classifier parameters and the filterbank coefficients. In the experimental section we show, using synthetic and real data, that the proposed method compares favorably with other classification/filtering approaches, without the need of parameter tuning. Pablo Ruiz 0002, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2014 | Fast iteratively reweighted least squares for lp regularized image deconvolution and reconstructionabstractIteratively reweighted least squares (IRLS) is one of the most effective methods to minimize the lpregularized linear inverse problem. Unfortunately, the regularizer is nonsmooth and nonconvex when 0 <; p <; 1. In spite of its properties and mainly due to its high computation cost, IRLS is not widely used in image deconvolution and reconstruction. In this paper, we first derive the IRLS method from the perspective of majorization minimization and then propose an Alternating Direction Method of Multipliers (ADMM) to solve the reweighted linear equations. Interestingly, the resulting algorithm has a shrinkage operator that pushes each component to zero in a multiplicative fashion. Experimental results on both image deconvolution and reconstruction demonstrate that the proposed method outperforms state-of-the-art algorithms in terms of speed and recovery quality. Xu Zhou 0005, Rafael Molina 0001, Fugen Zhou, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2014 | Automatic, fast, online calibration between depth and color cameras
Ilya Mikhelson, Philip Greggory Lee, Alan V. Sahakian, Ying Wu 0001, Aggelos K. Katsaggelos |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | Combining Poisson singular integral and total variation prior models in image restoration
Pablo Ruiz 0002, Hiram Madero Orozco, Javier Mateos, Osslan Osiris Vergara-Villegas, Rafael Molina 0001, Aggelos K. Katsaggelos |
Signal Process. | 6 |
| 2014 | Automated Recovery of Compressedly Observed Sparse Signals From Smooth BackgroundabstractWe propose a Bayesian based algorithm to recover sparse signals from compressed noisy measurements in the presence of a smooth background component. This problem is closely related to robust principal component analysis and compressive sensing, and is found in a number of practical areas. The proposed algorithm adopts a hierarchical Bayesian framework for modeling, and employs approximate inference to estimate the unknowns. Numerical examples demonstrate the effectiveness of the proposed algorithm and its advantage over the current state-of-the-art solutions. Zhaofu Chen, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Signal Process. Lett. | 3 |
| 2014 | Bayesian Active Remote Sensing Image ClassificationabstractIn recent years, kernel methods, in particular support vector machines (SVMs), have been successfully introduced to remote sensing image classification. Their properties make them appropriate for dealing with a high number of image features and a low number of available labeled spectra. The introduction of alternative approaches based on (parametric) Bayesian inference has been quite scarce in the more recent years. Assuming a particular prior data distribution may lead to poor results in remote sensing problems because of the specificities and complexity of the data. In this context, the emerging field of nonparametric Bayesian methods constitutes a proper theoretical framework to tackle the remote sensing image classification problem. This paper exploits the Bayesian modeling and inference paradigm to tackle the problem of kernel-based remote sensing image classification. This Bayesian methodology is appropriate for both finite- and infinite-dimensional feature spaces. The particular problem of active learning is addressed by proposing an incremental/active learning approach based on three different approaches: 1) the maximum differential of entropies; 2) the minimum distance to decision boundary; and 3) the minimum normalized distance. Parameters are estimated by using the evidence Bayesian approach, the kernel trick, and the marginal distribution of the observations instead of the posterior distribution of the adaptive parameters. This approach allows us to deal with infinite-dimensional feature spaces. The proposed approach is tested on the challenging problem of urban monitoring from multispectral and synthetic aperture radar data and in multiclass land cover classification of hyperspectral images, in both purely supervised and active learning settings. Similar results are obtained when compared to SVMs in the supervised mode, with the advantage of providing posterior estimates for classification and automatic parameter learning. Comparison with random sampling as well as standard active learning methods such as margin sampling and entropy-query-by-bagging reveals a systematic overall accuracy gain and faster convergence with the number of queries. Pablo Ruiz 0002, Javier Mateos, Gustau Camps-Valls, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2014 | Construction of Incoherent Unit Norm Tight Frames With Application to Compressed SensingabstractDespite the important properties of unit norm tight frames (UNTFs) and equiangular tight frames (ETFs), their construction has been proven extremely difficult. The few known techniques produce only a small number of such frames while imposing certain restrictions on frame dimensions. Motivated by the application of incoherent tight frames in compressed sensing (CS), we propose a methodology to construct incoherent UNTFs. When frame redundancy is not very high, the achieved maximal column correlation becomes close to the lowest possible bound. The proposed methodology may construct frames of any dimensions. The obtained frames are employed in CS to produce optimized projection matrices. Experimental results show that the proposed optimization technique improves CS signal recovery, increasing the reconstruction accuracy. Considering that the UNTFs and ETFs are important in sparse representations, channel coding, and communications, we expect that the proposed construction will be useful in other applications, besides the CS. Evaggelia Tsiligianni, Lisimachos P. Kondi, Aggelos K. Katsaggelos |
IEEE Trans. Inf. Theory | 3 |
| 2014 | Toward Dynamic Scene Understanding by Hierarchical Motion Pattern MiningabstractOur work addresses the problem of analyzing and understanding dynamic video scenes. A two-level motion pattern mining approach is proposed. At the first level, activities are modeled as distributions over patch-based features, including spatial location, moving direction, and speed. At the second level, traffic states are modeled as distributions over activities. Both patterns are shared among video clips. Compared to other works, one advantage of our method is that moving speed is considered to describe visual word. The other advantage is that traffic states are detected and assigned to every video frame. These enable finer semantic interpretation, more precise video segmentation, and anomaly detection. Specifically, every video frame is labeled by a certain traffic state, and the video is segmented frame by frame accordingly. Moving pixels in each frame, which do not belong to any activity or cannot exist in the corresponding traffic state, are detected as anomalies. We have successfully tested our approach on some challenging traffic surveillance sequences containing both pedestrian and vehicle motions. Zhongke Shi, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2014 | Variational Bayesian Methods For Multimedia ProblemsabstractIn this paper we present an introduction to Variational Bayesian (VB) methods in the context of probabilistic graphical models, and discuss their application in multimedia related problems. VB is a family of deterministic probability distribution approximation procedures that offer distinct advantages over alternative approaches based on stochastic sampling and those providing only point estimates. VB inference is flexible to be applied in different practical problems, yet is broad enough to subsume as its special cases several alternative inference approaches including Maximum A Posteriori (MAP) and the Expectation-Maximization (EM) algorithm. In this paper we also show the connections between VB and other posterior approximation methods such as the marginalization-based Loopy Belief Propagation (LBP) and the Expectation Propagation (EP) algorithms. Specifically, both VB and EP are variational methods that minimize functionals based on the Kullback-Leibler (KL) divergence. LBP, traditionally developed using graphical models, can also be viewed as a VB inference procedure. We present several multimedia related applications illustrating the use and effectiveness of the VB algorithms discussed herein. We hope that by reading this tutorial the readers will obtain a general understanding of Bayesian methods and establish connections among popular algorithms used in practice. Zhaofu Chen, S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 4 |
| 2013 | Binocular video object tracking with fast disparity estimationabstractThis paper presents a binocular PTU (pan-tilt unit) camera video object tracking scheme using the MeanShift algorithm and the runtime disparity estimation. The proposed method is to accommodate the requirement of 3D content generation and accurate tracking in more advanced video surveillance applications. The disparity estimation process for each stereoscopic pair is formulated as an energy minimization problem. The iterative solution procedure is implemented in a course-to-fine manner. The estimated disparity is used to scale the tracking window by the MeanShift algorithm, i.e. the size of the tracking area is adjustable according to its inner disparity, and thus the moving object can be better located by the camera. The program maintains the semi-real-time performance and acceptable accuracy as evaluated on a set of standard test data. In our experiment, two PointGrey cameras are controlled through a PTU device. The disparity estimation process on the recorded tracking video (640×480) achieves 6fps on an ordinary PC (2.66GHz CPU, 4GB RAM). Song Ci, Yanwei Liu 0001, Haohong Wang, Aggelos K. Katsaggelos |
AVSS | 5 |
| 2013 | Compressive sensing-based image denoising using adaptive multiple sampling and optimal error toleranceabstractIn this paper, we present a compressive sensing-based image denoising algorithm using spatially adaptive image representation and estimation of optimal error tolerance based on sparse signal analysis. The proposed method performs block-based multiple compressive sampling after decomposing the sparse signal into feature and non-feature regions using simple statistical analysis. For minimization of recovery error and number of iterations, the modified OMP method estimates the optimal error tolerance using the average variance in the recovery step. Experimental results demonstrate that the proposed denoising algorithm better removes noise without undesired artifacts than existing state-of-the-art methods in terms of both objective (PSNR/SSIM) and subjective measures. Processing time of the proposed method is 5 to 10 times faster than the standard OMP-based method. Wonseok Kang, Eunsung Lee, Eunjung Chea, Aggelos K. Katsaggelos, Joonki Paik |
ICASSP | 4 |
| 2013 | Forest hashing: Expediting large scale image retrievalabstractThis paper introduces a hybrid method for searching large image datasets for approximate nearest neighbor items, specifically SIFT descriptors. The basic idea behind our method is to create a serial system that first partitions approximate nearest neighbors using multiple kd-trees before calling upon locally designed spectral hashing tables for retrieval. This combination gives us the local approximate nearest neighbor accuracy of kd-trees with the computational efficiency of hashing techniques. Experimental results show that our approach efficiently and accurately outperforms previous methods designed to achieve similar goals. Jonathan Springer, Xin Xin 0009, Zhu Li 0001, Jeremy Watt, Aggelos K. Katsaggelos |
ICASSP | 5 |
| 2013 | Frequency-domain analysis of discrete wavelet transform coefficients and their adaptive shrinkage for anti-aliasingabstractWe present an antialiasing method using combined wavelet-Fourier transform and spatially adaptive shrinkage of the transform coefficients. Traditional antialiasing methods employ a simple low-pass filter onto the entire image, so the resulting image loses not only aliasing artifacts but also high-frequency components such as edges and ridges. The proposed algorithm analyzes the property of the LL subband of the discrete wavelet transform (DWT), and reduces aliasing artifacts using patch-adaptive shrinkage of the DWT coefficients. More specifically, an antialiased LL subband is obtained using adaptive patch-based aliasing reduction. To detect an aliased region, we subtract the discrete Fourier transform (DFT) coefficients of the LL subband from the DFT coefficients of antialiased LL subband. The detected aliasing artifacts in the LH, HL, and HH subbands are reduced by patch-wise adaptive shrinkage of the transform coefficients. The resulting antialiased image is obtained using the inverse DWT. The aliasing artifacts can be efficiently reduced by adaptively shrinking wavelet transform coefficients for preserving high-frequency image details. The proposed antialiasing algorithm is suitable for removing aliasing artifacts which frequently occur in imaging sensors with limited resolution. Eunjung Chae, Eunsung Lee, Wonseok Kang, Younghoon Lim, Junghoon Jung, Tae-Chan Kim 0002, Aggelos K. Katsaggelos, Joonki Paik |
ICIP | 7 |
| 2013 | Video compressive sensing using multiple measurement vectorsabstractCompressive Sensing (CS) suggests that, under certain conditions, a signal can be reconstructed using a small number of incoherent measurements. We propose a novel video CS framework based on Multiple Measurement Vectors (MMV) which is suitable for signals with temporal correlation such as video sequences. In addition, a CS circulant matrix is employed for fast reconstruction. Furthermore, the proposed framework allows the number of CS measurements associated with each frame to be chosen in the decoder rather than the encoder offering robustness compared to the multi-scale approaches. Experimental results on two video sequences exhibiting fast motion and occlusions, show the advantages of the proposed method over the current state-of-the-art in video CS. Michael Iliadis, Jeremy Watt, Leonidas Spinoulas, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2013 | Real-time super-resolution for digital zooming using finite kernel-based edge orientation estimation and truncated image restorationabstractThis paper presents a novel real-time super-resolution (SR) method using directionally adaptive image interpolation and image restoration. The proposed interpolation method estimates the edge orientation using steerable filters and performs edge refinement along the estimated edge orientation. Bi-linear and bi-cubic interpolation filters are then selectively used according to the estimated edge orientation for reducing jagging artifacts in slanting edge regions. The proposed restoration method can effectively remove image degradation caused by interpolation using the directionally adaptive truncated constrained least-squares (TCLS) filter. The proposed method provides high-quality magnified images which are similar to or better than the result of advanced interpolation or SR methods without high computational load. Experimental results indicate that the proposed system gives higher peak-to-peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) values than the state-of-the-art image interpolation methods. Wonseok Kang, Jaehwan Jeon, Eunsung Lee, Changhun Cho, Junghoon Jung, Tae-Chan Kim 0002, Aggelos K. Katsaggelos, Joonki Paik |
ICIP | 7 |
| 2013 | Drosophila eye nuclei segmentation based on graph cut and convex shape priorabstractThe rapid advance in three-dimensional (3D) confocal imaging technologies is rapidly increasing the availability of 3D cellular images. However, the lack of robust automated methods for the extraction of cell or organelle shapes from the images is hindering researchers ability to take full advantage of the increase in experimental output. The lack of appropriate methods is particularly significant when the density of the features of interest in high, such as in the developing eye of the fruit fly. Here, we present a novel and efficient nuclei segmentation algorithm based on the combination of graph cut and convex shape prior. The main characteristic of the algorithm is that it segments nuclei foreground using a graph cut algorithm and splits overlapping or touching cell nuclei by simple convex and concavity analysis, using a convex shape assumption for nuclei contour. We evaluate the performance of our method by applying it to a library of publicly-available two-dimensional (2D) images that were hand-labeled by experts. Our algorithm yields a substantial quantitative improvement over other methods for this benchmark. For example, our method achieves a decrease of 3.2 in the Hausdorff distance and an decrease of 1.8 per slice in the merged nuclei error. Nicolás Peláez, L. Rebay, Richard W. Carthew, Aggelos K. Katsaggelos, Luis A. Nunes Amaral |
ICIP | 6 |
| 2013 | Robust feature selection with self-matching scoreabstractWith the increasing power of mobile headsets and mobile networks, mobile visual search applications have gained popularity and became tractable. One of the key technologies to enable visual search are the robust and compact features, which are extracted from an image and are invariant to recapturing variations. One of the key factors for compact visual descriptors is the selection of local features. The size of the compact visual descriptors and the computational complexities of a visual search system increase with the number of features selected. In this sense, ranking the descriptors extracted from a single image according to their importance in terms of recapturing is very necessary and important. In this paper, we attack this problem by proposing a novel self-matching selection. In this method, we randomly apply an out-of-plane rotation to the target image and match the original features to the features that are extracted from the out-of-plane rotated image. The importance of the features is ranked according to the self-matching score. This method is proven to be better than other peak strength and edge strength based methods by 30% from experiments on a large database. Xin Xin 0009, Zhu Li 0001, Zhan Ma 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2013 | A multi-camera motion capture system for remote healthcare monitoringabstractThis paper presents a multi-camera motion capture system aiming to provide caregivers with timely access to the patient's health status through mobile communication devices. The major components include video capture, object detection, video coding and transmission, error concealment, and video analysis. Our contribution is twofold. First, several novel ideas are developed, including fast object detection, and content-aware and adaptive video coding and transmission. Second, all components are seamlessly integrated in a unified optimization framework dedicated for online data transmission. In the scenario, the subject walked on a treadmill with four tripod cameras capturing the video from different viewpoints. After video compression and transmission over a wireless sensor network, the remote receiver recovered the videos and performed multi-view motion capture for gait analysis. Experimental results show that the presented system design achieves better video quality than traditional video coding and transmission scheme, while the requirement for a low-cost, noninvasive and real-time healthcare monitoring system is accommodated. Song Ci, Aggelos K. Katsaggelos, Yanwei Liu 0001 |
ICME | 3 |
| 2013 | Multimedia multicast service provisioning in cognitive radio networksabstractIn this paper, we propose a design framework for achieving efficient multimedia multicast services in cognitive radio (CR) networks. The framework incorporates the characteristics of both heterogeneous network environment and the scalable video content. By adopting cooperative transmissions for the delivery of enhancement layer data, we can not only improve the achieved video quality but also protect the rights of subscribed secondary users. We also utilize network coding and superposition coding to achieve efficient multicast transmissions of the layered video packets in multi-channel CR networks. Numerical examples show the proposed framework can improve the average received data rate by up to 15%. When achieving the same video quality, the proposed framework can save 30% transmission time comparing with the scenario using direct transmission alone. Fen Hou, Zhaofu Chen, Jianwei Huang 0001, Zhu Li 0001, Aggelos K. Katsaggelos |
IWCMC | 5 |
| 2013 | Laplacian embedding and key points topology verification for large scale mobile visual identification
Xin Xin 0009, Zhu Li 0001, Aggelos K. Katsaggelos |
Signal Process. Image Commun. | 3 |
| 2013 | Compressive Blind Image DeconvolutionabstractWe propose a novel blind image deconvolution (BID) regularization framework for compressive sensing (CS) based imaging systems capturing blurred images. The proposed framework relies on a constrained optimization technique, which is solved by a sequence of unconstrained sub-problems, and allows the incorporation of existing CS reconstruction algorithms in compressive BID problems. As an example, a non-convex lp quasi-norm with is employed as a regularization term for the image, while a simultaneous auto-regressive regularization term is selected for the blur. Nevertheless, the proposed approach is very general and it can be easily adapted to other state-of-the-art BID schemes that utilize different, application specific, image/blur regularization terms. Experimental results, obtained with simulations using blurred synthetic images and real passive millimeter-wave images, show the feasibility of the proposed method and its advantages over existing approaches. Bruno Amizic, Leonidas Spinoulas, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2013 | Luma-Chroma Space Filter Design for Subpixel-Based Monochrome Image DownsamplingabstractIn general, subpixel-based downsampling can achieve higher apparent resolution of the down-sampled images on LCD or OLED displays than pixel-based downsampling. With the frequency domain analysis of subpixel-based downsampling, we discover special characteristics of the luma-chroma color transform choice for monochrome images. With these, we model the anti-aliasing filter design for subpixel-based monochrome image downsampling as a human visual system-based optimization problem with a two-term cost function and obtain a closed-form solution. One cost term measures the luminance distortion and the other term measures the chrominance aliasing in our chosen luma-chroma space. Simulation results suggest that the proposed method can achieve sharper down-sampled gray/font images compared with conventional pixel and subpixel-based methods, without noticeable color fringing artifacts. Lu Fang 0001, Oscar C. Au, Ngai-Man Cheung, Aggelos K. Katsaggelos, Houqiang Li, Feng Zou 0006 |
IEEE Trans. Image Process. | 4 |
| 2013 | Application-Aware Approach to Compression and Transmission of H.264 Encoded Video for Automated and Centralized Transportation SurveillanceabstractIn this paper, we present a transportation video coding and wireless transmission system specifically tailored to automated vehicle tracking applications. By taking into account the video characteristics and the lossy nature of the wireless channels, we propose video preprocessing and error control approaches to enhance tracking performance while conserving bandwidth resources and computational power at the transmitter. Compared with current state-of-the-art H.264-based implementations, our system is shown to yield over 80% bitrate savings for comparable tracking accuracy. Zhaofu Chen, Sotirios A. Tsaftaris, Eren Soyak, Aggelos K. Katsaggelos |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2012 | Bayesian Blind Deconvolution with General Sparse Image Priors
S. Derin Babacan, Rafael Molina 0001, Minh N. Do, Aggelos K. Katsaggelos |
ECCV (6) | 4 |
| 2012 | Single camera-based full depth map estimation using color shifting property of a multiple color-filter apertureabstractA multiple color-filter aperture (MCA) camera can provide depth information as well as color and intensity in the single-camera framework, where the MCA generates misalignment between color channels depending on the distance of a region-of-interest. In this paper, we present a single camera-based estimation of the full depth map using the color shifting property of the MCA. For estimating the color shifting vectors (CSVs) among red, green, and blue color channels, edges are extracted at each color channel. At the edge, we estimate CSVs using normalized cross correlation combined with color shifting mask map. A full depth map is then generated by depth interpolation using the matting Laplacian method from sparsely estimated CSVs at an edge location. Experimental results show that the proposed method can not only estimate the full depth map but also correct the misaligned color image to generate photorealistic color images using a single camera equipped with MCA. Monson H. Hayes III, Aggelos K. Katsaggelos, Joonki Paik |
ICASSP | 4 |
| 2012 | Compressive sampling with unknown blurring function: Application to passive millimeter-wave imagingabstractWe propose a novel blind image deconvolution (BID) regularization framework for compressive passive millimeter-wave (PMMW) imaging systems. The proposed framework is based on the variable-splitting optimization technique, which allows us to utilize existing compressive sensing reconstruction algorithms in compressive BID problems. In addition, a non-convex lpquasi-norm with 0 <; p <; 1 is employed as a regularization term for the image, while a simultaneous auto-regressive (SAR) regularization term is utilized for the blur. Furthermore, the proposed framework is very general and it can be easily adapted to other state-of-the-art BID approaches that utilize different image/blur regularization terms. Experimental results, obtained with simulations using a synthetic image and real PMMW images, show the advantage of the proposed approach compared to existing ones. Bruno Amizic, Leonidas Spinoulas, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2012 | A wireless video surveillance system with an active cameraabstractThis paper introduced a camera surveillance system in wireless communications. The system contains three major modules, PTU (pan-tilt unit) camera control for surveillance video capture, cross-layer control for data compression and transmission, and error concealment for video quality enhancement. Our contribution is twofold. First, a system design for data collection and transmission over wireless networks is presented and is evaluated with physical surveillance equipments. The camera is capable of following the moving target according to the control information. The end-to-end distortion estimation in the delay constrained video coding process takes into account the dynamic channel condition and physical layer modulation and coding scheme (MCS) to determine optimal coding and transmission parameters. Second, multiple error concealment strategies, including interleaving, boundary match and video up-sampling, are applied utilizing the special property of the PTU camera motion. Song Ci, Yanwei Liu 0001, Dalei Wu, Haohong Wang, Aggelos K. Katsaggelos |
VCIP | 6 |
| 2012 | A game theoretic approach to video streaming over peer-to-peer networks
Ehsan Maani, Zhaofu Chen, Aggelos K. Katsaggelos |
Signal Process. Image Commun. | 3 |
| 2012 | Special issue on advances in 2D/3D Video Streaming Over P2P Networks
Naeem Ramzan, Ebroul Izquierdo, Hyunggon Park, Aggelos K. Katsaggelos, Johan A. Pouwelse |
Signal Process. Image Commun. | 4 |
| 2012 | Compressive Light Field SensingabstractWe propose a novel design for light field image acquisition based on compressive sensing principles. By placing a randomly coded mask at the aperture of a camera, incoherent measurements of the light passing through different parts of the lens are encoded in the captured images. Each captured image is a random linear combination of different angular views of a scene. The encoded images are then used to recover the original light field image via a novel Bayesian reconstruction algorithm. Using the principles of compressive sensing, we show that light field images with a large number of angular views can be recovered from only a few acquisitions. Moreover, the proposed acquisition and recovery method provides light field images with high spatial resolution and signal-to-noise-ratio, and therefore is not affected by limitations common to existing light field camera designs. We present a prototype camera design based on the proposed framework by modifying a regular digital camera. Finally, we demonstrate the effectiveness of the proposed system using experimental results with both synthetic and real images. S. Derin Babacan, Reto Ansorge, Martin Luessi, Pablo Ruiz 0002, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 6 |
| 2012 | Antialiasing Filter Design for Subpixel Downsampling via Frequency-Domain AnalysisabstractIn this paper, we are concerned with image downsampling using subpixel techniques to achieve superior sharpness for small liquid crystal displays (LCDs). Such a problem exists when a high-resolution image or video is to be displayed on low-resolution display terminals. Limited by the low-resolution display, we have to shrink the image. Signal-processing theory tells us that optimal decimation requires low-pass filtering with a suitable cutoff frequency, followed by downsampling. In doing so, we need to remove many useful image details causing blurring. Subpixel-based downsampling, taking advantage of the fact that each pixel on a color LCD is actually composed of individual red, green, and blue subpixel stripes, can provide apparent higher resolution. In this paper, we use frequency-domain analysis to explain what happens in subpixel-based downsampling and why it is possible to achieve a higher apparent resolution. According to our frequency-domain analysis and observation, the cutoff frequency of the low-pass filter for subpixel-based decimation can be effectively extended beyond the Nyquist frequency using a novel antialiasing filter. Applying the proposed filters to two existing subpixel downsampling schemes called direct subpixel-based downsampling (DSD) and diagonal DSD (DDSD), we obtain two improved schemes, i.e., DSD based on frequency-domain analysis (DSD-FA) and DDSD based on frequency-domain analysis (DDSD-FA). Experimental results verify that the proposed DSD-FA and DDSD-FA can provide superior results, compared with existing subpixel or pixel-based downsampling methods. Lu Fang 0001, Oscar C. Au, Ketan Tang, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2012 | Shape Error Concealment Based on a Shape-Preserving Boundary ApproximationabstractIn objectbased video representation, video scenes are composed of several arbitrarily shaped video objects (VOs), defined by their texture, shape and motion. In errorprone communications, packet loss results in missing information at the decoder. The impact of transmission errors is minimised through error concealment. In this paper, we propose a spatial error concealment technique for recovering lost shape data. We consider a geometric shape representation consisting of the object boundary, which can be extracted from the -plane. Missing macroblocks result in a broken boundary. A Bspline curve is constructed to replace a missing boundary segment, based on a T spline representation of the received boundary. We use Tsplines because they produce shapepreserving approximations and do not change the characteristics of the original boundary. The representation ensures a good estimation of the first derivatives at the points touching the missing segment. Applying smoothing conditions, we manage to construct a new spline that joins smoothly with the received boundary, leading to successful concealment results. Experimental results on object shapes with different concealment difficulty demonstrate the performance of the proposed method. Comparisons with prior proposed methods are also presented. Evaggelia Tsiligianni, Lisimachos P. Kondi, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2012 | Discovering Thematic Objects in Image Collections and VideosabstractGiven a collection of images or a short video sequence, we define a thematic object as the key object that frequently appears and is the representative of the visual contents. Successful discovery of the thematic object is helpful for object search and tagging, video summarization and understanding, etc. However, this task is challenging because 1) there lacks a priori knowledge of the thematic objects, such as their shapes, scales, locations, and times of re-occurrences, and 2) the thematic object of interest can be under severe variations in appearances due to viewpoint and lighting condition changes, scale variations, etc. Instead of using a top-down generative model to discover thematic visual patterns, we propose a novel bottom-up approach to gradually prune uncommon local visual primitives and recover the thematic objects. A multilayer candidate pruning procedure is designed to accelerate the image data mining process. Our solution can efficiently locate thematic objects of various sizes and can tolerate large appearance variations of the same thematic object. Experiments on challenging image and video data sets and comparisons with existing methods validate the effectiveness of our method. Junsong Yuan 0001, Gangqiang Zhao, Yun Fu 0001, Zhu Li 0001, Aggelos K. Katsaggelos, Ying Wu 0001 |
IEEE Trans. Image Process. | 5 |
| 2012 | Noncontact Millimeter-Wave Real-Time Detection and Tracking of Heart Rate on an Ambulatory SubjectabstractThis paper presents a solution to an aiming problem in the remote sensing of vital signs using an integration of two systems. The problem is that to collect meaningful data with a millimeter-wave sensor, the antenna must be pointed very precisely at the subject's chest. Even small movements could make the data unreliable. To solve this problem, we attached a camera to the millimeter-wave antenna, and mounted this combined system on a pan/tilt base. Our algorithm initially finds a subject's face and then tracks him/her through subsequent frames, while calculating the position of the subject's chest. For each frame, the camera sends the location of the chest to the pan/tilt base, which rotates accordingly to make the antenna point at the subject's chest. This paper presents a system for concurrent tracking and data acquisition with results from some sample scenarios. Ilya Mikhelson, Philip Greggory Lee, Sasan Bakhtiari, Thomas W. Elmer, Aggelos K. Katsaggelos, Alan V. Sahakian |
IEEE Trans. Inf. Technol. Biomed. | 5 |
| 2012 | Joint Demosaicing and Subpixel-Based Down-Sampling for Bayer Images: A Fast Frequency-Domain Analysis ApproachabstractA portable device such as a digital camera with a single sensor and Bayer color filter array (CFA) requires demosaicing to reconstruct a full color image. To display a high resolution image on a low resolution LCD screen of the portable device, it must be down-sampled. The two steps, demosaicing and down-sampling, influence each other. On one hand, the color artifacts introduced in demosaicing may be magnified when followed by down-sampling; on the other hand, the detail removed in the down-sampling cannot be recovered in the demosaicing. Therefore, it is very important to consider simultaneous demosaicing and down-sampling. Lu Fang 0001, Oscar C. Au, Yan Chen 0007, Aggelos K. Katsaggelos, Hanli Wang |
IEEE Trans. Multim. | 4 |
| 2012 | Special Issue on Subspace and Manifold Learning for Image and Video Indexing and SearchabstractThe four papers in this special issue focus on the novel design and methodology of subspace and manifold learning for image and video indexing and search. Yun Fu 0001, Xian-Sheng Hua 0001, Zhu Li 0001, Aggelos K. Katsaggelos, Thomas S. Huang |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | Low-rank matrix completion by variational sparse Bayesian learningabstractThere has been a significant interest in the recovery of low-rank matrices from an incomplete of measurements, due to both theoretical and practical developments demonstrating the wide applicability of the problem. A number of methods have been developed for this recovery problem, however, a principled method for choosing the unknown target rank is generally missing. In this paper, we present a recovery algorithm based on sparse Bayesian learning (SBL) and automatic relevance determination principles. Starting from a matrix factorization formulation and enforcing the low-rank constraint in the estimates as a sparsity constraint, we develop an approach that is very effective in determining the correct rank while providing high recovery performance. We provide empirical results and comparisons with current state-of-the-art methods that illustrate the potential of this approach. S. Derin Babacan, Martin Luessi, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP | 4 |
| 2011 | Anti-aliasing filter for subpixel down-sampling based on frequency analysisabstractNowadays, digital pictures are usually captured at very high resolution ranged up to 12 mega-pixels. Limited by low-resolution display, we have to shrink the image. Signal processing theory tells us that optimal decimation requires low-pass filtering with a suit able cut-off frequency followed by down-sampling. In doing so, we need to remove lots of details. Subpixel-based down-sampling, taking advantage of the fact that each pixel on a color LCD is actually composed of individual red, green, and blue subpixel stripes, can provide apparent higher resolution. In this paper, we use frequency domain analysis to explain what happens in subpixel-based down sampling and why it is possible to achieve a higher apparent resolution. According to our frequency domain analysis and observation, the cut-off frequency of the low-pass filter for subpixel-based decimation can be effectively extended beyond the Nyquist frequency using a novel anti-aliasing filter. Experimental results verify that the proposed subpixel down-sampling scheme based on frequency analysis (SDSFA) can give superior results compared with existing pixel-based down-sampling methods. Lu Fang 0001, Ketan Tang, Oscar C. Au, Aggelos K. Katsaggelos |
ICASSP | 4 |
| 2011 | Compressive passive millimeter-wave imagingabstractIn this paper, we present a novel passive millimeter-wave (PMMW) imaging system designed using compressive sensing principles. We employ randomly encoded masks at the focal plane of the PMMW imager to acquire incoherent measurements of the imaged scene. We develop a Bayesian reconstruction algorithm to estimate the original image from these measurements, where the sparsity inherent to typical PMMW images is efficiently exploited. Comparisons with other existing reconstruction methods show that the proposed reconstruction algorithm provides higher quality image estimates. Finally, we demonstrate with simulations using real PMMW images that the imaging duration can be dramatically reduced by acquiring only a few measurements compared to the size of the image. S. Derin Babacan, Martin Luessi, Leonidas Spinoulas, Aggelos K. Katsaggelos, Nachappa Gopalsami, Thomas W. Elmer, Ryan Ahern, Shaolin Liao, Apostolos C. Raptis |
ICIP | 4 |
| 2011 | Tracking-optimized quantization for H.264 compression in transportation video surveillance applicationsabstractWe propose a tracking-aware system that removes video components of low tracking interest and optimizes the quantization during compression of frequency coefficients, particularly those that most influence trackers, significantly reducing bitrate while maintaining comparable tracking accuracy. We utilize tracking accuracy as our compression criterion in lieu of mean squared error metrics. The process of optimizing quantization tables suitable for automated tracking can be executed online or offline. The online implementation initializes the encoding procedure for a specific scene, but introduces delay. On the other hand, the offline procedure produces globally optimum quantization tables where the optimization occurs for a collection of video sequences. Our proposed system is designed with low processing power and memory requirements in mind, and as such can be deployed on remote nodes. Using H.264/AVC video coding and a commonly used state-of-the-art tracker we show that while maintaining comparable tracking accuracy our system allows for over 50% bitrate savings on top of existing savings from previous work. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2011 | Channel protection for H.264 compression in transportation video surveillance applicationsabstractThe compression of video and subsequent partial loss of the compressed bitstream can dramatically reduce the accuracy of automated tracking algorithms. This is problematic for centralized applications such as transportation surveillance systems, where remotely captured and compressed video is transmitted over lossy wireless links to a central location for tracking. We propose a low-complexity method for protecting compressed video against channel loss such that the tracking accuracy of decoded and concealed video is maximized. Our algorithm leverages a previous method of video processing that removes components of low tracking interest before compression to minimize bitrate, and uses some of the bitrate savings to introduce redundancy into the transmitted bitstream to reduce the probability of information loss. We show using a common tracker and loss concealment algorithm that our system allows for up to 100% increased tracking accuracy at a given bitrate, or 90% bitrate savings for comparable tracking quality. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2011 | Bayesian TV denoising of SAR imagesabstractSynthetic aperture radar (SAR) imagery suffers from the speckle phenomenon. Speckle gives rise to the presence of multiplicative noise which severely degrades the observed images. It is known that logarithmically transformed speckle can be well approximated by a Gaussian distribution. In this paper we propose an algorithm for despeckling images, within the log-transformed spatial domain, using a TV prior whose model parameter is automatically determined using the Evidence Analysis within the Hierarchical Bayesian Paradigm. The effectiveness of the proposed algorithm, over both synthetically speckled and real SAR images, is studied. Miguel Vega, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2011 | A novel iterative image restoration algorithm using nonstationary image priorsabstractIn this paper, we propose a novel algorithm for image restoration based on combining nonstationary edge-preserving priors. We develop a Bayesian modeling followed by an evidence analysis inference approach for deriving the foundations of the proposed iterative restoration algorithm. Simulation results over a variety of blurred and noisy standard test images indicate that the presented method outperforms current state-of-the-art image restoration algorithms. We finally present experimental results by digitally refocusing images captured with controlled defocus, successfully confirming the ability of the proposed restoration algorithm in recovering extra features and details, while still preserving edges. Esteban Vera, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2011 | Adaptive joint demosaicing and Subpixel-based Down-sampling for Bayer imageabstractA digital camera provided with a Bayer pattern single sensor needs color interpolation to reconstruct a full color image. To show high resolution image on a lower resolution display, it must then be down-sampled. These two steps influence each other, i.e., the color artifacts introduced in demosaicing may be magnified in subsequent down-sampling process and vice versa. Thanks to the fact that LCD displays are actually composed of separable subpixels, which can be individually addressed to achieve a higher effective apparent resolution. This paper presents an Adaptive Joint Demosaicing and Subpixel-based Down-sampling scheme (AJDSD) for single-sensor camera image, where the subpixel-based down-sampling is adaptively and directly applied in Bayer domain, without the process of demosaicing. Simulation results demonstrate that when compared with conventional “demosaicing-first and down-sampling-later” methods, AJDSD achieves superior performance improvement in terms of computational complexity. As for visual quality, AJDSD is more effective in preserving high frequency details, leading to much sharper and clearer results. Lu Fang 0001, Oscar C. Au, Aggelos K. Katsaggelos |
ICME | 3 |
| 2011 | Video retrieval using sparse Bayesian reconstructionabstractEvery day, a huge amount of video data is generated for different purposes and applications. Fast and accurate algorithms for efficient video search and retrieval are therefore essential. The interesting properties of sparse representation and the new sampling theory named Compressive Sensing (CS) constitute the core of the new approach to video representation and retrieval we are presenting in this paper. Once the representation (where sparsity is expected) has been chosen and the observations have been taken, the proposed approach utilizes Bayesian modeling and inference to tackle the retrieval problem. In order to speed up the inference process the use of Principal Components Analysis (PCA) to provide an alternative representation of the frames is analyzed. Experimental results validate the proposed approach whose robustness against noise is also examined. Pablo Ruiz 0002, S. Derin Babacan, Zhu Li 0001, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICME | 6 |
| 2011 | Understanding dynamic scenes by hierarchical motion pattern miningabstractOur work addresses the problem of analyzing and understanding dynamic video scenes. A two-level motion pattern mining approach is proposed. At the first level, single-agent motion patterns are modeled as distributions over pixel-based features. At the second level, interaction patterns are modeled as distributions over single-agent motion patterns. Both patterns are shared among video clips. Compared to other works, the advantage of our method is that interaction patterns are detected and assigned to every video frame. This enables a finer semantic interpretation and more precise anomaly detection. Specifically, every video frame is labeled by a certain interaction pattern and moving pixels in each frame which do not belong to any singleagent pattern or cannot exist in the corresponding interaction pattern are detected as anomalies. We have tested our approach on a challenging traffic surveillance sequence containing both pedestrian and vehicular motions and obtained promising results. Zhongke Shi, Aggelos K. Katsaggelos |
ICME | 4 |
| 2011 | Indexed spatio-temporal appearance models for query-driven video action recognitionabstractVideo action and event recognition is an important problem in video analysis research with many important applications, such as surveillance and video search. In this work, we deal with the appearance complexity in video action recognition by applying an indexing structure and partition in appearance space. The task requires spatio-temporal appearance modeling that can capture the discriminative information among different action classes. Traditional approaches are based on a global appearance model, which is not robust to local variations in background. In this work, we develop a query driven dynamic appearance modeling method and use a localized subspace to obtain a distance metric for appearance discrimination. Multiple localized models are constructed and utilized to measure the similarity between the trajectories and the sub-space metric is adaptive during the learning process. The processing is implemented based on an indexing scheme, which is very fast in computation. Simulation results demonstrate the effectiveness of the solution. Haomian Zheng, Zhu Li 0001, Aggelos K. Katsaggelos |
ICME | 3 |
| 2011 | Design of benchmark imagery for validating facility annotation algorithmsabstractThe design of benchmark imagery for validation of image an notation algorithms is considered. Emphasis is placed on imagery that contains industrial facilities, such as chemical re fineries. An application-level facility ontology is used as a means to define salient objects in the benchmark imagery. Instrinsic and extrinsic scene factors important for comprehensive validation are listed, and variability in the benchmarks discussed. Finally, the pros and cons of three forms of bench mark imagery: real, composite and synthetic, are delineated. Randy S. Roberts, Paul A. Pope, Ranga Raju Vatsavai, Ming Jiang 0005, Lloyd F. Arrowood, Timothy G. Trucano, Shaun S. Gleason, Anil M. Cheriyadat, Alexandre Sorokine, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas, Lucinda R. Gaines, Lawrence K. Chilton |
IGARSS | 10 |
| 2011 | Anomalous video event detection using spatiotemporal context
Junsong Yuan 0001, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
Comput. Vis. Image Underst. | 4 |
| 2011 | Low-Complexity Tracking-Aware H.264 Video Compression for Transportation SurveillanceabstractIn centralized transportation surveillance systems, video is captured and compressed at low processing power remote nodes and transmitted to a central location for processing. Such compression can reduce the accuracy of centrally run automated object tracking algorithms. In typical systems, the majority of communications bandwidth is spent on encoding temporal pixel variations such as acquisition noise or local changes to lighting. We propose a tracking-aware, H.264-compliant compression algorithm that removes temporal components of low tracking interest and optimizes the quantization of frequency coefficients, particularly those that most influence trackers, significantly reducing bitrate while maintaining comparable tracking accuracy. We utilize tracking accuracy as our compression criterion in lieu of mean squared error metrics. Our proposed system is designed with low processing power and memory requirements in mind, and as such can be deployed on remote nodes. Using H.264/AVC video coding and a commonly used state-of-the-art tracker we show that our algorithm allows for over 90% bitrate savings while maintaining comparable tracking accuracy. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Variational Bayesian Super ResolutionabstractIn this paper, we address the super resolution (SR) problem from a set of degraded low resolution (LR) images to obtain a high resolution (HR) image. Accurate estimation of the sub-pixel motion between the LR images significantly affects the performance of the reconstructed HR image. In this paper, we propose novel super resolution methods where the HR image and the motion parameters are estimated simultaneously. Utilizing a bayesian formulation, we model the unknown HR image, the acquisition process, the motion parameters and the unknown model parameters in a stochastic sense. Employing a variational bayesian analysis, we develop two novel algorithms which jointly estimate the distributions of all unknowns. The proposed framework has the following advantages: 1) Through the incorporation of uncertainty of the estimates, the algorithms prevent the propagation of errors between the estimates of the various unknowns; 2) the algorithms are robust to errors in the estimation of the motion parameters; and 3) using a fully bayesian formulation, the developed algorithms simultaneously estimate all algorithmic parameters along with the HR image and motion parameters, and therefore they are fully-automated and do not require parameter tuning. We also show that the proposed motion estimation method is a stochastic generalization of the classical Lucas-Kanade registration algorithm. Experimental results demonstrate that the proposed approaches are very effective and compare favorably to state-of-the-art SR algorithms. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2010 | Fast total variation image restoration with parameter estimation using bayesian inferenceabstractIn this paper we propose two fast Total Variation (TV) based algorithms for image restoration by utilizing variational posterior distribution approximation. The unknown image and the hyperparameters for the image and observation models are formulated and estimated simultaneously within a hierachical Bayesian framework, rendering the algorithms fully-automated without any free parameters. Experimental results demonstrate that the proposed algorithms provide restoration results competitive to existing methods in terms of image quality while achieving superior computational efficiency. Bruno Amizic, S. Derin Babacan, Michael Kwok-Po Ng, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP | 5 |
| 2010 | Symmetrical EEG/FMRI fusion with spatially adaptive priors using variational distribution approximationabstractIn this paper, we propose a symmetrical EEG/fMRI fusion algorithm which combines EEG and fMRI by means of a common generative model. The use of a total variation (TV) prior as well as spatially adaptive temporal priors enables adaptation to the local characteristics of the estimated responses. We utilize an approximate variational Bayesian framework and obtain a fully automatic fusion algorithm. Simulation results demonstrate that the proposed algorithm outperforms existing EEG/fMRI fusion methods. Martin Luessi, S. Derin Babacan, Rafael Molina 0001, James R. Booth, Aggelos K. Katsaggelos |
ICASSP | 5 |
| 2010 | Content-aware H.264 encoding for traffic video tracking applicationsabstractThe compression of video can reduce the accuracy of tracking algorithms, which is problematic for centralized applications that rely on remotely captured and compressed video for input. We show the effects of high compression on the features commonly used in real-time video object tracking. We propose a computationally efficient Region of Interest (ROI) extraction method, which is used during standard-compliant H.264 encoding to concentrate bitrate on regions in video most likely to contain objects of tracking interest (vehicles). This algorithm is shown to significantly increase tracking accuracy, which is measured by employing a commonly used automatic tracker. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2010 | Sparse Bayesian image restorationabstractIn this paper we propose a novel Bayesian algorithm for image restoration and parameter estimation. We utilize an image prior where Gaussian distributions are placed per pixel in the high-pass filter outputs of the image. By following the hierarchical Bayesian framework, we simultaneously estimate the unknown image and hyperparameters for both the image prior and the image degradation noise. We show that the proposed formulation is a special case of the popular lp-norm based formulations with p = 0, and therefore enforces sparsity to an high extent in the filtered image coefficients. Moreover, the proposed formulation results in a convex optimization problem, and therefore does not suffer from the robustness issues common with non-convex image priors. Experimental results demonstrate that the proposed algorithm provides superior performance compared to state-of-the-art restoration algorithms although no user-supervision is required. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2010 | Video anomaly detection in spatiotemporal contextabstractCompared to other approaches that analyze object trajectories, we propose to detect anomalous video events at three levels considering spatiotemporal context of video objects, i.e., point anomaly, sequential anomaly, and co-occurrence anomaly. A hierarchical data mining approach is proposed to achieve this task. At each level, the frequency based analysis is performed to automatically discover regular rules of normal events. The events deviating from these rules are detected as anomalies. Experiments on real traffic video prove that the detected video anomalies are hazardous or illegal according to the traffic rule. Junsong Yuan 0001, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2010 | In-sequence video duplicate detection with fast point-to-line matchingabstractA computational geometry approach is developed to detect video duplicate with mild transformations. We model the video sequence as a trajectory after scaling and projection. Through interpolation and equal curve length sampling, part of the frame points is selected. A simplified video representation is the line segment set connecting the left neighboring points. For a given query, match distortion is calculated by projecting the query frame points to the line segment set guided by the frame temporal relationship. Experiments demonstrate the effectiveness of the proposed approach. Bo Liu 0005, Zhu Li 0001, Meng Wang 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2010 | A game theoretic approach to video streaming over peer-to-peer networksabstractWe address the problem of content-aware, foresighted resource reciprocation for media streaming over peer-to-peer (P2P) networks. The envisioned P2P network consists of autonomous and self-interested peers trying to maximize their individual utilities. The resource reciprocation among such peers is modeled as a stochastic game and peers determine the optimal strategies for resource reciprocation using a Markov Decision Process (MDP) framework. Unlike existing solutions, this framework takes the content and the characteristics of the video signal into account by introducing an artificial currency in order to maximize the video quality in the entire network. Ehsan Maani, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2010 | Quantization optimized H.264 encoding for traffic video tracking applicationsabstractThe compression of video can reduce the accuracy of post-compression tracking algorithms. This is problematic for centralized applications such as traffic surveillance systems, where remotely captured and compressed video is transmitted to a central location for tracking. We propose a low complexity optimization framework that automatically identifies video features critical to tracking and concentrates bitrate on these features via quantization tables. Using the H.264 video coding standard and two commonly used state-of-the-art trackers we show that our algorithm allows for over 60% bitrate savings while maintaining comparable tracking accuracy. Eren Soyak, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2010 | Shape error concealment based on a shape-preserving boundary approximationabstractIn error-prone communications, packet loss results in missing information of shape, motion and texture of a video object (VO). Error concealment refers to the recovery of lost information at the decoder. In this paper, we propose a spatial shape error concealment technique. We consider a geometric representation of the shape of a VO consisting of its boundary, which can be extracted from the received a-plane. Some boundary parts are missing due to errors. We propose a method for modeling the received boundary based on a shape-preserving approximation that uses T-splines. Such an approximation provides a good estimation of the direction of a missing boundary segment, which we use to construct a concealment spline that joins smoothly with the received boundary parts. Evaggelia Tsiligianni, Lisimachos P. Kondi, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2010 | Using the Kullback-Leibler divergence to combine image priors in Super-Resolution image reconstructionabstractThis paper is devoted to the combination of image priors in Super Resolution (SR) image reconstruction. Taking into account that each combination of a given observation model and a prior model produces a different posterior distribution of the underlying High Resolution (HR) image, the use of variational posterior distribution approximation on each posterior will produce as many posterior approximations as priors we want to combine. A unique approximation is obtained here by finding the distribution on the HR image given the observations that minimizes a linear convex combination of the Kullback-Leibler divergences associated with each posterior distribution. We find this distribution in closed form and also relate the proposed approach to other prior combination methods in the literature. The estimated HR images are compared with images provided by other SR reconstruction methods. Salvador Villena, Miguel Vega, S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2010 | A novel image retrieval framework exploring inter cluster distanceabstractAn image could be described with local features like SIFT and with those features, images could be represented as “Bag-of-Visual-Words” (BVW). This representation has been widely used in content based image retrieval. Comparing BVW of two images is usually done in Euclidean space, like Euclidean distance or weighted variants. Neither of these methods consider the inter cluster relations. If there is a feature in one image without any match in all the clusters of another image's features, there will be no score for that feature. But, there are still some match in neighbor clusters. In this paper, we use dynamic programming to calculate full inter cluster distance map and with the distance, we can evaluate a feature in neighbor clusters. Our proposed method is evaluated in Caltech 101 database and experiments show that our method generally exceeds the method that don't consider inter cluster distance. Xin Xin 0009, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2010 | Part-based initialization for hand trackingabstractInitializing hand/finger articulation for tracking is a very challenging problem, mainly because hand articulation is complicated and it has a large number of degrees of freedom. Most existing algorithms initialize tracking manually, or use a nearest-neighbor search with restricting the number of possible hand gestures. This paper presents a new solution to this problem by increasing the dimensionality but taking advantage of the sparseness. The basic idea is to divide the set of phalange joint angles into many overlapping subsets. As each subset has a much smaller number of joint angles, it is much easier to design a smaller-scale articulation estimator. The estimation of the whole hand is done by the collaboration of a network of dependent smaller-scale estimators. This paper describes a novel way of designing the smaller-scale estimators as well as a principled way of fusing the estimates. A tracking system is also shown by using this initialization technique. Jiang Xu 0002, Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2010 | Geospatial image mining for nuclear proliferation detection: Challenges and new opportunitiesabstractWith increasing understanding and availability of nuclear technologies, and increasing persuasion of nuclear technologies by several new countries, it is increasingly becoming important to monitor the nuclear proliferation activities. There is a great need for developing technologies to automatically or semi-automatically detect nuclear proliferation activities using remote sensing. Images acquired from earth observation satellites is an important source of information in detecting proliferation activities. High-resolution remote sensing images are highly useful in verifying the correctness, as well as completeness of any nuclear program. DOE national laboratories are interested in detecting nuclear proliferation by developing advanced geospatial image mining algorithms. In this paper we describe the current understanding of geospatial image mining techniques and enumerate key gaps and identify future research needs in the context of nuclear proliferation. Ranga Raju Vatsavai, Budhendra L. Bhaduri, Anil M. Cheriyadat, Lloyd F. Arrowood, Eddie A. Bright, Shaun S. Gleason, Carl Diegert, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas, Reid B. Porter, Jim Bollinger, Barry Chen, Ryan Hohimer |
IGARSS | 8 |
| 2010 | Audio-visual anticipatory coarticulation modeling by human and machineabstractThe phenomenon of anticipatory coarticulation provides a ba-sis for the observed asynchrony between the acoustic and vi-sual onsets of phones in certain linguistic contexts. This type of asynchrony is typically not explicitly modeled in audio-visual speech models. In this work, we study within-word audio-visual asynchrony using manual labels of words in which theory suggests that audio-visual asynchrony should occur, and show that these hand labels confirm the theory. We then introduce a new statistical model of audio-visual speech, the asynchrony-dependent transition (ADT) model. This model allows asyn-chrony between audio and video states within word boundaries, where the audio and video state transitions depend not only on the state of that modality, but also on the instantaneous asyn-chrony. The ADT model outperforms a baseline synchronous model in mimicking the hand labels in a forced alignment task, and its behavior as parameters are changed conforms to our ex-pectations about anticipatory coarticulation. The same model could be used for speech recognition, although here we consider it only for the task of forced alignment for linguistic analysis. Index Terms: audio-visual speech recognition, audio-visual asynchrony, anticipatory coarticulation, dynamic Bayesian net-works 1. Louis H. Terry, Karen Livescu, Janet B. Pierrehumbert, Aggelos K. Katsaggelos |
INTERSPEECH | 4 |
| 2010 | Unequal Error Protection for Robust Streaming of Scalable Video Over Packet Lossy NetworksabstractEfficient bit stream adaptation and resilience to packet losses are two critical requirements in scalable video coding for transmission over packet-lossy networks. Various scalable layers have highly distinct importance, measured by their contribution to the overall video quality. This distinction is especially more significant in the scalable H.264/advanced video coding (AVC) video, due to the employed prediction hierarchy and the drift propagation when quality refinements are missing. Therefore, efficient bit stream adaptation and unequal protection of these layers are of special interest in the scalable H.264/AVC video. This paper proposes an algorithm to accurately estimate the overall distortion of decoder reconstructed frames due to enhancement layer truncation, drift/error propagation, and error concealment in the scalable H.264/AVC video. The method recursively computes the total decoder expected distortion at the picture-level for each layer in the prediction hierarchy. This ensures low computational cost since it bypasses highly complex pixel-level motion compensation operations. Simulation results show an accurate distortion estimation at various channel loss rates. The estimate is further integrated into a cross-layer optimization framework for optimized bit extraction and content-aware channel rate allocation. Experimental results demonstrate that precise distortion estimation enables our proposed transmission system to achieve a significantly higher average video peak signal-to-noise ratio compared to a conventional content independent system. Ehsan Maani, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Application-Centric Routing for Video Streaming Over MultiHop Wireless NetworksabstractMost existing works on routing for video transmission over multihop wireless networks only focus on how to satisfy the network-oriented quality-of-service (QoS), such as through put, delay, and packet loss rate rather than application-oriented QoS such as the user-perceived video quality. Although there are some research efforts which use application-centric video quality as the routing metric, they either calculate the video quality based on some predefined rate-distortion function or model without considering the impact of video coding and decoding (including error concealment) on routing, or use exhaustive search or heuristic methods to find the optimal path, leading to high computational complexity and/or suboptimal solutions. In this paper, we propose an application-centric routing framework for real-time video transmission in multihop wireless networks, where expected video distortion is adopted as the routing metric. The major contributions of this paper are: 1) the development of an efficient routing algorithm with the routing metric expressed in terms of the expected video distortion and being calculated on-the-fly, and 2) the development of a quality-driven cross-layer optimization framework to enhance the flexibility and robustness of routing by the joint optimization of routing path selection and video coding, thereby maximizing the user-perceived video quality under a given video playback delay constraint. Both theoretical and experimental results demonstrate that the proposed quality-driven application-centric routing approach can achieve a superior performance over existing network-centric routing approaches. Dalei Wu, Song Ci, Haohong Wang, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Bayesian Compressive Sensing Using Laplace PriorsabstractIn this paper, we model the components of the compressive sensing (CS) problem, i.e., the signal acquisition process, the unknown signal coefficients and the model parameters for the signal and noise using the Bayesian framework. We utilize a hierarchical form of the Laplace prior to model the sparsity of the unknown signal. We describe the relationship among a number of sparsity priors proposed in the literature, and show the advantages of the proposed model including its high degree of sparsity. Moreover, we show that some of the existing models are special cases of the proposed model. Using our model, we develop a constructive (greedy) algorithm designed for fast reconstruction useful in practical settings. Unlike most existing CS reconstruction methods, the proposed algorithm is fully automated, i.e., the unknown signal coefficients and all necessary parameters are estimated solely from the observation, and, therefore, no user-intervention is needed. Additionally, the proposed algorithm provides estimates of the uncertainty of the reconstructions. We provide experimental results with synthetic 1-D signals and images, and compare with the state-of-the-art CS reconstruction algorithms demonstrating the superior performance of the proposed approach. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2010 | Bayesian Blind Deconvolution From Differently Exposed Image PairsabstractPhotographs acquired under low-lighting conditions require long exposure times and therefore exhibit significant blurring due to the shaking of the camera. Using shorter exposure times results in sharper images but with a very high level of noise. In this paper we address the problem of utilizing two such images in order to obtain an estimate of the original scene and present a novel blind deconvolution algorithm for solving it. We formulate the problem in a hierarchical Bayesian framework by utilizing prior knowledge on the unknown image and blur, and also on the dependency between the two observed images. By incorporating a fully Bayesian analysis, the developed algorithm estimates all necessary model parameters along with the unknown image and blur, such that no user-intervention is needed. Moreover, we employ a variational Bayesian inference procedure, which allows for the statistical compensation of errors occurring at different stages of the restoration, and also provides uncertainties of the estimates. Experimental results with synthetic and real images demonstrate that the proposed method provides very high quality restoration results and compares favorably to existing methods even though no user supervision is needed. S. Derin Babacan, Jingnan Wang, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2010 | Maximum a Posteriori Video Super-Resolution Using a New Multichannel Image PriorabstractSuper-resolution (SR) is the term used to define the process of estimating a high-resolution (HR) image or a set of HR images from a set of low-resolution (LR) observations. In this paper we propose a class of SR algorithms based on the maximum a posteriori (MAP) framework. These algorithms utilize a new multichannel image prior model, along with the state-of-the-art single channel image prior and observation models. A hierarchical (two-level) Gaussian nonstationary version of the multichannel prior is also defined and utilized within the same framework. Numerical experiments comparing the proposed algorithms among themselves and with other algorithms in the literature, demonstrate the advantages of the adopted multichannel approach. Stefanos P. Belekos, Nikolas P. Galatsanos, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2010 | Variational Bayesian Image Restoration With a Product of Spatially Weighted Total Variation Image PriorsabstractIn this paper, a new image prior is introduced and used in image restoration. This prior is based on products of spatially weighted total variations (TV). These spatial weights provide this prior with the flexibility to better capture local image features than previous TV based priors. Bayesian inference is used for image restoration with this prior via the variational approximation. The proposed restoration algorithm is fully automatic in the sense that all necessary parameters are estimated from the data and is faster than previous similar algorithms. Numerical experiments are shown which demonstrate that image restoration based on this prior compares favorably with previous state-of-the-art restoration algorithms. Giannis K. Chantas, Nikolas P. Galatsanos, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2009 | Parameter Estimation in Bayesian Super-Resolution Image Reconstruction from Low Resolution Rotated and Translated Images
Salvador Villena, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
ACIVS | 4 |
| 2009 | Fast bayesian compressive sensing using Laplace priorsabstractIn this paper we model the components of the compressive sensing (CS) problem using the Bayesian framework by utilizing a hierarchical form of the Laplace prior to model sparsity of the unknown signal. This signal prior includes some of the existing models as special cases and achieves a high degree of sparsity. We develop a constructive (greedy) algorithm resulting from this formulation where necessary parameters are estimated solely from the observation and therefore no user-intervention is needed. We provide experimental results with synthetic 1D signals and images, and compare with the state-of-the-art CS reconstruction algorithms demonstrating the superior performance of the proposed approach. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2009 | Compressive sensing of light fieldsabstractWe propose a novel camera design for light field image acquisition using compressive sensing. By utilizing a randomly coded non-refractive mask in front of the aperture, incoherent measurements of the light passing through different regions are encoded in the captured images. A novel reconstruction algorithm is proposed to recover the original light field image from these acquisitions. Using the principles of compressive sensing, we demonstrate that light field images with high angular dimension can be captured with only a few acquisitions. Moreover, the proposed design provides images with high spatial resolution and signal-to-noise-ratio (SNR), and therefore does not suffer from limitations common to existing light-field camera designs. Experimental results demonstrate the efficiency of the proposed system. S. Derin Babacan, Reto Ansorge, Martin Luessi, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2009 | Bayesian blind deconvolution from differently exposed image pairsabstractPhotographs acquired under low-light conditions require long exposure times and therefore exhibit significant blurring due to the shaking of the camera. Using shorter exposure times results in sharper images but with a very high level of noise. In this paper we address this problem and present a novel blind deconvolution algorithm for a pair of differently exposed images. We formulate the problem in a hierarchical Bayesian framework by utilizing prior knowledge on the unknown image and blur, and also on the dependency between two observed images. By incorporating a fully Bayesian analysis, the developed algorithm estimates all necessary algorithm parameters along with the unknowns, such that no user-intervention is needed. Moreover, we employ a variational Bayesian inference procedure, which allows for the statistical compensation of errors occurring at different stages of the restoration, and also provides uncertainties of the estimates. Experimental results demonstrate the high restoration performance of the proposed algorithm. S. Derin Babacan, Jingnan Wang, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2009 | Maximum a posteriori super-resolution of compressed video using a new multichannel image priorabstractSuper-resolution (SR) algorithms for compressed video aim at recovering high-frequency information and estimating a high-resolution (HR) image or a set of HR images from a sequence of low-resolution (LR) video frames. In this paper we present a novel SR algorithm for compressed video based on the maximum a posteriori (MAP) framework. We utilize a new multichannel image prior model, along with the state-of-the art image prior and observation models. Moreover, relationship between model parameters and the decoded bitstream are established. Numerical experiments demonstrate the improved performance of the proposed method compared to existing algorithms for different compression ratios. Stefanos P. Belekos, Nikolas P. Galatsanos, S. Derin Babacan, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2009 | A video retrieval algorithm using random projectionsabstractIn this paper, we propose a fast and accurate video retrieval algorithm using random projections, an indexing structure and parallel computing. The video sequences are represented by low dimensional temporal trajectories in a set of low dimensional spaces through scaling and random projections. A kd-tree structure is used for efficient data access. We also develop an efficient retrieval algorithm. Simulation results demonstrate that the proposed algorithm is very fast and accurate in retrieval performance. Zhu Li 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2009 | Detecting contextual anomalies of crowd motion in surveillance videoabstractMany works have been proposed on detecting individual anomalies in crowd scenes, i.e., human behaviors anomalous with respect to the rest of the behaviors. In this paper, we introduce a new concept of contextual anomaly into the field of crowd analysis, i.e., the behaviors themselves are normal but they are anomalous in a specific context. Our system follows an unsupervised approach. It automatically discovers important contextual information from the crowd video and detects the blobs corresponding to contextually anomalous behaviors. Our experiments show that the approach works well in detecting contextual anomalies from crowd video with different motion contexts. Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2009 | Efficient motion compensated frame rate upconversion using multiple interpolations and median filteringabstractMethods for motion compensated frame rate upconversion exploit motion in order to generate interpolated frames that are temporally located between the available original frames of a video sequence. When the interpolated frames are inserted into the original sequence, the resulting sequence has a higher visual quality due to the increased frame rate. We propose a method for motion compensated frame rate upconversion which reduces upconversion artifacts by combining multiple intermediate interpolations utilizing median filtering. We use an efficient block based motion estimation method which makes use of motion vectors extracted from an H.264 bitstream. Low computational complexity makes the method suitable for real time operation. Martin Luessi, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2009 | Image restoration by mixture modelling of an overcomplete linear representationabstractWe present a new image restoration method based on modelling the coefficients of an overcomplete wavelet response to natural images with a mixture of two Gaussian distributions, having non-zero and zero mean respectively, and reflecting the assumption that this response is close to be sparse. Including the observation model, the resulting procedure iterates between image reconstruction from the hard-thresholding of the response to the current estimate and a fast blur compensation step. Results indicate that our method compares favorably with current wavelet-based restoration methods. Luis Mancera, S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2009 | Local Bayesian image restoration using variational methods and Gamma-Normal distributionsabstractIn this paper we present a new Bayesian methodology for the restoration of blurred and noisy images. Bayesian methods rely on image priors that encapsulate prior image knowledge and avoid the ill-posedness of image restoration problems. We use a spatially varying image prior utilizing a gamma-normal hyperprior distribution on the local precision parameters. This kind of hyperprior distribution, which to our knowledge has not been used before in image restoration, allows for the incorporation of information on local as well as global image variability, models correlation of the local precision parameters and is a conjugate hyperprior to the image model used in the paper. The proposed restoration technique is compared with other image restoration approaches, demonstrating its improved performance. Javier Mateos, Tom E. Bishop, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2009 | Two-dimensional channel coding for scalable H.264/AVC videoabstractEfficient bit stream adaptation and resilience to packet losses are two critical requirements in scalable video coding for transmission over packet-lossy networks. These requirements have a greater significance in scalable H.264/AVC video bit streams since missing refinement information in a layer propagates to all lower layers in the prediction hierarchy and causes substantial degradation in video quality. This work proposes an algorithm to accurately estimate the overall distortion of the reconstructed frames due to enhancement layer truncation, drift/error propagation, and error concealment in the scalable H.264/AVC video. This ensures low computational cost since it bypasses highly complex pixel-level motion compensation operations. Simulation results show an accurate distortion estimation at various channel loss rates. The estimate is further integrated into a cross-layer optimization framework for optimized bit extraction and content-aware channel rate allocation. Experimental results demonstrate that precise distortion estimation enables our proposed transmission system to achieve a significantly higher average video PSNR compared to a conventional content independent system. Ehsan Maani, Aggelos K. Katsaggelos |
PCS | 2 |
| 2009 | Application-Centric Routing for Video Streaming over Multi-hop Wireless NetworksabstractRouting for video transmissions over multi-hop wireless networks has gained increasing research interest in recent years. However, most existing works only focus on how to satisfy the network-oriented QoS, such as, throughput, delay, and packet loss rate rather than the user perceived quality. Although there are some research efforts which use application-centric video quality as the routing metric, the calculation of video quality is based on some predefined rate-distortion function or model without exploring the impact of video coding and decoding (including error concealment) on network path selection and the resulting received video quality. Moreover, unlike network- centric routing metrics, such as, hop count, average delay or average success probability of packet transmission, video distortion cannot be calculated either additively or multiplicatively in a hop-by-hop fashion due to the dependency among packets introduced by error concealment. As a result, most existing works use either exhaustive search or heuristic methods to find the optimal path, which leads to high computational complexity or suboptimal solutions to the routing problem of video transmission. In this paper, we propose an application- centric routing framework for real-time video transmission over multi-hop wireless networks, where expected video distortion is used as the routing metric. The major contributions of this work are: 1) the development of an efficient routing algorithm with the routing metric in terms of the expected video distortion being calculated on-the-fly, and 2) the development of a quality-driven cross-layer optimization framework to enhance the flexibility and robustness of routing by the joint optimization of routing path selection and video coding, thereby maximizing the user perceived video quality under a given video playback delay constraint. Both theoretical and experimental results demonstrate that the proposed quality-driven application-centric routing approach can achieve a superior performance over existing network-centric routing approaches. Dalei Wu, Song Ci, Haiyan Luo, Haohong Wang, Aggelos K. Katsaggelos |
SECON | 5 |
| 2009 | Guest EditorialabstractJournal Article Guest Editorial Get access Aggelos K. Katsaggelos, Aggelos K. Katsaggelos 1Department of Electrical Engineering and Computer Science, Northwestern University, Evanston, IL, USA Search for other works by this author on: Oxford Academic Google Scholar Rafael Molina Rafael Molina * 2Department of Computer Science and Artificial Intelligence, University of Granada, Granada, Spain *Corresponding author:[email protected] Search for other works by this author on: Oxford Academic Google Scholar The Computer Journal, Volume 52, Issue 4, July 2009, Pages 395–396, https://doi.org/10.1093/comjnl/bxp029 Published: 20 April 2009 Aggelos K. Katsaggelos, Rafael Molina 0001 |
Comput. J. | 1 |
| 2009 | Super-Resolution of Multispectral ImagesabstractIn this paper we propose and analyze a globally and locally adaptive super-resolution Bayesian methodology for pansharpening of multispectral images. The methodology incorporates prior knowledge on the expected characteristics of the multispectral images uses the sensor characteristics to model the observation process of both panchromatic and multispectral images and includes information on the unknown parameters in the model in the form of hyperprior distributions. Using real and synthetic data, the pansharpened multispectral images are compared with the images obtained by other pansharpening methods and their quality is assessed both qualitatively and quantitatively. Miguel Vega, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
Comput. J. | 4 |
| 2009 | An Efficient Video Indexing and Retrieval Algorithm Using the Luminance Field Trajectory ModelingabstractWith the phenomenal growth of the online and personal video repositories, an efficient and robust example-based video search solution is required to support applications like query by clip, query by capture, and repeated clip detection. In this letter, video sequences are represented as temporal trajectories via scaling and lower dimensional representation of the video frame luminance field, and a video trajectory indexing and matching scheme is developed to support video clip search. Simulation results demonstrate that the proposed approach achieves excellent performance in both response speed and precision-recall accuracy. Zhu Li 0001, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Variational Bayesian Blind Deconvolution Using a Total Variation PriorabstractIn this paper, we present novel algorithms for total variation (TV) based blind deconvolution and parameter estimation utilizing a variational framework. Using a hierarchical Bayesian model, the unknown image, blur, and hyperparameters for the image, blur, and noise priors are estimated simultaneously. A variational inference approach is utilized so that approximations of the posterior distributions of the unknowns are obtained, thus providing a measure of the uncertainty of the estimates. Experimental results demonstrate that the proposed approaches provide higher restoration performance than non-TV-based methods without any assumptions about the unknown hyperparameters. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2009 | SoftCuts: A Soft Edge Smoothness Prior for Color Image Super-ResolutionabstractDesigning effective image priors is of great interest to image super-resolution (SR), which is a severely under-determined problem. An edge smoothness prior is favored since it is able to suppress the jagged edge artifact effectively. However, for soft image edges with gradual intensity transitions, it is generally difficult to obtain analytical forms for evaluating their smoothness. This paper characterizes soft edge smoothness based on a novel SoftCuts metric by generalizing the Geocuts method . The proposed soft edge smoothness measure can approximate the average length of all level lines in an intensity image. Thus, the total length of all level lines can be minimized effectively by integrating this new form of prior. In addition, this paper presents a novel combination of this soft edge smoothness prior and the alpha matting technique for color image SR, by adaptively normalizing image edges according to their alpha-channel description. This leads to the adaptive SoftCuts algorithm, which represents a unified treatment of edges with different contrasts and scales. Experimental results are presented which demonstrate the effectiveness of the proposed method. Shengyang Dai, Wei Xu 0007, Ying Wu 0001, Yihong Gong, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 6 |
| 2009 | A Dynamic Hierarchical Clustering Method for Trajectory-Based Unusual Video Event DetectionabstractThe proposed unusual video event detection method is based on unsupervised clustering of object trajectories, which are modeled by hidden Markov models (HMM). The novelty of the method includes a dynamic hierarchical process incorporated in the trajectory clustering algorithm to prevent model overfitting and a 2-depth greedy search strategy for efficient clustering. Ying Wu 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2009 | Optimized Bit Extraction Using Distortion Modeling in the Scalable Extension of H.264/AVCabstractThe newly adopted scalable extension of H.264/AVC video coding standard (SVC) demonstrates significant improvements in coding efficiency in addition to an increased degree of supported scalability relative to the scalable profiles of prior video coding standards. Due to the complicated hierarchical prediction structure of the SVC and the concept of key pictures, content-aware rate adaptation of SVC bit streams to intermediate bit rates is a nontrivial task. The concept of quality layers has been introduced in the design of the SVC to allow for fast content-aware prioritized rate adaptation. However, existing quality layer assignment methods are suboptimal and do not consider all network abstraction layer (NAL) units from different layers for the optimization. In this paper, we first propose a technique to accurately and efficiently estimate the quality degradation resulting from discarding an arbitrary number of NAL units from multiple layers of a bitstream by properly taking drift into account. Then, we utilize this distortion estimation technique to assign quality layers to NAL units for a more efficient extraction. Experimental results show that a significant gain can be achieved by the proposed scheme. Ehsan Maani, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2009 | Special Issue on Quality-Driven Cross-Layer Design for Multimedia CommunicationsabstractThe 14 papers in this special issue are divided into four categories: theoretical frameworks and schemes; integration of media processing and network protocols; heterogeneity; and quality evaluation. Aggelos K. Katsaggelos, Song Ci, Haohong Wang, Qian Zhang 0001, Antonios Argyriou |
IEEE Trans. Multim. | 1 |
| 2009 | Decision-aided compensation of severe phase-impairment-induced inter-carrier interference in frequency-selective OFDMabstractA new, reduced complexity algorithm is proposed for compensating the Inter-Carrier Interference (ICI) caused by severe PHase Noise (PHN) and Residual Frequency Offset (RFO) in OFDM systems. The algorithm estimates and compensates the most significant terms of the frequency domain ICI process, which are optimally selected via a Minimum Mean Squared Error (MMSE) criterion. The algorithm requires minimal knowledge of the phase process statistics, the estimation of which is also considered. The scheme outperforms previously proposed compensation methods of similar complexity, when severe phase impairments are present. Konstantinos Nikitopoulos, Stelios Stefanatos, Aggelos K. Katsaggelos |
IEEE Trans. Wirel. Commun. | 3 |
| 2008 | Quality-Driven Optimization for Content-Aware Real-Time Video Streaming in Wireless Mesh NetworksabstractVideo transport over multi-hop wireless networks has received significant research interests recently. The majority of the research efforts in this field have been conducted taking the approach of cross-layer optimization. However, video content and user perceived quality have been largely ignored in existing work. In this paper, we integrate video content analysis into video transport over wireless mesh networks (WMN). A content-aware quality-driven cross-layer optimization framework is proposed to achieve the best end-to-end user perceived video quality. In our framework, the extracted video regions of interest (ROI) are discriminatingly coded, transmitted and protected in video encoding, network routing and packet scheduling by different network layers. We aim at the optimization of key parameters of each layer while focusing on their interactions across the holistic network protocol stack. The proposed framework is evaluated by H.264/AVC codec and WMN simulations. Experimental results demonstrate that the proposed framework can effectively provide a good user perceived video quality, especially when the delay requirement is stringent. Dalei Wu, Haiyan Luo, Song Ci, Haohong Wang, Aggelos K. Katsaggelos |
GLOBECOM | 5 |
| 2008 | Generalized Gaussian Markov random field image restoration using variational distribution approximationabstractIn this paper we propose novel algorithms for image restoration and parameter estimation with a Generalized Gaussian Markov Random Field (GGMRF) prior utilizing variational distribution approximation. The restored image and the unknown hyperparameters for both the image prior and the image degradation noise are simultaneously estimated within a hierarchical Bayesian framework. We develop two algorithms resulting from this formulation which provide approximations to the posterior distributions of the latent variables. Experimental results are provided to demonstrate the performance of the algorithms. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2008 | Abnormal event detection based on trajectory clustering by 2-depth greedy searchabstractClustering-based approaches for abnormal video event detection have been proven to be effective in the recent literature. Based on the framework proposed in our previous work [1], we have developed in this paper a new strategy for unsupervised trajectory clustering. More specifically, an information-based trajectory dissimilarity measure is proposed, based on the Bayesian information criterion (BIC). In order to minimize BIC, the agglomerative hierarchical clustering is applied using a 2-depth greedy search process. This strategy achieves better clustering results compared to the traditional 1-depth greedy search. The increased computational complexity is addressed with several bounds on the trajectory dissimilarity. Ying Wu 0001, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2008 | A Kd-Tree Based Dynamic Indexing Scheme for Video Retrieval and Geometry MatchingabstractEfficient indexing is a key in content-based video retrieval solutions. In this paper we propose a new dynamic indexing scheme based on the kd-tree structure. Video sequences are first represented as traces in an appropriate low dimensional space via luminance field scaling and PCA projection. Then, the indexing scheme is applied to give the video database a manageable structure. Being able to handle dynamic video clip insertions and deletions is an essential part of this solution. At the beginning, an ordinary kd-tree is created for the initial database. As new video traces are added to the database, they will be added to the indexing tree structure as well. A tree node will be split if its size exceeds a certain threshold. If the tree structure un-balance level exceeds a threshold, merging and re-splitting will be performed. Preliminary experiments showed that merging and re-splitting will ensure the efficiency of the indexing scheme. Zhu Li 0001, Aggelos K. Katsaggelos |
ICCCN | 3 |
| 2008 | Total variation super resolution using a variational approachabstractIn this paper we propose a novel algorithm for super resolution based on total variation prior and variational distribution approximations. We formulate the problem using a hierarchical Bayesian model where the reconstructed high resolution image and the model parameters are estimated simultaneously from the low resolution observations. The algorithm resulting from this formulation utilizes variational inference and provides approximations to the posterior distributions of the latent variables. Due to the simultaneous parameter estimation, the algorithm is fully automated so parameter tuning is not required. Experimental results show that the proposed approach outperforms some of the state-of-the-art super resolution algorithms. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2008 | Combination of MR surface coil images using weighted constrained least squaresabstractFor most MR imaging applications multiple surface coils are used to obtain images with high signal-to-noise ratios (SNR). However, signal intensity strongly diminishes with distance. Although there are a number of approaches to combine the surface coil images to obtain a high SNR and bias-free image, most of them are developed in an ad hoc manner and lack a systematic treatment. In this work we propose a new approach, an iterative weighted constrained least squares (WCLS) restoration method, for combining surface coil images. The algorithm is fully automated and outperforms approaches which appeared in the literature. S. Derin Babacan, Xiaoming Yin, Andrew C. Larson, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2008 | Local feature extraction for video copy detection in a databaseabstractIn this paper a new content-based copy identification method for video sequences is presented. It is robust to a number of image transformations and particulary robust to compression artifacts. A scale and rotation invariant local image descriptor for corner points in detected key frames is proposed based on a generalized Radon transform. In addition, a distance similarity metric is used that fuses intensity and geometry information to compare key frames extracted using a scene detection algorithm. Furthermore, to achieve low querying computational complexity a DP approach is employed. Experimental results demonstrate the effectiveness of our approach. Ehsan Maani, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2008 | Feature space video stream consistency estimation for dynamic stream weighting in audio-visual speech recognitionabstractMost current audio-visual automatic speech recognition (AV- ASR) systems use static weights to leverage between audio and visual information during information fusion. State of the art research has led to using audio reliability metrics for dynamically changing the fusion weights in order to successfully improve overall recognition results. So far, however, incorporating visual reliability metrics into these audio reliability metric based systems have not significantly improved performance. We introduce a new approach to this problem by inferring the "consistency" between the audio and visual information and leveraging the existing audio reliability metrics to create a video reliability metric. Our approach is formulated in the extracted feature space and, thus, does not rely on analyzing the actual video signal itself. The framework presented in this work competes with the audio-only reliability metric based systems and shows promise to consistently outperform. Louis H. Terry, Derek J. Shiell, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2008 | Vector quantization with memory and multi-labeling for isolated video-only automatic speech recognitionabstractWe describe a vector quantizer (VQ) with memory for automatic speech recognition (ASR) and compare the recognition performance results to those obtained with traditional memoryless VQ for ASR. Standard VQ for ASR quantizes the speech data independently of any past information. We introduce memory in a probabilistic framework for quantization state modeling. This is accomplished in the form of an ergodic hidden Markov model (HMM) in which the state occupied by the HMM represents the quantization label. We evaluate this approach in the context of video-only isolated digit ASR and implement both single stream (single labeling) and multi-stream (multi-labeling) systems. For single stream recognition, our approach increases the recognition rate from 62.67% to 66.95%. When using multi-labeling, our proposed vector quantizer with memory consistently outperforms the memoryless vector quantizer. Louis H. Terry, Derek J. Shiell, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2008 | A dynamic programming solution to tracking and elastically matching left ventricular walls in cardiac cine MRIabstractIn this paper an algorithm to detect and elastically match the contours of the epicardial walls of the left ventricle (LV) in cardiac phase-resolved 2-D magnetic resonance (MR) images is presented. For both tasks, dynamic programming (DP) is used. A mask conforming to the six segment model of the LV is fitted on a reference image and propagated utilizing the elastic matching information. At its present form the algorithm requires minimal parameter corrections among different sets of cine MRI images. Future extensions include comparisons with contours hand labeled by imaging experts. Sotirios A. Tsaftaris, Valentin Andermatt, Andre Schlegel, Aggelos K. Katsaggelos, Debiao Li, Rohan Dharmakumar |
ICIP | 4 |
| 2008 | Automated line flattening of Atomic Force Microscopy imagesabstractIn this paper, an automated algorithm to flatten lines from Atomic Force Microscopy (AFM) images is presented. Due to the mechanics of the AFM, there is a curvature distortion (bowing effect) present in the acquired images. At present, flattening such images requires human intervention to manually segment object data from the background, which is time consuming and highly inaccurate. The proposed method classifies the data into objects and background, and fits convex lines in an iterative fashion. Results on real images from DNA wrapped carbon nanotubes (DNA-CNTs) and synthetic experiments are presented, demonstrating the effectiveness of the proposed algorithm in increasing the resolution of the surface topography. Sotirios A. Tsaftaris, Jana Zujovic, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2008 | A phone-viseme dynamic Bayesian network for audio-visual automatic speech recognitionabstractThis work extends and improves a recently introduced (Dec. 2007) dynamic Bayesian network (DBN) based audio-visual automatic speech recognition (AV-ASR) system. That system models the audio and visual components of speech as being composed of the same sub-word units when, in fact, this is not psycholinguistically true. We extend the system to model the audio and visual streams as being composed of separate, yet related, sub-word units. We also introduce a novel stream weighting structure incorporated into the model itself. In doing so, our system makes improvements in word error rate (WER) and overall recognition accuracy in a large vocabulary continuous speech recognition task (LVCSR). The ldquobestrdquo performing proposed system attains a WER of 66.71%whereas the ldquobestrdquo baseline system performs at a WER of 64.30%. The proposed system also improves accuracy to 45.95% from 39.40%. Louis H. Terry, Aggelos K. Katsaggelos |
ICPR | 2 |
| 2008 | Optimized Bit Extraction Using Distortion Estimation in the Scalable Extension of H.264/AVCabstractThe newly adopted scalable extension of H.264/AVC video coding standard (SVC), demonstrates significant improvements in coding efficiency in addition to an increased degree of supported scalability relative to the scalable profiles of prior video coding standards. For efficient adaptation of SVC bit streams to intermediate bit rates, the concept of quality layers has been introduced in the design of the SVC. The concept of quality layers allow a rate distortion (RD) optimal bit extraction; However, existing Quality Layer assignment methods do not consider all network abstraction layer (NAL) units from different layers for the optimization. In this paper, we first propose a technique to accurately and efficiently estimate the quality degradation resulting from discarding an arbitrary number of NAL units from multiple layers of a bitstream. Then, we utilize this distortion estimation technique to assign quality layers to NAL units for a more efficient extraction. Experimental results show that a significant gain can be achieved by the proposed schemes. Ehsan Maani, Aggelos K. Katsaggelos |
ISM | 2 |
| 2008 | Super Resolution of Multispectral Images Using TV Image Models
Miguel Vega, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
KES (3) | 4 |
| 2008 | Locally adaptive subspace and similarity metric learning for visual data clustering and retrieval
Yun Fu 0001, Zhu Li 0001, Thomas S. Huang, Aggelos K. Katsaggelos |
Comput. Vis. Image Underst. | 4 |
| 2008 | Source fidelity over fading channels: performance of erasure and scalable codesabstractWe consider the transmission of a Gaussian source through a block fading channel. Assuming each block is decoded independently, the received distortion depends on the tradeoff between quantization accuracy and probability of outage. Namely, higher quantization accuracy requires a higher channel code rate, which increases the probability of outage. We first treat an outage as an erasure, and evaluate the received mean distortion with erasure coding across blocks as a function of the code length. We then evaluate the performance of scalable, or multi-resolution coding in which coded layers are superimposed within a coherence block, and the layers are sequentially decoded. Both the rate and power allocated to each layer are optimized. In addition to analyzing the performance with a finite number of layers, we evaluate the mean distortion at high signal-to-noise ratios as the number of layers becomes infinite. As the block length of the erasure code increases to infinity, the received distortion converges to a deterministic limit, which is less than the mean distortion with an infinite-layer scalable coding scheme. However, for the same standard deviation in received distortion, infinite layer scalable coding performs slightly better than erasure coding, and with much less decoding delay. Konstantinos E. Zachariadis, Michael L. Honig, Aggelos K. Katsaggelos |
IEEE Trans. Commun. | 3 |
| 2008 | Joint Source Adaptation and Resource Allocation for Multi-User Wireless Video StreamingabstractMulti-user video streaming over wireless channels is a challenging problem, where the demand for better video quality and small transmission delays needs to be reconciled with the limited and often time-varying communication resources. This paper presents a framework for joint network optimization, source adaptation, and deadline-driven scheduling for multi-user video streaming over wireless networks. We develop a joint adaptation, resource allocation and scheduling (JARS) algorithm, which allocates the communication resource based on the video users' quality of service, adapts video sources based on smart summarization, and schedules the transmissions to meet the frame delivery deadlines. The proposed algorithm leads to near full utilization of the network resources and satisfies the delivery deadlines for all video frames. Substantial performance improvements are achieved compared with heuristic schemes that do not take the interactions between multiple users into consideration. Jianwei Huang 0001, Zhu Li 0001, Mung Chiang, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Parameter Estimation in TV Image Restoration Using Variational Distribution ApproximationabstractIn this paper, we propose novel algorithms for total variation (TV) based image restoration and parameter estimation utilizing variational distribution approximations. Within the hierarchical Bayesian formulation, the reconstructed image and the unknown hyper parameters for the image prior and the noise are simultaneously estimated. The proposed algorithms provide approximations to the posterior distributions of the latent variables using variational methods. We show that some of the current approaches to TV-based image restoration are special cases of our framework. Experimental results show that the proposed approaches provide competitive performance without any assumptions about unknown hyper parameters and clearly outperform existing methods when additional information is included. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2008 | Resource Allocation for Downlink Multiuser Video Transmission Over Wireless Lossy NetworksabstractDemand for multimedia services, such as video streaming over wireless networks, has grown dramatically in recent years. The downlink transmission of multiple video sequences to multiple users over a shared resource-limited wireless channel, however, is a daunting task. Among the many challenges in this area are the time-varying channel conditions, limited available resources, such as bandwidth and power, and the different transmission requirements of different video content. This work takes into account the time-varying nature of the wireless channels, as well as the importance of individual video packets, to develop a cross-layer resource allocation and packet scheduling scheme for multiuser video streaming over lossy wireless packet access networks. Assuming that accurate channel feedback is not available at the scheduler, random channel losses combined with complex error concealment at the receiver make it impossible for the scheduler to determine the actual distortion of the sequence at the receiver. Therefore, the objective of the optimization is to minimize the expected distortion of the received sequence, where the expectation is calculated at the scheduler with respect to the packet loss probability in the channel. The expected distortion is used to order the packets in the transmission queue of each user, and then gradients of the expected distortion are used to efficiently allocate resources across users. Simulations show that the proposed scheme performs significantly better than a conventional content-independent scheme for video transmission. Ehsan Maani, Peshala V. Pahalawatta, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 5 |
| 2007 | Detector EnsembleabstractComponent-based detection methods have demonstrated their promise by integrating a set of part-detectors to deal with large appearance variations of the target. However, an essential and critical issue, i.e., how to handle the imperfectness of part-detectors in the integration, is not well addressed in the literature. This paper proposes a detector ensemble model that consists of a set of substructure-detectors, each of which is composed of several part-detectors. Two important issues are studied both in theory and in practice, (1) finding an optimal detector ensemble, and (2) detecting targets based on an ensemble. Based on some theoretical analysis, a new model selection strategy is proposed to learn an optimal detector ensemble that has a minimum number of false positives and satisfies the design requirement on the capacity of tolerating missing parts. In addition, this paper also links ensemble-based detection to the inference in Markov random field, and shows that the target detection can be done by a max-product belief propagation algorithm. Shengyang Dai, Ming Yang 0007, Ying Wu 0001, Aggelos K. Katsaggelos |
CVPR | 4 |
| 2007 | Optimal Mode Selection and Channel Coding for Video Transmission Over Wireless Channels using H.264/AVCabstractThis paper addresses the problem of joint encoder optimization and channel coding for realtime video transmission over wireless channels. An efficient solution is proposed to optimally select macroblock modes and quantizers as well as channel coding rates. The proposed optimization algorithm fully considers error resilience, forward error correction and error concealment. Experimental results demonstrate the effectiveness of the proposed approach. Ehsan Maani, Fan Zhai, Aggelos K. Katsaggelos |
ICASSP (1) | 3 |
| 2007 | Content-Aware Resource Allocation for Scalable Video Transmission to Multiple Users Over a Wireless NetworkabstractWireless video transmission is prone to unpredictable degradations due to time-varying channel conditions. Such degradations are difficult to overcome using conventional video coding techniques. Scalable video coding offers a flexible bitstream that can be dynamically adapted to fit the prevailing channel conditions. Within a scalable video coding framework, we develop simple packet prioritization strategies, which, when combined with a reasonable error concealment scheme and a content-aware resource allocation technique, provide for robust video transmission over time-varying channels. The packet prioritization as well as the calculation of the content-aware scheduling metric can be performed offline and signaled to the wireless scheduler. Peshala V. Pahalawatta, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICASSP (1) | 4 |
| 2007 | Packet Scheduling for Scalable Video Streaming Over Lossy Packet Access NetworksabstractVideo streaming applications have gained in popularity in recent years. The quality of service offered by such applications is limited by the available transmission rates as well as time-varying conditions, such as, channel fading and network congestion, which lead to packet losses. Scalable video coding techniques that allow for the flexible adaptation of temporal resolution as well as quality of an encoded bitstream can be immensely useful in developing video streaming applications that can adapt to time-varying network and channel conditions. Scalable coding techniques, however, are generally designed to offer progressive refinement, which introduces dependencies between encoded video packets. Therefore, when determining a packet scheduling technique for scalable coded video, the possibility of random packet losses, which might affect the decodability of subsequent packets, must be taken into account. In this paper, we take into account the available transmission rate, possibly time-varying channel conditions, and the possibility of random packet losses, to design a scheduling technique for video packets in a scalable bit-stream. Since the optimal solution to the scheduling problem requires an exhaustive, and therefore, intractable computation, we propose a greedy algorithm that will schedule the optimal packet for transmission at a given transmission opportunity based on the encoded content and the available channel state information. Simulation results show significant gains in performance when the proposed technique is compared to content and channel independent packet scheduling techniques. Ehsan Maani, Yijing Luo, Peshala V. Pahalawatta, Aggelos K. Katsaggelos |
ICCCN | 4 |
| 2007 | Total Variation Image Restoration and Parameter Estimation using Variational Posterior Distribution ApproximationabstractIn this paper we propose novel algorithms for total variation (TV) based image restoration and parameter estimation utilizing variational distribution approximations. By following the hierarchical Bayesian framework, we simultaneously estimate the reconstructed image and the unknown hyper parameters for both the image prior and the image degradation noise. Our algorithms provide an approximation to the posterior distributions of the unknowns so that both the uncertainty of the estimates can be measured and different values from these distributions can be used for the estimates. We also show that some of the current approaches to TV-based image restoration are special cases of our variational framework. Experimental results show that the proposed approaches provide competitive performance without any assumptions about unknown hyper parameters and clearly outperform existing methods when additional information is included. S. Derin Babacan, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 2007 | Abnormal Event Detection from Surveillance Video by Dynamic Hierarchical ClusteringabstractThe clustering-based approach for detecting abnormalities in surveillance video requires the appropriate definition of similarity between events. The HMM-based similarity defined previously falls short in handling the overfitting problem. We propose in this paper a multi-sample-based similarity measure, where HMM training and distance measuring are based on multiple samples. These multiple training data are acquired by a novel dynamic hierarchical clustering (DHC) method. By iteratively reclassifying and retraining the data groups at different clustering levels, the initial training and clustering errors due to overfitting will be sequentially corrected in later steps. Experimental results on real surveillance video show an improvement of the proposed method over a baseline method that uses single-sample-based similarity measure and spectral clustering. Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP (5) | 3 |
| 2007 | Resource Allocation for Downlink Multiuser Video Transmission Over Wireless Lossy NetworksabstractThe emergence of 3G and 4G wireless networks brings with it the possibility of streaming high quality video content on-demand to mobile users. Wireless video applications require appropriate scheduling techniques that make use of the specific characteristics of video content, as well as the well known gains from multiuser diversity. While fast and frequent channel feedback is available in the new generation of wireless networks, the channel estimates cannot be perfect, and channel losses should be taken into account in the packet scheduling and resource allocation. The proposed scheme is formulated as a joint optimization over the resource allocation and channel loss protection, in order to minimize the distortion of the received video sequences. The distortion is a function of the packets deliberately dropped at the transmission queue due to congestion, as well as of random channel losses. The scheme makes use of a packet prioritization strategy that orders video packets based on their contribution to reducing the expected distortion of the received video sequence. Simulation results show that the proposed technique significantly outperforms content-independent packet scheduling schemes. Ehsan Maani, Peshala V. Pahalawatta, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP (5) | 5 |
| 2007 | From Global to Local Bayesian Parameter Estimation in Image Restoration using Variational Distribution ApproximationsabstractIn this paper we present a new Bayesian methodology for the restoration of blurred and noisy images. Bayesian methods rely on image priors that encapsulate prior image knowledge and avoid the ill-posedness of the image restoration problems. Some of these priors depend on global variance parameters, unable to account for local characteristics. Here we first use variational methods to approximate probability posterior distributions for the global model to later use those distributions to define local and more realistic image models which lead to better restored images as it is shown in the experimental section. Rafael Molina 0001, Miguel Vega, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 2007 | DNA Microarray Image Intensity Extraction using EigenspotsabstractDNA microarrays are commonly used in the rapid analysis of gene expression in organisms. Image analysis is used to measure the average intensity of circular image areas (spots), which correspond to the level of expression of the genes. A crucial aspect of image analysis is the estimation of the background noise. Currently, background subtraction algorithms are used to estimate the local background noise and subtract it from the signal. In this paper we use principal component analysis (PCA) to de-correlate the signal from the noise, by projecting each spot on the space of eigenvectors, which we term eigenspots. PCA is well suited for such application due to the structural nature of the images. To compare the proposed method with other background estimation methods we use the industry standard signal-to-noise metric xdev. Sotirios A. Tsaftaris, Ramandeep Ahuja, Derek J. Shiell, Aggelos K. Katsaggelos |
ICIP (6) | 4 |
| 2007 | Joint Source Coding and Data Rate Adaptation for Multi-User Wireless Video TransmissionabstractMuch attention has been paid to the problem of optimally utilizing resources such as spectrum, power and time in order to achieve the best video delivery quality in wireless communications system, due to the fueling demand for such applications. In this work, we present a joint source coding and data adaptation scheme for downlink video transmission in a multi-user wireless network. We formulate a rate-distortion optimization problem, where the source coding and data rate are jointly designed according to the changing channel conditions. In addition, transmissions of video packets are optimally scheduled through exploiting the multi-user diversity. We solve the problem using a backward stochastic dynamical programming approach, and the simulation results have shown the advantage of the joint selection of source coding parameter and transmission rate coupled with optimal packet scheduling. Fan Zhai, Zhu Li 0001, Aggelos K. Katsaggelos |
ICME | 3 |
| 2007 | Content-Aware Resource Allocation and Packet Scheduling for Video Transmission over Wireless NetworksabstractA cross-layer packet scheduling scheme that streams pre-encoded video over wireless downlink packet access networks to multiple users is presented. The scheme can be used with the emerging wireless standards such as HSDPA and IEEE 802.16. A gradient based scheduling scheme is used in which user data rates are dynamically adjusted based on channel quality as well as the gradients of a utility function. The user utilities are designed as a function of the distortion of the received video. This enables distortion-aware packet scheduling both within and across multiple users. The utility takes into account decoder error concealment, an important component in deciding the received quality of the video. We consider both simple and complex error concealment techniques. Simulation results show that the gradient based scheduling framework combined with the content-aware utility functions provides a viable method for downlink packet scheduling as it can significantly outperform current content-independent techniques. Further tests determine the sensitivity of the system to the initial video encoding schemes, as well as to non-real-time packet ordering techniques. Peshala V. Pahalawatta, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
IEEE J. Sel. Areas Commun. | 4 |
| 2007 | Review of content-aware resource allocation schemes for video streaming over wireless networksabstractAbstract As wireless technology evolves towards its fourth generation (4G) of development, the prospect of offering multimedia services such as on‐demand video streaming and video conferencing to wireless mobile clients becomes increasingly more viable. The eventual success of such applications depends on the efficient management of the limited system resources while taking into account the time‐varying wireless channel conditions as well as the varying multimedia source content. In this paper, we review some of the recent advances in cross‐layer design schemes, which aim at providing significant gains in performance for video streaming systems through content‐aware resource allocation. Advances in both, real‐time video streaming, where the video is encoded and transmitted in real‐time, as well as, on‐demand video streaming, where the video is pre‐encoded in a media server, are considered. Copyright © 2007 John Wiley & Sons, Ltd. Peshala V. Pahalawatta, Aggelos K. Katsaggelos |
Wirel. Commun. Mob. Comput. | 2 |
| 2006 | Joint source-channel coding and power allocation for video transmission over wireless fading channelsabstractTechniques for modeling and simulating channel conditions play an essential role in efficient real-time video transmission over wireless networks. In this paper, we consider a Finite-State Markov Chain (FSMC) as the channel model to perform Joint Source Channel Coding with power allocation (JSCCPA). The optimization is done in an integrated manner and in one step. Computer simulations are performed to show the advantages of the proposed model. Ehsan Maani, Aggelos K. Katsaggelos |
CCNC | 2 |
| 2006 | Pricing Based Collaborative Multi-User Video Streamming Over Power Constrained Wireless DownlinkabstractVideo streaming is becoming an important application in wireless communications. In a typical scenario, a base station needs to serve multiple video users with a total transmitting power constraint. How to make appropriate video coding decisions and allocate limited transmitting power among users to achieve optimal total utility is an important problem. In this paper we develop a pricing based downlink power allocation scheme with collaborative video summarization among users. The scheme exploits the multiuser diversity in channel states and utility-resource tradeoff characteristics in video contents to achieve better resource utilization. The computational burden can be distributed among video sources and base station. Simulation results demonstrate the effectiveness of the proposed algorithm. Zhu Li 0001, Jianwei Huang 0001, Aggelos K. Katsaggelos |
ICASSP (5) | 3 |
| 2006 | DNA Hybridization as a Similarity Criterion for Querying Digital Signals Stored in DNA DatabasesabstractWe demonstrate via simulation that hybridization of DNA molecules can be used as a similarity criterion for retrieving digital signals encoded and stored in a synthesized DNA database. After introducing some necessary DNA terminology, we briefly explain how digital signals are transformed to DNA sequences. Since retrieval is achieved through hybridization of query and data carrying DNA molecules, we present a mathematical model to estimate hybridization efficiency (also known as selectivity annealing). We show that selectivity annealing is inversely proportional to the mean squared error (MSE) of the encoded signal values. In addition, we show that the concentration of the molecules plays the same role as the decision threshold employed in digital signal matching algorithms. Finally, similar to the digital domain, we define a DNA signal-to-noise ratio (SNR) measure to assess the performance of the DNA-based retrieval scheme. Simulations are presented to validate our arguments Sotirios A. Tsaftaris, Vassily Hatzimanikatis, Aggelos K. Katsaggelos |
ICASSP (2) | 3 |
| 2006 | Tracking Motion-Blurred Targets in VideoabstractMany emerging applications require tracking targets in video. Most existing visual tracking methods do not work well when the target is motion-blurred (especially due to fast motion), because the imperfectness of the target's appearances invalidates the image matching model (or the measurement model) in tracking. This paper presents a novel method to track motion-blurred targets by taking advantage of the blurs without performing image restoration. Unlike the global blur induced by camera motion, this paper is concerned with the local blurs that are due to target's motion. This is a challenging task because the blurs need to be identified blindly. The proposed method addresses this difficulty by integrating signal processing and statistical learning techniques. The estimated blurs are used to reduce the search range by providing strong motion predictions and to localize the best match accurately by modifying the measurement models. Shengyang Dai, Ming Yang 0007, Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2006 | Fast Video Shot Retrieval with Luminance Field Trace Indexing and Geometry MatchingabstractEfficient indexing is a key in content-based video retrieval solutions. In this paper we represent video sequences as traces in an appropriate low dimensional space via luminance field scaling and PCA projection, and develop an efficient indexing scheme and trace geometry matching algorithms for fast and robust video retrieval. Simulation results demonstrate that the proposed solution is very fast and accurate in retrieval performance. Zhu Li 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2006 | New Results on Efficient Optimal Multilevel Image ThresholdingabstractImage thresholding is one of the most common image processing operations, since almost all image processing schemes need some sort of separation of the pixels into different classes. In order to find the thresholds, almost all methods analyze the histogram of the image. In most cases, the optimal thresholds are found by either minimizing or maximizing an objective function, which depends on the positions of the thresholds. We identify two classes of objective functions for which the optimal thresholds can be found by algorithms with low time complexity. We show, that for example the method proposed by Otsu (1979) and other well known methods have objective functions belonging to these classes. By implementing the algorithms in ANSI C and comparing their execution times, we can make a quantitative statement about their performance. Martin Luessi, Marco Eichmann, Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2006 | Parameter Estimation in Bayesian Reconstruction of Multispectral Images using Super Resolution TechniquesabstractIn this paper we present a new super resolution Bayesian method for pansharpening of multispectral images which: a) incorporates prior knowledge on the expected characteristics of the multispectral images, b) uses the sensor characteristics to model the observation process of both panchromatic and multispectral images, and c) performs the estimation of all the unknown parameters in the model. Using real data, the pansharpened multispectral images are compared with the images obtained by other pansharpening methods and their quality is assessed both qualitatively and quantitatively. Rafael Molina 0001, Miguel Vega, Javier Mateos, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2006 | Resampling for Spatial ScalabilityabstractResampling is a fundamental issue in the design of a spatially scalable video codec. The resampling procedure is responsible for down-sampling the high-resolution video sequence to generate lower resolution data, as well as upsampling the transmitted lower resolution data to predict the original high-resolution frames. In both cases, the resampling operation must make trade-offs between coding efficiency, image quality and computational complexity. In this paper, we consider the resampling design problem within an optimization framework. C. Andrew Segall, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2006 | Locally Embedded Linear Subspaces for Efficient Video Indexing and RetrievalabstractEfficient Indexing is a key us content-based video retrieval solutions. In this paper we represent video sequences as traces via scaling and linear transformation of the frame luminance field. Then an appropriate lower dimensional subspace is identified for video trace indexing. We also develop a trace geometry matching algorithm for retrieval based on average projection distance with a locally embedded distance metric. Simulation results demonstrated the high accuracy and very fast retrieval speed for the proposed solution Zhu Li 0001, Aggelos K. Katsaggelos |
ICME | 3 |
| 2006 | Audio-Visual BiometricsabstractBiometric characteristics can be utilized in order to enable reliable and robust-to-impostor-attacks person recognition. Speaker recognition technology is commonly utilized in various systems enabling natural human computer interaction. The majority of the speaker recognition systems rely only on acoustic information, ignoring the visual modality. However, visual information conveys correlated and complimentary information to the audio information and its integration into a recognition system can potentially increase the system's performance, especially in the presence of adverse acoustic conditions. Acoustic and visual biometric signals, such as the person's voice and face, can be obtained using unobtrusive and user-friendly procedures and low-cost sensors. Developing unobtrusive biometric systems makes biometric technology more socially acceptable and accelerates its integration into every day life. In this paper, we describe the main components of audio-visual biometric systems, review existing systems and their performance, and discuss future research and development directions in this area Petar S. Aleksic, Aggelos K. Katsaggelos |
Proc. IEEE | 2 |
| 2006 | Automatic facial expression recognition using facial animation parameters and multistream HMMsabstractThe performance of an automatic facial expression recognition system can be significantly improved by modeling the reliability of different streams of facial expression information utilizing multistream hidden Markov models (HMMs). In this paper, we present an automatic multistream HMM facial expression recognition system and analyze its performance. The proposed system utilizes facial animation parameters (FAPs), supported by the MPEG-4 standard, as features for facial expression classification. Specifically, the FAPs describing the movement of the outer-lip contours and eyebrows are used as observations. Experiments are first performed employing single-stream HMMs under several different scenarios, utilizing outer-lip and eyebrow FAPs individually and jointly. A multistream HMM approach is proposed for introducing facial expression and FAP group dependent stream reliability weights. The stream weights are determined based on the facial expression recognition results obtained when FAP streams are utilized individually. The proposed multistream HMM facial expression system, which utilizes stream reliability weights, achieves relative reduction of the facial expression recognition error of 44% compared to the single-stream HMM system. Petar S. Aleksic, Aggelos K. Katsaggelos |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2006 | VAPOR: variance-aware per-pixel optimal resource allocationabstractCharacterizing the video quality seen by an end-user is a critical component of any video transmission system. In packet-based communication systems, such as wireless channels or the Internet, packet delivery is not guaranteed. Therefore, from the point-of-view of the transmitter, the distortion at the receiver is a random variable. Traditional approaches have primarily focused on minimizing the expected value of the end-to-end distortion. This paper explores the benefits of accounting for not only the mean, but also the variance of the end-to-end distortion when allocating limited source and channel resources. By accounting for the variance of the distortion, the proposed approach increases the reliability of the system by making it more likely that what the end-user sees, closely resembles the mean end-to-end distortion calculated at the transmitter. Experimental results demonstrate that variance-aware resource allocation can help limit error propagation and is more robust to channel-mismatch than approaches whose goal is to strictly minimize the expected distortion. Yiftach Eisenberg, Fan Zhai, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 5 |
| 2006 | Blind Deconvolution Using a Variational Approach to Parameter, Image, and Blur EstimationabstractFollowing the hierarchical Bayesian framework for blind deconvolution problems, in this paper, we propose the use of simultaneous autoregressions as prior distributions for both the image and blur, and gamma distributions for the unknown parameters (hyperparameters) of the priors and the image formation noise. We show how the gamma distributions on the unknown hyperparameters can be used to prevent the proposed blind deconvolution method from converging to undesirable image and blur estimates and also how these distributions can be inferred in realistic situations. We apply variational methods to approximate the posterior probability of the unknown image, blur, and hyperparameters and propose two different approximations of the posterior distribution. One of these approximations coincides with a classical blind deconvolution method. The proposed algorithms are tested experimentally and compared with existing blind deconvolution methods. Rafael Molina 0001, Javier Mateos, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2006 | Motion compensated shape error concealmentabstractThe introduction of Video Objects (VOs) is one of the innovations of MPEG-4. The alpha-plane of a VO defines its shape at a given instance in time and hence determines the boundary of its texture. In packet-based networks, shape, motion, and texture are subject to loss. While there has been considerable attention paid to the concealment of texture and motion errors, little has been done in the field of shape error concealment. In this paper we propose a post-processing shape error concealment technique that uses the motion compensated boundary information of the previously received alpha-plane. The proposed approach is based on matching received boundary segments in the current frame to the boundary in the previous frame. This matching is achieved by finding a maximally smooth motion vector field. After the current boundary segments are matched to the previous boundary, the missing boundary pieces are reconstructed by motion compensation. Experimental results demonstrating the performance of the proposed motion compensated shape error concealment method, and comparing it with the previously proposed weighted side matching method are presented. Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2006 | Joint source-channel coding for wireless object-based video communications utilizing data hidingabstractIn recent years, joint source-channel coding for multimedia communications has gained increased popularity. However, very limited work has been conducted to address the problem of joint source-channel coding for object-based video. In this paper, we propose a data hiding scheme that improves the error resilience of object-based video by adaptively embedding the shape and motion information into the texture data. Within a rate-distortion theoretical framework, the source coding, channel coding, data embedding, and decoder error concealment are jointly optimized based on knowledge of the transmission channel conditions. Our goal is to achieve the best video quality as expressed by the minimum total expected distortion. The optimization problem is solved using Lagrangian relaxation and dynamic programming. The performance of the proposed scheme is tested using simulations of a Rayleigh-fading wireless channel, and the algorithm is implemented based on the MPEG-4 verification model. Experimental results indicate that the proposed hybrid source-channel coding scheme significantly outperforms methods without data hiding or unequal error protection. Haohong Wang, Sotirios A. Tsaftaris, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2006 | Stochastic methods for joint registration, restoration, and interpolation of multiple undersampled imagesabstractUsing a stochastic framework, we propose two algorithms for the problem of obtaining a single high-resolution image from multiple noisy, blurred, and undersampled images. The first is based on a Bayesian formulation that is implemented via the expectation maximization algorithm. The second is based on a maximum a posteriori formulation. In both of our formulations, the registration, noise, and image statistics are treated as unknown parameters. These unknown parameters and the high-resolution image are estimated jointly based on the available observations. We present an efficient implementation of these algorithms in the frequency domain that allows their application to large images. Simulations are presented that test and compare the proposed algorithms. Nathan A. Woods, Nikolas P. Galatsanos, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2006 | Rate-distortion optimized hybrid error control for real-time packetized video transmissionabstractThe problem of application-layer error control for real-time video transmission over packet lossy networks is commonly addressed via joint source-channel coding (JSCC), where source coding and forward error correction (FEC) are jointly designed to compensate for packet losses. In this paper, we consider hybrid application-layer error correction consisting of FEC and retransmissions. The study is carried out in an integrated joint source-channel coding (IJSCC) framework, where error resilient source coding, channel coding, and error concealment are jointly considered in order to achieve the best video delivery quality. We first show the advantage of the proposed IJSCC framework as compared to a sequential JSCC approach, where error resilient source coding and channel coding are not fully integrated. In the USCC framework, we also study the performance of different error control scenarios, such as pure FEC, pure retransmission, and their combination. Pure FEC and application layer retransmissions are shown to each achieve optimal results depending on the packet loss rates and the round-trip time. A hybrid of FEC and retransmissions is shown to outperform each component individually due to its greater flexibility. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 5 |
| 2005 | Source fidelity over fading channels: erasure codes versus scalable codesabstractWe consider the transmission of a Gaussian source through a block fading channel. Assuming each block is decoded independently, the received distortion depends on the tradeoff between quantization accuracy and probability of outage. Namely, higher quantization accuracy requires a higher channel code rate, which increases the probability of outage. Here we evaluate the received mean distortion with erasure coding across blocks as a function of the code length. We also evaluate the performance of scalable, or multi-resolution coding in which coded layers are superimposed, and the layers are sequentially decoded. In addition to analyzing a finite number of layers, we evaluate the mean distortion at high signal-to-noise ratios as the number of layers becomes infinite. As the block length of the erasure code increases to infinity, the received distortion converges to a deterministic limit, which is less than the mean distortion with an infinite-layer scalable coding scheme. However, for the same standard deviation in received distortion, infinite layer scalable coding performs slightly better than erasure coding. Konstantinos E. Zachariadis, Michael L. Honig, Aggelos K. Katsaggelos |
GLOBECOM | 3 |
| 2005 | Comparison of MPEG-4 facial animation parameter groups with respect to audio-visual speech recognition performanceabstractIn this paper, we describe an audio-visual automatic speech recognition (AV-ASR) system that utilizes facial animation parameters (FAPs), supported by the MPEG-4 standard, for the visual representation of speech. We describe the visual feature extraction algorithms used for extracting FAPs, which control outer- and inner-lip movement. Principal component analysis (PCA) is performed on both inner- and outer-lip FAP vector in order to decrease their dimensionality and decorrelate them. The PCA-based projection weights of the extracted FAP vectors are used as visual features. Multi-stream hidden Markov models (HMMs) and a late integration approach are used to integrate audio and visual information and train a continuous AV-ASR system. We compare the performance of the developed AV-ASR system utilizing outer- and inner lip FAPs, individually and jointly. Experiments were performed for different dimensionalities of the visual features, at various SNRs (0-30dB) with additive white Gaussian noise, on a relatively large vocabulary (approximately 1000 words) database. The proposed system reduces the word error rate (WER) by 20% to 23% relatively to audio-only ASR WERs. Conclusions are drawn on the individual and combined effectiveness of the inner- and outer-lip FAPs, the trade off between the dimensionality of the visual features and the amount of speechreading information contained in them and its influence on the AV-ASR performance. Petar S. Aleksic, Aggelos K. Katsaggelos |
ICIP (3) | 2 |
| 2005 | A cross-layer approach for energy efficient transmission of progressively coded images over wireless channelsabstractMobility, made available by today's communication networks, imposes several limitations in the design of multimedia applications, due to high error rates, reduced bandwidth, strong variability, and mobile terminals lifetime. This paper investigates the use of an energy constrained cross-layer approach for the transmission of progressively coded images over packet-based wireless channels. We propose an optimum power allocation algorithm to enable unequal error protection of a pre-encoded image. Simulations are performed modeling transmission over a Rayleigh fading channel. The investigation focuses on JPEG2000, but it is applicable to other progressively coded bitstreams as well. Our experimental results demonstrate that it is possible to achieve a relevant performance enhancement with the proposed approach over uniform error protection. The optimal solution can also serve as a guideline for developing less computationally intensive empiric approaches. Cristina Emilia Costa, Fabrizio Granelli, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 2005 | Video summarization for multiple path communicationabstractFor video communications over wireless ad hoc networks, multiple paths with limited bandwidth are common. It therefore presents new challenges to the video encoding. In this paper, we formulate the problem as a multiple path video summarization problem under bit rate constraints, where video summaries are generated to satisfy each channel's rate constraint, while the combined summary at the receiving end achieves the minimum summarization distortion. The optimal solution (within a convex hull approximation) is found by Lagrangian relaxation and dynamic programming. Simulation results demonstrate the effectiveness of the approach. Zhu Li 0001, Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 2005 | Approximations of posterior distributions in blind deconvolution using variational methodsabstractIn this paper the blind deconvolution problem is formulated using the variational framework. With its use approximations of the involved probability distributions are developed resulting in two algorithms for the estimation of the posterior distributions of the hyperparameters, the blur, and the original image. The performance of the two proposed restoration algorithms is demonstrated experimentally. Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 2005 | Video Summarization and Transmission Power Adaptation for Very Low Bit Rate Multiuser Wireless Uplink Video CommunicationabstractIn this paper we consider the problem of efficiently serving multiple uplink video users in a wideband CDMA wireless communication system with mixed voice and video traffic. Very low available video bit rate (<64 kpbs) and multi-user interference are the major limiting factors in the system performance. Our solution is based on video summarization to achieve reasonable video quality at very low bit rate, and multi-user transmission adaptation/scheduling via summarization distortion-rate profile negotiation between base station and mobiles to meet power control constraints and minimize the interference among voice/video users. The simulation results demonstrate the effectiveness of the proposed approach Zhu Li 0001, Alan Q. Cheng, Aggelos K. Katsaggelos, Faisal Ishtiaq |
MMSP | 3 |
| 2005 | Rate-Distortion Optimization for Internet Video Summarization and TransmissionabstractThe goal of video summarization is to generate a shorter video sequence of a lengthy original sequence using only the key frames of the original sequence. We consider a video summarization scheme that generates a video summary that can be transmitted over an unreliable network such as the Internet with minimum distortion of the original video. We consider two methods of distortion measurement in our optimization scheme, and we apply the methods to a scenario in which feedback is available with the possibility of retransmitting lost packets. Simulation results showing the effectiveness of using the proposed schemes are presented Peshala V. Pahalawatta, Zhu Li 0001, Fan Zhai, Aggelos K. Katsaggelos |
MMSP | 4 |
| 2005 | Advances in Efficient Resource Allocation for Packet-Based Real-Time Video TransmissionabstractMultimedia applications involving the transmission of video over communication networks are rapidly increasing in popularity. Such applications can greatly benefit from adapting video coding parameters to network conditions as well as adapting network parameters to better support the application requirements. These two dimensions can both be viewed as allocating source and network resources to improve video quality. We highlight recent advances in optimal resource allocation for real-time video communications over unreliable and resource constrained communication channels. More specifically, we focus on point-to-point coding and delivery schemes in which the sequences are encoded on the fly. We present a high-level framework for resource-distortion optimization. The framework can be used for jointly considering factors across network layers, including source coding, channel resource allocation, and error concealment. For example, resources can take the form of transmission energy in a wireless channel, and transmission cost in a DiffServ-based Internet channel. This framework can be used to optimally trade off resource consumption with end-to-end video quality in packet-based video transmission. After giving an overview of this framework, we review recent work in two areas-energy efficient wireless video transmission and resource allocation for Internet-based applications. Aggelos K. Katsaggelos, Yiftach Eisenberg, Fan Zhai, Randall Berry, Thrasyvoulos N. Pappas |
Proc. IEEE | 1 |
| 2005 | Joint source-channel coding and power adaptation for energy efficient wireless video communications
Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
Signal Process. Image Commun. | 5 |
| 2005 | MINMAX optimal video summarizationabstractThe need for video summarization originates primarily from a viewing time constraint. A shorter version of the original video sequence is desirable in a number of applications. Clearly, a shorter version is also necessary in applications where storage, communication bandwidth and/or power are limited. In this paper, our work is based on a MINMAX optimization formulation with viewing time, frame skip and bit rate constraints. New metrics for missing frame and video summary distortions are introduced. Optimal algorithm based on dynamic programming is presented along with experimental results. Zhu Li 0001, Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Rate-distortion optimal bit allocation for object-based video codingabstractIn object-based video encoding, the encoding of the video data is decoupled into the encoding of shape, motion, and texture information, which enables certain functionalities, like content-based interactivity and content-based scalability. The fundamental problem, however, of how to jointly encode this separate information to reach the best coding efficiency has not been studied thoroughly. In this paper, we present an operational rate-distortion optimal scheme for the allocation of bits among shape, motion, and texture in object-based video encoding. Our approach is based on Lagrangian relaxation and dynamic programming. We implement our algorithm on the MPEG-4 video verification model, although it is applicable to any object-based video encoding scheme. The performance is accessed utilizing a proposed metric that jointly captures the distortion due to the encoding of the shape and texture. Experimental results demonstrate that the gains of lossy shape encoding depend on the percentage the shape bits occupy out of the total bit budget. This gain may be small or may be realized at very low bit rates for certain typical scenes. Haohong Wang, Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Cost-distortion optimized unequal error protection for object-based video communicationsabstractObject-based video coding is a relatively new technique to meet the fast growing demand for interactive multimedia applications. Compared with conventional frame-based video coding, it consists of two types of source data: shape information and texture information. Recently, joint source-channel coding for multimedia communications has gained increased popularity. However, very limited work has been conducted to address the problem of joint source-channel coding for object-based video. In this paper, we propose a cost-distortion optimal unequal error protection (UEP) scheme for object-based video communications. Our goal is to achieve the best video quality (minimum total expected distortion) with constraints on transmission cost and delay in a lossy network environment. The problem is solved using Lagrangian relaxation and dynamic programming. The performance of the proposed scheme is tested using simulations of a narrow-band block-fading wireless channel with additive white Gaussian noise and a simplified differentiated services Internet channel. Experimental results indicate that the proposed UEP scheme can significantly outperform equal error protection methods. Haohong Wang, Fan Zhai, Yiftach Eisenberg, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | Rate-Distortion Optimal Video Summary GenerationabstractThe need for video summarization originates primarily from a viewing time constraint. A shorter version of the original video sequence is desirable in a number of applications. Clearly, a shorter version is also necessary in applications where storage, communication bandwidth, and/or power are limited. The summarization process inevitably introduces distortion. The amount of summarization distortion is related to its "conciseness," or the number of frames available in the summary. If there are m frames in the original sequence and n frames in the summary, we define the summarization rate as m/n, to characterize this "conciseness". We also develop a new summarization distortion metric and formulate the summarization problem as a rate-distortion optimization problem. Optimal algorithms based on dynamic programming are presented and compared experimentally with heuristic algorithms. Practical constraints, like the maximum number of frames that can be skipped, are also considered in the formulation and solution of the problem. Zhu Li 0001, Guido M. Schuster, Aggelos K. Katsaggelos, Bhavan Gandhi |
IEEE Trans. Image Process. | 3 |
| 2005 | A multicamera setup for generating stereo panoramic videoabstractTraditional visual communication systems convey only two-dimensional (2-D) fixed field-of-view (FOV) video information. The viewer is presented with a series of flat, nonstereoscopic images, which fail to provide a realistic sense of depth. Furthermore, traditional video is restricted to only a small part of the scene, based on the director's discretion and the user is not allowed to "look around" in an environment. The objective of this work is to address both of these issues and develop new techniques for creating stereo panoramic video sequences. A stereo panoramic video sequence should be able to provide the viewer with stereo vision at any direction (complete 360-degree FOV) at video rates. In this paper, we propose a new technique for creating stereo panoramic video using a multicamera approach, thus creating a high-resolution output. We present a setup that is an extension of a previously known approach, developed for the generation of still stereo panoramas, and demonstrate that it is capable of creating high-resolution stereo panoramic video sequences. We further explore the limitations involved in a practical implementation of the setup, namely the limited number of cameras and the nonzero physical size of real cameras. The relevant tradeoffs are identified and studied. Stavros Tzavidas, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 2 |
| 2005 | Joint source coding and packet classification for real-time video transmission over differentiated services networksabstractDifferentiated Services (DiffServ) is one of the leading architectures for providing quality of service in the Internet. We propose a scheme for real-time video transmission over a DiffServ network that jointly considers video source coding, packet classification, and error concealment within a framework of cost-distortion optimization. The selections of encoding parameters and packet classification are both used to manage end-to-end delay variations and packet losses within the network. We present two dual formulations of the proposed scheme: the minimum distortion problem, in which the objective is to minimize the end-to-end distortion subject to cost and delay constraints, and the minimum cost problem, which minimizes the total cost subject to end-to-end distortion and delay constraints. A solution to these problems using Lagrangian relaxation and dynamic programming is given. Simulation results demonstrate the advantage of jointly adapting the source coding and packet classification in DiffServ networks. Fan Zhai, Carlos E. Luna, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 6 |
| 2004 | Estimation of High Resolution Images and Registration Parameters from Low Resolution Observations
Salvador Villena, Javier Abad, Rafael Molina 0001, Aggelos K. Katsaggelos |
CIARP | 4 |
| 2004 | Comparison of low- and high-level visual features for audio-visual continuous automatic speech recognitionabstractWe compare two different groups of visual features that can be used in addition to audio to improve automatic speech recognition (ASR), high- and low-level visual features. Facial animation parameters (FAPs), supported by the MPEG-4 standard for the visual representation of speech, are used as high-level visual features. Principal component analysis (PCA) based projection weights of the intensity images of the mouth area are used as low-level visual features. PCA is also applied on the FAPs. We develop an audio-visual ASR (AV-ASR) system and compare its performance for two different visual feature groups, following two approaches. The first approach assumes the same dimensionality for both high- and low-level visual features, while, in the second approach, the percentage of statistical variance described by the visual features used is the same. Multi-stream hidden Markov models (HMMs) and a late integration approach are used to integrate audio and visual information and perform continuous AV-ASR experiments. Experiments were performed at various SNRs (0-30 dB) with additive white Gaussian noise on a relatively large vocabulary database (approximately 1000 words). Conclusions are drawn on the trade off between the dimensionality of the visual features and the amount of speechreading information contained in them and its influence on the AV-ASR performance. Petar S. Aleksic, Aggelos K. Katsaggelos |
ICASSP (5) | 2 |
| 2004 | An adaptive coding scheme using affine motion model for MPEG P-VOPabstractBlock matching has been used for motion estimation and motion compensation in MPEG standards for years. While it has an acceptable performance in describing motion between frames, it requires quite a few bits to represent the motion vectors. In certain circumstances, the use of whole frame affine motion models would perform equally well or even better than block matching in terms of motion accuracy, while it results in the coding of only 6 parameters. In this paper, we modify an MPEG-4 codec by adding: (1) 6 affine model parameters to the frame header; and (2) mode selection among INTRA, SKIP, INTER-16/spl times/16, INTER-8/spl times/8, and GLOBAL-AFFINE modes by Lagrange optimal rate-distortion criteria. Simulation results demonstrate 10-20% decrease in bit-rate, compared to the MMS codec for an average coded P-frame with the same reconstruction PSNR. Xiaohuan Li 0003, Joel R. Jackson, Aggelos K. Katsaggelos, Russell M. Mersereau |
ICASSP (3) | 3 |
| 2004 | Rate-distortion optimal video summarization: a dynamic programming solutionabstractThe need for video summarization originates primarily from a viewing time constraint. A shorter version of the original video sequence is desirable in a number of applications. Clearly, a shorter version is also necessary in applications where storage, communication bandwidth and/or power are limited. Our work is based on a temporal rate-distortion optimization formulation for optimal summary generation. New metrics for video summary distortion are introduced. Optimal algorithms based on dynamic programming are presented along with the results from heuristic algorithms that can produce near optimal results in real time. Zhu Li 0001, Guido M. Schuster, Aggelos K. Katsaggelos, Bhavan Gandhi |
ICASSP (3) | 3 |
| 2004 | DNA-based matching of digital signalsabstractAdleman with his pioneering work set the stage for the new field of bio-computing research (Science, vol.266, p.1021-1024, 1994). His main idea was to use actual chemistry to solve problems that are either unsolvable by conventional computers, or require an enormous amount of computation. The main focus of our research is to consider the application of molecular computing to the domain of digital signal processing (DSP). In this paper, we consider matching problems that arise in signal processing applications and are amenable to a DNA-based solution. Digital data are encoded in DNA sequences using a sophisticated codeword set that satisfies the noise tolerance constraint (NTC) that we introduce. NTC, one of the main contributions of our work, takes into account the presence of noise in digital signals by exploiting the annealing between non-perfect complementary sequences. We propose an algorithm to map binary values into DNA codewords by satisfying a number of constraints, including the NTC. Using that algorithm, we retrieved 128 codewords that enables us to use a DNA based approach to digital signal matching. Sotirios A. Tsaftaris, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas, Eleftherios T. Papoutsakis |
ICASSP (5) | 2 |
| 2004 | Robust network-adaptive object-based video encodingabstractThe joint source-network encoding of object-based video is a very important and challenging research topic, which has not been adequately explored. In this paper, we propose a robust network-adaptive encoding approach for object-based video. The framework jointly considers source coding, packet loss during transmission, and error concealment at the decoding. The proposed method guarantees the minimum expected distortion for the decoded video, by optimally allocating the shape and texture coding parameters at the encoder. The resulting optimization problem is solved by Lagrangian relaxation and dynamic programming. Experimental results demonstrate that the proposed method has significant gains over the non network-adaptive method. Haohong Wang, Aggelos K. Katsaggelos |
ICASSP (3) | 2 |
| 2004 | Rate-distortion optimized product code forward error correction for video transmission over IP-based wireless networksabstractThe problem of encoding and transmitting a video sequence over an IP-based wireless network, consisting of both wired and wireless links, is addressed. To combat the different types of packet loss in the heterogeneous network, the use of a product code forward error correction (FEC) scheme capable of providing unequal error protection is considered. At the transport layer, Reed-Solomon (RS) coding is used to provide inter-packet protection. In addition, rate-compatible punctured convolutional (RCPC) coding is used at the link layer to provide unequal intra-packet protection. Optimal bit allocation is performed in a rate-distortion optimized joint source-channel coding and power allocation framework to achieve the best video quality. Simulation results illustrate the advantage of the proposed product code FEC scheme over previously studied approaches. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICASSP (5) | 5 |
| 2004 | Energy efficient wireless transmission of MPEG-4 fine granular scalable videoabstractFine granular scalability is a coding tool, recently introduced in the emerging MPEG-4 standard, which enables the creation of very flexible scalable video bitstreams. This paper investigates the transmission of fine granular scalable (FGS) video over wireless links, using power management for unequal error protection of the bitstream. In wireless systems, energy may be a limited resource, and a wise use of it is important for system efficiency. An algorithm is proposed which is able to optimally distribute the total available power for the transmission of the enhancement layer, given a distortion or energy constraint. Experimental results demonstrate the performance advantage of the proposed algorithm over fixed power schemes and heuristic approaches. Cristina Emilia Costa, Yiftach Eisenberg, Fan Zhai, Aggelos K. Katsaggelos |
ICC | 4 |
| 2004 | Rate-distortion optimized hybrid error control for real-time packetized video transmissionabstractIn this paper, hybrid error control for real-time video transmission is studied. The study is carried out using a proposed integrated joint source-channel coding framework, which jointly considers error resilient source coding, channel coding, and error concealment, in order to achieve the best video quality and focuses on the performance comparison of several error correction scenarios, such as forward error correction (FEC), retransmission, and the combination of both. Simulation results show that either FEC or retransmission can be optimal depending on the packet loss rates and network round trip time. The proposed hybrid FEC/retransmission scheme outperforms both. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICC | 5 |
| 2004 | A Hybrid Source-Channel Coding Scheme for Object-based Wireless Video CommunicationsabstractWe study the joint source-channel coding of object-based video, and propose a data hiding scheme that improves the video error resilience by adaptively embedding the shape and motion information in the texture data. Within a rate-distortion theoretical framework, the source coding, channel coding, data embedding, and decoder error concealment are jointly optimized based on the knowledge of transmission channel conditions. The problem is solved using Lagrangian relaxation and dynamic programming. Experimental results indicate that the proposed hybrid source-channel coding scheme significantly outperforms methods without data hiding or unequal error protection. Haohong Wang, Aggelos K. Katsaggelos |
ICCCN | 2 |
| 2004 | Motion estimation in high resolution image reconstruction from compressed video sequences
Luis D. Alvarez, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2004 | Fast video shot retrieval by trace geometry matching in principal component space
Zhu Li 0001, Aggelos K. Katsaggelos, Bhavan Gandhi |
ICIP | 2 |
| 2004 | Optimal video summarization with a bit budget constraintabstractThe need for video summarization originates primarily from a viewing time or a bit budget constraint. A shorter version of the original video sequence is desirable in a number of applications. Clearly, a shorter version is also necessary in applications where storage, communication bandwidth and/or power are limited, which translates into a bit budget constraint. Our work is based on a bit rate-summary distortion optimization formulation. New metrics for video summary distortion are introduced. An optimal algorithm based on Lagrangian relaxation and dynamic programming is presented. Zhu Li 0001, Guido M. Schuster, Aggelos K. Katsaggelos, Bhavan Gandhi |
ICIP | 3 |
| 2004 | Optimal sensor selection for video-based target tracking in a wireless sensor networkabstractThe use of wireless sensor networks for target tracking is an active area of research. Imaging sensors that obtain video-rate images of a scene can have a significant impact in such networks, as they can measure vital information on the identity, position, and velocity of moving targets. Since wireless networks must operate under stringent energy constraints, it is important to identify the optimal set of imagers to be used in a tracking scenario such that the network lifetime is maximized. We formulate this problem as one of maximizing the information utility gained from a set of sensors subject to a constraint on the average energy consumption in the network. We use an unscented Kalman filter framework to solve the tracking and data fusion problem with multiple imaging sensors in a computationally efficient manner, and use a lookahead algorithm to optimize the sensor selection based on the predicted trajectory of the target. Simulation results show the effectiveness of this method of sensor selection. Peshala V. Pahalawatta, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2004 | Motion compensated shape error concealmentabstractThe introduction of video objects (VOs) is one of the innovations of MPEG-4. The /spl alpha/-plane of a VO defines its shape at a given instance in time and hence determines the boundary of its texture. In packet-based networks, shape, motion, and texture are subject to loss. In this paper, we propose a post-processing shape error concealment technique that uses the motion compensated boundary information of the previously received /spl alpha/-plane. The proposed approach is based on matching received boundary segments in the current frame to the boundary in the previous frame. This matching is achieved by finding a maximally smooth motion vector field. After the current boundary segments are matched to the previous boundary, the missing boundary pieces are reconstructed by motion compensation. Experimental results demonstrating the performance of the proposed motion compensated shape error concealment method, and comparisons with our previously proposed spatial Hermite spline method are presented. Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2004 | Robust circle detection using a weighted mse estimator
Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2004 | Channel modeling and its effect on the end-to-end distortion in wireless video communicationsabstractA major limitation faced by a mobile user is their dependence on a limited battery supply. For wireless video communications, joint source coding and transmission power management (JSCPM) has recently been considered as a means of efficiently allocating transmission energy. In order to reduce complexity, the design of many of these adaptive resource allocation algorithms utilizes simplified channel models that do not account for the burstiness of the channel. We analyze the effects of such channel model simplifications on the end-to-end distortion. We present a channel model that is based on information theoretic considerations, which captures the bursty nature of wireless channels and accounts for packet lengths when calculating the probability of loss. Given the source coding and transmission parameters derived using a simplified channel model, our goal is to analyze how the end-to-end distortion is affected when a more realistic complex channel model is used to simulate losses. Experimental results suggest that the performance gain predictions for JSCPM using a simpler channel model are also valid when more sophisticated channel simulations are used, provided that a number of additional steps are taken after the optimization to account for the complex characteristics of wireless channels. Eren Soyak, Yiftach Eisenberg, Fan Zhai, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP | 6 |
| 2004 | Joint object-based video encoding and power management for energy efficient wireless video communicationsabstractIn this paper, we consider dynamic resource allocation for object-based wireless video communications. In object-based video coding, a video frame is comprised of objects that are described by their shape as well as their texture. By jointly considering source coding, error concealment and transmission power management at the physical layer, the proposed framework minimize the expect distortion at the receiver for given energy and delay constraints. In order to provide unequal error protection for the shape and texture information, a new video packetization scheme is proposed. Experimental results indicate that the proposed unequal error protection schemes significantly outperform equal error protection methods. Haohong Wang, Yiftach Eisenberg, Fan Zhai, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2004 | Optimal object-based video communications over differentiated services networksabstractIn this paper, we propose an optimal unequal error protection scheme for object-based video communications over differentiated services networks. Our goal is to achieve the best video quality (minimum total expected distortion) with constraints on transmission cost and delay. An end-to-end distortion estimation approach for object-based video is proposed, which can be used for different packetization schemes. The problem is solved using Lagrangian relaxation and dynamic programming. Experimental results indicate that the proposed unequal error protection schemes can significantly outperform equal error protection methods. Haohong Wang, Fan Zhai, Yiftach Eisenberg, Aggelos K. Katsaggelos |
ICIP | 4 |
| 2004 | An integrated joint source-channel coding framework for video transmission over packet lossy networksabstractThe problem of application-layer error control for real-time video transmission over packet lossy networks is commonly addressed by joint source-channel coding (JSCC). The traditional JSCC approaches solve this problem in a sequential manner, where source coding and channel coding are not fully integrated. In this paper, we present an integrated joint source-channel coding (IJSCC) framework, where error resilient source coding, channel coding and error concealment are jointly considered in an integrated manner. We show through both analysis and simulations the advantages of the proposed IJSCC approach, in comparison to a sequential JSCC approach. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2004 | Speech-to-video synthesis using MPEG-4 compliant visual featuresabstractThere is a strong correlation between the building blocks of speech (phonemes) and the building blocks of visual speech (visimes). In this paper, this correlation is exploited and an approach is proposed for synthesizing the visual representation of speech from a narrow-band acoustic speech signal. The visual speech is represented in terms of the facial animation parameters (FAPs), supported by the MPEG-4 standard. The main contribution of this paper is the development of a correlation hidden Markov model (CHMM) system, which integrates independently trained acoustic HMM (AHMM) and visual HMM (VHMM) systems, in order to realize speech-to-video synthesis. The proposed CHMM system allows for different model topologies for acoustic and visual HMMs. It performs late integration and reduces the amount of required training data compared to early integration modeling techniques. Temporal accuracy experiments, comparison of the synthesized FAPs to the original FAPs, and audio-visual automatic speech recognition (AV-ASR) experiments utilizing the synthesized visual speech were performed in order to objectively measure the performance of the system. The objective experiments demonstrated that the proposed approach reduces time alignment errors by 40.5% compared to the conventional temporal scaling method, that the synthesized FAP sequences are very similar to the original FAP sequences, and that synthesized FAP sequences contain visual speechreading information that can improve AV-ASR performance. Petar S. Aleksic, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Introduction to the Special Issue on Audio and Video Analysis for Multimedia Interactive Services
Ebroul Izquierdo, Aggelos K. Katsaggelos, Michael G. Strintzis |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Joint optimal object shape estimation and encodingabstractA major problem in object-oriented video coding and MPEG-4 is the encoding of object boundaries. Traditionally this problem is treated separately from the texture encoding problem. In this paper, we present a vertex-based shape coding method which is optimal in the operational rate-distortion sense and takes into account the texture information of the video frames. This is accomplished by utilizing a variable-width tolerance band whose width is a function of the texture profile. As an example, this width is inversely proportional to the magnitude of the image gradient. Thus, in areas where the confidence in the estimation of the boundary is low and/or coding errors in the boundary will not affect the application (e.g., object-oriented coding and MPEG-4) significantly, a larger boundary approximation error is allowed. We present experimental results which demonstrate the effectiveness of the proposed algorithm. Lisimachos P. Kondi, Gerry Melnikov, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Spatially adaptive high-resolution image reconstruction of DCT-based compressed imagesabstractThe problem of recovering a high-resolution image from a sequence of low-resolution DCT-based compressed observations is considered in this paper. The introduction of compression complicates the recovery problem. We analyze the DCT quantization noise and propose to model it in the spatial domain as a colored Gaussian process. This allows us to estimate the quantization noise at low bit-rates without explicit knowledge of the original image frame, and we propose a method that simultaneously estimates the quantization noise along with the high-resolution data. We also incorporate a nonstationary image prior model to address blocking and ringing artifacts while still preserving edges. To facilitate the simultaneous estimate, we employ a regularization functional to determine the regularization parameter without any prior knowledge of the reconstruction procedure. The smoothing functional to be minimized is then formulated to have a global minimizer in spite of its nonlinearity by enforcing convergence and convexity requirements. Experiments illustrate the benefit of the proposed method when compared to traditional high-resolution image reconstruction methods. Quantitative and qualitative comparisons are provided. Sung Cheol Park, Moon Gi Kang, C. Andrew Segall, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2004 | Shape error concealment using Hermite splinesabstractThe introduction of video objects (VOs) is one of the innovations of MPEG-4. The alpha-plane of a VO defines its shape at a given instance in time and hence determines the boundary of its texture. In packet-based networks, shape, motion, and texture are subject to loss. While there has been considerable attention paid to the concealment of texture and motion errors, little has been done in the field of shape error concealment. In this paper, we propose a post-processing shape error-concealment technique that uses geometric boundary information of the received alpha-plane. Second-order Hermite splines are used to model the received boundary in the neighboring blocks, while third order Hermite splines are used to model the missing boundary. The velocities of these splines are matched at the boundary point closest to the missing block. There exists the possibility of multiple concealing splines per group of lost boundary parts. Therefore, we draw every concealment spline combination that does not self-intersect and keep all possible results until the end. At the end, we select the concealment solution that results in one closed boundary. Experimental results demonstrating the performance of the proposed method and comparisons with prior proposed methods are presented. Guido M. Schuster, Xiaohuan Li 0003, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2004 | Bayesian resolution enhancement of compressed videoabstractSuper-resolution algorithms recover high-frequency information from a sequence of low-resolution observations. In this paper, we consider the impact of video compression on the super-resolution task. Hybrid motion-compensation and transform coding schemes are the focus, as these methods provide observations of the underlying displacement values as well as a variable noise process. We utilize the Bayesian framework to incorporate this information and fuse the super-resolution and post-processing problems. A tractable solution is defined, and relationships between algorithm parameters and information in the compressed bitstream are established. The association between resolution recovery and compression ratio is also explored. Simulations illustrate the performance of the procedure with both synthetic and nonsynthetic sequences. C. Andrew Segall, Aggelos K. Katsaggelos, Rafael Molina 0001, Javier Mateos |
IEEE Trans. Image Process. | 2 |
| 2003 | Multi-channel Reconstruction of Video Sequences from Low-Resolution and Compressed Observations
Luis D. Alvarez, Rafael Molina 0001, Aggelos K. Katsaggelos |
CIARP | 3 |
| 2003 | Parameter estimation in super-resolution image reconstruction problemsabstractWe consider the estimation of the unknown hyperparameters for the problem of reconstructing a high-resolution image from multiple undersampled, shifted, degraded frames with subpixel displacement errors. We derive mathematical expressions for the iterative calculation of the maximum likelihood estimate (MLE) of the unknown hyperparameters given the low resolution observed images. Experimental results are presented for evaluating the accuracy of the proposed method. Javier Abad, Miguel Vega, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP (3) | 4 |
| 2003 | Bayesian high resolution image reconstruction with incomplete multisensor low resolution systemsabstractWe consider the problem of reconstructing a high-resolution image from an incomplete set of undersampled, shifted, degraded frames with subpixel displacement errors. We derive mathematical expressions for the calculation of the maximum a posteriori (MAP) estimate of the high resolution image given the low resolution observed images. We also examine the role played by the prior model when an incomplete set of low resolution images is used. Finally, the proposed method is tested on real and synthetic images. Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP (3) | 3 |
| 2003 | A VQ-based blur identification algorithmabstractThe estimation of the point spread function (PSF) of the degradation system is often a necessary first step in the restoration of blurred images. A novel vector quantization (VQ)-based blur identification algorithm is presented. A number of codebooks are designed corresponding to various versions of the blurring function. Prototype images blurred by each candidate blur are used. Only the non-flat regions for specific frequency bands are represented by the entries in the codebooks. Given a noisy and blurred image, one of the codebooks is chosen based on a similarity measure, therefore providing the identification of the blur. Simulations are performed for various blurring functions and noise levels. The results demonstrate the effectiveness of the proposed algorithms. Ryo Nakagaki, Aggelos K. Katsaggelos |
ICASSP (3) | 2 |
| 2003 | Speech-to-video synthesis using facial animation parametersabstractThe presence of visual information in addition to audio could improve speech understanding in noisy environments. This additional information could be especially useful for people with impaired hearing who are able to speechread. This paper focuses on the problem of synthesizing the facial animation parameters (FAPs), supported by the MPEG-4 standard for the visual representation of speech, from a narrowband acoustic speech (telephone) signal. A correlation hidden Markov model (CHMM) system for performing visual speech synthesis is proposed. The CHMM system integrates an independently trained acoustic HMM (AHMM) system and a visual HMM (VHMM) system, in order to realize speech-to-video synthesis. Analyzing the synthesized FAPs and computing the time alignment errors perform objective experiments. Time alignment errors are reduced by 40.5% compared to the conventional temporal scaling method. Petar S. Aleksic, Aggelos K. Katsaggelos |
ICIP (3) | 2 |
| 2003 | Variance-aware distortion estimation for wireless video communicationsabstractThe problem of encoding and transmitting a video sequence over a wireless channel is considered. Our objective is to minimize the end-to-end distortion while using a limited amount of transmission energy and delay. In our approach, we jointly adapt the source-coding parameters and transmission power per packet. We introduce the concept of "variance-aware distortion estimation" (VADE), and present a framework for controlling both the expected value and the variance of the end-to-end distortion. This framework is based on knowledge of how the video is compressed, the probability of packet loss, and the concealment strategy. To the best of our knowledge, this paper is the first to address the trade-off between the mean and variance of the end-to-end distortion. Experimental results demonstrate the potential of the proposed approach. Yiftach Eisenberg, Fan Zhai, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (1) | 6 |
| 2003 | Efficient frame vector selection based on ordered setsabstractThe problem of finding the optimal set of quantized coefficients for a frame-based encoded signal is known to be of very high complexity. This paper presents an efficient method of finding the operational rate-distortion (RD) optimal set of coefficients. The major complexity reduction lies in the reformulation of the original RD-tradeoff problem, where a new set of coefficients is used as decision variables. These coefficients are connected to the orthogonalization of the set of selected frame vectors and not to the frame vectors themselves. By organizing all possible solutions as nodes in a solution tree, we use complexity saving techniques to find the optimal solution in an even more efficient way. Using an ordered vector selection process, the complexity can be again significantly reduced and efficient run-length encoding becomes feasible. Contrary to the original problem, the new problem can be solved optimally in a reasonable amount of time. Tom Ryen, Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP (3) | 3 |
| 2003 | Robust line detection using a weighted MSE estimatorabstractIn this paper we introduce a novel line detection algorithm based on a weighted minimum mean square error (MSE) formulation. This algorithm has been developed to enable an autonomous robot to follow a white line drawn on the floor, but is general in nature and widely applicable to line detection problems. Traditional approaches to line detections consist of two stages, an edge detection stage and a line detection stage using the edge detection result. There are several problems with this approach. First, the initial edge detection stage is sensitive to noise. Second, the second stage does not use all the information available in the image and there fore incorrect decisions made by the first stage cannot be corrected in the second stage. The proposed algorithm achieves its robustness by operating in one step, using all pixels of the image (correctly weighted) and not using any thresholds. The detected line is the solution of a weighted MSE problem. The following three questions are answered in the paper: (I) what mathematical model should be used for the line? (II) how should the weighted MSE problem be set up so that the optimal solution results in the parameters of the line model? And (III), how should the pixels in the image be weighted such that a weighted MSE optimal solution results in a robust line detection? Experimental results demonstrate the performance of the algorithm in noiseless and noisy conditions. Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP (1) | 2 |
| 2003 | Spline-based boundary loss concealmentabstractObject-based video coding requires the transmission of the object shape. This shape is sent as a binary /spl alpha/-plane. In lossy packet-based networks, such as the Internet, this information has a nonnegligible probability of not arriving at the receiver, and hence its loss needs to be concealed. In this paper we propose a shape concealment technique utilizing Hermite splines. The algorithm has the following steps: (I) the received boundary is detected and the lost boundary parts are grouped using the packet loss pattern. (II) for each of these lost boundary parts, the received boundary points that border the area of the lost boundary parts are collected. These boundary points are then modelled by a second order Hermite spline. This model is subsequently used to match the velocity along the received boundary with the velocity of the concealing cubic Hermite spline. (Ill) since in most cases there are more than one concealing splines we draw every spline combination that does not result in an intersection and keep all possible results until the end. (IV) if there are more than one possible solutions we select the one that results in one overall closed nonintersecting boundary and fill the interior of the boundary to get the concealed /spl alpha/-plane. Experimental results which demonstrate and compare the performance of the proposed concealment method are given at the end of the paper. Guido M. Schuster, Xiaohuan Li 0003, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 2003 | Bayesian parameter estimation in image reconstruction from subsampled blurred observationsabstractIn this paper we consider the estimation of the unknown hyperparameters for the problem of reconstructing a high-resolution image from multiple undersampled, shifted, blurred and degraded frames with subpixel displacement errors. We derive mathematical expressions for the iterative calculation of the maximum likelihood estimate (mle) of the unknown hyperparameters given the low resolution observed images. Finally, the proposed method is tested on a synthetic image. Miguel Vega, Javier Mateos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP (2) | 4 |
| 2003 | Object-based video compression scheme with optimal bit allocation among shape, motion and textureabstractIn object-based video, the encoding of the video data is decoupled into the encoding of shape, motion and texture information, which enables certain functionalities like content-based interactivity and scalability. However, the problem of how to jointly encode these separate signals to reach the best coding efficiency has never been solved thoroughly. In this paper, we present an operational rate-distortion optimal bit allocation scheme that provides a solution to this problem. Our approach is based on the Lagrangian relaxation and dynamic programming. Experimental results indicate that the proposed optimal encoding approach has considerable gains over an ad-hoc method without optimization. Furthermore the proposed algorithm is much more efficient than exhaustive search. Haohong Wang, Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP (3) | 3 |
| 2003 | EM-based simultaneous registration, restoration, and interpolation of super-resolved imagesabstractA maximum likelihood (ML) solution to the problem of obtaining high-resolution images from sequences of noisy, blurred, and low-resolution images is presented. In our formulation, the registration parameters of the low-resolution images, the degrading blur, and noise variance are unknown. Our algorithm has the advantage that all unknown parameters are obtained simultaneously using all of the available data. An efficient implementation is presented in the frequency domain, based on the expectation maximization (EM) algorithm. Simulations demonstrate the effectiveness of the algorithm. Nathan A. Woods, Nikolas P. Galatsanos, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 2003 | A novel cost-distortion optimization framework for video streaming over differentiated services networksabstractThis paper presents a novel framework for streaming video over a Differentiated Services (DiffServ) network that jointly considers video source coding, packet classification and error concealment within the scope of cost-distortion optimization. Our formulation incorporates the random network delay for each packet into the calculation of the probability of packet loss and manages the end-to-end packet delay by selecting the encoding parameters and packet priority. We formulate two approaches to evaluate the performance of the proposed framework: a minimum distortion approach and a minimum cost approach, in which the encoding mode and priority class for each packet are optimally selected so as to minimize the total distortion subject to cost constraints, or to minimize the total cost subject to end-to-end distortion constraints. Simulation results demonstrate the advantage of jointly adapting the source coding and packet classification. Fan Zhai, Yiftach Eisenberg, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (3) | 6 |
| 2003 | Product HMMs for audio-visual continuous speech recognition using facial animation parametersabstractThe use of visual information in addition to acoustic can improve automatic speech recognition. In this paper we compare different approaches for audio-visual information integration and show how they affect automatic speech recognition performance. We utilize facial animation parameters (FAPs), supported by the MPEG-4 standard for the visual representation as visual features. We use both single-stream and multi-stream hidden Markov models (HMM) to integrate audio and visual information. We performed both state and phone synchronous multi-stream integration. Product HMM topology is used to model the phone-synchronous integration. ASR experiments were performed under noisy audio conditions using a relatively large vocabulary (approximately 1000 words) audio-visual database. The proposed phone-synchronous system, which performed the best, reduces the word error rate (WER) by approximately 20% relatively to audio-only ASR (A-ASR) WERs, at various SNRs with additive white Gaussian noise. Petar S. Aleksic, Aggelos K. Katsaggelos |
ICME | 2 |
| 2003 | Temporal rate-distortion based optimal video summary generationabstractVideo summary work originates from a viewing time constraint. A shorter version of the original video sequence is desirable in some applications. Clearly, a shorter version is necessary in applications where storage or bandwidth is limited. Our work is based on visual significance analysis of frames in a video sequence and a video temporal rate-distortion optimization framework. Temporal rate defines the number of frames allowable into the video summary, while temporal distortion is computed based on MPEG-7 metrics between mismatched frames. Several one-pass and two-pass algorithms are proposed along with discussions on experimental results. Zhu Li 0001, Aggelos K. Katsaggelos, Bhavan Gandhi |
ICME | 2 |
| 2003 | Minmax optimal shape coding using skeleton decompositionabstractIn this paper, we consider the rate-distortion optimal encoding of shape information using a skeleton decomposition and the minimum maximum (minmax) distortion criterion. For bit budget constrained video communication applications, whose goal is to achieve as low as possible but almost constant distortion, the minmax criterion is the natural choice. We propose a 4D DAG (directed acyclic graph) shortest path algorithm implemented by dynamic programming to solve the minimum rate problem, and also provide a solution for the dual minimum distortion problem. Experimental results indicate that our algorithm has an outstanding performance compared with existing methods. Haohong Wang, Guido M. Schuster, Aggelos K. Katsaggelos |
ICME | 3 |
| 2003 | Operational rate-distortion optimal bit allocation between shape and texture for MPEG-4 video codingabstractMPEG-4 is the first multimedia standard that supports the decoupling of a video object into object shape and object texture information, which consequently brings up the optimal encoding problem for object-based video. In this paper, we present an operational rate-distortion optimal bit allocation scheme between shape and texture for MPEG-4 encoding. Our approach is based on the Lagrange multiplier method, while the adoption of dynamic programming techniques enables its higher efficiency over the exhaustive search algorithm. Our work will not only benefit the further study of joint shape and texture encoding, but also make possible the deeper study of optimal joint source-channel coding of object-based video. Haohong Wang, Guido M. Schuster, Aggelos K. Katsaggelos |
ICME | 3 |
| 2003 | A rate-distortion optimized error control scheme for scalable video streaming over the InternetabstractVideo streaming over the Internet is a challenging task due, in part to the wide range of bandwidth variations caused by network congestion. To deal with this challenge, we propose an optimal error control scheme for scalable video transmission over the Internet. The three major components of error controlerror resilience, forward error correction (FEC), and error concealment- are considered in the proposed framework. Rate-distortion (R-D) optimization is carried out to determine the encoding mode for each packet and the channel coding rates, in order to minimize the overall expected end-to-end distortion. Our simulation study demonstrates that the proposed approach is robust to the wide range channel bandwidth variations and greatly outperforms the classical R-D optimization scheme. Fan Zhai, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICME | 4 |
| 2003 | Joint source coding and data rate adaptation for energy efficient wireless video streamingabstractRapid growth in wireless networks is fueling demand for video services from mobile users. While the problem of transmitting video over unreliable channels has received some attention, the wireless network environment poses challenges such as transmission power management that have received little attention previously in connection with video. Transmission power management affects battery life in mobile devices, interference to other users, and network capacity. We consider energy efficient transmission of a video sequence under delay and quality constraints. The selection of source coding parameters is considered jointly with transmitter power and rate adaptation, and packet transmission scheduling. The goal is to transmit a video frame using the minimal required transmission energy under delay and quality constraints. Experimental results are presented that illustrate the advantages of the proposed approach. Carlos E. Luna, Yiftach Eisenberg, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
IEEE J. Sel. Areas Commun. | 5 |
| 2003 | Analysis and FPGA Implementation of Image Restoration under Resource ConstraintsabstractProgrammable logic is emerging as an attractive solution for many digital signal processing applications. In this work, we have investigated issues arising due to the resource constraints of FPGA-based systems. Using an iterative image restoration algorithm as an example we have shown how to manipulate the original algorithm to suit it to an FPGA implementation. Consequences of such manipulations have been estimated, such as loss of quality in the output image. We also present performance results from an actual implementation on a Xilinx FPGA. Our experiments demonstrate that, for different criteria, such as result quality or speed, the best implementation is different as well. Seda Ogrenci Memik, Aggelos K. Katsaggelos, Majid Sarrafzadeh |
IEEE Trans. Computers | 2 |
| 2003 | Maximizing user utility in video streaming applicationsabstractWe study some of the design tradeoffs of video streaming systems in networks with QoS guarantees. We approach this problem by using a utility function to quantify the benefit a user derives from the quality of the received video sequence. We also consider the cost to the network user for streaming the video sequence. We have formulated this utility maximization problem as a joint constrained optimization problem where we maximize the difference between the utility and the network cost, subject to the constraint that the decoder buffer does not underflow. In this manner, we can find the optimal tradeoff between video quality and network cost. We present a deterministic dynamic programming approach for both the constant bit rate and renegotiated constant bit rate service classes. Experimental results demonstrate the benefits and the performance of the proposed approach. Carlos E. Luna, Lisimachos P. Kondi, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2003 | Bayesian multichannel image restoration using compound Gauss-Markov random fieldsabstractIn this paper, we develop a multichannel image restoration algorithm using compound Gauss-Markov random fields (CGMRF) models. The line process in the CGMRF allows the channels to share important information regarding the objects present in the scene. In order to estimate the underlying multichannel image, two new iterative algorithms are presented and their convergence is established. They can be considered as extensions of the classical simulated annealing and iterative conditional methods. Experimental results with color images demonstrate the effectiveness of the proposed approaches. Rafael Molina 0001, Javier Mateos, Aggelos K. Katsaggelos, Miguel Vega |
IEEE Trans. Image Process. | 3 |
| 2003 | Parameter estimation in Bayesian high-resolution image reconstruction with multisensorsabstractIn this paper, we consider the estimation of the unknown parameters for the problem of reconstructing a high-resolution image from multiple undersampled, shifted, degraded frames with subpixel displacement errors. We derive mathematical expressions for the iterative calculation of the maximum likelihood estimate of the unknown parameters given the low resolution observed images. These iterative procedures require the manipulation of block-semi circulant (BSC) matrices, that is, block matrices with circulant blocks. We show how these BSC matrices can be easily manipulated in order to calculate the unknown parameters. Finally the proposed method is tested on real and synthetic images. Rafael Molina 0001, Miguel Vega, Javier Abad, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2003 | A VQ-based blind image restoration algorithmabstractIn this paper, learning-based algorithms for image restoration and blind image restoration are proposed. Such algorithms deviate from the traditional approaches in this area, by utilizing priors that are learned from similar images. Original images and their degraded versions by the known degradation operator (restoration problem) are utilized for designing the VQ codebooks. The codevectors are designed using the blurred images. For each such vector, the high frequency information obtained from the original images is also available. During restoration, the high frequency information of a given degraded image is estimated from its low frequency information based on the codebooks. For the blind restoration problem, a number of codebooks are designed corresponding to various versions of the blurring function. Given a noisy and blurred image, one of the codebooks is chosen based on a similarity measure, therefore providing the identification of the blur. To make the restoration process computationally efficient, the principal component analysis (PCA) and VQ-nearest neighbor approaches are utilized. Simulation results are presented to demonstrate the effectiveness of the proposed algorithms. Ryo Nakagaki, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2003 | An efficient rate-distortion optimal shape coding approach utilizing a skeleton-based decompositionabstractIn this paper, we present a new shape-coding approach, which decouples the shape information into two independent signal data sets; the skeleton and the boundary distance from the skeleton. The major benefit of this approach is that it allows for a more flexible tradeoff between approximation error and bit budget. Curves of arbitrary order can be utilized for approximating both the skeleton and distance signals. For a given bit budget for a video frame, we solve the problem of choosing the number and location of the control points for all skeleton and distance signals of all boundaries within a frame, so that the overall distortion is minimized. An operational rate-distortion (ORD) optimal approach using Lagrangian relaxation and a four-dimensional direct acyclic graph (DAG) shortest path algorithm is developed for solving the problem. To reduce the computational complexity from O(N(5)) to O(N(3)), where N is the number of admissible control points for a skeleton, a suboptimal greedy-trellis search algorithm is proposed and compared with the optimal algorithm. In addition, an even more efficient algorithm with computational complexity O(N(2)) that finds an ORD optimal solution using a relaxed distortion criterion is also proposed and compared with the optimal solution. Experimental results demonstrate that our proposed approaches outperform existing ORD optimal approaches, which do not follow the same decomposition of the source data. Haohong Wang, Guido M. Schuster, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 3 |
| 2002 | High-resolution image reconstruction of low-resolution DCT-based compressed imagesabstractThe problem of recovering a high-resolution image from a sequence of low-resolution DCT-based compressed images is considered in this paper. The presence of the compression system complicates the recovery problem, as the operation reduces the amount of frequency aliasing in the low-resolution frames and introduces a non-linear quantization process. The effect of the quantization error and resulting inaccurate sub-pixel motion information is modeled as a zero-mean additive correlated Gaussian noise. A regularization functional is introduced not only to reflect the relative amount of registration error in each low-resolution image but also to determine the regularization parameter without any prior knowledge in the reconstruction procedure. The effectiveness of the proposed algorithm is demonstrated experimentally. Sung Cheol Park, Moon Gi Kang, C. Andrew Segall, Aggelos K. Katsaggelos |
ICASSP | 4 |
| 2002 | A rate-distortion optimal coding alternative to matching pursuitabstractThis paper presents a method to find the operational rate-distortion optimal solution for an overcomplete signal decomposition. The idea of using overcomplete dictionaries, or frames, is to get a sparse representation of the signal. Traditionally, suboptimal algorithms, such as Matching Pursuit (MP), is used for this purpose. When using frames in a lossy compression scheme, the major issue is to find the best possible rate-distortion (RD) tradeoff. Given the frame and the Variable Length Code (VLC) table embedded in the entropy coder, the solution of the problem of establishing the best RD tradeoff has a very high complexity. The proposed approach reduces this complexity significantly by structuring the solution approach such that the dependent quantizer allocation problem reduces into an independent one. It is important to note that this large reduction in complexity is achieved without sacrificing optimality. The optimal rate-distortion solution depends on the VLC table embedded in the entropy coder. Thus, VLC optimization is part of this work. We show experimentally that the new approach outperforms Rate-Distortion Optimized Matching Pursuit, previously proposed in [1]. Tom Ryen, Guido M. Schuster, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2002 | Reconstruction of high-resolution image frames from a sequence of low-resolution and compressed observationsabstractA framework for recovering high-resolution information from a sequence of sub-sampled and compressed observations is presented. Compression schemes that describe a video sequence through a combination of motion vectors and transform coefficients are the focus (e.g. the MPEG and ITU family of standards), and we consider the influence of both the motion vectors and transform coefficients within the reconstruction algorithm. A Bayesian approach is utilized to incorporate the information, and results show a discemable improvement in resolution, as compared to standard interpolation methods. C. Andrew Segall, Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos |
ICASSP | 3 |
| 2002 | Audio-visual continuous speech recognition using MPEG-4 compliant visual featuresabstractWe utilize facial animation parameters (FAPs), supported by the MPEG-4 standard for the visual representation of speech, in order to improve automatic speech recognition (ASR) significantly. We describe a robust and automatic algorithm for extraction of FAPs from visual data that requires no hand labeling or extensive training procedures. Multi-stream hidden Markov models (HMM) are used to integrate audio and visual information. ASR experiments are performed under both clean and noisy audio conditions using a relatively large vocabulary (approximately 1000 words). The proposed system reduces the word error rate (WER) by 20% to 23% relative to audio-only ASR WERs, at various SNRs with additive white Gaussian noise, and by 19% relative to the audio-only ASR WER under clean audio conditions. Petar S. Aleksic, Jay J. Williams, Zhilin Wu, Aggelos K. Katsaggelos |
ICIP (1) | 4 |
| 2002 | Energy efficient wireless video communications for the digital set-top boxabstractIn the future, digital set-top boxes may serve as the primary access point for wireless home networks, enabling mobile users to use videoconferencing as well as streaming applications on hand-held devices. In this scenario, an important issue that must be addressed is the limited energy supply of a mobile device. This is of course a relevant issue for any wireless device. We focus on methods for efficiently utilizing transmission energy in wireless video communications. We present a general framework for the problem of minimizing the transmission energy required to provide an acceptable level of video quality. We discuss two special cases in which communication resources are adjusted simultaneously with the source coding parameters in order to provide (i) packet loss adaptation and (ii) transmission rate adaptation. Yiftach Eisenberg, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (2) | 5 |
| 2002 | A color vector quantization based video coderabstractColor vector quantization (VQ) has been an efficient still image compression scheme as well as a popular bitmap graphics format for display devices with limited color capability. In this paper we are proposing a color vector quantization-based video coder, exploiting the temporal stationary nature of color distribution among a group of pictures (GOP) over a short period. Color VQ is applied first to reduce the RGB image sequence into a single channel color index image. Motion estimation and compression is then performed in the index space, instead of the separate YCbCr channels. Initial results demonstrated that the proposed coder can provide good compression rates. By eliminating the need for an inverse DCT and color conversion, typical requirements in a JPEG/MPEG type of coders, the decoding is computationally very simple. This makes it suitable for certain applications like media playback, and visual communications with low-end mobile devices. Zhu Li 0001, Aggelos K. Katsaggelos |
ICIP (3) | 2 |
| 2002 | Optimal source coding and transmission power management using a min-max expected distortion approachabstractWe consider the problem of compressing a video sequence for transmission over a wireless channel. In our approach we jointly consider error resilience and concealment techniques, at the source coding level, and transmission power management at the physical layer. We formulate a minimum-maximum distortion problem, where our goal is to either (i) minimize the total transmission energy for a given maximum expected distortion, or (ii) minimize the maximum expected distortion at the receiver for a given maximum transmission energy. Experimental results show that simultaneously adjusting the source coding and transmission power is more energy efficient than considering these factors separately. Carlos E. Luna, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (1) | 5 |
| 2002 | A VQ-based image restoration algorithmabstractIn this paper, we develop a novel VQ-based image restoration algorithm. The mapping between high frequency information in the original images and low frequency information in the corresponding degraded ones is established and stored in the VQ codebooks. Prototype images are used, which belong to the same class of images. During restoration, the high frequency information of a given degraded image is estimated from its low frequency information based on the designed codebook. To make the restoration process computationally efficient, the principal component analysis (PCA) and VQ-nearest neighborhood approaches are utilized. Simulation results are presented to demonstrate the effectiveness of the proposed algorithm. Ryo Nakagaki, Aggelos K. Katsaggelos |
ICIP (1) | 2 |
| 2002 | Spatially adaptive high-resolution image reconstruction of low-resolution DCT-based compressed imagesabstractThe problem of recovering a high-resolution image from a sequence of low-resolution DCT-based compressed images is considered. The presence of the compression system complicates the recovery problem, as the operation reduces the amount of frequency aliasing in the low-resolution frames and introduces a non-linear quantization process, The effect of the quantization error and resulting inaccurate sub-pixel motion information is modeled as a zero-mean additive correlated Gaussian noise. A regularization functional is introduced, not only to reflect the relative amount of registration error in each low-resolution image, but also to determine the regularization parameter without any prior knowledge in the reconstruction procedure. The effectiveness of the proposed algorithm is demonstrated experimentally. Sung Cheol Park, Moon Gi Kang, C. Andrew Segall, Aggelos K. Katsaggelos |
ICIP (2) | 4 |
| 2002 | A recursive shape error concealment algorithmabstractThe encoding of shape information is a distinguishing feature of MPEG-4. In error prone communication networks, it is important and efficient to conceal shape errors spatially, so as to avoid propagation of errors in the video frames. The proposed system first defines the missing area's four rectangular neighbors. It then detects in the neighbors the lines that intersect the borders and redefines them in a geometric coordinate system. These lines are paired under some optimization objective so that they connect smoothly in the missing area. Each pair is extrapolated and corrected recursively until a close curve is formed. Test results on various sequences under various error rates are presented and compared with other shape concealment methods. Guido M. Schuster, Xiaohuan Li 0003, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 2002 | Disparity variation in stereo-panoramic videoabstractIn this paper we study the problem of generating stereo panoramic video. We propose a setup based on the theory of circular projections and we demonstrate that it is capable of generating stereo panoramic images at video rates. We further study some of the limitations involved in a practical implementation of the proposed setup, with an emphasis on the effects on the disparity, as perceived by human observers. Stavros Tzavidas, Aggelos K. Katsaggelos |
ICIP (3) | 2 |
| 2002 | Lip Tracking for MPEG-4 Facial AnimationabstractIt is very important to accurately track the mouth of a talking person for many applications, such as face recognition and human computer interaction. This is in general a difficult problem due to the complexity of shapes, colors, textures, and changing lighting conditions. We develop techniques for outer and inner lip tracking. From the tracking results FAPs are extracted which are used to drive an MPEG-4 decoder. A novel method consisting of a Gradient Vector Flow (GVF) snake with a parabolic template as an additional external force is proposed. Based on the results of the outer lip tracking, the inner lip is tracked using a similarity function and a temporal smoothness constraint. Numerical results are presented using the Bernstein database. Zhilin Wu, Petar S. Aleksic, Aggelos K. Katsaggelos |
ICMI | 3 |
| 2002 | SPECT Image Reconstruction Using Compound Prior ModelsabstractWe propose a new iterative method for Maximum a Posteriori (MAP) reconstruction of SPECT (Single Photon Emission Computed Tomography) images. The method uses Compound Gauss Markov Random Fields (CGMRF) as prior model and is stochastic for the line process and deterministic for the reconstruction. Synthetic and real images are used to compare the new method with existing ones. Antonio López, Rafael Molina 0001, Javier Mateos, Aggelos K. Katsaggelos |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2002 | Joint source coding and transmission power management for energy efficient wireless video communicationsabstractWe consider a situation where a video sequence is to be compressed and transmitted over a wireless channel. Our goal is to limit the amount of distortion in the received video sequence, while minimizing transmission energy. To accomplish this goal, we consider error resilience and concealment techniques at the source coding level, and transmission power management at the physical layer. We jointly consider these approaches in a novel framework. In this setting, we formulate and solve an optimization problem that corresponds to minimizing the energy required to transmit video under distortion and delay constraints. Experimental results show that simultaneously adjusting the source coding and transmission power is more energy efficient than considering these factors separately. Yiftach Eisenberg, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2002 | Joint source-channel coding for motion-compensated DCT-based SNR scalable videoabstractIn this paper, we develop an approach toward joint source-channel coding for motion-compensated DCT-based scalable video coding and transmission. A framework for the optimal selection of the source and channel coding rates over all scalable layers is presented such that the overall distortion is minimized. The algorithm utilizes universal rate distortion characteristics which are obtained experimentally and show the sensitivity of the source encoder and decoder to channel errors. The proposed algorithm allocates the available bit rate between scalable layers and, within each layer, between source and channel coding. We present the results of this rate allocation algorithm for video transmission over a wireless channel using the H.263 Version 2 signal-to-noise ratio (SNR) scalable codec for source coding and rate-compatible punctured convolutional (RCPC) codes for channel coding. We discuss the performance of the algorithm with respect to the channel conditions, coding methodologies, layer rates, and number of layers. Lisimachos P. Kondi, Faisal Ishtiaq, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2002 | A jointly optimal fractal/DCT compression schemeabstractIn this paper a hybrid fractal and discrete cosine transform (DCT) coder is developed. Drawing on the ability of DCT to remove inter-pixel redundancies and on the ability of fractal transforms to capitalize on long-range correlations within the image, the hybrid coder performs an operationally optimal, in the rate-distortion sense, bit allocation among coding parameters. An orthogonal basis framework is used within which an image segmentation and a hybrid block-based transform are selected jointly. The selection of coefficients in the DCT component of the overall block transform is made a part of the optimization procedure. A Lagrangian multiplier approach is used to optimize the hybrid transform parameters together with the segmentation. Differential encoding of the DC coefficient is employed, with the scanning path based on a 3rd-order Hilbert curve. Simulation results show a significant improvement in quality with respect to the JPEG standard, an approach based on optimization of DCT basis vectors, as well as, the purely fractal techniques. Gerry Melnikov, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 2 |
| 2002 | An HMM-based speech-to-video synthesizerabstractEmerging broadband communication systems promise a future of multimedia telephony, e.g. the addition of visual information to telephone conversations. It is useful to consider the problem of generating the critical information useful for speechreading, based on existing narrowband communications systems used for speech. This paper focuses on the problem of synthesizing visual articulatory movements given the acoustic speech signal. In this application, the acoustic speech signal is analyzed and the corresponding articulatory movements are synthesized for speechreading. This paper describes a hidden Markov model (HMM)-based visual speech synthesizer. The key elements in the application of HMMs to this problem are the decomposition of the overall modeling task into key stages and the judicious determination of the observation vector's components for each stage. The main contribution of this paper is a novel correlation HMM model that is able to integrate independently trained acoustic and visual HMMs for speech-to-visual synthesis. This model allows increased flexibility in choosing model topologies for the acoustic and visual HMMs. Moreover the propose model reduces the amount of training data compared to early integration modeling techniques. Results from objective experiments analysis show that the propose approach can reduce time alignment errors by 37.4% compared to conventional temporal scaling method. Furthermore, subjective results indicated that the purpose model can increase speech understanding. Jay J. Williams, Aggelos K. Katsaggelos |
IEEE Trans. Neural Networks | 2 |
| 2001 | Joint source-channel coding for scalable video using models of rate-distortion functionsabstractA joint source-channel coding scheme for scalable video is developed in this paper. An SNR scalable video coder is used and unequal error protection (UEP) is allowed for each scalable layer. Our problem is to allocate the available bit rate across scalable layers and, within each layer, between source and channel coding, while minimizing the end-to-end distortion of the received video sequence. The resulting optimization algorithm we propose utilizes universal rate-distortion characteristic plots. These plots show the contribution of each layer to the total distortion as a function of the source rate of the layer and the residual bit error rate (the error rate that remains after the use of channel coding). Models for these plots are proposed in order to reduce the computational complexity of the solution. Experimental results demonstrate the effectiveness of the proposed approach. Lisimachos P. Kondi, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2001 | SPECT image reconstruction using compound modelsabstractSPECT (single photon emission computed tomography) is used in nuclear medicine to determine the distribution of a radioactive isotope within a patient from tomographic views or projection data. These images are severely degraded due to the presence of noise and several physical factors like attenuation and scattering. We use, within the Bayesian framework, a compound Gauss Markov random field (CGMRF) as prior model to reconstruct such images. In order to find the maximum a posteriori (MAP) estimate we propose a new iterative method, which is stochastic for the line process and deterministic for the reconstruction. The proposed method is tested and compared with other reconstruction methods on both synthetic and real SPECT images. Antonio López, Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos |
ICASSP | 3 |
| 2001 | Maximizing user utility in video streaming applicationsabstractWe study the design tradeoffs involved in video streaming in networks with QoS guarantees. We approach this problem by using a utility function to quantify the benefit a user derives from the received video sequence. This benefit is expressed as a function of the total distortion. In addition, we also consider the cost, in network resources, of a video streaming system. The goal of the network user is then to obtain the most benefit for the smallest cost. We formulate this utility maximization problem as a joint constrained optimization problem. The difference between the utility and the network cost is maximized subject to the constraint that the decoder buffer does not underflow. We present a deterministic dynamic programming approach to find the optimal tradeoff for both the constant bit rate (CBR) and renegotiated constant bit rate (RCBR) service classes. Experimental results demonstrate the benefits and the performance of the proposed approach. Carlos E. Luna, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2001 | Rate distortion optimal signal compression using second order polynomial approximationabstractWe present a time domain signal compression algorithm based on the coding of line segments which are used to approximate the signal. These segments are fitted in a way that is optimal in the rate distortion sense. The approach is applicable to many types of signals, but in this paper we focus on the compression of electrocardiogram (ECG) signals. As opposed to traditional time-domain algorithms, where heuristics are used to extract representative signal samples from the original signal, an optimization algorithm is formulated for sample selection using graph theory, with linear interpolation applied to the reconstruction of the signal. In this paper the algorithm is generalized by using second order polynomial interpolation for the reconstruction of the signal from the extracted signal samples. The polynomials are fitted in a way that guarantees minimum reconstruction error given an upper bound on the number of bits. The method achieves good performance compared both to the case where linear interpolation is used in reconstruction of the signal and to other state-of-the-art ECG coders. Ranveig Nygaard, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2001 | A new constraint for the regularized enhancement of compressed videoabstractA novel fidelity constraint to the image enhancement problem is presented. With this constraint, we exploit the motion vectors of a compressed video bit-stream. These vectors establish a correspondence between image pixels across a series of frames, and we guarantee that processing the decoded sequence does not violate this correspondence. We develop the constraint within the context of MPEG-2 and incorporate the constraint into a regularized enhancement algorithm. Simulations are then performed. Quantitative and qualitative results illustrate an improvement in visual quality. C. Andrew Segall, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2001 | Minimizing transmission energy in wireless video communicationsabstractA key constraint in mobile communications is the reliance on a battery with a limited energy supply. Efficiently utilizing the available energy is therefore an important design consideration. We consider a situation where a video sequence is to be compressed and transmitted over a wireless channel. The goal is to limit the amount of distortion in the received video sequence while using the minimum required transmission energy. To accomplish this goal, we consider error resilience and concealment techniques, at the source coding level, as well as the dynamic allocation of physical layer communication resources. We consider these approaches jointly in a novel framework. We formulate an optimization problem that corresponds to minimizing the energy required to transmit a video frame with an acceptable level of distortion. We present methods for solving this problem and other extensions. Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (1) | 4 |
| 2001 | Joint source-channel coding for scalable video over DS-CDMA multipath fading channelsabstractWe extend our previous work on joint source-channel coding to scalable video transmission over wireless direct-sequence code-division-multiple-access (DS-CDMA) multipath fading channels. A SNR scalable video coder is used and unequal error protection (UEP) is allowed for each scalable layer. At the receiver-end an adaptive antenna array auxiliary-vector (AV) filter is utilized that provides space-time RAKE-type processing and multiple-access interference suppression. The choice of the AV receiver is dictated by realistic channel fading rates that limit the data record available for receiver adaptation and redesign. Our problem is to allocate the available bit rate of the user of interest between source and channel coding and across scalable layers, while minimizing the end-to-end distortion of the received video sequence. The optimization algorithm that we propose utilizes universal rate-distortion characteristic curves that show the contribution of each layer to the total distortion as a function of the source rate of the layer and the residual bit error rate (the error rate after channel coding). These plots can be approximated using appropriate functions to reduce the computational complexity of the solution. Lisimachos P. Kondi, Stella N. Batalama, Dimitris A. Pados, Aggelos K. Katsaggelos |
ICIP (1) | 4 |
| 2001 | Jointly optimal coding of texture and shapeabstractA major problem in object oriented video coding and MPEG-4 is the encoding of object boundaries. Traditionally, and within MPEG-4, the encoding of shape and texture information are separate steps (the extraction of shape is not considered by the standards). We present a vertex-based shape coding method which is optimal in the operational rate-distortion sense and takes into account the texture information of the video frames. This is accomplished by utilizing a variable-width tolerance band which is proportional to the degree of trust in the accuracy of the shape information at that location. Thus, in areas where the confidence in the estimation of the boundary is not high and/or coding errors in the boundary will not affect the application (object oriented coding, MPEG-4, etc.) significantly, a larger boundary approximation error is allowed. We present experimental results which demonstrate the effectiveness of the proposed algorithm. Lisimachos P. Kondi, Gerry Melnikov, Aggelos K. Katsaggelos |
ICIP (3) | 3 |
| 2001 | Bayesian high-resolution reconstruction of low-resolution compressed videoabstractA method for simultaneously estimating the high-resolution frames and the corresponding motion field from a compressed low-resolution video sequence is presented. The algorithm incorporates knowledge of the spatio-temporal correlation between low and high-resolution images to estimate the original high-resolution sequence from the degraded low-resolution observation. Information from the encoder is also exploited, including the transmitted motion vectors, quantization tables, coding modes and quantizer scale factors. Simulations illustrate an improvement in the peak signal-to-noise ratio when compared with traditional interpolation techniques and are corroborated with visual results. Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos, C. Andrew Segall |
ICIP (2) | 2 |
| 2001 | Application of the motion vector constraint to the regularized enhancement of compressed videoabstractWe present a novel fidelity constraint for the image enhancement problem by exploiting the motion vectors of a compressed video bit-stream. These vectors establish a correspondence between image pixels across a series of frames, and our goal is to maintain this relationship during processing. In our past work, we considered algorithms that relied on the sum-of-absolute differences as the match criteria. As we show in this paper, this metric is problematic for the enhancement problem. We then pose the constraint within the context of a sum-of-squared errors criterion for matching. This allows for a more rigorous treatment of the fidelity constraint. Finally, experimental results illustrate the performance of the new constraint. C. Andrew Segall, Aggelos K. Katsaggelos |
ICIP (1) | 2 |
| 2001 | A rate-distortion optimal video pre-processing algorithmabstractPre-processing algorithms improve the quality of a compression system by removing unimportant data before encoding. This enhances both the visual quality and coding efficiency of the system. We cast the pre-processing problem in the operational rate-distortion framework. Filtering the displaced frame difference is the focus, and the proposed method couples the choice of the quantization scale to the response of the prefilter. Coding errors are then addressed by penalizing significant differences between coded blocks. Finally, experimental results illustrate the efficacy of the method within the context of an MPEG-2 coding scenario. C. Andrew Segall, Passant V. Karunaratne, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 2001 | Rate-distortion optimal skeleton-based shape codingabstractWe present a new shape-coding approach, which decouples the shape information into two independent data sets, the skeleton and the distance of the boundary from the skeleton. The major benefit of this approach is that it allows a more flexible trade-off between accuracy of the approximation and bit-allocation cost, and thus, provides the possibility of better performance in the operational rate-distortion (ORD) optimal sense than other reported techniques. The characteristics of these data sets are studied and various approximation approaches are applied on each of them to reach an ORD optimal result. We apply, for example, polygonal approximation on both the skeleton and distance data. We demonstrate that the resulting approach outperforms existing ORD optimal approaches. Haohong Wang, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 2001 | Preprocessing of compressed digital video
C. Andrew Segall, Passant V. Karunaratne, Aggelos K. Katsaggelos |
VCIP | 3 |
| 2001 | An operational rate-distortion optimal single-pass SNR scalable video coderabstractIn this paper, we introduce a new methodology for signal-to-noise ratio (SNR) video scalability based on the partitioning of the DCT coefficients. The DCT coefficients of the displaced frame difference (DFD) for inter-blocks or the intensity for intra-blocks are partitioned into a base layer and one or more enhancement layers, thus, producing an embedded bitstream. Subsets of this bitstream can be transmitted with increasing video quality as measured by the SNR. Given a bit budget for the base and enhancement layers the partitioning of the DCT coefficients is done in a way that is optimal in the operational rate-distortion sense. The optimization is performed using Lagrangian relaxation and dynamic programming (DP). Experimental results are presented and conclusions are drawn. Lisimachos P. Kondi, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2001 | Resolution enhancement of monochrome and color video using motion compensationabstractWe propose an iterative algorithm for enhancing the resolution of monochrome and color image sequences. Various approaches toward motion estimation are investigated and compared. Improving the spatial resolution of an image sequence critically depends upon the accuracy of the motion estimator. The problem is complicated by the fact that the motion field is prone to significant errors since the original high-resolution images are not available. Improved motion estimates may be obtained by using a more robust and accurate motion estimator, such as a pel-recursive scheme instead of block matching, in processing color image sequences, there is the added advantage of having more flexibility in how the final motion estimates are obtained, and further improvement in the accuracy of the motion field is therefore possible. This is because there are three different intensity fields (channels) conveying the same motion information. In this paper, the choice of which motion estimator to use versus how the final estimates are obtained is weighed to see which issue is more critical in improving the estimated high-resolution sequences. Toward this end, an iterative algorithm is proposed, and two sets of experiments are presented. First, several different experiments using the same motion estimator but three different data fusion approaches to merge the individual motion fields were performed. Second, estimated high-resolution images using the block matching estimator were compared to those obtained by employing a pel-recursive scheme. Experiments were performed on a real color image sequence, and performance was measured by the peak signal to noise ratio (PSNR). Brian C. Tom, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2000 | FPGA implementation and analysis of image restorationabstractNo abstract available. F. S. Ogrenci, Aggelos K. Katsaggelos, Majid Sarrafzadeh |
FPGA | 2 |
| 2000 | Resolution enhancement of compressed low resolution videoabstractWe propose an iterative algorithm for the estimation of high resolution frames from a low resolution compressed video sequence. The algorithm exploits the existing correlation between the high and low resolution frames and the information provided by the encoder to obtain a high resolution frame. The performance of the algorithm is demonstrated experimentally. Javier Mateos, Aggelos K. Katsaggelos, Rafael Molina 0001 |
ICASSP | 2 |
| 2000 | A rate-distortion optimal scalable vertex based shape coding algorithmabstractWe present a rate-distortion (RD) optimized scalable vertex-based shape coding algorithm. Following the base layer, each successive enhancement layer refines a given shape approximation by optimally (within a layer) placing new vertices and perturbing existing vertices. An efficient low entropy distortion adaptive vertex coding strategy is employed to take advantage of information available from coarser layers. Based on the chosen vertex rate and distortion definitions, a resulting enhancement layer topology is solved by executing a directed acyclic graph (DAG) shortest path algorithm. Finally, an iterative VLC optimization scheme is employed to find both the optimized scalable code and the most efficient set of parameter VLC tables. Gerry Melnikov, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2000 | Multichannel image restoration using compound Gauss-Markov random fieldsabstractA solution to the multichannel image restoration problem is provided using compound Gauss-Markov random fields. For the single channel deblurring problem the convergence of the simulated annealing (SA) and iterative conditional mode (ICM) algorithms has not been established. We propose two new iterative multichannel restoration algorithms which can be considered as extensions of the classical SA and ICM approaches and whose convergence is established. Experimental results with color images demonstrate the effectiveness of the proposed algorithms. Rafael Molina 0001, Javier Mateos, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2000 | A hidden Markov model based visual speech synthesizerabstractThis paper describes a hidden Markov model (HMM) based visual synthesizer designed to assist persons with impaired hearing. This synthesizer builds on results in the area of audio-visual speech recognition. We describe how a correlation HMM can be used to integrate independent acoustic and visual HMMs for speech-to-visual synthesis. Our results show that an HMM correlating model can significantly improve synchronization errors versus techniques which compensate for rate differences through scaling. Jay J. Williams, Aggelos K. Katsaggelos, Mark A. Randolph |
ICASSP | 2 |
| 2000 | Simultaneous Motion Estimation and Resolution Enhancement of Compressed low Resolution VideoabstractWe propose an iterative algorithm for simultaneously estimating the motion field and high resolution frames from a compressed low resolution video sequence. The algorithm exploits the existing correlation between high and low resolution frames and information provided by the encoder, such as coding modes and motion vectors (when available), to obtain a higher resolution frame. The performance of the algorithm is demonstrated experimentally. Javier Mateos, Aggelos K. Katsaggelos, Rafael Molina 0001 |
ICIP | 2 |
| 2000 | Shape Approximation Through Recursive Scalable Layer GenerationabstractThis paper presents an efficient recursive algorithm for generating operationally optimal intra mode scalable layer decompositions of object contours. The problem is posed in terms of minimizing the shape distortion at full reconstruction subject to the total (for all scalable layers) bit budget constraint. Based on the chosen vertex-based representation, we solve the problem of determining the number and locations of approximating vertices for all scalable layers jointly and optimally. The number of scalable layers is not constrained, but, rather, is a by-product of the proposed optimization. The algorithm employs two different coding strategies: one for the base layer and one for the enhancement layers. By carefully defining scalable layer recursion and base layer segment costs the problem is solved by executing a directed acyclic graph (DAG) shortest path algorithm. Gerry Melnikov, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2000 | Enhancement of Compressed Video Using Visual Quality MeasurementsabstractThe enhancement of compressed video is considered. We present a general algorithm for processing the compressed data, with three variants of the algorithm having practical application. We then consider the algorithm within the context of MPEG-2. Assuming complete knowledge of the compressed bitstream, experiments compare the different realizations of the enhancement algorithm. Our comparisons stress improvements in visual quality, measured by models of the human visual system. Quantitative and qualitative results are provided. C. Andrew Segall, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2000 | Restoration of severely blurred high range images using stochastic and deterministic relaxation algorithms in compound Gauss?CMarkov random fields
Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos, Aurora Hermoso, C. Andrew Segall |
Pattern Recognit. | 2 |
| 2000 | A mathematical model for shape coding with B-splines
Fabian W. Meier, Guido M. Schuster, Aggelos K. Katsaggelos |
Signal Process. Image Commun. | 3 |
| 2000 | Shape coding using temporal correlation and joint VLC optimizationabstractThis paper investigates ways to explore the between frame correlation of shape information within the framework of an operationally rate-distortion (ORD) optimized coder. Contours are approximated both by connected second-order spline segments, each defined by three consecutive control points, and by segments of the motion-compensated reference contours. Consecutive control points are then encoded predictively using angle and run temporal contexts or by tracking the reference contour. We utilize a novel criterion for selecting global object motion vectors, which improves the efficiency. The problem is formulated as a Lagrangian minimization and solved using dynamic programming. Furthermore, we employ an iterative technique to remove dependency on a particular variable length code and jointly arrive at the ORD globally optimal solution and an optimized conditional parameter distribution. Gerry Melnikov, Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2000 | Sequential construction of 3-D-based scene descriptionabstractBinocular camera systems are commonly used to construct 3-D-based scene description. However, there is a tradeoff between the length of the camera baseline and the difficulty of the matching problem and the extent of the field of view of the 3-D scene. A large baseline system provides better depth resolution than a smaller baseline system at the expense of a narrower field of view. To increase the depth resolution without increasing the difficulty of the matching problem and decreasing the field of view of the 3-D scene, a sequential 3-D-based scene description technique is proposed. Multiple small-baseline 3-D scene descriptions from a single moving camera or an array of cameras are used to sequentially construct a large baseline 3-D scene description while maintaining the field of view of a small-baseline system. A Bayesian framework using a disparity-space image (DSI) technique for disparity estimation is presented. The cost function for large baseline image matching is designed based not only on the photometric matching error, the smoothness constraint, and the ordering constraint, but also on the previous disparity estimates from smaller baseline stereo image pairs as a prior model. Texture information is registered along the scan path of the camera(s). Experimental results demonstrate the effectiveness of this technique in visual communication applications. Chun-Jen Tsai, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Hierarchical Bayesian image restoration from partially known blursabstractIn this paper, we examine the restoration problem when the point-spread function (PSF) of the degradation system is partially known. For this problem, the PSF is assumed to be the sum of a known deterministic and an unknown random component. This problem has been examined before; however, in most previous works the problem of estimating the parameters that define the restoration filters was not addressed. In this paper, two iterative algorithms that simultaneously restore the image and estimate the parameters of the restoration filter are proposed using evidence analysis (EA) within the hierarchical Bayesian framework. We show that the restoration step of the first of these algorithms is in effect almost identical to the regularized constrained total least-squares (RCTLS) filter, while the restoration step of the second is identical to the linear minimum mean square-error (LMMSE) filter for this problem. Therefore, in this paper we provide a solution to the parameter estimation problem of the RCTLS filter. We further provide an alternative approach to the expectation-maximization (EM) framework to derive a parameter estimation algorithm for the LMMSE filter. These iterative algorithms are derived in the discrete Fourier transform (DFT) domain; therefore, they are computationally efficient even for large images. Numerical experiments are presented that test and compare the proposed algorithms. Nikolas P. Galatsanos, Vladimir Z. Mesarovic, Rafael Molina 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2000 | A Bayesian approach for the estimation and transmission of regularization parameters for reducing blocking artifactsabstractWith block-based compression approaches for both still images and sequences of images annoying blocking artifacts are exhibited, primarily at high compression ratios. They are due to the independent processing (quantization) of the block transformed values of the intensity or the displaced frame difference. We propose the application of the hierarchical Bayesian paradigm to the reconstruction of block discrete cosine transform (BDCT) compressed images and the estimation of the required parameters. We derive expressions for the iterative evaluation of these parameters applying the evidence analysis within the hierarchical Bayesian paradigm. The proposed method allows for the combination of parameters estimated at the coder and decoder. The performance of the proposed algorithms is demonstrated experimentally. Javier Mateos, Aggelos K. Katsaggelos, Rafael Molina 0001 |
IEEE Trans. Image Process. | 2 |
| 1999 | Inter mode vertex-based optimal shape codingabstractThis paper investigates the problem of optimal lossy encoding of object contours in the inter mode. Contours are approximated by connected second-order spline segments, each defined by three consecutive control points. Taking into account correlations in the temporal direction, control points are chosen optimally in the rate-distortion sense. Applying motion to contours in the reference frame followed by the temporal context extraction, we predict the next control point location, given the previously encoded one. Based on the chosen differential encoding scheme and an additive MPEG-4 based distortion metric, the problem is formulated as Lagrangian minimization. We utilize an iterative procedure to jointly find the optimal solution and the associated DPCM parameter probability mass functions. Gerry Melnikov, Guido M. Schuster, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 1999 | Bayesian image restoration using a wavelet-based subband decompositionabstractThe subband decomposition of a single channel image restoration problem is examined. The decomposition is carried out in the image model (prior model) in order to take into account the frequency activity of each band of the original image. The hyperparameters associated with each band together with the original image are rigorously estimated within the Bayesian framework. Finally, the proposed method is tested and compared with other methods on real images. Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Abad |
ICASSP | 2 |
| 1999 | Optical flow estimation from noisy data using differential techniquesabstractMany optical flow estimation techniques are based on the differential optical flow equation. These algorithms involve solving over-determined systems of optical flow equations. Least squares (LS) estimation is usually used to solve these systems even though the underlying noise does not conform to the model implied by LS estimation. To ameliorate this problem, work has been done using the total least squares (TLS) method instead. However, the noise model presumed by TLS is again different from the noise present in the system of optical flow equations. A proper way to solve the system of optical flow equations is the constrained total least squares (CTLS) technique. The derivation and analysis of the CTLS technique for optical flow estimation is presented in this paper. It is shown that CTLS outperforms TLS and LS optical flow estimation. Chun-Jen Tsai, Nikolas P. Galatsanos, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 1999 | A Rate Control Method for H.263 Temporal ScalabilityabstractIn this paper a novel methodology for temporal scalability within the H.263 video coding standard is presented. Temporal scalability is defined as the transmission of previously dropped frames in the form of a scalable enhancement layer to increase the overall encoded frame rate. The proposed methodology extends the base layer rate control to the enhancement layer and incorporates an adaptive technique for enhancement layer frame selection. Three criteria important in the selection of enhancement frames are identified, allowing us to adaptively choose frames such that the overall temporal resolution of the encoded video sequence is enhanced. Experimental results are provided and compared to a non-adaptive technique where enhancement frames are selected solely on the decay of the enhancement layer buffer. Faisal Ishtiaq, Aggelos K. Katsaggelos |
ICIP (4) | 2 |
| 1999 | An Optimal Single Pass SNR Scalable Video CoderabstractIn this paper, we introduce a new methodology for SNR video scalability which is based on the partitioning of the DCT coefficients. The partitioning is done in a way that is optimal in the rate-distortion sense. The optimization is performed using Lagrangian relaxation and dynamic programming (DP). Experimental results are presented and conclusions are drawn. Lisimachos P. Kondi, Aggelos K. Katsaggelos |
ICIP (1) | 2 |
| 1999 | Hyperparameter Estimation for Emission Computed Tomography DataabstractAlthough many statistical methods have been proposed for the restoration of tomographic images, their use in medical environments has been limited due to two important factors. These factors are the need for greater computational time than deterministic methods and the selection of the hyperparameters in the image models. Consequently, deterministic methods, like the classical filtered back-projection (FBP) and algebraic reconstruction (AR), are commonly used. In this work, we propose a method to estimate, from observed image data in emission tomography, the hyperparameters in a Generalized Gaussian Markov Random Field (GGMRF). We use the hierarchical Bayesian approach and evidence analysis to reconstruct the image and estimate the unknown hyperparameters. The method is tested on synthetic images. Antonio López, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 1999 | Jointly Optimal Inter-Mode Shape Coding and VLC SelectionabstractThis paper investigates how the between frame correlation of shape information can be exploited within the framework of an operationally rate-distortion (ORD) optimized coder. Contours are approximated either by connected second-order spline segments, each defined by three consecutive control points, or by segments of the motion-compensated reference contours. Consecutive control points are then encoded predictively using angle and run temporal contexts or by tracking the reference contour. We employ a novel criterion for selecting global object motion vectors, which further improves efficiency. The problem is formulated as Lagrangian minimization and solved using Dynamic Programming (DP). Furthermore, we employ an iterative technique to remove dependency on a particular VLC and jointly arrive at the OPD globally optimal solution and an optimized conditional parameter distribution. Gerry Melnikov, Aggelos K. Katsaggelos, Guido M. Schuster |
ICIP (2) | 2 |
| 1999 | Rate Distortion Optimal ECG Signal CompressionabstractSignal compression is an important problem encountered in many applications. Various techniques have been proposed over the years for addressing the problem. In this paper we present a time domain algorithm based on the coding of line segments which are used to approximate the signal. These segments are fit in a way that is optimal in the rate distortion sense. Although the approach is applicable to any type of signal, we focus, in this paper, on the compression of ElectroCardioGram (EGG) signals. ECG signal compression has traditionally been tackled by heuristic approaches. However, it has been demonstrated that exact optimization algorithms outperform these heuristic approaches by a wide margin with respect to reconstruction error. By formulating the compression problem as a graph theory problem, known optimization theory can be applied in order to yield optimal compression. In this paper we present an algorithm that will guarantee the smallest possible distortion among all methods applying linear interpolation given an upper bound on the number of bits. Compared to many other compression methods, we report superior performance for this method. Ranveig Nygaard, Gerry Melnikov, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 1999 | Sequential Construction of 3D-Based Scene DescriptionabstractA technique for constructing a sequential 3D-based scene description is proposed in this paper. Multiple small baseline 3D scene descriptions from a single moving camera or an array of cameras are used to sequentially construct a large baseline 3D scene description while maintaining the field of view of a small baseline system. A Bayesian framework using a disparity-space image (DSI) technique for disparity estimation is presented. The cost function for large baseline image matching is designed based not only on the photometric matching error, the smoothness constraint, and the ordering constraint, but also on the previous disparity estimates from smaller baseline stereo image pairs as a prior model. Texture information is registered along the scan path of the camera(s). Experimental results demonstrate the effectiveness of this technique in visual communication applications. Chun-Jen Tsai, Aggelos K. Katsaggelos |
ICIP (2) | 2 |
| 1999 | A Compressed Video Enhancement AlgorithmabstractThe problem of the enhancement of a low bit-rate compressed video sequence using the information provided by the encoder is investigated in this paper. The proposed algorithm is spatio-temporally adaptive and enforces different degrees of between-block, within-block, and temporal smoothness of the decompressed frames based on macroblock types. The algorithm uses the projections onto the sets that capture the information conveyed by the transmitted data to constrain the solution space. A partially automatic regularization parameter estimation algorithm is also proposed in this paper. An analysis of PSNR gain per macroblock type is presented to cast insight into the compressed video enhancement problem. The experimental results demonstrate the effectiveness of the proposed approach. Chun-Jen Tsai, Passant V. Karunaratne, Nikolas P. Galatsanos, Aggelos K. Katsaggelos |
ICIP (3) | 4 |
| 1999 | Context based optimal shape codingabstractThis paper investigates how context-based DPCM techniques can be used in conjunction with an operationally rate-distortion (ORD) optimized shape coder. Object contours are approximated by connected second-order spline segments, each defined by three consecutive control points, and, in the case of the inter mode, by segments of the motion-compensated reference contours. In scalable intra mode and non-scalable inter mode shape coding, consecutive control points are encoded predictively, using an angle and run framework, with respect to corresponding contexts in the reference frame and the base layer, respectively. We employ a novel criterion for selecting global object motion vectors in the inter mode, which further improves efficiency. Gerry Melnikov, Aggelos K. Katsaggelos, Guido M. Schuster |
MMSP | 2 |
| 1999 | Error concealment algorithms for compressed video
Min-Cheol Hong, Harald Schwab, Lisimachos P. Kondi, Aggelos K. Katsaggelos |
Signal Process. Image Commun. | 4 |
| 1999 | Bayesian and regularization methods for hyperparameter estimation in image restorationabstractIn this paper, we propose the application of the hierarchical Bayesian paradigm to the image restoration problem. We derive expressions for the iterative evaluation of the two hyperparameters applying the evidence and maximum a posteriori (MAP) analysis within the hierarchical Bayesian paradigm. We show analytically that the analysis provided by the evidence approach is more realistic and appropriate than the MAP approach for the image restoration problem. We furthermore study the relationship between the evidence and an iterative approach resulting from the set theoretic regularization approach for estimating the two hyperparameters, or their ratio, defined as the regularization parameter. Finally the proposed algorithms are tested experimentally. Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos |
IEEE Trans. Image Process. | 2 |
| 1999 | A Review of the Minimum Maximum Criterion for Optimal Bit Allocation Among Dependent QuantizersabstractIn this paper, we review a general framework for the optimal bit allocation among dependent quantizers based on the minimum maximum (MINMAX) distortion criterion. The pros and cons of this optimization criterion are discussed and compared to the well-known Lagrange multiplier method for the minimum average (MINAVE) distortion criterion. We argue that, in many applications, the MINMAX criterion is more appropriate than the more popular MINAVE criterion. We discuss the algorithms for solving the optimal bit allocation problem among dependent quantizers for both criteria and highlight the similarities and differences. We point out that any problem which can be solved with the MINAVE criterion can also be solved with the MINMAX criterion, since both approaches are based on the same assumptions. We discuss uniqueness of the MINMAX solution and the way both criteria can be applied simultaneously within the same optimization framework. Furthermore, we show how the discussed MINMAX approach can be directly extended to result in the lexicographically optimal solution. Finally, we apply the discussed MINMAX solution methods to still image compression, intermode frame compression of H.263, and shape coding applications. Guido M. Schuster, Gerry Melnikov, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 3 |
| 1999 | Dense Disparity Estimation with a Divide-and-Conquer Disparity Space Image TechniqueabstractA new divide-and-conquer technique for disparity estimation is proposed in this paper. This technique performs feature matching following the high confidence first principle, starting with the strongest feature point in the stereo pair of scanlines. Once the first matching pair is established, the ordering constraint in disparity estimation allows the original intra-scanline matching problem to be divided into two smaller subproblems. Each subproblem can then be solved recursively until there is no reliable feature point within the subintervals. This technique is very efficient for dense disparity map estimation for stereo images with rich features. For general scenes, this technique can be paired up with the disparity-space image (DSI) technique to compute dense disparity maps with integrated occlusion detection. In this approach, the divide-and-conquer part of the algorithm handles the matching of stronger features and the DSI-based technique handles the matching of pixels in between feature points and the detection of occlusions. An extension to the standard disparity-space technique is also presented to compliment the divide-and-conquer algorithm. Experiments demonstrate the effectiveness of the proposed divide-and-conquer DSI algorithm. Chun-Jen Tsai, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 2 |
| 1998 | Blind image restoration using local bound constraintsabstractA new method of incorporating local image characteristics into blind image restoration is proposed. The local variance of the degraded image is used as a measure of spatial activity, from which individual pixel bounds are determined. A parameter defined by the user controls the degree of smoothing. The local bounds define the solution more precisely than smoothness constraints on the image (including those that are spatially-adaptive), reducing the number of possible solutions and leading to a faster rate of convergence. Experimental results demonstrate the potential of this method as an alternative/supplement to smoothing constraints in blind image restoration. Kaaren L. May, Tania Stathaki, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 1998 | A non uniform segmentation optimal hybrid fractal/DCT image compression algorithmabstractIn this paper a hybrid fractal and discrete cosine transform (DCT) coder is developed. Drawing on the ability of DCT to remove inter-pixel redundancies and on the ability of fractal transforms to capitalize on long-range correlations in the image, the hybrid coder performs an optimal, in the rate-distortion sense, bit allocation among coding parameters. An orthogonal basis framework is used within which an image segmentation and a hybrid block-based transform are selected jointly. A Lagrangian multiplier approach is used to optimize the hybrid parameters and the segmentation. Differential encoding of the DC coefficient is employed, with the scanning path based on a 3rd-order Hilbert curve. Simulation results show a significant improvement in quality with respect to the JPEG standard. Gerry Melnikov, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 1998 | Hierarchical Bayesian image restoration from partially-known blursabstractA number of restoration filters have been proposed for the restoration problem from partially-known blurs. Previously we proposed the regularized constrained least-squares filter (RCTLS) and showed that it has a number of advantages over previous ones (Mesarovic et al. 1995). However, the problem of estimating the parameters that define the RCTLS filter has not yet been addressed. In this paper we propose a two-step algorithm based on the hierarchical Bayesian approach to simultaneously restore the image and estimate the parameters of the RCTLS restoration filter. The algorithm is derived in the DFT domain; thus, it is very efficient even for very large images. Vladimir Z. Mesarovic, Nikolas P. Galatsanos, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICASSP | 4 |
| 1998 | On Video SNR Scalability
Lisimachos P. Kondi, Faisal Ishtiaq, Aggelos K. Katsaggelos |
ICIP (3) | 3 |
| 1998 | Reduction of Blocking Artifacts in Block Transformed Compressed Color ImagesabstractWe use the information in the chrominance bands to reconstruct color block transformed compressed images. For the luminance and the two chrominance channels, we define a reconstruction problem and show how to estimate the unknown hyperparameters and reconstruct each band automatically. The method is tested on real images. Javier Mateos, Carlos Ilia Herráiz Montalvo, Blas C. Ruiz Jiménez, Rafael Molina 0001, Aggelos K. Katsaggelos |
ICIP (1) | 5 |
| 1998 | Iterative Determination of Local Bound Constraints in Iterative Image RestorationabstractIn this paper, the problem of how to better estimate spatially adaptive intensity bounds for image restoration is addressed. When the intensity bounds are estimated from a degraded image, blurring leads to underestimation of the bounds in the edge and texture regions. Therefore, an iterative implementation of the restoration algorithm has been proposed in which the intensity bounds are re-estimated from the current image estimate. However, direct update of the bounds leads to over-smoothing in regions where the bounds are active. Furthermore, the resulting algorithm exhibits slow convergence. In this paper, alternative methods of initially estimating and updating the bounds are proposed, and the results for the fixed- and updated-bound implementations are compared. A method for estimation of the bound tightness parameter is also proposed. Kaaren L. May, Tania Stathaki, Anthony G. Constantinides, Aggelos K. Katsaggelos |
ICIP (2) | 4 |
| 1998 | Simultaneous Optimal Boundary Encoding and Variable-Length Code SelectionabstractThis paper describes efficient and optimal encoding and representation of object contours. Contours are approximated by connected second-order spline segments, each defined by three consecutive control points. The placement of the control points is done optimally in the rate-distortion (RD) sense and jointly with their entropy encoding. We utilize a differential scheme for the rate and an additive area-based metric for the distortion to formulate the problem as a Lagrangian minimization. We investigate the sensitivity of the resulting operational RD curve on the variable length codes used and propose an iterative procedure arriving at the entropy representation of the original boundary for any given rate-distortion tradeoff. Gerry Melnikov, Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 1998 | Total Least Squares Estimation of Stereo Optical FlowabstractWe propose a new method for disparity assisted stereo optical flow estimation. This method is based on the linearization of the round-about compatibility constraint, which converts a stereo optical flow estimation problem to a single channel optical flow estimation problem. An over-determined system of optical flow equations can then be constructed for estimating the flow fields. The total least squares, instead of the traditional least squares is used to solve the estimation problem. We also investigate the extension of the locally constant flow model across the time domain. The experiments presented demonstrate that the proposed technique performs very well. Chun-Jen Tsai, Nikolas P. Galatsanos, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 1998 | Regularized Restoration of Partial-Response Distortions in Sporadically Degraded Images
Damon L. Tull, Aggelos K. Katsaggelos |
ICIP (3) | 2 |
| 1998 | MPEG-4 and rate-distortion-based shape-coding techniquesabstractWe address the problem of the efficient encoding of object boundaries. This problem is becoming increasingly important in applications such as content-based storage and retrieval, studio and television postproduction, and mobile multimedia applications. The MPEG-4 visual standard will allow the transmission of arbitrarily shaped video objects. The techniques developed for shape coding within the MPEG-4 standardization effort are described and compared first. A framework for the representation of shapes using their contours is presented next. Such representations are achieved using curves of various orders, and they are optimal in the rate-distortion sense. Finally, conclusions are drawn. Aggelos K. Katsaggelos, Lisimachos P. Kondi, Fabian W. Meier, Jörn Ostermann, Guido M. Schuster |
Proc. IEEE | 1 |
| 1998 | Guest Editorial Applications Of Artificial Neural Networks To Image Processing
Rama Chellappa, Kunihiko Fukushima, Aggelos K. Katsaggelos, Sun-Yuan Kung, Yann LeCun, Nasser M. Nasrabadi, Tomaso A. Poggio |
IEEE Trans. Image Process. | 3 |
| 1998 | Hybrid image segmentation using watersheds and fast region mergingabstractA hybrid multidimensional image segmentation algorithm is proposed, which combines edge and region-based techniques through the morphological algorithm of watersheds. An edge-preserving statistical noise reduction approach is used as a preprocessing stage in order to compute an accurate estimate of the image gradient. Then, an initial partitioning of the image into primitive regions is produced by applying the watershed transform on the image gradient magnitude. This initial segmentation is the input to a computationally efficient hierarchical (bottom-up) region merging process that produces the final segmentation. The latter process uses the region adjacency graph (RAG) representation of the image regions. At each step, the most similar pair of regions is determined (minimum cost RAG edge), the regions are merged and the RAG is updated. Traditionally, the above is implemented by storing all RAG edges in a priority queue. We propose a significantly faster algorithm, which additionally maintains the so-called nearest neighbor graph, due to which the priority queue size and processing time are drastically reduced. The final segmentation provides, due to the RAG, one-pixel wide, closed, and accurately localized contours/surfaces. Experimental results obtained with two-dimensional/three-dimensional (2-D/3-D) magnetic resonance images are presented. Kostas Haris, Serafim N. Efstratiadis, Nicos Maglaveras, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 1998 | An optimal polygonal boundary encoding scheme in the rate distortion senseabstractIn this paper, we present fast and efficient methods for the lossy encoding of object boundaries that are given as eight-connect chain codes. We approximate the boundary by a polygon, and consider the problem of finding the polygon which leads to the smallest distortion for a given number of bits. We also address the dual problem of finding the polygon which leads to the smallest bit rate for a given distortion. We consider two different classes of distortion measures. The first class is based on the maximum operator and the second class is based on the summation operator. For the first class, we derive a fast and optimal scheme that is based on a shortest path algorithm for a weighted directed acyclic graph. For the second class we propose a solution approach that is based on the Lagrange multiplier method, which uses the above-mentioned shortest path algorithm. Since the Lagrange multiplier method can only find solutions on the convex hull of the operational rate distortion function, we also propose a tree-pruning-based algorithm that can find all the optimal solutions. Finally, we present results of the proposed schemes using objects from the Miss America sequence. Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 1998 | An optimal quadtree-based motion estimation and motion-compensated interpolation scheme for video compressionabstractWe propose an optimal quadtree (QT)-based motion estimator for video compression. It is optimal in the sense that for a given bit budget for encoding the displacement vector field (DVF) and the QT segmentation, the scheme finds a DVF and a QT segmentation which minimizes the energy of the resulting displaced frame difference (DFD). We find the optimal QT decomposition and the optimal DVF jointly using the Lagrangian multiplier method and a multilevel dynamic program. We introduce a new, very fast convex search for the optimal Lagrangian multiplier lambda(*), which results in a very fast convergence of the Lagrangian multiplier method. The resulting DVF is spatially inhomogeneous, since large blocks are used in areas with simple motion and small blocks in areas with complex motion. We also propose a novel motion-compensated interpolation scheme which uses the same mathematical tools developed for the QT-based motion estimator. One of the advantages of this scheme is the globally optimal control of the tradeoff between the interpolation error energy and the DVF smoothness. Another advantage is that no interpolation of the DVF is required since we directly estimate the DVF and the QT-segmentation for the frame which needs to be interpolated. We present results with the proposed QT-based motion estimator which show that for the same DFD energy the proposed estimator uses about 25% fewer bits than the commonly used block matching algorithm. We also experimentally compare the interpolated frames using the proposed motion compensated interpolation scheme with the reconstructed original frames. Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 1997 | A Bayesian approach to blind deconvolution based on Dirichlet distributionsabstractThis paper deals with the simultaneous identification of the blur and the restoration of a noisy and blurred image. We propose the use of Dirichlet distributions to model our prior knowledge about the blurring function together with smoothness constraints on the restored image to solve the blind deconvolution problem. We show that the use of Dirichlet distributions offers a lot of flexibility in incorporating vague or very precise knowledge about the blurring process into the blind deconvolution process. The proposed MAP estimator offers additional flexibility in modeling the original image. Experimental results demonstrate the performance of the proposed algorithm. Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Abad, Javier Mateos |
ICASSP | 2 |
| 1997 | Optimal bit allocation among dependent quantizers for the minimum maximum distortion criterionabstractIn this paper we introduce an optimal bit allocation scheme for dependent quantizers for the minimum maximum distortion criterion. First we show how minimizing the bit rate for a given maximum distortion can be achieved in a dependent coding framework using dynamic programming (DP). Then we employ an iterative algorithm to minimize the maximum distortion for a given bit rate, which invokes the DP scheme. We prove that it converges to the optimal solution. Finally we present a comparison between the minimum total distortion criterion and the minimum maximum distortion criterion for the encoding of an H.263 Intra frame. In this comparison we also point out the similarities between the proposed minimum maximum distortion approach and the Lagrangian multiplier based minimum total distortion approach. Guido M. Schuster, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 1997 | An Iterative Weighted Regularized Algorithm for improving the resolution of Video SequencesabstractThis paper introduces an iterative regularized approach to increase the resolution of a video sequence. A multiple input smoothing convex functional is defined and used to obtain a globally optimal high resolution video sequence. A mathematical model of multiple inputs is described by using the point spread function between the original and bilinearly interpolated images in the spatial domain, and motion estimation between frames in the temporal domain. An iterative algorithm is utilized for obtaining the solution. The regularization parameter is updated at each iteration step from the partially restored video sequence. Experimental results demonstrate the capability of the proposed approach. Min-Cheol Hong, Moon Gi Kang, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 1997 | A Mixed Norm Image RestorationabstractIn this paper, we propose an iterative mixed norm image restoration algorithm. A functional which combines the least mean squares (LMS) and the least mean fourth (LMF) functionals is proposed. A function of the kurtosis is used to determine the relative importance between the LMS and the LMF functionals. An iterative algorithm is utilized for obtaining a solution and its convergence is analyzed. Experimental results demonstrate the capability of the proposed approach. Min-Cheol Hong, Tania Stathaki, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 1997 | An Efficient Boundary Encoding Scheme which is Optimal in the Rate-Distortion SenseabstractA major problem in object oriented video coding is the efficient encoding of the shape information of arbitrarily shaped objects. Efficient shape coding schemes are also needed in encoding the shape information of video object planes (VOP) in the MPEG-4 standard. In this paper, we present an efficient method for the lossy encoding of object shapes which are given as 8-connect chain codes (Meier et al., 1997). We approximate the object shape by a second order B-spline curve and consider the problem of finding the curve with the lowest bit rate for a given distortion. The presented scheme is optimal, efficient and offers complete control over the trade-off between bit-rate and distortion. We present results with the proposed scheme using objects shapes of different sizes. Fabian W. Meier, Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 1997 | Model-Based Synthetic View Generation from a Monocular Video SequenceabstractIn this paper a model-based multi-view image generation system for video conferencing is presented. The system assumes that a 3-D model of the person in front of the camera is available. It extracts texture from speaking person sequence images and maps it to the static 3-D model during the videoconference session. Since only the incrementally updated texture information is transmitted during the whole session, the bandwidth requirement is very small. Based on the experimental results one can conclude that the proposed system is very promising for practical applications. Chun-Jen Tsai, Aggelos K. Katsaggelos, Peter Eisert, Bernd Girod |
ICIP (1) | 2 |
| 1997 | Frame rate and viseme analysis for multimedia applicationsabstractIn the future, multimedia technology will be able to provide video frame rates equal to or better than 30 frames per second (FPS). Until that time the hearing impaired community will be using band limited communication systems over unshielded twisted pair copper wiring. As a result, multimedia communication systems will use a coder/decoder (CODEC) to compress the video and audio signals for transmission. For these systems to be usable by the hearing impaired community, the algorithms within the CODEC have to be designed to account for the perceptual boundaries of the hearing impaired. We investigate the perceptual boundaries of speech reading and multimedia technology, which are the constraints that effect speech reading performance. We analyze and draw conclusions on the relationship between viseme groupings, accuracy of viseme recognition, and presentation rate. These results are critical in the design of multimedia systems for the hearing impaired. Jay J. Williams, Janet C. Rutledge, Dean C. Garstecki, Aggelos K. Katsaggelos |
MMSP | 4 |
| 1997 | A Theory for the Optimal Bit Allocation Between Displacement Vector Field and Displaced Frame DifferenceabstractWe address the fundamental problem of optimally splitting a video sequence into two sources of information, the displaced frame difference (DFD) and the displacement vector field (DVF). We first consider the case of a lossless motion-compensated video coder (MCVC), and derive a general dynamic programming (DP) formulation which results in an optimal tradeoff between the DVF and the DFD. We then consider the more important case of a lossy MCVC, and present an algorithm which solves the tradeoff between the rate and the distortion. This algorithm is based on the Lagrange multiplier method and the DP approach introduced for the lossless MCVC. We then present an H.263-based MCVC which uses the proposed optimal bit allocation, and compare its results to H.263. As expected, the proposed coder is superior in the rate-distortion sense. In addition to this, it offers many advantages for a rate control scheme. The presented theory can be applied to build new optimal coders, and to analyze the heuristics employed in existing coders. In fact, whenever one changes an existing coder, the proposed theory can be used to evaluate how the change affects its performance. Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE J. Sel. Areas Commun. | 2 |
| 1997 | Simultaneous multichannel image restoration and estimation of the regularization parametersabstractIn this correspondence, a constrained least-squares multichannel image restoration approach is proposed, in which no prior knowledge of the noise variance at each channel or the degree of smoothness of the original image is required. The regularization functional for each channel is determined by incorporating both within-channel and cross-channel information. It is shown that the proposed smoothing functional has a global minimizer. Moon Gi Kang, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 1997 | A video compression scheme with optimal bit allocation among segmentation, motion, and residual errorabstractWe present a theory for the optimal bit allocation among quadtree (QT) segmentation, displacement vector field (DVF), and displaced frame difference (DFD). The theory is applicable to variable block size motion-compensated video coders (VBSMCVC), where the variable block sizes are encoded using the QT structure, the DVF is encoded by first-order differential pulse code modulation (DPCM), the DFD is encoded by a block-based scheme, and an additive distortion measure is employed. We derive an optimal scanning path for a QT that is based on a Hilbert curve. We consider the case of a lossless VBSMCVC first, for which we develop the optimal bit allocation algorithm using dynamic programming (DP). We then consider a lossy VBSMCVC, for which we use Lagrangian relaxation, and show how an iterative scheme, which employs the DP-based solution, can be used to find the optimal solution. We finally present a VBSMCVC, which is based on the proposed theory, which employs a DCT-based DFD encoding scheme. We compare the proposed coder with H.263. The results show that it outperforms H.263 significantly in the rate distortion sense, as well as in the subjective sense. Guido M. Schuster, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 1996 | Detection and encoding of occluded areas in very low bit rate video codingabstractOne of the challenging problems that most existing video codecs face today is the encoding of the information pertaining to the occluded areas, i.e., the areas which are covered or uncovered by moving objects. The existing techniques necessitate the transmission of the position information of the occluded areas following the detection process, which can constitute a large overhead in bandwidth consumption. In addition, these detection techniques fail under noisy conditions. On the other hand, no effort was made to incorporate the spatio-temporal correlation that exists between motion fields of consecutive frames. A new method to detect the occluded areas is described. In this method, the decoder is given additional intelligence to extract the position information of the occluded areas, thus reducing the bandwidth needed to transmit the occlusion information significantly. According to the proposed method, the temporal correlation of the motion fields is exploited. The proposed method is robust under noisy conditions and provides a computationally simple solution. Taner Özcelik, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 1996 | A video compression scheme with optimal bit allocation between displacement vector field and displaced frame differenceabstractWe address the fundamental problem of optimally splitting a video sequence into two sources of information, the displaced frame difference (DFD) and the displacement vector field (DVF). We first consider the case of a lossless motion compensated video coder (MCVC) and derive a general dynamic programming (DP) formulation which results in an optimal tradeoff between the DVF and the DFD. We then consider the more important case of a lossy MCVC and present an algorithm which solves optimally the bit allocation between the rate and the distortion. This algorithm is based on Lagrangian relaxation and the DP approach introduced for the lossless MCVC. We then present an H.263-based MCVC which uses the proposed optimal bit allocation scheme and compare its results to H.263. As expected, the proposed coder is superior in the rate-distortion sense. Guido M. Schuster, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 1996 | Recursive map displacement field estimation and its applicationsabstractWe briefly describe some of our work on the use of stochastic models to describe the displacement vector field (DVF) in an image sequence. Specifically, autoregressive models are used which describe the abrupt transitions in the DVF with the use of a line process, but also result in spatio-temporally recursive structures. The use of such models in developing maximum a posteriori estimators for the DVF and the line process is subsequently described. Finally, the extension and application of the resulting estimator to the problems of object tracking, video compression and restoration of video sequences is reviewed. James C. Brailean, Aggelos K. Katsaggelos |
ICIP (1) | 2 |
| 1996 | Restoration of severely blurred high range images using compound modelsabstractWe examine the use of compound Gauss Markov random fields (CGMRF) to restore severely blurred high range images. For this deblurring problem, the convergence of the simulated annealing (SA) and iterative conditional mode (ICM) algorithms has not been established. We propose two new iterative restoration algorithms which extend the classical SA and ICM approaches. Their convergence is established and they are tested on real and synthetic images. Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos, Javier Abad |
ICIP (2) | 2 |
| 1996 | An efficient boundary encoding scheme which is optimal in the rate-distortion senseabstractIn this paper, we present a fast and optimal method for the lossy encoding of object boundaries which are given as 8-connect chain codes. We approximate the boundary by a polygon and consider the problem of finding the polygon which can be encoded with the smallest number of bits for a given maximum distortion. The presented scheme is an extension of the approaches introduced in [l, 2], in that the admissible polygon vertices belong to a band around the original boundary and a new vertex encoding scheme is proposed. Guido M. Schuster, Aggelos K. Katsaggelos |
ICIP (2) | 2 |
| 1996 | Resolution enhancement of video sequences using motion compensationabstractImproving the spatial resolution of an image sequence critically depends upon the accuracy of the motion estimator. The problem is complicated by the fact that the motion field is prone to significant errors since original high resolution images are not available. In earlier work, bilinearly interpolated images were used as initial conditions for the proposed algorithm. In this paper, the use of other initial conditions, such as previously estimated images, is experimentally investigated. Furthermore, various forms of the iterative video resolution enhancement algorithm are studied experimentally. Brian C. Tom, Aggelos K. Katsaggelos |
ICIP (1) | 2 |
| 1996 | Regularized blur-assisted displacement field estimationabstractDue to the finite acquisition time of practical cameras, objects can move during image acquisition, therefore introducing motion blur degradations. Traditionally, these degradations are treated as undesirable artifacts that should be removed before further processing. In this work, we consider the use of motion blur as an indication of scene motion. ,Ve present two robust regu- larized motion estimation algorithns that consider the use of (motion) blur in their formulation. The first algorithm uses motion blur as prior knowledge for the estimation of the motion field. The second algorithm considers the joint estimation of the motion and motion blur. Each approach results in a motion blur point spread field, a motion field and a restored image in an approach that is different from previous work. Preliminarv results are presented. Damon L. Tull, Aggelos K. Katsaggelos |
ICIP (3) | 2 |
| 1996 | Multichannel Regularized Iterative Restoration of Motion Compensated Image SequencesabstractRestoration of image sequences is an important problem that can be encountered in many image processing applications, such as visual communications, robot guidance, and target tracking. The independent restoration of each frame in an image sequence is a suboptimal approach because the between-frame correlations are not explicitly taken into consideration. In this paper we address this problem by proposing a multichannel restoration approach. The multiple time-frames (channels) of the image sequence are restored simultaneously by using a multichannel regularized least-squares formulation of the problem. The regularization operator captures both within- and between-frame (channel) properties of the image sequence with the explicit use of the displacement vector field. We propose a number of different approaches to obtain the multichannel regularization operator, as well as an algorithm to iteratively compute the restored images. We present experiments that demonstrate the value of the proposed multichannel approach. Mungi Choi, Nikolas P. Galatsanos, Aggelos K. Katsaggelos |
J. Vis. Commun. Image Represent. | 3 |
| 1996 | Spatially adaptive wavelet-based multiscale image restorationabstractIn this paper, we present a new spatially adaptive approach to the restoration of noisy blurred images, which is particularly effective at producing sharp deconvolution while suppressing the noise in the flat regions of an image. This is accomplished through a multiscale Kalman smoothing filter applied to a prefiltered observed image in the discrete, separable, 2-D wavelet domain. The prefiltering step involves constrained least-squares filtering based on optimal choices for the regularization parameter. This leads to a reduction in the support of the required state vectors of the multiscale restoration filter in the wavelet domain and improvement in the computational efficiency of the multiscale filter. The proposed method has the benefit that the majority of the regularization, or noise suppression, of the restoration is accomplished by the efficient multiscale filtering of wavelet detail coefficients ordered on quadtrees. Not only does this lead to potential parallel implementation schemes, but it permits adaptivity to the local edge information in the image. In particular, this method changes filter parameters depending on scale, local signal-to-noise ratio (SNR), and orientation. Because the wavelet detail coefficients are a manifestation of the multiscale edge information in an image, this algorithm may be viewed as an "edge-adaptive" multiscale restoration approach. Mark R. Banham, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 1995 | Noise robust spatial gradient estimation for use in displacement estimationabstractAn important component of any spatial temporal gradient motion estimation algorithm is the accuracy by which spatial gradients are calculated. When an image sequence is corrupted by noise, the problem of determining these spatial gradients becomes extremely difficult. This is immediately apparent, since the magnitude response of the derivative operator is |/spl omega/|. In other words, the components of an image are amplified upon differentiation in proportion to their frequency value. Thus, high-frequency noise terms will dominate any low-frequency features in the differentiated image. If this corrupted differentiated image is then used within a spatio-temporal gradient motion estimator, the noise will erroneously influence the estimated motion vector. The problem of estimating the spatial gradient is treated as an inverse problem with noise. Formulating the problem in this manner results in a recursive gradient estimator that suppresses the effects of noise. James C. Brailean, Aggelos K. Katsaggelos |
ICIP | 2 |
| 1995 | Motion field prediction and restoration for low bit-rate video codingabstractMotion vector field (MVF) prediction methods are presented followed by a restoration method. These methods combined with a proposed motion compensated (MC) video coding scheme are suitable for low bit rate transmission. An expression is derived for the initial estimate of the working MVF based on the preceding MVF. Spatio-temporally adaptive regularization is applied using neighborhood information. The output MVF is used as the initial prediction estimate for a Kalman MVF restoration approach. By applying this method to both the encoder and decoder, the resulting MC MVF and image intensity temporal updates are coded and transmitted. The restoration method produces accurate estimates of the MVF, thus resulting in a significant transmission cost reduction. Experiments with standard video-conference image sequences demonstrate the improved performance of the proposed scheme. Serafim N. Efstratiadis, Michael G. Strintzis, Aggelos K. Katsaggelos |
ICIP | 3 |
| 1995 | Reconstruction of a high-resolution image by simultaneous registration, restoration, and interpolation of low-resolution imagesabstractIn this paper a solution is provided to the problem of obtaining a high resolution image from several low resolution images that have been subsampled and displaced by different amounts of sub-pixel shifts. In its most general form, this problem can be broken up into three sub-problems: registration, restoration, and interpolation. Previous work has either solved all three sub-problems independently, or more recently, solved either the first two steps (registration and restoration) or the last two steps together. However, none of the existing methods solve all three sub-problems simultaneously. This paper poses the low resolution to high resolution problem as a maximum likelihood (ML) problem which is solved by the expectation-maximization (EM) algorithm. By exploiting the structure of the matrices involved, the problem ran be solved in the discrete frequency domain. The ML problem is then the estimation of the sub-pixel shifts, the noise variances of each image, the power spectra of the high resolution image, and the high resolution image itself. Experimental results are shown which demonstrate the effectiveness of this approach. Brian C. Tom, Aggelos K. Katsaggelos |
ICIP | 2 |