EDBT 2026 Demo / reviewers in the wild / expert
A. Murat Tekalp
dblp:00/5029 · also Ahmet Murat Tekalp
· DBLP profile ↗
283ranked-venue papers
17as first author
26since 2021 · last 2026
0000-0003-1465-8121ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 257 · 13 first-author · 25 since 2021Artificial intelligence and machine learning · 18 · 2 since 2021Computer networks · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OracleGS: Grounding Generative Priors for Sparse-View Gaussian SplattingabstractSparse-view novel view synthesis is fundamentally ill-posed due to severe geometric ambiguity. Current methods are caught in a trade-off: regressive models are geometrically faithful but incomplete, whereas generative models can complete scenes but often introduce structural inconsistencies. We propose OracleGS, a novel framework that reconciles generative completeness with regressive fidelity for sparse view Gaussian Splatting. Instead of using generative models to patch incomplete reconstructions, our "propose-and-validate" framework first leverages a pre-trained 3D-aware diffusion model to synthesize novel views to propose a complete scene. We then repurpose a multi-view stereo (MVS) model as a 3D-aware oracle to validate the 3D uncertainties of generated views, using its attention maps to reveal regions where the generated views are well-supported by multi-view evidence versus where they fall into regions of high uncertainty due to occlusion, lack of texture, or direct inconsistency. This uncertainty signal directly guides the optimization of a 3D Gaussian Splatting model via an uncertainty-weighted loss. Our approach conditions the powerful generative prior on multi-view geometric evidence, filtering hallucinatory artifacts while preserving plausible completions in under-constrained regions, outperforming state-of-the-art methods on datasets including Mip-NeRF 360 and NeRF Synthetic. Atakan Topaloglu, Kunyi Li, Michael Niemeyer, Nassir Navab, A. Murat Tekalp, Federico Tombari |
WACV | 5 |
| 2026 | Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion ModelsabstractSuper-resolution (SR) is an ill-posed inverse problem with many feasible solutions that are consistent with a given low-resolution image. On one hand, regressive SR models aim to balance fidelity and perceptual quality to yield a single solution; but this trade-off often leads to artifacts that introduce ambiguity in information-critical applications such as identifying digits or letters. On the other hand, diffusion models generate a diverse set of SR images; but now selecting the most trustworthy solution out of this set becomes a challenge. This paper introduces a robust, automated framework for identifying the most trustworthy SR sample from a diffusion-generated set by leveraging the semantic reasoning capabilities of vision-language models (VLMs). Specifically, VLMs such as BLIP-2, GPT-4o, and their variants are prompted with structured queries to evaluate semantic correctness, visual quality, and the presence of artifacts. The top-ranked SR candidates are then ensembled to yield a single trustworthy output in a cost-effective manner. To rigorously assess the validity of VLM-selected samples, we propose a novel Trustworthiness Score (TWS)—a hybrid metric that quantifies SR reliability based on three complementary components: semantic similarity using CLIP embeddings, structural integrity via SSIM on edge maps, and artifact sensitivity measured through a multi-level wavelet decomposition. We empirically demonstrate that TWS correlates strongly with human preference in both ambiguous and natural images, and that VLM-guided selections consistently yield high TWS values. Compared to conventional metrics like PSNR, LPIPS, and DISTS—which fail to reflect information fidelity—our approach offers a principled, scalable, and generalizable solution for navigating the uncertainty of the diffusion SR space. By aligning model outputs with human expectations and semantic correctness, this work sets a new benchmark for trustworthiness in generative SR tasks. Cansu Korkmaz, A. Murat Tekalp, Zafer Dogan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Padé Neurons for Efficient Neural ModelsabstractNeural networks commonly employ the McCulloch-Pitts neuron model, which is a linear model followed by a point-wise non-linear activation. Various researchers have already advanced inherently non-linear neuron models, such as quadratic neurons, generalized operational neurons, generative neurons, and super neurons, which offer stronger non-linearity compared to point-wise activation functions. In this paper, we introduce a novel and better non-linear neuron model called Padé neurons ( $\mathrm {\textit {Paon}}$ s), inspired by Padé approximants. $\mathrm {\textit {Paon}}$ s offer several advantages, such as diversity of non-linearity, since each $\mathrm {\textit {Paon}}$ learns a different non-linear function of its inputs, and layer efficiency, since $\mathrm {\textit {Paon}}$ s provide stronger non-linearity in much fewer layers compared to piecewise linear approximation. Furthermore, $\mathrm {\textit {Paon}}$ s include all previously proposed neuron models as special cases, thus any neuron model in any network can be replaced by $\mathrm {\textit {Paon}}$ s. We note that there has been a proposal to employ the Padé approximation as a generalized point-wise activation function, which is fundamentally different from our model. To validate the efficacy of $\mathrm {\textit {Paon}}$ s, in our experiments, we replace classic neurons in some well-known neural image super-resolution, compression, and classification models based on the ResNet architecture with $\mathrm {\textit {Paon}}$ s. Our comprehensive experimental results and analyses demonstrate that neural models built by $\mathrm {\textit {Paon}}$ s provide better or equal performance than their classic counterparts with a smaller number of layers. The PyTorch implementation code for $\mathrm {\textit {Paon}}$ is open-sourced at https://github.com/onur-keles/Paon. Onur Keles, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 2025 | Distributed virtual selective-forwarding units and SDN-assisted edge computing for optimization of multi-party WebRTC videoconferencing
Riza Arda Kirmizioglu, A. Murat Tekalp, Burak Gorkemli |
Signal Process. Image Commun. | 2 |
| 2024 | Training Generative Image Super-Resolution Models by Wavelet-Domain Losses Enables Better Control of ArtifactsabstractSuper-resolution (SR) is an ill-posed inverse problem, where the size of the set of feasible solutions that are consistent with a given low-resolution image is very large. Many algorithms have been proposed to find a “good” solution among the feasible solutions that strike a balance between fidelity and perceptual quality. Unfortunately, all known methods generate artifacts and hallucinations while trying to reconstruct high-frequency (HF) image details. A fundamental question is: Can a model learn to distinguish genuine image details from artifacts? Although some recent works focused on the differentiation of details and artifacts, this is a very challenging problem and a satisfactory solution is yet to be found. This paper shows that the characterization of genuine HF details versus artifacts can be better learned by training GAN-based SR models using wavelet-domain loss functions compared to RGB-domain or Fourier-space losses. Although wavelet-domain losses have been used in the literature before, they have not been used in the context of the SR task. More specifically, we train the discriminator only on the HF wavelet sub-bands instead of on RGB images and the generator is trained by a fidelity loss over wavelet subbands to make it sensitive to the scale and orientation of structures. Extensive experimental results demonstrate that our model achieves better perception-distortion trade-off according to multiple objective measures and visual evaluations. Cansu Korkmaz, A. Murat Tekalp, Zafer Dogan |
CVPR | 2 |
| 2024 | Saliency-Aware End-to-End Learned Variable-Bitrate 360-Degree Image CompressionabstractEffective compression of 360° images, also referred to as omnidirectional images (ODIs), is of high interest for various virtual reality (VR) and related applications. 2D image compression methods ignore the equator-biased nature of ODIs and fail to address oversampling near the poles, leading to inefficient compression when applied to ODI. We present a new learned saliency-aware 360° image compression architecture that prioritizes bit allocation to more significant regions, considering the unique properties of ODIs. By assigning fewer bits to less important regions, significant data size reduction can be achieved while maintaining high visual quality in the significant regions. To the best of our knowledge, this is the first study that proposes an end-to-end variable-rate model to compress 360° images leveraging saliency information. The results show significant bit-rate savings over the state-of-the-art learned and traditional ODI compression methods at similar perceptual visual quality. Supplementary materials are available at [Supplementary URL]. Oguzhan Güngördü, A. Murat Tekalp |
ICIP | 2 |
| 2024 | Paon: A New Neuron Model Using Padé ApproximantsabstractConvolutional neural networks (CNN) are built upon the classical McCulloch-Pitts neuron model, which is essentially a linear model, where the nonlinearity is provided by a separate activation function. Several researchers have proposed enhanced neuron models, including quadratic neurons, generalized operational neurons, generative neurons, and super neurons, with stronger nonlinearity than that provided by the pointwise activation function. There has also been a proposal to use Padé approximation as a generalized activation function. In this paper, we introduce a brand new neuron model called Padé neurons (Paons1), inspired by the Padé approximants, which is the best mathematical approximation of a transcendental function as a ratio of polynomials with different orders. We show that Paons are a super set of all other proposed neuron models. Hence, the basic neuron in any known CNN model can be replaced by Paons. In this paper, we extend the well-known ResNet to PadeNet (built by Paons) to demonstrate the concept. Our experiments on the single-image super-resolution task show that PadeNets can obtain better results than competing architectures.1https://github.com/onur-keles/Paon Onur Keles, A. Murat Tekalp |
ICIP | 2 |
| 2024 | Trustworthy Sr: Resolving Ambiguity In Image Super-Resolution Via Diffusion Models And Human FeedbackabstractSuper-resolution (SR) is an ill-posed inverse problem with a large set of feasible solutions that are consistent with a given low-resolution image. Various deterministic algorithms aim to find a single solution that balances fidelity and perceptual quality; however, this trade-off often causes visual artifacts that bring ambiguity in information-centric applications. On the other hand, diffusion models (DMs) excel in generating a diverse set of feasible SR images that span the solution space. The challenge is then how to determine the most likely solution among this set in a trustworthy manner. We observe that quantitative measures, such as PSNR, LPIPS, DISTS, are not reliable indicators to resolve ambiguous cases. To this effect, we propose employing human feedback, where we ask human subjects to select a small number of likely samples and we ensemble the averages of selected samples. This strategy leverages the high-quality image generation capabilities of DMs, while recognizing the importance of obtaining a single trustworthy solution, especially in use cases, such as identification of specific digits or letters, where generating multiple feasible solutions may not lead to a reliable outcome. Experimental results demonstrate that our proposed strategy provides more trustworthy solutions when compared to state-of-the art SR methods. Cansu Korkmaz, Ege Çirakman, A. Murat Tekalp, Zafer Dogan |
ICIP | 3 |
| 2024 | Motion-Adaptive Inference for Flexible Learned B-Frame CompressionabstractWhile the performance of recent learned intra and sequential video compression models exceed that of respective traditional codecs, the performance of learned B-frame compression models generally lag behind traditional B-frame coding. The performance gap is bigger for complex scenes with large motions. This is related to the fact that the distance between the past and future references vary in hierarchical B-frame compression depending on the level of hierarchy, which causes motion range to vary. The inability of a single B-frame compression model to adapt to various motion ranges causes loss of performance. As a remedy, we propose controlling the motion range for flow prediction during inference (to approximately match the range of motions in the training data) by downsampling video frames adaptively according to amount of motion and level of hierarchy in order to compress all B-frames using a single flexible-rate model. We present state-of-the-art BD rate results to demonstrate the superiority of our proposed single-model motion-adaptive inference approach to all existing learned B-frame compression models.1.1The models and instructions to reproduce our results will be released at https://github.com/KUIS-AI-Tekalp-Research-Group/video-compression/tree/master/ICIP2024 Mustafa Akin Yilmaz, O. Ugur Ulas, Ahmet Bilican, A. Murat Tekalp |
ICIP | 4 |
| 2024 | A new multi-picture architecture for learned video deinterlacing and demosaicing with parallel deformable convolution and self-attention blocks
Ronglei Ji, A. Murat Tekalp |
Image Vis. Comput. | 2 |
| 2023 | Spatio-Temporal Perception-Distortion Trade-Off in Learned Video SRabstractPerception-distortion trade-off is well-understood for single-image super-resolution. However, its extension to video super-resolution (VSR) is not straightforward, since popular perceptual measures only evaluate naturalness of spatial textures and do not take naturalness of flow (temporal coherence) into account. To this effect, we propose a new measure of spatio-temporal perceptual video quality emphasizing naturalness of optical flow via the perceptual straightness hypothesis (PSH) for meaningful spatio-temporal perception-distortion trade-off. We also propose a new architecture for perceptual VSR (PSVR) to explicitly enforce naturalness of flow to achieve realistic spatio-temporal perception-distortion trade-off according to the proposed measures. Experimental results with PVSR support the hypothesis that a meaningful perception-distortion tradeoff for video should account for the naturalness of motion in addition to naturalness of texture. Nasrin Rahimi, A. Murat Tekalp |
ICIP | 2 |
| 2023 | Multi-Scale Deformable Alignment and Content-Adaptive Inference for Flexible-Rate Bi-Directional Video CompressionabstractThe lack of ability to adapt the motion compensation model to video content is an important limitation of current end-to-end learned video compression models. This paper advances the state-of-the-art by proposing an adaptive motion-compensation model for end-to-end rate-distortion optimized hierarchical bi-directional video compression. In particular, we propose two novelties: i) a multi-scale deformable alignment scheme at the feature level combined with multi-scale conditional coding, ii) motion-content adaptive inference. In addition, we employ a gain unit, which enables a single model to operate at multiple rate-distortion operating points. We also exploit the gain unit to control bit allocation among intra-coded vs. bi-directionally coded frames by fine tuning corresponding models for truly flexible-rate learned video coding. Experimental results demonstrate state-of-the-art rate-distortion performance exceeding those of all prior art in learned video coding1. Mustafa Akin Yilmaz, O. Ugur Ulas, A. Murat Tekalp |
ICIP | 3 |
| 2022 | Flexible-Rate Learned Hierarchical Bi-Directional Video Compression with Motion Refinement and Frame-Level Bit AllocationabstractThis paper presents improvements and novel additions to our recent work on end-to-end optimized hierarchical bidirectional video compression [1] to further advance the state-of-the-art in learned video compression. As an improvement, we combine motion estimation and prediction modules and compress refined residual motion vectors for improved rate-distortion performance. As novel addition, we adapted the gain unit proposed for image compression to flexible-rate video compression in two ways: first, the gain unit enables a single encoder model to operate at multiple rate-distortion operating points; second, we exploit the gain unit to control bit allocation among intra-coded vs. bi-directionally coded frames by fine tuning corresponding models for truly flexible-rate learned video coding. Experimental results demonstrate that we obtain state-of-the-art rate-distortion performance exceeding those of all prior art in learned video coding. Eren Çetin, Mustafa Akin Yilmaz, A. Murat Tekalp |
ICIP | 3 |
| 2022 | Multi-Field De-Interlacing Using Deformable Convolution Residual Blocks and Self-AttentionabstractAlthough deep learning has made significant impact on image/video restoration and super-resolution, learned deinterlacing has so far received less attention in academia or industry. This is despite deinterlacing is well-suited for supervised learning from synthetic data since the degradation model is known and fixed. In this paper, we propose a novel multi-field full frame-rate deinterlacing network, which adapts the state-of-the-art superresolution approaches to the deinterlacing task. Our model aligns features from adjacent fields to a reference field (to be deinterlaced) using both deformable convolution residual blocks and self attention. Our extensive experimental results demonstrate that the proposed method provides state-of-the-art deinterlacing results in terms of both numerical and perceptual performance. At the time of writing, our model ranks first in the Full FrameRate LeaderBoard at https://videoprocessing.ai/benchmarks/deinterlacer.html Ronglei Ji, A. Murat Tekalp |
ICIP | 2 |
| 2022 | MMSR: Multiple-Model Learned Image Super-Resolution Benefiting from Class-Specific Image PriorsabstractAssuming a known degradation model, the performance of a learned image super-resolution (SR) model depends on how well the variety of image characteristics within the training set matches those in the test set. As a result, the performance of an SR model varies noticeably from image to image over a test set depending on whether characteristics of specific images are similar to those in the training set or not. Hence, in general, a single SR model cannot generalize well enough for all types of image content. In this work, we show that training multiple SR models for different classes of images (e.g., for text, texture, etc.) to exploit class-specific image priors and employing a post-processing network that learns how to best fuse the outputs produced by these multiple SR models surpasses the performance of state-of-the-art generic SR models. Experimental results clearly demonstrate that the proposed multiple-model SR (MMSR) approach significantly outperforms a single pre-trained state-of-the-art SR model both quantitatively and visually. It even exceeds the performance of the best single class-specific SR model trained on similar text or texture images. Cansu Korkmaz, A. Murat Tekalp, Zafer Dogan |
ICIP | 2 |
| 2022 | Perception-Distortion Trade-Off in the SR Space Spanned by Flow ModelsabstractFlow-based generative super-resolution (SR) models learn to produce a diverse set of feasible SR solutions, called the SR space. Diversity of SR solutions increases with the temperature (τ) of latent variables, which introduces random variations of texture among sample solutions, resulting in visual artifacts and low fidelity. In this paper, we present a simple but effective image ensembling/fusion approach to obtain a single SR image eliminating random artifacts and improving fidelity without significantly compromising perceptual quality. We achieve this by benefiting from a diverse set of feasible photorealistic solutions in the SR space spanned by flow models. We propose different image ensembling and fusion strategies which offer multiple paths to move sample solutions in the SR space to more desired destinations in the perception-distortion plane in a controllable manner depending on the fidelity vs. perceptual quality requirements of the task at hand. Experimental results demonstrate that our image ensembling/fusion strategy achieves more promising perception-distortion trade-off compared to sample SR images produced by flow models and adversarially trained models in terms of both quantitative metrics and visual quality. Cansu Korkmaz, A. Murat Tekalp, Zafer Dogan, Erkut Erdem, Aykut Erdem |
ICIP | 2 |
| 2022 | Flexible luma-chroma bit allocation in learned image compression for high-fidelity sharper imagesabstractHigh-fidelity learned image/video compression solutions are typically optimized with respect to l1 or l2 loss in RGB 444 format and evaluated by RGB PSNR. It is well-known that optimization of a fidelity criterion results in blurry images, which is typically alleviated by adding a content-based and/or adversarial loss terms. However, such conditional generative models result in loss of fidelity. In this paper, we propose a simple solution to obtain sharper images without losing fidelity based on learned flexible-rate coding using gained variational auto-encoder (gained-VAE) in the luma-chroma (YCrCb 444) domain. This allows us to implement image-adaptive luma-chroma bit allocation during inference, i.e., to increase Y PSNR at the expense of slightly lower chroma PSNR to obtain sharper images without introducing color artifacts based on the observation that Y PSNR correlates with image sharpness better than RGB PSNR. We note that the proposed inference-time image-adaptive luma-chroma bit allocation strategy can be incorporated into any VAE-based image compression model. Experimental results show that sharper images with better VMAF and Y PSNR can be obtained by optimizing models for YCrCb MSE with the proposed image-adaptive luma-chroma bit/quality allocation compared to state-of-the-art models optimizing RGB MSE at the same bpp. O. Ugur Ulas, A. Murat Tekalp |
PCS | 2 |
| 2022 | End-to-End Rate-Distortion Optimized Learned Hierarchical Bi-Directional Video CompressionabstractConventional video compression (VC) methods are based on motion compensated transform coding, and the steps of motion estimation, mode and quantization parameter selection, and entropy coding are optimized individually due to the combinatorial nature of the end-to-end optimization problem. Learned VC allows end-to-end rate-distortion (R-D) optimized training of nonlinear transform, motion and entropy model simultaneously. Most works on learned VC consider end-to-end optimization of a sequential video codec based on R-D loss averaged over pairs of successive frames. It is well-known in conventional VC that hierarchical, bi-directional coding outperforms sequential compression because of its ability to use both past and future reference frames. This paper proposes a learned hierarchical bi-directional video codec (LHBDC) that combines the benefits of hierarchical motion-compensated prediction and end-to-end optimization. Experimental results show that we achieve the best R-D results that are reported for learned VC schemes to date in both PSNR and MS-SSIM. Compared to conventional video codecs, the R-D performance of our end-to-end optimized codec outperforms those of both x265 and SVT-HEVC encoders ("veryslow" preset) in PSNR and MS-SSIM as well as HM 16.23 reference software in MS-SSIM. We present ablation studies showing performance gains due to proposed novel tools such as learned masking, flow-field subsampling, and temporal flow vector prediction. The models and instructions to reproduce our results can be found in https://github.com/makinyilmaz/LHBDC/. Mustafa Akin Yilmaz, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 2021 | Self-Organized Residual Blocks For Image Super-ResolutionabstractIt has become a standard practice to use the convolutional networks (ConvNet) with RELU non-linearity in image restoration and super-resolution (SR). Although the universal approximation theorem states that a multi-layer neural network can approximate any non-linear function with the desired precision, it does not reveal the best network architecture to do so. Recently, operational neural networks (ONNs) that choose the best non-linearity from a set of alternatives, and their “self-organized” variants (Self-ONN) that approximate any non-linearity via Taylor series have been proposed to address the well-known limitations and drawbacks of conventional ConvNets such as network homogeneity using only the McCulloch-Pitts neuron model. In this paper, we propose the concept of self-organized operational residual (SOR) blocks, and present hybrid network architectures combining regular residual and SOR blocks to strike a balance between the benefits of stronger non-linearity and the overall number of parameters. The experimental results demonstrate that the proposed architectures yield performance improvements in both PSNR and perceptual metrics. Onur Keles, A. Murat Tekalp, Junaid Malik, Serkan Kiranyaz |
ICIP | 2 |
| 2021 | Two-Stage Domain Adapted Training For Better Generalization In Real-World Image Restoration And Super-ResolutionabstractIt is well-known that in inverse problems, end-to-end trained networks overfit the degradation model seen in the training set, i.e., they do not generalize to other types of degradations well. Recently, an approach to first map images downsampled by unknown filters to bicubicly downsampled look-alike images was proposed to successfully super-resolve such images. In this paper, we show that any inverse problem can be formulated by first mapping the input degraded images to an intermediate domain, and then training a second network to form output images from these intermediate images. Furthermore, the best intermediate domain may vary according to the task. Our experimental results demonstrate that this two-stage domain-adapted training strategy does not only achieve better results on a given class of unknown degradations but can also generalize to other unseen classes of degradations better. Cansu Korkmaz, A. Murat Tekalp, Zafer Dogan |
ICIP | 2 |
| 2021 | Self-Organized Variational Autoencoders (Self-Vae) For Learned Image CompressionabstractIn end-to-end optimized learned image compression, it is standard practice to use a convolutional variational autoencoder with generalized divisive normalization (GDN) to transform images into a latent space. Recently, Operational Neural Networks (ONNs) that learn the best non-linearity from a set of alternatives, and their “self-organized” variants, Self-ONNs, that approximate any non-linearity via Taylor series have been proposed to address the limitations of convolutional layers and a fixed nonlinear activation. In this paper, we propose to replace the convolutional and GDN layers in the variational autoencoder with self-organized operational layers, and propose a novel self-organized variational autoencoder (Self-VAE) architecture that benefits from stronger non-linearity. The experimental results demonstrate that the proposed Self-VAE yields improvements in both rate-distortion performance and perceptual image quality. Mustafa Akin Yilmaz, Onur Keles, Hilal Güven, A. Murat Tekalp, Junaid Malik, Serkan Kiranyaz |
ICIP | 4 |
| 2021 | DFPN: Deformable Frame Prediction NetworkabstractLearned frame prediction is a current problem of interest in computer vision and video processing/compression. Although several deep network architectures have been proposed for learned frame prediction, to the best of our knowledge, there is no work based on using deformable convolutions for frame prediction. To this effect, we propose a deformable frame prediction network (DFPN) for task-oriented implicit motion modeling and next frame prediction. Experimental results demonstrate that the proposed DFPN model achieves state of the art results in next frame prediction in sequences with global motion. Our models and results are available at https://github.com/makinyilmaz/DFPN. Mustafa Akin Yilmaz, A. Murat Tekalp |
ICIP | 2 |
| 2021 | On the Computation of PSNR for a Set of Images or VideoabstractWhen comparing learned image/video restoration and compression methods, it is common to report peak-signal to noise ratio (PSNR) results. However, there does not exist a generally agreed upon practice to compute PSNR for sets of images or video. Some authors report average of individual image/frame PSNR, which is equivalent to computing a single PSNR from the geometric mean of individual image/frame mean-square error (MSE). Others compute a single PSNR from the arithmetic mean of frame MSEs for each video. Furthermore, some compute the MSE/PSNR of Y-channel only, while others compute MSE/PSNR for RGB channels. This paper investigates different approaches to computing PSNR for sets of images, single video, and sets of video and the relation between them. We show the difference between computing the PSNR based on arithmetic vs. geometric mean of MSE depends on the distribution of MSE over the set of images or video, and that this distribution is task-dependent. In particular, these two methods yield larger differences in restoration problems, where the MSE is exponentially distributed and smaller differences in compression problems, where the MSE distribution is narrower. We hope this paper will motivate the community to clearly describe how they compute reported PSNR values to enable consistent comparison. Onur Keles, Mustafa Akin Yilmaz, A. Murat Tekalp, Cansu Korkmaz, Zafer Dogan |
PCS | 3 |
| 2021 | A Practical Approach for Rate-Distortion-Perception Analysis in Learned Image CompressionabstractRate-distortion optimization (RDO) of codecs, where distortion is quantified by the mean-square error, has been a standard practice in image/video compression over the years. RDO serves well for optimization of codec performance for evaluation of the results in terms of PSNR. However, it is well known that the PSNR does not correlate well with perceptual evaluation of images; hence, RDO is not well suited for perceptual optimization of codecs. Recently, rate-distortion-perception trade-off has been formalized by taking the Kullback-Leibler (KL) divergence between the distributions of the original and reconstructed images as a perception measure. Learned image compression methods that simultaneously optimize rate, mean-square loss, VGG loss, and an adversarial loss were proposed. Yet, there exists no easy approach to fix the rate, distortion or perception at a desired level in a practical learned image compression solution to perform an analysis of the trade-off between rate, distortion and perception measures. In this paper, we propose a practical approach to fix the rate to carry out perception-distortion analysis at a fixed rate in order to perform perceptual evaluation of image compression results in a principled manner. Experimental results provide several insights for practical rate-distortion-perception analysis in learned image compression. Ogun Kirmemis, A. Murat Tekalp |
PCS | 2 |
| 2021 | Learned Multi-Field De-Interlacing with Feature Alignment via Deformable Residual Convolution BlocksabstractDeinterlacing continues to be an important problem of interest since many digital TV broadcasts and catalog content are still in interlaced format. Although deep learning has had huge impact in all forms of image/video processing, learned deinterlacing has not received much attention in the industry or academia. In this paper, we propose a novel multi-field deinterlacing network that aligns features from adjacent fields to a reference field (to be deinterlaced) using deformable residual convolution blocks. To the best of our knowledge, this paper is the first to propose fusion of multi-field features that are aligned via deformable convolutions for deinterlacing. We demonstrate through extensive experimental results that the proposed method provides state-of-the-art deinterlacing results in terms of both PSNR and perceptual quality. Ronglei Ji, A. Murat Tekalp |
VCIP | 2 |
| 2021 | Controlling P2P-CDN Live Streaming Services at SDN-Enabled Multi-Access Edge DatacentersabstractRecognizing the shortcomings of current hybrid peer-to-peer (P2P) content-distribution network (CDN) video solutions and the potential of emerging multi-access edge datacenters, we propose a novel P2P-CDN service model that is hosted at software defined networks (SDN)-enabled multi-access edge datacenters operated by network service providers (NSP). An important feature of the proposed service architecture is that both CDN access by peers and P2P video streaming between peers within edge access networks are fully controlled by cooperation of the video content provider (VCP) and NSP to optimize video service key performance indicators (KPI). The proposed fully controlled P2P-CDN architecture with P2P group formation and chunk scheduling managed at edge datacenters reduces the load on CDN servers while overcoming quality of experience (QoE) fluctuations per flow and unfairness between multiple heterogeneous video-resolution clients over reserved access network slices. Other advantages of this service include: i) better video quality and lower delay for clients; ii) better use of edge network resources; iii) avoiding illegal, unauthorized P2P content sharing. To the best of our knowledge, there are no solutions in the literature that address P2P-CDN services managed at NSP-edge datacenters combining P2P-assisted CDN, SDN-assisted edge computing, and premium service over reserved slices. Experimental results show that the proposed P2P-CDN service deployed at SDN-enabled edge datacenters provides excellent service KPI compared to other state-of-the-art solutions. Selin Nacakli, A. Murat Tekalp |
IEEE Trans. Multim. | 2 |
| 2020 | Shrinkage as Activation for Learned Image CompressionabstractWith recent advances in learned entropy and context models, the rate-distortion performance of deep learned image compression methods reached or surpassed those of conventional codecs. However, learned image compression is currently more complex and slower than conventional image compression. Learned image and video compression methods almost exclusively employ the generalized divisive normalization (GDN) activation function. This paper investigates the effect of activation function on the performance of image compression in terms of both objective and subjective criteria as well as runtime. In particular, we show that the distribution of latents produced by hard shrinkage fits a Laplacian better, and it is possible to achieve similar rate-distortion and better visual performance using hard shrinkage with lower complexity. Ogun Kirmemis, A. Murat Tekalp |
ICIP | 2 |
| 2020 | End-to-End Rate-Distortion Optimization for Bi-Directional Learned Video CompressionabstractConventional video compression methods employ a linear transform and block motion model, and the steps of motion estimation, mode and quantization parameter selection, and entropy coding are optimized individually due to combinatorial nature of the end-to-end optimization problem. Learned video compression allows end-to-end rate-distortion optimized training of all nonlinear modules, quantization parameter and entropy model simultaneously. While previous work on learned video compression considered training a sequential video codec based on end-to-end optimization of cost averaged over pairs of successive frames, it is well-known in conventional video compression that hierarchical, bi-directional coding outperforms sequential compression. In this paper, we propose for the first time end-to-end optimization of a hierarchical, bi-directional motion compensated learned codec by accumulating cost function over fixed-size groups of pictures (GOP). Experimental results show that the rate-distortion performance of our proposed learned bi-directional GOP coder outperforms the state-of-the-art end-to-end optimized learned sequential compression as expected. Mustafa Akin Yilmaz, A. Murat Tekalp |
ICIP | 2 |
| 2020 | Realizing a Low-Power Head-Mounted Phase-Only Holographic Display by Light-Weight CompressionabstractHead-mounted holographic displays (HMHD) are projected to be the first commercial realization of holographic video display systems. HMHDs use liquid crystal on silicon (LCoS) spatial light modulators (SLM), which are best suited to display phase-only holograms (POH). The performance/watt requirement of a monochrome, 60 fps Full HD, 2-eye, POH HMHD system is about 10 TFLOPS/W, which is orders of magnitude higher than that is achievable by commercially available mobile processors. To mitigate this compute power constraint, display-ready POHs shall be generated on a nearby server and sent to the HMHD in compressed form over a wireless link. This paper discusses design of a feasible HMHD-based augmented reality system, focusing on compression requirements and per-pixel rate-distortion trade-off for transmission of display-ready POH from the server to HMHD. Since the decoder in the HMHD needs to operate on low power, only coding methods that have low-power decoder implementation are considered. Effects of 2D phase unwrapping and flat quantization on compression performance are also reported. We next propose a versatile PCM-POH codec with progressive quantization that can adapt to SLM-dynamic-range and available bitrate, and features per-pixel rate-distortion control to achieve acceptable POH quality at target rates of 60-200 Mbit/s that can be reliably achieved by current wireless technologies. Our results demonstrate feasibility of realizing a low-power, quality-ensured, multi-user, interactive HMHD augmented reality system with commercially available components using the proposed adaptive compression of display-ready POH with light-weight decoding. Burak Soner, Erdem Ulusoy, A. Murat Tekalp, Hakan Urey |
IEEE Trans. Image Process. | 3 |
| 2020 | Multi-Party WebRTC Services Using Delay and Bandwidth Aware SDN-Assisted IP Multicasting of Scalable Video Over 5G NetworksabstractAt present, multi-party WebRTC videoconferencing between peers with heterogenous network resources and terminals is enabled over the best-effort Internet using a central selective forwarding unit (SFU), where each peer sends a scalable encoded video stream to the SFU. This connection model avoids the upload bandwidth bottleneck associated with mesh connections; however, it increases peer delay and overall network load (resource consumption) in addition to requiring investment in servers since all video traffic must go through SFU servers. To this effect, we propose a new multi-party WebRTC service model over future 5G networks, where a video service provider (VSP) collaborates with a network service providers (NSP) to offer an NSP-managed service to stream scalable video layers using software-defined networking (SDN)-assisted Internet protocol (IP) multicasting between peers using NSP infrastructure. In the proposed service model, each peer sends a scalable coded video upstream, which is selectively duplicated and forwarded as layer streams at SDN switches in the network, instead of at a central SFU, in a multi-party WebRTC session managed by multicast trees maintained by the SDN controller. Experimental results show that the proposed SDN-assisted IP multicast service architecture is more efficient than the SFU model in terms of end-to-end service delay and overall network resource consumption, while avoiding peer upload bandwidth bottleneck and distributing traffic more evenly across the network. The proposed architecture enables efficient provisioning of premium managed WebRTC services over bandwidth-reserved SDN slices to provide videoconferencing experience with guaranteed video quality over 5G networks. Riza Arda Kirmizioglu, A. Murat Tekalp |
IEEE Trans. Multim. | 2 |
| 2019 | Effect of Architectures and Training Methods on the Performance of Learned Video Frame PredictionabstractWe analyze the performance of feedforward vs. recurrent neural network (RNN) architectures and associated training methods for learned frame prediction. To this effect, we trained a residual fully convolutional neural network (FCNN), a convolutional RNN (CRNN), and a convolutional long short-term memory (CLSTM) network for next frame prediction using the mean square loss. We performed both stateless and stateful training for recurrent networks. Experimental results show that the residual FCNN architecture performs the best in terms of peak signal to noise ratio (PSNR) at the expense of higher training and test (inference) computational complexity. The CRNN can be trained stably and very efficiently using the stateful truncated backpropagation through time procedure, and it requires an order of magnitude less inference runtime to achieve near real-time frame prediction with an acceptable performance. Mustafa Akin Yilmaz, A. Murat Tekalp |
ICIP | 2 |
| 2019 | SDN-enabled distributed open exchange: Dynamic QoS-path optimization in multi-operator services
K. Tolga Bagci, A. Murat Tekalp |
Comput. Networks | 2 |
| 2019 | Motion-Based Rate Adaptation in WebRTC Videoconferencing Using Scalable Video CodingabstractThis paper proposes methods for rate adaptation by motion-based spatial and temporal resolution selection in both mesh-connected and selective-forwarding-unit (SFU) connected WebRTC videoconferencing using scalable video coding. In the mesh-connected case, the proposed motion-adaptive spatial/temporal layer selection allows each peer to send video to different peers with different terminal types and network rates at different rates using a single encoder. In the SFU-connected case, motion-adaptive rate control is used both at peers to adapt to the network rate between the sending peer and SFU by spatio-temporal resolution adaptation and at the SFU by layer selection to adapt to the network rate between the SFU and receiving peer. Experimental results show that our proposed motion-based rate adaptation achieves better perceptual video quality with sufficiently high frame rates and lower quantization parameter for video with high motion; and high spatial resolution and lower quantization parameter for video with low motion compared to simple rate-distortion model-based layer selection that does not use motion complexity, at the same rate. Gonca Bakar, Riza Arda Kirmizioglu, A. Murat Tekalp |
IEEE Trans. Multim. | 3 |
| 2018 | Multi-Party Webrtc Videoconferencing Using Scalable Vp9 Video: From Best-Effort Over-The-Top To Managed Value-Added ServicesabstractWe propose architectures and implementations for WebRTC videoconferencing services using scalable VP9 video coding with motion-adaptive rate control as either a best-effort over-the-top service or as a managed value-added service with rate reservation over software-defined networks. In the best-effort service, clients perform motion-adaptive layer selection according to the available bandwidth in order to achieve the best overall visual video quality. In the value-added managed service, the service manager reserves bandwidth between mesh-connected clients according to rates agreed by them, and clients perform motion-adaptive layer selection to adapt their send rates to the bandwidths reserved between the end points. The proposed framework has been demonstrated to yield excellent results for both point-to-point two-party and mesh-connected multi-party videoconferencing. Riza Arda Kirmizioglu, B. Can Kaya, A. Murat Tekalp |
ICME | 3 |
| 2018 | Dynamic Control Plane for SDN at ScaleabstractAs SDN migrates to wide area networks and 5G core networks, a scalable, highly reliable, low latency distributed control plane becomes a key factor that differentiates operator solutions for network control and management. In order to meet the high reliability and low latency requirements under time-varying volume of control traffic, the distributed control plane, consisting of multiple controllers and a combination of out-of-band and in-band control channels, needs to be managed dynamically. To this effect, we propose a novel programmable distributed control plane architecture with a dynamically managed in-band control network, where in-band mode switches communicate with their controllers over a virtual overlay to the data plane with dynamic topology. We dynamically manage the number of controllers, switches, and control flows assigned to each controller as well as traffic over control channels achieving both controller and control traffic load-balancing. We introduce “control flow table” (rules embedded in the flow table of a switch to manage in-band control flows) in order to implement the proposed distributed dynamic control plane. We propose methods for off-loading congested controllers and congested in-band control channels using control flow tables. A validation test-bed and experimental results over multiple topologies are presented to demonstrate the scalability and performance improvements achieved by the proposed dynamic control plane management procedures when the controller CPU and/or availability or throughput of in-band control channels becomes bottlenecks. Burak Gorkemli, Sinan Tatlicioglu, A. Murat Tekalp, Seyhan Civanlar, Erhan Lokman |
IEEE J. Sel. Areas Commun. | 3 |
| 2018 | Dynamic Resource Allocation by Batch Optimization for Value-Added Video Services Over SDNabstractWe propose a video service architecture and a novel resource allocation optimization framework to enable network service providers (NSP) to offer value-added video services (VAVS) over software-defined networking including different service levels, service-level awareness of users, and associated business models. To this effect, we introduce a new batch-optimization framework, where resource (path, bitrate, and admission control) allocations for a small group of flows (consisting of new service requests and some existing ones) are performed simultaneously as the number of new service requests and network conditions vary. The optimization problem becomes NP-complete when path computations are jointly (re-)optimized as a group in order to accommodate all service requests to the extent possible, to best utilize entire network resources in a fair manner, and maximize network service provider's revenue. In order to compute dynamic resource allocations online, we propose a heuristic group-constrained-shortest path procedure that aims for a fair allocation of resources among a group of requests with the same service level, while maximizing the total NSP revenue. Experimental results demonstrate the feasibility of the proposed method for possible deployment by NSP to offer future VAVS, and that the proposed solution is close to the optimal solution, which is approximately computed using a divide-and conquer strategy, for varying network size and traffic load conditions. In particular, we show that processing service requests in batches significantly improves total revenue and fairness in congested mode of operation. K. Tolga Bagci, A. Murat Tekalp |
IEEE Trans. Multim. | 2 |
| 2017 | Distributed-collaborative managed dash video servicesabstractWe propose a new distributed-collaborative managed DASH video service architecture over software defined networks (SDN) that enables fair and stable video quality to heterogeneous resolution clients. The proposed service is managed by the video service provider (VSP) in collaboration with the network service provider (NSP), where groups of clients sharing a network slice with a reserved throughput collaborate with each other to compute their own fair-share bitrates. Our novel distributed service architecture allows each client to share its buffer status with other clients in the same collaboration group so that each client can estimate a group-buffer-status aware fair-share bitrate, enforce this rate by TCP receive-window size control over a network slice reserved for the group, and perform application-level DASH video rate adaptation that is consistent with this enforced fair bitrate. Experimental results show that the proposed collaborative video service outperforms the traditional competitive DASH clients in terms of (i) minimizing quality fluctuations per client, (ii) fairness among heterogeneous DASH clients, and (iii) maximizing the total goodput of reserved network slice. Kemal E. Sahin, K. Tolga Bagci, A. Murat Tekalp |
CNSM | 3 |
| 2017 | Motion-Based Adaptive Streaming in WebRTC Using Spatio-Temporal Scalable VP9 Video CodingabstractWebRTC has become a popular platform for real-time communications over the best-effort Internet. It employs the Google Congestion Control algorithm to obtain an estimate of the state of the network. The default configuration employs single-layer CBR video encoding given the available network rate with rate control achieved by varying the quantization parameter and video frame rate. Recently, some open-source WebRTC platforms provided support for VP9 encoding with spatial scalable encoding option. The main contribution of this paper is to incorporate motion- based spatial resolution adaptation for adaptive streaming rate control and evaluate the use of single-layer (non- scalable) VP9 encoding vs. motion-based mixed spatio- temporal scalable VP9 encoding in point-to-point RTC between two parties in the presence of network congestion. Our results show that, during intervals of high motion activity, spatial resolution reduction with sufficiently high frame rates and reasonable quantization parameter values yield more pleasing video quality compared to the standard rate control scheme employed in open-source WebRTC implementations, which uses only quantization parameter and frame rate for rate control. Gonca Bakar, Riza Arda Kirmizioglu, A. Murat Tekalp |
GLOBECOM | 3 |
| 2017 | Emerging 3-D Imaging and Display TechnologiesabstractWe have become an information-centric society vastly dependent on the collection, communication, and presentation of information. At any given moment, it is likely that we are in the vicinity of some form of a display as displays play a prominent role in a variety of devices and applications. Three-dimensional imaging and display technologies are important components for presentation and visualization of information and for creating real-world-like environments in communication. There are broad applications of 3-D imaging and display technologies in computers, communication, mobile devices, TV, video, entertainment, robotics, metrology, security and defense, healthcare, and medicine. Bahram Javidi, A. Murat Tekalp |
Proc. IEEE | 2 |
| 2017 | Adaptive Multiview Video Delivery Using Hybrid NetworkingabstractMultiview entertainment is the next step in 3D immersive media networking owing to its improved depth perception and free-viewpoint viewing capability whereby users can observe the scene from the desired viewpoint. This paper outlines a delivery system for multiview plus depth video, combining the broadcast and broadband networks. The digital video broadcast (DVB) network is used along with adaptive peer-to-peer (P2P) distribution over the Internet to deliver high-volume multimedia to users. The DVB network has been used to deliver part of the 3D service, owing to its robustness and wide availability, as a mechanism to guarantee the minimum 3D quality of experience. The developed system brings key contributions in the P2P transport for real-time multimedia delivery, including a user preference-aware adaptation mechanism, adaptive redundant chunk scheduling for robustness, incentives to decrease the load on the content server for improved system scalability, and resynchronization capability with the DVB transmission. The introduced features are compared with those of some other well-known P2P solutions to highlight the quantitative gains. A subjective testing campaign has also been organized on the developed hybrid platform, which proves the effectiveness of user-aware adaptation over network-based adaptation on a mean opinion score scale. Erhan Ekmekcioglu, Cihat Goktug Gurler, Ahmet M. Kondoz, A. Murat Tekalp |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Compete or Collaborate: Architectures for Collaborative DASH Video Over Future NetworksabstractDynamic adaptive streaming over HTTP (DASH) clients compete with each other over one or more bottleneck links in a network, which results in fluctuations in TCP throughput and QoE, QoE unfairness among clients, and underutilization of the network capacity. We propose centralized and distributed architectures for collaboration between network service provider (NSP), video service provider (VSP), and users (DASH clients) to provide NSP-managed or VSP-managed DASH services over software-defined networks (SDN) with quality-of-service (QoS) reserved network slices. We show that QoS reservation alone is not sufficient to overcome QoE fluctuations per client and unfairness between heterogeneous video clients, and clients also need to employ TCP receive-window adaptation knowing their fair-share bitrate. To this effect, we propose two collaborative streaming service models to inform clients about their fair-share bitrates. We first present an NSP-managed service model with centralized collaboration between the NSP, VSP, and the users, where a traffic engineering manager at the NSP assigns a fair-share bitrate to each DASH client. We then present a VSP-managed service model with centralized or distributed collaboration architectures, where in the former the VSP determines the fair-share bitrate for each client over a reserved network slice and in the latter a group of DASH clients sharing a reserved network slice collaborate among themselves. In the novel distributed collaboration framework, collaboration groups are identified by the VSP, and clients within a group share critical parameters with each other so that each client can estimate its fair-share bitrate. Experimental results demonstrate that collaboration rather than competition between clients not only helps them achieve a smooth goodput near their fair-share bitrate, but also improves the total goodput over the reserved slice. K. Tolga Bagci, Kemal E. Sahin, A. Murat Tekalp |
IEEE Trans. Multim. | 3 |
| 2016 | Queue-allocation optimization for adaptive video streaming over software defined networks with multiple service-levelsabstractInternet service providers (ISP) are deploying software defined networking (SDN), which enables them to better utilize their resources and increase their revenues by offering on-demand differentiated services. In particular, SDN makes provisioning of dynamically managed video services with multiple levels of service viable, since the controller has complete vision of network resources and can change flow paths dynamically by continuously optimizing routes for each flow according to network state. This paper proposes a new method for flow-path computation by queue allocation with the objective of maximizing ISP revenues under per-flow service-level constraints. We formulate the optimization problem and implement its solution in the path computation unit of an SDN controller. Our experiments support that customers requesting improved quality video service receive significantly better QoE compared to best effort services to justify their paying extra for these services. K. Tolga Bagci, Kemal E. Sahin, A. Murat Tekalp |
ICIP | 3 |
| 2014 | Distributed QoS Architectures for Multimedia Streaming Over Software Defined NetworksabstractThis paper presents novel QoS extensions to distributed control plane architectures for multimedia delivery over large-scale, multi-operator Software Defined Networks (SDNs). We foresee that large-scale SDNs shall be managed by a distributed control plane consisting of multiple controllers, where each controller performs optimal QoS routing within its domain and shares summarized (aggregated) QoS routing information with other domain controllers to enable inter-domain QoS routing with reduced problem dimensionality. To this effect, this paper proposes (i) topology aggregation and link summarization methods to efficiently acquire network topology and state information, (ii) a general optimization framework for flow-based end-to-end QoS provision over multi-domain networks, and (iii) two distributed control plane designs by addressing the messaging between controllers for scalable and secure inter-domain QoS routing. We apply these extensions to streaming of layered videos and compare the performance of different control planes in terms of received video quality, communication cost and memory overhead. Our experimental results show that the proposed distributed solution closely approaches the global optimum (with full network state information) and nicely scales to large networks. Hilmi E. Egilmez, A. Murat Tekalp |
IEEE Trans. Multim. | 2 |
| 2013 | Scalable vs. multiple-description video coding for adaptive streaming over peer-to-peer networksabstractBoth scalable video coding (SVC) and multiple description coding (MDC) have been evaluated in the literature for adaptive streaming over peer-to-peer (P2P) networks. However, these evaluations have used either unrealistic or too specific P2P topologies or omitted considering key video coding parameters such as compression efficiency and redundancy. In this study, we evaluate the performances of SVC, MDC and scalable multiple description coding (SMDC) with network-level adaptation schemes under varying peer bandwidth, average end-to-end delay, and peer exit ratio, as well as MDC redundancy levels. Extensive test results help us to reach conclusions regarding under which conditions SVC, MDC or SMDC should be preferred to achieve a reliable video service over mesh P2P overlays. K. Tolga Bagci, Cihat Goktug Gurler, A. Murat Tekalp |
ICIP | 3 |
| 2013 | An Optimization Framework for QoS-Enabled Adaptive Video Streaming Over OpenFlow NetworksabstractOpenFlow is a programmable network protocol and associated hardware designed to effectively manage and direct traffic by decoupling control and forwarding layers of routing. This paper presents an analytical framework for optimization of forwarding decisions at the control layer to enable dynamic Quality of Service (QoS) over OpenFlow networks and discusses application of this framework to QoS-enabled streaming of scalable encoded videos with two QoS levels. We pose and solve optimization of dynamic QoS routing as a constrained shortest path problem, where we treat the base layer of scalable encoded video as a level-1 QoS flow, while the enhancement layers can be treated as level-2 QoS or best-effort flows. We provide experimental results which show that the proposed dynamic QoS framework achieves significant improvement in overall quality of streaming of scalable encoded videos under various coding configurations and network congestion scenarios. Hilmi E. Egilmez, Seyhan Civanlar, A. Murat Tekalp |
IEEE Trans. Multim. | 3 |
| 2012 | A distributed QoS routing architecture for scalable video streaming over multi-domain OpenFlow networksabstractThis paper proposes a new Quality of Service (QoS) optimized routing architecture for video streaming over large-scale multi-domain OpenFlow networks managed by a distributed control plane, where each controller performs optimal routing within its domain and shares summarized intra-domain routing data with other controllers to reduce problem dimensionality for calculating inter-domain routing. We apply the proposed architecture to streaming of scalable (layered) videos, where the base layer routes are dynamically optimized to fulfill a required QoS level, while enhancement layers follow traditional shortest path. We show that the proposed solution approaches the expensive non-scalable globally optimal solution (single controller for the whole network) in terms of received video quality under various congestion scenarios. Hilmi E. Egilmez, Seyhan Civanlar, A. Murat Tekalp |
ICIP | 3 |
| 2012 | Variable chunk size and adaptive scheduling window for P2P streaming of scalable videoabstractInspired by the success of BitTorrent in peer-to-peer (P2P) file sharing, many research groups have proposed extensions of BitTorrent to enable P2P video delivery over the Internet in a time-sensitive manner. However, fixed sized chunks, a feature of the Torrent protocol, leads to inefficiency in application layer framing for error resilient video streaming, since video slices are variable length when encoded at constant quality. This paper proposes two modifications to the Torrent protocol, variable chunk size and adaptive scheduling window, for efficient, error-resilient, adaptive P2P streaming of scalable video. Experimental results show that the proposed modifications yield superior results in terms of number of decoded frames, hence superior quality of experience, in P2P video streaming. Cihat Goktug Gurler, S. Sedef Savas, A. Murat Tekalp |
ICIP | 3 |
| 2012 | Quality of experience aware adaptation strategies for multi-view video over P2P networksabstractThis paper addresses quality of experience (QoE) aware adaptive streaming of multi-view video (MVV) over P2P networks. First, we investigate the effect of different adaptation methods over the QoE using subjective tests, and recommend an MVV adaptation decision chart according to peer buffer status. Second, we propose a simulcast MVV encoding scheme that uses a combination of both standard H.264/AVC and Scalable Video Coding (SVC) extension to achieve the best coding efficiency, while maintaining high adaptation capability. Third, we propose a chunk picking policy that is applicable to any P2P streaming architecture to enable QoE aware adaptive streaming. Finally, we perform P2P video streaming tests over a controlled LAN environment to demonstrate / validate the proposed solution. Cihat Goktug Gurler, S. Sedef Savas, A. Murat Tekalp |
ICIP | 3 |
| 2012 | Evaluation of adaptation methods for multi-view videoabstractMulti-view video (MVV) is the next step in the evaluation of 3DTV. Using IP networks as the transport medium seems to be the most promising solution because MVV has flexible bitrate requirements that can change based on the number of views requested from the receiver. With the advanced streaming technologies like scalable video coding, the capability of video services over IP has been greatly enhanced. However, a successful MVV delivery service cannot be achieved without properly addressing the perceived quality of experience (QoE) of MVV. QoE is an important issue especially in adaptive video streaming in which the quality of the content varies to match the available channel capacity. This study evaluates the effect of different scaling methods some of which are unique to MVV and propose a novel systematic adaptation strategy in order to deliver the best QoE under diverse network conditions. Extensive subjective tests are conducted to compare different scaling methods on MVV by using high definition contents. S. Sedef Savas, Cihat Goktug Gurler, A. Murat Tekalp |
ICIP | 3 |
| 2012 | Resilient peer-to-peer streaming of scalable video over hierarchical multicast trees with backup parent pools
Müge Sayit, Emrullah Turhan Tunali, A. Murat Tekalp |
Signal Process. Image Commun. | 3 |
| 2012 | Adaptation strategies for MGS scalable video streaming
Burak Gorkemli, A. Murat Tekalp |
Signal Process. Image Commun. | 2 |
| 2012 | Adaptive streaming of multi-view video over P2P networks
S. Sedef Savas, Cihat Goktug Gurler, A. Murat Tekalp, Erhan Ekmekcioglu, Stewart Worrall 0001, Ahmet M. Kondoz |
Signal Process. Image Commun. | 3 |
| 2012 | Learn2Dance: Learning Statistical Music-to-Dance Mappings for Choreography SynthesisabstractWe propose a novel framework for learning many-to-many statistical mappings from musical measures to dance figures towards generating plausible music-driven dance choreographies. We obtain music-to-dance mappings through use of four statistical models: 1) musical measure models, representing a many-to-one relation, each of which associates different melody patterns to a given dance figure via a hidden Markov model (HMM); 2) exchangeable figures model, which captures the diversity in a dance performance through a one-to-many relation, extracted by unsupervised clustering of musical measure segments based on melodic similarity; 3) figure transition model, which captures the intrinsic dependencies of dance figure sequences via an n-gram model; 4) dance figure models, capturing the variations in the way particular dance figures are performed, by modeling the motion trajectory of each dance figure via an HMM. Based on the first three of these statistical mappings, we define a discrete HMM and synthesize alternative dance figure sequences by employing a modified Viterbi algorithm. The motion parameters of the dance figures in the synthesized choreography are then computed using the dance figure models. Finally, the generated motion parameters are animated synchronously with the musical audio using a 3-D character model. Objective and subjective evaluation results demonstrate that the proposed framework is able to produce compelling music-driven choreographies. Ferda Ofli, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
IEEE Trans. Multim. | 4 |
| 2012 | Correction to "Learn2Dance: Learning Statistical Music-to-Dance Mappings for Choreography Synthesis"abstractIn the above paper (ibid., vol. 14, no. 3, pp. 747-759, June 2012, p. 750), 44 was printed in error rather than 144. The corrected sentence and equation (6) are presented here. Ferda Ofli, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
IEEE Trans. Multim. | 4 |
| 2011 | Scalable video streaming over OpenFlow networks: An optimization framework for QoS routingabstractOpenFlow is a clean-slate Future Internet architecture that decouples control and forwarding layers of routing, which has recently started being deployed throughout the world for research purposes. This paper presents an optimization framework for the OpenFlow controller in order to provide QoS support for scalable video streaming over an OpenFlow network. We pose and solve two optimization problems, where we route the base layer of SVC encoded video as a lossless-QoS flow, while the enhancement layers can be routed either as a lossy-QoS flow or as a best effort flow, respectively. The proposed approach differs from current QoS architectures since we provide dynamic rerouting capability possibly using non-shortest paths for lossless and lossy QoS flows. We show that dynamic rerouting of QoS flows achieves significant improvement on the video's overall PSNR under network congestion. Hilmi E. Egilmez, Burak Gorkemli, A. Murat Tekalp, Seyhan Civanlar |
ICIP | 3 |
| 2011 | Flexible Transport of 3-D Video Over NetworksabstractThree-dimensional (3-D) video is the next natural step in the evolution of digital media technologies. Recent 3-D autostereoscopic displays can display multiview video with up to 200 views. While it is possible to broadcast 3-D stereo video (two views) over digital TV platforms today, streaming over Internet Protocol (IP) provides a more flexible approach for distribution of stereo and free-view 3-D media to home and mobile with different connection bandwidths and different 3-D displays. Here, flexible transport refers to rate-scalable, resolution-scalable, and view-scalable transport over different channels including digital video broadcasting (DVB) and/or IP. In this paper, we first briefly review the state of the art in 3-D video formats, coding methods for different transport options and video formats, IP streaming protocols, and streaming architectures. We then take a look at beyond the state of the art in 3-D video transport research, including asymmetric stereoscopic video streaming, adaptive and peer-to-peer (P2P) streaming of multiview video, view-selective streaming and future directions in broadcast of 3-D media over IP and jointly over DVB and IP. Cihat Goktug Gurler, Burak Gorkemli, Gorkem Saygili, A. Murat Tekalp |
Proc. IEEE | 4 |
| 2010 | Multi-modal analysis of dance performances for music-driven choreography synthesisabstractWe propose a framework for modeling, analysis, annotation and synthesis of multi-modal dance performances. We analyze correlations between music features and dance figure labels on training dance videos in order to construct a mapping from music measures (segments) to dance figures towards generating music-driven dance choreographies. We assume that dance figure segment boundaries coincide with music measures (audio boundaries). For each training video, figure segments are manually labeled by an expert to indicate the type of dance motion. Chroma features of each measure are used for music analysis. We model temporal statistics of such chroma features corresponding to each dance figure label to identify different rhythmic patterns for that dance motion. The correlations between dance figures and music measures, as well as, correlations between consecutive dance figures are used to construct a mapping for music-driven dance choreography synthesis. Experimental results demonstrate the success of proposed music-driven choreography synthesis framework. Ferda Ofli, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
ICASSP | 4 |
| 2010 | Effects of MGS fragmentation, slice mode and extraction strategies on the performance of SVC with medium-grained scalabilityabstractThis paper presents a comparison of a wide set of MGS fragmentation configurations of SVC in terms of their PSNR performance, with the slice mode on or off, using multiple extraction methods. We also propose a priority-based hierarchical extraction method which outperforms other extraction schemes for most MGS configurations. Experimental results show that splitting the MGS layer into more than five fragments, when the slice mode is on, may result in noticeable decrease in the average PSNR. It is also observed that for videos with large key frame enhancement NAL units, MGS fragmentation and/or slice mode have positive impact on the PSNR of the extracted video at low bitrates. While using slice mode without MGS fragmentation may improve the PSNR performance at low rates, it may result in uneven video quality within frames due to varying quality of slices. Therefore, we recommend combined use of up to five MGS fragments and slice mode, especially for low bitrate video applications. Burak Gorkemli, Yalcin Sadi, A. Murat Tekalp |
ICIP | 3 |
| 2010 | Adaptation strategies for streaming SVC videoabstractThis paper aims to determine the best rate adaptation strategy to maximize the received video quality when streaming SVC video over the Internet. Different bandwidth estimation techniques are implemented for different transport protocols, such as using the TFRC rate when available or calculating the packet transmission rate otherwise. It is observed that controlling the rate of packets dispatched to the transport queue to match the video extraction rate resulted in oscillatory behavior in DCCP CCID3, decreasing the received video quality. Experimental results show that video should be sent at the maximum available network rate rather than at the extraction rate, provided that receiver buffer does not overflow. When the network is over-provisioned, the packet dispatch rate may also be limited with the maximum extractable video rate, to decrease the retransmission traffic without affecting the received video quality. Burak Gorkemli, A. Murat Tekalp |
ICIP | 2 |
| 2010 | Adaptive stereoscopic 3D video streamingabstractThis paper presents a comparative analysis of scalable stereoscopic video coding strategies for adaptive streaming. In particular, we compare scalable simulcast coding of both views using SVC with scalable coding of one view with SVC and non-scalable coding of the other view using H.264/AVC, and benchmark them against non-scalable dependent coding of both views using the MVC. All of these coding options allow both symmetric and asymmetric coding of stereo videos. In addition, we propose a lightweight and periodic feedback mechanism for rate estimation and a strategy to adapt the total stereo source rate using SNR scalability option of SVC, while minimizing the loss rate of non-discardable packets. Experimental results show that dynamic rate scaling of only one view provides sufficient rate adaptation capability and better overall compression efficiency compared to scaling both of the views. Cihat Goktug Gurler, K. Tolga Bagci, A. Murat Tekalp |
ICIP | 3 |
| 2010 | Viterbi-like joint optimization of stereo extraction for on-line rate adaptation in scalable multiview video codingabstractThe concept of Quality Layers (QL) has been adopted in the SVC standard in order to ensure optimal rate adaptation of pre-coded video in the rate-distortion sense. We have previously extended QL to multiview scalable video for efficient transport of 3DTV over the Internet. However, it is not possible to use the QL method in applications that require real-time encoding since priority determination process assumes the availability of the whole stereo video sequence. In this work, a Viterbi-like on-line, joint optimization of right and left view rate-adaptation is proposed for real-time scalable stereo video coding (with one GoP delay). We assume that the encoder/extractor is aware of the available dynamic network bandwidth in order to perform rate-distortion optimized MGS layer selection for each GoP. Experimental results show that the performance of proposed on-line method is comparable to that of QL that would require the whole stereo sequence. Nükhet Özbek, A. Murat Tekalp |
ICIP | 2 |
| 2010 | Quality assessment of asymmetric stereo video codingabstractIt is well known that the human visual system can perceive high frequencies in 3D, even if that information is present in only one of the views. Therefore, the best 3D stereo quality may be achieved by asymmetric coding where the reference (right) and auxiliary (left) views are coded at unequal PSNR. However, the questions of what should be the level of this asymmetry and whether asymmetry should be achieved by spatial resolution reduction or SNR (quality) reduction are open issues. Extensive subjective tests indicate that when the reference view is encoded at sufficiently high quality, the auxiliary view can be encoded above a tow-quality threshold without a noticeable degradation on the perceived stereo video quality. This low-quality threshold may depend on the 3D display; e.g., it is about 31 dB for a parallax barrier display and 33 dB for a polarized projection display. Subjective tests show that, above this PSNR threshold value, users prefer SNR reduction over spatial resolution reduction on both parallax barrier and polarized projection displays. It is also observed that, if the auxiliary view is encoded below this threshold value, symmetric coding starts to perform better than asymmetric coding in terms of perceived 3D video quality. Gorkem Saygili, Cihat Goktug Gurler, A. Murat Tekalp |
ICIP | 3 |
| 2010 | 3DTV and 3d video communicationsabstractWith wider availability of low cost multi-view cameras, 3D displays, and broadband communication options, 3D media is destined to move from the movie theater to home and mobile platforms. In the near term, popular 3D media will most likely be in the form of stereoscopic video with associated spatial audio. Recent trials indicate that consumers are willing to watch stereoscopic 3D media on their TVs, laptops, and mobile phones. While it is possible to broadcast 3D stereoscopic media (two-views) over digital TV platforms today, streaming over IP will provide a more flexible approach for distribution of 3D media to users with different connection bandwidths and different 3D displays. In the intermediate term, free-view 3D video and 3DTV with multi-view capture are next steps in the evolution of 3D media technology. Recent free-view 3D auto-stereoscopic displays can display multi-view video, ranging from 5 to 200 views. Transmission of multi-view 3D media, via broadcast or on-demand, to end users with varying 3D display terminals and bandwidths is one of the biggest challenges to realize the vision of bringing 3D media experience to the home and mobile devices. This requires flexible rate-scalable, resolution-scalable, view-scalable, view-selective, and packet-loss resilient transport methods. In this talk, first I will briefly review the state of the art in 3D video formats, coding methods, IP streaming protocols and streaming architectures. We will then take a look at 3D video transport options. There are two main platforms for 3D broadcasting: standard digital television (DTV) platforms and the IP platform. I will summarize the approach of European project DIOMEDES which is developing novel methods for adaptive streaming of multi-view video over a combination of DVB and IP platforms. I will also summarize additional challenges associated with real-time interactive 3D video communications for applications such as 3D telepresence. Finally, open research challenges for the long term vision of haptic video and holographic 3D video will be presented. A. Murat Tekalp |
MSWiM | 1 |
| 2010 | Immersive haptic interaction with mediaabstractNew 3D video representations enable new modalities of interaction, such as haptic interaction, with 2D and 3D video for truly immersive media applications. Haptic interaction with video includes haptic structure and haptic motion for new immersive experiences. It is possible to compute haptic structure signals from 3D scene geometry or depth information. This paper introduces the concept of haptic motion, as well as new methods to compute haptic structure and motion signals for 2D video-plus-depth representation. The resulting haptic signals can be rendered using a haptic cursor attached to a 2D or 3D video display. Nuray Dindar, A. Murat Tekalp, Cagatay Basdogan |
VCIP | 2 |
| 2010 | Architectures for multi-threaded MVC-compliant multi-view video decoding and benchmark tests
Cihat Goktug Gurler, Anil Aksay, Gozde Bozdagi Akar, A. Murat Tekalp |
Signal Process. Image Commun. | 4 |
| 2009 | Optimization of encoding configuration in scalable multiple description coding for rate-adaptive P2P video multicastingabstractIt is well-known that in peer-to-peer (P2P) streaming from a single source, single point of failure can be avoided by multiple description coding over multiple multicast trees since it provides path diversity. In this scenario, we propose using scalable multiple description coding (SMDC), where each description is scalable so that all descriptions can be efficiently adapted to the available rate of each link for effective congestion control. We also propose a multiple objective optimization (MOO) framework for selection of the best encoding configuration for SMDC from a set of candidates, which will strike the best balance between minimizing average end-to-end rate-distortion performance of each description given a set of packet loss probabilities, while minimizing overall redundancy and maximizing the range of extraction points of each scalable description. The optimization variables are some SVC encoding parameters and MD generation alternatives that result in different levels of redundancy at a fixed total rate for all descriptions. The framework can be used for optimization over other SVC encoding variables and MD generation methods if desired. Results of Monte-Carlo simulation of SMDC streaming of videos demonstrate the performance of the proposed method. Tenzile Berkin Abanoz, A. Murat Tekalp |
ICIP | 2 |
| 2009 | Bandwidth-aware multiple multicast tree formation for P2P scalable video streaming using hierarchical clustersabstractPeer-to-peer (P2P) video streaming is a promising method for multimedia distribution over the Internet, yet many problems remain to be solved such as providing the best quality of service to each peer in proportion to its available resources, low-delay, and fault tolerance. In this paper, we propose a new bandwidth-aware multiple multicast tree formation procedure built on top of a hierarchical cluster based P2P overlay architecture for scalable video (SVC) streaming. The tree formation procedure considers number of sources, SVC layer rates available at each source, as well as delay and available bandwidth over links in an attempt to maximize the quality of received video at each peer. Simulations are performed on NS2 with 500 nodes to demonstrate that the overall performance of the system in terms of average received video quality of all peers is significantly better if peers with higher available bandwidth are placed higher up in the trees and peers with lower bandwidth are near the leaves. Müge Sayit, Emrullah Turhan Tunali, A. Murat Tekalp |
ICIP | 3 |
| 2009 | 3D display dependent quality evaluation and rate allocation using scalable video codingabstractIt is well known that the human visual system can perceive high frequency content in 3D, even if that information is present in only one of the views. Then, the best 3D perception quality may be achieved by allocating the rates of the reference (right) and auxiliary (left) views asymmetrically. However the question of whether the rate reduction for the auxiliary view should be achieved by spatial resolution reduction (coding a downsampled version of the video followed by upsampling after decoding) or quality (QP) reduction is an open issue. This paper shows that which approach should be preferred depends on the 3D display technology used at the receiver. Subjective tests indicate that users prefer lower quality (larger QP) coding of the auxiliary view over lower resolution coding if a “full spatial resolution” 3D display technology (such as polarized projection) is employed. On the other hand, users prefer lower resolution coding of the auxiliary view over lower quality coding if a “reduced spatial resolution” 3D display technology (such as parallax barrier - autostereoscopic) is used. Therefore, we conclude that for 3D IPTV services, while receiving full quality/resolution reference view, users should subscribe to differently scaled versions of the auxiliary view depending on their 3D display technology. We also propose an objective 3D video quality measure that takes the 3D display technology into account. Gorkem Saygili, Cihat Goktug Gurler, A. Murat Tekalp |
ICIP | 3 |
| 2009 | Multi-threaded architectures and benchmark tests for real-time multi-view video decodingabstract3D video based on multi-view representations is becoming widely popular. Real-time encoding/decoding of such video is an important concern as the number and resolution of views increase. We present systematic methods for design and optimization of real-time multi-view video encoding/decoding algorithms using multi-core processors and provide benchmark results. The proposed multi-core decoding architectures are fully compliant with the current JVT-MVC international standard, and enable multi-threaded processing with negligible loss of encoding efficiency. Benchmark results show that multi-core processors and multi-threading decoding is necessary for real-time multiview video decoding and display. Cihat Goktug Gurler, Anil Aksay, Gozde Bozdagi Akar, A. Murat Tekalp |
ICME | 4 |
| 2009 | Quality Layers in scalable multi-view video codingabstractThe quality layers concept is proposed and adopted in the JSVM in order to ensure an optimal adaptation in a rate-distortion sense. In this paper it is extended to multi-view case for scalable coding of multi-view videos (MVV). Rate-visual-distortion optimized rate allocation among the views is necessary for efficient transport of multi-view data over the Internet. This paper addresses a standard compatible approach to solve the problem. Experimental results are presented that demonstrate the effectiveness of the proposed method for several multiview videos. Nükhet Özbek, A. Murat Tekalp |
ICME | 2 |
| 2009 | SVC-based scalable multiple description video coding and optimization of encoding configuration
Tenzile Berkin Abanoz, A. Murat Tekalp |
Signal Process. Image Commun. | 2 |
| 2008 | Audio-driven human body motion analysis and synthesisabstractThis paper presents a framework for audio-driven human body motion analysis and synthesis. We address the problem in the context of a dance performance, where gestures and movements of the dancer are mainly driven by a musical piece and characterized by the repetition of a set of dance figures. The system is trained in a supervised manner using the multiview video recordings of the dancer. The human body posture is extracted from multiview video information without any human intervention using a novel marker-based algorithm based on annealing particle filtering. Audio is analyzed to extract beat and tempo information. The joint analysis of audio and motion features provides a correlation model that is then used to animate a dancing avatar when driven with any musical piece of the same genre. Results are provided showing the effectiveness of the proposed algorithm. Ferda Ofli, Cristian Canton, Joëlle Tilmanne, Yasemin Demir, Elif Bozkurt, Yücel Yemez, Engin Erzin, A. Murat Tekalp |
ICASSP | 8 |
| 2008 | Video streaming over wireless DCCPabstractIt is envisioned that access networks will be mostly wireless in the future. Hence, it is of interest to consider extensions of the Datagram Congestion Control Protocol (DCCP) for wireless networks. This paper focuses on the problems of video streaming over DCCP in the wireless domain and proposes a cross-layer solution in which the wireless packet loss information available in the Medium Access (MAC) layer is utilized by DCCP to distinguish congestion losses from wireless losses and behave accordingly. Tests performed with our modified DCCP confirm that using cross-layer loss information prevents unnecessary rate decreases and results in better video streaming experiences. Burak Gorkemli, M. Oguz Sunay, A. Murat Tekalp |
ICIP | 3 |
| 2008 | Unsupervised dance figure analysis from video for dancing Avatar animationabstractThis paper presents a framework for unsupervised video analysis in the context of dance performances, where gestures and 3D movements of a dancer are characterized by repetition of a set of unknown dance figures. The system is trained in an unsupervised manner using Hidden Markov Models (HMMs) to automatically segment multi-view video recordings of a dancer into recurring elementary temporal body motion patterns to identify the dance figures. That is, a parallel HMM structure is employed to automatically determine the number and the temporal boundaries of different dance figures in a given dance video. The success of the analysis framework has been evaluated by visualizing these dance figures on a dancing avatar animated by the computed 3D analysis parameters. Experimental results demonstrate that the proposed framework enables synthetic agents and/or robots to learn dance figures from video automatically. Ferda Ofli, Engin Erzin, Yücel Yemez, A. Murat Tekalp, Çigdem Eroglu Erdem, A. Tanju Erdem, Tolga Abaci, Mehmet K. Özkan |
ICIP | 4 |
| 2008 | Rate-visual-distortion optimized extraction with Quality Layers for scalable coding of stereo videosabstractThe quality layers concept is proposed and adopted in the JSVM in order to ensure an optimal adaptation in a rate-distortion sense. In this paper it is extended to multiview case for scalable coding of stereo videos. Rate-visual-distortion optimized rate allocation among the views is necessary for efficient transport of 3DTV data over the Internet. This paper addresses a standard compatible approach to solve the problem. Experimental results are presented that demonstrate the effectiveness of the proposed method for several stereo videos. Nükhet Özbek, A. Murat Tekalp |
ICIP | 2 |
| 2008 | Unequal inter-view rate allocation using scalable stereo video coding and an objective stereo video quality measureabstractIn stereoscopic 3D video, it is well-known that humans can perceive high quality 3D video provided that one of the views is in high quality. Hence, in stereo video encoding, the best overall rate vs. perceived-distortion performance may be achieved by reduction of the spatial, temporal, and/or quantization resolution of the second view, while keeping the first view in full resolution. In this paper, we address the best selection of unequal inter-view rate allocation strategy depending on the content of the video for a scalable multi-view video codec (SMVC) (Ozbek et al., 2006). Since the perceived 3D video quality does not correlate well with the average PSNR of the two views, we propose a new quantitative measure using a weighted combination of two PSNR values and a jerkiness measure. We verified that unequal rate allocation between the left and right views results in better perceived stereo video quality. Nükhet Özbek, A. Murat Tekalp |
ICME | 2 |
| 2008 | Analysis of Head Gesture and Prosody Patterns for Prosody-Driven Head-Gesture AnimationabstractWe propose a new two-stage framework for joint analysis of head gesture and speech prosody patterns of a speaker towards automatic realistic synthesis of head gestures from speech prosody. In the first stage analysis, we perform Hidden Markov Model (HMM) based unsupervised temporal segmentation of head gesture and speech prosody features separately to determine elementary head gesture and speech prosody patterns, respectively, for a particular speaker. In the second stage, joint analysis of correlations between these elementary head gesture and prosody patterns is performed using Multi-Stream HMMs to determine an audio-visual mapping model. The resulting audio-visual mapping model is then employed to synthesize natural head gestures from arbitrary input test speech given a head model for the speaker. In the synthesis stage, the audio-visual mapping model is used to predict a sequence of gesture patterns from the prosody pattern sequence computed for the input test speech. The Euler angles associated with each gesture pattern are then applied to animate the speaker head model. Objective and subjective evaluations indicate that the proposed synthesis by analysis scheme provides natural looking head gestures for the speaker with any input test speech, as well as in "prosody transplant" and gesture transplant" scenarios. Mehmet Emre Sargin, Yücel Yemez, Engin Erzin, A. Murat Tekalp |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Watermarking and Streaming Compressed VideoabstractIn this paper, we propose a novel method for watermarking compressed video for streaming. The proposed method performs motion compensated watermarking. There are two consequences of this; first we see that the impact of the watermark on the final bitrate is negligible, and the proposed method has inherent robustness to mild packet losses in the channel. Experimental results show that the proposed method operates at lower distortion levels and has higher watermark detection rates compared to current watermarking methods. We observe 20% bitrate reduction for the same watermark detection rate. We also show that when data partitioning is used over a packet loss channel the proposed method performs at acceptable watermark detection rates. Oztan Harmanci, Mehmet Kivanç Mihçak, A. Murat Tekalp |
ICASSP (1) | 3 |
| 2007 | Rate Allocation Between Views in Scalable Stereo Video Coding using an Objective Stereo Video Quality MeasureabstractIt is well-known that in stereoscopic 3D video systems humans perceive good quality 3D video as long as one of the eyes sees a high quality view. Hence, in stereo video encoding/streaming, best rate allocation between views can be addressed by reduction of the spatial resolution, frame rate, and/or quantization parameter of the second view with respect to the first view. In this paper, we address selection of the rate allocation strategy between views for our recently developed scalable multi-view video codec (SMVC) (N. Ozbek and A. M. Tekalp, 2006) to obtain the best rate-distortion performance. Since 3D video quality perception does not correlate well with the overall PSNR of the two views, we propose a new quantitative measure for stereo video quality as weighted combination of two PSNR values and a jerkiness measure. The weights are determined by means of correlating subjective quality test results and the objective measure scores on a set of test videos. DSCQS test methodology is used for subjective evaluation of stereo videos. Experimental results are presented to demonstrate how the objective and subjective 3D video quality varies for different choices of rate allocation between the views. Nükhet Özbek, A. Murat Tekalp, Emrullah Turhan Tunali |
ICASSP (1) | 2 |
| 2007 | Prosody-Driven Head-Gesture AnimationabstractWe present a new framework for joint analysis of head gesture and speech prosody patterns of a speaker towards automatic realistic synthesis of head gestures from speech prosody. The proposed two-stage analysis aims to "learn" both elementary prosody and head gesture patterns for a particular speaker, as well as the correlations between these head gesture and prosody patterns from a training video sequence. The resulting audio-visual mapping model is then employed to synthesize natural head gestures from arbitrary input test speech given a head model for the speaker. Objective and subjective evaluations indicate that the proposed synthesis by analysis scheme provides natural looking head gestures for the speaker with any input test speech. Mehmet Emre Sargin, Engin Erzin, Yücel Yemez, A. Murat Tekalp, A. Tanju Erdem, Çigdem Eroglu Erdem, Mehmet K. Özkan |
ICASSP (2) | 4 |
| 2007 | Optimal Selection of Encoding Configuration for Scalable Video CodingabstractIt is well-known that the wider the range of extraction points a scalable bitstream supports, the lower the compression efficiency at these extraction points. Moreover, this compression efficiency generally varies according to what combination of scalability types are used to support this range of extraction points as specified by the encoding configuration. Hence, we propose some objective criteria as a measure of coverage, compression efficiency and rate-distortion performance of a configuration, and then present a multiple-objective optimization formulation to select the best encoding configuration for scalable video coding, given a range of bitstreams that must be supported. The method is demonstrated by experimental results. Tenzile Berkin Abanoz, A. Murat Tekalp |
ICIP (2) | 2 |
| 2007 | Selective Streaming of Multi-View Video for Head-Tracking 3D DisplaysabstractWe present a novel client-driven multi-view video streaming system that allows a user watch 3-D video interactively with significantly reduced bandwidth requirements by transmitting a small number of views selected according to his/her head position. The proposed scheme can be used to efficiently stream a dense set of multi-view sequences (light-fields) or wider baseline multi-view sequences together with depth information. The user's head position is tracked and predicted into the future to select the views that best match the user's current viewing angle dynamically. Prediction of future head positions is needed so that views matching the predicted head positions can be requested from the server ahead of time in order to account for delays due to network transport and stream switching. Highly compressed, lower quality versions of some other views are also requested in order to provide protection against having to display the wrong view when the current user viewpoint differs from the predicted viewpoint. The proposed system makes use of multi-view coding (MVC) and scalable video coding (SVC) concepts together to obtain improved compress ion efficiency while providing flexibility in bandwidth allocation to the selected views. Rate-distortion performance of the proposed system is demonstrated under different experimental conditions. Engin Kurutepe, M. Reha Civanlar, A. Murat Tekalp |
ICIP (3) | 3 |
| 2007 | Estimation and Analysis of Facial Animation Parameter PatternsabstractWe propose a framework for estimation and analysis of temporal facial expression patterns of a speaker. The proposed system aims to learn personalized elementary dynamic facial expression patterns for a particular speaker. We use head-and-shoulder stereo video sequences to track lip, eye, eyebrow, and eyelid motion of a speaker in 3D. MPEG-4 Facial Definition Parameters (FDPs) are used as the feature set, and temporal facial expression patterns are represented by the MPEG-4 Facial Animation Parameters (FAPs). We perform Hidden Markov Model (HMM) based unsupervised temporal segmentation of upper and lower facial expression features separately to determine recurrent elementary facial expression patterns for a particular speaker. These facial expression patterns coded by FAP sequences, which may not be tied with prespecified emotions, can be used for personalized emotion estimation and synthesis of a speaker. Experimental results are presented. Ferda Ofli, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
ICIP (4) | 4 |
| 2007 | Adaptive Streaming of Scalable Stereoscopic Video Over DCCPabstractWe propose a new adaptive streaming model that utilizes DCCP in order to efficiently stream stereoscopic video over the Internet for 3DTV transport. The model allocates the available channel bandwidth, which is calculated by the DCCP, among the views according to the suppression theory of human vision. The video rate is adapted to the DCCP rate for each group of pictures (GoP) by adaptive extraction of layers from a scalable multi-view bitstream. The objective of the streaming model is to maximize perceived quality of the received 3D video while minimizing the number of possible display interrupts. Experimental results successfully demonstrate stereo video streaming over DCCP on wide area network. Nükhet Özbek, Burak Gorkemli, A. Murat Tekalp, Emrullah Turhan Tunali |
ICIP (6) | 3 |
| 2007 | Multicamera Audio-Visual Analysis of Dance FiguresabstractWe present an automated system for multicamera motion capture and audio-visual analysis of dance figures. The multiview video of a dancing actor is acquired using 8 synchronized cameras. The motion capture technique is based on 3D tracking of the markers attached to the person's body in the scene, using stereo color information without need for an explicit 3D model. The resulting set of 3D points is then used to extract the body motion features as 3D displacement vectors whereas MFC coefficients serve as the audio features. In the first stage of multimodal analysis, we perform Hidden Markov Model (HMM) based unsupervised temporal segmentation of the audio and body motion features, separately, to determine the recurrent elementary audio and body motion patterns. Then in the second stage, we investigate the correlation of body motion patterns with audio patterns, that can be used for estimation and synthesis of realistic audio-driven body animation. Ferda Ofli, Yasemin Demir, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
ICME | 5 |
| 2007 | A New Scalable Multi-View Video Coding Configuration for Robust Selective Streaming of Free-Viewpoint TVabstractFree viewpoint TV (FTV) is a new media format that allows a user to change his/her viewpoint freely. To this effect, multi-view video must be coded to satisfy two conflicting requirements: (i) achieve high compression efficiency, and (ii) allow view switching with low delay. This paper proposes a new encoding configuration for scalable multi-view video coding, which achieves a compromise between the two requirements. In the new scalable multi-view configuration, the base layer is encoded with inter-view prediction at a minimum acceptable quality, while enhancement layers for each view only depend on their respective base layers (with no interview prediction). Thus, the base layer shall be served to all users, while enhancement layers shall be served selectively to users depending on their channel bandwidth and viewing direction. We compare the compression efficiency of the proposed method with those of non-scalable multi-view coding (MVC) and simulcast (H.264/AVC of each view independently) solutions. Nükhet Özbek, A. Murat Tekalp, Emrullah Turhan Tunali |
ICME | 2 |
| 2007 | Cross-Layer Optimized Rate Adaptation and Scheduling for Multiple-User Wireless Video StreamingabstractWe present a cross-layer optimized video rate adaptation and user scheduling scheme for multi-user wireless video streaming aiming for maximum quality of service (QoS) for each user,, maximum system video throughput, and QoS fairness among users. These objectives are jointly optimized using a multi-objective optimization (MOO) framework that aims to serve the user with the least remaining playback time, highest delivered video seconds per transmission slot and maximum video quality. Experiments with the IS-856 (1timesEV-DO) standard numerology and ITU pedestrian A and vehicular B environments show significant improvements over the state-of- the-art wireless schedulers in terms of user QoS, QoS fairness, and the system throughput. Tanir Ozcelebi, M. Oguz Sunay, A. Murat Tekalp, M. Reha Civanlar |
IEEE J. Sel. Areas Commun. | 3 |
| 2007 | End-to-end stereoscopic video streaming with content-adaptive rate and format control
Anil Aksay, Selen Pehlivan, Engin Kurutepe, Cagdas Bilen, Tanir Ozcelebi, Gozde Bozdagi Akar, M. Reha Civanlar, A. Murat Tekalp |
Signal Process. Image Commun. | 8 |
| 2007 | Transport Methods in 3DTV - A SurveyabstractWe present a survey of transport methods for 3-D video ranging from early analog 3DTV systems to most recent digital technologies that show promise in designing 3DTV systems of tomorrow. Potential digital transport architectures for 3DTV include the DVB architecture for broadcast and the Internet Protocol (IP) architecture for wired or wireless streaming. There are different multiview representation/compression methods for delivering the 3-D experience, which provide a tradeoff between compression efficiency, random access to views, and ease of rate adaptation, including the "video-plus-depth" compressed representation and various multiview video coding (MVC) options. Commercial activities using these representations in broadcast and IP streaming have emerged, and successful transport of such data has been reported. Motivated by the growing impact of the Internet protocol based media transport technologies, we focus on the ubiquitous Internet as the network infrastructure of choice for future 3DTV systems. Current research issues in unicast and multicast mode multiview video streaming include network protocols such as DCCP and peer-to-peer protocols, effective congestion control, packet loss protection and concealment, video rate adaptation, and network/service scalability. Examples of end-to-end systems for multiview video streaming have been provided. Gozde Bozdagi Akar, A. Murat Tekalp, Christoph Fehn, M. Reha Civanlar |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Client-Driven Selective Streaming of Multiview Video for Interactive 3DTVabstractWe present a novel client-driven multiview video streaming system that allows a user to watch 3D video interactively with significantly reduced bandwidth requirements by transmitting a small number of views selected according to his/her head position. The user's head position is tracked and predicted into the future to select the views that best match the user's current viewing angle dynamically. Prediction of future head positions is needed so that views matching the predicted head positions can be prefetched in order to account for delays due to network transport and stream switching. The system allocates more bandwidth to the selected views in order to render the current viewing angle. Highly compressed, lower quality versions of some other views are also prefetched for concealment if the current user viewpoint differs from the predicted viewpoint. An objective measure based on the abruptness of the head movements and delays in the system is introduced to determine the number of additional lower quality views to be prefetched. The proposed system makes use of multiview coding (MVC) and scalable video coding (SVC) concepts together to obtain improved compression efficiency while providing flexibility in bandwidth allocation to the selected views. Rate-distortion performance of the proposed system is demonstrated under different experimental conditions. Engin Kurutepe, M. Reha Civanlar, A. Murat Tekalp |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | A Stochastic Framework for Rate-Distortion Optimized Video Coding Over Error-Prone NetworksabstractThis paper proposes a complete stochastic framework for RD optimal encoder design for video over error-prone networks, which applies to any motion-compensated predictive video codec. The distortion measure has been taken as the mean square error over an ensemble of channels given an estimate of the instantaneous packet loss probability. We show that 1) the optimal motion compensated prediction, in the MSE sense, requires computation of the expected value of the reference frames, and 2) calculation of the MSE (distortion measure) requires computation of the second moment of the reference frames. We propose a recursive procedure for the computation of both the expected value and second moment of the reference frames, which are together called the stochastic frame buffer. Furthermore, we propose a stochastic RD optimization method for selection of the optimal macroblock mode and motion vectors given the instantaneous packet loss probability. If available, channel feedback can also be incorporated into the proposed stochastic framework. However, the proposed framework does not require a feedback channel to exist, and when it exists, it does not have to be lossless. In the absence of any packet losses, the proposed stochastic framework reduces to the well-known deterministic RD optimization procedures. One possible application of the optimal stochastic framework would be for multicast streaming to an ensemble of receivers. Experimental results indicate that the proposed framework outperforms other available error tracking and control schemes. Oztan Harmanci, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 2007 | Rate-Distortion Optimal Video Transport Over IP Allowing Packets With Bit ErrorsabstractWe propose new models and methods for rate-distortion (RD) optimal video delivery over IP, when packets with bit errors are also delivered. In particular, we propose RD optimal methods for slicing and unequal error protection (UEP) of packets over IP allowing transmission of packets with bit errors. The proposed framework can be employed in a classical independent-layer transport model for optimal slicing, as well as in a cross-layer transport model for optimal slicing and UEP, where the forward error correction (FEC) coding is performed at the link layer, but the application controls the FEC code rate with the constraint that a given IP packet is subject to constant channel protection. The proposed method uses a novel dynamic programming approach to determine the optimal slicing and UEP configuration for each video frame in a practical manner, that is compliant with the AVC/H.264 standard. We also propose new rate and distortion estimation techniques at the encoder side in order to efficiently evaluate the objective function for a slice configuration. The cross-layer formulation option effectively determines which regions of a frame should be protected better; hence, it can be considered as a spatial UEP scheme. We successfully demonstrate, by means of experimental results, that each component of the proposed system provides significant gains, up to 2.0 dB, compared to competitive methods. Oztan Harmanci, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 2007 | Delay-Distortion Optimization for Content-Adaptive Video StreamingabstractWe propose a new pre-roll delay-distortion optimization (DDO) framework that allows determination of the minimum pre-roll delay and distortion while ensuring continuous playback for on-demand content-adaptive video streaming over limited bitrate networks. The input video is first divided into temporal segments, which are assigned a relevance weight and a maximum distortion level, called relevance-distortion policy, which may be specified by the user. The system then encodes the input video according to the specified relevance-distortion policy, whereby the optimal spatial and temporal resolutions and quantization parameters, also called encoding parameters, are selected for each temporal segment. The optimal encoding parameters are computed using a novel, multi-objective optimization formulation, where a relevance weighted distortion measure and pre-roll delay are jointly minimized under maximum allowable buffer size, continuous playback, and maximum allowable distortion constraints. The performance of the system has been demonstrated for on-demand streaming of soccer videos with substantial improvement in the weighted distortion without any increase in pre-roll delay over a very low-bitrate network using AVC/H.264 encoding Tanir Ozcelebi, A. Murat Tekalp, M. Reha Civanlar |
IEEE Trans. Multim. | 2 |
| 2007 | Audiovisual Synchronization and Fusion Using Canonical Correlation AnalysisabstractIt is well-known that early integration (also called data fusion) is effective when the modalities are correlated, and late integration (also called decision or opinion fusion) is optimal when modalities are uncorrelated. In this paper, we propose a new multimodal fusion strategy for open-set speaker identification using a combination of early and late integration following canonical correlation analysis (CCA) of speech and lip texture features. We also propose a method for high precision synchronization of the speech and lip features using CCA prior to the proposed fusion. Experimental results show that i) the proposed fusion strategy yields the best equal error rates (EER), which are used to quantify the performance of the fusion strategy for open-set speaker identification, and ii) precise synchronization prior to fusion improves the EER; hence, the best EER is obtained when the proposed synchronization scheme is employed together with the proposed fusion strategy. We note that the proposed fusion strategy outperforms others because the features used in the late integration are truly uncorrelated, since they are output of the CCA analysis. Mehmet Emre Sargin, Yücel Yemez, Engin Erzin, A. Murat Tekalp |
IEEE Trans. Multim. | 4 |
| 2006 | Multimodal Speaker Identification Using Canonical Correlation AnalysisabstractIn this work, we explore the use of canonical correlation analysis to improve the performance of multimodal recognition systems that involve multiple correlated modalities. More specifically, we consider the audiovisual speaker identification problem, where speech and lip texture (or intensity) modalities are fused in an open-set identification framework. Our motivation is based on the following observation. The late integration strategy, which is also referred to as decision or opinion fusion, is effective especially in case the contributing modalities are uncorrelated and thus the resulting partial decisions are statistically independent. Early integration techniques on the other hand can be favored only if a couple of modalities are highly correlated. However, coupled modalities such as audio and lip texture also consist of some components that are mutually independent. Thus we first perform a cross-correlation analysis on the audio and lip modalities so as to extract the correlated part of the information, and then employ an optimal combination of early and late integration techniques to fuse the extracted features. The results of the experiments testing the performance of the proposed system are also provided Mehmet Emre Sargin, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
ICASSP (1) | 4 |
| 2006 | An Analysis of Constant Bitrate and Constant PSNR Video Encoding for Wireless NetworksabstractIn wireless networks, transmission of constant-quality, high bitrate video is a challenging task due to channel capacity and buffer limitations. Content adaptive rate control, is used as a solution to this problem. Instead of transmitting all of the video content at low quality, the most important content can be transmitted at high quality while still preserving an acceptable quality for the remaining segments. Furthermore, the rate control strategy inside the individual temporal segments plays a key role for the network performance and viewing quality. Although constant quality video encoding inside the temporal segments is preferable for the best viewing experience, it causes more network packet losses due to adverse bitrate fluctuations in the video stream. In cases when the network is too much loaded, it may be better to employ constant bitrate encoding for network friendliness. In this paper, a performance analysis of constant bitrate and constant peak signal-to-noise ratio encoding for content adaptive rate controlled video streaming over wireless networks is presented. Experimental results obtained using AVC/H.264 encoding in a CDMA/HDR multi-user environment with cross-layer optimized scheduling show performance comparisons of CBR and CPSNR encoding. Tanir Ozcelebi, Fabio De Vito, A. Murat Tekalp, M. Reha Civanlar, M. Oguz Sunay, Juan Carlos De Martin |
ICC | 3 |
| 2006 | Adaptive Peer-To-Peer Video Streaming with Optimized Flexible Multiple Description CodingabstractEfficient peer-to-peer (P2P) video streaming is a challenging task due to time-varying nature of both the number of available peers and network/channel conditions. To this effect, we propose a receiver driven P2P streaming system which utilizes a flexible scalable multiple description coding method, where the number of base and enhancement descriptions, and the rate and redundancy level of each description can be adapted on the fly. The optimization of the parameters of the proposed MDC scheme according to network conditions is discussed within the context of the proposed adaptive P2P streaming framework, where the number and quality of available streaming peers/paths are a priori unknown and vary in time. Experimental results, by means of NS-2 network simulation of a P2P video streaming system, show that adaptation of the number, type, and rate of descriptions and the redundancy level of each description according to network conditions yields significantly superior performance when compared to MDC schemes using a fixed number of descriptions/layers with fixed rate and redundancy level. Emrah Akyol, A. Murat Tekalp, M. Reha Civanlar |
ICIP | 2 |
| 2006 | Multi-View Image Registration for Wide-Baseline Visual Sensor NetworksabstractWe present a new dense multi-view registration technique for wide-baseline video/images that integrates a parametric optical flow-based approach with a sparse set of feature correspondences, based on a locally planar approximation of a nonplanar scene. The proposed method can deal with illuminance variations between the views, which is critically important for wide-baseline applications. It differs from existing work on wide-baseline image registration in that it requires only image information and provides dense matching without computing any camera calibration matrices or performing any prior scene segmentation. These characteristics render the method suitable for practical deployment in visual sensor networks, towards which the current work is directed. We demonstrate the performance of the proposed method on simulated multi-view images of a virtual 3D world composed of piece-wise smooth textured surfaces, as well as real wide-baseline images of nonplanar textured surfaces. Gulcin Caner, A. Murat Tekalp, Gaurav Sharma 0001, Wendi B. Heinzelman |
ICIP | 2 |
| 2006 | Hierarchical Representation and Coding of 3D Mesh GeometryabstractHierarchical mesh representation and mesh simplification have been addressed in computer graphics for adaptive level-of-detail rendering of 3D objects. In this paper, we propose a hierarchical system suitable for progressive transmission using hierarchical 3D meshes such that each mesh level has Delaunay topology. The Delaunay topology constraint on each mesh layer not only helps to design meshes with desired geometric properties, but also enables efficient compression of the mesh data. The hierarchical compression technique is based on a nearest-neighbor ordering of mesh node points. This ordering serves to define the mesh boundary as well as a spatial prediction relation on the nodes, which is employed for differential node point location. The compression method allows progressive transmission and quality scalability. Isil Celasun, Serkan Eroksuz, Rizwan A. Siddiqui, A. Murat Tekalp |
ICIP | 4 |
| 2006 | Rate-Distortion Optimal Video Transport Over IP with Bit ErrorsabstractIn this paper we propose a method for video delivery over bit error channels. In particular, we propose a rate distortion optimal method for slicing and unequal error protection (UEP) of packets over bit error channels. The proposed method performs full frame based search using a novel dynamic programming approach to determine the optimal slicing configuration in a practically short time. Also we propose a rate and distortion estimation technique that decreases the time to evaluate the objective function for a slice configuration. The proposed method can perform rate-distortion UEP that can be used over forward error correction (FEC) capable channels. We show that the proposed method successfully exploit the local dynamics of a video frame and perform more than 1 dB better than common methods. Oztan Harmanci, A. Murat Tekalp |
ICIP | 2 |
| 2006 | Application-Layer QoS Fairness in Wireless Video SchedulingabstractIn mobile video transmission systems, the initial delay for pre-fetching video at the client buffer needs to be short due to buffer limitations and application-layer user convenience. Therefore, an effective cross-layer wireless design is required that considers both physical and application layer aspects of such a system. We present a cross-layer optimized multi-user video adaptation and scheduling scheme for wireless video communication, where quality-of-service (QoS) fairness among users is provided while maximizing user convenience and video throughput. Application and physical layer aspects are jointly optimized using a multi-objective optimization (MOO) framework that tries to schedule the user with the least remaining playback time and the highest video throughput (delivered video seconds per transmission slot) with maximum video quality. Experiments with the IS-856 (1xEV-DO) standard and ITU pedestrian A and vehicular B environments show the improvements over today's schedulers in terms of QoS fairness and user utility. Tanir Ozcelebi, M. Oguz Sunay, M. Reha Civanlar, A. Murat Tekalp |
ICIP | 4 |
| 2006 | Scalable Multi-View Video Coding for Interactive 3DTVabstractA standard for scalable video coding (SVC) is currently being worked on by the ISO MPEG group. Work on standardization of multiple-view video coding (MVC) has also recently started under the ISO MPEG. Although there are many approaches published on SVC and MVC, there is no current work reported on scalable multi-view video coding (SMVC). This paper presents new coding structures for scalable stereo and multi-view video coding. The proposed structures are implemented as extensions to the JSVM software and resulting bitrates and PSNR are demonstrated. SMVC can be used for transport of multiview video over IP for interactive 3DTV by dynamic adaptive combination of temporal, spatial, and SNR scalability according to network conditions Nükhet Özbek, A. Murat Tekalp |
ICME | 2 |
| 2006 | Combined Gesture-Speech Analysis and Speech Driven Gesture SynthesisabstractMultimodal speech and speaker modeling and recognition are widely accepted as vital aspects of state of the art human-machine interaction systems. While correlations between speech and lip motion as well as speech and facial expressions are widely studied, relatively little work has been done to investigate the correlations between speech and gesture. Detection and modeling of head, hand and arm gestures of a speaker have been studied extensively and these gestures were shown to carry linguistic information. A typical example is the head gesture while saying "yes/no". In this study, correlation between gestures and speech is investigated. In speech signal analysis, keyword spotting and prosodic accent event detection has been performed. In gesture analysis, hand positions and parameters of global head motion are used as features. The detection of gestures is based on discrete pre-designated symbol sets, which are manually labeled during the training phase. The gesture-speech correlation is modeled by examining the co-occurring speech and gesture patterns. This correlation can be used to fuse gesture and speech modalities for edutainment applications (i.e. video games, 3-D animations) where natural gestures of talking avatars are animated from speech. A speech driven gesture animation example has been implemented for demonstration Mehmet Emre Sargin, Oya Aran, Alexey Karpov 0001, Ferda Ofli, Yelena Yasinnik, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
ICME | 9 |
| 2006 | Semantic multimedia analysis for content-adaptive video streamingabstractThis paper provides a survey of recent approaches for semantic content analysis using multiple modalities, such as video, audio, and on-screen text. It also introduces new methods for shot (GoP) level rate allocation for content-adaptive multimedia streaming in order to achieve the best user utility given relevance-distortion policy and available average channel bandwidth of a user. Some examples are provided. A. Murat Tekalp |
ISCAS | 1 |
| 2006 | Multimodal speaker/speech recognition using lip motion, lip texture and audio
Hasan Ertan Çetingül, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
Signal Process. | 4 |
| 2006 | Local Image Registration by Adaptive FilteringabstractWe propose a new adaptive filtering framework for local image registration, which compensates for the effect of local distortions/displacements without explicitly estimating a distortion/displacement field. To this effect, we formulate local image registration as a two-dimensional (2-D) system identification problem with spatially varying system parameters. We utilize a 2-D adaptive filtering framework to identify the locally varying system parameters, where a new block adaptive filtering scheme is introduced. We discuss the conditions under which the adaptive filter coefficients conform to a local displacement vector at each pixel. Experimental results demonstrate that the proposed 2-D adaptive filtering framework is very successful in modeling and compensation of both local distortions, such as Stirmark attacks, and local motion, such as in the presence of a parallax field. In particular, we show that the proposed method can provide image registration to: a) enable reliable detection of watermarks following a Stirmark attack in nonblind detection scenarios, b) compensate for lens distortions, and c) align multiview images with nonparametric local motion. Gulcin Caner, A. Murat Tekalp, Gaurav Sharma 0001, Wendi B. Heinzelman |
IEEE Trans. Image Process. | 2 |
| 2006 | Lossless watermarking for image authentication: a new framework and an implementationabstractWe present a novel framework for lossless (invertible) authentication watermarking, which enables zero-distortion reconstruction of the un-watermarked images upon verification. As opposed to earlier lossless authentication methods that required reconstruction of the original image prior to validation, the new framework allows validation of the watermarked images before recovery of the original image. This reduces computational requirements in situations when either the verification step fails or the zero-distortion reconstruction is not needed. For verified images, integrity of the reconstructed image is ensured by the uniqueness of the reconstruction procedure. The framework also enables public(-key) authentication without granting access to the perfect original and allows for efficient tamper localization. Effectiveness of the framework is demonstrated by implementing the framework using hierarchical image authentication along with lossless generalized-least significant bit data embedding. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 2006 | Discriminative Analysis of Lip Motion Features for Speaker Identification and Speech-ReadingabstractThere have been several studies that jointly use audio, lip intensity, and lip geometry information for speaker identification and speech-reading applications. This paper proposes using explicit lip motion information, instead of or in addition to lip intensity and/or geometry information, for speaker identification and speech-reading within a unified feature selection and discrimination analysis framework, and addresses two important issues: 1) Is using explicit lip motion information useful, and, 2) if so, what are the best lip motion features for these two applications? The best lip motion features for speaker identification are considered to be those that result in the highest discrimination of individual speakers in a population, whereas for speech-reading, the best features are those providing the highest phoneme/word/phrase recognition rate. Several lip motion feature candidates have been considered including dense motion features within a bounding box about the lip, lip contour motion features, and combination of these with lip shape features. Furthermore, a novel two-stage, spatial, and temporal discrimination analysis is introduced to select the best lip motion features for speaker identification and speech-reading applications. Experimental results using an hidden-Markov-model-based recognition system indicate that using explicit lip motion information provides additional performance gains in both applications, and lip motion features prove more valuable in the case of speech-reading application. Hasan Ertan Çetingül, Yücel Yemez, Engin Erzin, A. Murat Tekalp |
IEEE Trans. Image Process. | 4 |
| 2005 | Cross-layer design for real-time video streaming over 1xEV-DO using multiple objective optimizationabstractIn wireless packet transmission systems, it is crucial to provide fairness in service while maximizing user utility and channel throughput. This is possible via intelligent allocation of system resources. For real time video applications, a pre-roll time for pre-fetching data at the client buffer is needed in order to compensate for channel variations that cause client buffer under/overflows, hence facilitating continuous playout of the video. In this paper, a novel multiple objective optimized (MOO) opportunistic multiple access scheme for optimal scheduling of users in a 1xEV-DO (IS-856) system is presented. At each time slot, the user that experiences the best compromise between the least buffer occupancy and the best channel condition is served. Experiments conducted in ITU Pedestrian A and Vehicular B environments show that our algorithm treats each user fairly while its channel throughput performance is very close to the ideal case, in which the user with the best channel characteristics is always served with no buffer constraints. Tanir Ozcelebi, M. Oguz Sunay, A. Murat Tekalp, M. Reha Civanlar |
GLOBECOM | 3 |
| 2005 | An adaptive filtering framework for image registrationabstractImage registration is a fundamental task in both image processing and computer vision. We present a novel method for local image registration based on adaptive filtering techniques. We utilize an adaptive filter to estimate and track correspondences among multiple images containing overlapping views of common scene regions. Image pixels are traversed in an order established by space-filling curves, to preserve the contiguity and hence track locally varying registration changes. The algorithm differs from pre-existing work on image registration in that it requires only local information and relatively low computational effort. These characteristics render the method suitable for deployment in imaging sensor networks, toward which the current work is directed. We evaluate the performance of the proposed algorithm using images captured with a digital camera in various real-world scenarios. Experimental results show that the proposed method can significantly improve accuracy and robustness over a global 2D parametric registration and can also outperform the local registration algorithm based on the Lucas-Kanade optical flow technique (Lucas, B. and Kanade, T., 1981). Gulcin Caner, A. Murat Tekalp, Gaurav Sharma 0001, Wendi B. Heinzelman |
ICASSP (2) | 2 |
| 2005 | Pitch and Duration Modification for Speech WatermarkingabstractWe propose a speech watermarking algorithm based on the modification of the pitch (fundamental frequency) and duration of the quasi-periodic speech segments. Natural variability of these speech features allows watermarking modifications to be imperceptible to the human observer. On the other hand, the significance of these features makes the system robust to common signal processing operations and low data-rate source excitation based speech coders. This class of coders is particularly obstructive for conventional audio watermarking algorithms when applied to speech signals. A pitch synchronous overlap and add (PSOLA) algorithm is used for pitch and duration modifications in the watermark embedding phase. Experiments with multiple speech codecs show very good robustness with low data-rate (5-8 kbps) speech coders. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
ICASSP (2) | 3 |
| 2005 | Robust Lip-Motion Features For Speaker IdentificationabstractThe paper addresses the selection of robust lip-motion features for an audio-visual open-set speaker identification problem. We consider two alternatives for initial lip motion representation. In the first alternative, the feature vector is composed of the 2D-DCT coefficients of the motion vectors estimated within the detected rectangular mouth region, whereas in the second, lip boundaries are tracked over the video frames, and only the motion vectors around the lip contour are taken into account along with the shape of the lip boundary. Experimental results of the HMM-based identification system are included for performance comparison of the two lip motion representation alternatives. Hasan Ertan Çetingül, Yücel Yemez, Engin Erzin, A. Murat Tekalp |
ICASSP (1) | 4 |
| 2005 | A Zero Error Propagation Extension to H264 for Low Delay Video Communications Over Lossy ChannelsabstractIn this paper, we introduce a method for video transmission over lossy channels. The proposed method is similar to NEWPRED and uses feedback information to stop error propagation. Main operational difference is replication of packet loss and error concealment process at the encoder and using the possibly error concealed frames as reference frames in the motion prediction loop. We have studied the optimization of the proposed method under moderate error rates where ACK/NACK mode NEWPRED does not efficiently solve the problem. The proposed solution is an adaptive method that optimizes the operation based on the expected error propagation. We present comparison results with NEWPRED under various channel conditions and video sources. Oztan Harmanci, A. Murat Tekalp |
ICASSP (2) | 2 |
| 2005 | Object Recognition by Partial Shape Matching Guided SearchabstractWe retrieve samples from large image/video databases by means of learning by example. We propose a fast partial shape matching guided 2D object recognition algorithm to significantly accelerate recognition/matching of full (non-occluded) or partially occluded objects. The significant increase in speed comes from the fact that the search space is reduced to only those combinations of regions in the neighborhood of potential partial matches, as opposed to all combinations of regions as was done in our prior work (Xu et al. (2003)). Theoretical calculations and experimental results are provided to demonstrate the effectiveness of the proposed algorithm on real images. Eli Saber, Yaowu Xu, A. Murat Tekalp |
ICASSP (2) | 3 |
| 2005 | Scalable multiple description video coding with flexible number of descriptionsabstractMultiple description video coding mitigates the effects of packet losses introduced by congestion and/or bit errors. In this paper, we propose a novel multiple description video coding technique, based on fully scalable wavelet video coding, which allows post encoding adaptation of the number of descriptions, the redundancy level of each description, and bitrate of each description by manipulation of the encoded bitstream. We demonstrate that the proposed method provides excellent coding efficiency, outperforming most other multiple description methods proposed so far. We also provide experimental results to show that varying the number of descriptions according to network conditions is superior to using a fixed number of descriptions, by means of NS-2 network simulation of a peer-to-peer video streaming system. Emrah Akyol, A. Murat Tekalp, M. Reha Civanlar |
ICIP (3) | 2 |
| 2005 | Stochastic frame buffers for rate distortion optimized loss resilient video communicationsabstractIn this paper we propose an error control scheme for video communications over lossy channels. The proposed algorithm uses stochastic frame buffers (SFB) to determine the expected decoder error by using channel simulated frames during motion vector (MV) and reference frame search, macroblock (MB) mode decisions and residual selection. Unlike previous methods, our algorithm also focuses on optimizing the MV and prediction frame selection and residual selection over error prone channels. The algorithm minimizes a Lagrangian to achieve the optimal result for mode and MVs. If there is channel feedback, the system can utilize it to update the SFB. We perform experiments that compare our system with intra refreshes, an improved NEWPRED and optimal mode switching algorithms such as ROPE under various conditions such as existence of a feedback channel or not. We show that the proposed system outperforms others. Oztan Harmanci, A. Murat Tekalp |
ICIP (1) | 2 |
| 2005 | FAST H.264/AVC video encoding with multiple frame referencesabstractWe focus on the question of how we can select the best multiple reference pictures for enhanced H.264 video encoding by a fast, computationally efficient method. We propose a simple histogram-similarity based method for selecting the best set of multiple reference pictures. Out-of-order coding of these frames is implemented by means of pyramid encoding. Experimental results show that the proposed approach can provide encoding time saving up to 23% with similar picture quality and bitrate for selected video sequences. Nükhet Özbek, A. Murat Tekalp |
ICIP (1) | 2 |
| 2005 | Minimum delay content adaptive video streaming over variable bitrate channels with a novel stream switching solutionabstractWe present a new channel adaptive stream switching solution for variable bitrate video transmission, where the receiver buffer status is used for making switching decisions and instantaneous transmission rate is determined by means of explicit channel feedback or TCP friendly rate control. Receiver buffer status is used in selecting from a set of pre-encoded bitstreams avoiding buffer underflows and overflows. A pre-roll delay to buffer data at the receiving side is necessary in order to compensate for variations in the channel throughput and the encoding bitrate. For each stream, content dependent video coding parameters are chosen by means of dynamic programming such that the maximum overall video quality and minimum pre-roll delay are achieved for a finite number of transmission rates. Experimental results show that buffer violations are successfully avoided using the proposed stream switching framework as opposed to a regular non-adaptive streaming case. Tanir Ozcelebi, M. Reha Civanlar, A. Murat Tekalp |
ICIP (1) | 3 |
| 2005 | Partial shape recognition by sub-matrix matching for partial matching guided image labeling
Eli Saber, Yaowu Xu, A. Murat Tekalp |
Pattern Recognit. | 3 |
| 2005 | Lossless generalized-LSB data embeddingabstractWe present a novel lossless (reversible) data-embedding technique, which enables the exact recovery of the original host signal upon extraction of the embedded information. A generalization of the well-known least significant bit (LSB) modification is proposed as the data-embedding method, which introduces additional operating points on the capacity-distortion curve. Lossless recovery of the original is achieved by compressing portions of the signal that are susceptible to embedding distortion and transmitting these compressed descriptions as a part of the embedded payload. A prediction-based conditional entropy coder which utilizes unaltered portions of the host signal as side-information improves the compression efficiency and, thus, the lossless data-embedding capacity. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp, Eli Saber |
IEEE Trans. Image Process. | 3 |
| 2005 | Multimodal speaker identification using an adaptive classifier cascade based on modality reliabilityabstractWe present a multimodal open-set speaker identification system that integrates information coming from audio, face and lip motion modalities. For fusion of multiple modalities, we propose a new adaptive cascade rule that favors reliable modality combinations through a cascade of classifiers. The order of the classifiers in the cascade is adaptively determined based on the reliability of each modality combination. A novel reliability measure, that genuinely fits to the open-set speaker identification problem, is also proposed to assess accept or reject decisions of a classifier. A formal framework is developed based on probability of correct decision for analytical comparison of the proposed adaptive rule with other classifier combination rules. The proposed adaptive rule is more robust in the presence of unreliable modalities, and outperforms the hard-level max rule and soft-level weighted summation rule, provided that the employed reliability measure is effective in assessment of classifier decisions. Experimental results that support this assertion are provided. Engin Erzin, Yücel Yemez, A. Murat Tekalp |
IEEE Trans. Multim. | 3 |
| 2004 | Semantic object segmentation by dynamic learning from multiple examplesabstractWe present a novel "dynamic learning" approach for an intelligent image database system to automatically improve object segmentation and labeling without user intervention, as new examples become available, for object-based indexing. The proposed approach is an extension of our earlier work on "learning by example", which addressed labeling of similar objects in a set of database images based on a single example (Saber et al. (2003)). It utilizes multiple example object templates to improve the accuracy of existing object segmentations and labels. We also propose to use Normalized Area of Symmetric Differences (NASD) as the similarity metric in "dynamic learning", due to its robustness to boundary noise that results from automatic image segmentation. The performance of the dynamic learning concept is demonstrated by experimental results. Yaowu Xu, Eli Saber, A. Murat Tekalp |
ICASSP (3) | 3 |
| 2004 | Motion-compensated temporal filtering within the H.264/AVC standardabstractWe propose an adaptive motion-compensated temporal filtering (MCTF) structure to provide efficient temporal scalability within the H.264/AVC video compression standard. MCTF has traditionally been considered within fully scalable wavelet video coders. However, motion-compensated simple 5/3 lifted temporal wavelet filtering suffers at scene changes, as well as occlusion regions. We note that the bi-directional motion compensation mode in the H.264 standard is best equipped with the state of the art adaptive features such as adaptive block size, mode switching between forward, backward and bidirectional prediction and in-loop deblocking filter. Hence, we propose a GOP structure to implement block-based adaptive MCTF within the H.264 syntax using stored B-pictures, similar to the motion-compensated 5/3 wavelet filtering. We provide experimental results to compare the results of our proposed codec with those of other scalable wavelet video coders which use MCTF. It is also possible to employ the proposed adaptive MCTF structure within fully scalable wavelet video codecs. Emrah Akyol, A. Murat Tekalp, M. Reha Civanlar |
ICIP | 2 |
| 2004 | Discriminative lip-motion features for biometric speaker identification
Hasan Ertan Çetingül, Yücel Yemez, Engin Erzin, A. Murat Tekalp |
ICIP | 4 |
| 2004 | Optimization of h264 for low delay video communications over lossy channelsabstractIn this paper, we study the data partitioning (DP) and its optimization for H264 video coding standard. H264 does not include DP in baseline profile, which is the most suitable profile for low delay, low complexity, and loss prone environments. To analyze the optimization of DP, we first introduce the concept of subchannels to abstract the physical layer. This allows us to move channel coding from application layer to physical layer. Then, we build the video encoder system around NEWPRED (1996) so that error propagation and its analysis is eliminated. Finally, we provide macroblock and slice level optimizations that result in optimal mode decisions and unequal error protection (LTEP) rates for data partitions. Experimental results show about 0.5 dB performance increase as compared to no data partitioning. Oztan Harmanci, A. Murat Tekalp |
ICIP | 2 |
| 2004 | Optimal rate and input format control for content and context adaptive video streaming
Tanir Ozcelebi, A. Murat Tekalp, M. Reha Civanlar |
ICIP | 2 |
| 2004 | Adaptive classifier cascade for multimodal speaker identificationabstractWe present a multimodal open-set speaker identification system that integrates information coming from audio, face and lip motion modalities. For fusion of multiple modalities, we propose a new adaptive cascade rule that favors reliable modality combinations through a cascade of classifiers. The order of the classifiers in the cascade is adaptively determined based on the reliability of each modality combination. A novel reliability measure, that genuinely fits to the open-set speaker identification problem, is also proposed to assess accept or reject decisions of a classifier. The proposed adaptive rule is more robust in the presence of unreliable modalities, and outperforms the hard-level max rule and soft-level weighted summation rule, provided that the employed reliability measure is effective in assessment of classifier decisions. Experimental results that support this assertion are provided. Engin Erzin, Yücel Yemez, A. Murat Tekalp |
INTERSPEECH | 3 |
| 2004 | On optimal selection of lip-motion features for speaker identificationabstractThis paper addresses the selection of best lip motion features for biometric open-set speaker identification. The best features are those that result in the highest discrimination of individual speakers in a population. We first detect the face region in each video frame. The lip region for each frame is then segmented following the registration of successive face regions by global motion compensation. The initial lip feature vector is composed of the 2D-DCT coefficients of the optical flow vectors within the lip region at each frame. We propose to select the most discriminative features from the full set of transform coefficients by using a probabilistic measure that maximizes the ratio of intra-class and inter-class probabilities. The resulting discriminative feature vector with reduced dimension is expected to maximize the identification performance. Experimental results are also included to demonstrate the performance. Hasan Ertan Çetingül, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
MMSP | 4 |
| 2004 | Optimal rate and input format control for content and context adaptive streaming of sports videosabstractA novel dynamic programming based technique for optimal selection of input video format and compression rate for video streaming based on "relevancy" of the content and user context is presented. The technique uses context dependent content analysis to divide the input video into temporal segments. User selected relevance levels (weights) are refined by using audio information and assigned to these segments. The weights are used in formulating a constrained optimization problem, which is solved using dynamic programming. The technique minimizes a weighted distortion measure and the initial waiting time for continuous playback under maximum acceptable distortion constraints. Spatial resolution, frame rate and average audio volume of the temporal segments and the DCT quantization parameters are used as optimization variables. Tanir Ozcelebi, A. Murat Tekalp, M. Reha Civanlar |
MMSP | 2 |
| 2004 | Dynamic learning from multiple examples for semantic object segmentation and search
Yaowu Xu, Eli Saber, A. Murat Tekalp |
Comput. Vis. Image Underst. | 3 |
| 2004 | Collusion-resilient fingerprinting by random pre-warpingabstractFingerprinting of audio-visual content using digital watermarks is an effective means of determining originators of unauthorized/pirated copies. Watermarks embedded in content can trace the traitor responsible for piracy. Multiple users may, however, collude and collectively escape identification by creating an average of their individually watermarked copies that appears unwatermarked. We propose a novel collusion-resilience mechanism, wherein the host signal is warped randomly prior to watermarking. As each copy undergoes a distinctive warp, collusion through averaging either yields low-quality results or requires substantial computational resources to undo random warps. The method is independent of the watermarking scheme used and imposes no restrictions on the watermark signal. We demonstrate the effectiveness of this approach on digital images. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
IEEE Signal Process. Lett. | 3 |
| 2004 | Performance measures for video object segmentation and trackingabstractWe propose measures to evaluate quantitatively the performance of video object segmentation and tracking methods without ground-truth (GT) segmentation maps. The proposed measures are based on spatial differences of color and motion along the boundary of the estimated video object plane and temporal differences between the color histogram of the current object plane and its predecessors. They can be used to localize (spatially and/or temporally) regions where segmentation results are good or bad; and/or they can be combined to yield a single numerical measure to indicate the goodness of the boundary segmentation and tracking results over a sequence. The validity of the proposed performance measures without GT have been demonstrated by canonical correlation analysis with another set of measures with GT on a set of sequences (where GT information is available). Experimental results are presented to evaluate the segmentation maps obtained from various sequences using different segmentation approaches. Çigdem Eroglu Erdem, Bülent Sankur, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 2004 | Integrated semantic-syntactic video modeling for search and browsingabstractVideo processing and computer vision communities usually employ shot-based or object-based structural video models and associate low-level (color, texture, shape, and motion) and semantic descriptions (textual annotations) with these structural (syntactic) elements. Database and information retrieval communities, on the other hand, employ entity-relation or object-oriented models to model the semantics of multimedia documents. This paper proposes a new generic integrated semantic-syntactic video model to include all of these elements within a single framework to enable structured video search and browsing combining textual and low-level descriptors. The proposed model includes semantic entities (video objects and events) and the relations between them. We introduce a new "actor" entity to enable grouping of object roles in specific events. This context-dependent classification of attributes of an object allows for more efficient browsing and retrieval. The model also allows for decomposition of events into elementary motion units and elementary reaction/interaction units in order to access mid-level semantics and low-level video features. The instantiations of the model are expressed as graphs. Users can formulate flexible queries that can be translated into such graphs. Alternatively, users can input query graphs by editing an abstract model (model template). Search and retrieval is accomplished by matching the query graph with those instantiated models in the database. Examples and experimental results are provided to demonstrate the effectiveness of the proposed integrated modeling and querying framework. Ahmet Ekin, A. Murat Tekalp, Rajiv Mehrotra |
IEEE Trans. Multim. | 2 |
| 2003 | Level-embedded lossless image compressionabstractA level-embedded lossless compression method for continuous-tone still images is presented. Level (bit-plane) scalability is achieved by separating the image into two layers before compression and excellent compression performance is obtained by exploiting both spatial and inter-level correlations. A comparison of the proposed scheme with a number of scalable and non-scalable lossless image compression algorithms is performed to benchmark its performance. The results indicate that the level-embedded compression incurs only a small penalty in compression efficiency. Mehmet Utku Celik, A. Murat Tekalp, Gaurav Sharma 0001 |
ICASSP (3) | 2 |
| 2003 | Stochastic modeling of motion tracking failuresabstractThis research introduces a new and effective method of predicting motion tracking failures and demonstrates its application towards the analysis of gait and human motion. We define a tracking failure as an event and describe its temporal characteristics using a hidden Markov model (HMM). This stochastic model is trained using previous examples of tracking failures and is applied to the Kalman-based tracking of a parametric, structural model of the human body. With an observation sequence derived from the noise covariance matrices of the structural model parameters, we show a causal relationship between the conditional output probability of the HMM and imminent tracking failures. Results are demonstrated on a variety of multi-view sequences of complex human motion. Shiloh L. Dockstader, Nikita S. Imennov, A. Murat Tekalp |
ICASSP (3) | 3 |
| 2003 | Shot type classification by dominant color for sports video segmentation and summarizationabstractThis paper introduces a novel generic framework for sports video processing by using the common feature of most sports: the dominant color of the field. This dominant field color is automatically detected and updated to compensate for the lighting and weather changes by a robust dominant color region detection algorithm. We also introduce new shot type classification algorithms for soccer and basketball. Finally, we use shot-based low-level features for domain-specific high-level applications. Specifically, the system detects soccer goals and summarizes soccer games in real-time, and it enables basketball fans to skip fouls, free throws, and time-out events by segmenting a basketball game into plays and breaks. Ahmet Ekin, A. Murat Tekalp |
ICASSP (3) | 2 |
| 2003 | Joint audio-video processing for biometric speaker identificationabstractWe present a bimodal audio-visual speaker identification system. The objective is to improve the recognition performance over conventional unimodal schemes. The proposed system exploits not only the temporal and spatial correlations existing in the speech and video signals of a speaker, but also the cross-correlation between these two modalities. Lip images extracted from each video frame are transformed onto an eigenspace. The obtained eigenlip coefficients are interpolated to match the rate of the speech signal and fused with Mel frequency cepstral coefficients (MFCC) of the corresponding speech signal. The resulting joint feature vectors are used to train and test a hidden Markov model (HMM) based identification system. Experimental results are included to demonstrate the system performance. Alper Kanak, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
ICASSP (2) | 4 |
| 2003 | Markov-Based Failure Prediction for Human Motion AnalysisabstractThis paper presents a new method of detecting and predicting motion tracking failures with applications in human motion and gait analysis. We define a tracking failure as an event and describe its temporal characteristics using a hidden Markov model (HMM). This stochastic model is trained using previous examples of tracking failures. We derive vector observations for the HMM using the noise covariance matrices characterizing a tracked, 3D structural model of the human body. We show a causal relationship between the conditional output probability of the HMM, as transformed using a logarithmic mapping function, and impending tracking failures. Results are illustrated on several multi-view sequences of complex human motion. Shiloh L. Dockstader, Nikita S. Imennov, A. Murat Tekalp |
ICCV | 3 |
| 2003 | Collusion-resilient fingerprinting using random prewarpingabstractFingerprinting of audio-visual content using digital watermarks is an effective means of determining the originators of unauthorized copies and fighting piracy in digital distribution networks. In particular, watermarks embedded within the content help trace the traitor responsible for the piracy. A group of users may, however, collude and collectively escape identification by creating an average of their individually watermarked copies that appears unwatermarked. We propose a novel collusion-resilience mechanism, wherein the host signal is warped randomly prior to watermarking. As each copy undergoes a distinctive warp, collusion through averaging either yields low-quality results or requires substantial computational resources to undo random warps. The proposed method is independent of the watermarking scheme used and does not impose any restrictions on the watermark signal that are required by some collusion resistant watermarking schemes. We demonstrate the effectiveness of this approach on digital images. Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
ICIP (1) | 3 |
| 2003 | Fault-tolerant tracking for gait analysisabstractThis research introduces a method of predicting tracking failures and applies it to the robust analysis of human gait. The body is represented using a multicomponent structural model. For each component, the proposed approach extracts features from tracked noise covariance matrices and uses them to construct an observation sequence for a hidden Markov model (HMM) trained to detect tracking failures. When transformed with a logarithmic function, the conditional output probability of the HMM is shown to have a causal relationship with imminent tracking failures. This fusion of multiple structural models with a reliable means of failure prediction facilitates the successful tracking and extraction of gait variables. Results are demonstrated on numerous video sequences. Shiloh L. Dockstader, Nikita S. Imennov, Michel J. Berg, A. Murat Tekalp |
ICIP (2) | 4 |
| 2003 | A robust Bayesian network for articulated motion classificationabstractWe introduce a new approach to motion-based recognition that combines the temporally descriptive abilities of a hidden Markov model (HMM) with the inferential power of a Bayesian belief network. We define activities using a collection of multiple Markov models, each associated with a unique set of body model parameters or gait variables. A single Bayesian network integrates the models by operating on virtual evidence derived from the HMM conditional output probabilities. We introduce both fundamental and auxiliary models for characterizing events and tracking failures, respectively. We demonstrate the system using multi-view video sequences corrupted by occlusion, noise, and entirely missing observations. Shiloh L. Dockstader, Nikita S. Imennov, A. Murat Tekalp |
ICIP (3) | 3 |
| 2003 | Robust dominant color region detection and color-based applications for sports videoabstractThis paper proposes a novel automatic dominant color region detection algorithm that is robust to temporal variations in the dominant color due to field, weather, and lighting conditions throughout a sports video. The algorithm automatically learns the dominant color statistics of the field independent of the sports type, and updates color statistics throughout a sporting event by using two color spaces, a control space and a primary space. The robustness of the algorithm results from adaptation of the statistics of the dominant color in the primary space with drift protection using the control space, and fusion of the information from two spaces. We also propose novel and generic color-based algorithms for referee, player-of-interest, and play-break event detection in sports video. The efficiency of the proposed algorithms is demonstrated over a dataset of various sports video, including basketball, football, golf, and soccer video. Ahmet Ekin, A. Murat Tekalp |
ICIP (1) | 2 |
| 2003 | Multimodal speaker identification with audio-video processingabstractIn this paper we present a multimodal audio-visual speaker identification system. The objective is to improve the recognition performance over conventional unimodal schemes. The proposed system decomposes the information existing in a video stream into three components: speech, face texture and lip motion. Lip motion between successive frames is first computed in terms of optical flow vectors and then encoded as a feature vector in a magnitude direction histogram domain. The feature vectors obtained along the whole stream are then interpolated to match the rate of the speech signal and fused with mel frequency cepstral coefficients (MFCC) of the corresponding speech signal. The resulting joint feature vectors are used to train and test a Hidden Markov Model (HMM) based identification system. Face texture images are treated separately in eigenface domain and integrated to the system through decision-fusion. Experimental results are also included for demonstration of the system performance. Yücel Yemez, Alper Kanak, Engin Erzin, A. Murat Tekalp |
ICIP (3) | 4 |
| 2003 | Super resolution recovery for multi-camera surveillance imagingabstractIn many surveillance video applications, it is of interest to recognize an object or a person, which occupies a small portion of a low-resolution, noisy video. This paper addresses the problem of super-resolution recovery of a region of interest from more than one low-resolution view of a scene recorded by multiple cameras. The multiple camera scenario alleviates the difficulty in registration of multiple frames of video that contain non-rigid or multiple object motion in the single camera case. With proper temporal registration of multiple videos, arbitrary scene motion can be handled. The success of super-resolution recovery from multiple views in real applications vitally depends on two factors: i) the accuracy of multiple view registration results, and ii) the accuracy of the camera and data acquisition model. We propose a system, which consists of a method for sub-pixel accurate spatio-temporal alignment of multiple video sequences for view registration and the projections onto convex sets method for super-resolution recovery. Experiments were implemented using two commercial analog video cameras, which do not perform on-board compression. Experimental results show that the super resolution recovery of dynamic scenes can be achieved as long as the multiple views of the scene can be registered with sub-pixel accuracy. Gulcin Caner, A. Murat Tekalp, Wendi B. Heinzelman |
ICME | 2 |
| 2003 | Generic play-break event detection for summarization and hierarchical sports video analysisabstractThis paper proposes a single generic real-time (or near real-time) play-break event detection algorithm for multiple sports, which include football, tennis, basketball, and soccer. The proposed algorithm only uses shot-based generic cinematic features, such as shot type and shot length. Detected play-break events are employed for two purposes: 1) all plays in certain sports, such as football and tennis, are presented as summaries, and 2) play-break events, as part of a hierarchical event detection scheme, determine the segments-of-interest for the other event detection algorithms. An example of such event detection algorithms is given for soccer goal events, where the proposed soccer goal detection algorithm exploits the common cinematic techniques that are employed during the breaks that follow the goal plays. We demonstrate the genericity of the proposed play-break detection algorithm over football, tennis, basketball video and the effectiveness of the proposed soccer goal detection algorithm over a large data set. Ahmet Ekin, A. Murat Tekalp |
ICME | 2 |
| 2003 | Joint audio-video processing for biometric speaker identificationabstractIn this paper we present a bimodal audio-visual speaker identification system. The objective is to improve the recognition performance over conventional unimodal schemes. The proposed system exploits not only the temporal and spatial correlations existing in speech and video signals of a speaker, but also the cross-correlation between these two modalities. Lip images extracted for each video frame are transformed onto an eigenspace. The obtained eigenlip coefficients are interpolated to match the rate of the speech signal and fused with mel frequency cepstral coefficients (MFCC) of the corresponding speech signal. The resulting joint feature vectors are used to train and test a hidden Markov model (HMM) based identification system. Experimental results are also included for demonstration of the system performance. Alper Kanak, Engin Erzin, Yücel Yemez, A. Murat Tekalp |
ICME | 4 |
| 2003 | Performance measures for video object segmentation and tracking
Çigdem Eroglu Erdem, Bülent Sankur, A. Murat Tekalp |
VCIP | 3 |
| 2003 | Object-based image labeling through learning by example and multi-level segmentation
Yaowu Xu, Pinar Duygulu, Eli Saber, A. Murat Tekalp, Fatos T. Yarman-Vural |
Pattern Recognit. | 4 |
| 2003 | Gray-level-embedded lossless image compression
Mehmet Utku Celik, Gaurav Sharma 0001, A. Murat Tekalp |
Signal Process. Image Commun. | 3 |
| 2003 | Bi-directional 2-D mesh representation for video object rendering, editing and superresolution in the presence of occlusion
P. Erhan Eren, A. Murat Tekalp |
Signal Process. Image Commun. | 2 |
| 2003 | Video object tracking with feedback of performance measuresabstractPresents a scalable object tracking framework, which is capable of tracking the contour of nonrigid objects in the presence of occlusion. The framework consists of open-loop boundary prediction and closed-loop boundary correction parts. The open-loop prediction block adaptively divides the object contour into subcontours, and estimates the mapping parameters for each subsegment. The closed-loop boundary correction block employs a suitably weighted combination of low-level features such as color edge, color segmentation, motion models, and motion segmentation for each subcontour. Performance evaluation measures are used in a feedback loop to evaluate the goodness of the segmentation/tracking in order to adjust the weights assigned to each of these low-level features for each subcontour at each frame. The framework is scalable because it can be adapted to track a coarse estimate of the boundary of selected objects in real-time, as well as pixel-accurate boundary tracking in off-line mode. The proposed method does not depend on any single motion or shape model, and does not need training. Experimental results demonstrate that the algorithm is able to track the object boundaries under significant occlusion and background clutter. Çigdem Eroglu Erdem, Bülent Sankur, A. Murat Tekalp |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2003 | Stochastic kinematic modeling and feature extraction for gait analysisabstractThis research presents a new model-based approach toward the three-dimensional (3-D) tracking and extraction of gait and human motion. We suggest the use of a hierarchical, structural model of the human body that introduces the concept of soft kinematic constraints. These constraints take the form of a priori, stochastic distributions learned from previous configurations of the body exhibited during specific activities; they are used to supplement an existing motion model limited by hard kinematic constraints. We use time-varying parameters of the structural model to measure gait velocity, stance width, stride length, stance times, and other gait variables with multiple degrees of accuracy and robustness. To characterize tracking performance, we also introduce a novel geometric model of expected tracking failures. We demonstrate and quantify the performance of the suggested models using multi-view, video sequences of human movement captured in a complex home environment. Shiloh L. Dockstader, Michel J. Berg, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 2003 | Automatic soccer video analysis and summarizationabstractWe propose a fully automatic and computationally efficient framework for analysis and summarization of soccer videos using cinematic and object-based features. The proposed framework includes some novel low-level processing algorithms, such as dominant color region detection, robust shot boundary detection, and shot classification, as well as some higher-level algorithms for goal detection, referee detection, and penalty-box detection. The system can output three types of summaries: i) all slow-motion segments in a game; ii) all goals in a game; iii) slow-motion segments classified according to object-based features. The first two types of summaries are based on cinematic features only for speedy processing, while the summaries of the last type contain higher-level semantics. The proposed framework is efficient, effective, and robust. It is efficient in the sense that there is no need to compute object-based features when cinematic features are sufficient for the detection of certain events, e.g., goals in soccer. It is effective in the sense that the framework can also employ object-based features when needed to increase accuracy (at the expense of more computation). The efficiency, effectiveness, and robustness of the proposed framework are demonstrated over a large data set, consisting of more than 13 hours of soccer video, captured in different countries and under different conditions. Ahmet Ekin, A. Murat Tekalp, Rajiv Mehrotra |
IEEE Trans. Image Process. | 2 |
| 2003 | Object segmentation and labeling by learning from examplesabstractWe propose a system that employs low-level image segmentation followed by color and two-dimensional (2-D) shape matching to automatically group those low-level segments into objects based on their similarity to a set of example object templates presented by the user. A hierarchical content tree data structure is used for each database image to store matching combinations of low-level regions as objects. The system automatically initializes the content tree with only "elementary nodes" representing homogeneous low-level regions. The "learning" phase refers to labeling of combinations of low-level regions that have resulted in successful color and/or 2-D shape matches with the example template(s). These combinations are labeled as "object nodes" in the hierarchical content tree. Once learning is performed, the speed of second-time retrieval of learned objects in the database increases significantly. The learning step can be performed off-line provided that example objects are given in the form of user interest profiles. Experimental results are presented to demonstrate the effectiveness of the proposed system with hierarchical content tree representation and learning by color and 2-D shape matching on collections of car and face images. Yaowu Xu, Eli Saber, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 2003 | Two-stage hierarchical video summary extraction to match low-level user browsing preferencesabstractA compact summary of video that conveys visual content at various levels of detail enhances user interaction significantly. In this paper, we propose a two-stage framework to generate MPEG-7-compliant hierarchical key frame summaries of video sequences. At the first stage, which is carried out off-line at the time of content production, fuzzy clustering and data pruning methods are applied to given video segments to obtain a nonredundant set of key frames that comprise the finest level of the hierarchical summary. The number of key frames allocated to each shot or segment is determined dynamically and without user supervision through the use of cluster validation techniques. A coarser summary is generated on-demand in the second stage by reducing the number of key frames to match the low-level browsing preferences of a user. The proposed method has been validated by experimental results on a collection of video programs. A. Müfit Ferman, A. Murat Tekalp |
IEEE Trans. Multim. | 2 |
| 2002 | Semantics of multimedia in MPEG-7abstractIn this paper, we present the tools standardized by MPEG-7 for describing the semantics of multimedia. In particular, we focus on the abstraction model, entities, attributes and relations of MPEG-7 semantic descriptions. MPEG-7 tools can describe the semantics of specific instances of multimedia such as one image or one video segment but can also generalize these descriptions either to multiple instances of multimedia or to a set of semantic descriptions. The key components of MPEG-7 semantic descriptions are semantic entities such as objects and events, attributes of these entities such as labels and properties, and, finally, relations of these entities such as an object being the patient of an event. The descriptive power and usability of these tools has been demonstrated in numerous experiments and applications, these make them key candidates to enable intelligent applications that deal with multimedia at human levels. Ana B. Benitez, Hawley K. Rising, Corinne Jörgensen, Riccardo Leonardi, Alessandro Bugatti, Kôiti Hasida, Rajiv Mehrotra, A. Murat Tekalp, Ahmet Ekin, Toby Walker |
ICIP (1) | 8 |
| 2002 | Reversible data hidingabstractWe present a novel reversible (lossless) data hiding (embedding) technique, which enables the exact recovery of the original host signal upon extraction of the embedded information. A generalization of the well-known LSB (least significant bit) modification is proposed as the data embedding method, which introduces additional operating points on the capacity-distortion curve. Lossless recovery of the original is achieved by compressing portions of the signal that are susceptible to embedding distortion, and transmitting these compressed descriptions as a part of the embedded payload. A prediction-based conditional entropy coder which utilizes static portions of the host as side-information improves the compression efficiency, and thus the lossless data embedding capacity. Mehmet Utku Celik, Gaurav Sharma 0001, Eli Saber, A. Murat Tekalp |
ICIP (2) | 4 |
| 2002 | Integrated semantic-syntactic video event modeling for search and retrievalabstractThis paper proposes to integrate text-based database models (commonly used in the database and information retrieval literature) and low-level video features, such as object-based motion features (commonly used in image processing and computer vision), within a single framework to describe video events for fast and effective browsing and retrieval using a mixture of textual and low-level descriptors. The instantiations of the model are expressed as graphs. Users can formulate flexible queries that can be translated into such graphs. Alternatively, users can input query graphs by editing an abstract model (model template) derived either from a model database or from an example scene in the media database. Search and retrieval is accomplished by matching the query graph with those instantiated models in the database. The proposed approach allows for integrated textual and low-level, descriptor-based video search, where the integration is enabled by the model. Examples and experimental results are provided to demonstrate the effectiveness of the proposed integrated modeling and querying framework. Ahmet Ekin, Rajiv Mehrotra, A. Murat Tekalp |
ICIP (1) | 3 |
| 2002 | Performance analysis of a kinematic human motion modelabstractWe describe a thorough quantitative analysis of a novel, kinematic model of the human body used for the tracking and extraction of gait variables. The model includes multiple levels of structural complexity coupled with absolute and probabilistic kinematic constraints; while the relevant variables include velocity, stride length, stance width, and stance times. The analysis is based on the accuracy of extracted gait variables for various permutations of the proposed kinematic body model. Using patterns of complex human motion collected from a multi-camera video monitoring system, we demonstrate the precise costs and benefits associated with various model components and kinematic constraints. Shiloh L. Dockstader, Michel J. Berg, A. Murat Tekalp |
ICME (1) | 3 |
| 2002 | Framework for tracking and analysis of soccer video
Ahmet Ekin, A. Murat Tekalp |
VCIP | 2 |
| 2002 | Interactive Optimization of 3D Shape and 2D Correspondence Using Multiple Geometric Constraints via POCSabstractThe traditional approach of handling motion tracking and structure from motion (SFM) independently in successive steps exhibits inherent limitations in terms of achievable precision and incorporation of prior geometric constraints about the scene. This paper proposes a projections onto convex sets (POCS) framework for iterative refinement of the measurement matrix in the well-known factorization method to incorporate multiple geometric constraints about the scene, thereby improving the accuracy of both 2D feature point tracking and 3D structure estimates. Regularities in the scene, such as points on line and plane and parallel lines and planes, that can be interactively identified and marked at each POCS iteration, enforce rank and parallelism constraints on appropriately defined local measurement matrices, one for each constraint. The POCS framework allows for the integration of the information in each of these local measurement matrices into a single measurement matrix that is "closest" to the initial observed measurement matrix in Frobenius norm, which is then factored in the usual manner. Experimental results demonstrate that the proposed interactive POCS framework consistently improves both 2D correspondences and 3D shape/motion estimates and similar results cannot be achieved by enforcing these constraints as either post or preprocessing. Zhaohui Sun, A. Murat Tekalp, Nassir Navab, Visvanathan Ramesh |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Robust watermarking of fingerprint images
Bilge Günsel, Umut Uludag, A. Murat Tekalp |
Pattern Recognit. | 3 |
| 2002 | Hierarchical watermarking for secure image authentication with localizationabstractSeveral fragile watermarking schemes presented in the literature are either vulnerable to vector quantization (VQ) counterfeiting attacks or sacrifice localization accuracy to improve security. Using a hierarchical structure, we propose a method that thwarts the VQ attack while sustaining the superior localization properties of blockwise independent watermarking methods. In particular, we propose dividing the image into blocks in a multilevel hierarchy and calculating block signatures in this hierarchy. While signatures of small blocks on the lowest level of the hierarchy ensure superior accuracy of tamper localization, higher level block signatures provide increasing resistance to VQ attacks. At the top level, a signature calculated using the whole image completely thwarts the counterfeiting attack. Moreover, "sliding window" searches through the hierarchy enable the verification of untampered regions after an image has been cropped. We provide experimental results to demonstrate the effectiveness of our method. Mehmet Utku Celik, Gaurav Sharma 0001, Eli Saber, A. Murat Tekalp |
IEEE Trans. Image Process. | 4 |
| 2002 | Robust color histogram descriptors for video segment retrieval and identificationabstractEffective and efficient representation of color features of multiple video frames or pictures is an important yet challenging task for visual information management systems. Key frame-based methods to represent the color features of a group of frames (GoF) are highly dependent on the selection criterion of the representative frame(s), and may lead to unreliable results. We present various histogram-based color descriptors to reliably capture and represent the color properties of multiple images or a GoF. One family of such descriptors, called alpha-trimmed average histograms, combine individual frame or image histograms using a specific filtering operation to generate robust color histograms that can eliminate the adverse effects of brightness/color variations, occlusion, and edit effects on the color representation. We show the efficacy of the alpha-trimmed average histograms for video segment retrieval applications, and illustrate how they consistently outperform key frame-based methods. Another color histogram descriptor that we introduce, called the intersection histogram, reflects the number of pixels of a given color that is common to all the frames in the GoF. We employ the intersection histogram to develop a fast and efficient algorithm for identification of the video segment to which a query frame belongs. The proposed color histogram descriptors have been included in the ISO standard MPEG-7 after extensive evaluation experiments. A. Müfit Ferman, A. Murat Tekalp, Rajiv Mehrotra |
IEEE Trans. Image Process. | 2 |
| 2002 | Temporal segmentation of video objects for hierarchical object-based motion descriptionabstractThis paper describes a hierarchical approach for object-based motion description of video in terms of object motions and object-to-object interactions. We present a temporal hierarchy for object motion description, which consists of low-level elementary motion units (EMU) and high-level action units (AU). Likewise, object-to-object interactions are decomposed into a hierarchy of low-level elementary reaction units (ERU) and high-level interaction units (IU). We then propose an algorithm for temporal segmentation of video objects into EMUs, whose dominant motion can be described by a single representative parametric model. The algorithm also computes a representative (dominant) affine model for each EMU. We also provide algorithms for identification of ERUs and for classification of the type of ERUs. Experimental results demonstrate that segmenting the life-span of video objects into EMUS and ERUs facilitates the generation of high-level visual summaries for fast browsing and navigation. At present, the formation of high-level action and interaction units is done interactively. We also provide a set of query-by-example results for low-level EMU retrieval from a database based on similarity of the representative dominant affine models. Ahmet Ekin, A. Murat Tekalp, Rajiv Mehrotra |
IEEE Trans. Image Process. | 3 |
| 2001 | Non-Rigid Object Tracking using Performance Evaluation Measures as FeedbackabstractWe present a scalable object tracking framework which is capable of tracking the contour of rigid and non-rigid objects in the presence of occlusion. The method adaptively divides the object contour into sub-contours, and employs several low-level features such as color edge, color segmentation, motion models, motion segmentation, and shape continuity information in a feedback loop to track each sub-contour. We also introduce some novel performance evaluation measures to evaluate the goodness of the segmentation and tracking. The results of these performance measures are utilized in a feedback loop to adjust the weights assigned to each of these low-level features for each sub-contour at each frame. The framework is scalable because it can be adapted to roughly track simple objects in real-time as well as pixel-accurate tracking of more complex objects in offline mode. The proposed method does not depend on any single motion or shape model, and does not need training. Experimental results demonstrate that the algorithm is able to track the object boundaries accurately under significant occlusion and background clutter. Çigdem Eroglu Erdem, Bülent Sankur, A. Murat Tekalp |
CVPR (2) | 3 |
| 2001 | Extraction of semantic description of events using Bayesian networksabstractWe use Bayesian belief networks to statistically model the trends for event detection. We automatically detect non-rigid object trajectories for object motion units. Then, we use dominant and secondary trajectories of a single object in several consecutive motion units to understand semantic actions or those of more than one object to recognize semantic interactions between objects. We demonstrate sample Bayesian networks to detect events and extract the event descriptions, such as "catch the ball", "throw the ball" and "walk". Ahmet Ekin, A. Murat Tekalp, Rajiv Mehrotra |
ICASSP | 2 |
| 2001 | A hierarchical image authentication watermark with improved localization and securityabstractSeveral fragile watermarking schemes presented in the literature are either vulnerable to vector quantization (VQ) counterfeiting attacks or sacrifice localization accuracy to improve security. Using a hierarchical structure, we propose a method that thwarts the VQ attack while sustaining the superior localization properties of blockwise independent watermarking methods. In particular, we propose dividing the image into blocks in a multi-level hierarchy and calculating block signatures in this hierarchy. While signatures of small blocks on the lowest level of the hierarchy ensure superior accuracy of tamper localization, higher level block signatures provide increasing resistance to VQ attacks. At the top level, a signature calculated using the whole image completely thwarts the counterfeiting attack. Moreover, "sliding window" searches through the hierarchy enable the verification of untampered regions after an image has been cropped. Mehmet Utku Celik, Gaurav Sharma 0001, Eli Saber, A. Murat Tekalp |
ICIP (2) | 4 |
| 2001 | Multi-view spatial integration and tracking with Bayesian networksabstractWe present a novel method for the spatial integration of multiple views as a means for tracking point features in the presence of occlusion. The proposed technique employs a dynamic, multi-dimensional Bayesian network to combine information from multiple views. To achieve real-time performance, the system is implemented in a distributed fashion; the two-dimensional tracking for each view, as well as the spatial integration, occurs on a dedicated processor. We demonstrate the efficacy of the proposed spatial integration on the multi-view tracking of a person in a home environment. Our results show a considerable increase in the accuracy of tracking features throughout periods of occlusion. Shiloh L. Dockstader, A. Murat Tekalp |
ICIP (1) | 2 |
| 2001 | Automatic extraction of low-level object motion descriptorsabstractIn our previous work, we developed a method to extract low-level elementary motion units (EMU) and elementary reaction units (ERU) for object-based event description, assuming perfect video object segmentation and tracking. This paper features the following contributions. (1) We propose a novel object tracking algorithm for a specific domain. (2) We evaluate the performance of our EMU and ERU extraction system with automatically-computed but imperfect video object information obtained by the proposed tracker. It is assumed that objects for which descriptors are sought for are interactively marked on the initial frame. (3) We extend the description of ERUs to enable better description of low-level reactions. Experimental results are provided to demonstrate each contribution. Ahmet Ekin, Rajiv Mehrotra, A. Murat Tekalp |
ICIP (2) | 3 |
| 2001 | Metrics for performance evaluation of video object segmentation and tracking without ground-truthabstractWe present metrics to evaluate the performance of video object segmentation and tracking methods quantitatively when ground-truth segmentation maps are not available. The proposed metrics are based on the color and motion differences along the boundary of the estimated video object plane and the color histogram differences between the current object plane and its temporal neighbors. These metrics can be used to localize (spatially and/or temporally) regions where segmentation results are good or bad; or combined to yield a single numerical measure to indicate the goodness of the boundary segmentation and tracking results. Experimental results are presented to evaluate the segmentation map of the "Man" object in the "Hall Monitor" sequence both in terms of a single numerical measure, as well as localization of the good and bad segments of the boundary. Çigdem Eroglu Erdem, Bülent Sankur, A. Murat Tekalp |
ICIP (2) | 3 |
| 2001 | Error Characterization of the Factorization Method
Zhaohui Sun, Visvanathan Ramesh, A. Murat Tekalp |
Comput. Vis. Image Underst. | 3 |
| 2001 | Multiple camera tracking of interacting and occluded human motionabstractWe propose a distributed, real-time computing platform for tracking multiple interacting persons in motion. To combat the negative effects of occlusion and articulated motion we use a multiview implementation, where each view is first independently processed on a dedicated processor. This monocular processing uses a predictor-corrector filter to weigh reprojections of three-dimensional (3-D) position estimates, obtained by the central processor, against observations of measurable image motion. The corrected state vectors from each view provide input observations to a Bayesian belief network, in the central processor, with a dynamic, multidimensional topology that varies as a function of scene content and feature confidence. The Bayesian net fuses independent observations from multiple cameras by iteratively resolving independency relationships and confidence levels within the graph, thereby producing the most likely vector of 3-D state estimates given the available data. To maintain temporal continuity, we follow the network with a layer of Kalman filtering that updates the 3-D state estimates. We demonstrate the efficacy of the proposed system using a multiview sequence of several people in motion. Our experiments suggest that, when compared with data fusion based on averaging, the proposed technique yields a noticeable improvement in tracking accuracy. Shiloh L. Dockstader, A. Murat Tekalp |
Proc. IEEE | 2 |
| 2001 | 2-D mesh-based video object segmentation and tracking with occlusion resolution
Isil Celasun, A. Murat Tekalp, Mete H. Gökçetekin, Derin M. Harmanci |
Signal Process. Image Commun. | 2 |
| 2000 | Adaptive motion estimation using local measures of texture and similarityabstractTraditional approaches to the estimation of motion in video sequences have relied on the appropriate selection of various algorithm parameters. This dependence becomes a prohibitive drawback in applications where automation is desirable or necessary or in sequences where a single set of parameters can not achieve sufficiently accurate results. We investigate a number of techniques for locally adapting both the spatio-temporal filters and the hierarchical structure used in the estimation of optical flow. The surviving technique utilizes projected active contours and gradient-based Chamfer distance images to adapt the filters and a temporally-based Kolmogorov-Smirnov metric to locally adapt the hierarchical structure. The advantages of using these adaptive variations are demonstrated on articulated and self-occluding motion. Shiloh L. Dockstader, A. Murat Tekalp |
ICASSP | 2 |
| 2000 | 2D mesh-based detection and representation of an occluding object for object-based videoabstractIn this study, an algorithm for mesh-based detection of occlusion caused by a newly entering object into the scene, which covers the information present in the current frame, and mesh-based representation of it is proposed. A 2D Delaunay triangulated dynamic mesh is initially designed on the first frame of the sequence. The motion of each node is then compared to its average motion. Frames with nodes of high activity, with different directions and forming a region are selected to be analyzed for detection of newly entering object(s) into the scene. A region formed by detection of bad motion vectors is enlarged using a distance criterion. The detected frame and the preceding one are range filtered, The luminance components of these two frames are formed. The difference of the range filtered frames and of their respective luminance components are taken into account. The differences are checked with respect to a threshold value inside the formed region. Pixels exceeding this threshold form the newly entering object. Since there may be separate regions formed by these pixels, mesh-based merging of these regions is then accomplished. The detected newly entering object is then meshed and tracked as a new object in the scene in accordance with the occluded object. The proposed 2D mesh-based occlusion detection and representation method can be applied in object-based video coding, storage and manipulation. Mete H. Gökçetekin, Isil Celasun, A. Murat Tekalp |
ICASSP | 3 |
| 2000 | Object based image retrieval based on multi-level segmentationabstractCurrently, image retrieval systems are based on low-level features of color, texture and shape, not on the semantic descriptions that are common to humans, such as objects, people, and place. In order to narrow down the gap between the low level and semantic level, object-based content analysis, which segments the semantically meaningful objects of images, is an essential step. In this study, we propose a learning process in order to perform effective automatic off-line analysis on a multi-level segmented image stack. Meaningful objects are extracted given certain user search patterns and interest profiles. Color and/or shape information of the objects is stored in the hierarchical content representations of the images. This information is utilized by a hierarchical matching scheme to improve the retrieval speed in the subsequent searches. Yaowu Xu, Pinar Duygulu, Eli Saber, A. Murat Tekalp, Fatos T. Yarman-Vural |
ICASSP | 4 |
| 2000 | Parametric Description of Object Motion Using EMUsabstractLarge scale deployment of digital multimedia applications requires effective and efficient image and video representations for indexing. We present a hierarchical representation of video objects based on their motion. The proposed representation enables both low and semantic level description of object motion. At the low level, we consider a parametric description of the dominant motion of video objects, and define elementary motion units (EMU) where the dominant motion of the object is coherent and does not undergo considerable change. A group of EMUs forms an action unit which carries semantic information. Interactions between objects are also considered by calculating the relative motion between objects in an object-based framework. Experimental results for retrieval of EMUs based on similarity of dominant object motion are demonstrated. Ahmet Ekin, Rajiv Mehrotra, A. Murat Tekalp |
ICIP | 3 |
| 2000 | Group-of-Frames/Pictures Color Histogram Descriptors for Multimedia ApplicationsabstractJoint representation of color-based features for multiple images or a video segment is an important task for visual information management systems. Generally, a key-frame or key-image is selected from such a group, and the color-related features of the entire collection are represented with those of the chosen sample. Such methods are highly dependent on the quality of the representative sample, and may lead to unreliable results. We present a set of histogram-based descriptors that reliably capture the color content of multiple images or video frames. These descriptors are defined for a group-of-frames (GoF) or a group-of-pictures (GoP). A single representation for the entire collection is obtained by combining individual frame or image histograms in various ways. We demonstrate the efficacy of GoF-histograms for video segment retrieval and the GoP-histograms for fast image search. This descriptor has been accepted to the Working Draft of MPEG-7, the evolving ISO standard for multimedia content description. A. Müfit Ferman, Santhana Krishnamachari, A. Murat Tekalp, Mohamed Abdel-Mottaleb, Rajiv Mehrotra |
ICIP | 3 |
| 2000 | Mesh-Based Segmentation and Update for Object-Based VideoabstractThere is no normative segmentation method in the emerging MPEG-4 standard. In this paper, we propose mesh-based segmentation method for object-based video which fuses mesh-based motion and color information using region approaches instead of pixel-based ones suitable for MPEG-4. A 2D Delaunay triangulated dynamic content-based mesh is initially designed on the first frame of the sequence with an optimal number of nodes. Nodes motion values and motion values of their neighbors are used to form a test region of triangles for the determination of the boundary of the video object. The formed test region is then remeshed and the whole mesh is updated. Color information of the neighboring triangles on one step layers provides us with a refinement of the boundary of the video object. The inside of the refined boundary is optimally meshed and a search mechanism is used through layers. The mesh based system can handle multiple video objects and the related results are given. The proposed 2D mesh based segmentation, occlusion detection and representation method suitable for progressive transmission can be applied in object-based video coding, storage and manipulation. Mete H. Gökçetekin, Mehmet Derin Harmanci, Isil Celasun, A. Murat Tekalp |
ICIP | 4 |
| 2000 | Image Retrieval Through Shape Matching of Partially Occluded Objects Using Hierarchical Content DescriptionabstractThis paper proposes a new contour based shape matching approach capable of recognizing partially occluded objects in images. The process can be divided into the following steps: region formation from low level color image segmentation, B-spline filtering and feature point extraction, correspondence determination, least square estimation of affine transformation parameters, and similarity measuring. Once similarity is established, the information is retained using a hierarchical content description scheme, enabling expedient object based image retrieval at a later time. Yaowu Xu, Eli Saber, A. Murat Tekalp |
ICIP | 3 |
| 2000 | Face and 2-D mesh animation in MPEG-4abstractThis paper presents an overview of some of the synthetic visual objects supported by MPEG-4 version-1, namely animated faces and animated arbitrary 2D uniform and Delaunay meshes. We discuss both specification and compression of face animation and 2D-mesh animation in MPEG-4. Face animation allows to animate a proprietary face model or a face model downloaded to the decoder. We also address integration of the face animation tool with the text-to-speech interface (TTSI), so that face animation can be driven by text input. A. Murat Tekalp, Jörn Ostermann |
Signal Process. Image Commun. | 1 |
| 2000 | Optimal 2-D hierarchical content-based mesh design and update for object-based videoabstractRepresentation of video objects (VOs) using hierarchical 2-D content-based meshes for accurate tracking and level of detail (LOD) rendering have been previously proposed, where a simple suboptimal hierarchical mesh design algorithm was employed. However, it was concluded that the performance of the tracking and rendering very much depends on how well each level of the hierarchical mesh structure fits the VO under consideration. To this effect, this paper proposes an optimized design of hierarchical 2-D content-based meshes with a shape-adaptive simplification and a temporal update mechanism for object-based video. Particular contributions of this work are: (1) analysis of optimal number of nodes for the initial fine level-of-detail mesh design; (2) adaptive shape simplification across hierarchy levels; (3) optimization of the interior-node decimation method to remove only a maximal independent set to preserve Delaunay topology across hierarchy levels for better bitrate versus quality performance; and (4) a mesh-update mechanism which serves to update a temporally 2-D dynamic mesh in case of occlusion due to 3-D motion and self-occlusion. The proposed optimized and temporally updated hierarchical mesh representation can be applied in object-based video coding, retrieval, and manipulation. Isil Celasun, A. Murat Tekalp |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Semi-automatic video object segmentation in the presence of occlusionabstractWe describe a semi-automatic approach for segmenting a video sequence into spatio-temporal video objects in the presence of occlusion. Motion and shape of each video object is represented by a 2-D mesh. Assuming that the boundary of an object of interest is interactively marked on some keyframes, the proposed method finds the boundary of the object in all other frames automatically by tracking the 2-D mesh representation of the object in both forward and backward directions. A key contribution of the proposed method is automatic detection of covered and uncovered regions at each frame, and assignment of pixels in the uncovered regions to the object or background based on color and motion similarity. Experimental results are presented on two MPEG-4 test sequences and the resulting segmentations are evaluated both visually and quantitatively. Candemir Toklu, A. Murat Tekalp, A. Tanju Erdem |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Tracking visible boundary of objects using occlusion adaptive motion snakeabstractWe propose a novel technique for tracking the visible boundary of a video object in the presence of occlusion. Starting with an initial contour that is interactively specified by the user and may be automatically refined by using intra-energy terms, the proposed technique employs piecewise contour prediction using local motion and color information on both sides of the contour segment, and contour snapping using scale-invariant intra-frame and inter-frame energy terms. The piecewise (segmented) nature of the contour prediction scheme and modeling of the motion on both sides of each contour segment enable accurate determination of whether and where the tracked boundary is occluded by another object. The proposed snake energy terms are associated with contour segments (as opposed to node points) and they are scale/resolution independent to allow multi-resolution contour tracking without the need to retune the weights of the energy terms at each resolution level. This facilitates contour prediction at coarse resolution and snapping at fine resolution with high accuracy. Experimental results are provided to illustrate the performance of the proposed occlusion detection algorithm and the novel snake energy terms that enable visible boundary tracking in the presence of occlusion. A. Tanju Erdem, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 2000 | Guest editorial introduction to the special issue on image and video processing for digital libraries
B. S. Manjunath, Thomas S. Huang, A. Murat Tekalp, HongJiang Zhang |
IEEE Trans. Image Process. | 3 |
| 2000 | Two-dimensional mesh-based mosaic representation for manipulation of video objects with occlusionabstractWe present a two-dimensional (2-D) mesh-based mosaic representation, consisting of an object mesh and a mosaic mesh for each frame and a final mosaic image, for video objects with mildly deformable motion in the presence of self and/or object-to-object (external) occlusion. Unlike classical mosaic representations where successive frames are registered using global motion models, we map the uncovered regions in the successive frames onto the mosaic reference frame using local affine models, i.e., those of the neighboring mesh patches. The proposed method to compute this mosaic representation is tightly coupled with an occlusion adaptive 2-D mesh tracking procedure, which consist of propagating the object mesh frame to frame, and updating of both object and mosaic meshes to optimize texture mapping from the mosaic to each instance of the object. The proposed representation has been applied to video object rendering and editing, including self transfiguration, synthetic transfiguration, and 2-D augmented reality in the presence of self and/or external occlusion. We also provide an algorithm to determine the minimum number of still views needed to reconstruct a replacement mosaic which is needed for synthetic transfiguration. Experimental results are provided to demonstrate both the 2-D mesh-based mosaic synthesis and two different video object editing applications on real video sequences. Candemir Toklu, A. Tanju Erdem, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 1999 | Keyframe-Based Bi-Directional 2-D Mesh Representation for Video Object Tracking and ManipulationabstractWe propose a new bi-directional 2-D mesh representation of video objects, which utilizes multiple keyframes with forward and backward tracking. Experimental results on use of this representation for video object tracking in the presence of self occlusion are presented. P. Erhan Eren, A. Murat Tekalp |
ICIP (2) | 2 |
| 1999 | Probabilistic Analysis and Extraction of Video ContentabstractIn this paper we present a probabilistic framework for mapping low-level visual features into a specific set of semantic descriptors. Specifically, we employ hidden Markov models (HMMs) and Bayesian belief networks (BBNs) at various stages to characterize content domains and extract the relevant semantic information. HMMs are utilized at the shot and sequence levels to model the sequentially-varying structure of video sequences and delineate the video stream in terms of the constituent shots. BBNs, on the-other hand, act on and within each shot, to provide more detailed descriptions of shot content using the physical features of video objects. The semantic content extraction problem is thus addressed at all physical (shot and object) levels, within a consistent representation and processing framework. A. Müfit Ferman, A. Murat Tekalp |
ICIP (2) | 2 |
| 1999 | Object Formation by Learning in Visual Databases Using Hierarchical Content DescriptionabstractThis paper proposes a self-learning content-based image indexing and retrieval system that employs a hierarchical content representation (consisting of objects and regions) and a hierarchical content matching method for effective and efficient image/object retrieval. The “learning” behavior is enabled by our proposed hierarchical content representation which allows easy storage of combinations of regions that have resulted in successful matches to objects of interest as determined by user search patterns and profiles. The learning step effectively performs an automatic off-line analysis of database images into meaningful objects. Once the learning phase is complete, the speed of shape based retrieval of the learned objects in the database increases significantly. Experimental results are presented to show the effectiveness of the proposed hierarchical content representation, hierarchical matching, and the learning behavior on collections of car images. Yaowu Xu, Eli Saber, A. Murat Tekalp |
ICIP (2) | 3 |
| 1999 | Hierarchical 2-D mesh representation, tracking, and compression for object-based videoabstractThis paper proposes methods for designing, tracking and coding hierarchical two-dimensional (2-D) content-based mesh representations. The design procedure consists of constructing a fine-to-coarse hierarchy of Delaunay meshes, using image- and shape-based criteria for mesh geometry simplification. Hierarchical tracking employs a coarse-to-fine strategy with mesh-based motion vector optimization. We introduce new techniques to maintain the initial mesh hierarchy and topology during tracking by imposing certain constraints at each stage of the procedure. The hierarchical compression technique is based on a nearest neighbor ordering of mesh node points. This ordering serves to identify the mesh boundary nodes as well as establish spatial predictors for differential coding of node coordinates and motion vectors. The proposed hierarchical mesh representation, which has applications in object-based video manipulation, investing, and compression, provides improved tracking performance (compared to a nonhierarchical representation) and allows progressive (scalable) transmission of the object geometry (including shape) and motion information, as well as variable level-of-detail rendering. Experimental results are presented to compare the tracking and compression performance of hierarchical versus nonhierarchical mesh representations and to demonstrate the tradeoff between image quality and mesh bit rate for 2-D mesh-based video object rendering. Peter J. L. van Beek, A. Murat Tekalp, Ning Zhuang, Isil Celasun, Minghui Xia |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | End-to-end color printer calibration by total least squares regressionabstractNeugebauer modeling plays an important role in obtaining end-to-end device characterization profiles for halftone color printer calibration. This paper proposes total least square (TLS) regression methods to estimate the parameters of various Neugebauer models. Compared to the traditional least squares (LS) based methods, the TLS approach is physically more appropriate for the printer modeling problem because it accounts for errors in the measured reflectance of both the primaries and the modeled samples. A TLS method based on print measurements from single-colorant step-wedges is first developed. The method is then extended to incorporate multicolorant print measurements using an iterative algorithm. The LS and TLS techniques are compared through tests performed on two color printers, one employing conventional rotated halftone screens and the other using a dot-on-dot halftone screen configuration. Our experiments indicate that the TLS methods yield a consistent and significant improvement over the LS-based techniques for model parameter estimation. The gains from the TLS method are particularly significant when the number of patches for which measured data is available is limited. Minghui Xia, Eli Saber, Gaurav Sharma 0001, A. Murat Tekalp |
IEEE Trans. Image Process. | 4 |
| 1998 | Parametric motion modeling based on trilinear constraints for object-based video compressionabstractWe propose a new parametric motion model based on the so-called "trifocal tensor" representation, which captures rigid 3D motion of static scenes with a depth of field. The estimation of the trifocal tensor requires solution of a set of linear equations given at least seven point correspondences across three frames. The proposed parametric representation, called the trilinear model, is superior to other forms such as translational, affine, perspective, and bilinear models, because it can implicitly encode the depth of the scene and 3D motion of the scene/camera under perspective projection unlike others. A video object can thus be represented by its first VOP, a set of trifocal tensors and the corresponding prediction residues. Motion estimation and compensation based on the new parametric model are incorporated into the MPEG-4 Video Verification Model to compare its efficacy for object-based video compression with the state-of-the art motion compensation methods. Experimental results are provided to demonstrate the performance of the trilinear model for object-based video compression. Zhaohui Sun, A. Murat Tekalp |
ICASSP | 2 |
| 1998 | Optimal Hierarchical Design of 2D Dynamic Meshes for VideoabstractThis paper proposes methods for designing hierarchical 2D dynamic meshes, for representation of object-based video. This representation consists of a hierarchy of Delaunay meshes, obtained by recursive simplification of the initial fine level-of-detail mesh geometry. Nodes in the initial fine level-of-detail mesh are selected using an edge and corner detector. A dynamic programming-like approach is employed in the mesh simplification procedure to obtain an optimal hierarchical design. Mesh simplification entails removal of mesh nodes to reduce the level of detail. The selection of nodes to be removed is achieved by associating a cost with each mesh node. The hierarchical mesh representation can be applied in object-based video coding, storage and manipulation. Isil Celasun, Erdal Ilgaz, A. Murat Tekalp, Peter J. L. van Beek, Ning Zhuang |
ICIP (2) | 3 |
| 1998 | Effective Content Representation for Video
A. Müfit Ferman, A. Murat Tekalp, Rajiv Mehrotra |
ICIP (3) | 2 |
| 1998 | Occlusion Adaptive Motion Snake
A. Tanju Erdem, A. Murat Tekalp |
ICIP (3) | 3 |
| 1998 | Content-based Video Abstraction
Bilge Günsel, A. Murat Tekalp |
ICIP (3) | 2 |
| 1998 | A High-Performance Shot Boundary Detection Algorithm using Multiple CuesabstractA central step in content-based video retrieval is the temporal segmentation of video. An application independent approach to video segmentation is to detect temporally contiguous segments without significant content change between successive frames. Each such segment is termed a shot. A high-performance shot boundary detection-based video segmentation algorithm is proposed. The technique uses unsupervised clustering on a multiple feature input space, followed by a heuristic elimination process to detect, with almost perfect accuracy, shot boundaries in the video. With an extremely high accuracy coupled with a very small number of false positives, this algorithm outperforms most of the existing techniques. Milind R. Naphade, Rajiv Mehrotra, A. Müfit Ferman, Jim Warnick, Thomas S. Huang, A. Murat Tekalp |
ICIP (1) | 6 |
| 1998 | Image Registration using a 3-D Scene RepresentationabstractWe propose a trifocal motion model, which captures rigid scene/camera motion and 3-D scene structure under perspective projection, to register multiple images without camera calibration. Three images/frames are processed at a time to register the current frame F/sub i/ with a common reference frame R by using a common correspondence frame C involving only linear operations. The registration method requires estimation of 27 trifocal tensor parameters between every tripled of frames (R, C, F/sub i/), and a dense correspondence field between the frames R and C. The proposed image registration method can be employed in mosaic synthesis, superresolution, and video standard conversion. The superiority of the proposed trifocal image registration over affine (6 parameters) and perspective (8 parameters) registration is demonstrated by experimental results. Zhaohui Sun, A. Murat Tekalp |
ICIP (1) | 2 |
| 1998 | Total Least Square Techniques in Color Printer CharacterizationabstractThe Neugebauer model is a powerful tool in obtaining end-to-end device characterization profiles for halftone color printer calibration. In this paper, we propose total least square (TLS) regression methods to estimate the parameters of Neugebauer models. Compared to the traditional least squares (LS) based methods, the TLS approach is a physically more appropriate procedure, because it accounts for errors in the measured reflectance of both the selected primaries and the modeled reflectance. The proposed TLS techniques are tested on a Xerox color printer with rotated halftone screen, and the results are compared with the LS based algorithms. Our experiments indicate that the TLS methods yield a significant improvement over the LS based techniques for model parameter estimation. Minghui Xia, Eli Saber, Gaurav Sharma 0001, A. Murat Tekalp, Ronald Sosinski |
ICIP (2) | 4 |
| 1998 | Interactive object-based analysis and manipulation of digital videoabstractWith the advent of MPEG-4, object-based natural/synthetic hybrid multimedia content is becoming more ubiquitous. In this paper, we address object-based interactive analysis of natural video for editing/authoring natural/synthetic hybrid content. Boundary and local motion of video objects are described by snake and 2-D mesh representations, respectively. The 2-D mesh modeling in effect performs a mapping of a natural video object into a computer graphics representation, namely geometry with motion and a texture map; thus allowing for easy editing of natural video objects using tools already developed in computer graphics. This paper presents the components of a tool designed for interactive video object analysis and editing using graphical models whose syntax and semantics conform with VRML and MPEG-4 standards. The analysis tool is developed using Java and C++, while rendering and editing are performed within VRML or MPEG-4 browsers and authoring tools. Demonstrations of the video analysis and editing are included. P. Erhan Eren, Ning Zhuang, A. Murat Tekalp |
MMSP | 4 |
| 1998 | Region-Based Parametric Motion Segmentation Using Color InformationabstractThis paper presents pixel-based and region-based parametric motion segmentation methods for robust motion segmentation with the goal of aligning motion boundaries with those of real objects in a scene. We first describe a two-step iterative procedure for parametric motion segmentation by either motion-vector or motion-compensated intensity matching. We next present a region-based extension of this method, whereby all pixels within a predefined spatial region are assigned the same motion label. These predefined regions may be fixed- or variable-size blocks or arbitrary-shaped areas defined by color or texture uniformity. A particular combination of these pixel-based and region-based methods is then proposed as a complete algorithm to obtain the best possible segmentation results on a variety of image sequences. Experimental results showing the benefits of the proposed scheme are provided. Yücel Altunbasak, P. Erhan Eren, A. Murat Tekalp |
Graph. Model. Image Process. | 3 |
| 1998 | Efficient Filtering and Clustering Methods for Temporal Video Segmentation and Visual SummarizationabstractAutomatic temporal segmentation and visual summary generation methods that require minimal user interaction are key requirements in video information management systems. Clustering presents an ideal method for achieving these goals, as it allows direct integration of multiple information sources. This paper proposes a clustering-based framework to achieve these tasks automatically and with a minimum of user-defined parameters. The use of multiple frame difference features and short-time techniques are presented for efficient detection of cut-type shot boundaries. Generic temporal filtering methods are used to process the signals used in shot boundary detection, resulting in better suppression of false alarms. Clustering is also extended to the key frame extraction problem: Color-based shot representations are provided by average and intersection histograms, which are then used in a clustering scheme to identify reference key frames within each slot. The technique achieves good compaction with a minimum number of visually nonredundant key frames. A. Müfit Ferman, A. Murat Tekalp |
J. Vis. Commun. Image Represent. | 2 |
| 1998 | Special Issue On Multimedia Signal Processing, Part I [Scanning the Issue]abstract10.1109/JPROC.1998.664271 A. Murat Tekalp |
Proc. IEEE | 1 |
| 1998 | Two-dimensional mesh-based visual-object representation for interactive synthetic/natural digital videoabstractThis paper first provides an overview of two-dimensional (2-D) and three-dimensional mesh models for digital video processing. It then introduces 2-D mesh-based modeling of video objects as a compact representation of motion and shape for interactive, synthetic/natural video manipulation, compression, and indexing. The 2-D mesh representation and the mesh geometry and motion compression have been included in the visual tools of the upcoming MPEG-4 standard. Functionalities enabled by 2-D mesh-based visual-object representation include animation of still texture maps, transfiguration of video overlays, video morphing, and shape-and motion-based retrieval of video objects. A. Murat Tekalp, Peter J. L. van Beek, Candemir Toklu, Bilge Günsel |
Proc. IEEE | 1 |
| 1998 | Shape similarity matching for query-by-exampleabstractThis paper describes a unified approach for two-dimensional (2-D) shape matching and similarity ranking of objects by means of a modal representation. In particular, we propose a new shape-similarity metric in the eigenshape space for object/image retrieval from a visual database via query-by-example. This differs from prior work which performed point correspondence determination and similarity ranking of shapes in separate steps. The proposed method employs selected boundary and/or contour points of an object as a coarse-to-fine shape representation, and does not require extraction of connected boundaries or silhouettes. It is rotation-, translation- and scale-invariant, and can handle mild deformations of objects (e.g. due to partial occlusions or pose variations). Results comparing the unified method with an earlier two-step approach using B-spline-based modal matching and Hausdorff distance ranking are presented on retail and museum catalog style still-image databases. Bilge Günsel, A. Murat Tekalp |
Pattern Recognit. | 2 |
| 1998 | Frontal-view face detection and facial feature extraction using color, shape and symmetry based cost functionsabstractWe describe an algorithm for detecting human faces and facial features, such as the location of the eyes, nose and mouth. First, a supervised pixel-based color classifier is employed to mark all pixels that are within a prespecified distance of “skin color”, which is computed from a training set of skin patches. This color-classification map is then smoothed by Gibbs random field model-based filters to define skin regions. An ellipse model is fit to each disjoint skin region. Finally, we introduce symmetry-based cost functions to search the center of the eyes, tip of nose, and center of mouth within ellipses whose aspect ratio is similar to that of a face. Eli Saber, A. Murat Tekalp |
Pattern Recognit. Lett. | 2 |
| 1998 | Content-based access to video objects: Temporal Segmentation, visual summarization, and feature extractionabstractThe classical approach to content-based video access has been ‘frame-based’, consisting of shot boundary detection, followed by selection of key frames that characterize the visual content of each shot, and then clustering of the camera shots to form story units. However, in an object-based multimedia environment, content-based random access to individual video objects becomes a desirable feature. To this effect, this paper introduces an ‘object-based’ approach to temporal video partitioning and content-based indexing, where the basic indexing unit is ‘lifespan of a video object’, rather than a ‘camera shot’ or a ‘story unit’. We propose to represent each video object by an adaptive 2D triangular mesh. A mesh-based object tracking scheme is then employed to compute the motion trajectories of all mesh node points until the object exits the field of view. A new similarity measure that is based on motion discontinuities and shape changes of the tracked object is defined to detect content changes, resulting in temporal lifespan segments. A set of ‘key snapshots’ which constitute a visual summary of the lifespan of the object is automatically selected. These key snapshots are then used to animate objects of interest using tracked motion trajectories for a moving visual representation. The proposed scheme provides such functionalities as object-based search/browsing for interactive video retrieval, surveillance video analysis, and object-based content manipulation/editing for studio postprocessing and desktop multimedia authoring. The approach is applicable to any video data where the initial appearance of object(s) can be specified, and the object motion can be modeled by a piecewise affine transformation. The system is demonstrated using different types of video: virtual studio productions (composited video), surveillance video, and TV broadcast video. Der klassische Ansatz zu inhaltsorientiertem Videozugriff war “frame-basiert” und bestand aus der Detektion der Grenzen kurzer Aufnahmeabschnitte, gefolgt von der Auswahl wichtiger Frames, die den visuellen Inhalt jeder dieser Aufnahmeabschnitte charakterisieren, und dem Zusammenfügen dieser Kameraaufnahmen, um Filmeinheiten zu bilden. In einer objektorientierten Multimedia-Umgebung wird jedoch ein willkürlicher, inhaltsorientierter Zugriff auf individuelle Videoobjekte ein wünschenswertes Merkmal. Dazu wird in dieser Arbeit ein objektbasierter Ansatz zur zeitlichen Videopartitionierung und inhaltsbasierten Katalogisierung eingeführt, bei der die zugrundeliegende Katalogisierungseinheit die “Lebensdauer eines Videoobjektes” anstatt des “Aufnahmeabschnitts”oder der “Filmeinheit”ist. Wir schlagen vor, jedes Videoobjekt durch ein adaptives 2D Dreieckgitter zu repräsentieren. Es wird dann ein gitterbasiertes Verfahren zur Objektverfolgung verwendet, mit dem die Bewegungsbahnen aller Knotenpunkte des Gitters berechnet werden, bis das Objekt den Betrachtungsbereich verläßt. Eines neues Ähnlichkeitsmaß zum Erkennen von Änderungen des Inhalts, das auf Bewegungsdiskontinuitäten und Änderung der Gestalt des verfolgten Objekts basiert, wird definiert und führt zu Videosegmenten verschiedener Lebensdauer. Eine Anzahl von “Schlüsselaufnahmen”, die eine visuelle Zusammenfassung der Lebensdauer des Objektes darstellen, wird automatischen ausgewählt. Für eine visuelle Bewegungsdarstellung werden diese Schlüsselaufnahmen zur Animation interessierender Objekte durch die Benutzung der verfolgten Bewegungsbahnen verwendet. Das vorgeschlagene Verfahren stellt Funktionen wie objektorientiertes Suchen/Umsehen für interaktiven Videoempfang und Analyse von Überwachungsvideos, und objektorientierte Veränderung/Editierung des Inhalts für die Nachbearbeitung im Studio und zur Erstellung von Multimedia-Produkten am Rechner zur Verfügung. Der Ansatz läßt sich auf beliebige Videodaten anwenden, bei denen das erstmalige Auftreten eines Objekts festgelegt und die Bewegung des Objekts durch eine stückweise affine Transformation modelliert werden kann. Das System wird mit verschiedenen Videodaten vorgestellt: virtuelle Studioproduktionen (künstliches Video), Aufnahmen von Überwachungskameras und Fernschaufnahmen. L’approche classique de l’accès vidéo qui s’appuie sur le contenu est “basé sur des trames”, consistant en la détection de frontières de prises de vue suivic par la sélection de trames-clés qui caractérisent le contenu visuel de chaque prise, pour finir avec la coalescence des prises de vue pour former des unités de scène. Toutefois, dans un environnement multimédia basé sur des objets, l’accès aléatoire basé sur le contenu à des objets vidéo individuels devient une caractéristique souhaitable. A cet effet, ce papier introduit une approche “basée sur des objets” au partitionnement vidéo temporel et un indexage basé sur le contenu, où l’unité d’indexage de base est la “durée de vie d’un objet vidéo”, plutôt qu’une “prise de vue” ou une “unité de scène”. On propose de représenter chaque objet vidéo par un réseau triangulaire 2D adaptatif. Une méthode de poursuite d’objet basée sur un réseau est ensuite employée pour calculer les trajectoires des mouvements de chaque noeud du réseau jusqu’à ce que l’objet quitte le champ de vision. Une nouvelle mesure de similarité basée sur les discontinuités de mouvement et les changements de forme de l’objet poursuivi est définie de manière à détecter les changements de contenu et résultant en segments de durée de vie temporels. Un ensemble de “prises de vue-clés” constituant un résumé visuel de la durée de vie de l’objet est sélectionné automatiquement. Ces prises de vue-clés sont alors utilisées pour animer des objets d’intérêt en faisant appel aux trajectoires de mouvement poursuivies pour une repésentation visuelle mobile. La méthode proposée fournit de telles fonctionnalités en tant qu’outil de recherche/parcours pour la récupération vidéo interactive, l’analyse de vidéo-surveillance et la manipulation/édition de contenu basé sur des objets pour le post-traitement en studio et la classification multimédia. L’approche cst applicable à n’importe quelle séquence vidéo où l’apparence initiale des objets peut être spécifiée et le mouvement des objets modélisé par une transformation affine par morceaux. Le système est testé en utilisant différents types de vidéo: productions en studio virtuelles (vidéo composite), vidéo-surveillance et transmission télévisée. Bilge Günsel, A. Murat Tekalp, Peter J. L. van Beek |
Signal Process. | 2 |
| 1998 | Trifocal motion modeling for object-based video compression and manipulationabstractFollowing an overview of two-dimensional (2-D) parametric motion models commonly used in video manipulation and compression, we introduce trifocal transfer, which is an image-based scene representation used in computer vision, as a motion compensation method that uses three frames at a time to implicitly capture camera/scene motion and scene depth. Trifocal transfer requires a trifocal tensor that is computed by matching image features across three views and a dense correspondence between two of the three views. We propose approximating the dense correspondence between two of the three views by a parametric model in order to apply the trifocal transfer for object-based video compression and background mosaic generation. Backward, forward, and bidirectional motion compensation methods based on trifocal transfer are presented. The performance of the proposed motion compensation approaches using the trifocal model has been compared with various other compensation methods, such as dense motion, block motion, and global affine transform on several video sequences. Finally, video compression and mosaic synthesis based on the trifocal motion model are implemented within the MPEG-4 Video Verification Model (VM), and the results are compared with those of the standard MPEG-4 video VM. Experimental results show that the trifocal motion model is superior to block and affine models when there is depth variation and camera translation. Zhaohui Sun, A. Murat Tekalp |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | A new motion-compensated reduced-order model Kalman filter for space-varying restoration of progressive and interlaced videoabstractWe propose a new approach for motion-compensated, reduced order model Kalman filtering for restoration of progressive and interlaced video. In the case of interlaced inputs, the proposed filter also performs deinterlacing. In contrast to the literature, both motion-compensation and reduced-order state modeling are achieved by augmenting the observation equation, as opposed to modifying the state-transition equation. The proposed modeling, which includes the two-dimensional (2-D) reduced order model Kalman filtering (ROMKF) of Angwin and Kaufman as a special case, results in significant performance improvement in fixed-lag Kalman filtering of space-varying blurred images. This is demonstrated by experimental results. Andrew J. Patti, A. Murat Tekalp, M. Ibrahim Sezan |
IEEE Trans. Image Process. | 2 |
| 1997 | Object-Based Video Indexing for Virtual Studio ProductionsabstractThis paper introduces an object-based approach for temporal video partitioning and content-based indexing, where the basic indexing unit is "lifespan of a video object", rather than a "camera shot" or "story unit". We propose a system to extract content-based features of video objects (VOs), based on a compact 2D triangular mesh representation of them. An adaptive mesh-based video object tracking scheme is then employed to compute the motion trajectories of all node points. A set of "key snapshots" which constitute a visual summary of the lifespan of the object are automatically selected using motion and shape information. The system provides direct access to the VOs and gives the functionalities such as object-based search, manipulation, animation, and tracking. Bilge Günsel, A. Murat Tekalp, Peter J. L. van Beek |
CVPR | 2 |
| 1997 | Region-based affine motion segmentation using color informationabstractThis paper presents region-based affine motion parameter clustering methods by motion-vector and intensity matching to provide improved robustness and alignment of motion boundaries with real object boundaries. Regions may be formed as fixed- or variable-size rectangular blocks, triangular patches, or arbitrary-shaped areas defined by color or texture uniformity. A particular combination of these affine clustering methods is then proposed to obtain the best segmentation results on a variety of image sequences. Experimental results showing the benefits of the proposed scheme are provided. P. Erhan Eren, Yücel Altunbasak, A. Murat Tekalp |
ICASSP | 3 |
| 1997 | Motion and shape signatures for object-based indexing of MPEG-4 compressed videoabstractThe emerging MPEG-4 standard enables direct access to individual objects in the video stream, along with boundary/shape, texture, and motion information about each object. This paper proposes an object-based video indexing method that is directly applicable to the MPEG-4 compressed video bitstreams. The method aims to provide object-based content-interactivity; thus, defines the audio-visual object as the indexing unit. The scheme involves object-based temporal segmentation of the video bit-stream, selection of key-frames and key-video-object-planes, and characterization of the motion and/or shape of each video object including the background object. We also propose syntax and semantics for an indexing field to meet the content-based access requirement of MPEG-4. Experimental results are shown on two MPEG-4 test sequences. A. Müfit Ferman, Bilge Günsel, A. Murat Tekalp |
ICASSP | 3 |
| 1997 | 2-D mesh-based synthetic transfiguration of an object with occlusionabstractThis paper addresses 2-D mesh-based object tracking and mesh-based object mosaic construction for synthetic transfiguration of deformable video objects with deformable boundaries in the presence of another occluding object and/or self-occlusion. In particular, we update the 2-D triangular mesh model of a video object incrementally to account for the newly uncovered parts of the object as they are detected during the tracking process. Then, the minimum number of reference views (still images of a replacement object) needed to perform the synthetic transfiguration (object replacement and animation) is determined (depending on the complexity of the motion of the object-to-be-replaced), and the transfiguration of the replacement object is accomplished by 2-D mesh-based texture mapping in between these reference views. The proposed method is demonstrated by replacing an orange juice bottle by a cranberry juice bottle in a real video clip. Candemir Toklu, A. Tanju Erdem, A. Murat Tekalp |
ICASSP | 3 |
| 1997 | 2-D Mesh Geometry and Motion Compression for Efficient Object-Based Video RepresentationabstractMethods for object-based compression and composition of natural and synthetic video content are currently emerging in standards such as MPEG-4 and VRML. This paper describes novel techniques for compression of 2-D triangular mesh geometry and motion, enabling efficient representation and manipulation of video content. Specifically, mesh geometry is compressed by predictive coding of mesh node locations. Mesh node motion vectors are compressed by predictive techniques as well. Preliminary results show that the mesh data can be coded at a fraction of the bits used to code a typical video object. Peter J. L. van Beek, A. Murat Tekalp, Atul Puri |
ICIP (3) | 2 |
| 1997 | Special Effects Authoring Using 2-D Mesh ModelsabstractWe propose a video manipulation formalism that includes several special effects authoring tools such as animation, transfiguration, and augmented reality, based on a 2-D mesh-based video object representation. Texture maps of video objects are registered on a reference mesh, and then transfigured (rendered) by warping them using 2-D mesh mappings that are obtained by tracking the original sequence of video objects. The formalism also allows alpha blending of a moving object with another object using an alpha map sequence and a moving texture map. A VRML browser is used as an interactive user-interface to implement these manipulation tools. A Web demo is also available. P. Erhan Eren, Candemir Toklu, A. Murat Tekalp |
ICIP (1) | 3 |
| 1997 | Moving Visual Representations of Video Objects for Content-Based Search and BrowsingabstractThis paper proposes object-based moving visual representations for quick browsing of video content. These representations are hierarchical, such that at the coarse level a sequence of alpha planes provides a moving representation of object shape and motion information for object contours. Alternatively, a 2D mesh representation provides a complete visual representation of object motion and shape. The finest level visual representation can be obtained by texture mapping onto the moving meshes. The paper also discusses trade-offs between each representation in terms of the amount of indexing information that needs to be stored, the robustness of the representation, and the accuracy of the representation. Bilge Günsel, A. Murat Tekalp, Peter J. L. van Beek |
ICIP (2) | 2 |
| 1997 | Simultaneous Alpha Map Generation and 2-D Mesh Tracking for Multimedia ApplicationsabstractWe propose a semi-automatic alpha-map generation method, where video object boundaries that are marked interactively on some key-frames are tracked in all other frames using only YUV data to generate alpha-maps. We use a 2-D triangular mesh to represent and track video objects. One of the major problems in alpha plane generation by object tracking is the detection and classification of uncovered regions due to object motion and self-occlusion (especially in the presence of out-of-plane rotations). The proposed method can identify uncovered regions and then classify them as belonging to the object of interest or not based on color and motion similarities. Candemir Toklu, A. Murat Tekalp, A. Tanju Erdem |
ICIP (1) | 2 |
| 1997 | Object-based video manipulation and composition using 2D meshes in VRMLabstractThis paper deals with manipulation of natural video objects using graphical models and composition of processed objects within a VRML browser. In particular, 2-D mesh-based representation of video objects and resulting functionalities are integrated within VRML 2.0 for real-time rendering using IndexedFaceSet and CoordinateInterpolator nodes of VRML 2.0. A software demonstration of this representation and functionalities provided is shown. P. Erhan Eren, Candemir Toklu, A. Murat Tekalp |
MMSP | 3 |
| 1997 | Fusion of color and edge information for improved segmentation and edge linkingabstractWe propose a new method for combined color image segmentation and edge linking. The image is first segmented based on color information only. The segmentation map is modeled by a Gibbs random field, to ensure formation of spatially contiguous regions. Next, spatial edge locations are determined using the magnitude of the gradient of the 3-channel image vector field. Finally, regions in the segmentation map are split and merged by a region-labeling procedure to enforce their consistency with the edge map. The boundaries of the final segmentation map constitute a linked edge map. Experimental results are reported. Eli Saber, A. Murat Tekalp, Gozde Bozdagi Akar |
Image Vis. Comput. | 2 |
| 1997 | Region-Based Shape Matching for Automatic Image Annotation and Query-by-ExampleabstractWe present a method for automatic image annotation and retrieval based on query-by-example by region-based shape matching. The proposed method consists of two parts: region selection and shape matching. In the first part, the image is partitioned into disjoint, connected regions with more-or-less uniform color, whose boundaries coincide with spatial edge locations. Each region or valid combinations of neighboring regions constitute “potential objects.” In the second part, the shape of each potential object is tested to determine whether it matches one from a set of given templates. To this effect, we propose a new shape matching method, which is translation-, rotation-, and isotropic scale-invariant, where the boundary of each potential object, as well as of each template, is represented by a B-spline. We, then, identify correspondences between the joint points of the B-splines of potential objects and templates by using a modal matching method. These correspondences are used to estimate the parameters of an affine mapping to register the object with the template. A proximity measure is then computed between the two contours based on the Hausdorff distance. We demonstrate the performance of the proposed method on a variety of images. Eli Saber, A. Murat Tekalp |
J. Vis. Commun. Image Represent. | 2 |
| 1997 | Robust methods for high-quality stills from interlaced video in the presence of dominant motionabstractWe present robust algorithms which combine global motion compensation and motion adaption for deinterlacing in the presence of both dominant motion, such as camera zoom, pan, or jitter, and local motion, such as object motion. The dominant motion is modeled by a global affine warping and estimated by a gradient-based estimation method. Two alternative algorithms are proposed for compensation of the dominant motion: a bilinear interpolation based on the affine model, and a projections onto convex sets (POCS) based method that takes into account blurring in the image formation. It is important to note that the latter must be used if the blurring is severe enough to act as an anti-alias filter, which imposes an irreversible limit on the resolution improvement ability of any motion-compensated filter. Global motion-compensated images are then input to a motion-adaptive filter to detect and correct for those pixels where there exists local motion. A dynamic thresholding for motion detection is presented, with weighted directional-filtering for regions where motion is detected, to obtain the best results. Experimental results with application to obtaining high quality stills from video camcorders demonstrate the effectiveness of the proposed methods. Andrew J. Patti, M. Ibrahim Sezan, A. Murat Tekalp |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1997 | Two- versus three-dimensional object-based video compressionabstractThis paper compares two-dimensional (2-D) and three-dimensional (3-D) object modeling in terms of their capabilities and performance (peak signal-to-noise-ratio and visual image quality) for very low bitrate video coding. We show that 2-D object-based coding with affine/perspective transformations and triangular mesh models can simulate almost all capabilities of 3-D object-based approaches using wireframe models at a fraction of the computational cost. Furthermore, experiments indicate that a 2-D mesh-based coder-decoder performs favorably compared to the new H.263 standard in terms of visual quality. A. Murat Tekalp, Yücel Altunbasak, Gozde Bozdagi Akar |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1997 | Closed-form connectivity-preserving solutions for motion compensation using 2-D meshesabstractMotion compensation using two-dimensional (2-D) mesh models requires computation of the parameters of a spatial transformation for each mesh element (patch). It is well known that the parameters of an affine (bilinear or perspective) mapping can be uniquely estimated from three (four) point correspondences (at the vertices of a triangular or quadrilateral mesh element). On the other hand, overdetermined solutions using more than the required minimum number of point correspondences provide increased robustness against correspondence-estimation errors, however, this necessitates special consideration to preserve mesh-connectivity. This paper presents closed-form, overdetermined solutions for least squares estimation of affine motion parameters for a triangular mesh, which preserve mesh-connectivity using patch-based or node-based connectivity constraints. In particular, four new algorithms are presented: patch-constrained methods using point correspondences or spatio-temporal intensity gradients, and node-constrained methods using point correspondences or spatio-temporal intensity gradients. The methods using point correspondences can be viewed as postprocessing of a dense motion field for best representation in terms of a set of irregularly spaced samples. The methods that are based on spatio-temporal intensity gradients offer closed-form solutions for direct estimation of the best node-point motion vectors (equivalently the best transformation parameters). We show that the performance of the proposed closed-form solutions are comparable to those of the alternative search-based solutions at a fraction of the computational cost. Yücel Altunbasak, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 1997 | Occlusion-adaptive, content-based mesh design and forward tracking abstractTwo-dimensional (2-D) mesh-based motion compensation preserves neighboring relations (through connectivity of the mesh) as well as allowing warping transformations between pairs of frames; thus, it effectively eliminates blocking artifacts that are common in motion compensation by block matching. However, available 2-D mesh models, whether uniform or non-uniform, enforce connectivity everywhere within a frame, which is clearly not suitable across occlusion boundaries. To this effect, we hereby propose an occlusion-adaptive forward-tracking mesh model, where connectivity of the mesh elements (patches) across covered and uncovered region boundaries are broken. This is achieved by allowing no node points within the background to be covered (BTBC) and refining the mesh structure within the model failure (MF) region(s) at each frame. The proposed content-based mesh structure enables better rendition of the motion (compared to a uniform or a hierarchical mesh), while tracking is necessary to avoid transmission of all node locations at each frame. Experimental results show successful motion compensation and tracking. Yücel Altunbasak, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 1997 | Motion segmentation by multistage affine classificationabstractWe present a multistage affine motion segmentation method that combines the benefits of the dominant motion and block-based affine modeling approaches. In particular, we propose two key modifications to a recent motion segmentation algorithm developed by Wang and Adelson (1994). 1) The adaptive k-means clustering step is replaced by a merging step, whereby the affine parameters of a block which has the smallest representation error, rather than the respective cluster center, is used to represent each layer; and 2) we implement it in multiple stages, where pixels belonging to a single motion model are labeled at each stage. Performance improvement due to the proposed modifications is demonstrated on real video frames. George Borshukov, Gozde Bozdagi Akar, Yücel Altunbasak, A. Murat Tekalp |
IEEE Trans. Image Process. | 4 |
| 1997 | Simultaneous motion estimation and segmentationabstractWe present a Bayesian framework that combines motion (optical flow) estimation and segmentation based on a representation of the motion field as the sum of a parametric field and a residual field. The parameters describing the parametric component are found by a least squares procedure given the best estimates of the motion and segmentation fields. The motion field is updated by estimating the minimum-norm residual field given the best estimate of the parametric field, under the constraint that motion field be smooth within each segment. The segmentation field is updated to yield the minimum-norm residual field given the best estimate of the motion field, using Gibbsian priors. The solution to successive optimization problems are obtained using the highest confidence first (HCF) or iterated conditional mode, (ICM) optimization methods. Experimental results on real video are shown. Michael M. Chang, A. Murat Tekalp, M. Ibrahim Sezan |
IEEE Trans. Image Process. | 2 |
| 1997 | Robust, object-based high-resolution image reconstruction from low-resolution videoabstractWe propose a robust, object-based approach to high-resolution image reconstruction from video using the projections onto convex sets (POCS) framework. The proposed method employs a validity map and/or a segmentation map. The validity map disables projections based on observations with inaccurate motion information for robust reconstruction in the presence of motion estimation errors; while the segmentation map enables object-based processing where more accurate motion models can be utilized to improve the quality of the reconstructed image. Procedures for the computation of the validity map and segmentation map are presented. Experimental results demonstrate the improvement in image quality that can be achieved by the proposed methods. P. Erhan Eren, M. Ibrahim Sezan, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 1997 | Superresolution video reconstruction with arbitrary sampling lattices and nonzero aperture timeabstractPrinting from an NTSC source and conversion of NTSC source material to high-definition television (HDTV) format are some of the applications that motivate superresolution (SR) image and video reconstruction from low-resolution (LR) and possibly blurred sources. Existing methods for SR image reconstruction are limited by the assumptions that the input LR images are sampled progressively, and that the aperture time of the camera is zero, thus ignoring the motion blur occurring during the aperture time. Because of the observed adverse effects of these assumptions for many common video sources, this paper proposes (i) a complete model of video acquisition with an arbitrary input sampling lattice and a nonzero aperture time, and (ii) an algorithm based on this model using the theory of projections onto convex sets to reconstruct SR still images or video from an LR time sequence of images. Experimental results with real video are provided, which clearly demonstrate that a significant increase in the image resolution can be achieved by taking the motion blurring into account especially when there exists large interframe motion. Andrew J. Patti, M. Ibrahim Sezan, A. Murat Tekalp |
IEEE Trans. Image Process. | 3 |
| 1996 | Occlusion-adaptive 2-D mesh trackingabstractAvailable 2-D mesh models, whether uniform or adaptive, preserve neighboring relationships between the patches (connectivity) everywhere within the frame, which is clearly not suitable across motion and occlusion boundaries. To this effect, this paper proposes an occlusion-adaptive, content-based mesh design and forward tracking algorithm, where no node points are placed within the background to be covered (BTBC), and node points along the boundary of the uncovered background (UB) are allowed to split. Furthermore, the mesh is refined within the UB region at every frame. Experimental results show successful motion compensation and tracking. Yücel Altunbasak, A. Murat Tekalp |
ICASSP | 2 |
| 1996 | Fusion of color and edge information for improved segmentation and edge linkingabstractWe propose a new method for combined color image segmentation and edge linking. The image is first segmented based on color information only. The segmentation map is modeled by a Gibbs random field, to ensure formation of spatially contiguous regions. Next, spatial edge locations are determined using the magnitude of the gradient of the 3-channel image vector field. Finally, regions in the segmentation map are split and merged by a region-labeling procedure to enforce their consistency with the edge map. The boundaries of the final segmentation map constitute a linked edge map. Experimental results are reported. Eli Saber, A. Murat Tekalp, Gozde Bozdagi Akar |
ICASSP | 2 |
| 1996 | Robust region-based high-resolution image reconstruction from low-resolution videoabstractWe propose a region-based approach to high-resolution image reconstruction from video using the projections onto convex sets framework. The region-based formalism allows the use of a validity map and/or a segmentation map. The validity map disables projections based on observations with inaccurate motion information, for robust reconstruction in the presence of motion estimation errors; while the segmentation map enables object-based processing where more accurate motion models can be utilized to improve the reconstruction quality of image object. A procedure for the computation of the validity map is presented. Experimental results demonstrate the quality improvement that can be achieved by the proposed region-based approach. P. Erhan Eren, M. Ibrahim Sezan, A. Murat Tekalp |
ICIP (1) | 3 |
| 1996 | Integration of color, shape, and texture for image annotation and retrievalabstractWe present algorithms for automatic image annotation and retrieval based on pixel- or region-based color; region-based shape; and block- or region-based texture features and several schemes for integrating them. Automatic region selection is accomplished by integrating color and spatial edge features. Color, shape, and texture indexing may be knowledge-based (using appropriate training sets) or by example. The multi-feature integration algorithms are designed to: (1) offer the user a wide range of options and flexibilities to enhance the outcome of the search and retrieval operations, and (2) provide a compromise between accuracy and computational complexity. One of the novel features of this system is its ability to automatically segment images into meaningful regions, and provide the option of region (object)-based image indexing without user interaction. Eli Saber, A. Murat Tekalp |
ICIP (3) | 2 |
| 1996 | 2-D mesh-based tracking of deformable objects with occlusionabstractMany interactive video and multimedia applications, such as interactive TV, augmented reality and bit stream editing, demand object-based video modeling. These applications require accurate tracking of the boundary, local motion and intensity variations of each object. This paper proposes a new approach to track deformable objects in video sequences, which combines deformable contours and 2-D mesh models while addressing partial occlusion, including self-occlusion. Experimental results demonstrate successful tracking of objects with deformable boundaries in the presence of self-occlusion. Candemir Toklu, A. Murat Tekalp, A. Tanju Erdem, M. Ibrahim Sezan |
ICIP (1) | 2 |
| 1996 | Similarity analysis for shape retrieval by exampleabstractThis work proposes a query-by-example algorithm for shape retrieval from a color image database. The main contribution of this work is an unified approach for shape matching and similarity ranking using a modal representation. Prior work mostly performed point correspondence determination and similarity ranking of shapes in two distinct steps or performed one of them. We define a new shape-similarity metric and attempt to address the question "how similar two shapes in an image database are", which currently is an important problem in shape-based retrieval systems. Unlike most reported methods, the presented approach does not require extraction of connected boundaries or silhouettes. It is rotation-, and scale-invariant, and can handle mild deformations of objects. The results are promising for using the algorithm in the context of database query by image content. Bilge Günsel, A. Murat Tekalp |
ICPR | 2 |
| 1996 | Face detection and facial feature extraction using color, shape and symmetry-based cost functionsabstractThis paper describes an algorithm for detecting human faces and subsequently localizing the eyes, nose, and mouth. First, we locate the face based on color and shape information. To this effect, a supervised pixel-based color classifier is used to mark all pixels which are within a prespecified distance of "skin color". This color-classification map is then subject to smoothing employing either morphological operations or filtering using a Gibbs random field model. The eigenvalues and eigenvectors computed from the spatial covariance matrix are utilized to fit an ellipse to the skin region under analysis. The Hausdorff distance is employed as a means for comparison, yielding a measure of proximity between the shape of the region and the ellipse model. Then, we introduce symmetry-based cost functions to locate the center of the eyes, tip of nose, and center of mouth within the facial segmentation mask. The cost functions are designed to take advantage of the inherent symmetries associated with facial patterns. We demonstrate the performance of our algorithm on a variety of images. Eli Saber, A. Murat Tekalp |
ICPR | 2 |
| 1996 | Video indexing through integration of syntactic and semantic featuresabstractThis paper proposes a content-based video indexing system which provides the functionalities necessary for automatic management of video data through integration of syntactic and semantic features. The proposed system has been applied to detection, classification and then indexing of news programs collected from different TV channels. Although the paper focuses on news programs, the same methods can be used to extent-based index and search other TV programs with distinct semantic structure. Bilge Günsel, A. Müfit Ferman, A. Murat Tekalp |
WACV | 3 |
| 1996 | Automatic Image Annotation Using Adaptive Color ClassificationabstractWe describe a system which automatically annotates images with a set of prespecified keywords, based on supervised color classification of pixels intoNprespecified classes using simple pixelwise operations. The conditional distribution of the chrominance components of pixels belonging to each class is modeled by a two-dimensional Gaussian function, where the mean vector and the covariance matrix for each class are estimated from appropriate training sets. Then, a succession of binary hypothesis tests with image-adaptive thresholds has been employed to decide whether each pixel in a given image belongs to one of the predetermined classes. To this effect, a universal decision threshold is first selected for each class based on receiver operating characteristics (ROC) curves quantifying the optimum “true positive” vs “false positive” performance on the training set. Then, a new method is introduced for adapting these thresholds to the characteristics of individual input images based on histogram cluster analysis. If a particular pixel is found to belong to more than one class, a maximuma posterioriprobability (MAP) rule is employed to resolve the ambiguity. The performance improvement obtained by the proposed adaptive hypothesis testing approach over using universal decision thresholds is demonstrated by annotating a database of 31 images. Eli Saber, A. Murat Tekalp, Reiner Eschbach, Keith T. Knox |
CVGIP Graph. Model. Image Process. | 2 |
| 1996 | Tracking Motion and Intensity Variations Using Hierarchical 2-D Mesh Modeling for Synthetic Object TransfigurationabstractWe propose a method for tracking the motion and intensity variations of a 2-D mildly deformable image object using a hierarchical 2-D mesh model. The proposed method is applied to synthetic object transfiguration, namely, replacing an object in a real video clip with another synthetic or natural object via digital postprocessing. Successful transfiguration requires accurate tracking of both motion and intensity (contrast and brightness) variations of the object-to-be-replaced so that the replacement object can be rendered in exactly the same way from a single still picture. The proposed method is capable of tracking image regions corresponding to scene objects with nonplanar and/or mildly deforming surfaces, accounting for intensity variations, and is shown to be effective with real image sequences. Candemir Toklu, A. Tanju Erdem, M. Ibrahim Sezan, A. Murat Tekalp |
CVGIP Graph. Model. Image Process. | 4 |
| 1995 | Simultaneous stereo-motion fusion and 3-D motion trackingabstractPresents a new framework for combining maximum likelihood (ML) stereo-motion fusion with adaptive iterated extended Kalman filtering (IEKF) for 3-D motion tracking. The ML stereo-fusion step, with two stereo-pairs, generates observations of 3-D feature matches to be used by the IEKF step. The IEKF step, in turn, computes updated 3-D motion parameter estimates to be used by the ML stereo-motion fusion step. The covariance of the observation noise process is regulated by the value of the ML cost function to address occlusion related problems. The proposed simultaneous approach is compared with performing the 3-D feature correspondence estimation and the Kalman filtering separately using simulated stereo imagery. Yücel Altunbasak, A. Murat Tekalp, Gozde Bozdagi Akar |
ICASSP | 2 |
| 1995 | High resolution standards conversion of low resolution videoabstractWith the advent of frame grabbers capable of acquiring multiple video frames, a great deal of attention is being directed at creating high-resolution (hi-res) imagery from interlaced or low-resolution (low-res) video. This is a multi-faceted problem, which generally necessitates standards conversion and hi-res reconstruction. Standards conversion is the problem of converting from one spatio-temporal sampling lattice to another, while hi-res image reconstruction involves increasing the spatial sampling density. Also of interest is removing degradations that occur during the image acquisition process. These tasks have all received considerable, yet separate, treatment in the literature. A unifying video formation model is presented which addresses these problems simultaneously. Then, a POCS-based algorithm for generating high-resolution imagery from video is delineated. Results with real imagery are included. Andrew J. Patti, M. Ibrahim Sezan, A. Murat Tekalp |
ICASSP | 3 |
| 1995 | Two-dimensional object-based coding using a content-based mesh and affine motion parameterizationabstractWe present a complete system for 2-D object-based video compression with a method for 2-D content-based triangular mesh design, two connectivity preserving affine motion parameterization schemes, two methods for temporal mesh propagation, a polygon-based adaptive model failure detection/coding scheme, and bit rate control strategies. The feasibility of the proposed methods has been demonstrated by experimental results. Yücel Altunbasak, A. Murat Tekalp, Gozde Bozdagi Akar |
ICIP | 2 |
| 1995 | 2-D mesh tracking for synthetic transfigurationabstractSynthetic object transfiguration is the problem of replacing an existing object in an image sequence with a new one via digital post-processing. Successful transfiguration requires accurate tracking of both motion and intensity variations of the existing object so that the new object can be rendered in exactly the same way. We propose a 2-D mesh tracking method that estimates, on a local basis, the motion and intensity variations of a region of interest. The proposed method is capable of accounting for illumination changes, tracking image regions corresponding to scene objects with non-planar and/or mildly deforming surfaces, and is shown to be effective in real-life situations. Candemir Toklu, A. Tanju Erdem, M. Ibrahim Sezan, A. Murat Tekalp |
ICIP (3) | 4 |
| 1995 | Out-of-plane motion compensation in multislice spin-echo MRIabstractIn magnetic resonance imaging (MRI), it is well-known that patient motion plays a significant role in the degradation of image quality. Although the case of translational in-plane motion (x-y-motion) has been studied by several researchers, the effect of rigid, translational out-of-plane motion (z-motion) has not yet been completely analyzed due to its more complex nature. Out-of-plane motion introduces blurring along the slice-selection direction in addition to motion artifacts. Here, the authors present a model to represent the effect of out-of-plane motion on multislice MR data. The inversion of this model not only results in the correction of the artifacts due to out-of-plane motion, but also reduces blurring in the slice-selection direction, yielding higher resolution images. Because of the shift-varying nature of the authors' model, they propose to use a nonlinear postprocessing method, projection onto convex sets (POCS), for its inversion, provided that the motion kernel and the slice-selection profile are known. The proposed method has been tested on simulated data and then applied to actual MR data to demonstrate the feasibility of the technique in real imaging situations. Jonathan K. Riek, A. Murat Tekalp, Warren E. Smith, Edmund Kwok |
IEEE Trans. Medical Imaging | 2 |
| 1994 | Simultaneous 3-D motion estimation and wire-frame model adaptation including photometric effects for knowledge-based video codingabstractWe address the problem of 3-D motion estimation in the context of knowledge-based coding of facial image sequences. The proposed method handles the global and local motion estimation and the adaptation of a generic wire-frame to a particular speaker simultaneously within an optical flow based framework including the photometric effects of motion. We use a flexible wire-frame model whose local structure is characterized by the normal vectors of the patches which are related to the coordinates of the nodes. Geometrical constraints that describe the propagation of the movement of the nodes are introduced, which are then efficiently utilized to reduce the number of independent structure parameters. A stochastic relaxation algorithm has been used to determine optimum global motion estimates and the parameters describing the structure of the wire-frame model. For the initialization of the motion and structure parameters, a modified feature based algorithm is used. Experimental results with simulated facial image sequences are given.> Gozde Bozdagi Akar, A. Murat Tekalp, Levent Onural |
ICASSP (5) | 2 |
| 1994 | An algorithm for simultaneous motion estimation and scene segmentationabstractWe present an algorithm that simultaneously performs motion estimation and scene segmentation in image sequences. This algorithm improves upon the traditional approach of using the motion estimates that are obtained separately as input for scene segmentation. Based on the maximum a posteriori probability (MAP) criterion, the interdependence of constraints for motion estimation and scene segmentation are expressed in a Gibbs distribution. The solution is obtained based on the highest confidence first (HCF) and iterated conditional mode (ICM) optimization methods. Experimental results suggest that this algorithm is of potential use for low-bit rate video coding.> Michael M. Chang, M. Ibrahim Sezan, A. Murat Tekalp |
ICASSP (5) | 3 |
| 1994 | Digital video standards conversion in the presence of accelerated motionabstractWe address the standards conversion of digital video containing accelerated motion. The spectral characterization of such video is performed by means of short time spectral analysis. The support of the Short Time Spectrum (STS) of video with accelerated motion is derived, and is compared to the STS of linear time-varying (LTV) filters that operate along an accelerated motion trajectory. By utilizing the multiplication property of the STS, it is shown that filtering along the accelerated motion trajectory potentially yields the smallest amount of aliasing in standards conversion applications.> Andrew J. Patti, M. Ibrahim Sezan, A. Murat Tekalp |
ICASSP (5) | 3 |
| 1994 | Simultaneous Motion-Disparity Estimation and Segmentation from StereoabstractBoth motion/structure estimation from monocular video and disparity estimation from still-frame stereo are known to be ill-posed problems. Further, because disparity varies by depth, and motion parameters are different for independently moving objects, they can both benefit from scene segmentation. To this effect, we present a framework for simultaneous motion and disparity estimation including scene segmentation in stereo video. In this formulation, pairs of disparity and motion parameter vector values are segmented into K regions, where within each region a single set of motion parameters is defined, and the disparity field is allowed to vary smoothly. The algorithm iterates between computing the maximum a posteriori probability (MAP) estimates of the disparity and segmentation fields conditioned on the present motion parameter estimates, and the maximum likelihood (hit) estimates of the motion parameters via simulated annealing (SA). Simulation results are provided.> Yücel Altunbasak, A. Murat Tekalp, Gozde Bozdagi Akar |
ICIP (3) | 2 |
| 1994 | High-Resolution Image Reconstructionfrom a Low-Resolution Image Sequence in the Presence of Time-Varying Motion BlurabstractWe address the problem of reconstruction of a high-resolution image from a sequence of low-resolution images containing arbitrary relative motion, excluding occlusion effects. We develop a formulation that simultaneously takes into account blurring due to relative sensor-object motion, sensor integration, and additive noise. We propose a POCS-based algorithm for performing the high-resolution reconstruction, and provide experimental results.> Andrew J. Patti, M. Ibrahim Sezan, A. Murat Tekalp |
ICIP (1) | 3 |
| 1994 | Digital video standards conversion in the presence of accelerated motion
Andrew J. Patti, M. Ibrahim Sezan, A. Murat Tekalp |
Signal Process. Image Commun. | 3 |
| 1994 | DCT coding of nonrectangularly sampled imagesabstractDiscrete cosine transform (DCT) coding is widely used for compression of rectangularly sampled images. We address efficient DCT coding of nonrectangularly sampled images. To this effect, we discuss an efficient method for the computation of the DCT on nonrectangular sampling grids using the Smith-normal decomposition. Simulation results are provided.> Emre Gündüzhan, A. Enis Çetin, A. Murat Tekalp |
IEEE Signal Process. Lett. | 3 |
| 1994 | 3-D motion estimation and wireframe adaptation including photometric effects for model-based coding of facial image sequencesabstractProposes a novel formulation where 3D global and local motion estimation and the adaptation of a generic wireframe model to a particular speaker are considered simultaneously within an optical flow based framework including the photometric effects of the motion. We use a flexible wireframe model whose local structure is characterized by the normal vectors of the patches which are related to the coordinates of the nodes. Geometrical constraints that describe the propagation of the movement of the nodes are introduced, which are then efficiently utilized to reduce the number of independent structure parameters. A stochastic relaxation algorithm has been used to determine optimum global motion estimates and the parameters describing the structure of the wireframe model. Results with both simulated and real facial image sequences are provided.> Gozde Bozdagi Akar, A. Murat Tekalp, Levent Onural |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1994 | An improvement to MBASIC algorithm for 3-D motion and depth estimationabstractIn model-based coding of facial images, the accuracy of motion and depth parameter estimates strongly affects the coding efficiency. MBASIC (model-based analysis-synthesis image coding) is a simple and effective iterative algorithm recently proposed by Aizawa et el. (see Signal Processing: Image Communication, no.1, p.139-52, 1989) for 3-D motion and depth estimation when the initial depth estimates are relatively accurate. In this correspondence, we analyze its performance in the presence of errors in the initial depth estimates and propose a modification to MBASIC algorithm that significantly improves its robustness to random errors with only a small increase in the computational load. Gozde Bozdagi Akar, A. Murat Tekalp, Levent Onural |
IEEE Trans. Image Process. | 2 |
| 1994 | POCS-based restoration of space-varying blurred imagesabstractWe propose a new method for space-varying image restoration using the method of projection onto convex sets (POCS). The formulation allows the use of a different blurring function at each pixel of the image in a computationally efficient manner. We illustrate the performance of the proposed approach by comparing the new results with those of the ROMKF method on simulated images. We also present results on a real-life image with unknown space-varying out-of-focus blur. Mehmet K. Özkan, A. Murat Tekalp, M. Ibrahim Sezan |
IEEE Trans. Image Process. | 2 |
| 1993 | Motion-field segmentation using an adaptive MAP criterion
Michael M. Chang, A. Murat Tekalp, M. Ibrahim Sezan |
ICASSP (5) | 2 |
| 1993 | New approaches for space-invariant image restoration
Andrew J. Patti, Mehmet K. Özkan, A. Murat Tekalp, M. Ibrahim Sezan |
ICASSP (5) | 3 |
| 1993 | Effect of Z-motion in the phase of the K-space MRI data and identification of periodic Z-motion kernels
Jonathan K. Riek, A. Murat Tekalp, Warren E. Smith |
ICASSP (5) | 2 |
| 1993 | Adaptive Bayesian approach for color image segmentationabstractA Bayesian segmentation algorithm to separate color images into regions of distinct colors is presented. The algorithm takes into account the local color variations in the image in an adaptive manner. A Gibbs random field (GRF) is used as the a priori probability model for the segmentation process to impose a spatial connectivity constraint. We study the performance of the proposed algorithm in different color spaces and its application in reduced data rendering of color images. Experimental results and discussion are included. Michael M. Chang, Andrew J. Patti, M. Ibrahim Sezan, A. Murat Tekalp |
VCIP | 4 |
| 1993 | Adaptive motion-compensated filtering of noisy image sequencesabstractThe authors propose a novel adaptive spatiotemporal filter, called the adaptive weighted averaging (AWA) filter, for effective noise suppression in image sequences without introducing visually disturbing blurring artifacts. Filtering is performed by computing the weighted average of image values within a spatiotemporal support along the estimated motion trajectory at each pixel. The weights are determined by optimizing a well defined mathematical criterion, which provides an implicit mechanism for deemphasizing the contribution of the outlier pixels within the spatiotemporal filter support to avoid blurring. The AWA filter is therefore particularly well suited for filtering sequences that contain segments with abruptly changing scene content due to, for example, rapid zooming and changes in the view of the camera. The performance of the proposed AWA filter is compared with that of the spatiotemporal, local linear minimum mean square error (LMMSE) filtering. The results demonstrate that the proposed AWA filter-outperforms the LMMSE filter, especially in the cases of low signal-to-noise ratios and abruptly varying scene content.> Mehmet K. Özkan, M. Ibrahim Sezan, A. Murat Tekalp |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1992 | Modeling and correction of artifacts due to z-motion in 2-D MRIabstractIn 2D Fourier magnetic resonance imaging (MRI), patient motion, including respiratory, cardiac, and involuntary motion, plays a significant role in degradation of the image quality. The in-plane motion (x-y motion) results only in phase distortions in the raw (Fourier domain) MRI data, which can be easily corrected once the motion kernel is known. However, the effect of motion in the slice-selection direction (z-motion) has not yet been completely analyzed due to its significantly more complex nature. The authors propose a novel model to represent the effect of z-motion in the data acquisition process, and then develop two postprocessing methods for the correction of z-motion artifacts provided that the z-motion kernel is known. The model takes the form of a multi-input single-output, shift-variant system. Thus, the inversion of this system is treated by using some nonlinear methods such as the projection onto convex sets (POCS) technique and a Monte Carlo method.> Jonathan K. Riek, A. Murat Tekalp, Warren E. Smith |
ICASSP | 2 |
| 1992 | High-resolution image reconstruction from lower-resolution image sequences and space-varying image restorationabstractThe authors address the problem of reconstruction of a high-resolution image from a number of lower-resolution (possibly noisy) frames of the same scene where the successive frames are uniformly based versions of each other at subpixel displacements. In particular, two previously proposed methods, a frequency-domain method and a method based on projections onto convex sets (POCSs), are extended to take into account the presence of both sensor blurring and observation noise. A new two-step procedure is proposed, and it is shown that the POCS formulation presented for the high-resolution image reconstruction problem can also be used as a new method for the restoration of spatially invariant blurred images. Some simulation results are provided.> A. Murat Tekalp, Mehmet K. Özkan, M. Ibrahim Sezan |
ICASSP | 1 |
| 1992 | Editorial
A. Murat Tekalp |
Signal Process. | 1 |
| 1992 | The seventh MDSP workshop
A. Murat Tekalp |
Signal Process. | 1 |
| 1992 | Efficient multiframe Wiener restoration of blurred and noisy image sequencesabstractComputationally efficient multiframe Wiener filtering algorithms that account for both intraframe (spatial) and interframe (temporal) correlations are proposed for restoring image sequences that are degraded by both blur and noise. One is a general computationally efficient multiframe filter, the cross-correlated multiframe (CCMF) Wiener filter, which directly utilizes the power and cross power spectra of only NxN matrices, where N is the number of frames used in the restoration. In certain special cases the CCMF lends itself to a closed-form solution that does not involve any matrix inversion. A special case is the motion-compensated multiframe (MCMF) filter, where each frame is assumed to be a globally shifted version of the previous frame. In this case, the interframe correlations can be implicitly accounted for using the estimated motion information. Thus the MCMF filter requires neither explicit estimation of cross correlations among the frames nor matrix inversion. Performance and robustness results are given. Mehmet K. Özkan, A. Tanju Erdem, M. Ibrahim Sezan, A. Murat Tekalp |
IEEE Trans. Image Process. | 4 |
| 1992 | Maximum likelihood parametric blur identification based on a continuous spatial domain modelabstractA formulation for maximum-likelihood (ML) blur identification based on parametric modeling of the blur in the continuous spatial coordinates is proposed. Unlike previous ML blur identification methods based on discrete spatial domain blur models, this formulation makes it possible to find the ML estimate of the extent, as well as other parameters, of arbitrary point spread functions that admit a closed-form parametric description in the continuous coordinates. Experimental results are presented for the cases of 1-D uniform motion blur, 2-D out-of-focus blur, and 2-D truncated Gaussian blur at different signal-to-noise ratios. Gordana Pavlovic, A. Murat Tekalp |
IEEE Trans. Image Process. | 2 |
| 1991 | Matching-extrapolation of bicumulants of one-D signals using two-D AR modelingabstractA 2-D autoregressive (AR) model is proposed which is not based on an underlying linear process representation, for matching-extrapolation of arbitrary bicumulant sequences of one-dimensional signals. This model employs the same number of coefficients as the number of given bicumulant samples, thus, makes it possible to extrapolate a given finite set of bicumulant samples to infinity while being consistent with, i.e., matching, all of the given samples. The computation of the model coefficients requires solving a system of linear equations. An algorithm is developed for the extrapolation of the given bicumulant samples based on the proposed 2-D AR model. An example is presented to compare the performance of the proposed method with other approaches for the extrapolation of a given set of bicumulant samples.> A. Tanju Erdem, A. Murat Tekalp |
ICASSP | 2 |
| 1991 | Maximum likelihood parametric blur identification based on a continuous spatial domain modelabstractA new formulation is proposed for maximum likelihood (ML) blur identification that is based on a parametric description of the blur in the continuous spatial coordinates. The aim of this formulation is to find the ML estimate of the extent of certain point spread functions (PSF). It is shown that this can be achieved by formulating the problem in the continuous spatial coordinates, as opposed to using the conventional discrete spatial domain model. Experimental results are presented for the cases of uniform motion blur, out of focus blur and truncated Gaussian blur at various signal-to-noise ratios.> Gordana Pavlovic, A. Murat Tekalp |
ICASSP | 2 |
| 1991 | On modeling the focus blur in image restorationabstractThe point spread function (PSF) of a defocused lens, in the continuous spatial coordinates, can be modeled using the principles of either geometrical optics or physical optics. The authors compare the PSFs and the corresponding optical transfer functions obtained on the basis of these two models under different amounts of blurs and different imaging sensor resolutions, in the discrete spatial domain. Different approximations for the discretization of the PSF are considered. The effect of using geometrical versus physical optics-based models on the quality of restored images is investigated experimentally in the case of images recorded by a charge-coupled device (CCD) camera.> M. Ibrahim Sezan, Gordana Pavlovic, A. Murat Tekalp, A. Tanju Erdem |
ICASSP | 3 |
| 1991 | A composite signal model that simultaneously realizes arbitrary polynomial bispectra and rational power spectraabstractA novel signal representation, called a system with multiplicity K (SWM-K) is reviewed for exact modeling of a given polynomial bispectrum. An extension of this model is presented such that the extended model can simultaneously realize an arbitrary polynomial bispectrum and a rational power spectrum. In order to simultaneously realize arbitrary polynomial bispectra and rational power spectra, another linear shift-invariant system which is driven by a Gaussian input independent of the other K inputs is added to the SWM-K. Conditions for the existence of this composite model are stated. Simulation results are provided for an image processing example.> A. Murat Tekalp, A. Tanju Erdem, Michael M. Chang |
ICASSP | 1 |
| 1991 | Comparative study of some statistical and set-theoretic methods for image restorationabstractIn the last decade, many researchers have devoted considerable effort to the problem of image restoration. However, no recent study has been undertaken for a comparative evaluation of these techniques under conditions where a user may have different kinds of a priori information about the ideal image. To this effect, we briefly survey some recent techniques and compare the performance of a linear space-invariant (LSI) maximum a posteriori (MAP) filter, an LSI reduced update Kalman filter (RUKF), an edge-adaptive RUKF, and an adaptive convex-type constraint-based restoration implemented via the method of projection onto convex sets (POCS). The mean square errors resulting from the LSI algorithms are compared with that of the finite impulse response Wiener filter, which is the theoretical limit in this case. We also compare the results visually in terms of their sharpness and the appearance of artifacts. As expected, the space-variant restoration methods which are adaptive to local image properties obtain the best results. A. Murat Tekalp, H. Joel Trussell |
CVGIP Graph. Model. Image Process. | 1 |
| 1990 | Blur identification using bispectrumabstractBlur identification methods based on the bispectrum of an observed noisy and blurred image are proposed. Blur identification is addressed, in the case of uniform blurs, by zero-crossing detection in a particular cross section of the bispectrum of the observed image. The identification of general finite-impulse-response blurs is considered. The blurred image is represented by a nonminimum phase ARMA (autoregressive moving average) model whose parameters are identified using the complex bispectrum of the observed image. These methods are superior to previous methods which are based on the second-order statistics of the image, since the bispectrum is insensitive to additive, Gaussian noise in theory, and it contains the phase of the blur transfer function. Simulation results are provided.> A. Tanju Erdem, A. Murat Tekalp |
ICASSP | 2 |
| 1990 | Restoration in the presence of multiplicative noise with application to scanned photographic imagesabstractA linear minimum mean square error (LMMSE) deconvolution filter is derived in the presence of a multiplicative noise in order to incorporate the nonlinear sensor characteristics into the restoration of noisy and blurred scanned photographic images. Image restoration is proposed in the exposure domain where a linear convolutional relationship between the original and the observed images can be established at the expense of having a multiplicative observation noise. Examples of photographically blurred images are provided to demonstrate that images can be identified and restored in the exposure domain which cannot be identified or restored in the optical density domain.> Gordana Pavlovic, A. Murat Tekalp |
ICASSP | 2 |
| 1990 | Image modeling using higher-order statistics with application to predictive image codingabstractThe mathematical framework to develop parametric image models based on higher-order statistics (HOS) is discussed. For non-Gaussian images, the model parameters will depend on the HOS if the model has a nonlinear form or if a fidelity criterion other than the mean square error (MSE) is utilized. Three new image models are proposed: (i) a nonlinear model based on the MSE criterion, (ii) a linear model under a criterion other than the MSE, and (iii) a nonlinear model using a criterion other than the MSE. These models are applied to predictive image coding, and the results are compared with those obtained by a linear model based on the MSE criterion.> A. Murat Tekalp, Mehmet K. Özkan, A. Tanju Erdem |
ICASSP | 1 |
| 1990 | Tutorial review of recent developments in digital image restorationabstractIn this review, we consider the three fundamental aspects of image restoration: (i) modeling, (ii) model identification methods, and (iii) restoration methods. Modeling refers to determining a model of the relationship between the ideal image and the observed degraded image, as well as modeling the ideal image itself, on the basis of a priori information. Model parameters are determined by various identification methods. Restoration algorithms are discussed in two categories: general algorithms and specialized algorithms. We also furnish a brief discussion of present and future research directions. M. Ibrahim Sezan, A. Murat Tekalp |
VCIP | 2 |
| 1989 | Higher-order spectrum factorization with applicationsabstractThe authors address the problem of modeling a given higher order spectrum as that of the output of a linear time-invariant system driven by a higher order white random signal. This can be posed as a higher order spectrum factorization problem. The authors provide a theorem concerning the existence of such a factorization. A fast algorithm for efficient implementation of the factorization, if it exists, is then proposed. As applications, the authors present the problems of identification of non-minimum-phase linear time-invariant systems and phase reconstruction.> A. Murat Tekalp, A. Tanju Erdem |
ICASSP | 1 |
| 1988 | Iterative image restoration with ringing suppression using the method of POCSabstractDevelops an iterative, space-variant, adaptive restoration algorithm incorporating both regularization and ringing suppression constraints. The algorithm is formulated using the method of projections onto convex sets (POCS). In the proposed method, regularized restoration is implemented by projecting at each iteration onto the extended Wiener solution set. Projection onto this set forces the solution to be equal to the Wiener solution at a set of frequencies over which the magnitude of the degradation transfer function is above a predetermined threshold. Ringing suppression is achieved by adaptively bounding the regional image variance in an attempt to reduce the energy of the ringing artifacts. Three corresponding convex constraint sets are defined to impose tight, moderate, and loose bounds on image variances over the uniform, texture and edge regions, respectively.> M. Ibrahim Sezan, A. Murat Tekalp |
ICASSP | 2 |
| 1988 | Decision-directed segmentation for the restoration of images degraded by a class of space-variant blursabstractA decision-directed filtering algorithm has been developed for model-based segmentation and restoration of images degraded by a class of space-variant blurs. It is assumed that the space-variant blur can be represented by a collection of L distinct point-spread functions, where L is a predetermined integer, so that at each pixel one of the functions will be more or less matched to the observed data. A multiple-model Kalman filtering procedure with online model detection based on maximum a posteriori probability decision was used to restore the image. The result of the decision process constitutes a model-based segmentation of the degraded image into regions of spatially invariant blurs. There are several applications of the proposed algorithm. Among these, automatic blur detection, i.e. segmentation of a partially blurred image into regions of blur and no-blur, is demonstrated as an example.> A. Murat Tekalp, Howard Kaufman, John W. Woods |
ICASSP | 1 |
| 1988 | Comparative study of some recent statistical and set-theoretic methods for image restorationabstractThe performance is compared of a linear space-invariant (LSI) maximum a posteriori filter, an LSI reduced update Kalman filter (RUKF), an edge-adaptive RUKF, and an adaptive convex-type constraint-based restoration implemented via the method of projection onto convex sets. The finite impulse response Wiener filter is taken as a benchmark in this comparison. In image restoration, the LSI techniques are found to have some important drawbacks, such as producing ringing artifacts. As expected, the space-variant restoration methods which are adaptive to local image properties provide the best results.> A. Murat Tekalp, H. Joel Trussell |
ICASSP | 1 |
| 1987 | Effects of constraints, initialization, and finite-word length in blind deblurring of images by convex projectionsabstractReconstruction from phase information via convex projections is considered for deblurring smeared images when the blurring function does not introduce any phase. The effects of initialization and constraints in the algorithm are investigated. It is shown by examples that the number of quantization levels when the smeared image is digitized determines the quality of the restored image. Cheng-Tie Chen, M. Ibrahim Sezan, A. Murat Tekalp |
ICASSP | 3 |
| 1987 | Regularized signal restoration using the theory of convex projectionsabstractA solution to the regularized signal restoration problem is formulated in the context of projections onto convex sets. An extended Wiener set is defined based on the Wiener solution in the frequency domain. A constrained Wiener solution is determined as a point in the intersection of the extended Wiener set and other deterministic constraint sets. The relationship between the proposed technique and the constrained Miller-Tikhonov regularization developed earlier is established. M. Ibrahim Sezan, A. Murat Tekalp, Cheng-Tie Chen |
ICASSP | 2 |
| 1985 | Identification of image and blur parameters for the restoration of noncausal blursabstractAn optimal statistical parameter estimation technique is presented for the identification of unknown image and blur-model parameters. Maximum likelihood estimates of the unknown parameters are derived both in the absence and in the presence of observation noise. The proposed algorithms are able to locate the zeros of the observed image spectrum on the entire Z1-Z2plane and therefore, unlike previous algorithms, are not restricted to the Fourier domain. Images restored with the proposed algorithms are shown as examples. A. Murat Tekalp, Howard Kaufman, John W. Woods |
ICASSP | 1 |
| 1985 | Boundary value problem in image restorationabstractRecursive spatial domain filtering is often used in image restoration. The techniques range from deterministic inverse filter algorithms [1] to stochastic Kalman filters [2,3]. However, all the techniques have in common the problem of choosing the appropriate boundary values. It is the purpose of this paper to demonstrate the importance of the boundary values in image restoration and to show how we can improve the transient response of the steady-state Kalman filters in [2] and [3] with better choices for the boundary values. John W. Woods, Jan Biemond, A. Murat Tekalp |
ICASSP | 3 |
| 1983 | A multiple model algorithm for the adaptive restoration of imagesabstractWe present a multiple model image restoration technique with on-line edge detection over noisy and blurred images. Four edge models corresponding to major correlation directions and an isotropic model are used to represent the image. Unlike previous results, a decision-directed approach is used to adaptively estimate the edge orientation which then defines the appropriate deconvolution model. The overall effect is space-variant deconvolution implemented by efficient reduced update filters. Processed images are shown as examples. A. Murat Tekalp, John W. Woods, Howard Kaufman |
ICASSP | 1 |