Ivan V. Bajic

dblp:37/3765 · DBLP profile ↗
← Back
126ranked-venue papers
16as first author
34since 2021 · last 2025
0000-0003-3154-5743ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 107 · 13 first-author · 28 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Computer networks · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Semantics-Guided Generative Image Compression
abstract
Advancements in text-to-image generative AI with large multimodal models are spreading into the field of image compression, creating high-quality representation of images at extremely low bit rates. This work introduces novel components to the existing multimodal image semantic compression (MISC) approach, enhancing the quality of the generated images in terms of PSNR and perceptual metrics. The new components include semantic segmentation guidance for the generative decoder, as well as content-adaptive diffusion, which controls the number of diffusion steps based on image characteristics. The results show that our newly introduced methods significantly improve the baseline MISC model while also decreasing the complexity. As a result, both the encoding and decoding time are reduced by more than 36%. Moreover, the proposed compression framework outperforms mainstream codecs in terms of perceptual similarity and quality. The code and visual examples are available.1
Hyomin Choi, Ivan V. Bajic
ICIP3
2025 How Universal Are SAM2 Features?
Masoud Khairi Atani, Alon Harell, Hyomin Choi, Runyu Yang, Fabien Racapé, Ivan V. Bajic
PCS6
2025 Bit Allocation Transfer for Perceptual Quality Enhancement of VVC Intra Coding
Runyu Yang, Ivan V. Bajic
PCS2
2025 Efficient Signed Graph Sampling via Balancing & Gershgorin Disc Perfect Alignment
abstract
A basic premise in graph signal processing (GSP) is that a graph encoding pairwise (anti-)correlations of the targeted signal as edge weights is leveraged for graph filtering. Existing fast graph sampling schemes are designed and tested only for positive graphs describing positive correlations. However, there are many real-world datasets exhibiting strong anti-correlations, and thus a suitable model is a signed graph, containing both positive and negative edge weights. In this paper, we propose the first linear-time method for sampling signed graphs, centered on the concept of balanced signed graphs. Specifically, given an empirical covariance data matrix , we first learn a sparse inverse matrix , interpreted as a graph Laplacian corresponding to a signed graph . We approximate with a balanced signed graph via fast edge weight augmentation in linear time, where the eigenpairs of Laplacian for are graph frequencies. Next, we select a node subset for sampling to minimize the error of the signal interpolated from samples in two steps. We first align all Gershgorin disc left-ends of Laplacian at the smallest eigenvalue via similarity transform , leveraging a recent linear algebra theorem called Gershgorin disc perfect alignment (GDPA). We then perform sampling on using a previous fast Gershgorin disc alignment sampling (GDAS) scheme. Experiments show that our signed graph sampling method outperformed fast sampling schemes designed for positive graphs on various datasets with anti-correlations.
Chinthaka Dinesh, Gene Cheung, Saghar Bagheri, Ivan V. Bajic
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Rate-Distortion Theory in Coding for Machines and Its Applications
abstract
Recent years have seen a tremendous growth in both the capability and popularity of automatic machine analysis of media, especially images and video. As a result, a growing need for efficient compression methods optimised for machine vision, rather than human vision, has emerged. To meet this growing demand, significant developments have been made in image and video coding for machines. Unfortunately, while there is a substantial body of knowledge regarding rate-distortion theory for human vision, the same cannot be said of machine analysis. In this paper, we greatly extend the current rate-distortion theory for machines, providing insight into important design considerations of machine-vision codecs. We then utilise this newfound understanding to improve several methods for learned image coding for machines. Our proposed methods achieve state-of-the-art rate-distortion performance on several computer vision tasks - classification, instance and semantic segmentation, and object detection.
Alon Harell, Yalda Foroutan, Nilesh A. Ahuja, Parual Datta, Bhavya Kanzariya, V. Srinivasa Somayazulu, Omesh Tickoo, Anderson de Andrade, Ivan V. Bajic
IEEE Trans. Pattern Anal. Mach. Intell.9
2024 A Fast Four-Parameter Affine Motion Compensation Algorithm for Video Coding
abstract
This paper proposes a fast four-parameter Affine Motion Compensation (AMC) algorithm. As shown in Fig. 1, the translation Motion Vector (MV) is derived by reusing the AMC sub-block MV derivation method firstly, which is used to conduct translation pre-transform. Secondly, a coordinate system whose coordinate origin is located on its top-left control point is established for the transformed block. Finally, the geometric relationship between two control point motion vectors (CPMVs) of the transformed block can be described as follows,\begin{equation*}\delta = \left| {\left({m{v_{0x}} - m{v_{1x}}}\right) \times H - \left({m{v_{0y}} + m{v_{1y}}}\right) \times W} \right| = 0\tag{1}\end{equation*}
Jiaqi Zhang 0007, Ivan V. Bajic, Shanshe Wang, Songlin Sun
DCC3
2024 Learned Compression of Encoding Distributions
abstract
The entropy bottleneck introduced by Ballé et al. [1] is a common component used in many learned compression models. It encodes a transformed latent representation using a static distribution whose parameters are learned during training. However, the actual distribution of the latent data may vary wildly across different inputs. The static distribution attempts to encompass all possible input distributions, thus fitting none of them particularly well. This unfortunate phenomenon, sometimes known as the amortization gap, results in suboptimal compression. To address this issue, we propose a method that dynamically adapts the encoding distribution to match the latent data distribution for a specific input. First, our model estimates a better encoding distribution for a given input. This distribution is then compressed and transmitted as an additional side-information bitstream. Finally, the decoder reconstructs the encoding distribution and uses it to decompress the corresponding latent data. Our method achieves a Bjøntegaard-Delta (BD)-rate gain of -7.10% on the Kodak test dataset when applied to the standard fully-factorized architecture. Furthermore, considering computational complexity, the transform used by our method is an order of magnitude cheaper in terms of Multiply-Accumulate (MAC) operations compared to related side-information methods such as the scale hyperprior.
Mateen Ulhaq, Ivan V. Bajic
ICIP2
2024 Learned Multimodal Compression for Autonomous Driving
abstract
Autonomous driving sensors generate an enormous amount of data. In this paper, we explore learned multimodal compression for autonomous driving, specifically targeted at 3D object detection. We focus on camera and LiDAR modalities and explore several coding approaches. One approach involves joint coding of fused modalities, while others involve coding one modality first, followed by conditional coding of the other modality. We evaluate the performance of these coding schemes on the nuScenes dataset. Our experimental results indicate that joint coding of fused modalities yields better results compared to the alternatives.
Hadi Hadizadeh, Ivan V. Bajic
MMSP2
2024 Scalable Human-Machine Point Cloud Compression
abstract
Due to the limited computational capabilities of edge devices, deep learning inference can be quite expensive. One remedy is to compress and transmit point cloud data over the network for server-side processing. Unfortunately, this approach can be sensitive to network factors, including available bitrate. Luckily, the bitrate requirements can be reduced without sacrificing inference accuracy by using a machine task-specialized codec. In this paper, we present a scalable codec for point-cloud data that is specialized for the machine task of classification, while also providing a mechanism for human viewing. In the proposed scalable codec, the “base” bitstream supports the machine task, and an “enhancement” bitstream may be used for better input reconstruction performance for human viewing. We base our architecture on PointNet++, and test its efficacy on the ModelNet40 dataset. We show significant improvements over prior non-specialized codecs.
Mateen Ulhaq, Ivan V. Bajic
PCS2
2024 Privacy-Preserving Autoencoder for Collaborative Object Detection
abstract
Privacy is a crucial concern in collaborative machine vision where a part of a Deep Neural Network (DNN) model runs on the edge, and the rest is executed on the cloud. In such applications, the machine vision model does not need the exact visual content to perform its task. Taking advantage of this potential, private information could be removed from the data insofar as it does not significantly impair the accuracy of the machine vision system. In this paper, we present an autoencoder-style network integrated within an object detection pipeline, which generates a latent representation of the input image that preserves task-relevant information while removing private information. Our approach employs an adversarial training strategy that not only removes private information from the bottleneck of the autoencoder but also promotes improved compression efficiency for feature channels coded by conventional codecs like VVC-Intra. We assess the proposed system using a realistic evaluation framework for privacy, directly measuring face and license plate recognition accuracy. Experimental results show that our proposed method is able to reduce the bitrate significantly at the same object detection accuracy compared to coding the input images directly, while keeping the face and license plate recognition accuracy on the images recovered from the bottleneck features low, implying strong privacy protection. Our code is available at https://github.com/bardia-az/ppa-code.
Bardia Azizian, Ivan V. Bajic
IEEE Trans. Image Process.2
2023 Smart Split-Federated Learning over Noisy Channels for Embryo Image Segmentation
abstract
Split-Federated (SplitFed) learning is an extension of federated learning that places minimal requirements on the clients’ computing infrastructure, since only a small portion of the overall model is deployed on the clients’ hardware. In SplitFed learning, feature values, gradient updates, and model updates are transferred across communication channels. In this paper, we study the effects of noise in the communication channels on the learning process and the quality of the final model. We propose a smart averaging strategy for SplitFed learning with the goal of improving resilience against channel noise. Experiments on a segmentation model for embryo images shows that the proposed smart averaging strategy is able to tolerate two orders of magnitude stronger noise in the communication channels compared to conventional averaging, while still maintaining the accuracy of the final model.
Zahra Hafezi Kafshgari, Ivan V. Bajic, Parvaneh Saeedi
ICASSP2
2023 Base Layer Efficiency in Scalable Human-Machine Coding
abstract
A basic premise in scalable human-machine coding is that the base layer is intended for AUTOMATED machine analysis and is therefore more compressible than the same content would be for human viewing. Use cases for such coding include video surveillance and traffic monitoring, where the majority of the content will never be seen by humans. Therefore, base layer efficiency is of paramount importance because the system would most frequently operate at the base-layer rate. In this paper, we analyze the coding efficiency of the base layer in a state-of-the-art scalable human-machine image codec, and show that it can be improved. In particular, we demonstrate that gains of 20-40% in BD-Rate compared to the currently best results on object detection and instance segmentation are possible.
Yalda Foroutan, Alon Harell, Anderson de Andrade, Ivan V. Bajic
ICIP4
2023 Grad-FEC: Unequal Loss Protection of Deep Features in Collaborative Intelligence
abstract
Collaborative intelligence (CI) involves dividing an artificial intelligence (AI) model into two parts: front-end, to be deployed on an edge device, and back-end, to be deployed in the cloud. The deep feature tensors produced by the front-end are transmitted to the cloud through a communication channel, which may be subject to packet loss. To address this issue, in this paper, we propose a novel approach to enhance the resilience of the CI system in the presence of packet loss through Unequal Loss Protection (ULP). The proposed ULP approach involves a feature importance estimator, which estimates the importance of feature packets produced by the front-end, and then selectively applies Forward Error Correction (FEC) codes to protect important packets. Experimental results demonstrate that the proposed approach can significantly improve the reliability and robustness of the CI system in the presence of packet loss.
Korcan Uyanik, S. Faegheh Yeganli, Ivan V. Bajic
ICIP3
2023 Multi-Task Learning for Screen Content Image Coding
abstract
With the rise of remote work and collaboration, compression of screen content images (SCI) is becoming increasingly important. While there are efficient codecs for natural images, as well as codecs for purely-synthetic images, those SCIs that contain both synthetic and natural content pose a particular challenge. In this paper, we propose a learning-based image coding model developed for such SCIs. By training an encoder to provide a latent representation suitable for two tasks – input reconstruction and synthetic/natural region segmentation – we create an effective SCI image codec whose strong performance is verified through experiments. Once trained, the second task (segmentation) need not be used; the codec still benefits from the segmentation-friendly latent representation.
Rashid Zamanshoar Heris, Ivan V. Bajic
ISCAS2
2023 Metaverse: A Young Gamer's Perspective
abstract
When developing technologies for the Metaverse, it is important to understand the needs and requirements of end users. Relatively little is known about the specific perspectives on the use of the Metaverse by the youngest audience: children ten and under. This paper explores the Metaverse from the perspective of a young gamer. It examines their understanding of the Metaverse in relation to the physical world and other technologies they may be familiar with, looks at some of their expectations of the Metaverse, and then relates these to the specific multimedia signal processing (MMSP) research challenges. The perspectives presented in the paper may be useful for planning more detailed subjective experiments involving young gamers, as well as informing the research on MMSP technologies targeted at these users.
Ivan V. Bajic, Teo Saeedi-Bajic, Kai Saeedi-Bajic
MMSP1
2023 Learned Point Cloud Compression for Classification
abstract
Deep learning is increasingly being used to perform machine vision tasks such as classification, object detection, and segmentation on 3D point cloud data. However, deep learning inference is computationally expensive. The limited computational capabilities of end devices thus necessitate a codec for transmitting point cloud data over the network for server-side processing. Such a codec must be lightweight and capable of achieving high compression ratios without sacrificing accuracy. Motivated by this, we present a novel point cloud codec that is highly specialized for the machine task of classification. Our codec, based on PointNet, achieves a significantly better rate-accuracy trade-off in comparison to alternative methods. In particular, it achieves a 94% reduction in BD-bitrate over non-specialized codecs on the ModelNet40 dataset. For low-resource end devices, we also propose two lightweight configurations of our encoder that achieve similar BD-bitrate reductions of 93% and 92% with 3% and 5% drops in top-1 accuracy, while consuming only 0.470 and 0.048 encoder-side kMACs/point, respectively. Our codec demonstrates the potential of specialized codecs for machine analysis of point clouds, and provides a basis for extension to more complex tasks and datasets in the future.
Mateen Ulhaq, Ivan V. Bajic
MMSP2
2023 Point Cloud Sampling via Graph Balancing and Gershgorin Disc Alignment
abstract
Point cloud (PC)—a collection of discrete geometric samples of a 3D object’s surface—is typically large, which entails expensive subsequent operations. Thus, PC sub-sampling is of practical importance. Previous model-based sub-sampling schemes are ad-hoc in design and do not preserve the overall shape sufficiently well, while previous data-driven schemes are trained for specific pre-determined input PC sizes and sub-sampling rates and thus do not generalize well. Leveraging advances in graph sampling, we propose a fast PC sub-sampling algorithm of linear time complexity that chooses a 3D point subset while minimizing a global reconstruction error. Specifically, to articulate a sampling objective, we first assume a super-resolution (SR) method based on feature graph Laplacian regularization (FGLR) that reconstructs the original high-res PC, given points chosen by a sampling matrix${\mathbf H}$. We prove that minimizing a worst-case SR reconstruction error is equivalent to maximizing the smallest eigenvalue$\lambda _{\min }$of matrix${\mathbf H}^{\top } {\mathbf H}+ \mu {\boldsymbol{\mathcal{L}}}$, where${\boldsymbol{\mathcal{L}}}$is a symmetric, positive semi-definite matrix derived from a neighborhood graph connecting the 3D points. To arrive at a fast algorithm, instead of maximizing$\lambda _{\min }$, we maximize a lower bound$\lambda ^-_{\min }({\mathbf H}^{\top } {\mathbf H}+ \mu {\boldsymbol{\mathcal{L}}})$via selection of${\mathbf H}$—this translates to a graph sampling problem for a signed graph${\mathcal G}$with self-loops specified by graph Laplacian${\boldsymbol{\mathcal{L}}}$. We tackle this general graph sampling problem in three steps. First, we approximate${\mathcal G}$with a balanced graph${\mathcal G}_B$specified by Laplacian${\boldsymbol{\mathcal{L}}}_B$. Second, leveraging a recent linear algebraic theorem called Gershgorin disc perfect alignment (GDPA), we perform a similarity transform${\boldsymbol{\mathcal{L}}}_p~=~{\mathbf S}{\boldsymbol{\mathcal{L}}}_B {\mathbf S}^{-1}$, so that all Gershgorin disc left-ends of${\boldsymbol{\mathcal{L}}}_p$are aligned exactly at$\lambda _{\min }({\boldsymbol{\mathcal{L}}}_B)$. Finally, we choose samples on${\mathcal G}_B$using a previous graph sampling algorithm to maximize$\lambda ^-_{\min }({\mathbf H}^{\top } {\mathbf H}+ \mu {\boldsymbol{\mathcal{L}}}_p)$in linear time. Experimental results show that 3D points chosen by our algorithm outperformed competing schemes both numerically and visually in reconstruction quality.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Linear-Time Sampling on Signed Graphs Via Gershgorin Disc Perfect Alignment
abstract
In graph signal processing (GSP), an appropriate underlying graph encodes pairwise (anti-)correlations of targeted discrete signals as edge weights. However, existing fast graph sampling schemes are designed and tested for positive graphs describing only positive correlations. In this paper, we show that for datasets with inherent strong anti-correlations, a suitable graph structure is instead a signed graph with both positive and negative edge weights, and in response, we propose a linear-time signed graph sampling method. Specifically, given an empirical covariance data matrix ${\mathbf{\bar C}}$, we first employ graphical lasso to learn a sparse inverse matrix $\mathcal{L}$, interpreted as a generalized graph Laplacian for signed graph $\mathcal{G}$. We then propose a fast signed graph sampling scheme containing three steps: i) augment $\mathcal{G}$ to a balanced graph ${\mathcal{G}_B}$, ii) align all Gershgorin disc left-ends of corresponding Laplacian ${\mathcal{L}_B}$ at smallest eigenvalue ${\lambda _{\min }}\left( {{\mathcal{L}_B}} \right)$ via similarity transform ${\mathcal{L}_p} = {\mathbf{S}}{\mathcal{L}_B}{{\mathbf{S}}^{ - 1}}$, leveraging a recent linear algebra theorem called Gershgorin disc perfect alignment (GDPA), and iii) perform sampling on ${\mathcal{L}_p}$ using a previous fast Gershgorin disc alignment sampling scheme (GDAS). Experimental results show that our signed graph sampling method outperformed existing fast sampling schemes noticeably on two political voting datasets.
Chinthaka Dinesh, Saghar Bagheri, Gene Cheung, Ivan V. Bajic
ICASSP4
2022 Does Video Compression Impact Tracking Accuracy?
abstract
Everyone “knows” that compressing a video will degrade the accuracy of object tracking. Yet, a literature search on this topic reveals that there is very little documented evidence for this presumed fact. Part of the reason is that, until recently, there were no object tracking datasets for uncompressed video, which made studying the effects of compression on tracking accuracy difficult. In this paper, using a recently published dataset that contains tracking annotations for uncompressed videos, we examined the degradation of tracking accuracy due to video compression using rigorous statistical methods. Specifically, we examined the impact of quantization parameter (QP) and motion search range (MSR) on Multiple Object Tracking Accuracy (MOTA). The results show that QP impacts MOTA at the 95% confidence level, while there is insufficient evidence to claim that MSR impacts MOTA. Moreover, regression analysis allows us to derive a quantitative relationship between MOTA and QP for the specific tracker used in the experiments.
Takehiro Tanaka, Alon Harell, Ivan V. Bajic
ISCAS3
2022 Scalable Video Coding for Humans and Machines
abstract
Video content is watched not only by humans, but increasingly also by machines. For example, machine learning models analyze surveillance video for security and traffic moni-toring, search through YouTube videos for inappropriate content, and so on. In this paper, we propose a scalable video coding framework that supports machine vision (specifically, object detection) through its base layer bitstream and human vision via its enhancement layer bitstream. The proposed framework includes components from both conventional and Deep Neural Network (DNN)-based video coding. The results show that on object detection, the proposed framework achieves 13–19% bit savings compared to state-of-the-art video codecs, while remaining competitive in terms of MS-SSIM on the human vision task.
Hyomin Choi, Ivan V. Bajic
MMSP2
2022 Privacy-Preserving Feature Coding for Machines
abstract
Automated machine vision pipelines do not need the exact visual content to perform their tasks. Therefore, there is a potential to remove private information from the data without significantly affecting the machine vision accuracy. We present a novel method to create a privacy-preserving latent representation of an image that could be used by a downstream machine vision model. This latent representation is constructed using adversarial training to prevent accurate reconstruction of the input while preserving the task accuracy. Specifically, we split a Deep Neural Network (DNN) model and insert an autoencoder whose purpose is to both reduce the dimensionality as well as remove information relevant to input reconstruction while minimizing the impact on task accuracy. Our results show that input reconstruction ability can be reduced by about 0.8 dB at the equivalent task accuracy, with degradation concentrated near the edges, which is important for privacy. At the same time, 30% bit savings are achieved compared to coding the features directly.
Bardia Azizian, Ivan V. Bajic
PCS2
2022 Rate-Distortion in Image Coding for Machines
abstract
In recent years, there has been a sharp increase in transmission of images to remote servers specifically for the purpose of computer vision. In many applications, such as surveillance, images are mostly transmitted for automated analysis, and rarely seen by humans. Using traditional compression for this scenario has been shown to be inefficient in terms of bit-rate, likely due to the focus on human based distortion metrics. Thus, it is important to create specific image coding methods for joint use by humans and machines. One way to create the machine side of such a codec is to perform feature matching of some intermediate layer in a Deep Neural Network performing the machine task. In this work, we explore the effects of the layer choice used in training a learnable codec for humans and machines. We prove, using the data processing inequality, that matching features from deeper layers is preferable in the sense of rate-distortion. Next, we confirm our findings empirically by re-training an existing model for scalable human-machine coding. In our experiments we show the trade-off between the human and machine sides of such a scalable model, and discuss the benefit of using deeper layers for training in that regard.
Alon Harell, Anderson de Andrade, Ivan V. Bajic
PCS3
2022 Scalable Image Coding for Humans and Machines
abstract
At present, and increasingly so in the future, much of the captured visual content will not be seen by humans. Instead, it will be used for automated machine vision analytics and may require occasional human viewing. Examples of such applications include traffic monitoring, visual surveillance, autonomous navigation, and industrial machine vision. To address such requirements, we develop an end-to-end learned image codec whose latent space is designed to support scalability from simpler to more complicated tasks. The simplest task is assigned to a subset of the latent space (the base layer), while more complicated tasks make use of additional subsets of the latent space, i.e., both the base and enhancement layer(s). For the experiments, we establish a 2-layer and a 3-layer model, each of which offers input reconstruction for human vision, plus machine vision task(s), and compare them with relevant benchmarks. The experiments show that our scalable codecs offer 37%-80% bitrate savings on machine vision tasks compared to best alternatives, while being comparable to state-of-the-art image codecs in terms of input reconstruction.
Hyomin Choi, Ivan V. Bajic
IEEE Trans. Image Process.2
2022 Point Cloud Video Super-Resolution via Partial Point Coupling and Graph Smoothness
abstract
Point cloud (PC) is a collection of discrete geometric samples of a physical object in 3D space. A PC video consists of temporal frames evenly spaced in time, each containing a static PC at one time instant. PCs in adjacent frames typically do not have point-to-point (P2P) correspondence, and thus exploiting temporal redundancy for PC restoration across frames is difficult. In this paper, we focus on the super-resolution (SR) problem for PC video: increase point density of PCs in video frames while preserving salient geometric features consistently across time. We accomplish this with two ideas. First, we establish partial P2P coupling between PCs of adjacent frames by interpolating interior points in a low-resolution PC patch in frame t and translating them to a corresponding patch in frame t+1 , via a motion model computed by iterative closest point (ICP). Second, we promote piecewise smoothness in 3D geometry in each patch using feature graph Laplacian regularizer (FGLR) in an easily computable quadratic form. The two ideas translate to an unconstrained quadratic programming (QP) problem with a system of linear equations as solution-one where we ensure the numerical stability by upper-bounding the condition number of the coefficient matrix. Finally, to improve the accuracy of the ICP motion model, we re-sample points in a super-resolved patch at time t to better match a low-resolution patch at time t+1 via bipartite graph matching after each SR iteration. Experimental results show temporally consistent super-resolved PC videos generated by our scheme, outperforming SR competitors that optimized on a per-frame basis, in two established PC metrics.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic
IEEE Trans. Image Process.3
2021 Collaborative Intelligence: Challenges and Opportunities
abstract
This paper presents an overview of the emerging area of collaborative intelligence (CI). Our goal is to raise awareness in the signal processing community of the challenges and opportunities in this area of growing importance, where key developments are expected to come from signal processing and related disciplines. The paper surveys the current state of the art in CI, with special emphasis on signal processing-related challenges in feature compression, error resilience, privacy, and system-level design.
Ivan V. Bajic, Weisi Lin, Yonghong Tian 0001
ICASSP1
2021 Latent Space Motion Analysis for Collaborative Intelligence
abstract
When the input to a deep neural network (DNN) is a video signal, a sequence of feature tensors is produced at the intermediate layers of the model. If neighboring frames of the input video are related through motion, a natural question is, "what is the relationship between the corresponding feature tensors?" By analyzing the effect of common DNN operations on optical flow, we show that the motion present in each channel of a feature tensor is approximately equal to the scaled version of the input motion. The analysis is validated through experiments utilizing common motion models.
Mateen Ulhaq, Ivan V. Bajic
ICASSP2
2021 Latent Space Inpainting for Loss-Resilient Collaborative Object Detection
abstract
Edge devices, such as cameras and mobile units, are increasingly capable of performing sophisticated computation in addition to their traditional roles in sensing and communicating signals. The focus of this paper is on collaborative object detection, where deep features computed on the edge device from input images are transmitted to the cloud for further processing. We consider the impact of packet loss on the transmitted features and examine several ways for recovering the missing data. In particular, through theory and experiments, we show that methods for image inpainting based on partial differential equations work well for the recovery of missing features in the latent space. The obtained results represent the new state of the art for missing data recovery in collaborative object detection.
Ivan V. Bajic
ICC1
2021 Latent-Space Scalability for Multi-Task Collaborative Intelligence
abstract
We investigate latent-space scalability for multi-task collaborative intelligence, where one of the tasks is object detection and the other is input reconstruction. In our proposed approach, part of the latent space can be selectively decoded to support object detection while the remainder can be decoded when input reconstruction is needed. Such an approach allows reduced computational resources when only object detection is required, and this can be achieved without reconstructing input pixels. By varying the scaling factors of various terms in the training loss function, the system can be trained to achieve various trade-offs between object detection accuracy and input reconstruction quality. Experiments are conducted to demonstrate the adjustable system performance on the two tasks compared to the relevant benchmarks.
Hyomin Choi, Ivan V. Bajic
ICIP2
2021 CALTEC: Content-Adaptive Linear Tensor Completion For Collaborative Intelligence
abstract
In collaborative intelligence, an artificial intelligence (AI) model is typically split between an edge device and the cloud. Feature tensors produced by the edge sub-model are sent to the cloud via an imperfect communication channel. At the cloud side, parts of the feature tensor may be missing due to packet loss. In this paper we propose a method called Content-Adaptive Linear Tensor Completion (CALTeC) to recover the missing feature data. The proposed method is fast, data-adaptive, does not require pre-training, and produces better results than existing methods for tensor data recovery in collaborative intelligence.
Ashiv Dhondea, Robert A. Cohen, Ivan V. Bajic
ICIP3
2021 Scalable Privacy in Multi-Task Image Compression
abstract
Learning-based compression systems have shown great potential for multi-task inference from their latent-space representation of the input image. In such systems, the decoder is supposed to be able to perform various analyses of the input image, such as object detection or segmentation, besides decoding the image. At the same time, privacy concerns around visual ana-lytics have grown in response to the increasing capabilities of such systems to reveal private information. In this paper, we propose a method to make latent-space inference more privacy-friendly using mutual information-based criteria. In particular, we show how organizing and compressing the latent representation of the image according to task-specific mutual information can make the model maintain high analytics accuracy while becoming less able to reconstruct the input image and thereby reveal private information.
Saeed Ranjbar Alvar, Ivan V. Bajic
VCIP2
2021 DFTS2: Deep Feature Transmission Simulation for Collaborative Intelligence
abstract
In edge-cloud collaborative intelligence (CI), an unreliable transmission channel exists in the information path of the AI model performing the inference. It is important to be able to simulate the performance of the CI system across an imperfect channel in order to understand system behavior and develop appropriate error control strategies. In this paper we present a simulation framework called DFTS2, which enables researchers to define the components of the CI system in TensorFlow 2, select a packet-based channel model with various parameters, and simulate system behavior under various channel conditions and error/loss control strategies. Using DFTS2, we also present the most comprehensive study to date of the packet loss concealment methods for collaborative image classification models.
Ashiv Dhondea, Robert A. Cohen, Ivan V. Bajic
VCIP3
2021 Pareto-Optimal Bit Allocation for Collaborative Intelligence
abstract
In recent studies, collaborative intelligence (CI) has emerged as a promising framework for deployment of Artificial Intelligence (AI)-based services on mobile/edge devices. In CI, the AI model (a deep neural network) is split between the edge and the cloud, and intermediate features are sent from the edge sub-model to the cloud sub-model. In this article, we study bit allocation for feature coding in multi-stream CI systems. We model task distortion as a function of rate using convex surfaces similar to those found in distortion-rate theory. Using such models, we are able to provide closed-form bit allocation solutions for single-task systems and scalarized multi-task systems. Moreover, we provide analytical characterization of the full Pareto set for 2-stream k -task systems, and bounds on the Pareto set for 3-stream 2-task systems. Analytical results are examined on a variety of DNN models from the literature to demonstrate wide applicability of the results.
Saeed Ranjbar Alvar, Ivan V. Bajic
IEEE Trans. Image Process.2
2021 Affine Transformation-Based Deep Frame Prediction
abstract
We propose a neural network model to estimate the current frame from two reference frames, using affine transformation and adaptive spatially-varying filters. The estimated affine transformation allows for using shorter filters compared to existing approaches for deep frame prediction. The predicted frame is used as a reference for coding the current frame. Since the proposed model is available at both encoder and decoder, there is no need to code or transmit motion information for the predicted frame. By making use of dilated convolutions and reduced filter length, our model is significantly smaller, yet more accurate, than any of the neural networks in prior works on this topic. Two versions of the proposed model - one for unidirectional, and one for bi-directional prediction - are trained using a combination of Discrete Cosine Transform (DCT)-based ℓ1-loss with various transform sizes, multi-scale Mean Squared Error (MSE) loss, and an object context reconstruction loss. The trained models are integrated with the HEVC video coding pipeline. The experiments show that the proposed models achieve about 7.3%, 5.4%, and 4.2% bit savings for the luminance component on average in the Low delay P, Low delay, and Random access configurations, respectively.
Hyomin Choi, Ivan V. Bajic
IEEE Trans. Image Process.2
2021 Soft Video Multicasting Using Adaptive Compressed Sensing
abstract
Recently, soft video multicasting has gained a lot of attention, especially in broadcast and mobile scenarios where the bit rate supported by the channel may differ across receivers, and may vary quickly over time. Unlike the conventional designs that force the source to use a single bit rate according to the receiver with the worst channel quality, soft video delivery schemes transmit the video such that the video quality at each receiver is commensurate with its specific instantaneous channel quality. In this paper, we present a soft video multicasting system using an adaptive block-based compressed sensing (BCS) method. The proposed system consists of an encoder, a transmission system, and a decoder. At the encoder side, each block in each frame of the input video is adaptively sampled with a rate that depends on the texture complexity and visual saliency of the block. The obtained BCS samples are then placed into several packets, and the packets are transmitted via a channel-aware OFDM (orthogonal frequency division multiplexing) transmission system with a number of subchannels. At the decoder side, the received BCS samples are first used to build an initial approximation of the transmitted frame. To further improve the reconstruction quality, an iterative BCS reconstruction algorithm is then proposed that uses an adaptive transform and an adaptive soft-thresholding operator, which exploits the temporal similarity between adjacent frames to achieve better reconstruction quality. The extensive objective and subjective experimental results indicate the superiority of the proposed system over the state-of-the-art soft video multicasting systems.
Hadi Hadizadeh, Ivan V. Bajic
IEEE Trans. Multim.2
2020 Bit Allocation for Multi-Task Collaborative Intelligence
abstract
Recent studies have shown that collaborative intelligence (CI) is a promising framework for deployment of Artificial Intelligence (AI)-based services on mobile devices. In CI, a deep neural network is split between the mobile device and the cloud. Deep features obtained at the mobile are compressed and transferred to the cloud to complete the inference. So far, the methods in the literature focused on transferring a single deep feature tensor from the mobile to the cloud. Such methods are not applicable to some recent, high-performance networks with multiple branches and skip connections. In this paper, we propose the first bit allocation method for multi-stream, multi-task CI. We first establish a model for the joint distortion of the multiple tasks as a function of the bit rates assigned to different deep feature tensors. Then, using the proposed model, we solve the rate-distortion optimization problem under a total rate constraint to obtain the best rate allocation among the tensors to be transferred. Experimental results illustrate the efficacy of the proposed scheme compared to several alternative bit allocation methods.
Saeed Ranjbar Alvar, Ivan V. Bajic
ICASSP2
2020 Back-And-Forth Prediction for Deep Tensor Compression
abstract
Recent AI applications such as Collaborative Intelligence with neural networks involve transferring deep feature tensors between various computing devices. This necessitates tensor compression in order to optimize the usage of bandwidth-constrained channels between devices. In this paper we present a prediction scheme called Back-and-Forth (BaF) prediction, developed for deep feature tensors, which allows us to dramatically reduce tensor size and improve its compressibility. Our experiments with a state-of-the-art object detector demonstrate that the proposed method allows us to significantly reduce the number of bits needed for compressing feature tensors extracted from deep within the model, with negligible degradation of the detection performance and without requiring any retraining of the network weights. We achieve a 62% and 75% reduction in tensor size while keeping the loss in accuracy of the network to less than 1% and 2%, respectively.
Hyomin Choi, Robert A. Cohen, Ivan V. Bajic
ICASSP3
2020 Super-Resolution of 3D Color Point Clouds Via Fast Graph Total Variation
abstract
3D point clouds acquired by low-cost sensors are often in lower spatial resolutions than desired for rendering images on high-resolution displays. In this paper, we propose a fast super-resolution (SR) algorithm for color 3D point clouds. We first populate a target low-res point cloud with added interior points. We refine the newly added 3D coordinates and their RGB values by minimizing a graph total variation (GTV) term of connected points' surface normals and RGB values respectively. Unlike non-local methods that require computation-intensive searches of similar patches in a large defined space, our algorithm is inherently local and performs smoothing of newly inserted points only with respect to neighboring points. Moreover, differing from our previous GTV-based SR algorithm that employs gradient descent procedures with sensitive step size parameters due to GTV's non-smooth l1-norm, we rewrite the l1objective into a linear proxy, so that together with constraints on surface normals / RGB values, it can be solved efficiently as a parameter-free linear program (LP). Experimental results show that our algorithm outperforms competing non-graph-based point cloud SR schemes, and is significantly faster than our previous graph-based SR method.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic
ICASSP3
2020 Sampling Of 3d Point Cloud Via Gershgorin Disc Alignment
abstract
Point cloud-a collection of geometric samples of a physical object in 3D space-can be very large in size, which entails a large computation cost for many imaging applications. In this paper, we reduce the size of a point cloud towards a more compact representation via optimal sub-sampling. Specifically, we first derive a sampling objective that maximizes the stability (maximizes the smallest eigenvalue λmin(B) of a coefficient matrix B = HTH + μL) of a linear system super-resolving a sub-sampled point cloud. To circumvent eigen-decomposition, we maximize instead a lower bound λmin-(B) using a fast graph sampling scheme called Gershgorin disc alignment (GDA) based on the well-known Gershgorin circle theorem. However, GDA requires that the disc left-ends of real matrix L are initially aligned at the same value, which is not the case for point clouds. Orthogonally, we recently derived a matrix theorem proving that disc left-ends of a generalized graph Laplacian matrix for a balanced and irreducible signed graph can be perfectly aligned via a similarity transform using the matrix's first eigenvector. Leveraging this work, we first interpret L as a generalized graph Laplacian matrix and balance the underlying graph. We then align disc left-ends of the resulting generalized graph Laplacian of the balanced graph using its first eigenvector, in order to employ GDA for point cloud sampling. Experiments show that our sampling method outperforms competing methods in super-resolved point cloud quality.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic
ICIP4
2020 Lightweight Compression Of Neural Network Feature Tensors For Collaborative Intelligence
abstract
In collaborative intelligence applications, part of a deep neural network (DNN) is deployed on a relatively low-complexity device such as a mobile phone or edge device, and the remainder of the DNN is processed where more computing resources are available, such as in the cloud. This paper presents a novel lightweight compression technique designed specifically to code the activations of a split DNN layer, while having a low complexity suitable for edge devices and not requiring any retraining. We also present a modified entropy-constrained quantizer design algorithm optimized for clipped activations. When applied to popular object-detection and classification DNNs, we were able to compress the 32-bit floating point activations down to 0.6 to 0.8 bits, while keeping the loss in accuracy to less than 1%. When compared to HEVC, we found that the lightweight codec consistently provided better inference accuracy, by up to 1.3%. The performance and simplicity of this lightweight compression technique makes it an attractive option for coding a layer's activations in split neural networks for edge/cloud applications.
Robert A. Cohen, Hyomin Choi, Ivan V. Bajic
ICME3
2020 Light field all-in-focus image fusion based on spatially-guided angular information
Yingchun Wu, Yumei Wang, Jie Liang 0001, Ivan V. Bajic, Anhong Wang
J. Vis. Commun. Image Represent.4
2020 Deep Frame Prediction for Video Coding
abstract
We propose a novel frame prediction method using a deep neural network (DNN), with the goal of improving the video coding efficiency. The proposed DNN makes use of decoded frames, at both the encoder and decoder to predict the textures of the current coding block. Unlike conventional inter-prediction, the proposed method does not require any motion information to be transferred between the encoder and the decoder. Still, both the uni-directional and bi-directional predictions are possible using the proposed DNN, which is enabled by the use of the temporal index channel, in addition to the color channels. In this paper, we developed a jointly trained DNN for both uni-directional and bi-directional predictions, as well as separate networks for uni-directional and bi-directional predictions, and compared the efficacy of both the approaches. The proposed DNNs were compared with the conventional motion-compensated prediction in the latest video coding standard, High Efficiency Video Coding (HEVC), in terms of the BD-bitrate. The experiments show that the proposed joint DNN (for both uni-directional and bi-directional predictions) reduces the luminance bitrate by about 4.4%, 2.4%, and 2.3% in the low delay $P$ , low delay, and random access configurations, respectively. In addition, using the separately trained DNNs brings further bit savings of about 0.3%-0.5%.
Hyomin Choi, Ivan V. Bajic
IEEE Trans. Circuits Syst. Video Technol.2
2020 Point Cloud Denoising via Feature Graph Laplacian Regularization
abstract
Point cloud is a collection of 3D coordinates that are discrete geometric samples of an object's 2D surfaces. Imperfection in the acquisition process means that point clouds are often corrupted with noise. Building on recent advances in graph signal processing, we design local algorithms for 3D point cloud denoising. Specifically, we design a signal-dependent feature graph Laplacian regularizer (SDFGLR) that assumes surface normals computed from point coordinates are piecewise smooth with respect to a signal-dependent graph Laplacian matrix. Using SDFGLR as a signal prior, we formulate an optimization problem with a general 'p-norm fidelity term that can explicitly remove only two types of additive noise: small but non-sparse noise like Gaussian (using '2 fidelity term) and large but sparser noise like Laplacian (using '1 fidelity term). To establish a linear relationship between normals and 3D point coordinates, we first perform bipartite graph approximation to divide the point cloud into two disjoint node sets (red and blue). We then optimize the red and blue nodes' coordinates alternately. For '2-norm fidelity term, we iteratively solve an unconstrained quadratic programming (QP) problem, efficiently computed using conjugate gradient with a bounded condition number to ensure numerical stability. For '1-norm fidelity term, we iteratively minimize an '1-'2 cost function using accelerated proximal gradient (APG), where a good step size is chosen via Lipschitz continuity analysis. Finally, we propose simple mean and median filters for flat patches of a given point cloud to estimate the noise variance given the noise type, which in turn is used to compute a weight parameter trading off the fidelity term and signal prior in the problem formulation. Extensive experiments show state-of-the-art denoising performance among local methods using our proposed algorithms.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic
IEEE Trans. Image Process.3
2019 Wavenilm: A Causal Neural Network for Power Disaggregation from the Complex Power Signal
abstract
Non-intrusive load monitoring (NILM) helps meet energy conservation goals by estimating individual appliance power usage from a single aggregate measurement. Deep neural networks have become increasingly popular in attempting to solve NILM problems; however, many of them are not causal which is important for real-time application. We present a causal 1-D convolutional neural network inspired by WaveNet for NILM on low-frequency data. We also study using various components of the complex power signal for NILM, and demonstrate that using all four components available in a popular NILM dataset (current, active power, reactive power, and apparent power) we achieve faster convergence and higher performance than state-of-the-art results for the same dataset.
Alon Harell, Stephen Makonin, Ivan V. Bajic
ICASSP3
2019 Multi-Task Learning with Compressible Features for Collaborative Intelligence
abstract
A promising way to deploy Artificial Intelligence (AI)-based services on mobile devices is to run a part of the AI model (a deep neural network) on the mobile itself, and the rest in the cloud. This is sometimes referred to as collaborative intelligence. In this framework, intermediate features from the deep network need to be transmitted to the cloud for further processing. We study the case where such features are used for multiple purposes in the cloud (multi-tasking) and where they need to be compressible in order to allow efficient transmission to the cloud. To this end, we introduce a new loss function that encourages feature compressibility while improving system performance on multiple tasks. Experimental results show that with the compression-friendly loss, one can achieve around 20% bitrate reduction without sacrificing the performance on several vision-related tasks.
Saeed Ranjbar Alvar, Ivan V. Bajic
ICIP2
2019 3D Point Cloud Super-Resolution via Graph Total Variation on Surface Normals
abstract
Point cloud is a collection of 3D coordinates that are discrete geometric samples of an object's 2D surfaces. Using a low-cost 3D scanner to acquire data means that point clouds are often in lower resolution than desired for rendering on high-resolution displays. Building on recent advances in graph signal processing, we design a local algorithm for 3D point cloud super-resolution (SR). First, we initialize new points at centroids of local triangles formed using the low-resolution point cloud, and connect all points using a k-nearest-neighbor graph. Then, to establish a linear relationship between surface normals and 3D point coordinates, we perform bipartite graph approximation to divide all nodes into two disjoint sets, which are optimized alternately until convergence. For each node set, to promote piecewise smooth (PWS) 2D surfaces, we design a graph total variation (GTV) objective for nearby surface normals, under the constraint that coordinates of the original points are preserved. We pursue an augmented Lagrangian approach to tackle the optimization, and solve the unconstrained equivalent using the alternating method of multipliers (ADMM). Extensive experiments show that our proposed point cloud SR algorithm outperforms competing schemes objectively and subjectively for a large variety of point clouds.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic
ICIP3
2019 3D Point Cloud Color Denoising Using Convex Graph-Signal Smoothness Priors
abstract
Point cloud is a collection of 3D coordinates and associated color information, which are discrete samples of an object's 2D surfaces. Imperfection in the acquisition process means that point clouds are often corrupted with noise in both geometric and color spaces. Building on recent advances in graph signal processing, we design two algorithms for 3D point cloud color denoising. Specifically, we develop a smoothness notion for 3D color point clouds using graph Laplacian regularizer (GLR) or graph total variation (GTV) priors defined on the RGB space to promote piecewise smoothness (PWS) of RGB values. Using GLR or GTV as signal prior, we formulate the point cloud color denoising problem as a maximum a posteriori (MAP) estimation problem. For the GLR prior, the MAP formulation leads to an unconstrained quadratic programming (QP) problem, which can be efficiently computed using conjugate gradient (CG). For the GTV prior, the MAP formulation results in an ℓ1-ℓ2cost function; we minimize it using alternating direction method of multipliers (ADMM) and proximal gradient descent, where a good step size is chosen via Lipschitz continuity analysis. Extensive experiments show satisfactory denoising performance using our proposed algorithms.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic
MMSP3
2019 A Perceptual Distinguishability Predictor For JND-Noise-Contaminated Images
abstract
Just noticeable difference (JND) models are widely used for perceptual redundancy estimation in images and videos. A common method for measuring the accuracy of a JND model is to inject random noise in an image based on the JND model, and check whether the JND-noise-contaminated image is perceptually distinguishable from the original image or not. Also, when comparing the accuracy of two different JND models, the model that produces the JND-noise-contaminated image with better quality at the same level of noise energy is the better model. But in both of these cases, a subjective test is necessary, which is very time consuming and costly. In this paper, we present a full-reference metric called PDP (perceptual distinguishability predictor), which can be used to determine whether a given JND-noise-contaminated image is perceptually distinguishable from the reference image. The proposed metric employs the concept of sparse coding, and extracts a feature vector out of a given image pair. The feature vector is then fed to a multilayer neural network for classification. To train the network, we built a public database of 999 natural images with distinguishbility thresholds for four different JND models obtained from an extensive subjective experiment. The results indicated that PDD achieves high classification accuracy of 97.1%. The proposed method can be used to objectively compare various JND models without performing any subjective test. It can also be used to obtain proper scaling factors to improve the JND thresholds estimated by an arbitrary JND model.
Hadi Hadizadeh, Ahmad Reza Heravi, Ivan V. Bajic, Parastoo Karami
IEEE Trans. Image Process.3
2018 Can you Find a Face in a HEVC Bitstream?
abstract
Finding faces in images is one of the most important tasks in computer vision, with applications in biometrics, surveillance, human-computer interaction, and other areas. In our earlier work, we demonstrated that it is possible to tell whether or not an image contains a face by only examining the HEVC syntax, without fully reconstructing the image. In the present work we move further in this direction by showing how to localize faces in HEVC-coded images, without full reconstruction. We also demonstrate the benefits that such approach can have in privacy-friendly face localization.
Saeed Ranjbar Alvar, Hyomin Choi, Ivan V. Bajic
ICASSP3
2018 High Efficiency Compression for Object Detection
abstract
Image and video compression has traditionally been tailored to human vision. However, modern applications such as visual analytics and surveillance rely on computers “seeing” and analyzing the images before (or instead of) humans. For these applications, it is important to adjust compression to computer vision. In this paper we present a bit allocation and rate control strategy that is tailored to object detection. U sing the initial convolutional layers of a state-of-the-art object detector, we create an importance map that can guide bit allocation to areas that are important for object detection. The proposed method enables bit rate savings of 7% or more compared to default HEVC, at the equivalent object detection rate.
Hyomin Choi, Ivan V. Bajic
ICASSP2
2018 Deep Feature Compression for Collaborative Object Detection
abstract
Recent studies have shown that the efficiency of deep neural networks in mobile applications can be significantly improved by distributing the computational workload between the mobile device and the cloud. This paradigm, termed collaborative intelligence, involves communicating feature data between the mobile and the cloud. The efficiency of such approach can be further improved by lossy compression of feature data, which has not been examined to date. In this work we focus on collaborative object detection and study the impact of both near-lossless and lossy compression of feature data on its accuracy. We also propose a strategy for improving the accuracy under lossy feature compression. Experiments indicate that using this strategy, the communication overhead can be reduced by up to 70% without sacrificing accuracy.
Hyomin Choi, Ivan V. Bajic
ICIP2
2018 MV-YOLO: Motion Vector-Aided Tracking by Semantic Object Detection
abstract
Object tracking is the cornerstone of many visual analytics systems. While considerable progress has been made in this area in recent years, robust, efficient, and accurate tracking in real-world video remains a challenge. In this paper, we present a hybrid tracker that leverages motion information from the compressed video stream and a general-purpose semantic object detector acting on decoded frames to construct a fast and efficient tracking engine. The proposed approach is compared with several well-known recent trackers on the OTB tracking dataset. The results indicate advantages of the proposed method in terms of speed and/or accuracy. Other desirable features of the proposed method are its simplicity and deployment efficiency, which stems from the fact that it reuses the resources and information that may already exist in the system for other reasons.
Saeed Ranjbar Alvar, Ivan V. Bajic
MMSP2
2018 Near-Lossless Deep Feature Compression for Collaborative Intelligence
abstract
Collaborative intelligence is a new paradigm for efficient deployment of deep neural networks across the mobile-cloud infrastructure. By dividing the network between the mobile and the cloud, it is possible to distribute the computational workload such that the overall energy and/or latency of the system is minimized. However, this necessitates sending deep feature data from the mobile to the cloud in order to perform inference. In this work, we examine the differences between the deep feature data and natural image data, and propose a simple and effective near-lossless deep feature compressor. The proposed method achieves up to 5% bit rate reduction compared to HEVC-Intra and even more against other popular image codecs. Finally, we suggest an approach for reconstructing the input image from compressed deep features that could serve to supplement the inference performed by the deep model.
Hyomin Choi, Ivan V. Bajic
MMSP2
2018 Local 3D Point Cloud Denoising via Bipartite Graph Approximation & Total Variation
abstract
Acquired 3D point cloud data, whether from active sensors directly or from stereo-matching algorithms indirectly, typically contain non-negligible noise. To address the point cloud denoising problem, we propose a local graph-based algorithm. Specifically, given a $k$ -nearest-neighbor graph of the 3D points, we first approximate it with a bipartite graph (independent sets of red and blue nodes) using a KL divergence criterion. For each partite of nodes (say red), we first define surface normal of each red node using 3D coordinates of neighboring blue nodes, so that red node normals n can be written as a linear function of red node coordinates p. We then formulate a convex optimization problem, with a quadratic fidelity term $\Vert \mathbf{p}-\mathbf{q}\Vert_{2}^{2}$ given noisy observed red coordinates q and a graph total variation (GTV) regularization term for surface normals of neighboring red nodes. We minimize the resulting $l_{2}-l_{1}$-norm using alternating direction method of multipliers (ADMM) and proximal gradient descent. The two partites of nodes are alternately optimized until convergence. Experimental results show that compared to state-of-the-art schemes with similar complexity, our proposed algorithm achieves the best overall denoising performance objectively and subjectively.
Chinthaka Dinesh, Gene Cheung, Ivan V. Bajic, Cheng Yang 0003
MMSP3
2018 Speech Intelligibility of Microphone Arrays in Reverberant Environments with Interference
abstract
It is known that speech intelligibility degrades with additive noise and reverberation, and that quantitative parameters such as fidelity and signal-to-noise ratio can be improved by using microphone arrays with various beamforming algorithms. However, it is not clear how the array configuration impacts the intelligibility of speech. Numerical experiments, using widely-used models, provide the most convenient comparison, and the approach allows rapid assessment of parameters such as the array configuration, the number and spacing of the elements, and modeled features such as room reflection coefficients. For a typical reverberant room with a single wanted source and two unwanted sources (interferers), we compare the performance of two ceiling-mounted configurations - the uniform linear array (ULA) and a uniform circular array (UCA). The microphones are taken as omnidirectional and equispaced along the array loci, and we use a standard gain-constrained power minimization beamformer. In this study, a limiting performance is presented by emphasizing the early reflections over the late ones for the prior steering vector. Under this steering vector condition, for the same number of elements, the UCA easily outperforms the ULA on known quality and intelligibility metrics. For both arrays in this room scenario, all the metrics increase with an increasing number of microphones, although for one intelligibility metric, diminishing returns set in at about 12 microphones.
Elham Ideli, Rodney G. Vaughan, Ivan V. Bajic
MMSP3
2018 Adaptive Nonrigid Inpainting of Three-Dimensional Point Cloud Geometry
abstract
In this letter, we introduce several algorithms for geometry inpainting of three-dimensional (3-D) point clouds with large holes. The algorithms are exemplar based. Hole filling is performed iteratively using templates near the hole boundary to find the best matching regions elsewhere in the cloud, from where existing points are transferred to the hole. We propose two improvements over the previous work on exemplar-based hole filling. The first one is adaptive template size selection in each iteration, which simultaneously leads to higher accuracy and lower execution time. The second improvement is a nonrigid transformation to better align the candidate set of points with the template before the point transfer, which leads to even higher accuracy. We demonstrate the algorithm's ability to fill holes that are difficult or impossible to fill by existing methods.
Chinthaka Dinesh, Ivan V. Bajic, Gene Cheung
IEEE Signal Process. Lett.2
2018 Full-Reference Objective Quality Assessment of Tone-Mapped Images
abstract
In this paper we present a novel method for full-reference image quality assessment (IQA) of tone-mapped images displayed on standard low dynamic range (LDR) displays. Due to the dynamic range compression caused by the tone-mapping process a mixture of several artifacts and distortions may be produced in the tone-mapped images. This makes the quality assessment of the tone-mapped images very challenging. Due to the diversity of such artifacts and distortions we propose a “bag of features” (BOF) approach to tackle this problem. Specifically in the proposed method a number of different perceptually relevant quality-related features are first extracted from a given tone-mapped image and its reference HDR image. These features are designed such that they capture different aspects and attributes of the tone-mapped image such as its structural fidelity naturalness and overall brightness. A support vector regressor is then trained based on the extracted features and it is used for measuring the visual quality of a tone-mapped image. Our experimental results indicate that the proposed method achieves high accuracy as compared to several existing methods.
Hadi Hadizadeh, Ivan V. Bajic
IEEE Trans. Multim.2
2017 Corner proposals from HEVC bitstreams
abstract
Corner-like features are important in computer vision problems such as object matching, tracking, recognition, and retrieval. Most corner detectors operate in the pixel domain, which means that they require image or video to be fully decoded and reconstructed before detection can start. In this paper we describe a method for generating corner proposals from compressed HEVC bitstreams without full decoding. Specifically, we utilize HEVC syntax and intra prediction directions to find the locations that are likely to contain corners. The proposed method is lightweight and can be applied to intra-coded frames or still images coded by HEVC. Experimental results illustrate that the proposed method is able to identify most regions where conventional pixel-domain corner detectors would find corners.
Hyomin Choi, Ivan V. Bajic
ISCAS2
2017 Exemplar-based framework for 3D point cloud hole filling
abstract
Holes can arise in 3D point clouds due to a number of reasons such as incomplete scans, occlusions, and packet loss. We present an exemplar-based framework for hole filling in 3D point clouds, which exploits non-local self similarity to provide plausible reconstruction even for large holes and complex surfaces. Points along the hole boundary are given priority that determines the order in which they are processed. Hole filling is performed iteratively and uses templates near the hole boundary to find the best matching regions elsewhere in the cloud, from where existing points are transferred to the hole. The proposed method has been compared with several existing methods and has shown superior results, both visually and in terms of the Hausdorff distance.
Chinthaka Dinesh, Ivan V. Bajic, Gene Cheung
VCIP2
2017 Compressed-domain visual saliency models: a comparative study
Seyed Hossein Khatoonabadi, Ivan V. Bajic, Yufeng Shan
Multim. Tools Appl.2
2017 Saliency-Guided Just Noticeable Distortion Estimation Using the Normalized Laplacian Pyramid
abstract
The human visual system (HVS), like any other physical system, has limitations. For instance, it is known that the HVS can only sense the content changes that are larger than the so-called just noticeable distortion (JND) threshold. Also, to reduce the computational load on the brain, the visual attention mechanism is deployed such that regions with higher visual saliency are processed with higher priority than other less-salient regions. It is also known that visual saliency has a modulatory effect on JND thresholds. In this letter, we present a novel pixel-wise JND estimation method that considers the interplay between visual saliency and JND thresholds. In the proposed method, the largest JND thresholds of a given image are found such that the perceptual distance between the image and its JND noise-contaminated version is minimized in a perceptual space defined by the coefficients of the image in a normalized Laplacian pyramid. Experimental results indicate that the proposed method outperforms four of the latest JND models for static images.
Hadi Hadizadeh, Atiyeh Rajati, Ivan V. Bajic
IEEE Signal Process. Lett.3
2017 A Simulation Study of a Three-Dimensional Sound Field Reproduction System for Immersive Communication
abstract
Immersive communication systems promise a greatly improved user experience through the use of advanced technologies tailored to various human senses, such as sight, hearing, and touch. This paper focuses on a simulation study of three-dimensional (3-D) sound field reproduction (SFR) for immersive communication. At the transmitting end, the incident sound field is captured via a microphone array, active talkers are detected, and a clean version of the signals corresponding to active talkers are found by existing methods. The captured information is transmitted to the receiving end, where a 3-D sound field from virtual sources corresponding to active talkers is synthesized around listeners' heads by the proposed SFR method. In our system, the radiation patterns of the higher order (directive) loudspeakers are optimized by the constrained matching pursuit algorithm. Implementation of the directive loudspeaker patterns is not addressed here, and their deployment is assumed. Simulation results quantify the benefit of higher order loudspeakers for speech sound field synthesis in reverberant rooms.
Hanieh Khalilian, Ivan V. Bajic, Rodney G. Vaughan
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Online MoCap Data Coding With Bit Allocation, Rate Control, and Motion-Adaptive Post-Processing
abstract
With the advancements in methods for capturing 3D object motion, motion capture (MoCap) data are starting to be used beyond their traditional realm of animation and gaming in areas such as the arts, rehabilitation, automotive industry, remote interactions, and so on. As the amount of MoCap data increase, compression becomes crucial for further expansion and adoption of these technologies. In this paper, we extend our previous work on low-delay MoCap data compression by introducing two improvements. The first improvement is the bit allocation to long-term and short-term reference MoCap frames, which provides a 10–15% reduction in coded bit rate at the same quality. The second improvement is the post-processing in the form of motion-adaptive temporal low-pass filtering, which is able to provide another 9–13% savings in the bit rate. The experimental results also indicate that the proposed online MoCap codec is competitive with several state-of-the-art offline codecs. Overall, the proposed techniques integrate into a highly effective online MoCap codec that is suitable for low-delay applications, whose implementation is provided alongside this paper to aid further research in the field.
Choong-Hoon Kwak, Ivan V. Bajic
IEEE Trans. Multim.2
2016 Robust Domain-Filling Plumb-Line Lens Distortion Correction
abstract
A new robust plumb-line lens distortion correction algorithm is proposed, which uses a forward-transform model. The proposed method offers three principal advantages over conventional methods. First, our approach requires fewer training samples to estimate correction parameters, which is an important practical advantage for cameras that are already installed and in use. Second, our method shows a greater degree of resistance to image noise. Finally, once the correction parameters are estimated, the corrected image completely fills the image domain and does not require any post processing, which allows for lower run-time complexity compared to conventional plumb-line methods.
Mehdi P. Stapleton, Ivan V. Bajic
ISM2
2016 No-reference image quality assessment using statistical wavelet-packet features
Hadi Hadizadeh, Ivan V. Bajic
Pattern Recognit. Lett.2
2016 Color Gaussian Jet Features For No-Reference Quality Assessment of Multiply-Distorted Images
abstract
In this letter we present a novel no-reference image quality assessment (NR-IQA) method for the visual quality prediction of multiply-distorted color images. In the proposed method, to describe the image structure, a number of feature maps are first calculated based on the color Gaussian jet of the image. The popular local binary pattern operator is then applied on the computed feature maps to measure any potential structural degradations caused by multiple distortions. The performance of the proposed method was compared with 14 prominent full-reference IQA methods as well as 11 NR-IQA methods on two multidistortion IQA databases. The results indicate that the proposed method outperforms all the compared methods with a high accuracy at a moderate complexity.
Hadi Hadizadeh, Ivan V. Bajic
IEEE Signal Process. Lett.2
2016 Comparison of Loudspeaker Placement Methods for Sound Field Reproduction
abstract
This paper presents a comparison between several loudspeaker placement methods for sound field reproduction (SFR). The goal of these placement methods is to reduce the SFR error under a power constraint by selecting suitable locations for the loudspeakers. The first method is based on singular value decomposition of the acoustic transfer function (ATF) matrix. Depending on the configuration, an ideal ATF matrix is created and, then, approximated by selecting the appropriate locations for the loudspeakers. Another method is based on the constrained matching pursuit (CMP) algorithm, in which candidate locations of the loudspeakers are selected iteratively to minimize the approximation error of the desired sound field. The third method is based on sparsity-promoting sound field approximation using the least absolute shrinkage and selection operator. Loudspeaker placements obtained using these methods are compared against benchmark configuration of uniformly distributed loudspeakers. The comparison indicates that for constrained power, the CMP-based placement has the least reproduction error.
Hanieh Khalilian, Ivan V. Bajic, Rodney G. Vaughan
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 How many bits does it take for a stimulus to be salient?
abstract
Visual saliency has been shown to depend on the unpredictability of the visual stimulus given its surround. Various previous works have advocated the equivalence between stimulus saliency and uncompressibility. We propose a direct measure of this quantity, namely the number of bits required by an optimal video compressor to encode a given video patch, and show that features derived from this measure are highly predictive of eye fixations. To account for global saliency effects, these are embedded in a Markov random field model. The resulting saliency measure is shown to achieve state-of-the-art accuracy for the prediction of fixations, at a very low computational cost. Since most modern cameras incorporate video encoders, this paves the way for in-camera saliency estimation, which could be useful in a variety of computer vision applications.
Seyed Hossein Khatoonabadi, Nuno Vasconcelos, Ivan V. Bajic, Yufeng Shan
CVPR3
2015 Joint optimization of loudspeaker placement and radiation patterns for Sound Field Reproduction
abstract
A new method is presented for optimizing loudspeaker placement, radiation patterns, and excitations for Sound Field Reproduction (SFR). A power constraint is included which is to help control the sound increase in the regions away from listening volume. For a known primary source, the loudspeaker locations and patterns are jointly optimized by Constrained Matching Pursuit (CMP), and this can be undertaken offline, i.e., before system operation. The excitations of the designed loudspeakers are then determined by conventional Lagrangian optimization. Simulations for free space conditions show that the new method yields a lower reproduction error under a power constraint than other SFR methods.
Hanieh Khalilian, Ivan V. Bajic, Rodney G. Vaughan
ICASSP2
2015 Compressed-domain correlates of human fixations in dynamic scenes
Seyed Hossein Khatoonabadi, Ivan V. Bajic, Yufeng Shan
Multim. Tools Appl.2
2015 Constant Modulus Blind Adaptive Beamforming Based on Unscented Kalman Filtering
abstract
An unscented Kalman filter-based constant modulus adaptation algorithm (UKF-CMA) is proposed for blind uniform linear beamforming. The proposed algorithm is obtained by first developing a model of the constant modulus (CM) criterion and then fitting that model into the Kalman filter-style state space model by using an auxiliary parameter. The proposed algorithm does not require a priori information about the process noise and measurement noise covariance matrices and hence it can be applied readily. Simulation results demonstrate that the proposed algorithm offers improved performance compared to the recursive least square-based CM (RLS-CMA) and least-mean square-based CM (LMS-CMA) algorithms for adaptive blind beamforming.
M. Zulfiquar A. Bhotto, Ivan V. Bajic
IEEE Signal Process. Lett.2
2014 Hybrid compression of dynamic 3D mesh data
abstract
Geometric representation of objects and surfaces in terms of 3D meshes is becoming increasingly important in a variety of applications. As the amount of data in this format increases, the problem of compression becomes vital for the further development of the field. In this paper we present a codec for dynamic 3D mesh data that utilizes the “hybrid” framework from video coding, built upon temporal prediction and spatial transform. We discuss various features of the codec, including unrestricted quantization and two-stage entropy coding, and investigate its compression efficiency on a variety of test material. A discussion of various prediction structures and their impact on error resilience is also provided.
Choong-Hoon Kwak, Ivan V. Bajic
ICASSP2
2014 Low-saliency prior for disocclusion hole filling in DIBR-synthesized images
abstract
Although images as viewed from intermediate virtual viewpoints can be synthesized using texture and depth maps from nearby camera views via depth-image-based rendering (DIBR), the rendered images contain disocclusion holes - spatial regions that were not visible in the reference views due to foreground object occlusion - that requires proper filling. In this paper, we introduce a new signal prior into the hole filling problem formulation: given disocclusion holes are part of the background and background tends to have low visual saliency, the extrapolated signal into the holes must also be of low saliency. Mathematically, we add a low-saliency prior to an exemplar-based inpainting algorithm, so that the best-matched block has both small matching cost and is of low visual saliency. Moreover, we compute a suitable Lagrange multiplier value for the saliency cost term via analysis of the reference images. Experimental results show that using a low-saliency prior can improve performance by 0.5 dB over a previous hole filling scheme.
Bruno Macchiavello, Camilo C. Dorea, Edson M. Hung, Gene Cheung, Ivan V. Bajic
ICASSP5
2014 Comparison of visual saliency models for compressed video
abstract
Visual saliency modeling is an increasingly important research problem. While most saliency models for dynamic scenes operate on raw video, several models have also been developed for compressed video. This paper compares the accuracy of nine such models on a common eye-tracking dataset. The results indicate that a reasonably accurate saliency estimation is possible even using only motion vectors from the compressed bitstream. Successful strategies in compressed-domain saliency modeling are highlighted, and certain challenges are identified for future improvement.
Seyed Hossein Khatoonabadi, Ivan V. Bajic, Yufeng Shan
ICIP2
2014 Saliency-Aware Video Compression
abstract
In region-of-interest (ROI)-based video coding, ROI parts of the frame are encoded with higher quality than non-ROI parts. At low bit rates, such encoding may produce attention-grabbing coding artifacts, which may draw viewer's attention away from ROI, thereby degrading visual quality. In this paper, we present a saliency-aware video compression method for ROI-based video coding. The proposed method aims at reducing salient coding artifacts in non-ROI parts of the frame in order to keep user's attention on ROI. Further, the method allows saliency to increase in high quality parts of the frame, and allows saliency to reduce in non-ROI parts. Experimental results indicate that the proposed method is able to improve visual quality of encoded video relative to conventional rate distortion optimized video coding, as well as two state-of-the art perceptual video coding methods.
Hadi Hadizadeh, Ivan V. Bajic
IEEE Trans. Image Process.2
2013 Towards optimal loudspeaker placement for sound field reproduction
abstract
With a view to the suppression of unwanted sound, a planar array of loudspeakers is used to recreate a sound field in a nearby cubic listening area. Using free-space propagation, the formulation for selecting optimal locations of loudspeakers is developed so that numerical experiments can give a feel for the best possible suppression. First, to provide a benchmark, a target Acoustic Transfer Function (ATF) matrix is found that minimizes the reproduction error for a number of uniformly placed loudspeakers. Then the same number of loudspeakers is positioned by selecting from densely placed candidates so that their ATF matrix best approximates the target sound field. For recreating the field of a point source located in the direction of the array, the loudspeaker array with selected locations is shown to improve the reproduction accuracy over a reasonable bandwidth.
Hanieh Khalilian, Ivan V. Bajic, Rodney G. Vaughan
ICASSP2
2013 MoCap data coding with unrestricted quantization and rate control
abstract
Motion Capture (MoCap) technology is becoming increasingly popular in gaming, entertainment, and multimedia industries. Interactive systems usingMoCap technology require low-delay MoCap data compression. In this paper, we extend previous work on low-delay MoCap compression by introducing several useful features, such as unrestricted quantization, more efficient entropy coding, as well as encoder rate control. Experimental results show that the proposed rate control provides better than 99% accuracy in controlling encoder's output bitrate. At the same time, improvements in quantization and entropy coding provide over 20% reduction in bit rate for the same reconstruction quality, compared to the current state of the art.
Choong-Hoon Kwak, Ivan V. Bajic
ICASSP2
2013 QP initialization and interview MAD prediction for rate control in HEVC-based multi-view video coding
abstract
Rate control is an important component of an end-to-end video communication system. Currently, there are several proposals for rate control in the upcoming High Efficiency Video Coding (HEVC) standard, but none specifically for the multi-view extension of the standard. In this paper, we apply one of the HEVC single-view rate control schemes to the multi-view scenario, and propose two improvements. One improvement deals with Quantization Parameter (QP) initialization, and the other deals with interview Mean Absolute Difference (MAD) prediction. Experimental results demonstrate increased accuracy of rate control, reduced fluctuation of instantaneous bitrate, as well as a reduction in PSNR degradation compared to the existing rate control algorithm.
Woong Lim, Ivan V. Bajic, Dong-Gyu Sim
ICASSP2
2013 3D motion in visual saliency modeling
abstract
Visual saliency is a probabilistic estimate of how likely a given spatial area in an image or video is to attract human visual attention relative to other areas. Bottom-up saliency models aggregate low-level image features like luminance and color contrast, flicker, 2D motion, etc. to construct a plausible saliency map. In this paper, we introduce 3D motion (object movements towards or away from the observer) into bottom-up video saliency modeling. Given availability of per-pixel depth maps, we first propose a novel algorithm to estimate 3D motion vectors (3DMVs) for arbitrarily shaped sub-blocks in texture-plus-depth videos. We then derive two feature channels from 3DMVs to be incorporated into a widely accepted bottom-up saliency model. Experiments on subjective quality of Region-of-Interest (ROI) based video coding show that our enriched saliency model with 3DMV channels is more accurate in estimating human visual attention.
Pengfei Wan 0001, Yunlong Feng, Gene Cheung, Ivan V. Bajic, Oscar C. Au, Yusheng Ji
ICASSP4
2013 Guiding visual attention by manipulating orientation in images
abstract
Visual attention plays an important role in directing our gaze to potentially interesting areas in images. Our attention is involuntarily drawn to areas that are perceptually different from their immediate surroundings. Such areas are labeled “salient.” They originate from variations in principal visual features such as color, intensity, and orientation. In this study, we analyze how manipulating the orientation of a particular region of an image affects human visual attention. Statistical Hough transform is applied on a selected region in an image to construct the edge distribution of that region over a range of orientations. The remainder of the image is analyzed using a weighted statistical Hough transform to obtain the edge distribution in the region's surroundings. We measure the dissimilarity between these two distributions as the region is rotated and show that the region becomes more salient as the dissimilarity is increased. This model also allows us to predict the angle of rotation at which the selected region becomes most salient, which enables us to manipulate the image so that the selected region's saliency is maximized. We apply our method to a set of natural images and verify its effectiveness through eye-tracking.
Victor A. Mateescu, Ivan V. Bajic
ICME2
2013 Frame rate up-conversion using global and local higher-order motion
abstract
A region-based frame rate up-conversion (FRUC) algorithm based on higher-order global and local motion is proposed in this paper. First, perspective global motion parameters are estimated so that backgrounds of neighboring frames can be aligned. Then, the foreground is separated from the background using structural similarity, and iteratively decomposed into regions with homogeneous motion that fit a local affine or perspective model. Also, regions entering or leaving the scene are identified. Finally, an intermediate frame is synthesized by warping the background and foreground regions towards the desired frame location using their associated motion parameters. Experimental results show that the proposed algorithm outperforms several recent FRUC methods in both PSNR and SSIM, and produces more aesthetically pleasing frames, free of blocking artifacts.
Chun Qian, Ivan V. Bajic
ICME2
2013 3-D Motion Estimation for Visual Saliency Modeling
abstract
Visual saliency is a probabilistic estimate of how likely a spatial area in an image or video frame is to attract human visual attention relative to other areas. When existing bottom-up saliency models aggregate low-level features to construct a plausible saliency map, only 2-D motion cues are used as motion features, even though videos typically capture dynamic 3-D scenes. In this paper, we introduce 3-D motion into bottom-up saliency modeling for texture-plus-depth videos. We first propose an efficient 3-D motion estimation algorithm, which computes a 3-D motion vector (3DMV) for each sub-block in the frame. Using the computed 3DMVs, we then derive several saliency channels (called 3DMV channels), which are incorporated into a bottom-up saliency model to obtain enhanced saliency maps. Experiments tracking human gaze show that incorporating our 3DMV channels into bottom-up saliency model significantly improves the accuracy of derived saliency maps.
Pengfei Wan 0001, Yunlong Feng, Gene Cheung, Ivan V. Bajic, Oscar C. Au
IEEE Signal Process. Lett.4
2013 Video Watermarking With Empirical PCA-Based Decoding
abstract
A new method for video watermarking is presented in this paper. In the proposed method, data are embedded in the LL subband of wavelet coefficients, and decoding is performed based on the comparison among the elements of the first principal component resulting from empirical principal component analysis (PCA). The locations for data embedding are selected such that they offer the most robust PCA-based decoding. Data are inserted in the LL subband in an adaptive manner based on the energy of high frequency subbands and visual saliency. Extensive testing was performed under various types of attacks, such as spatial attacks (uniform and Gaussian noise and median filtering), compression attacks (MPEG-2, H. 263, and H. 264), and temporal attacks (frame repetition, frame averaging, frame swapping, and frame rate conversion). The results show that the proposed method offers improved performance compared with several methods from the literature, especially under additive noise and compression attacks.
Hanieh Khalilian, Ivan V. Bajic
IEEE Trans. Image Process.2
2013 Video Object Tracking in the Compressed Domain Using Spatio-Temporal Markov Random Fields
abstract
Despite the recent progress in both pixel-domain and compressed-domain video object tracking, the need for a tracking framework with both reasonable accuracy and reasonable complexity still exists. This paper presents a method for tracking moving objects in H.264/AVC-compressed video sequences using a spatio-temporal Markov random field (ST-MRF) model. An ST-MRF model naturally integrates the spatial and temporal aspects of the object's motion. Built upon such a model, the proposed method works in the compressed domain and uses only the motion vectors (MVs) and block coding modes from the compressed bitstream to perform tracking. First, the MVs are preprocessed through intracoded block motion approximation and global motion compensation. At each frame, the decision of whether a particular block belongs to the object being tracked is made with the help of the ST-MRF model, which is updated from frame to frame in order to follow the changes in the object's motion. The proposed method is tested on a number of standard sequences, and the results demonstrate its advantages over some of the recent state-of-the-art methods.
Seyed Hossein Khatoonabadi, Ivan V. Bajic
IEEE Trans. Image Process.2
2013 Video Error Concealment Using a Computation-Efficient Low Saliency Prior
abstract
Error concealment in packet-loss-corrupted streaming video is inherently an under-determined problem, as there are insufficient number of well-defined criteria to recover the missing blocks perfectly. When a Region-of-Interest (ROI) based unequal error protection (UEP) scheme is deployed during video streaming-i.e., more visually salient regions are strongly protected-a lost block is likely to be of low saliency in the original frame. In this paper, we propose to add a low-saliency prior to the error concealment problem as a regularization term. It serves two purposes. First, in ROI-based UEP video streaming, low-saliency prior provides the correct side information for the client to identify the correct replacement blocks for concealment. Second, in the event that a perfectly matched block cannot be unambiguously identified, the low-saliency prior reduces viewer's visual attention on the loss-stricken region, resulting in higher overall subjective quality. We study the effectiveness of a low-saliency prior in the context of a previously proposed RECAP error concealment system. RECAP transmits a low-resolution (LR) version of an image alongside the original high-resolution (HR) version, so that if blocks in the HR version are lost, the correctly-received LR version can serve as a template for matching of suitable replacement blocks from a previously correctly-decoded HR frame. We add a low-saliency prior to the block identification process, so that only replacement candidate blocks with good match and low saliency can be selected. Further, we develop a low-complexity convex approximation to the well known Itti-Koch-Niebur saliency model, which enables the low-saliency error concealment problem to be solved efficiently. Experimental results show that: i) PSNR of the error-concealed frames can be increased dramatically (up to 3.6 dB over the original RECAP), showing the effectiveness of a low-saliency prior in the under-determined error concealment problem; and ii) subjective quality of the repaired video using our proposal, as confirmed by an extensive user study, is better than the original RECAP.
Hadi Hadizadeh, Ivan V. Bajic, Gene Cheung
IEEE Trans. Multim.2
2012 Saliency-Cognizant Error Concealment in Loss-Corrupted Streaming Video
abstract
Error concealment in packet-loss-corrupted streaming video is inherently an under-determined problem, as there are insufficient number of well-defined criteria to recover the missing blocks perfectly. When a Region-of-Interest (ROI) based unequal error protection (UEP) scheme is deployed during video streaming -- i.e., more visually salient regions are strongly protected -- a %(e.g., using strong Forward Error Correction (FEC) codes) -- a lost block is likely to be of low saliency in the original frame. In this paper, we propose to add a low-saliency prior to the error concealment problem as a regularization term. It serves two purposes. First, in ROI-based UEP video streaming, low-saliency prior provides the right side information for the client to identify the correct replacement blocks for concealment. Second, in the event that a perfectly matched block cannot be unambiguously identified, the low-saliency prior reduces viewer's visual attention on the loss-stricken region, resulting in higher overall subjective quality. We study the effectiveness of a low-saliency prior in the context of a previously proposed RECAP[1] error concealment system. RECAP transmits a low-resolution (LR) version of an image alongside the original high-resolution (HR) version, so that if blocks in the HR version are lost, the correctly-received LR version can serve as a template for matching of suitable replacement blocks from a previously correctly-decoded HR frame. We add a low-saliency prior to the block identification process, so that only replacement candidate blocks with good match and low saliency can be selected. Further, we design and apply four saliency reduction operators iteratively in a loop, in order to reduce the saliency of candidate blocks. Experimental results show that: i) PSNR of the error-concealed frames can be increased dramatically (up to $3.2$dB over the original RECAP), showing the effectiveness of a low-saliency prior in the under-determined error concealment problem, and ii) subjective quality of the repaired video using our proposal, as confirmed by an extensive user study, is better than the original RECAP.
Hadi Hadizadeh, Ivan V. Bajic, Gene Cheung
ICME2
2012 Eye-Tracking Database for a Set of Standard Video Sequences
abstract
This correspondence describes a publicly available database of eye-tracking data, collected on a set of standard video sequences that are frequently used in video compression, processing, and transmission simulations. A unique feature of this database is that it contains eye-tracking data for both the first and second viewings of the sequence. We have made available the uncompressed video sequences and the raw eye-tracking data for each sequence, along with different visualizations of the data and a preliminary analysis based on two well-known visual attention models.
Hadi Hadizadeh, Mario J. Enriquez, Ivan V. Bajic
IEEE Trans. Image Process.3
2011 Good-looking green images
abstract
In this paper we present a novel perceptually-based algorithm for color quantization that produces images that consume less energy than conventionally quantized images when displayed on modern energy-adaptive displays. To evaluate the performance of the proposed algorithm, we performed a subjective study on a standard Kodak color image database. Experimental results indicate that the proposed algorithm is able to reduce the energy consumption by 4.25% on average, while achieving the same or better subjective image quality as conventional color quantization.
Hadi Hadizadeh, Ivan V. Bajic, Parvaneh Saeedi, Scott Daly
ICIP2
2011 Saliency-preserving video compression
abstract
In region-of-interest (ROI) video coding, the part of the frame designated as ROI is encoded with higher quality relative to the rest of the frame. At low bit rates, coding artifacts in non-ROI parts of the frame may become salient and draw user's attention away from ROI, thereby degrading visual quality. In this paper we propose a saliency-preserving framework for ROI video coding. This approach aims at reducing attention-grabbing visual artifacts in non-ROI parts of the frame in order to keep user's attention on ROI. Experimental results indicate that the proposed method is able to improve the visual quality of ROI video at low bit rates.
Hadi Hadizadeh, Ivan V. Bajic
ICME2
2011 Hybrid low-delay compression of motion capture data
abstract
Motion Capture (MoCap) is becoming important in many areas of technology, science, and art, including graphics, visualization, gaming, and medical applications. In parallel with its increased use and abundance, compression of this kind of data is becoming more important. In this paper, we propose a hybrid low-delay compression scheme for MoCap data that is particularly suitable for interactive applications such as online gaming or telemedicine. Experimental results confirm the superiority of the proposed approach against state-of-the-art methods for MoCap compression in both compression efficiency and delay, making it suitable for interactive applications.
Choong-Hoon Kwak, Ivan V. Bajic
ICME2
2011 Error concealment strategies for Motion Capture data streaming
abstract
Motion Capture (MoCap) is becoming important in many areas of technology, science, and art, including graphics, visualization, gaming, and medical applications. MoCap data streaming faces many challenges found in other kinds of media communications, an important one being the degradation of decoded data quality due to packet loss. In this paper, we propose several strategies for error concealment of MoCap data based on the principles of mechanics. In particular, forward, backward and bi-directional error concealment strategies are presented. Experimental results verify the effectiveness of the proposed methods and demonstrate the advantage of bi-directional error concealment over forward and backward approaches.
Choong-Hoon Kwak, Ivan V. Bajic
ICME2
2011 A Testbed and Methodology for Comparing Live Video Frame Rate Control Methods
abstract
Most existing methods for video frame rate control were developed for prerecorded video. With the proliferation of mobile audiovisual communication devices, it is becoming important to develop frame rate control methods for live (i.e., not prerecorded) video as well. In this letter, we present a testbed and the corresponding methodology suitable for comparing such methods. We illustrate the methodology and the use of the testbed through comparison of an existing state-of-the-art frame rate control method against a simple new control method developed specifically for live video.
Ivan V. Bajic, Xiaonan Ma
IEEE Signal Process. Lett.1
2011 A Joint Approach to Global Motion Estimation and Motion Segmentation From a Coarsely Sampled Motion Vector Field
abstract
In many content-based video processing systems, the presence of moving objects limits the accuracy of global motion estimation (GME). On the other hand, the inaccuracy of global motion parameter estimates affects the performance of motion segmentation. In this paper, we introduce a procedure for simultaneous object segmentation and GME from a coarsely sampled (i.e., block-based) motion vector (MV) field. The procedure starts with removing MV outliers from the MV field, and then performs GME to obtain an estimate of global motion parameters. Using these estimates, global motion is removed from the MV field, and moving region segmentation is performed on this compensated MV field. MVs in the moving regions are treated as outliers in the context of GME in the next round of processing. Iterating between GME and motion segmentation helps improve both GME and segmentation accuracy. Experimental results demonstrate the advantage of the proposed approach over state-of-the-art methods on both synthetic motion fields and MVs from real video sequences.
Yue-Meng Chen, Ivan V. Bajic
IEEE Trans. Circuits Syst. Video Technol.2
2011 Rate-Distortion Optimized Pixel-Based Motion Vector Concatenation for Reference Picture Selection
abstract
Reference picture selection (RPS) is a powerful error-control technique for video streaming. Previously, two fast block-based motion vector concatenation (MVC) algorithms were proposed for video transcoding based on forward dominant vector selection (FDVS) and activity dominant vector selection (ADVS). In this paper, we cast these algorithms in the rate-distortion (RD) optimization framework. We also present two novel RD optimized pixel-based MVC schemes for RPS. Experimental results indicate that the proposed methods provide higher video quality compared to both FDVS and ADVS. In addition, we study the complexity of various RPS algorithms as a function of the loss rate and round trip time.
Hadi Hadizadeh, Ivan V. Bajic
IEEE Trans. Circuits Syst. Video Technol.2
2011 Burst-Loss-Resilient Packetization of Video
abstract
In video transmission over packet-based networks, packet losses often occur in bursts. In this paper, we present a novel packetization method for increasing the robustness of compressed video against bursty packet losses. The proposed method is based on creating a coding order of macroblocks (MBs) so that the blocks that are close to each other in the coding order end up being far from each other in the frame. We formulate this idea as a discrete optimization problem, prove its NP-hardness, and discuss several possible solution methods. Experimental results indicate that the proposed method improves the quality of reconstructed frames under burst loss by several decibels compared to conventional flexible MB ordering techniques, and about 0.7 dB compared to the state-of-the-art method called explicit chessboard wipe.
Hadi Hadizadeh, Ivan V. Bajic
IEEE Trans. Image Process.2
2011 Moving Region Segmentation From Compressed Video Using Global Motion Estimation and Markov Random Fields
abstract
In this paper, we propose an unsupervised segmentation algorithm for extracting moving regions from compressed video using global motion estimation (GME) and Markov random field (MRF) classification. First, motion vectors (MVs) are compensated from global motion and quantized into several representative classes, from which MRF priors are estimated. Then, a coarse segmentation map of the MV field is obtained using a maximum a posteriori estimate of the MRF label process. Finally, the boundaries of segmented moving regions are refined using color and edge information. The algorithm has been validated on a number of test sequences, and experimental results are provided to demonstrate its advantages over state-of-the-art methods.
Yue-Meng Chen, Ivan V. Bajic, Parvaneh Saeedi
IEEE Trans. Multim.2
2010 NAL-SIM: An Interactive Simulator for H.264/AVC Video Coding and Transmission
abstract
In this paper we present a high-level graphical simulator for video coding and transmission using the H.264/AVC standard. The main objective of the developed simulator is to build an overall video communication system model including source encoding, channel modeling and decoding in order to interactively investigate the performance of various H.264/AVC coding schemes in the face of bandwidth constraints and channel errors. The developed simulator can be employed as a useful research or educational tool for video communication systems based on the H.264/AVC video coding standard.
Hadi Hadizadeh, Ivan V. Bajic
CCNC2
2010 Predictive Video Decoding Based on Ordinal Depth of Moving Regions
abstract
Predictive video decoding is a technique for reducing end-to-end delay and/or concealing frame loss in video communications. In this paper, we propose an improved predictive decoding technique that incorporates the estimation of the ordinal depth of moving regions, and uses it to compose the predicted frame. Ordinal depth information enables occlusion prediction in the future frame, allowing the predictor to synthesize a more realistic frame. Experimental results show that the proposed method achieves better subjective and objective quality of synthesized frames compared to a state-of-the-art block-based frame prediction approach.
Yue-Meng Chen, Ivan V. Bajic
ICC2
2010 Burst Loss Resilient Packetization of Video
abstract
In video transmission over packet-based networks, packet losses usually occur in bursts. In this paper, we present a novel packetization method for increasing the robustness of compressed video against bursty packet losses. The proposed method is based on creating a coding order of macroblocks so that the blocks that are close to each other in the coding order end up being far from each other in the frame. Experimental results indicate that the proposed method improves the quality of reconstructed frames under burst loss by several dB compared to conventional FMO techniques, and about 0.7 dB compared to the state-of-the-art method called Explicit Chessboard-Wipe (ECW).
Hadi Hadizadeh, Ivan V. Bajic
ICC2
2010 Motion segmentation in compressed video using Markov Random Fields
abstract
In this paper, we propose an unsupervised segmentation algorithm for extracting moving objects/regions from compressed video using Markov Random Field (MRF) classification. First, motion vectors (MVs) are quantized into several representative classes, from which MRF priors are estimated. Then, a coarse segmentation map of the MV field is obtained using a maximum a posteriori estimate of the MRF label process. Finally, the boundaries of segmented moving regions are refined using color and edge information. The algorithm has been validated on a number of test sequences, and experimental results are provided to demonstrate its superiority over state-of-the-art methods.
Yue-Meng Chen, Ivan V. Bajic, Parvaneh Saeedi
ICME2
2010 Pixel-based motion vector concatenation for Reference Picture Selection
abstract
Reference Picture Selection (RPS) is a powerful error control technique for video streaming. Previously, two fast block-based motion vector concatenation (MVC) algorithms were proposed for video transcoding based on forward dominant vector selection (FDVS) and activity dominant vector selection (ADVS). In this paper, we examine the use of these algorithms in RPS, in the context of video transmission. We also present a novel pixel-based MVC scheme for RPS. Experimental results indicate that the proposed method provides higher video quality compared to both the FDVS and ADVS. In addition, we study the complexity of various RPS algorithms as a function of the loss rate and round trip time.
Hadi Hadizadeh, Ivan V. Bajic
ICME2
2010 Motion Vector Outlier Rejection Cascade for Global Motion Estimation
abstract
Global motion estimation (GME) from motion vector (MV) field in compressed domain greatly reduces the complexity of conventional pixel-based GME. However, outlier MVs, caused by noise or foreground objects, may reduce the accuracy of MV-based GME. In this paper, we propose a cascade-of-rejectors approach for removing MV outliers to achieve efficient and accurate GME. Experimental results show that the proposed MV outlier rejection cascade significantly lowers the complexity MV-based GME, with an accuracy close to or better than state-of-the-art methods.
Yue-Meng Chen, Ivan V. Bajic
IEEE Signal Process. Lett.2
2010 Region-Based Predictive Decoding of Video
abstract
This letter presents a region-based predictive decoding technique for end-to-end delay reduction in video communications. Video frames are predicted from past video data and displayed before they arrive at the decoder, thereby reducing the end-to-end delay. This problem is similar to whole-frame concealment, so we compare our method to a state-of-the-art whole-frame concealment algorithm. As demonstrated in the letter, our method yields better visual results with peak signal-to-noise ratio gains up to 2.5 dB on sequences with complex and high-intensity motion. These gains are mainly due to motion segmentation and region-based motion prediction.
Yue-Meng Chen, Ivan V. Bajic
IEEE Trans. Circuits Syst. Video Technol.2
2010 Joint Decoding of Unequally Protected JPEG2000 Bitstreams and Reed-Solomon Codes
abstract
In this paper we present joint decoding of JPEG2000 bitstreams and Reed-Solomon codes in the context of unequal loss protection. Using error resilience features of JPEG2000 bitstreams, the joint decoder helps to restore the erased symbols when the Reed-Solomon decoder fails to retrieve them on its own. However, the joint decoding process might become time-consuming due to a search through the set of possible erased symbols. We propose the use of smaller codeblocks and transmission of a relatively small amount of side information with high reliability as two approaches to accelerate the joint decoding process. The accelerated joint decoder can deliver essentially the same quality enhancement as the nonaccelerated one, while operating several times faster.
Sohail Bahmani, Ivan V. Bajic, Atousa Hajshirmohammadi
IEEE Trans. Image Process.2
2009 Subset Selection in Type-II Hybrid ARQ/FEC for Video Multicast
abstract
This paper proposes an error control scheme that minimizes the total distortion experienced by the receivers using a new version of Type-II hybrid ARQ/FEC. Based on the feedback information about the losses in the previous Group of Pictures (GOP), the server sends parity packets for a subset of the frames from that GOP with the aim of minimizing the total distortion experienced by the receivers. The subset selection problem is NP-hard, so we propose a suboptimal method to solve it based on simulated annealing. Experimental results for the case when a single parity packet is used per group of 16 packets show that the proposed subset selection improves the plain Type-II hybrid ARQ/FEC by over 4 dB in decoded video PSNR, and achieves a 1-1.5 dB gain compared to a state-of-the-art error control method based on rate-distortion optimized frame retransmission.
S. Mohsen Amiri, Ivan V. Bajic
ICC2
2009 Improved Joint Source-Channel Decoding of JPEG2000 Images and Reed-Solomon Codes
abstract
In this paper we present improvements to the recently-proposed joint decoding of JPEG2000 bitstreams and Reed-Solomon codes in the context of unequal loss protection. Using error resilience features of JPEG2000 bitstreams, the joint decoder helps to restore the erased symbols when the Reed-Solomon decoder fails to retrieve them on its own. We make use of the ability of the JPEG2000 decoder to provide rough error localization within a coding pass to speed up the search for erased symbol values. In addition, we show how transmitting a relatively small amount of side information with high reliability may help the joint decoder by reducing the size of the search space and bypassing some of the JPEG2000 decoding iterations needed to verify the correctness of the restored source information. The improved joint decoder is up to 20 times faster compared to the previous one.
Sohail Bahmani, Ivan V. Bajic, Atousa Hajshirmohammadi
ICC2
2009 Scalable Video Streaming With Fine-Grain Adaptive Forward Error Correction
abstract
In this paper, we investigate a fine-grain adaptive forward error correction (FGA-FEC) coding scheme for scalable video bitstreams. In our work, both the embedded source bitstream and the error-control codes are granularly adapted at block level in intermediate overlay nodes to satisfy heterogeneous users with both different video frame-rate/spatial resolution/quality preferences and different network connections. The proposed FGA-FEC scheme encodes and adapts the embedded source-coded bitstream in such a way that if part of the video source data is actively dropped, parity bits protecting that piece of data are also removed, yielding an efficient result without any transcoding.
Yufeng Shan, Ivan V. Bajic, John W. Woods, Shivkumar Kalyanaraman
IEEE Trans. Circuits Syst. Video Technol.2
2008 Joint source-chanel decoding of JPEG2000 images with unequal loss protection
abstract
This paper presents a method for joint decoding of JPEG2000 bistreams and Reed-Solomon codes in the context of unequal loss protection. When the Reed-Solomon decoder is unable to retrieve the erased source symbols, the proposed joint decoder searches through the set of possible erased source symbols, making use of error resilience features of JPEG2000 to retrieve correct symbols. The joint decoder can be used as an add-on module to some of the existing schemes for unequal loss protection, and can improve the PSNR of decoded images by over 10 dB in some cases.
Sohail Bahmani, Ivan V. Bajic, Atousa Hajshirmohammadi
ICASSP2
2008 A Novel Noncausal Whole-Frame Concealment Algorithm for Video Streaming
abstract
Error concealment is very important for video communication as an application-layer error control mechanism which can be used independently of the underlying communication infrastructure. In this paper, we propose a noncausal whole-frame concealment algorithm to conceal lost frames in the video sequence. The proposed algorithm consist of two parts. First, the preliminary version of the lost frame is generated using motion vectors from three different sources: from previous frames, upcoming frames, and, in case of RPS-NACK, the reset frame as well. Then, to fill the remaining empty areas, the algorithm uses spatiotemporal boundary matching. Experimental results show that the complexity of the proposed algorithm is low enough to run in real time, while achieving higher PSNR and providing better visual quality than the competing state-of-the-art method.
S. Mohsen Amiri, Ivan V. Bajic
ISM2
2007 Predictive Decoding for Delay Reduction in Video Communications
abstract
Low delay is critically important for interactive video communication. This paper presents several predictive decoding techniques for delay reduction. Video frames are predicted from past video data, and displayed before they arrive at the decoder. This enables the user to choose the proper trade-off between quality and delay. In this way, it is possible to reduce the perceived end-to-end communication delay by about 100 ms while maintaining reasonable video quality.
Yue-Meng Chen, Ivan V. Bajic
GLOBECOM2
2007 Error Concealment for Scalable Motion-Compensated Subband/Wavelet Video Coders
abstract
In this paper, we present two error-concealment algorithms developed for scalable motion-compensated subband/wavelet video coders. These algorithms exploit the properties of motion-compensated temporal filtering to recover lost video data by motion compensation from correctly received previous and future video frames. Our experiments indicate that backward motion-compensated prediction outperforms replacement from neighboring correctly received frames by up to 3 dB in terms of PSNR. In addition, a bidirectional algorithm tops the unidirectional one by up to 1 dB. Also, visual improvements are often higher that PSNR improvements would suggest.
Ivan V. Bajic, John W. Woods
IEEE Trans. Circuits Syst. Video Technol.1
2006 Non-causal error control for wireless video streaming with noncoherent signaling
abstract
The existence of large receiver-side buffers in typical video streaming applications allows us to implement non-causal algorithms for video processing at the receiver. We have recently proposed an error control scheme that takes advantage of this fact, and uses backward error concealment to recover lost video data. It was demonstrated that such a scheme can be several times more energy efficient than ARQ over a wireless channel modeled in the capacity-outage framework. In this paper we examine the behavior of non-causal error control (NCEC) with noncoherent signaling over a Rayleigh fading channel. We demonstrate that NCEC remains to be superior to ARQ even in this more realistic setting.
Ivan V. Bajic
ISCAS1
2006 Efficient Error Control for Wireless Video Multicast
abstract
In this paper we develop joint source rate selection, power management, and error control schemes for wireless multicast of SNR scalable video. In particular, two error control schemes are analyzed and compared-type 2 hybrid ARQ/FEC (T2HA/F) and non-causal error control (NCEC). T2HA/F is based on incremental redundancy chosen to recover losses and provide the required level of video quality for all users. On the other hand, NCEC relies on error concealment from future video data whose quality is adaptively adjusted to bring the video quality up to the desired level for all users. Experimental results indicate that NCEC can be significantly more efficient than T2HA/F
Ivan V. Bajic
MMSP1
2006 Adaptive MAP error concealment for dispersively packetized wavelet-coded images
abstract
In this paper, we present an adaptive maximum a posteriori (MAP) error concealment algorithm for dispersively packetized wavelet-coded images. We model the subbands of a wavelet-coded image as Markov random fields, and use the edge characteristics in a particular subband, and regularity properties of subband/wavelet samples across scales, to adapt the potential functions locally. The resulting adaptive MAP estimation gives PSNR advantages of up to 0.7 dB compared to the competing algorithms. The advantage is most evident near the edges, which helps improve the visual quality of the reconstructed images.
Ivan V. Bajic
IEEE Trans. Image Process.1
2006 Noncausal Error Control for Video Streaming Over Wireless Packet Networks
abstract
Video streaming systems usually employ a reasonably large receiver buffer to deal with jitter. The existence of such a buffer allows us to implement noncausal algorithms for video processing at the receiver. In this paper, we describe an error control scheme that takes advantage of this fact, and employs a noncausal error-concealment algorithm. We use the knowledge of the packet-loss realization for the previously transmitted data to adjust the transmission strategy for the future data, knowing that this future data will be used to conceal the data in the previously lost packets. We demonstrate that such a strategy can be several times more energy efficient than ARQ.
Ivan V. Bajic
IEEE Trans. Multim.1
2005 Joint source-network error control coding for scalable overlay video streaming
abstract
In this paper, we propose a joint source-network error control coding (JSNC) scheme which efficiently integrates scalable video coding, error control coding and overlay infrastructure to stream video to heterogeneous users. The distributed overlay nodes adapt both the video bitstream and error control coding based on both user requirements and available bandwidth. A novel fine granular adaptive FEC (FGA-FECcheme, a generalization of MD-FEC, is proposed for error recovery during video transmission to heterogeneous users. Encoding once, the FGA-FEC can satisfy multiple heterogeneous users simultaneously without decoding/re-encoding FEC at intermediate nodes.
Yufeng Shan, Shivkumar Kalyanaraman, John W. Woods, Ivan V. Bajic
ICIP (1)4
2005 Performance analysis of the efficacy of packet-level FEC in improving video transport over networks
abstract
Packet video transport over networks is expected to experience packet losses due to congestion, link failures and timeouts. Packet-level FEC is often proposed to combat these packet losses and thus improve received video quality. However, the redundant packets associated with the FEC coding will increase the effective network load and, as a result, further exacerbate the network packet-loss rates. In this work, we propose a model framework to describe FEC-protected packet video network transport systems and analytically investigate the overall efficacy of packet-level FEC in improving the end-to-end video quality. We consider a simplified network scenario, where the network congestion performance can be described in terms of a single bottleneck node, modeled as a multiplexer. Our results show that packet-level FEC can significantly improve the end-to-end video quality provided that the coding block size is large enough and the coding rate is appropriately chosen, assuming other system parameters are likewise appropriately chosen. We also demonstrate that, when multiple sources share the multiplexer, FEC coding can achieve significant multiplexing gain in terms of the number of video sources that can be supported for a given end-to-end performance.
Xunqi Yu, James W. Modestino, Ivan V. Bajic
ICIP (2)3
2005 Overlay multi-hop FEC scheme for video streaming
Yufeng Shan, Ivan V. Bajic, Shivkumar Kalyanaraman, John W. Woods
Signal Process. Image Commun.2
2004 On the effects of path correlation in multi-path video communications using FEC over lossy packet networks
abstract
In this paper, we investigate the effect of path correlation on video communications when path-diversity routing techniques are employed together with forward error correction (FEC). We study the statistical properties of packet losses when correlation between paths exists and demonstrate that the effects of path correlation on received video quality depend on various factors, such as the effectiveness of error control strategies (passive error concealment, FEC, etc.) and the existing channel conditions. We demonstrate that, contrary to common belief, path correlation need not be detrimental to the system performance. The results are expected to have some impact on the development of explicit multi-path routing strategies.
Qi Qu, Ivan V. Bajic, Xusheng Tian, James W. Modestino
GLOBECOM2
2004 Optimal subsampling of circularly bandlimited images
abstract
In this paper, we consider the problem of optimal subsampling of circularly bandlimited images. We show that the two facets of this problem can be formulated as constrained sphere packings, and we describe the algorithm that may be used to solve them. We also compare the resulting subsampling patterns to the more conventional rectangular subsampling and illustrate the potential savings in subsampling density.
Ivan V. Bajic
ICASSP (3)1
2004 Overlay multi-hop fec scheme for video streaming over peer-to-peer networks
abstract
Overlay networks offer promising capabilities for video streaming, due to their support for application-layer processing at the overlay forwarding nodes. In this paper we propose a novel overlay multi-hop FEC (OM-FEC) scheme that provides FEC encoding/decoding capabilities at intermediate nodes in an overlay path. Based on the current network conditions, the end-to-end overlay path is partitioned into segments, and appropriate FEC codes are applied over those segments. We evaluate our work in a real-world scenario and illustrate that the proposed OM-FEC can outperform a pure end-to-end strategy by 10-15 dB in terms of video PSNR.
Yufeng Shan, Ivan V. Bajic, Shivkumar Kalyanaraman, John W. Woods
ICIP2
2003 Integrated end-to-end buffer management and congestion control for scalable video communications
abstract
In this paper we present a video communication system that integrates end-to-end buffer management and congestion control at the source with the playout adjustment mechanism at the receiver. While each component of the system has been considered independently in the literature, our focus in this work is their integration. The proposed system exploits the fact that when congestion control is implemented at the source, most of the loss occurs at the source and not within the network. Based on this observation, we design the buffer management to trade off random loss for controlled loss of visually less important data. Frame rate is adjusted at the receiver to maximize the visual quality of the displayed video based on the overall loss. We tested our system with both H.26L and a subband/wavelet video coder, and found that it significantly improves the received video quality in both cases.
Ivan V. Bajic, Omesh Tickoo, Anand Balan, Shivkumar Kalyanaraman, John W. Woods
ICIP (3)1
2003 EZBC video streaming with channel coding and error concealment
Ivan V. Bajic, John W. Woods
VCIP1
2003 Domain-based multiple description coding of images and video
abstract
In this paper, we present a method of creating domain-based multiple descriptions of images and video. These descriptions are created by partitioning the transform domain of the signal into sets whose points are maximally separated from each other. This property enables simple error concealment methods to produce good estimates of lost signal samples. We present the approach in the context of Internet transmission of subband/wavelet-coded images and scalable motion compensated three-dimensional (3D) subband/wavelet-coded video, but applications are not limited to these scenarios. The results indicate that the proposed methods offer improvements over similar competing methods by up to 1 dB for images, and several decibels for video. Visual quality is also improved.
Ivan V. Bajic, John W. Woods
IEEE Trans. Image Process.1
2003 Maximum minimal distance partitioning of the Z2 lattice
abstract
We study the problem of dividing the /spl Zopf//sup 2/ lattice into partitions so that minimal intra-partition distance between the points is maximized. We show that this problem is analogous to the problem of sphere packing. An upper bound on the achievable intra-partition distances for a given number of partitions follows naturally from this observation, since the optimal sphere packing in two dimensions is achieved by the hexagonal lattice. Specific instances of this problem, when the number of partitions is 2/sup m/, were treated in trellis-coded modulation (TCM) code design by Ungerboeck (1982) and others. It is seen that methods previously used for set partitioning in TCM code design are asymptotically suboptimal as the number of partitions increases. We propose an algorithm for solving the /spl Zopf//sup 2/ lattice partitioning problem for an arbitrary number of partitions.
Ivan V. Bajic, John W. Woods
IEEE Trans. Inf. Theory1
2002 Concatenated multiple description coding of frame-rate scalable video
abstract
In this paper we present a method for concatenated multiple description (MD) coding of frame-rate scalable video. The proposed method combines domain-based MD coding and FEC-based MD coding. We find that the combined system benefits from both of its components and is significantly better than either of them at higher packet loss rates.
Ivan V. Bajic, John W. Woods
ICIP (2)1
2002 Domain-based multiple description coding of images and video
Ivan V. Bajic, John W. Woods
VCIP1