Shahram Shirani

dblp:64/2629 · DBLP profile ↗
← Back
113ranked-venue papers
11as first author
18since 2021 · last 2025
0000-0002-7217-3467ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 94 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 5 since 2021Computer networks · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 SFI-Swin: symmetric face inpainting with swin transformer by distinctly learning face components distributions
Mohammad H. Givkashi, Mohammadreza Naderi, Nader Karimi, Shahram Shirani, Shadrokh Samavi
Multim. Tools Appl.4
2025 A parametric rate-distortion model for video transcoding
Maedeh Jamali, Nader Karimi, Shadrokh Samavi, Shahram Shirani
Multim. Tools Appl.4
2024 Medical Knowledge-Guided Semi-Supervised Bi-Ventricular Segmentation
abstract
Pseudo-labeling is a well-studied semi-supervised learning approach that generates artificial labels for unlabeled data based on the predictions of an initial model trained on labeled data. Although pseudo-labeling is an effective approach for a wide range of tasks, incorporating physical knowledge concentrated on a particular organ (such as cardiac structures), can outperform the general strategy employed in pseudo-labeling. In this work, we propose to integrate different physical (medical) properties into the semi-supervised bi-ventricular segmentation task. We incorporate this knowledge as regularization terms in the loss function, uncertainty criteria for assessing the predictions, and pseudo-label modification methods. Our extracted properties are based on the physical characteristics of the ventricles and are robust to modality changes. We validated our method using the ACDC and SCD datasets. Numerical measurements confirm the success of the proposed approach. The implementation of our work is available at “https://github.com/behnamrahmati/MedicalGuided”
Behnam Rahmati, Shahram Shirani, Zahra Keshavarz-Motamed
ICIP2
2024 Redundant co-training: Semi-supervised segmentation of medical images using informative redundancy
abstract
Pseudo-labeling, consistency regularization, and co-training are common paradigms for semi-supervised learning. In this paper, we propose a novel method based on co-training and pseudo-labeling for the semi-supervised segmentation of the left ventricle. Our co-training strategy is novel and unlike most previous works does not rely on using multiple-view datasets, performing weak/strong augmentations on the input images or perturbations on the networks. We proposed creating redundant labels by utilizing the provided ground-truths and training networks segmenting different overlapping regions corresponding to the created labels. Although the new labels seem to be redundant, we demonstrated that they provide valuable information to the networks. The predictions of the redundant networks (which are trained on the redundant labels) can be used in the pixels where the primary network’s predictions are not reliable. This enables extracting a secondary source of information without requiring any additional ground-truths. The common practice in pseudo-labeling is using the reliable predictions of the unlabeled data and discarding the unreliable ones. However, we proposed utilizing predictions from the redundant networks to generate pseudo-labels for the unreliable pixels in the primary network’s predictions, rather than simply discarding them. We validated our method on two left ventricle segmentation datasets, and it surpassed the state-of-the-art semi-supervised learning approaches. Furthermore, we conducted extensive studies to analyze the proposed method from different aspects. Implementation of our work is available at https://github.com/behnam-rahmati/redundant-cotraining .
Behnam Rahmati, Shahram Shirani, Zahra Keshavarz-Motamed
Neurocomputing2
2024 Semi-supervised segmentation of medical images focused on the pixels with unreliable predictions
Behnam Rahmati, Shahram Shirani, Zahra Keshavarz-Motamed
Neurocomputing2
2024 Supervised deep learning for content-aware image retargeting with Fourier Convolutions
Mohammad H. Givkashi, Mohammadreza Naderi, Nader Karimi, Shahram Shirani, Shadrokh Samavi
Multim. Tools Appl.4
2024 Aesthetic-aware image retargeting based on foreground-background separation and PSO optimization
Mohammadreza Naderi, Mohammad H. Givkashi, Nader Karimi, Shahram Shirani, Shadrokh Samavi
Multim. Tools Appl.4
2023 Deep CNN-Based Pre-Encoding Perceptual Quality Control and Prediction
abstract
Inevitable utilization of lossy compression methods results in distortion and degrading of video perceptual quality. Predicting perceptual quality before compression is essential to optimize compression parameters, e.g., quantization parameter (QP) and assigning optimized bandwidth. This paper presents three Intra frame (I-frame) perceptual quality prediction methods. The proposed methods work based on deep CNN structure. The proposed methods are integrated with high efficiency video coding (HEVC, H.265) reference codec. The VMAF index has been utilized to measure the perceptual quality of video samples. An end-to-end CNN network performs spatial feature extraction for perceptual quality prediction. The proposed methods are designed based on our experimental observations. The proposed methods are evaluated with 17 video samples, and the results show a reliable, accurate performance of approaches.
Maryam Jenab, Shahram Shirani
ICIP2
2023 Segmentation of the Left Ventricle for the Cardiac Phases between End-Diastole and End-Systole
abstract
Automatic segmentation of cardiac structures, including the left ventricle (LV), is a crucial step for evaluating cardiac function. While deep learning is the leading approach for this task, the scarcity of labeled data presents a significant challenge. In scenarios with very limited labeled data, even semi-supervised learning methods may not be effective as they rely on the performance of an initially trained network. In this work, we propose a novel method for segmenting the LV contours for unlabeled cardiac phases between end-diastole (ED) and end-systole (ES). Our method leverages temporal coherence from the LV volume-time and shape information from the labeled cardiac phases to transform the ED and ES ground truth labels to the intervening cardiac phases, resulting in a larger set of labeled data. We evaluate our approach using three methods: 1- visual inspection, 2- using our method as data augmentation for training a CNN with limited ground truths, and 3- comparing results with fully supervised segmentation networks. Our method outperforms in all validation methods. Implementation is available at: "https://github.com/behnam-rahmati/LV-label-propagation".
Behnam Rahmati, Shahram Shirani, Zahra Keshavarz-Motamed
ICIP2
2023 A Novel Weakly Supervised Segmentation Approach for Rapid Left Ventricle Annotation
abstract
In the field of medical image segmentation, convolutional neural networks stand out as a successful method. However, in order to perform well, several labeled images are required. Manual pixel-level annotation of medical images requires the presence of a well-trained expert, is time-consuming, and is expensive. Weakly supervised learning approaches aim to address these challenges.In this work, we propose a novel weakly supervised segmentation approach specifically designed for the left ventricle. By utilizing the circular shape of the left ventricle, we introduce a weak annotation framework based on concentric circles, representing pixels inside and outside the ventricle. Our weakly supervised learning approach incorporates a loss function that ignores unannotated pixels and incorporates total variation regularization for smooth predictions. Our method significantly reduces annotation time from several minutes to 5-10 seconds per scan and eliminates the need for expert presence. We validated our method on the Sunnybrook dataset and our method reached around 0.96 of the accuracy of the networks trained in fully supervised manner (with pixel-level annotations). The implementation of our work is available at "https://github.com/behnam-rahmati/LV-weaklysupervised"
Behnam Rahmati, Shahram Shirani, Zahra Keshavarz-Motamed
ICIP2
2023 Sequence-to-Sequence Multi-Modal Speech In-Painting
Mahsa Kadkhodaei Elyaderani, Shahram Shirani
INTERSPEECH2
2023 Robust watermarking using diffusion of logo into auto-encoder feature maps
Maedeh Jamali, Nader Karimi, Pejman Khadivi, Shahram Shirani, Shadrokh Samavi
Multim. Tools Appl.4
2023 Correction: Robust watermarking using diffusion of logo into auto-encoder feature maps
Maedeh Jamali, Nader Karimi, Pejman Khadivi, Shahram Shirani, Shadrokh Samavi
Multim. Tools Appl.4
2022 Atmospheric Turbulence Removal in Long-Range Imaging Using a Data-Driven-Based Approach
Hamid R. Fazlali, Shahram Shirani, Michael BradforSd, Thia Kirubarajan
Int. J. Comput. Vis.2
2022 Single image rain/snow removal using distortion type information
Hamid R. Fazlali, Shahram Shirani, Michael Bradford, Thia Kirubarajan
Multim. Tools Appl.2
2022 Neural network solution for a real-time no-reference video quality assessment of H.264/AVC video bitstreams
Yasamin Fazliani, Ernesto Andrade, Shahram Shirani
Multim. Tools Appl.3
2022 Feature Aggregation Networks Based on Dual Attention Capsules for Visual Object Tracking
abstract
Tracking-by-detection algorithms have considerably enhanced tracking performance with the introduction of recent convolutional neural networks (CNNs). However, most trackers directly exploit standard scalar-output CNN features, which may not capture enough feature encoding information, instead of aggregated CNN features of vector-output form. In this paper, we propose an end-to-end feature aggregation capsule framework. First, based on the existing CNN network, we aggregate a certain number of similar position-aware CNN features into a capsule to model the feature similarity. The acquired vector-level feature capsules (rather than previous scalar-level pointwise features) are utilized for differentiation learning. We then propose a group attention module to better model the entity representation between different capsule groups thus optimizes total discriminative capability. Third, to reduce the prediction interference resulted by the side effect of dimension rising within capsules, we propose a penalty attention module. Such strategy could dynamically adjust values of neurons by estimating whether they are beneficial or harmful to tracking. Experimental results on five representative benchmarks (UAVDT, DTB70, UAV123, VOT2016 and VOT2018) demonstrate the excellent tracking performance of our dual attention based capsule tracker (DACapT). Specially, it exceeds the previous top tracker by 4.6%/1.9% in precision/success evaluations on UAVDT.
Yi Cao 0003, Hongbing Ji, Wenbo Zhang 0007, Shahram Shirani
IEEE Trans. Circuits Syst. Video Technol.4
2021 Classification of diabetic retinopathy using unlabeled data and knowledge distillation
Sajjad Abbasi, Mohsen Hajabdollahi, Pejman Khadivi, Nader Karimi, Roshanak Roshandel, Shahram Shirani, Shadrokh Samavi
Artif. Intell. Medicine6
2020 Convolutional Neural Network Pruning Using Filter Attenuation
abstract
Filters are the essential elements in convolutional neural networks (CNNs). Filters generate feature maps and form the main part of the computational and memory requirements of the convolutional networks. In filter pruning methods, a filter with all of its components, including channels and connections, are removed. The removal of a filter can cause a drastic change in the network's performance. Also, the removed filters cannot come back to the network structure. We want to address these problems in this paper. We propose a CNN pruning method based on filter attenuation in which weak filters are not abruptly removed. Instead, weak filters are attenuated and gradually removed. In the proposed attenuation approach, there is a chance for weak filters to return to the network. The filter attenuation method is assessed using the VGG model for the Cifar10 image classification task. Simulation results show that the filter attenuation works well based on different pruning criteria, and better results are obtained in comparison with the conventional pruning methods.
Morteza Mousa Pasandi, Mohsen Hajabdollahi, Nader Karimi, Shadrokh Samavi, Shahram Shirani
ICIP5
2020 Image Watermarking with Region of Interest Determination Using Deep Neural Networks
abstract
Watermarking is a popular technique used in various applications, such as copyright protection of digital media, including audio, video, and image files. Proper watermarking should satisfy multiple criteria, such as robustness and transparency. While a successful watermarking needs to meet these criteria, there is a tradeoff between the two opposing criteria of robustness and transparency. This paper proposes a method for determining the appropriate locations for embedding watermarks with high strength factors. For this purpose, a deep neural network, known as Mask R-CNN, is used, which is pre-trained on the COCO dataset. This neural network finds a good strength factor for those sub-blocks of the host image selected for embedding. The proposed technique can be used in conjunction with most DWT and DCT based semi-blind watermarking approaches. Experiments show that the proposed method is robust against different attacks and demonstrates good transparency.
Mahnoosh Bagheri, Majid Mohrekesh, Nader Karimi, Shadrokh Samavi, Shahram Shirani, Pejman Khadivi
ICMLA5
2020 Aerial image dehazing using a deep convolutional autoencoder
Hamid R. Fazlali, Shahram Shirani, Michael McDonald 0001, Daly Brown, Thia Kirubarajan
Multim. Tools Appl.2
2020 Cloud/haze detection in airborne videos using a convolutional neural network
Hamid R. Fazlali, Shahram Shirani, Michael McDonald 0001, Thia Kirubarajan
Multim. Tools Appl.2
2019 Extremely Tiny Siamese Networks with Multi-level Fusions for Visual Object Tracking
Yi Cao 0003, Hongbing Ji, Wenbo Zhang 0007, Shahram Shirani
FUSION4
2019 Learning Based Hybrid No-Reference Video Quality Assessment of Compressed Videos
abstract
A near real-time no-reference video quality assessment method is proposed for videos encoded by H.264/AVC codec. A fully connected neural network is trained with features extracted from both bit-stream and pixel domains along with their respective subjective quality scores. Feature selection procedure is designed in a manner to address spatial as well as temporal artifacts of the encoded sequences, while minimizing the overall run-time in order to adapt this work to applications in live streaming. The performance of our method is verified by applying it on H.264-encoded sequences from the LIVE video dataset and the correlation with the differential mean opinion scores (DMOS) from the subjective tests are presented. Our framework outperforms widely-used NR VQA methods and a number of state-of-art full-reference VQA methods.
Yasamin Fazliani, Ernesto Andrade, Shahram Shirani
ISCAS3
2018 Robust image watermarking scheme using bit-plane of hadamard coefficients
Elham Etemad, Shadrokh Samavi, S. Mohamad R. Soroushmehr, Nader Karimi, Mohammad Etemad, Shahram Shirani, Kayvan Najarian
Multim. Tools Appl.6
2018 Hierarchical watermarking framework based on analysis of local complexity variations
Majid Mohrekesh, Shekoofeh Azizi, Shahram Shirani, Nader Karimi, Shadrokh Samavi
Multim. Tools Appl.3
2018 An Adaptive Patch-Based Reconstruction Scheme for View Synthesis by Disparity Estimation Using Optical Flow
abstract
Due to the rapid growth of technology and the dropping cost of cameras, multiview imaging applications have attracted many researchers in recent years. Free viewpoint and 3D Televisions are among these interesting applications. One of the problems that should be solved to realize such applications is rendering. In this paper, we propose an optical flow-assisted adaptive patch-based view synthesis algorithm. This patch-based scheme reduces the size and number of holes during reconstruction. The size of patch is determined in response to edge information for better reconstruction, especially near the boundaries. In the first stage of the algorithm, disparity is obtained using optical flow estimation. Then, a reconstructed version of the left and right views is generated using our adaptive patch-based algorithm. The mismatches between each view and its reconstructed version are obtained in the mismatch detection steps. This stage results in two masks as outputs, which help with the refinement of disparities and the selection of the best patches for final synthesis. Finally, the remaining holes are filled using our simple hole-filling scheme and the refined disparities. The objective and subjective performances of the proposed algorithm are compared with recent methods. The results show that the proposed algorithm achieves an improvement of 2.14 dB on average.
Hoda Rezaee Kaviani, Shahram Shirani
IEEE Trans. Circuits Syst. Video Technol.2
2017 Adaptive blind image watermarking using edge pixel concentration
Hamid R. Fazlali, Shadrokh Samavi, Nader Karimi, Shahram Shirani
Multim. Tools Appl.4
2017 Framework for robust blind image watermarking based on classification of attacks
M. Heidari, Shadrokh Samavi, S. Mohamad R. Soroushmehr, Shahram Shirani, Nader Karimi, Kayvan Najarian
Multim. Tools Appl.4
2017 Image retargeting using depth assisted saliency map
F. Shafieyan, Nader Karimi, Behzad Mirmahboub, Shadrokh Samavi, Shahram Shirani
Signal Process. Image Commun.5
2016 Toward practical guideline for design of image compression algorithms for biomedical applications
Nader Karimi, Shadrokh Samavi, S. Mohamad R. Soroushmehr, Shahram Shirani, Kayvan Najarian
Expert Syst. Appl.4
2016 Frame Rate Upconversion Using Optical Flow and Patch-Based Reconstruction
abstract
In this paper, we present a frame rate upconversion method using optical flow motion estimation and a patch-based reconstruction scheme. First, forward and backward motion vectors (MVs) are obtained using an optical flow algorithm, and reconstructed versions of the current and previous frames are generated by our patch-based reconstruction scheme. Using the original and reconstructed versions of the current and previous frames, two mismatch masks are obtained. Then, two versions of the middle frame are generated using a patch-based scheme, with estimated MVs and the current and previous frames. Finally, a middle mask, which identifies the mismatch areas of the two middle frames, is reconstructed. Using these three masks, the best candidates for interpolation are selected and fused to obtain the final middle frame. Due to the patch-based nature of our reconstruction scheme, most of the holes and cracks will be filled. Although there is always a probability of having holes, the size and number of such holes are much smaller than those that would be generated using pixel-based mapping. The rare holes are filled using existing hole-filling algorithms. The experimental results and a comparison of our method with existing algorithms show that our method performs better in terms of both objective and subjective quality of the final interpolated frames. The average peak signal-to-noise ratio (PSNR) improvement of our method is 1-2 dB.
Hoda Rezaee Kaviani, Shahram Shirani
IEEE Trans. Circuits Syst. Video Technol.2
2015 Ocrapose: An indoor positioning system using smartphone/tablet cameras and OCR-aided stereo feature matching
abstract
In this paper, we propose an image-based localization system, applicable for a number of indoor scenarios including office buildings, airports, chain stores, etc. In such applications, text/numbers are suitable distinctive landmarks for localization. The proposed system takes advantage of OCR to read the text/numbers and provide a rough estimate using the floor plan. Next, it performs OCR-aided stereo feature matching to refine the estimate by solving a PnP problem. Experiments show that this system achieves a median localization error of less than 50 cm for test positions located as far as 7 meters from a 20cm by 30cm number plate using different test devices in a university building scenario.
Hamed Sadeghi, Shahrokh Valaee, Shahram Shirani
ICASSP3
2015 Iterative mask generation method for handling occlusion in optical flow assisted view interpolation
abstract
Having a key role in 3D and free viewpoint TV applications, view interpolation techniques have attracted many researchers in recent years. In this paper we present a patch-based reconstruction scheme for view interpolation using optical flow disparity estimation. In the first step of the algorithm, reconstructed versions of the left and right views are obtained using the proposed patch-based scheme. Then mismatch masks are generated to find the best patches for final reconstruction. Finally the intermediate view is obtained by fusing the selected patches and the remaining holes are filled with a simple hole-filling algorithm. The performance of the proposed algorithm is compared with recent methods in terms of objective and subjective quality. The results show that the proposed method achieves an improvement of 3.3 dB on average.
Hoda Rezaee Kaviani, Shahram Shirani
ICIP2
2015 Multifocus image fusion based on surface area analysis
abstract
Multifocus image fusion is an important research area in image processing and machine vision applications. Due to use of optical lenses, captured images are not usually focused everywhere in the image. Therefore objects near the focal range have evident details while other objects appear blurry. Multifocus image fusion algorithm takes several images with different focal ranges and combines them to produce an image that is focused everywhere. To identify focused regions in each of input images, generally spatial domain and transform domain methods are used. These methods usually suffer from artifacts such as blockiness or ringing. In this paper we propose a new criteria to determine focused pixels in an image. We find the points that have the same intensities in input images and segment the input images based on them. Subsequently we calculate surface area of pixels inside every segment based on intensity variations over rows and columns. Segment with more surface area of input images selected as the focus segment. Experimental results reveal the superiority of our method in comparison to compared algorithms.
Iman Roosta, Nader Karimi, Shadrokh Samavi, Shahram Shirani
ICIP4
2015 Use of symmetry in prediction-error field for lossless compression of 3D MRI images
Nader Karimi, Shadrokh Samavi, Somaieh Amraee, Shahram Shirani
Multim. Tools Appl.4
2014 Frame rate up-conversion using nonparametric estimator
abstract
In some applications such as digital video broadcasting, video is transmitted over a low capacity channel with lower frame rates. The lower the frame rate, the jerkier or unevener the video motion would be noticed. To solve this problem, frame rate up conversion (FRUC) is employed to increase the frame rate. In this paper, we propose a new FRUC method using the nonlocal-means estimator. In this method, a pixel is reconstructed as a weighted linear combination of pixel pairs in its adjacent frames. The pixels of each pair are temporally symmetric from the view point of the pixel being interpolated. The weights are calculated based on the self-similarity assumption. To reduce the computational complexity, we calculate the weights of linear combination for each super-pixel. Experimental results show the superior performance of our proposed method in comparison to the existing methods.
Roozbeh Dehghannasiri, S. Mohamad R. Soroushmehr, Shahram Shirani
ICIP3
2014 Multi-view video super-resolution for hybrid cameras using modified NLM and adaptive thresholding
abstract
Dual-mode (hybrid) cameras are able to simultaneously shoot two types of video streams: a high resolution with low frame rate stream and similarly a low resolution with high frame rate stream. There are some works on super-resolving the single camera video sequences. In this paper we propose a method for multi-view video super-resolution, utilizing sequences generated from a hybrid camera. In the proposed method we exploit the self-similarity in the spatial and temporal domains to reconstruct a pixel. We modify the nonlocal means method to be applied in our method. Through a combination of techniques including adaptive thresholds and specialized candidate pixel selection schemes, the proposed method reconstructs a high fidelity video stream with considerably improved performance.
Robert Lengyel, S. Mohamad R. Soroushmehr, Shahram Shirani
ICIP3
2014 Strategic image denoising using a support vector machine with seam energy and saliency features
abstract
We propose a method of using a support vector machine (SVM) to select between multiple well-performing contemporary denoising algorithms for each pixel of a noisy image. We describe a number of novel and pre-existing features based on seam energy, local colour, and saliency which are used as inputs to the SVM. Our SVM strategic image de-noising (SVMSID) results demonstrate better image quality than either candidate denoising algorithm, as measured using the perceptually-based quaternion structural similarity image metric (QSSIM).
Laura McCrackin, Shahram Shirani
ICIP2
2014 Image seam carving using depth assisted saliency map
abstract
Retargeting algorithms are needed to transfer an image from a device to another with different size and resolution. The goal is to preserve the best visual quality for important objects of the original image. In order to reduce image size, pixels should be removed from less important parts of the image. Therefore, we need an energy function to select less important pixels in seam carving. Various energy functions have been proposed in previous works to minimize the distortion in salient objects. In this paper we combine three different importance maps to form a new energy map. We first use both gradient and depth maps to highlight the values in the saliency map, eventually generates the final energy map. Experimental results using the proposed energy map show better visual appearance in comparison to previous algorithms even at high resizing percentage. The visual artifacts that cause shape deformation in salient objects and deteriorates geometrical consistency of the scene are considerably reduced in our proposed algorithm.
F. Shafieyan, Nader Karimi, Behzad Mirmahboub, Shadrokh Samavi, Shahram Shirani
ICIP5
2014 Semi-supervised logo-based indoor localization using smartphone cameras
abstract
In this paper, we propose a homography-aware semi-supervised formulation for the logo-based indoor localization problem using smartphone cameras. Our method labels unmatched feature points detected inside the logo parts of query images with their estimated 3D coordinates. The 3D coordinates are computed using the homography estimated from the matched features. We demonstrate the accuracy improvement and lower localization error variance resulted from our semi-supervised approach via experiments in an indoor scenario.
Hamed Sadeghi, Shahrokh Valaee, Shahram Shirani
PIMRC3
2014 Simple and efficient motion estimation algorithm by continuum search
S. Mohamad R. Soroushmehr, Shadrokh Samavi, Shahram Shirani
Multim. Tools Appl.3
2014 A MAP-Based Image Interpolation Method via Viterbi Decoding of Markov Chains of Interpolation Functions
abstract
A new method of image resolution up-conversion (image interpolation) based on maximum a posteriori sequence estimation is proposed. Instead of making a hard decision about the value of each missing pixel, we estimate the missing pixels in groups. At each missing pixel of the high resolution (HR) image, we consider an ensemble of candidate interpolation methods (interpolation functions). The interpolation functions are interpreted as states of a Markov model. In other words, the proposed method undergoes state transitions from one missing pixel position to the next. Accordingly, the interpolation problem is translated to the problem of estimating the optimal sequence of interpolation functions corresponding to the sequence of missing HR pixel positions. We derive a parameter-free probabilistic model for this to-be-estimated sequence of interpolation functions. Then, we solve the estimation problem using a trellis representation and the Viterbi algorithm. Using directional interpolation functions and sequence estimation techniques, we classify the new algorithm as an adaptive directional interpolation using soft-decision estimation techniques. Experimental results show that the proposed algorithm yields images with higher or comparable peak signal-to-noise ratios compared with some benchmark interpolation methods in the literature while being efficient in terms of implementation and complexity considerations.
Farhang Vedadi, Shahram Shirani
IEEE Trans. Image Process.2
2013 Visual sensor network lifetime maximization by prioritized scheduling of nodes
Mohsen Hooshmand, S. Mohamad R. Soroushmehr, Pejman Khadivi, Shadrokh Samavi, Shahram Shirani
J. Netw. Comput. Appl.5
2013 An adaptive LSB matching steganography based on octonary complexity measure
Vajiheh Sabeti, Shadrokh Samavi, Shahram Shirani
Multim. Tools Appl.3
2013 De-Interlacing Using Nonlocal Costs and Markov-Chain-Based Estimation of Interpolation Methods
abstract
A new method of de-interlacing is proposed. De-interlacing is revisited as the problem of assigning a sequence of interpolation methods (interpolators) to a sequence of missing pixels of an interlaced frame (field). With this assumption, our de-interlacing algorithm (de-interlacer), undergoes transitions from one interpolation method to another, as it moves from one missing pixel position to the horizontally adjacent missing pixel position in a missing row of a field. We assume a discrete countable-state Markov-chain model on the sequence of interpolators (Markov-chain states) which are selected from a user-defined set of candidate interpolators. An estimation of the optimum sequence of interpolators with the aforementioned Markov-chain model requires the definition of an efficient cost function as well as a global optimization technique. Our algorithm introduces for the first time using a nonlocal cost (NLC) scheme. The proposed algorithm uses the NLC to not only measure the fitness of an interpolator at a missing pixel position, but also to derive an approximation for transition matrix (TM) of the Markov-chain of interpolators. The TM in our algorithm is a frame-variate matrix, i.e., the algorithm updates the TM for each frame automatically. The algorithm finally uses a Viterbi algorithm to find the global optimum sequence of interpolators given the cost function defined and neighboring original pixels in hand. Next, we introduce a new MAP-based formulation for the estimation of the sequence of interpolators this time not by estimating the best sequence of interpolators but by successive estimations of the best interpolator at each missing pixel using Forward-Backward algorithm. Simulation results prove that, while competitive with each other on different test sequences, the proposed methods (one using Viterbi and the other Forward-Backward algorithm) are superior to state-of-the-art de-interlacing algorithms proposed recently. Finally, we propose motion compensated versions of our algorithm based on optical flow computation methods and discuss how it can improve the proposed algorithm.
Farhang Vedadi, Shahram Shirani
IEEE Trans. Image Process.2
2013 Lossless Compression of RNAi Fluorescence Images Using Regional Fluctuations of Pixels
abstract
RNA interference (RNAi) is considered one of the most powerful genomic tools which allows the study of drug discovery and understanding of the complex cellular processes by high-content screens. This field of study, which was the subject of 2006 Nobel Prize of medicine, has drastically changed the conventional methods of analysis of genes. A large number of images have been produced by the RNAi experiments. Even though a number of capable special purpose methods have been proposed recently for the processing of RNAi images but there is no customized compression scheme for these images. Hence, highly proficient tools are required to compress these images. In this paper, we propose a new efficient lossless compression scheme for the RNAi images. A new predictor specifically designed for these images is proposed. It is shown that pixels can be classified into three categories based on their intensity distributions. Using classification of pixels based on the intensity fluctuations among the neighbors of a pixel a context-based method is designed. Comparisons of the proposed method with the existing state-of-the-art lossless compression standards and well-known general-purpose methods are performed to show the efficiency of the proposed method.
Nader Karimi, Shadrokh Samavi, Shahram Shirani
IEEE J. Biomed. Health Informatics3
2012 Elevating watermark robustness by data diffusion in Contourlet coefficients
abstract
As concerns about copyright protection increased amongst multimedia owners in recent years, many watermarking algorithms proposed to protect copyright of digital images. These methods are either spatial or frequency domain techniques. It is essential for a watermarking method to have acceptable robustness. That is why many existing methods try to improve their robustness against signal processing modifications. In this paper a block based watermarking scheme is proposed that embeds a binary logo into Contourlet coefficients of image. To increase robustness, embedding is done in two scales and watermark is inserted into DCT coefficients of Contourlet blocks to diffuse the effects of the embedding throughout the coefficients. Experimental results, and comparison with a robust Contourlet domain method, show that the proposed scheme has better robustness against some image processing attacks, while improving fidelity. Furthermore the proposed algorithm has the advantage of having a blind extraction phase.
Hoda Rezaee Kaviani, Shadrokh Samavi, Nader Karimi, Shahram Shirani
ICC4
2012 Image resolution up-conversion via maximum a posteriori interpolator sequence estimation and Viterbi algorithm
abstract
A new method of image resolution up-conversion based on maximum a posteriori sequence estimation is proposed. At each missing pixel of the high resolution (HR) image we consider an ensemble of candidate interpolation methods (interpolator). The interpolators are interpreted as states of a finite-state machine (FSM). Accordingly, the up-scaling problem is converted to the problem of estimating the optimal sequence of interpolators corresponding to the sequence of missing HR pixel positions. We derive a parameter-free probabilistic model for this FSM to solve the estimation problem using trellis diagrams and Viterbi algorithm. The experimental results prove that the proposed algorithm results sharper HR images and higher peak signal-to-noise ratios (PSNR) comparing to many algorithms in this domain.
Farhang Vedadi, Shahram Shirani
ICIP2
2012 View-Invariant Fall Detection System Based on Silhouette Area and Orientation
abstract
Population of old generation that live alone is growing in most countries. Surveillance systems help them stay home and reduce the burden on the healthcare system. Automatic visual surveillance systems have advantages over wearable devices. They extract features from video sequences and use them for event classification. But these features are dependent on the position of cameras relative to the person. Therefore they need multi-camera for more accuracy that increases cost and complexity. In this paper we propose using silhouette area combined with inclination angle as robust features that can be measured using only one camera with an arbitrary direction. Through rigorous simulations on a publicly available dataset the error rate of the system is found to be less than 1%.
Behzad Mirmahboub, Shadrokh Samavi, Nader Karimi, Shahram Shirani
ICME4
2011 Compression of 3D MRI images based on symmetry in prediction-error field
abstract
Three dimensional MRI images which are power tools for diagnosis of many diseases require large storage space. A number of lossless compression schemes exist for this purpose. In this paper we propose a new approach for the compression of these images which exploits the inherent symmetry that exists in the 3D MRI images. A block matching routine is employed to work on the symmetrical characteristics of these images. Another type of block matching is also applied to eliminate the inter-slice temporal correlations. The obtained results outperform the existing standard compression techniques.
Somaieh Amraee, Nader Karimi, Shadrokh Samavi, Shahram Shirani
ICME4
2011 Size-Controllable Region-of-Interest in Scalable Image Representation
abstract
Differentiating region-of-interest (ROI) from non-ROI in an image in terms of relative size as well as fidelity becomes an important functionality for future visual communication environment with a variety of display devices. In this paper, we propose a scalable image representation with the ROI functionality in the spatial domain, which allows us to generate a hierarchy of images with arbitrary sizes. The ROI functionality of our scalable representation is a result of a nonuniform grid transformation in the spatial domain, where only the center of ROI and an expansion parameter are to be known. Our grid transformation guarantees no loss of information within the area of ROI.
Chee Sun Won, Shahram Shirani
IEEE Trans. Image Process.2
2010 Compressive sensing with modified Total Variation minimization algorithm
abstract
In this paper, the reconstruction problem of compressive sensing algorithm that is exploited for image compression, is investigated. Considering the Total Variation (TV) minimization algorithm, and by adding some new constraints compatible with typical image properties, the performance of the reconstruction is improved. Using DCT and contourlet transforms, sparse expansion of the image are exploited to provide new constraints to remove irrelevant vectors from the feasible set of the optimization problem while keeping the problem as a standard Second Order Cone Programming (SOCP) one. Experimental results show that, the proposed method, with new constraints, outperforms the conventional TV minimization method by up to 2 dB in PSNR.
Mohammadreza Dadkhah, Shahram Shirani, M. Jamal Deen
ICASSP2
2010 Adaptive Modification of Transform Coefficients for Image Compression
Nader Karimi, Shadrokh Samavi, Shahram Shirani
ICASSP3
2010 Multi-Layered image compression using structure tensor for texture identification
abstract
Compression of images using transform methods has been of interest for many years. In this paper we propose a new multilayer image compression method which uses wavelet and contourlet transforms. We used structure tensor for identifying texture regions of the image by producing a binary mask. Then we apply wavelet to smooth regions and use contourlet transform for texture area. The proposed method avoids the redundancy of contourlet which has been a bottleneck for low bit rate compression purposes. We showed that images that are compressed and reconstructed by our method at low bit rates have good qualities both visually and in terms of the produced PSNRs.
Hossein Talebi Esfandarani, Nader Karimi, Shadrokh Samavi, Shahram Shirani
ICME4
2010 Real time fractal image coder based on characteristic vector matching
Shadrokh Samavi, Mehdi Habibi, Shahram Shirani, Narges Rowshanbin
Image Vis. Comput.3
2010 Steganalysis and payload estimation of embedding in pixel differences using neural networks
Vajiheh Sabeti, Shadrokh Samavi, Shahram Shirani
Pattern Recognit.4
2010 Unequal Erasure Protection Technique for Scalable Multistreams
abstract
This paper presents a novel unequal erasure protection (UEP) strategy for the transmission of scalable data, formed by interleaving independently decodable and scalable streams, over packet erasure networks. The technique, termed multistream UEP (M-UEP), differs from the traditional UEP strategy by: 1) placing separate streams in separate packets to establish independence and 2) using permuted systematic Reed-Solomon codes to enhance the distribution of message symbols amongst the packets. M-UEP improves upon UEP by ensuring that all received source symbols are decoded. The R-D optimal redundancy allocation problem for M-UEP is formulated and its globally optimal solution is shown to have a time complexity of O(2(N)N(L+1)(N+1)) , where N is the number of packets and L is the packet length. To address the high complexity of the globally optimal solution, an efficient suboptimal algorithm is proposed which runs in O(N(2)L(2)) time. The proposed M-UEP algorithm is applied on SPIHT coded images in conjunction with an appropriate grouping of wavelet coefficients into streams. The experimental results reveal that M-UEP consistently outperforms the traditional UEP reaching peak improvements of 0.6 dB. Moreover, our tests show that M-UEP is more robust than UEP in adverse channel conditions.
Sorina Dumitrescu, Geoffrey Rivers, Shahram Shirani
IEEE Trans. Image Process.3
2009 A block based encoding algorithm for matching pursuit image coding
abstract
A new encoding algorithm is proposed for encoding matching pursuit parameters in image compression applications. The new algorithm is a block based coding approach that takes advantage of correlations in atom positions and inner product coefficients in a rate distortion optimal manner. Simulation results show the proposed algorithm significantly improves the performance of the existing encoding algorithms and outperforms JPEG2000 at low bit rates.
Alireza Shoa, Shahram Shirani
ICIP2
2009 Block matching algorithm based on local codirectionality of blocks
abstract
In this paper based on statistical analysis performed on a number of video sequences with different motion characteristics, we show that strong correlation exists between the direction of motion of a block and those of its neighboring blocks. Secondly, we show that a block moves in the same direction as a neighboring block that its prediction vector produces least distortion. Based on these findings an algorithm is proposed which adaptively determines the direction of motion of a block based on the motion direction of its neighboring blocks. The algorithm almost always avoids trapping into local minima. Test results show that in terms of speed and performance the proposed algorithm is superior to many of the existing fast algorithms.
S. Mohamad R. Soroushmehr, Shadrokh Samavi, Shahram Shirani
ICME3
2009 Optimized Atom Position and Coefficient Coding for Matching Pursuit-Based Image Compression
abstract
In this paper, we propose a new encoding algorithm for matching pursuit image coding. We show that coding performance is improved when correlations between atom positions and atom coefficients are both used in encoding. We find the optimum tradeoff between efficient atom position coding and efficient atom coefficient coding and optimize the encoder parameters. Our proposed algorithm outperforms the existing coding algorithms designed for matching pursuit image coding. Additionally, we show that our algorithm results in better rate distortion performance than JPEG 2000 at low bit rates.
Alireza Shoa, Shahram Shirani
IEEE Trans. Image Process.2
2008 Near lossless image compression by local packing of histogram
abstract
In this paper a low complexity algorithm is proposed for near lossless compression of images. The reconstructed near lossless image can differ from the original one within a pixelwise error tolerance. This property is used to convert the histogram of the original image, by the proposed algorithm, to a new histogram which is proved to have minimum entropy. Hence, a new image is formed which has minimum entropy and high spatial correlation among its pixels and can efficiently be compressed. Simulation results show the effectiveness of this compression algorithm.
Ebrahim Nasr-Esfahani, Shadrokh Samavi, Nader Karimi, Shahram Shirani
ICASSP4
2008 Sub-optimal MMSE based joint source/channel decoding of a matching pursuit coded image bit-stream over a memoryless noisy channel
abstract
Joint source/channel (JSC) decoding based on using the residual redundancy in a source coder output stream is an interesting bandwidth efficient method of reducing the effects of a noisy channel. In this paper, we consider the problem of JSC decoding of a matching pursuit (MP) based image transmission over a memory-less noisy channel. This problem is solved by means of MMSE decoding formulation for a sequence of received analysis indices, and a sub-optimal solution which yields high quality error concealment is devised. The proposed method exploits the residual redundancy that exists in neighboring image blocks as well as the neighboring MP analysis stages to overcome the channel noise degradations with no increase in the required bandwidth.
Abbas Ebrahimi-Moghadam, Shahram Shirani
ICME2
2008 Multiple description coding of audio using phase scrambling
abstract
In this paper, we proposed a method to decrease the effects of data loss in an audio segment using phase scrambling. Phase scrambling is used to spread the data of each block of an audio segment over all other blocks of the scrambled audio, followed by coding. In our experiments we observed the effects of different cases of loss and also used a customized recovery method for lost segments of data. The results obtained by employing this method shows great improvements compared to cases of data loss without exploiting phase scrambling. This technique can be readily used in transmission of audio segments over unreliable networks such as VoIP.
Seyed-Parsa Hojjat, Kaveh F. Sadri, Shahram Shirani
ICME3
2008 Novel R-D optimized uneven erasure-protection strategy for scalable data formed of multiple code streams
abstract
This paper presents a novel uneven erasure-protection (UEP) strategy for scalable data formed by interleaving independently decodable and scalable sub-streams. The main differences between UEP and the proposed novel strategy are: 1) the placement of separate sub-streams in separate packets to establish independence and 2) the use of permuted systematic Reed-Solomon (RS) codes to enhance the distribution of message symbols amongst the packets. To maximize the performance of this transmission scheme, we optimize the erasure protection assignment in a rate-distortion sense. Our experiments demonstrate the net superiority of the new UEP strategy over the traditional UEP for moderate numbers of packets (up to 42). Peak performance improvements (0.3-0.7 dB) occur between 4-10 packets for various transmission rates and packet erasure rates.
Geoffrey Rivers, Sorina Dumitrescu, Shahram Shirani
MMSP3
2008 Adaptive Rate-Distortion Optimal In-Loop Quantization for Matching Pursuit
abstract
In this paper, an adaptive in-loop quantization technique is proposed for quantizing inner product coefficients in matching pursuit. For each matching pursuit (MP) stage a different quantizer is used based on the probability distribution of MP coefficients. The quantizers are optimized for a given rate budget constraint. Additionally, our proposed adaptive quantization scheme finds the optimal quantizers for each stage based on the already encoded inner product coefficients. Experimental results show that our proposed adaptive quantization scheme outperforms existing quantization methods used in matching pursuit image coding.
Alireza Shoa, Shahram Shirani
IEEE Trans. Image Process.2
2007 A Robust Method of Determining Context-Coder Solutions for Encoding Affine Motion Vectors
abstract
The translational motion model can't effectively model complex motion such as scaling, shearing and rotation. That is why more complex motion models like the affine model have been proposed. However, affine motion vectors (AMV)s are more complicated to encode than the translational motion vectors. In this paper we propose several context-coder solutions based on our novel context type limited exhaustive search simulation. As a result the average compression gains of 7.9%, 5.9%, and 9.6%, for mobile, cost guard, and modified mobile video sequences were respectively realized, with peek compression improvements of 13%. In addition, our simulation is described and successfully compared with J. Vaisey and Jin Tong (2002).
Roman C. Kordasiewicz, Michael Gallant, Shahram Shirani
ICASSP (1)3
2007 Affine Prediction as a Post Processing Stage
abstract
Translational motion vectors (MV)s and macro block (MB) frame partitioning are the predominant means of motion estimation (ME) and motion compensation (MC). However, the translational motion model does not describe sufficiently complex motion such as rotation, zoom or shearing. To remedy this one can start computing more advanced motion parameters and/or partition the frame differently. However these approaches are either very computationally expensive and/or have limited search ranges. Thus, in this paper we propose a novel post processing stage which can be easily incorporated into most of the current coders. This stage generates the predictor for each inter MB, based on an affine motion model using translational motion vectors. Our approach has very low computational complexity, however average PSNR gains of up to 0.6 dB were realized for video sequences with complex motion.
Roman C. Kordasiewicz, Michael Gallant, Shahram Shirani
ICASSP (1)3
2007 Near-Lossless Image Compression Based on Maximization of Run Length Sequences
abstract
In this paper an algorithm is proposed which performs near-lossless image compression. For each pixel in a row of the image a group of value-states are considered, which have values close to that of the pixel. A trellis is constructed for every row of the image where the nodes of the trellis are the states of the pixels of that row. The goal of the algorithm is to find a path on this trellis that creates a sequence which can be efficiently coded using run length encoding (RLE). For sections of the pixels of the row that suitable RLE cannot be achieved then minimization of the entropy is employed to complete a path on the trellis. The application of the algorithm to a wide range of standard images shows that the scheme, while having low computational complexity, is competitive with other near-lossless image compression methods.
Ebrahim Nasr-Esfahani, Shadrokh Samavi, Nader Karimi, Shahram Shirani
ICIP (4)4
2007 Lossless Microarray Image Compression using Region Based Predictors
abstract
Microarray image technology is a powerful tool for monitoring the expression of thousands of genes simultaneously. Each microarray experiment produces large amount of image data, hence efficient compression routines that exploit microarray image structures are required. In this paper we introduce a lossless image compression method which segments the pixels of the image into three categories of background, foreground, and spot edges. The segmentation is performed by finding a threshold value which minimizes the weighted sum of the standard deviations of the foreground and background pixels. Each segment of the image is compressed using a separate predictor. The results of the implementation of the method show its superiority compared to the well-known microarray compression schemes as well as to the general lossless image compression standards.
Abbas Neekabadi, Shadrokh Samavi, S. A. Razavi, Nader Karimi, Shahram Shirani
ICIP (2)5
2007 Distributed Parameter Estimation with Side Information: A Factor Graph Approach
abstract
In this paper, a low complexity algorithm for distributed maximum likelihood estimation of a binary symmetric source (BSS) using side-information is proposed. The estimation is formulated as an incomplete-data problem and is solved by the expectation-maximization (EM) algorithm. A low-complexity implementation of the algorithm using coset codes and LDPC-based syndrome decoding with message passing over factor-graph is also proposed. The algorithm is a generalization of the LDPC-based syndrome decoding algorithm for the case when the probability distribution of the source is not known a-priori. Hence, the algorithm may be considered as a tool for achieving the corner points of the Slepian-Wolf (SW) region in distributed coding when the correlation channel information is not available. The estimation efficiency is studied by comparing the mean square error with the achievable Fisher information.
Amin Zia, James P. Reilly, Shahram Shirani
ISIT3
2007 Modeling Quantization of Affine Motion Vector Coefficients
abstract
Affine motion compensated prediction (AMCP) is an advanced tool which may be incorporated into future video compression standards. There are numerous coders already using AMCP . However the increased number of motion vector components is a disadvantage and quantizing these components can have significant consequences on the difference macro blocks (DMBs). This paper examines the quantization of affine motion vector (AMV) coefficients, by deriving a quadratic relationship between DMB energy and AMV quantization step size. Mathematical derivations and simulations are provided, including two literature comparisons demonstrating the benefits of this work. In the first comparison, the quantization of orthogonalized AMVs in is compared with quantization guided by the novel quadratic model. In the second comparison, Nokia's MVC coder is modified to use the quadratic model to generate quantization step sizes for various granularities; sequence, frame, and quarter-frame, demonstrating up to 8.7% bit rate reductions. Model driven AMV quantization step size choices are shown to be very close to and even outperform limited exhaustive search AMV quantization step size choices, at a quarter of the computational cost
Roman C. Kordasiewicz, Michael Gallant, Shahram Shirani
IEEE Trans. Circuits Syst. Video Technol.3
2007 Affine Motion Prediction Based on Translational Motion Vectors
abstract
In all of the video coding standards like H.26X and MPEG-X, much of the compression comes from motion compensated prediction (MCP). Translational motion vectors (MVs) poorly model complex motion and thus coders using polynomial or affine MVs have been proposed in the past. In this paper, we demonstrate a novel affine predictor stage which can be easily incorporated into current codecs greatly increasing MCP quality. If used passively to generate the final prediction, gains of up to 0.7 and 1.6 dB were easily realized for ldquomobilerdquo and ldquoflower gardenrdquo video sequences, respectively. In addition, when the translational MVs are refined, gains of up to 0.98 and 1.88 dB for ldquomobilerdquo and ldquoflower gardenrdquo video sequences were respectively realized.
Roman C. Kordasiewicz, Michael Gallant, Shahram Shirani
IEEE Trans. Circuits Syst. Video Technol.3
2007 Matching Pursuit-Based Region-of-Interest Image Coding
abstract
Matching pursuit (MP) is a multiresolution signal analysis method and can be used to render a selected region of an image with a specific quality. A novel, scalable, and progressive MP-based region-of-interest image-coding scheme is presented. The method is capable of providing a trade off between rate, distortion, and complexity. The method also provides an interactive way of information refinement for regions of an image with higher receiver's priority. By selecting a proper subset of the huge initial MP dictionary, using the method described in this paper, the complexity burden of MP analysis can be adapted to the computational power of the image encoder.
Abbas Ebrahimi-Moghadam, Shahram Shirani
IEEE Trans. Image Process.2
2007 Encoding of Affine Motion Vectors
abstract
An affine motion model provides better motion representation than a translational motion model. Therefore, it is a good candidate for advanced video compression algorithms, requiring higher compression efficiency than current algorithms. One disadvantage of the affine motion model is the increased number of motion vector parameters, therefore increased motion vector bit rate. We develop and analyze several simulation based approaches of entropy coding for orthonormalized affine motion vector (AMV) coefficients, by considering various context-types and coders. In our work we expand the traditional idea of a context type by introducing four new context types. We compare our method of contexts-type and coder selection with context quantization. The best of our contexts-type and coder solutions produces 4% to 15% average AMV bit-rate reductions over the original VLC approach. For more difficult content AMV bit rate reduction up to 26% is reported.
Roman C. Kordasiewicz, Michael Gallant, Shahram Shirani
IEEE Trans. Multim.3
2006 Distortion of Matching Pursuit: Modeling and Optimization
abstract
Summary form only given. The distortion of matching pursuit is expressed in terms of MP encoder parameters for uniformly distributed signals and dictionaries. Under certain conditions, the distortion caused by matching pursuit decomposition of the signal can be calculated in terms of the norm of the signal, signal dimension, dictionary size, and the number of matching pursuit stages. The distortion caused by quantization of inner product coefficients can be calculated in terms of the number of quantization levels and the norm of the signal assuming the coefficients are uniformly distributed in the range of the quantizer. The distortion of the matching pursuit for random signals and dictionaries was accurately predicted based on simulation results. The optimized matching pursuit encoder shows optimum performance for non-uniform signal and dictionary distributions
Alireza Shoa, Shahram Shirani
DCC2
2006 A Coding Theorem for Multiterminal Estimation
abstract
In this paper a coding theorem for multiterminal estimation is presented. The theorem is a generalization of the distributed coding theorem first proved by Slepian and Wolf (1973), where the goal is to estimate the joint probability distribution of correlated sources, rather than to reconstruct them at the receiver. For this, it is shown first that the joint-type of the received sequences is a sufficient statistic for estimation. Then, it is proved that for sufficiently large sequences, only a sum-rate lower bounded by the mutual information of the correlated sources is "sufficient" to perfectly reconstruct the sufficient statistic at the receiver. Simulation results for the special case of estimation with side information at the receiver is provided
Amin Zia, James P. Reilly, Timothy R. Field, Shahram Shirani
ICASSP (4)4
2006 Novel Progressive Region of Interest Image Coding Based on Matching Pursuits
abstract
A progressive and scalable, region of interest (ROl) image coding scheme based on matching pursuits (MP) is presented. Matching pursuit is a multi-resolutional signal analysis tool and can be employed in order to progressively refine the quality of a set of selected regions of an image up to a specific grade. The computational complexity of this analysis method can be reduced by decreasing the size of MP dictionary. Thus, the proposed method provides a trade off between complexity, rate, and quality. By the suggested scheme, regions of an image with higher receiver's priority are refined in an interactive manner. The transmitter sends an initial coarse version of the image. Then, the receiver transmits its preferred ROI parameters. Afterwards, the reconstructed image is refined according to the ROl parameters, in a progressive way
Abbas Ebrahimi-Moghadam, Shahram Shirani
ICME2
2006 Optimization of Matching Pursuit Encoder Based on Analytical Approximation of Matching Pursuit Distortion
abstract
Distortion of matching pursuit is calculated in terms of matching pursuit encoder parameters for uniformly distributed signals and dictionaries. Then, the MP encoder is optimized using the analytically derived approximation for MP distortion. Our simulation results show that this optimized MP encoder exhibits optimum performance for nonuniform signal and dictionary distributions as well
Alireza Shoa, Shahram Shirani
ICME2
2006 Adaptive Quantization for Matching Pursuit
abstract
We propose an adaptive quantization scheme for matching pursuit. Different quantizers are used in different matching pursuit (MP) stages based on the probability distribution of MP coefficients. The quantizers are optimized for a given rate budget constraint. Additionally, our proposed adaptive quantization scheme finds the optimal quantizers for each stage based on the information known to both encoder and decoder. Our quantization scheme outperforms existing MP encoders
Alireza Shoa, Shahram Shirani
MMSP2
2006 On Interpolation and Resampling of Discrete Data
abstract
This letter introduces a new representation of discrete signals based on the mathematical notions of functionals and continuous dual spaces. A new and more general sampling theorem is also suggested. Next, the problems of interpolating and resampling discrete signals are addressed; and a general solution using functional interpolation-which is applicable to many different settings-is proposed. Families of resampling filters dubbed de Boor-Ron filters that use de Boor-Ron interpolation are introduced, and their numerical realization is discussed. Some applications of this research are suggested
Pouya Dehghani Tafti, Shahram Shirani, Xiaolin Wu 0001
IEEE Signal Process. Lett.2
2006 Real-time processing and compression of DNA microarray images
abstract
In this paper, we present a pipeline architecture specifically designed to process and compress DNA microarray images. Many of the pixilated image generation methods produce one row of the image at a time. This property is fully exploited by the proposed pipeline that takes in one row of the produced image at each clock pulse and performs the necessary image processing steps on it. This will remove the present need for sluggish software routines that are considered a major bottleneck in the microarray technology. Moreover, two different structures are proposed for compressing DNA microarray images. The proposed architecture is proved to be highly modular, scalable, and suited for a standard cell VLSI implementation.
Shadrokh Samavi, Shahram Shirani, Nader Karimi
IEEE Trans. Image Process.2
2006 Content-based multiple description image coding
abstract
The multiple description coding method proposed in this paper provides the least amount of degradation, caused by loss of descriptors, for those areas of the image which are of greater interest. This is achieved by employing a nonlinear geometrical transform to add redundancy mainly to the area of interest followed by a partitioning of the transformed image into subimages which are coded and transmitted separately. Simulations show that this approach yields acceptable performance even when only one descriptor is received.
Shahram Shirani
IEEE Trans. Multim.1
2005 Multi-dimensional average-interpolating refinement on arbitrary lattices
abstract
Multi-dimensional datasets containing local averages of a function arise in many applications such as processing of CCD captures and medical images. Motivated by this fact we introduce multi-dimensional average-interpolating refinement on arbitrary lattices in arbitrary dimensions. Our refinement algorithm results in smooth scaling functions of compact support. This method forms a basis for multi-dimensional multi-resolution analysis and subdivision on datasets obtained by locally averaging a smooth function. As an example, we present two-dimensional polynomial average-interpolating subdivision on the quincunx lattice and show that the resulting scaling functions are highly regular in the sense of Sobolev.
Pouya Dehghani Tafti, Shahram Shirani, Xiaolin Wu 0001
ICASSP (4)2
2005 ASIC and FPGA implementations of H.264 DCT and quantization blocks
abstract
In the search for ever better and faster video compression standards H.264 was created. With it arose the need for hardware acceleration of its very computationally intensive parts. To address this need, this paper proposes two sets of architectures for the integer discrete transform (DCT) and quantization blocks from H.264. The first set of architectures for the DCT and quantization were optimized for area, which resulted in transform and quantizer blocks that occupy 294 and 1749 gates respectively. The second set of speed optimized designs has a throughput anywhere from 11 to 2552 M pixels/s. All of the designs were synthesized for Xilinx Virtex 2-Pro and 0.18/spl mu/m TSMC CMOS technology, as well as the combined DCT and quantization blocks went through comprehensive place and route flow.
Roman C. Kordasiewicz, Shahram Shirani
ICIP (3)2
2005 Modelling the effect of quantizing affine motion vectors on rate and energy of difference macroblocks
abstract
This paper derives the relationship between the energy of difference macroblock and affine motion vector quantization step size. This is an important step in the analysis of advanced motion models for future video coding methods, as it facilitates rate optimization for affine motion vectors. The derived model shows that the difference macroblock energy has a squared relationship with the affine motion vector quantization step size. In addition, experimental results are shown validating this result and providing further insight.
Roman C. Kordasiewicz, Shahram Shirani, Michael Gallant
ICIP (1)2
2005 Tree structure search for matching pursuit
abstract
Matching pursuit has found many applications recently especially in very low bit rate video coding. In this paper we show how the complexity of matching pursuit can be reduced from O(N) to O(log(2N)) using tree structured dictionaries. Moreover, we show how tree structured dictionaries provide an efficient coding strategy that is more resilient to error than random coding. Our simulation results showed an improvement of about 3 dB in PSNR can be achieved using the code provided by the tree structured dictionary.
Alireza Shoa, Shahram Shirani
ICIP (3)2
2005 Area of surface as a basis for vertex removal based mesh simplification
abstract
A new, area-based mesh simplification algorithm is described. The proposed algorithm removes the center vertex of a polygon which consists of n/spl ges/3 faces and represents that polygon with n-2 faces. A global search method is introduced that iteratively determines which vertex is to be removed using the proposed area-based distortion measurement. Various re-triangulations are also considered to improve the perceptual quality of the final approximation. Experimental results demonstrate the performance of the proposed algorithm for data reduction while maintaining the quality of the rendered objects.
Insu Park, Shahram Shirani, David W. Capson
ICME2
2005 Error Concealment for Facial Animation Based on Prediction of Muscle Data
abstract
An error concealment algorithm is proposed based on flow of facial expression to improve communication of animated facial data over a limited bandwidth and error prone channel. Facial expression flow is tracked using dominant muscles which are those with maximum change between two successive frames. The receiver uses linear interpolation and the information on facial expression flow to interpolate the erroneous facial animation data. Experimental results are provided to show that the proposed error concealment method improves the quality of an animated face communications.
Insu Park, Shahram Shirani, David W. Capson
ICME2
2005 Error-resilient region-of-interest video coding
abstract
An error resilient video coding method is proposed. It yields less degradation, caused by data loss, in areas of the frame which are of greater interest compared to the rest of the frame. This is achieved by employing a nonlinear transform that duplicates macroblocks inside the region of interest (ROI). This increased redundancy for ROI yields a high quality reconstruction of ROI even with missing data. Simulations show that this approach has acceptable performance. Moreover, the method proposed can be implemented through pre- and post-processing of the video sequence, without modification to the source codecs (e.g., H.263, MPEG4, H.264).
Ali Jerbi, Shahram Shirani
IEEE Trans. Circuits Syst. Video Technol.3
2005 Progressive scalable interactive region-of-interest image coding using vector quantization
abstract
We have developed novel progressive scalable region-of-interest (ROI) image compression schemes with rate-distortion-complexity tradeoff based on vector quantization. Residual vector quantization (RVQ) equips the encoder with a multi-resolution apparatus which is useful for rate-distortion tradeoff. Having all advantages of RVQ, jointly suboptimized RVQ provides a distortion-complexity adjustment. The systems are unbalanced in the sense that the decoder has less computational requirements than the encoder. The proposed jointly suboptimized RVQ method provides an interactive tool for fast ROI-based browsing from image archives.
Abbas Ebrahimi-Moghadam, Shahram Shirani
IEEE Trans. Multim.2
2004 Lossless and Lossy Compression of DNA Microarray Images
abstract
This paper presents a new approach for lossless and lossy compression of DNA microarray images. A DNA microarray is a single stranded DNA fragments, arranged on a glass or nylon slide where the hybridized spots can be detected and extracted by laser scanning of the microarray. The spatial optimization techniques and transforms are employed to achieve excellent compression ratio.
Naser Faramarzpour, Shahram Shirani
Data Compression Conference2
2004 Multiple description coding of images using phase scrambling
abstract
We discuss phase scrambling which spreads the information in each pixel of an image among virtually all the pixels of the resulting scrambled image. This property can be exploited in multiple description coding of images where the loss of one or many descriptions is a common case. We employ phase scrambling, as a form of all-pass filtering to mix the information of each pixel with all the pixels of the image, followed by decomposing the scrambled image into multiple descriptions. Our experiments show that this technique is competitive with other proposed methods, such as lapped orthogonal transforms. Phase scrambling does not produce localized visual artifacts, such as ringing and blocking effects, and does not require complex post-filtering to yield acceptable reconstruction quality. Another advantage of phase scrambling is that the scrambling can be implemented in hardware and performed in real time.
Kaveh F. Sadri, Shahram Shirani
ICASSP (3)2
2004 An information geometric approach to channel identification
abstract
The semi-blind MIMO channel identification problem is modelled as a stochastic maximum likelihood estimation problem and an iterative method, called information geometric identification (IGID), for channel identification and tracking is presented. The method is developed based on the results from information geometry; specifically, the alternating projections theorem first proved by I. Csiszar and G. Tusnady (see Statistics and Decisions, Suppl. Issue, no.1, p.205-37, 1984). It is demonstrated that the proposed method has similar performance compared to a recently reported method based on the expectation maximization (EM) algorithm (Aldana, C.H. and Cioffi, J., IEEE Int. Conf. on Commun., 2001). Since the IGID method has an analytical solution, the proposed algorithm can be implemented much faster, while having a similar performance. The method can be considered as a generalization of all the methods developed based on the EM algorithm.
Amin Zia, James P. Reilly, Shahram Shirani
ICASSP (4)3
2004 Information geometric approach to channel identification: a comparison with EM-MCMC algorithm
abstract
After reviewing the information geometric channel identification algorithm (IGID) (A. Zia et al., 2003), the application of the algorithm for semi-blind identification of the MIMO channel with Gaussian input sources is discussed. The method is developed based on the results from information geometry; specifically, the alternating projections theorem first proved by Csiszar and G. Tusnady (1984) which provides an iterative method for minimizing the distance between two sets of probability distributions. Also, an EM-type identification algorithm (EM-MCMC) for which the necessary expectation computations are performed using Markov-chain Monte-Carlo (MCMC) method is introduced. The comparative analysis of channel identification using two methods for MIMO systems with ISI-free flat-fading channels is given. It is shown that the IGID method has a similar performance while benefiting from an analytical solution. Thus, complex multidimensional integrations usually necessary in similar EM-type methods are avoided. This characteristic provides very fast computation times relative to previous EM-type algorithms.
Amin Zia, James P. Reilly, Shahram Shirani
ICC3
2003 IMAGE FOVEATION BASED ON VECTOR QUANTIZATION
abstract
Summary form only given. The perceptual resolution of vision is greatly space variant and is highest at the point of fixation and decreases rapidly away from this point. Novel unstructured and structured vector quantization (VQ) schemes are proposed to take advantage of this property of the human visual system (HVS) by providing the best image quality around the fixation point. As a foveation technique, the unstructured VQ does not enjoy progressive transmission in the sense that any request of improvement in the ROI is always replied by complete retransmission of image vectors with a higher resolution VQ scheme. The structured VQ, which is based on a residual vector quantizer (RVQ), yields an embedded progressive bit stream. The main idea for the RVQ foveation is to use more quantization stages for image vectors closer to the fixation point. This method allows the receiver to change its fixation point without any waste of transmitted information and to have multiple fixation points. This quantization strategy enables gradual image resolution change from the ROI to the background.
Abbas Ebrahimi-Moghadam, Shahram Shirani
DCC2
2003 SCO link sharing in Bluetooth voice access networks
Terry Todd 0001, Shahram Shirani
J. Parallel Distributed Comput.3
2003 Application of nonlinear pre- and post-processing in low bit rate, error resilient image communication
Shahram Shirani, Ali Jerbi
Signal Process. Image Commun.1
2002 Low bit rate, error resilient image communication using nonlinear pre and post-processing and progressive image transmission
abstract
The low bit rate, error resilience image coding method proposed here takes into account the content of an image and yields the least amount of degradation, caused by data loss, in those areas of the image which are of greater interest. This is achieved by employing a non-linear geometrical transform to add redundancy mainly to the region of interest (ROI). The nonlinear transform shifts most of the high frequency components of the image inside ROI to lower frequencies therefore, most of the high frequency components can be discarded. This is achieved effectively by employing the progressive mode of JPEG for encoding the nonlinear transformed image. Simulations show that this approach has acceptable performance. Moreover, the method proposed can be implemented through pre- and post-processing of the image data, without modification to the source codecs (e.g., JPEG).
Shahram Shirani
ICIP (3)1
2001 Standard-compliant multiple description video coding
abstract
We address the problem of robust video transmission over unreliable networks. Our approach employs the principle of multiple-descriptions, through pre- and post-processing of the video data, without modification to the source or channel codecs. We employ oversampling to add redundancy to the original video data followed by a decomposition of the oversampled video frames into "equal" sub-images which can be coded and transmitted over separate channels. Simulations using two descriptions show that this approach maintains excellent reconstructed video quality when only one description is received.
Shahram Shirani, Michael Gallant, Faouzi Kossentini
ICIP (1)1
2001 A robust multimedia watermarking technique using Zernike transform
abstract
A new watermarking method using rotation-invariant Zernike moments is introduced. The watermark signal is embedded in the Zernike moments of the input image. The watermarked image does not show any quality degradation. Tests shows that this method is robust to additive noise, JPEG compression and rotation.
Masoud Farzam, Shahram Shirani
MMSP2
2000 Error concealment for MPEG-4 video communication in an error prone environment
abstract
We propose an error concealment method for shape information in the MPEG-4 coded video sequences. A maximum a posteriori estimator, which employs an adaptive Markov random field, is used to restore the missing shape information. The proposed concealment method successfully reconstructs the missing shape data, with good computation-performance tradeoffs.
Shahram Shirani, Berna Erol, Faouzi Kossentini
ICASSP1
2000 An Efficient, Similarity-Based Error Concealment Method for Block-Based Coded Images
abstract
We propose an efficient, similarity-based error concealment method for block-based coded images. In a hierarchical matching procedure, the image is first searched at a lower resolution to find the best match of a layer of pixels around a missing block. Then, the search is performed on the full resolution image. A fast search algorithm which follows a diamond-shaped search area is employed in both resolutions. The missing block is replaced with the block connected to the layer that yields the best match. Moreover, the proposed matching criterion takes into account the geometrical structure extracted from the surrounding pixels of a lost block. The fast search matching method needs significantly less number of computations required by a full search matching method and achieves almost the same reconstruction quality.
Ismaeil R. Ismaeil, Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
ICIP2
2000 Optimal mode selection and synchronization for robust video communications over error-prone networks
abstract
We describe an effective method for increasing error resilience of video transmission over bit error prone networks. Rate-distortion optimized mode selection and synchronization marker insertion algorithms are introduced. The resulting video communication system takes into account the channel condition and the error concealment method used by the decoder, to optimize video coding mode selection and placement of synchronization markers in the compressed bit stream. The effects of mismatch between the parameters used by the encoder and the parameters associated with the actual channel condition and the decoder error concealment method are evaluated. Results for the binary symmetric channel and wideband code division multiple access mobile network models are presented in order to illustrate the advantages of the proposed method.
Guy Côté, Shahram Shirani, Faouzi Kossentini
IEEE J. Sel. Areas Commun.2
2000 A concealment method for video communications in an error-prone environment
abstract
In this paper, we propose a two-stage error-concealment method for block-based compressed video which was transmitted in an error-prone environment. In the first stage, we obtain initial estimates of the missing blocks. If the motion vectors associated with the missing blocks are available, motion compensation is used to provide good estimates. Otherwise, a novel algorithm which preserves image continuity is used to estimate the blocks. In the second stage, a maximum a posteriori (MAP) estimator, which employs an adaptive Markov random field (MRF) as the image a priori model is used to improve the video reconstruction quality. The adaptive model enables the estimation to incorporate information embedded not only in the immediate neighborhood pixels but also in a wider neighborhood into the reconstruction procedure without increasing the order of the MRF model. The proposed concealment method achieves very good computation-performance tradeoffs, as demonstrated via experimental results.
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
IEEE J. Sel. Areas Commun.1
2000 Reconstruction of baseline JPEG coded images in error prone environments
abstract
In this paper, a two-stage method for the reconstruction of missing data in the transmission of baseline JPEG coded images in error prone environments is proposed. In the first stage, we estimate the values of the missing DC coefficients. As effects of errors in estimating the missing DC values will appear as a number of stripes across the image, a technique for removing such stripes is also developed. In the second stage, the data of missing blocks is reconstructed by exploiting the correlation between adjacent blocks. Simulation results intricate that our reconstruction method performs very well. The two key contributions of our method are that it does not assume nondifferential encoding of the DC coefficients, and that it performs well in the reconstruction of diagonal edges.
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
IEEE Trans. Image Process.1
2000 A Concealment Method for Shape Information in MPEG-4 Coded Video Sequences
abstract
We propose a new method for error concealment of shape information in MPEG-4 video bit streams that are transmitted over error prone channels. The proposed method employs a MAP estimator with a Markov random field (MRF) as the image a priori model. The MRF is designed for binary shape information and its parameters are adapted based on the information of neighboring blocks. Our experimental results show that the proposed concealment method restores missing shape blocks with high accuracy. Compared to the median filtering method, our method restores 20% more missing shape data, with a much greater subjective improvement. The proposed algorithm requires a relatively small number of integer multiplications and additions and simple logic operations, making it suitable for real-time implementations.
Shahram Shirani, Berna Erol, Faouzi Kossentini
IEEE Trans. Multim.1
1999 An adaptive Markov random field based error concealment method for video communication in an error prone environment
abstract
Loss of coded data during its transmission can affect a decoded video sequence to a large extent, making concealment of errors caused by data loss a serious issue. Previous work in spatial error concealment exploiting MRF models used a single pixel wide region around the erroneous area to achieve a reconstruction based on an optimality measure. This practically restricts the amount of available information that is used in a concealment procedure to a small region around the missing area. Incorporating more pixels usually means a higher order model and this is expensive as the complexity grows exponentially with the order of the MRF model. Using previously proposed approaches, the damaged area is reconstructed fairly well in very low frequency portions of the image. However, the reconstruction process yields blurry results with a significant loss of details in high frequency, or edge portions of the image. In our proposed approach, a MRF is used as the image a priori model. More available information is incorporated in the reconstruction procedure not by increasing the order of the model but instead by adaptively adjusting the model parameters. Adaptation is done based on the image characteristics determined in a large region around the damaged area. Thus, the reconstruction procedure can make use of information embedded in not only immediate neighborhood pixels but also in a wider neighborhood without a dramatic increase in computational complexity. The proposed method outperforms the previous methods in the reconstruction of missing edges.
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
ICASSP1
1999 Robust H263 Video Communication over Mobile Channels
abstract
In this paper, we describe an effective method for increasing error resilience of video transmission over bit error prone networks. Rate-distortion optimized mode selection and synchronization marker insertion algorithms are introduced. The resulting video communication system takes into account the channel condition and the error concealment method used by the decoder, to optimize video coding mode selection and placement of synchronization markers in the compressed bit stream. Results for the binary symmetric channel and wideband code division multiple access mobile network models are presented in order to illustrate the advantages of the proposed method.
Guy Côté, Shahram Shirani, Faouzi Kossentini
ICIP (2)2
1999 Content-Based Retrieval of Video Sequences Under Partial Occlusion
abstract
In this paper, we present a spatio-temporal segmentation method for content-based video retrieval. Our method is effective in the presence of partial occlusion or when new areas become exposed. In addition, periodic spatial segmentation is not performed. Rather, it is invoked as needed depending on whether or not new regions have been detected. A discussion of how this method can be used for content-based retrieval is also provided. Finally, experimental results that demonstrate the performance of the proposed spatio-temporal segmentation method are presented.
Shahram Shirani, Ali Jerbi, Faouzi Kossentini, Rabab K. Ward, Q. M. Jonathan Wu
ICIP (3)1
1999 On RD optimized progressive image coding using JPEG
abstract
Among the many different modes of operations allowed in the current JPEG standard, the sequential and progressive modes are the most widely used. While the sequential JPEG mode yields essentially the same level of compression performance for most encoder implementations, the performance of progressive JPEG depends highly upon the designed encoder structure. This is due to the flexibility the standard leaves open in designing progressive JPEG encoders. In this work, a rate-distortion (RD) optimized JPEG compliant progressive encoder is presented that produces a sequence of scans, ordered in terms of decreasing importance. Our encoder outperforms an optimized sequential JPEG encoder in terms of compression efficiency, substantially at low and high bit rates. Moreover, unlike existing JPEG compliant encoders, our encoder can achieve precise rate/distortion control. Substantially better compression performance and precise rate control, provided by our progressive JPEG compliant encoding algorithm, are two highly desired features currently sought for the emerging JPEG-2000 standard.
Jaehan In, Shahram Shirani, Faouzi Kossentini
IEEE Trans. Image Process.2
1998 JPEG compliant efficient progressive image coding
abstract
Among the different modes of operations allowed in the current JPEG standard, the sequential and progressive modes are the most widely used. While the sequential JPEG mode yields essentially the same level of compression performance for most encoder implementations, the performance of progressive JPEG depends highly upon the designed encoder structure. This is due to the flexibility the standard leaves open in designing progressive JPEG encoders. In this paper, a rate-distortion optimized JPEG compliant progressive encoder is presented that produces a sequence of bit scans, ordered in terms of decreasing importance. Our encoder outperforms a baseline sequential JPEG encoder in terms of compression, significantly at medium bit rates, and substantially at low and high bit rates. Moreover, unlike baseline JPEG encoders, ours can achieve precise rate/distortion control. Good rate-distortion performance at low bit rates and precise rate control, provided by our JPEG compliant progressive encoder, are two highly desired features currently sought for JPEG-2000.
Jaehan In, Shahram Shirani, Faouzi Kossentini
ICASSP2
1998 Reconstruction of Motion Vector Missing Macroblocks in263 Encoded Video Transmission over Lossy Networks
Shahram Shirani, Faouzi Kossentini, Rabab K. Ward
ICIP (3)1