Mark R. Pickering

dblp:75/4744 · also Mark Pickering · DBLP profile ↗
← Back
97ranked-venue papers
10as first author
7since 2021 · last 2023
0000-0001-6736-3859ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 9 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 2 since 2021Artificial intelligence and machine learning · 2Security and privacy · 2Systems, architecture and hardware · 1Computer networks · 1Theory of computation · 1
YearPublicationVenuePosition
2023 A Two-Step Discrete Cosine Basis Oriented Motion Modeling Approach for Enhanced Motion Compensation
abstract
Video coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. Modern video coding systems are block-based wherein commonality modeling is carried out only from the perspective of the block that need be coded next. In this work, we argue for a commonality modeling approach that can provide a seamless blending between global and local homogeneity information in terms of motion. For this purpose, at first a prediction of the current frame, the frame that need be coded, is generated by performing a two-step discrete cosine basis oriented (DCO) motion modeling. The DCO motion model is employed rather than traditional translational or affine motion model since it has the ability to efficiently model complex motion fields by providing a smooth and sparse representation. Moreover, the proposed two-step motion modeling approach can yield better motion compensation at a reduced computational complexity since an informed guess is designed for initializing the motion search procedure. After that the current frame is partitioned into rectangular regions and the conformance of these regions to the learned motion model is investigated. Depending on the non-conformance to the estimated global motion model, an additional DCO motion model is introduced to increase the local motion homogeneity. In this way, the proposed approach generates a motion compensated prediction of the current frame through the minimization of both global and local motion commonality. Experimental results show an improved rate-distortion performance of a reference high efficiency video coding (HEVC) encoder, specifically up to around 9% savings in bit rate, that employs the DCO prediction frame as a reference frame for encoding the current frame. When compared to the more recent video coding standard, the versatile video coding (VVC) encoder, a bit rate savings of 2.37% is reported.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering
IEEE Trans. Image Process.3
2022 An Edge Aware Motion Modeling Technique Leveraging on the Discrete Cosine Basis Oriented Motion Model and Frame Super Resolution
abstract
To capture motion homogeneity between successive frames, the edge position difference (EPD) measure based motion modeling (EPD-MM) has shown good motion compensation capabilities. The EPD-MM technique is underpinned by the fact that from one frame to next, edges map to edges and such mapping can be captured by an appropriate motion model. An example of such a motion model is the discrete cosine basis oriented (DCO) motion model, which can capture complex motion and has a smooth and sparse representation. However, for higher resolution video sequences, the baseline EPD-MM approach equipped with the DCO motion model, may fail to approximate the underlying motion field accurately. This is due to the difficulty in fitting motion model parameters by incorporating significantly large number of moving edge pixels. Observing the fact that in lower resolution version of the current frame$C$, the same scene structure is present although scaled down moving objects contain smaller number of edge pixels; in this paper we propose to carry out the EPD-MM technique, augmented by the DCO motion model, over lower resolution form of$C$. The resultant edge motion compensated prediction is then upsampled back to the original resolution of$C$, employing single image super resolution (SISR) technique. Experimental results show an improved prediction PSNR of 1.85 dB, on average, from the proposed approach compared to that of the baseline EPD-MM. Moreover, if this predicted frame is employed as an additional reference frame to encode$C$, bit rate savings of up to 7.90% is achievable over a HEVC reference.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering, Andrew J. Lambert
DCC3
2022 An Enhanced Video Coding Technique Leveraging on Edge Aware Motion Modeling and Frame Super Resolution
abstract
To capture motion homogeneity between successive frames, the edge position difference (EPD) measure based motion modeling (EPD-MM) has shown good motion compensation capabilities. The EPD-MM technique is underpinned by the fact that from one frame to next, edges map to edges and such mapping can be captured by an appropriate motion model. However, for higher resolution video sequences, the baseline EPD-MM approach, may fail to approximate the underlying motion field accurately. This is due to the difficulty in fitting motion model parameters by incorporating significantly large number of moving edge pixels. Observing the fact that in lower resolution version of the current frame$C$, the same scene structure is present although scaled down moving objects contain smaller number of edge pixels; in this paper we propose to carry out the EPD-MM over lower resolution form of$C$. The resultant edge motion compensated prediction is then upsampled back to the original resolution of$C$employing single image super resolution (SISR) technique. Experimental results show an improved prediction PSNR of 1.71 dB from the proposed approach compared to that of the baseline EPD-MM. Moreover, if this predicted frame is employed as an additional reference frame to encode$C$, bit rate savings of up to 8.48% is achievable over a HEVC reference and 0.2% is achievable over a VVC reference.
Ashek Ahmmed, Mark R. Pickering, Andrew J. Lambert
MMSP2
2022 Discrete Cosine Basis Oriented Motion Modeling With Cuboidal Applicability Regions For Versatile Video Coding
abstract
The relentless expansion of video based applications is underpinned by video coding technologies. The latest video coding standard i.e. versatile video coding (VVC) can provide superior compression performance than its predecessors. In this regard, motion modeling plays a central role. Experimental results showed that the discrete cosine basis oriented motion model can describe complex motion better than an affine motion model, adopted in the VVC. Hence, in this paper we propose to augment the VVC motion modeling technique with a set of discrete cosine basis oriented motion models and the applicability region of each such motion model is determined by non-overlapping rectangular regions, known as cuboids. Experimental results show a bit rate savings of up to 2.37% is achievable with respect to a VVC reference.
Ashek Ahmmed, Wassim Hamidouche, Andrew J. Lambert, Mark R. Pickering, M. Manzur Murshed
PCS4
2022 Dynamic Mesh Commonality Modeling Using the Cuboidal Partitioning
abstract
For 3D object representation, volumetric contents like meshes and point clouds provide suitable formats. However, a dynamic mesh sequence may require significantly large amount of data because it consists of information that varies with time. Hence, for the facilitation of storage and transmission of such content, efficient compression technologies are required. MPEG has started standardization activities aiming to develop a mesh compression standard that would be able to handle dynamic meshes with time varying connectivity information and time varying attribute maps. The attribute maps are features associated with the mesh surface and stored as 2D images/videos. In this paper, we propose to capture the commonality information in the dynamic mesh attribute maps using the cuboidal partitioning algorithm. This algorithm is capable of modeling both the global and local commonality within an image in a compact and computationally efficient way. Experimental results show that the proposed approach can outperform the anchor HEVC codec, suggested by MPEG to encode such sequences, with a bit rate savings of up to 3.66%.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, Mark R. Pickering
VCIP4
2021 Dynamic Point Cloud Texture Video Compression using the Edge Position Difference Oriented Motion Model
abstract
Immersive media representation format based on point clouds has underpinned significant opportunities for extended reality applications. Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression (V-PCC) technique projects a dynamic point cloud into geometry and texture video sequences. The projected texture video is then coded using modern video coding standard like HEVC. Since the properties of projected texture video frames are different from traditional video frames, HEVC-based commonality modeling can be inefficient. An improved commonality modeling technique is proposed that employs edge position difference oriented motion model. Experimental results show that the proposed commonality modeling technique can yield savings in bit rate of up to 3.15% over the V-PCC HEVC reference encoder.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering
DCC3
2021 Camcording-Resistant Forensic Watermarking Fallback System Using Secondary Watermark Signal
abstract
Forensic watermarking is used to track down digital pirates after they illegally redistribute video content. Although existing algorithms often resist common signal processing attacks, they are not always robust against camcording attacks. As a solution in the state of the art, registration methods are used to align the attacked video to the original one. However, watermark detection still fails when the quality is sufficiently decreased or when exposed to targeted attacks. Therefore, this paper proposes a novel fallback system that aims to detect the watermark when traditional methods fail. More concretely, we demonstrate that a primary watermark embedded by a traditional scheme indirectly creates a secondary watermark signal during video encoding. This secondary watermark consists of compression artifacts and is detected by the fallback system. Additionally, the proposed system incorporates video registration to cope with camcording attacks. The experimental results indicate that the fallback system has a striking increase in robustness compared to the existing methods. For example, the observed false-negative rate for targeted attacks improves from 100% to 0%. Moreover, the fallback is camcording resistant even when the traditional method combined with registration is not. In conclusion, the proposed system can be used as a fallback when traditional detection fails.
Hannes Mareen, Martijn Courteaux, Johan De Praeter, Md. Asikuzzaman, Glenn Van Wallendael, Mark R. Pickering, Peter Lambert
IEEE Trans. Circuits Syst. Video Technol.6
2020 Edge Oriented Hierarchical Motion Estimation For Video Coding
abstract
Efficient video compression relies heavily on mitigating the temporal redundancy that exists between successive video frames. This is achieved through effective motion modelling. In conventional video coding standards, the motion of the current frame is modelled from the neighbouring frames using block-based motion estimation techniques. However, as the motion discontinuities are tied to the moving objects in a video frame, the block-based techniques are unable to model the actual motion of individual objects. In this paper, an object-based hierarchical motion estimation and prediction technique for high-efficiency video coding (HEVC) is proposed. We use an edge position difference (EPD) similarity measure, which has the ability to align the largest object in the frames, to estimate the motion of the object in the current frame from the neighbouring one. In other words, it estimates the largest object's motion instead of the whole frame's global motion. The proposed method gradually models all of the objects' motions and establishes a prediction of the current frame. The predicted frame is then exploited as an additional reference frame in the HEVC compression algorithm. Our experimental results demonstrate that our proposed approach achieves a bit rate savings with a peak signal to noise ratio (PSNR) gain over the HEVC standard.
Md. Asikuzzaman, Ashek Ahmmed, Mark R. Pickering, Thomas Sikora
ICIP3
2020 Object-Oriented Motion Estimation using Edge-Based Image Registration
abstract
Video data storage and transmission cost can be reduced by minimizing the temporally redundant information among frames using an appropriate motion-compensated prediction technique. In the current video coding standard, the neighbouring frames are exploited to predict the motion of the current frame using global motion estimation-based approaches. However, the global motion estimation of a frame may not produce the actual motion of individual objects in the frame as each of the objects in a frame usually has its own motion. In this paper, an edge-based motion estimation technique is presented that finds the motion of each object in the frame rather than finding the global motion of that frame. In the proposed method, edge position difference (EPD) similarity measure-based image registration between the two frames is applied to register each object in the frame. A superpixel search is then applied to segment the registered object. Finally, the proposed edge-based image registration technique and Demons algorithm are applied to predict the objects in the current frame. Our experimental analysis demonstrates that the proposed algorithm can estimate the motions of individual objects in the current frame accurately compared to the existing global motion estimation-based approaches.
Md. Asikuzzaman, Deepak Rajamohan, Mark R. Pickering
MMSP3
2019 Urdu-Text: A Dataset and Benchmark for Urdu Text Detection and Recognition in Natural Scenes
abstract
Multi-lingual text in natural scene images conveys useful information and is a fundamental tool for tourists to interact with their environment. Multi-lingual text detection and recognition in natural scenes, therefore, has become a challenging problem for researchers in the last few years. Recently, a large-scale multi-lingual dataset for scene text detection and script identification is published by the ICDAR which, contains scene images with text in six different scripts including Arabic. This paper presents a novel dataset and benchmark for Urdu text in natural scenes. Currently, no dataset for Urdu text in natural scenes is publicly available. Urdu is a type of cursive language, which is derived from Arabic script and uses many similar alphabet characters. Therefore, the proposed dataset could be helpful for multi-lingual text detection, recognition and script identification. The aim of this dataset is to help the research community for algorithm development and evaluation of Urdu text in natural scenes. The Urdu-Text dataset contains 1400 complete scene images and 8200-segmented words. The images in the dataset contain a broad variety of text instances in multi-orientations with small and large font sizes. The dataset contains ground truths in the form of bounding boxes at the word level, the script of the text and the text-transcription. The performance of three deep neural networks is evaluated to measure the robustness of the Urdu-Text dataset.
Asghar Ali, Mark R. Pickering
ICDAR2
2019 Leveraging the Discrete Cosine Basis for Better Motion Modelling in Highly Textured Video Sequences
abstract
Motion modelling plays a central role in video compression. This role is even more critical in highly textured video sequences, whereby a small error can produce large residuals that are costly to compress. While the translational motion model employed by existing coding standards, such as HEVC, is sufficient in most cases, using higher order models is beneficial; for this reason, the upcoming video coding standard, VVC, employs a 4-parameter affine model. In this work, we explore the use of the discrete cosine basis for motion modelling in highly textured video sequences, and show that this is beneficial. In particular, we use a single high-order model to describe a frame's motion; we employ this motion to produce an extra prediction reference, which is added to the HEVC list of references. Experimental results show that a median delta bit rate of 4.44% is achievable over conventional HEVC if this extra reference frame is used in addition to the temporal references offered by HEVC.
Ashek Ahmmed, Aous Thabit Naman, Mark R. Pickering
ICIP3
2019 Urban-Rural Fringe Recognition with the Integration of Optical and Nighttime Lights Data
abstract
Spatial identification of urban-rural fringes (URF) is crucial for monitoring urban sprawl and mapping out urban management planning. This paper proposes an efficient approach for extracting URF, integrating optical and nighttime lights (NTL) data. Results illustrate that the proposed approach is an effective and practical algorithm for URF identification. The findings highlight the potential of combining optical and NTL data for earth observation, which provide opportunities for new applications.
Xiuping Jia, Mark R. Pickering
IGARSS3
2018 Improving Impervious Surface Estimation by Integrating Multispectral and Nighttime Light Images
abstract
The extension of impervious surface serves as a key indicator of urbanization and environmental quality. This paper explores the integrated use of Landsat imagery and nighttime light (NTL) imagery for impervious surface mapping. An optical and NTL based spectral mixture analysis (ON_SMA) at sub-pixel level is proposed where local endmembers are used and adaptive endmember sets are selected based on the information provided by both sensors. Results illustrate that comparing with optical based spectral mixture analysis, the proposed approach yields effective improvement in impervious surface detection due to the contribution of the NTL data. The findings highlight the potential of combining optical and NTL data for earth observation, which provide opportunities for new applications.
Xiuping Jia, Mark R. Pickering, Genyun Sun
IGARSS3
2018 Quantitative Monitoring of Complete Rice Growing Seasons Using Sentinel 2 Time Series Images
abstract
The payload Multispectral Instrument (MSI) on the satellite Sentinel- 2A provides data with strong spectral information, reasonable spatial resolution and good revisit time, which make them suitable for crop monitoring. With the availability of the image data over the complete rice growing season over two consecutive years, spectral time series analysis of rice crops is conducted in this study for the Riverina region of Coleambally, New South Wales, Australia. Vegetation and water indices are adopated to compare the growing patterns of rice over the 2015/2016 and 2016/2017 growing seasons. Rice crops of different varieties are identified and examined. Different seed sowing methods are also compared in terms of their effect on the latter season. The results show that the growth pattern of rice follows a particular trend that can be distinguished from other types of vegetation. This is facilitated by a property unique to rice where water sensitive indices produce a higher reading than vegetation indices during the initial flooding period of the season, after which, the crop growth reverses this and the biomass sensitive index becomes larger. Spectral analysis of rice crops planted via various methods reveals that the drill sowing method did not produce this unique characteristic as the late flooding time results in a very short submersion period for the rice seedlings, which is a valuable finding.
Emma Madigan, Yiqing Guo, Mark R. Pickering, Alex Held, Xiuping Jia
IGARSS3
2018 Object-Based Motion Estimation Using the EPD Similarity Measure
abstract
Effective motion compensated prediction plays a significant role in efficient video compression. Image registration can be used to estimate the motion of the scene in a frame by finding the geometric transformation which automatically aligns reference and target images. In the video coding literature, image registration has been applied to find the global motion in a video frame. However, if the motion of individual objects in a frame is inconsistent across time, the global motion may provide a very inefficient representation of the true motion present in the scene. In this paper we propose a motion estimation algorithm for video coding using a new similarity measure called the edge position difference (EPD). This technique estimates the motion of the individual objects based on matching the edges of objects rather than estimating the motion using the pixel values in the frame. Experimental results demonstrate that the proposed edge-based similarity measure approach achieves superior motion compensated prediction for objects in a scene when compared to the approach which only considers the pixel values of the frame.
Md. Asikuzzaman, Mark R. Pickering
PCS2
2018 An Overview of Digital Video Watermarking
abstract
The illegal distribution of a digital movie is a common and significant threat to the film industry. With the advent of high-speed broadband Internet access, a pirated copy of a digital video can now be easily distributed to a global audience. A possible means of limiting this type of digital theft is digital video watermarking whereby additional information, called a watermark, is embedded in the host video. This watermark can be extracted at the decoder and used to determine whether the video content is watermarked. This paper presents a review of the digital video watermarking techniques in which their applications, challenges, and important properties are discussed, and categorizes them based on the domain in which they embed the watermark. It then provides an overview of a few emerging innovative solutions using watermarks. Protecting a 3D video by watermarking is an emerging area of research. The relevant 3D video watermarking techniques in the literature are classified based on the image-based representations of a 3D video in stereoscopic, depth-image-based rendering, and multi-view video watermarking. We discuss each technique, and then present a survey of the literature. Finally, we provide a summary of this paper and propose some future research directions.
Md. Asikuzzaman, Mark R. Pickering
IEEE Trans. Circuits Syst. Video Technol.2
2017 Modeling and performance evaluation of stealthy false data injection attacks on smart grid in the presence of corrupted measurements
Adnan Anwar, Abdun Naser Mahmood, Mark R. Pickering
J. Comput. Syst. Sci.3
2016 Motion Hint Field with Content Adaptive Motion Model for High Efficiency Video Coding (HEVC)
abstract
Traditional video coding standards employ block-based translational motion modelwhere all the pixels inside the current block are assigned a single motion vector. Thisuniformity of motion within a block assumption does not hold if the block containsa motion discontinuity. To improve the coding gain, modern video codecs partitionblocks around object boundaries into smaller square or rectangular sub-blocks. The prediction residual energy of the current frame is minimized at the expense of increasing the bit rate to code motion data. The inspiration behind motion hints is to move away from this redundant approach of using the motion model to describe object boundaries, since the spatial structure of previously-decoded frames can be exploited to infer appropriate boundaries for the future ones.A motion hint provides a global description of motion over a specific domain and is related to the foreground-background segmentation where the foreground and background motions are the hints. A bi-directional motion hints based coding paradigm was proposed in [1, 2] that carries out segmentation in the reference frames. The segmented foreground and background regions are then mapped (motion compensated) and fused together to generate a prediction for the current frame. In this paper, the motion hint model is tuned according to the motion hint field's complexity for superior motion compensation, where the candidate motion model set is affine, elastic[3].
Ashek Ahmmed, Mark R. Pickering
DCC2
2016 Spectral unmixing for fire smoke detection and removal
abstract
Optical remote sensing images are often contaminated by smoke from forest fires or biomass burning. In this paper, a smoke removal method is proposed based on a spectral unmixing technique. Smoke components are detected first by generating subpixel smoke fraction masks with spectral mixture analysis. The smoke component is then subtracted from each smoky pixel. Finally, the attenuated signal is restored by rescaling the abundances of the endmember classes present in each pixel. The proposed method has high feasibility and is not dependent on other auxiliary data. The experiments have been conducted on an AVIRIS data set and the results show that the proposed method is effective for smoke detection and data correction.
Meng Xu 0002, Xiuping Jia, Mark R. Pickering, Dar A. Roberts
IGARSS3
2016 Homogeneous motion discovery oriented reference frame for high efficiency video coding
abstract
Traditional video coding uses the motion model to approximate geometric boundaries of moving objects where motion discontinuities occur. Motion hints based inter-frame prediction paradigm moves away from this redundant approach and employs an innovative framework consisting of motion hint fields that are continuous and invertible, at least, over their respective domains. However, estimation of motion hint is computationally demanding, in particular for high resolution video sequences. In this paper, we propose to discover motion models and their associated masks over the current frame and then use these models and masks to form a prediction of the current frame. The prediction process is computationally simpler and experimental results show that a savings in bit rate of 2.3% is achievable over standalone HEVC if this predicted frame is used as an additional reference frame.
Ashek Ahmmed, David S. Taubman, Aous Thabit Naman, Mark R. Pickering
PCS4
2016 Cloud Removal Based on Sparse Representation via Multitemporal Dictionary Learning
abstract
Cloud covers, which generally appear in optical remote sensing images, limit the use of collected images in many applications. It is known that removing these cloud effects is a necessary preprocessing step in remote sensing image analysis. In general, auxiliary images need to be used as the reference images to determine the true ground cover underneath cloud-contaminated areas. In this paper, a new cloud removal approach, which is called multitemporal dictionary learning (MDL), is proposed. Dictionaries of the cloudy areas (target data) and the cloud-free areas (reference data) are learned separately in the spectral domain. The removal process is conducted by combining coefficients from the reference image and the dictionary learned from the target image. This method could well recover the data contaminated by thin and thick clouds or cloud shadows. Our experimental results show that the MDL method is effective in removing clouds from both quantitative and qualitative viewpoints.
Meng Xu 0002, Xiuping Jia, Mark R. Pickering, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2016 Thin Cloud Removal Based on Signal Transmission Principles and Spectral Mixture Analysis
abstract
Cloud removal is an important goal for enhancing the utilization of optical remote sensing satellite images. Clouds dynamically affect the signal transmission due to their different shapes, heights, and distribution. In the case of thick opaque clouds, pixel replacement has been commonly adopted. For thin clouds, pixel correction techniques allow the effects of thin clouds to be removed while retaining the remaining information in the contaminated pixels. In this paper, we develop a new method based on signal transmission and spectral mixture analysis for pixel correction which makes use of a cloud removal model that considers not only the additive reflectance from the clouds but also the energy absorption when solar radiation passes through them. Data correction is achieved by subtracting the product of the cloud endmember signature and the cloud abundance and rescaling according to the cloud thickness. The proposed method has no requirement for meteorological data and does not rely on reference images. Our experimental results indicate that the proposed approach is able to perform effective removal of thin clouds in different scenarios.
Meng Xu 0002, Mark R. Pickering, Antonio Plaza, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.2
2016 Light Field Multi-View Video Coding With Two-Directional Parallel Inter-View Prediction
abstract
Light field (LF) technology has been popularly adopted by a wide range of conventional industries. However, one problem when dealing with LFs is the sheer size of data volume. There have been many multi-view video coding (MVC)-based LF video coding methods reported in the literature, aiming at finding the best prediction structure for LF video coding. It is clear that the number of possible prediction structures is unlimited, and it is also observed that the coding bit-rate can be reduced by increasing the number of bi-directionally encoded views in the prediction structure. However, none work has been conducted to analyze the relationship of the prediction structure with its coding performance. In light of this observation, we first design a new LF-MVC prediction structure by extending the inter-view prediction into a two-directional parallel structure. Analytical models for source coding rate and encoding time are developed to analyze their relationships with the prediction structure, and are proven to be well-matched to our experimental results. Experimental evaluation of two LF video sequences demonstrates that the proposed LF-MVC prediction structure can achieve a factor of 26% bit-rate reduction against the conventional MVC prediction structure for an LF video with 5×5 views, and a further 34% bit-rate reduction for an LF video with a larger 10×10 views. Compared with the state-of-the-art MVC-based LF video coding prediction structures in the literature, LF-MVC can achieve the best coding performance, and with its high encoding efficiency, is well suited for deployment in practical LF-based 3D systems.
Eric Wang 0001, Wei Xiang 0001, Mark R. Pickering, Chang Wen Chen
IEEE Trans. Image Process.3
2016 Robust DT CWT-Based DIBR 3D Video Watermarking Using Chrominance Embedding
abstract
The popularity of 3D video is increasing daily due to the availability of low-cost 3D televisions and high-speed Internet access. However, currently the contents of 3D video can be distributed illegally without any protection. For views generated using a depth-image-based rendering technique, not only the left and right views can be distributed as 3D content, but also the center, left, or right views individually as 2D content. As digital video watermarking is a possible way of protecting these views from unauthorized distribution, in this paper, we propose a digital watermarking method for depth-image-based rendered 3D video. In this method, the watermark is embedded in both of the chrominance channels of a YUV representation of the center view using the dual-tree complex wavelet transform. Then, the left and right views are generated from the watermarked center view and depth map using a depth-image based rendering technique. Finally, the watermark can be extracted from the center, left, and right views in a blind fashion without using the original unwatermarked center, left, or right views. This watermark is robust to geometric distortions, such as upscaling, rotation and cropping, downscaling to an arbitrary resolution, and the most common video distortions, including lossy compression and additive noise. Due to the approximate shift invariance characteristic of the dual-tree complex wavelet transform, the technique is robust against distortions in the left and right views generated using depth-image based rendering. The proposed method can also survive baseline distance adjustment and both 2D and 3D camcording.
Md. Asikuzzaman, Md. Jahangir Alam 0005, Andrew J. Lambert, Mark R. Pickering
IEEE Trans. Multim.4
2015 Cloud effects removal via sparse representation
abstract
Optical remote sensing images are often contaminated by the presence of clouds. The development of cloud effect removal techniques can maximize the usefulness of multispectral or hyperspectral images collected in the spectral range from visible to mid infrared. This paper presents a new data reconstruction technique, via dictionary learning and sparse representation, to remove the cloud effects. Dictionaries of the cloudy data (target data) and the cloud free data (reference data) are learned separately in the spectral domain, where each atom represents a fine ground cover component under the two imaging conditions. In this study, it is found that the sparse coefficients of the reference data are the true weightings of each atom, which can be used to replace the cloud affected coefficients to achieve data correction. Experiments were conducted using Landsat 8 OLI data sets downloaded from the USGS website. The testing results show that clouds of various thickness and cloud shadows can be removed effectively using the proposed method.
Meng Xu 0002, Xiuping Jia, Mark R. Pickering
IGARSS3
2015 Motion hints mode for macroblock coding in bi-predictive slices
abstract
Recent advances in motion modelling have largely focused on careful partitioning of motion blocks in the vicinity of object boundaries. The need for such fine partitioning can be avoided by using motion hints which provide a global description of motion over specific domains. Experimental results show that, with a hybrid setting, more than 50% of the motion discontinuity macroblocks are coded using the motion hints mode in low bit rate cases. The use of this mode leads to a gain of prediction PSNR of 1.11 dB, or equivalently 17.05% savings in bit rate, when compared to the H.264/AVC reference and considering both low and high bit rate applications.
Ashek Ahmmed, Md. Jahangir Alam 0005, Aous Thabit Naman, Mark R. Pickering, David S. Taubman
PCS4
2015 A blind and robust video watermarking scheme in the DT CWT and SVD domain
abstract
The piracy of a digital movie is a significant problem for movie studios and producers but can be prevented by digital video watermarking. In existing watermarking algorithms, robustness to several attacks on the watermark has been improved. However, none of these existing techniques are robust to a combination of the common geometric distortions of scaling, rotation, cropping and downscaling in resolution with other attacks such as video compression. In this paper, a blind video watermarking algorithm is proposed where the watermark is embedded in the singular values of the dual-tree complex wavelet transform coefficients of the chrominance channel. As distortion in the chrominance channel is less sensitive to the human eye, the original video quality is maintained. The singular value decomposition is used due to the good stability of its singular values while the approximate shift invariance characteristic of the dual-tree complex wavelet transform ensures robustness to geometric attacks. The proposed scheme is robust to upscaling, rotation, cropping, downscaling to an arbitrary resolution, aspect ratio change, noise addition and H.264/AVC compression.
Md. Asikuzzaman, Md. Jahangir Alam 0005, Mark R. Pickering
PCS3
2015 Speckle reduction and deblurring of ultrasound images using artificial neural network
abstract
Ultrasound (US) imaging is widely used in clinical diagnostics as it is an economical, portable, painless, comparatively safe, and non-invasive real-time tool. However, the image quality of US imaging is severely affected by the presence of speckle noise during the acquisition process. It is essential to achieve speckle-free high resolution US imaging for better clinical diagnosis. In this paper, we propose a speckle and blur reduction algorithm for US imaging based on artificial neural networks (ANNs). Here, speckle noise is modelled as a multiplicative noise following a Rayleigh distribution, whereas blur is modelled as a Gaussian blur function. The noise and blur variances are estimated by a cascade-forward back propagation (CFBP) neural network using a set of intensity and wavelet features of the US image. The estimated noise and blur variances are then used for speckle reduction by solving the inverse Rayleigh function, and for de-blurring, using the Lucy-Richardson algorithm. The proposed approach gives improved results for both qualitative and quantitative measures.
Muhammad Shahin Uddin, Kalyan Kumar Halder, Murat Tahtali, Andrew J. Lambert, Mark R. Pickering
PCS5
2014 A blind watermarking scheme for depth-image-based rendered 3D video using the dual-tree complex wavelet transform
abstract
The amount of unauthorized distribution of 3D video is increasing day by day due to the availability of high speed Internet and low cost 3D TV. Note that, not only both left and right views generated using depth-image-based rendering can be distributed as 3D content but also the centre, left or right view individually as 2D content. Video watermarking is a possible way to protect this type of illegal distribution. In this paper, we propose a digital watermarking method for depth-image-based rendered 3D video to protect each centre, left, and right view. In this method, the watermark is embedded into the centre view using the dual-tree complex wavelet transform. At the receiver, the left and the right views are generated from the centre view and the depth map using depth-image-based rendering. Finally, we extract the watermark from the centre, left and right views in a blind fashion. This scheme is robust to the most common video distortions which include geometric attacks such as scaling, rotation and cropping as well as lossy JPEG compression and additive noise.
Md. Asikuzzaman, Md. Jahangir Alam 0005, Andrew J. Lambert, Mark R. Pickering
ICIP4
2014 A one-bit approach for image registration
abstract
Motion estimation or optic flow computation for automatic navigation and obstacle avoidance programs running on Unmanned Aerial Vehicles (UAVs) is a challenging task. These challenges come from the requirements of real-time processing speed and small light-weight image processing hardware with very limited resources (especially memory space) embedded on the UAVs. Solutions towards both simplifying computation and saving hardware resources have recently received much interest. This paper presents an approach for image registration using binary images which addresses these two requirements. This approach uses translational information between two corresponding patches of binary images to estimate global motion. These low bit-resolution images require a very small amount of memory space to store them and allow simple logic operations such as XOR and AND to be used instead of more complex computations such as subtractions and multiplications.
An Hung Nguyen, Mark R. Pickering, Andrew J. Lambert
ICMV2
2014 Automatic cloud removal for Landsat 8 OLI images using cirrus band
abstract
The detection of cirrus cloud has historically been difficult due to the lack of a 1.375μm wavelength band in earlier Landsat ETM+ imagery. Landsat 8 OLI has addressed this problem by adding a new cirrus band at this wavelength. This paper presents a study of the effectiveness of utilizing the newly available data for this purpose. An image-based method is developed for cirrus cloud contamination correction. The relationship between a visible or infrared band and the cirrus band is estimated via a linear regression using the data in a homogenous land cover area. The key issue of how to automatic identify homogenous background from the cirrus contaminated data is addressed. The images corrected by our method show satisfactory quality.
Meng Xu 0002, Xiuping Jia, Mark R. Pickering
IGARSS3
2014 Robust rigid registration of CT to MRI brain volumes using the SCV similarity measure
abstract
Multi-modal medical image registration is an important processing step for extracting the maximum amount of information from multi-modal medical images. In this paper, to perform image registration of CT and MRI data volumes, we use the sum-of-conditional variance (SCV) similarity measure which utilizes the joint probability distribution of two images and allows Gauss-Newton optimization to be used. We compare the results from experiments on clinical CT and MRI datasets obtained using the SCV similarity measure, the entropy images on sum-of-squared-difference (eSSD) method and the mutual information (MI) approach. Our results indicate that the proposed SCV approach outperforms the eSSD and MI similarity measure approaches.
Mst. Nargis Aktar, Md. Jahangir Alain, Mark R. Pickering
VCIP3
2014 Subspace Detection Using a Mutual Information Measure for Hyperspectral Image Classification
abstract
Finding a subspace which consists of the most informative features for reliable hyperspectral image classification is a challenging task. Feature reduction is often achieved via feature selection and feature extraction techniques. In this letter, a hybrid approach which combines both treatments is proposed. Principal Component Analysis (PCA) is applied as a preprocessing step so that each of the new features is generated from the complete set of the original spectral bands. Feature selection is then performed effectively using a normalized Mutual Information (nMI) measure with two constraints to maximize general relevance and minimize redundancy in the selected subspace. The proposed algorithm (PCA-nMI) is tested on hyperspectral images and the experimental results show that the modifications give significant improvement in terms of classification accuracy.
Md. Ali Hossain, Xiuping Jia, Mark R. Pickering
IEEE Geosci. Remote. Sens. Lett.3
2014 Imperceptible and Robust Blind Video Watermarking Using Chrominance Embedding: A Set of Approaches in the DT CWT Domain
abstract
Illegal distribution of a digital movie is a significant threat to the film industries. With the advent of high-speed broadband Internet access, a pirated copy of a digital video can be easily distributed to a global audience. Digital video watermarking is a possible means of limiting this type of digital distribution. In existing watermarking methods, the watermark is usually embedded into the luminance channel of a video frame, which affects imperceptibility. In addition, none of the existing techniques are robust to the combination of commonly used attacks, such as compression, upscaling, rotation, cropping, downscaling in resolution, frame rate conversion, and camcording. In this paper, we initially propose a basic blind digital video watermarking algorithm, where the watermark is embedded into one level of the dual-tree complex wavelet transform of the chrominance channel to provide high quality watermarked video and extracted using the same key that was used for embedding. This algorithm is robust to compression, upscaling, rotation, and cropping. An extension of this method extracts the watermark from any level(s) of the dual-tree complex wavelet transform depending on the resolution of the downscaled version of the watermarked frame rather than only from the embedding level to survive downscaling to an arbitrary resolution. Finally, the watermark of a frame is extracted from the information of that frame without using the key that was used during watermark embedding to provide robustness to temporal synchronization attacks, such as frame rate conversion. This scheme is also robust to compression, camcording, watermark estimation remodulation, temporal frame averaging, multiple watermark embedding, downscaling in resolution, and other geometric attacks, such as upscaling, rotation, and cropping.
Md. Asikuzzaman, Md. Jahangir Alam 0005, Andrew J. Lambert, Mark R. Pickering
IEEE Trans. Inf. Forensics Secur.4
2013 Incorporating spatial properties in subspace detection
abstract
The aim of this analysis is to develop a subspace detection technique using a hybrid approach which combines nonlinear feature extraction and feature selection for the task of hyperspectral image classification. In the proposed approach Kernel Principal Component Analysis (KPCA) is applied at the first step to generate the new features from the original data. Then pixel based spatial correlation is measured for each of the KPCA images to rank them based on their spatial objects/contents. These KPCA and spatial correlation based ranking scores are combined to obtain an informative subset of features. The experimental analysis conducted on a real hyperspectral image acquired by the AVIRIS sensor shows the advantage of the proposed approach in terms of classification accuracy.
Md. Ali Hossain, Xiuping Jia, Mark R. Pickering
IGARSS3
2013 Motion segmentation initialization strategies for bi-directional inter-frame prediction
abstract
Experimental results and the latest standards have proved that segmentation based video coding systems can outperform the traditional block-based video coding systems. However, this approach requires the simultaneous estimation of both the shape and motion of moving objects in a video scene. In most of the cases neither the shape nor the motion are known initially. Another critical aspect of this tightly-coupled relationship is that inaccurate motion estimation may cause poor segmentation and erroneous segmentation may negatively impact motion estimation. While some of the existing approaches require user intervention and some use clues such as depth, colour or occlusion to separate the foreground from the background, we propose to use motion reliability information for this purpose. This is because the ingredients necessary for the calculation of motion reliability are the by-product of block-based motion estimation and compensation between the reference frames. Therefore, they require very little or no increase in the computational overhead. In this paper, we explore several motion segmentation initialization strategies based on motion reliability. The performances of these initialization approaches are investigated, in terms of the PSNR, for the predicted inter-frames.
Ashek Ahmmed, Rui Xu 0030, Aous Thabit Naman, Md. Jahangir Alam 0005, Mark R. Pickering, David S. Taubman
MMSP5
2013 Motion hints based inter-frame prediction for hybrid video coding
abstract
Experimental results and the latest standards have proved video coding systems with the ability to adapt the size and shape of the motion estimation area to the objects in the scene can outperform the traditional block-based video coding systems. In this paper, a segmentation-based coding strategy that employs bi-directional motion hints for interframe prediction is proposed. The appealing thing about motion hints is that they are continuous and invertible, even though the observed motion field for a frame will be discontinuous and non-invertible. The proposed scheme outperforms the rate-distortion performance of H.264/AVC reference by 1.1 dB and a bit rebate of 26.6% is achieved.
Ashek Ahmmed, Md. Jahangir Alam 0005, Mark R. Pickering, Rui Xu 0030, Aous Thabit Naman, David S. Taubman
PCS3
2013 A blind high definition videowatermarking scheme robust to geometric and temporal synchronization attacks
abstract
Due to the availability of high speed online streaming sites, a pirated copy of a digital video can be easily distributed to a global audience. This paper proposes a digital video watermarking technique based on the dual-tree complex wavelet transform that can protect this pirated digital video content. In this scheme, the watermark is embedded into the chrominance channel of the video frames to provide a high quality watermarked video. The watermark is detectable without reference video content as well as the original watermark which makes this method robust to temporal synchronization attacks such as frame dropping and frame rate conversion. The proposed method is also robust to geometric attacks such as arbitrary downscaling in resolution, rotation, upscaling, and cropping.
Md. Asikuzzaman, Md. Jahangir Alam 0005, Andrew J. Lambert, Mark R. Pickering
VCIP4
2012 Super resolution OF 3D MRI images using a Gaussian scale mixture model constraint
abstract
In multi-slice magnetic resonance imaging (MRI) the resolution in the slice direction is usually reduced to allow faster acquisition times and to reduce the amount of noise in each 2-D slice. In this paper, a novel image super resolution (SR) algorithm is presented that is used to improve the resolution of the 3D MRI volumes in the slice direction. The proposed SR algorithm uses a complex wavelet-based de-blurring approach with a Gaussian scale mixture model sparseness constraint. The algorithm takes several multi-slice volumes of the same anatomical region captured at different angles and combines these low-resolution images together to form a single 3D volume with much higher resolution in the slice direction. Our results show that the 3D volumes reconstructed using this approach have higher quality than volumes produced by the best previously proposed approaches.
Rafiqul Islam 0001, Andrew J. Lambert, Mark R. Pickering
ICASSP3
2012 Modified SIFT for multi-modal remote sensing image registration
abstract
The scale invariant feature transform (SIFT) is a widely used method for image registration and object recognition. The SIFT method is well known for its ability to identify objects at varying scales and rotations among clutter and occlusion with very fast processing time. The application of SIFT on multi-modal remote sensing images for image registration purposes, however, often results in inaccurate and sometimes incorrect matching. Commonly a very large number of feature points are generated from a remote sensing image but a very small number of feature points are matched giving a high false alarm rate. This paper proposes a method containing several modifications to improve the feature matching performance of the SIFT algorithm by adapting it to suit the characteristics of remote sensing images. The proposed method leads to more matching points with a significantly higher rate of correct matches.
Mahmudul Hasan 0002, Mark R. Pickering, Xiuping Jia
IGARSS2
2012 Improved feature selection based on a mutual information measure for hyperspectral image classification
abstract
Hyperspectral images contain a large amount of information which presents a major challenge for efficient classification. In this paper the information content of each spectral band is analyzed and an improved feature selection technique is proposed for the minimization of dependent information while maximizing the relevancy based on normalized mutual information (NMI). Experimental results are provided for comparisons among some relevant and recentmethods for hyperspectral feature selection in terms of their classification accuracy using real hyperspectral images.
Md. Ali Hossain, Xiuping Jia, Mark R. Pickering
IGARSS3
2012 Computationally efficient global motion estimation using a multi-pass image interpolation algorithm
abstract
The computational complexity of motion estimation between video frames for video coding remains a significant challenge even with current computing power. An important recent advance in the development of efficient motion estimation algorithms is the use of image registration in the estimation of global motion parameters for object-based video coding. However, the main disadvantage of this approach is the increased computational complexity required to estimate the parameters which define the more complex motion models. In this paper, we propose a multi-patch based low complexity global motion estimation (GME) algorithm which uses the relatively new Image Interpolation Algorithm (I2A). Experimental results show that our proposed algorithm achieves the same registration accuracy as the standard GME approach but with significantly less iterations required.
Md. Nazmul Haque, Moyuresh Biswas, Mark R. Pickering
PCS3
2012 Video coding using fast geometry-adaptive partitioning and an elastic motion model
Abdullah Al Muhit, Mark R. Pickering, Michael R. Frater, John F. Arnold
J. Vis. Commun. Image Represent.2
2012 A Low-Complexity Image Registration Algorithm for Global Motion Estimation
abstract
An important recent application of image registration is the estimation of global motion parameters for object-based video coding. However, the main disadvantage of standard approaches to global motion estimation (GME) is the increased computational complexity with the higher degree of motion models when compared to block-based local motion estimation approaches. In this paper, we propose a new low complexity GME algorithm. In our proposed algorithm, full-precision images are replaced with 1 bit-per-pixel images which allows many of the arithmetic operations in the standard GME approach to be replaced with logic operations. Experimental results show that our proposed algorithm achieves the same registration accuracy as the standard GME approach but with significantly reduced computational complexity. Our results also demonstrate the superior performance of the proposed algorithm when compared with previously proposed low-complexity GME approaches.
Md. Nazmul Haque, Moyuresh Biswas, Mark R. Pickering, Michael R. Frater
IEEE Trans. Circuits Syst. Video Technol.3
2012 Robust Automatic Registration of Multimodal Satellite Images Using CCRE With Partial Volume Interpolation
abstract
One of the most important steps in data fusion is image registration. Automatic image-to-image registration for images captured by different sensors traditionally requires the use of information-theoretic similarity measures such as mutual information. Recently, a new similarity measure known as cross-cumulative residual entropy (CCRE) has been proposed for multimodal image registration in medical imaging applications. In this paper, we investigate the use of CCRE for multisensor registration of remote sensing imagery. In particular, we investigate the extreme case of registering synthetic aperture radar images to optical images. We also propose a novel extension to the Parzen-window optimization approach proposed by Thévenaz which involves applying partial volume interpolation in the calculation of the gradients of the similarity measure. Our experimental results show that our proposed approach which uses CCRE as the similarity measure and partial volume interpolation in the optimization procedure provides superior performance to other approaches investigated.
Mahmudul Hasan 0002, Mark R. Pickering, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.2
2011 Accurate depth estimation using structured light and passive stereo disparity estimation
abstract
In this paper we propose a new approach to depth estimation which combines structured light and passive stereo disparity estimation techniques to generate accurate disparity maps in both textured and textureless areas of a scene. We project a structured light pattern with adaptive colors onto the scene and simultaneously capture stereo images with a pair of cameras. By matching points in the projected pattern and the left image we first acquire a sparse disparity map. This sparse disparity map is interpolated and used to initialize a passive stereo disparity algorithm to improve disparity accuracy in textured areas. Finally, this map is interpolated in the textureless areas. By comparing our final map with efficient be lief propagation and the initial interpolated disparity map, we show our approach performs better than using passive-only or active-only techniques.
Qiang Li 0003, Moyuresh Biswas, Mark R. Pickering, Michael R. Frater
ICIP3
2011 A new similarity measure for multi-modal image registration
abstract
Multi-modal similarity measures are required to register images of the same object using different sensors. This registration is often required for medical images of the same patient captured using different imaging modalities such as MRI, CT and PET. In this paper, a new multi-modal similarity measure is proposed which is based on calculating the sum-of-conditional variances from the joint histogram of the two images to be registered. The formulation of this new similarity measure allows the standard Gauss-Newton optimization procedure to be used. Our experimental results show that this new approach is more accurate and robust than the most common and best performing alternative and is also more computationally efficient.
Mark R. Pickering
ICIP1
2011 Unsupervised feature extraction based on a mutual information measure for hyperspectral image classification
abstract
Finding the most informative features from high dimensional space for reliable class data modeling is one of the most challenging problems in hyperspectral image classification. The problem can be address using two basic techniques: feature selection and feature extraction. One of the most popular feature extraction methods is Principal Component Analysis (PCA), however its components are not always suitable for classification. In this paper, we present a feature reduction method (MI-PCA) which uses a nonparametric mutual information (MI) measure on the components obtained via PCA. Supervised classification results using a hyperspectral data set confirm that the new MI-PCA technique provides better classification accuracy by selecting more relevant features than when using either PCA or MI on the original data.
Md. Ali Hossain, Mark R. Pickering, Xiuping Jia
IGARSS2
2011 Image segmentation from scale and rotation invariant texture features from the double dyadic dual-tree complex wavelet transform
Edward H. S. Lo, Mark R. Pickering, Michael R. Frater, John F. Arnold
Image Vis. Comput.2
2011 Automatic, robust global motion estimation using clustering
Nafisa Tarannum, Mark R. Pickering, Michael R. Frater
Signal Process. Image Commun.2
2010 Improved H.264-based video coding using an adaptive transform
abstract
In block-based video coding, the Discrete Cosine Transform (DCT) has been adopted for signal decorrelation in state-of-the-art standards. Although the Karhunen Loeve Transform (KLT) is known to achieve optimal energy compaction, it has been reported to offer only moderate compression as the KLT basis functions are source dependent and hence require the transform itself to be coded. This paper describes a technique for prediction-error block coding using the KLT. The proposed method does not require coding of the KLT bases. Instead the basis functions can be derived at the decoder in a manner similar to the encoder. The proposed method is incorporated into a standard H.264 video codec using an adaptive transform selection approach. Our experiments show that the Peak Signal-to-Noise Ratio (PSNR) improvement of up to 0.9 dB is achieved with the proposed technique when compared with the standard H.264 codec.
Moyuresh Biswas, Mark R. Pickering, Michael R. Frater
ICIP2
2010 Regisration of hyperspectral and trichromatic images via cross cumulative residual entropy maximisation
abstract
In this paper we address the problem of image fusion between imagery acquired by trichromatic sensors and hyperspectral imagers. We do this by presenting a method aimed at registering a high-resolution trichromatic image with lower resolution hyperspectral data. The method presented here maps the hyperspectral image into the grayscale image so as to employ the cross cumulative residual entropy for purposes of multimodal registration. We illustrate the utility of our approach by presenting registration results on a set of surveillance image pairs consisting of a set of high-oblique colour and hyperspectral images.
Mahmudul Hasan 0002, Mark R. Pickering, Antonio Robles-Kelly, Jun Zhou 0001, Xiuping Jia
ICIP2
2010 Spatial noise shaping using convex optimization for perceptual image coding
abstract
In this paper we propose a new convex optimization framework for precise spatial noise shaping. The effectiveness of this new technique is demonstrated in the application of perceptual coding of images. A modified JPEG 2000 codec is implemented using the proposed new framework and compared with existing perceptual coding algorithms. Results of subjective tests show that the new framework can provide a significant improvement in bitrate savings compared to the best performing wavelet domain technique. The algorithm allows much more precise control of distortion than existing spatial domain techniques and is fully compliant with part 1 of the JPEG 2000 standard.
Mark R. Pickering, Junyong You, Touradj Ebrahimi, Andrew Perkis
ICIP1
2010 Multi-spectral remote sensing image registration via spatial relationship analysis on sift keypoints
abstract
Multi-sensor image registration is a challenging task in remote sensing. Considering the fact that multi-sensor devices capture the images at different times, multi-spectral image registration is necessary for data fusion of the images. Several conventional methods for image registration suffer from poor performance due to their sensitivity to scale and intensity variation. The scale invariant feature transform (SIFT) is widely used for image registration and object recognition to address these problems. However, directly applying SIFT to remote sensing image registration often results in a very large number of feature points or keypoints but a small number of matching points with a high false alarm rate. We argue that this is due to the fact that spatial information is not considered during the SIFT-based matching process. This paper proposes a method to improve SIFT-based matching by taking advantage of neighborhood information. The proposed method generates more correct matching points as the relative structure in different remote sensing images are almost static.
Mahmudul Hasan 0002, Xiuping Jia, Antonio Robles-Kelly, Jun Zhou 0001, Mark R. Pickering
IGARSS5
2010 Transform-domain super resolution for improved motion-compensated prediction
abstract
This paper introduces a new approach to motion compensated prediction for video compression. A two-way motion compensation approach is utilized in which the motion compensated prediction is generated either from a super resolution mosaic accumulated using a number of previously decoded frames or from a low-resolution reference frame. The super resolution mosaic is constructed using our improved global motion estimation technique. Experimental results show that the two-way motion compensation provides a better coding gain compared with a standard motion compensation approach and moreover, in the process of super resolution mosaic construction, our robust global motion estimation process outperforms other recent motion estimation techniques when utilizing the same compression framework.
Nafisa Tarannum, Mark R. Pickering, Michael R. Frater, John F. Arnold
ISCAS2
2010 An adaptive low-complexity global motion estimation algorithm
abstract
One important recent application of image registration has been in the estimation of global motion parameters for object-based video coding. A limitation of current global motion estimation approaches is the additional complexity of the gradient-descent optimization that is typically required to calculate the optimal set of global motion parameters. In this paper we propose a new low-complexity algorithm for global motion estimation. The complexity of the proposed algorithm is reduced by performing the majority of the operations in the gradient-descent optimization using logic operations rather than full-precision arithmetic operations. This use of logic operations means that the algorithm can be implemented much more easily in hardware platforms such as field programmable gate arrays (FPGAs). Experimental results show that the execution time for software implementations of the new algorithm is reduced by a factor of almost four when compared to existing fast implementations without any significant loss in registration accuracy.
Md. Nazmul Haque, Moyuresh Biswas, Mark R. Pickering, Michael R. Frater
PCS3
2010 Video Coding Using Elastic Motion Model and Larger Blocks
abstract
Motion-compensated prediction is the key to high-performance video coding. Previous works have explored alternatives to the classical translational motion model in video coding, but the cumulative rate-distortion performance has not been significant enough to see such approaches adopted in mainstream standards. In this paper, we propose a new extended prediction strategy that incorporates non-translational motion prediction. This method uses an elastic motion model with 2-D cosine basis functions to estimate non-translational motion between the blocks. To achieve superior performance, the proposed scheme takes advantage of larger blocks with multi-level partitioning. Experimental results show that this combined framework outperforms the existing techniques, including those available in the recent H.264 standard.
Abdullah Al Muhit, Mark R. Pickering, Michael R. Frater, John F. Arnold
IEEE Trans. Circuits Syst. Video Technol.2
2009 Resilient transmission of motion data in multiple description coding of video
abstract
In the most common form of wavelet video coding, a three-dimensional discrete wavelet transform (DWT) is applied on a group of video frames (GOF). This approach suffers in compression performance because no motion compensation is used. Other approaches to motion compensated wavelet video coding fails to achieve good video quality in packet loss network due to the absence of a resilient transmission method for the motion data. This paper describes a resilient transmission method for the motion parameters. A multiple description (MD) codec is used to encode the video for lossy network. Motion coding and transmission are incorporated in the MD framework. The resulting codec achieves good compression performance in lossless network as well as good resilience in unreliable network. Experimental results also show encouraging performance of the proposed codec compared to state-of-the-art video codec.
Moyuresh Biswas, Michael R. Frater, Mark R. Pickering, John F. Arnold
ICIP3
2009 Motion compensation using geometry and an elastic motion model
abstract
To progress the compression performance of standard video coding algorithms, emerging motion compensation techniques will need to be integrated with the current standard techniques such as those used in the H.264. Higher order motion models, geometry-adaptive partitioning and motion-assisted merging are such techniques that can be considered for the next generation of video coders. In this paper, we examine how geometry information can benefit the use of elastic motion models to accomplish better prediction. Relative complexity issues are also discussed which is important in the standardization process. Experimental results suggest that geometry-adaptive block partitioning can add to the performance of elastic motion models to a certain extent, although the increased complexity is of some concern for real-time coding applications.
Abdullah Al Muhit, Mark R. Pickering, Michael R. Frater
ICIP2
2009 Improved packetization of motion parameters for error resilient transmission of wavelet video
abstract
Resilient delivery of the motion parameters is critical in video transmission over unreliable networks. Several previous approaches to motion compensated wavelet video coding have failed to achieve good video quality due to the absence of a resilient transmission method for the motion vectors (MV). A codec employing a wavelet transform with no motion compensation is also not attractive, since it suffers from poor compression performance. This paper therefore investigates packetization methods for the motion parameters. A multiple description (MD) codec is used to encode the video for lossy networks. The motion packetization techniques are incorporated in the MD framework. We compare the performance of the investigated techniques in both lossless and lossy networks and conclude by proposing the best packet structure for motion transmission. Our proposed approach provides better performance than a recently proposed MD variant of H.264.
Moyuresh Biswas, Michael R. Frater, Mark R. Pickering, John F. Arnold
PCS3
2009 A fast approach for geometry-adaptive block partitioning
abstract
State-of-the-art video compression standards such as H.264 employ tree-structured motion compensation by splitting macro-blocks into fixed square or rectangular sub-blocks. Although, this approach leads to improved compression performance, recent studies have shown that further gain can be achieved via slicing blocks with arbitrary line segments to better match the boundaries between moving objects. However, finding the best partition remains an extremely computationally-intensive task. In this paper, we propose a fast method to identify efficient partitions using a two-step search of the radius and angle of the line segment. Experimental results show that this scheme is able to identify efficient motion boundaries using less iterations than existing techniques while maintaining comparable performance.
Abdullah Al Muhit, Mark R. Pickering, Michael R. Frater
PCS2
2009 Improved Resilience for Video Over Packet Loss Networks With MDC and Optimized Packetization
abstract
We investigate the problem of robust video transmission over lossy packet networks. A resilient video coding framework is important to ensure quality video over these networks. We propose a combination of a rate-distortion optimized multiple description codec and an integrated packetization method to constitute the error resilient codec. Multiple description (MD) coding techniques utilize path-diversity (multiple paths between sender and receiver) in networks by sending the descriptions along different paths. Two different optimization controls for the MD codec are proposed that are suited to variable rates of packet loss for multipath transmission of the MD coded video. An optimized strategy for packetizing the descriptions is also proposed which guarantees that each packet is self-contained and efficient. Simulations done under various packet loss scenarios show the need for two different optimization strategies and also that the developed MD codec achieves significantly improved video quality when compared with similar techniques.
Moyuresh Biswas, Michael R. Frater, John F. Arnold, Mark R. Pickering
IEEE Trans. Circuits Syst. Video Technol.4
2008 Improved resilience for video over packet loss networks with MDC and optimized packetization
abstract
We investigate the problem of robust video transmission over packet loss networks. A resilient video coding framework is important to ensure quality video over these networks. We propose a combination of multiple description coding and an optimized packetization method to constitute the error resilient codec. Multiple description coding technique utilizes path-diversity (multiple paths between sender and receiver) in networks by sending the descriptions along different paths. An optimized strategy for packetizing the descriptions is also proposed which guarantees that each packet is self-contained and efficient. Experimental results show that the proposed method achieves improved video quality over lossy networks.
Moyuresh Biswas, Michael R. Frater, John F. Arnold, Mark R. Pickering
ICIP4
2008 An optimized Multiple Description video codec for lossy packet networks
abstract
The problem of resilient video transmission over lossy packet networks is addressed in this paper. We propose a rate-distortion optimized multiple description (MD) codec. Two different optimization controls of the codec are described that are suited to rates of packet loss, including the case where packets can travel over multiple paths through the network, with each path-dependent packet-loss probabilities. A packetization method optimized to work seamlessly with the proposed MD codec is also proposed. Simulations performed under various packet loss scenarios show the importance of the two optimizations and also that the proposed framework achieves significantly improved video quality when compared with similar techniques.
Moyuresh Biswas, Michael R. Frater, John F. Arnold, Mark R. Pickering
MMSP4
2008 Extended motion compensation using larger blocks and an elastic motion model
abstract
Motion estimation and compensation contribute greatly to video compression efficiency. Previous work has explored the effects of higher-order motion models with standard block sizes on video coding, but the impact on compression efficiency has not been sufficient to see such approaches adopted in mainstream video-compression standards. In this paper, we put forward a new strategy to augment the performance of extended motion models in the perspective of classical block-matching algorithms. To accomplish superior compression, the proposed scheme takes advantage of larger blocks with multi-level partitioning and a higher order elastic motion model. Performance evaluations show that these two features, when combined under a common framework, can significantly outperform the existing techniques including those available in the recent H.264 standard.
Abdullah Al Muhit, Mark R. Pickering, Michael R. Frater, John F. Arnold
MMSP2
2008 An automatic and robust approach for global motion estimation
abstract
In recent years, global motion estimation (GME) has become an important tool in the fields of video coding, video compression and computer vision. Estimating the correct global motion is often more difficult when the video scene contains large foreground objects. Previous approaches to addressing this problem have required the application of algorithm parameters that are sequence dependent or give inconsistent results for different video sequences. In this paper, we propose a fully automatic approach that can successfully estimate global motion in the presence of large foreground objects. The proposed algorithm determines an initial estimate of the foreground pixels and then reduces the effect of the remaining foreground by using a modified Lorentzian estimator. Experimental results show the proposed method produces superior and more consistent performance than some recent approaches for a wide range of sequences.
Nafisa Tarannum, Mark R. Pickering, Michael R. Frater
MMSP2
2008 Generalized framework for reduced precision global motion estimation between digital images
abstract
The efficiency of real-time digital image processing operations has an important impact on the cost and realizability of complex algorithms. Global motion estimation is an example of such a complex algorithm. Most digital image processing is carried out with a precision of 8 bits per pixel, however there has always been interest in low-complexity algorithms. One way of achieving low complexity is through low precision, such as might be achieved by quantization of each pixel to a single bit. Previous approaches to one-bit motion estimation have achieved quantization through a combination of spatial filtering/averaging and threshold setting. In this paper we present a generalized framework for precision reduction. Motivated by this framework, we show that bit-plane selection provides higher performance, with lower complexity, than conventional approaches to quantization.
Michael R. Frater, Elanor Huntington, Mark R. Pickering, John F. Arnold
MMSP4
2008 A Video Watermarking Scheme Based on the Dual-Tree Complex Wavelet Transform
abstract
A watermarking scheme that discourages theater camcorder piracy through the enforcement of playback control is presented. In this method, the video is watermarked so that its display is not permitted if a compliant video player detects the watermark. A watermark that is robust to geometric distortions (rotation, scaling, cropping) and lossy compression is required in order to block access to media content that has been re-recorded with a camera inside a movie theater. We introduce a new video watermarking algorithm for playback control that takes advantage of the properties of the dual-tree complex wavelet transform. This transform offers the advantages of the regular and the complex wavelets (perfect reconstruction, shift invariance, and good directional selectivity). Our method relies on these characteristics to create a watermark that is robust to geometric distortions and lossy compression. The proposed scheme is simple to implement and outperforms comparable methods when tested against geometric distortions.
Lino Coria-Mendoza, Mark R. Pickering, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Inf. Forensics Secur.2
2007 Image Segmentation using Invariant Texture Features from the Double Dyadic Dual-Tree Complex Wavelet Transform
abstract
In this paper we propose a new texture segmentation technique that produces segmentation results which more closely match the manual segmentation that would be performed by a human operator. To perform this type of segmentation, we propose a new texture feature based on the double dyadic dual-tree complex wavelet transform (D3T-CWT) which provides the ability to analyse a signal at and between dyadic scales. This new texture feature is invariant to shift, rotation and scale and hence can group the texture features in a single object (which may have different sizes and orientations) into a single more meaningful segment. When compared with other texture segmentation approaches, the proposed approach provides segmentation results which more closely match the semantically meaningful objects in the scene.
Edward H. S. Lo, Mark R. Pickering, Michael R. Frater, John F. Arnold
ICASSP (1)2
2006 Enhanced Motion Compensation Using Elastic Image Registration
abstract
In this paper we propose a method for extending the standard motion estimation algorithms used in video compression algorithms, to incorporate motion parameters that describe non-translational motion. This method uses an elastic image registration algorithm with two-dimensional discrete cosine basis functions to estimate the motion between block partitions in two video frames. A rate and a distortion are then determined for the elastic motion compensation and these are included in the mode decision calculations of the encoder algorithm. This allows the encoder to trade off bits to describe a more complex motion with bits to describe a larger residual error from a simpler translational motion model. Experimental results show an improvement in compression efficiency of 10-20% can be obtained with this enhanced motion compensation algorithm.
Mark R. Pickering, Michael R. Frater, John F. Arnold
ICIP1
2006 Network traffic demand patterns for video over relative differentiated services networks
James Macnicol, Mark R. Pickering, Michael R. Frater, John F. Arnold
Signal Process. Image Commun.2
2006 Efficient Streaming Packet Video Over Differentiated Services Networks
abstract
We investigate streaming video over Differentiated Services (Diffserv) networks that can provide a number of aggregated traffic classes with increasing quality guarantee. We propose a method to measure the loss impact of a video packet on the quality of the decoded video images. We show how the optimal Quality-of-Service (QoS) mapping from the video packets into a set of traffic classes depends on the loss rates of the classes and the pricing model, and we develop an algorithm to accurately find the optimal QoS mapping. The performance of our algorithm is evaluated through computer simulations and compares favorably to an existing algorithm.
James Macnicol, Mark R. Pickering, Michael R. Frater, John F. Arnold
IEEE Trans. Multim.3
2005 Arobust approach to super-resolution sprite generation
abstract
In this paper, we propose a robust approach to super-resolution static sprite generation from multiple low-resolution images. Considering both short-term and long-term motion influences, a hybrid global motion estimation technique is first presented for sprite generation. An iterative super-resolution reconstruction algorithm is then proposed for the super-resolution sprite construction. This algorithm is robust to outliers and results in a sprite image with high visual quality. In addition, the generated super-resolution sprite can provide better reconstructed images than a conventional low-resolution sprite. Experimental results demonstrate the effectiveness of the proposed methods.
Getian Ye, Mark R. Pickering, Michael R. Frater, John F. Arnold
ICIP (1)2
2005 Efficient multi-image registration with illumination and lens distortion correction
abstract
In this paper, we propose some extensions of an efficient gradient-based image registration method called the inverse compositional algorithm. Specifically, these extensions include cumulative multi-image registration and incorporations of illumination change and lens distortion correction. By combining these extensions, we propose efficient cumulative multi-image registration methods with illumination and lens distortion correction. It is shown that high efficiency can still be achieved for multiple images using the proposed methods. Some experimental results show the efficacy of the proposed methods.
Getian Ye, Mark R. Pickering, Michael R. Frater, John F. Arnold
ICIP (3)2
2005 Tile-boundary artifact reduction using odd tile size and the low-pass first convention
abstract
It is well known that tile-boundary artifacts occur in wavelet-based lossy image coding. However, until now, their cause has not been understood well. In this paper, we show that boundary artifacts are an inescapable consequence of the usual methods used to choose tile size and the type of symmetric extension employed in a wavelet-based image decomposition system. This paper presents a novel method for reducing these tile-boundary artifacts. The method employs odd tile sizes (2N + 1 samples) rather than the conventional even tile sizes (2N samples). It is shown that, for the same bit rate, an image compressed using an odd tile length low-pass first (OTLPF) convention has significantly less boundary artifacts than an image compressed using even tile sizes. The OTLPF convention can also be incorporated into the JPEG 2000 image compression algorithm using extensions defined in Part 2 of this standard.
Jianxin Wei 0002, Mark R. Pickering, Michael R. Frater, John F. Arnold, John A. Boman, Wenjun Zeng 0001
IEEE Trans. Image Process.2
2004 Quality tradeoffs for packet video over differentiated services networks
abstract
Differentiated services networks can only function effectively if an appropriate mix of traffic can be maintained in all service classes using a practical pricing model. Previous work has noted that the same video sequence being distributed to clients who are prepared to pay different amounts for the service results in traffic distributions that vary considerably between those clients. It is shown that this variation can be reduced by trading off coding distortion and the amount of protection given to packets. As a result, a simple pricing model provides stable traffic distributions over a wide range of network conditions and client budgets.
James Macnicol, Mark R. Pickering, Michael R. Frater, John F. Arnold
ICASSP (5)2
2004 Scale and rotation invariant texture features from the dual-tree complex wavelet transform
abstract
Image segmentation can be viewed as the process of classifying regions in a picture into groups with common properties (e.g., texture). A difficulty arising is that a common texture can be classified differently when viewed at different scales and rotated viewpoints. The paper presents a feature vector based on the DT-CWT (dual-tree complex wavelet transform) (Kingsbury, N., Applied and Computational Harmonic Anal., vol.10, p.234-53, 2001) that is invariant to scale and rotation. The promising image segmentation results (without cleaning misclassified regions) demonstrate the suitability of this feature vector in representing texture.
Edward H. S. Lo, Mark R. Pickering, Michael R. Frater, John F. Arnold
ICIP2
2004 A motion confidence measure from phase information
abstract
The reliance of many video processing techniques on motion estimation requires a good motion confidence measure, in order to ensure that unreliable estimations are not accepted into a system. This paper proposes a new measure of motion confidence based on phase information from Fourier transform of a video frame. Using the measure, blocks associated with unreliable motion estimation are detected efficiently and accurately. The merit of the technique is a much improved sensitivity over circumstances where other measures show ambiguity. This improvement is demonstrated through our experimental and analytical results.
Long To, Mark R. Pickering, Michael R. Frater, John F. Arnold
ICIP2
2003 Optimal QoS mapping for streaming video over Differentiated Services networks
abstract
We investigate streaming packet video over relative Differentiated Services (DiffServ) networks which can provide a number of aggregated traffic classes, ordered in a way such that class q+l is better or at least no worse than class q in terms of packet loss. We propose an algorithm for optimal quality of service (QoS) mapping from the video packets to a set of available DiffServ classes. The performance of our algorithm is evaluated through experimental tests and compares favorably to previous works.
Mark R. Pickering, Michael R. Frater, John F. Arnold
ICASSP (5)2
2003 Simultaneous tracking and registration in a multisensor surveillance system
abstract
In this paper, we present a multisensor surveillance system that consists of an optical sensor and an infrared sensor. In this system, a background subtraction method based on zero-order statistics is utilized for the moving object segmentation. Additionally, we propose a generic approach to simultaneous object tracking and multisensor image registration. An efficient face detection system is shown as an application that will have enhanced performance from the registration and fusion of the information from the two sensors. Experimental results show the efficacy of the proposed system.
Getian Ye, Jianxin Wei 0002, Mark R. Pickering, Michael R. Frater, John F. Arnold
ICIP (1)3
2003 A new approach to controlling compression-induced distortion of hyperspectral images
abstract
Images compressed by current lossy compression techniques suffer from distortion that is uniformly distributed spatially and spectrally. We demonstrate that classification accuracy changes as a function of pixel class and spectral location (spectral bands) of the distortion and that the uniform distribution of distortion is therefore not an optimal approach to controlling distortion within an image. We show that a superior approach involves locating areas within the image whose classification accuracies are relatively insensitive to distortion and limiting the application of distortion to these areas. By following this approach, we show that substantial levels of compression-induced distortion can be tolerated without a significant reduction in subsequent classification accuracy. Hyperspectral images are prime candidates for data compression due to their inherent size. Lossy compression algorithms are attractive as they typically provide the most impressive levels of compression (compression ratio) but result in some distortion of the original image. The acceptability of this distortion depends on the end use of the data set that, in the case of hyperspectral images, invariably involves the use of computer-based tools to process the images into a set of classes representing the ground coverage or conditions present in the data. Compression-induced distortion tends to reduce the accuracy associated with the classification process, but the relationship between distortion and classification accuracy varies across different classes of data within the same image and is not particularly predictable. We demonstrate an alternative to the uniform application of distortion during compression that aims to locate spectral and spatial areas within an image where sensitivity to distortion is likely to be reduced. We then restrict the application of the compression-induced distortion to these areas of low sensitivity and show that the subsequent classification accuracies are superior to uniformly distorted images.
R. Ian Faulconbridge, Mark R. Pickering, Michael J. Ryan, Xiuping Jia
IGARSS2
2001 A new method for boundary artefact reduction in JPEG 2000
abstract
It is well known that tile boundary artefacts occur in wavelet-based lossy image coding. However, until now, their cause has not been well understood. We show that boundary artefacts are an intrinsic feature associated with the common method used to choose the tile size and the type of symmetric extension employed in wavelet-based image decomposition. A novel method of reducing tile boundary artefacts is presented. This method has recently been adopted as part of the JPEG 2000 verification model. In this technique, odd tile sizes of 2/sup N/+1 are chosen rather than the conventional even tile sizes of 2/sup N/. We show that, for the same bit-rate, an image compressed using odd tile length and low pass first convention (OTLPF) has significantly fewer boundary artefacts than an image compressed using even tile sizes.
Jianxin Wei 0002, Mark R. Pickering, Michael R. Frater, John A. Boman, John F. Arnold
ICIP (3)2
2001 Efficient spatial-spectral compression of hyperspectral data
abstract
Mean-normalized vector quantization (M-NVQ) has been demonstrated to be the preferred technique for lossless compression of hyperspectral data. In this paper, a jointly optimized spatial M-NVQ/spectral DCT technique is shown to produce compression ratios significantly better than those obtained by the optimized spatial M-NVQ technique alone.
Mark R. Pickering, Michael J. Ryan
IEEE Trans. Geosci. Remote. Sens.1
2000 Compression of Hyperspectral Data Using Vector Quantisation and the Discrete Cosine Transform
abstract
Mean-normalised vector quantization (M-NVQ) has been demonstrated to be the preferred vector quantization technique for application to the lossless compression of hyperspectral data. This work optimises M-NVQ parameters for application to lossy compression and slight improvement is shown to be gained by the implementation of spatial and spectral discrete cosine transform (DCT) techniques for coding of the M-NVQ residuals. Much more efficient compression is shown to be obtained by optimising the M-NVQ and DCT techniques simultaneously, rather than sequentially. Optimised spatial M-NVQ/spectral DCT is shown to produce compression ratios of between 1.5 and 2.5 times better than those obtained by the spatial M-NVQ technique alone. Compression ratios of up to 43:1 are achieved without significant loss in classification accuracy.
Mark R. Pickering, Michael J. Ryan
ICIP1
2000 New method for reducing boundary artifacts in block-based wavelet image compression
Jianxin Wei 0002, Mark R. Pickering, Michael R. Frater, John F. Arnold
VCIP2
2000 Error concealment in video coding of arbitrarily shaped objects
Michael R. Frater, W. S. Lee, Mark R. Pickering, John F. Arnold
Signal Process. Image Commun.3
2000 A robust codec for transmission of very low bit-rate video over channels with bursty errors
abstract
We describe a robust codec for the transmission of very low bit-rate video over channels with a variety of errors, including random and bursty bit errors and packet loss. The codec exploits adaptivity to give good performance with a low overhead. By only protecting macroblocks which would otherwise be poorly concealed by the decoder the codec allows adaptive selection of the parts of video to protect. For protection, it uses multiple description codes which indirectly provide frequency-based adaptivity by protecting the more significant DCT coefficients. Simulations show significant improvements in the performance of the codec when compared to codecs which use intra macroblock updating (raster scan and random) at the same overhead. The codec is efficient in its use of bits and has good error resilience properties both objectively and subjectively over a wide range of conditions. Further, transcoding of the received bit stream to the standard H.263 syntax is relatively easy.
Wee Sun Lee, Mark R. Pickering, Michael R. Frater, John F. Arnold
IEEE Trans. Circuits Syst. Video Technol.2
1998 Spatial Temporal Concealment of Lost Blocks in Coded Video
Wee Sun Lee, Michael R. Frater, Mark R. Pickering, John F. Arnold
ICIP (3)3
1998 Error Concealment for Arbitrarily Shaped Video Objects
Wee Sun Lee, Michael R. Frater, Mark R. Pickering, John F. Arnold
ICIP (3)3
1997 Robustness of multiplexing protocols for audio-visual services over wireless networks
abstract
The protocols used to multiplex the various streams in an audio-visual service (such as MPEG 2 systems and ITU-T H.223) have an important impact on the quality of the service. The fact that wireless channels are often associated with high bit-error rates which are often bursty means that the multiplexing protocols should be designed to be robust against such errors. We suggest two departures from traditional practice which can improve significantly the performance of a packet-oriented multiplex for use with audio-visual services: (1) where the packet header is protected against errors using forward error correcting codes, the use of a synchronization codeword (or flag) is often unnecessary, and (2) for channels where errors tend to occur in bursts, the occurrence of errors in the header can be reduced by transmitting the header information twice (at both ends of the packet) with a low level of FEC protection instead of using the same number of bits to obtain an increased level of protection using a single code word. The value of these approaches is confirmed by simulations.
Wee Sun Lee, Michael R. Frater, Mark R. Pickering, John F. Arnold
ICIP (3)3
1997 A diversity-based scheme for reducing error propagation in video
abstract
We describe a robust codec for the transmission of very low bit-rate video over channels with bursty errors. The codec uses a diversity-based method in the form of multiple description codes to reduce the effect of errors in the video bit-stream. To improve the efficiency for transmission of very low bit-rate video, the codec also exploits adaptivity by only protecting macroblocks (using multiple description codes) which would otherwise be poorly concealed by the decoder. Simulations show significant improvements in the performance of the codec when compared to codecs which use intra macroblock updating (raster scan and random) with the same overhead for reducing the effects of errors.
Wee Sun Lee, Michael R. Frater, Mark R. Pickering, John F. Arnold
ICIP (3)3
1997 Error Resilience in Video and Multiplexing Layers for Very Low Bit-Rate Video Coding Systems
abstract
The transmission of audio-visual services on low-bit-rate, wireless telecommunications systems requires the use of coding techniques that are both efficient in their use of bits and robust against errors introduced in transmission. In this paper, we present efficient techniques for improving the error resilience of audio-visual services. These techniques are based on coding simultaneously for synchronization and error protection or detection. We apply the techniques to improve the performance of the multiplexing protocol (which combines the video and audio streams so that they can be transmitted on a single circuit), and also to improve the robustness of the coded video. We show through simulations that the techniques are efficient in their use of bits and effective against bursty errors common in wireless channels. For a simulation of the DECT channel at a bit-error rate of 10/sup -3/, the techniques give an order of magnitude improvement in the probability of lost packets in the multiplexer layer over more conventional techniques. In the video layer, the techniques give an improvement of between 1-2 dB over ITU-T Recommendation H.263. The techniques proposed for the video layer also have the advantage of permitting simple transcoding with bit streams complying with H.263.
Wee Sun Lee, Mark R. Pickering, Michael R. Frater, John F. Arnold
IEEE J. Sel. Areas Commun.2
1997 A VBR rate control algorithm for MPEG-2 video coders with perceptually adaptive quantisation and traffic shaping
Mark R. Pickering, John F. Arnold, Michael C. Cavenor
Signal Process. Image Commun.1
1997 An adaptive search length algorithm for block matching motion estimation
abstract
This paper presents a new fast search algorithm for block matching motion estimation called the adaptive search length (ASL) algorithm. The ASL algorithm adaptively varies the number of positions searched for each block while still maintaining control of the average number of searches per block for each frame. Experimental results show that the peak signal-to-noise ratio (PSNR) of decoded sequences which were coded using the ASL algorithm is within 0.25 dB of the PSNR of decoded sequences which were coded using the full search block matching algorithm. It is also shown that the ASL algorithm requires only 10% of the computations required by the full search algorithm to achieve this level of decoded image quality.
Mark R. Pickering, John F. Arnold, Michael R. Frater
IEEE Trans. Circuits Syst. Video Technol.1
1996 An adaptive block matching algorithm for efficient motion estimation
abstract
A new fast search algorithm for block matching motion estimation is presented in this paper. The new algorithm, called the adaptive search length (ASL) algorithm, allows the maximum number of searches for each block of the image to vary according to the difficulty in finding the optimum motion vector. The PSNR of the decoded images produced by a video coder operating with the ASL algorithm were compared with those produced by a coder operating with the full search block matching algorithm. The results presented show that, for a PSNR of within .25 dB of the full search PSNR, the ASL algorithm requires only 10% of the searches required by the full search algorithm.
Mark R. Pickering, John F. Arnold, Michael R. Frater
ICIP (3)1
1996 An error concealment technique in the spatial frequency domain
Mark R. Pickering, Michael R. Frater, John F. Arnold, M. W. Grigg
Signal Process.1
1994 A perceptually efficient VBR rate control algorithm
abstract
This paper describes a rate control algorithm for a variable bit-rate (VBR) video coder. The algorithm described varies the quantizer step size of the coder according to properties of an image sequence that affect the perception of errors. The algorithm also limits the output bit-rate of the coder without the use of buffers to more efficiently use network bandwidth. It is shown that a VBR encoder using this algorithm will provide decoded image sequences with a consistent perceived quality that is comparable with, or better than, the perceived quality of images coded with a CBR encoder.
Mark R. Pickering, John F. Arnold
IEEE Trans. Image Process.1