Mark R. Pickering

dblp:75/4744 · also Mark Pickering · DBLP profile ↗
← Back
4ranked-venue papers in the field
0as first author
2since 2021 · last 2022
0000-0001-6736-3859ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3Other / Interdisciplinary · 1
YearPublicationVenuePosition
2022 An Edge Aware Motion Modeling Technique Leveraging on the Discrete Cosine Basis Oriented Motion Model and Frame Super Resolution
abstract
To capture motion homogeneity between successive frames, the edge position difference (EPD) measure based motion modeling (EPD-MM) has shown good motion compensation capabilities. The EPD-MM technique is underpinned by the fact that from one frame to next, edges map to edges and such mapping can be captured by an appropriate motion model. An example of such a motion model is the discrete cosine basis oriented (DCO) motion model, which can capture complex motion and has a smooth and sparse representation. However, for higher resolution video sequences, the baseline EPD-MM approach equipped with the DCO motion model, may fail to approximate the underlying motion field accurately. This is due to the difficulty in fitting motion model parameters by incorporating significantly large number of moving edge pixels. Observing the fact that in lower resolution version of the current frame$C$, the same scene structure is present although scaled down moving objects contain smaller number of edge pixels; in this paper we propose to carry out the EPD-MM technique, augmented by the DCO motion model, over lower resolution form of$C$. The resultant edge motion compensated prediction is then upsampled back to the original resolution of$C$, employing single image super resolution (SISR) technique. Experimental results show an improved prediction PSNR of 1.85 dB, on average, from the proposed approach compared to that of the baseline EPD-MM. Moreover, if this predicted frame is employed as an additional reference frame to encode$C$, bit rate savings of up to 7.90% is achievable over a HEVC reference.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering, Andrew J. Lambert
DCC3
2021 Dynamic Point Cloud Texture Video Compression using the Edge Position Difference Oriented Motion Model
abstract
Immersive media representation format based on point clouds has underpinned significant opportunities for extended reality applications. Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression (V-PCC) technique projects a dynamic point cloud into geometry and texture video sequences. The projected texture video is then coded using modern video coding standard like HEVC. Since the properties of projected texture video frames are different from traditional video frames, HEVC-based commonality modeling can be inefficient. An improved commonality modeling technique is proposed that employs edge position difference oriented motion model. Experimental results show that the proposed commonality modeling technique can yield savings in bit rate of up to 3.15% over the V-PCC HEVC reference encoder.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering
DCC3
2019 Urdu-Text: A Dataset and Benchmark for Urdu Text Detection and Recognition in Natural Scenes
abstract
Multi-lingual text in natural scene images conveys useful information and is a fundamental tool for tourists to interact with their environment. Multi-lingual text detection and recognition in natural scenes, therefore, has become a challenging problem for researchers in the last few years. Recently, a large-scale multi-lingual dataset for scene text detection and script identification is published by the ICDAR which, contains scene images with text in six different scripts including Arabic. This paper presents a novel dataset and benchmark for Urdu text in natural scenes. Currently, no dataset for Urdu text in natural scenes is publicly available. Urdu is a type of cursive language, which is derived from Arabic script and uses many similar alphabet characters. Therefore, the proposed dataset could be helpful for multi-lingual text detection, recognition and script identification. The aim of this dataset is to help the research community for algorithm development and evaluation of Urdu text in natural scenes. The Urdu-Text dataset contains 1400 complete scene images and 8200-segmented words. The images in the dataset contain a broad variety of text instances in multi-orientations with small and large font sizes. The dataset contains ground truths in the form of bounding boxes at the word level, the script of the text and the text-transcription. The performance of three deep neural networks is evaluated to measure the robustness of the Urdu-Text dataset.
Asghar Ali, Mark R. Pickering
ICDAR2
2016 Motion Hint Field with Content Adaptive Motion Model for High Efficiency Video Coding (HEVC)
abstract
Traditional video coding standards employ block-based translational motion modelwhere all the pixels inside the current block are assigned a single motion vector. Thisuniformity of motion within a block assumption does not hold if the block containsa motion discontinuity. To improve the coding gain, modern video codecs partitionblocks around object boundaries into smaller square or rectangular sub-blocks. The prediction residual energy of the current frame is minimized at the expense of increasing the bit rate to code motion data. The inspiration behind motion hints is to move away from this redundant approach of using the motion model to describe object boundaries, since the spatial structure of previously-decoded frames can be exploited to infer appropriate boundaries for the future ones.A motion hint provides a global description of motion over a specific domain and is related to the foreground-background segmentation where the foreground and background motions are the hints. A bi-directional motion hints based coding paradigm was proposed in [1, 2] that carries out segmentation in the reference frames. The segmented foreground and background regions are then mapped (motion compensated) and fused together to generate a prediction for the current frame. In this paper, the motion hint model is tuned according to the motion hint field's complexity for superior motion compensation, where the candidate motion model set is affine, elastic[3].
Ashek Ahmmed, Mark R. Pickering
DCC2