Ashek Ahmmed

dblp:137/2372 · DBLP profile ↗
← Back
23ranked-venue papers
22as first author
10since 2021 · last 2023
0000-0002-6750-2342ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 22 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2023 A Two-Step Discrete Cosine Basis Oriented Motion Modeling Approach for Enhanced Motion Compensation
abstract
Video coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. Modern video coding systems are block-based wherein commonality modeling is carried out only from the perspective of the block that need be coded next. In this work, we argue for a commonality modeling approach that can provide a seamless blending between global and local homogeneity information in terms of motion. For this purpose, at first a prediction of the current frame, the frame that need be coded, is generated by performing a two-step discrete cosine basis oriented (DCO) motion modeling. The DCO motion model is employed rather than traditional translational or affine motion model since it has the ability to efficiently model complex motion fields by providing a smooth and sparse representation. Moreover, the proposed two-step motion modeling approach can yield better motion compensation at a reduced computational complexity since an informed guess is designed for initializing the motion search procedure. After that the current frame is partitioned into rectangular regions and the conformance of these regions to the learned motion model is investigated. Depending on the non-conformance to the estimated global motion model, an additional DCO motion model is introduced to increase the local motion homogeneity. In this way, the proposed approach generates a motion compensated prediction of the current frame through the minimization of both global and local motion commonality. Experimental results show an improved rate-distortion performance of a reference high efficiency video coding (HEVC) encoder, specifically up to around 9% savings in bit rate, that employs the DCO prediction frame as a reference frame for encoding the current frame. When compared to the more recent video coding standard, the versatile video coding (VVC) encoder, a bit rate savings of 2.37% is reported.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering
IEEE Trans. Image Process.1
2022 An Edge Aware Motion Modeling Technique Leveraging on the Discrete Cosine Basis Oriented Motion Model and Frame Super Resolution
abstract
To capture motion homogeneity between successive frames, the edge position difference (EPD) measure based motion modeling (EPD-MM) has shown good motion compensation capabilities. The EPD-MM technique is underpinned by the fact that from one frame to next, edges map to edges and such mapping can be captured by an appropriate motion model. An example of such a motion model is the discrete cosine basis oriented (DCO) motion model, which can capture complex motion and has a smooth and sparse representation. However, for higher resolution video sequences, the baseline EPD-MM approach equipped with the DCO motion model, may fail to approximate the underlying motion field accurately. This is due to the difficulty in fitting motion model parameters by incorporating significantly large number of moving edge pixels. Observing the fact that in lower resolution version of the current frame$C$, the same scene structure is present although scaled down moving objects contain smaller number of edge pixels; in this paper we propose to carry out the EPD-MM technique, augmented by the DCO motion model, over lower resolution form of$C$. The resultant edge motion compensated prediction is then upsampled back to the original resolution of$C$, employing single image super resolution (SISR) technique. Experimental results show an improved prediction PSNR of 1.85 dB, on average, from the proposed approach compared to that of the baseline EPD-MM. Moreover, if this predicted frame is employed as an additional reference frame to encode$C$, bit rate savings of up to 7.90% is achievable over a HEVC reference.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering, Andrew J. Lambert
DCC1
2022 An Enhanced Video Coding Technique Leveraging on Edge Aware Motion Modeling and Frame Super Resolution
abstract
To capture motion homogeneity between successive frames, the edge position difference (EPD) measure based motion modeling (EPD-MM) has shown good motion compensation capabilities. The EPD-MM technique is underpinned by the fact that from one frame to next, edges map to edges and such mapping can be captured by an appropriate motion model. However, for higher resolution video sequences, the baseline EPD-MM approach, may fail to approximate the underlying motion field accurately. This is due to the difficulty in fitting motion model parameters by incorporating significantly large number of moving edge pixels. Observing the fact that in lower resolution version of the current frame$C$, the same scene structure is present although scaled down moving objects contain smaller number of edge pixels; in this paper we propose to carry out the EPD-MM over lower resolution form of$C$. The resultant edge motion compensated prediction is then upsampled back to the original resolution of$C$employing single image super resolution (SISR) technique. Experimental results show an improved prediction PSNR of 1.71 dB from the proposed approach compared to that of the baseline EPD-MM. Moreover, if this predicted frame is employed as an additional reference frame to encode$C$, bit rate savings of up to 8.48% is achievable over a HEVC reference and 0.2% is achievable over a VVC reference.
Ashek Ahmmed, Mark R. Pickering, Andrew J. Lambert
MMSP1
2022 Discrete Cosine Basis Oriented Motion Modeling With Cuboidal Applicability Regions For Versatile Video Coding
abstract
The relentless expansion of video based applications is underpinned by video coding technologies. The latest video coding standard i.e. versatile video coding (VVC) can provide superior compression performance than its predecessors. In this regard, motion modeling plays a central role. Experimental results showed that the discrete cosine basis oriented motion model can describe complex motion better than an affine motion model, adopted in the VVC. Hence, in this paper we propose to augment the VVC motion modeling technique with a set of discrete cosine basis oriented motion models and the applicability region of each such motion model is determined by non-overlapping rectangular regions, known as cuboids. Experimental results show a bit rate savings of up to 2.37% is achievable with respect to a VVC reference.
Ashek Ahmmed, Wassim Hamidouche, Andrew J. Lambert, Mark R. Pickering, M. Manzur Murshed
PCS1
2022 Dynamic Mesh Commonality Modeling Using the Cuboidal Partitioning
abstract
For 3D object representation, volumetric contents like meshes and point clouds provide suitable formats. However, a dynamic mesh sequence may require significantly large amount of data because it consists of information that varies with time. Hence, for the facilitation of storage and transmission of such content, efficient compression technologies are required. MPEG has started standardization activities aiming to develop a mesh compression standard that would be able to handle dynamic meshes with time varying connectivity information and time varying attribute maps. The attribute maps are features associated with the mesh surface and stored as 2D images/videos. In this paper, we propose to capture the commonality information in the dynamic mesh attribute maps using the cuboidal partitioning algorithm. This algorithm is capable of modeling both the global and local commonality within an image in a compact and computationally efficient way. Experimental results show that the proposed approach can outperform the anchor HEVC codec, suggested by MPEG to encode such sequences, with a bit rate savings of up to 3.66%.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, Mark R. Pickering
VCIP1
2022 A Commonality Modeling Framework for Enhanced Video Coding Leveraging on the Cuboidal Partitioning Based Representation of Frames
abstract
Video coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. Modern video coding systems are block-based wherein commonality modeling is carried out only from the perspective of the block that need be coded next. In this work, we argue for a commonality modeling approach that can provide a seamless blending between global and local homogeneity information. For this purpose, at first the frame that need be coded, is recursively partitioned into rectangular regions based on the homogeneity information of the entire frame. After that each obtained rectangular region’s feature descriptor is taken to be the average value of all the pixels’ intensities encompassing the region. In this way, the proposed approach generates a coarse representation of the current frame by minimizing both global and local commonality. This coarse frame is computationally simple and has a compact representation. It attempts to preserve important structural properties of the current frame which can be viewed subjectively as well as from improved rate-distortion performance of a reference scalable HEVC coder that employs the coarse frame as a reference frame for encoding the current frame.
Ashek Ahmmed, M. Manzur Murshed, Manoranjan Paul, David S. Taubman
IEEE Trans. Multim.1
2021 Dynamic Point Cloud Texture Video Compression using the Edge Position Difference Oriented Motion Model
abstract
Immersive media representation format based on point clouds has underpinned significant opportunities for extended reality applications. Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression (V-PCC) technique projects a dynamic point cloud into geometry and texture video sequences. The projected texture video is then coded using modern video coding standard like HEVC. Since the properties of projected texture video frames are different from traditional video frames, HEVC-based commonality modeling can be inefficient. An improved commonality modeling technique is proposed that employs edge position difference oriented motion model. Experimental results show that the proposed commonality modeling technique can yield savings in bit rate of up to 3.15% over the V-PCC HEVC reference encoder.
Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering
DCC1
2021 Dynamic Point Cloud Compression Using A Cuboid Oriented Discrete Cosine Based Motion Model
abstract
Immersive media representation format based on point clouds has underpinned significant opportunities for extended reality applications. Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression technique projects a dynamic point cloud into geometry and texture video sequences. The projected texture video is then coded using modern video coding standard like HEVC. Since the properties of projected texture video frames are different from traditional video frames, HEVC-based commonality modeling can be inefficient. An improved commonality modeling technique is proposed that employs discrete cosine basis oriented motion models and the domains of such models are approximated by homogeneous regions called cuboids. Experimental results show that the proposed commonality modeling technique can yield savings in bit rate of up to 4.17%.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
ICASSP1
2021 Human-Machine Collaborative Video Coding Through Cuboidal Partitioning
abstract
Video coding algorithms encode and decode an entire video frame while feature coding techniques only preserve and communicate the most critical information needed for a given application. This is because video coding targets human perception, while feature coding aims for machine vision tasks. Recently, attempts are being made to bridge the gap between these two domains. In this work, we propose a video coding framework by leveraging on to the commonality that exists between human vision and machine vision applications using cuboids. This is because cuboids, estimated rectangular regions over a video frame, are computationally efficient, has a compact representation and object centric. Such properties are already shown to add value to traditional video coding systems. Herein cuboidal feature descriptors are extracted from the current frame and then employed for accomplishing a machine vision task in the form of object detection. Experimental results show that a trained classifier yields superior average precision when equipped with cuboidal features oriented representation of the current test frame. Additionally, this representation costs 7% less in bit rate if the captured frames are need be communicated to a receiver.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
ICIP1
2021 Dynamic Point Cloud Geometry Compression using Cuboid based Commonality Modeling Framework
abstract
Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression (V-PCC) technique projects a dynamic point cloud into geometry and texture video sequences. The projected geometry and texture video frames are then encoded using modern video coding standard like HEVC. However, HEVC encoder is unable to exploit the global commonality that exists within a geometry frame and between successive geometry frames to a greater extent. This is because in HEVC, the current frame partitioning starts from a rigid $64 \times 64$ pixels level without considering the structure of the scene need be coded. In this paper, an improved commonality modeling framework is proposed, by leveraging on cuboid-based frame partitioning, to encode point cloud geometry frames. The associated frame-partitioning scheme is based on statistical properties of the current geometry frame and therefore yields a flexible block partitioning structure composed of cuboids. Additionally, the proposed commonality modeling approach is computationally efficient and has a compact representation. Experimental results show that if the V-PCC reference encoder is augmented by the proposed commonality modeling technique, a bit rate savings of 2.71% and 4.25% are achieved for full body and upper body of human point clouds’ geometry sequences respectively.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
ICIP1
2020 Leveraging Cuboids for Better Motion Modeling in High Efficiency Video Coding
abstract
In conventional video compression systems, motion model is used to approximate the geometry of moving object boundaries. It is possible to relieve motion model from describing discontinuities in the underlying motion field, by incorporating motion hint that can predict the spatial structure of future frames using the structure of reference frames. However, formation of highly accurate motion hint is computationally demanding, in particular for high resolution video sequences. Cuboids, rectangular regions derived using statistical features, attempt to separate out different objects present in the scene; they are computationally efficient and have sparse representation. Leveraging on the advantages of cuboids, in this paper, we propose to discover homogeneous motion regions and their associated motion based on cuboids. Afterwards, the estimated motion models and their domains are applied to form a prediction of the current frame. Experimental results show that a savings in bit rate of 3.96% is achievable over standalone HEVC reference, if this predicted frame is used as an additional reference frame for the current frame.
Ashek Ahmmed, M. Manzur Murshed, Manoranjan Paul
ICASSP1
2020 Edge Oriented Hierarchical Motion Estimation For Video Coding
abstract
Efficient video compression relies heavily on mitigating the temporal redundancy that exists between successive video frames. This is achieved through effective motion modelling. In conventional video coding standards, the motion of the current frame is modelled from the neighbouring frames using block-based motion estimation techniques. However, as the motion discontinuities are tied to the moving objects in a video frame, the block-based techniques are unable to model the actual motion of individual objects. In this paper, an object-based hierarchical motion estimation and prediction technique for high-efficiency video coding (HEVC) is proposed. We use an edge position difference (EPD) similarity measure, which has the ability to align the largest object in the frames, to estimate the motion of the object in the current frame from the neighbouring one. In other words, it estimates the largest object's motion instead of the whole frame's global motion. The proposed method gradually models all of the objects' motions and establishes a prediction of the current frame. The predicted frame is then exploited as an additional reference frame in the HEVC compression algorithm. Our experimental results demonstrate that our proposed approach achieves a bit rate savings with a peak signal to noise ratio (PSNR) gain over the HEVC standard.
Md. Asikuzzaman, Ashek Ahmmed, Mark R. Pickering, Thomas Sikora
ICIP2
2020 A Coarse Representation of Frames Oriented Video Coding By Leveraging Cuboidal Partitioning of Image Data
abstract
Video coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. In this work, we form a coarse representation of the current frame by minimizing commonality within that frame while preserving important structural properties of the frame. The building blocks of this coarse representation are rectangular regions called cuboids, which are computationally simple and has a compact description. Then we propose to employ the coarse frame as an additional source for predictive coding of the current frame. Experimental results show an improvement in bit rate savings over a reference codec for HEVC, with minor increase in the codec computational complexity.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
MMSP1
2019 Leveraging the Discrete Cosine Basis for Better Motion Modelling in Highly Textured Video Sequences
abstract
Motion modelling plays a central role in video compression. This role is even more critical in highly textured video sequences, whereby a small error can produce large residuals that are costly to compress. While the translational motion model employed by existing coding standards, such as HEVC, is sufficient in most cases, using higher order models is beneficial; for this reason, the upcoming video coding standard, VVC, employs a 4-parameter affine model. In this work, we explore the use of the discrete cosine basis for motion modelling in highly textured video sequences, and show that this is beneficial. In particular, we use a single high-order model to describe a frame's motion; we employ this motion to produce an extra prediction reference, which is added to the HEVC list of references. Experimental results show that a median delta bit rate of 4.44% is achievable over conventional HEVC if this extra reference frame is used in addition to the temporal references offered by HEVC.
Ashek Ahmmed, Aous Thabit Naman, Mark R. Pickering
ICIP1
2019 Discrete Cosine Basis Oriented Motion Modeling for Fisheye and 360 Degree Video Coding
abstract
Motion modeling plays a central role in video compression. This role is even more critical in fisheye video sequences since the wide-angle fisheye imagery has special characteristics as in exhibiting radial distortion. While the translational motion model employed by modern video coding standards, such as HEVC, is sufficient in most cases, using higher order models is beneficial; for this reason, the upcoming video coding standard, VVC, employs a 4-parameter affine model. Discrete cosine basis has the ability to efficiently model complex motion fields. In this work, we investigate the motion modeling behaviour of the discrete cosine basis equipped with higher frequency cosine vectors. In particular, the developed discrete cosine basis is used as a single high-order model to describe a fisheye frame's motion; we employ this motion to produce an extra prediction reference, which is added to the HEVC list of references. Experimental results show an increase in delta bit rate, over conventional HEVC, when higher frequency cosine vectors are added in the motion modeling process. Then leveraging on this modified discrete cosine basis, we propose to employ it for predicting the motion in 360 degree video frames because of their resemblance with fisheye images. In this case, a delta bit rate of 2% is achieved, over conventional HEVC.
Ashek Ahmmed, Manoranjan Paul
MMSP1
2019 Discrete Cosine Basis Oriented Homogeneous Motion Discovery for 360-Degree Video Coding
Ashek Ahmmed, Manoranjan Paul
PSIVT1
2018 Enhanced Homogeneous Motion Discovery Oriented Prediction for Key Intermediate Frames
abstract
Conventional video compression systems use motion model to approximate the geometry of moving object boundaries. Motion model can be relieved from describing discontinuities in the underlying motion field, by employing motion hint that exploits the spatial structure of reference frames to infer appropriate boundaries for the future ones. However, estimation of highly accurate motion hint is computationally demanding, in particular for high resolution video sequences. Leveraging on the advantages of homogeneous motion discovery oriented prediction, in this paper, we propose to tune the intra-domain motion uniformity for B-frames as per the frame's reference utility. Experimental results show an improved bit rate savings compared to the approach where no such selective tuning is enforced.
Ashek Ahmmed, Aous Thabit Naman, David S. Taubman
PCS1
2016 Motion Hint Field with Content Adaptive Motion Model for High Efficiency Video Coding (HEVC)
abstract
Traditional video coding standards employ block-based translational motion modelwhere all the pixels inside the current block are assigned a single motion vector. Thisuniformity of motion within a block assumption does not hold if the block containsa motion discontinuity. To improve the coding gain, modern video codecs partitionblocks around object boundaries into smaller square or rectangular sub-blocks. The prediction residual energy of the current frame is minimized at the expense of increasing the bit rate to code motion data. The inspiration behind motion hints is to move away from this redundant approach of using the motion model to describe object boundaries, since the spatial structure of previously-decoded frames can be exploited to infer appropriate boundaries for the future ones.A motion hint provides a global description of motion over a specific domain and is related to the foreground-background segmentation where the foreground and background motions are the hints. A bi-directional motion hints based coding paradigm was proposed in [1, 2] that carries out segmentation in the reference frames. The segmented foreground and background regions are then mapped (motion compensated) and fused together to generate a prediction for the current frame. In this paper, the motion hint model is tuned according to the motion hint field's complexity for superior motion compensation, where the candidate motion model set is affine, elastic[3].
Ashek Ahmmed, Mark R. Pickering
DCC1
2016 Fisheye video coding using elastic motion compensated reference frames
abstract
Fisheye cameras have become extremely popular in applications where the goal is to capture large fields of view with only one camera. However, the wide-angle fisheye imagery has special characteristics that may not be very well suited for modern video codecs that employ block-based translational motion model. This model fails to describe complex deformable motion which is often present in fisheye videos. In this paper, we advocate for the usage of elastic motion model in compensating such a complex motion. The presented design enables the re-use of existing codecs, such as HEVC, without modifications in low-level coding tools. Experimental results show that a savings in bit rate of up to 6.54% is achievable over standalone HEVC if the elastic motion compensated prediction is used as an additional reference frame.
Ashek Ahmmed, Miska M. Hannuksela, Moncef Gabbouj
ICIP1
2016 Homogeneous motion discovery oriented reference frame for high efficiency video coding
abstract
Traditional video coding uses the motion model to approximate geometric boundaries of moving objects where motion discontinuities occur. Motion hints based inter-frame prediction paradigm moves away from this redundant approach and employs an innovative framework consisting of motion hint fields that are continuous and invertible, at least, over their respective domains. However, estimation of motion hint is computationally demanding, in particular for high resolution video sequences. In this paper, we propose to discover motion models and their associated masks over the current frame and then use these models and masks to form a prediction of the current frame. The prediction process is computationally simpler and experimental results show that a savings in bit rate of 2.3% is achievable over standalone HEVC if this predicted frame is used as an additional reference frame.
Ashek Ahmmed, David S. Taubman, Aous Thabit Naman, Mark R. Pickering
PCS1
2015 Motion hints mode for macroblock coding in bi-predictive slices
abstract
Recent advances in motion modelling have largely focused on careful partitioning of motion blocks in the vicinity of object boundaries. The need for such fine partitioning can be avoided by using motion hints which provide a global description of motion over specific domains. Experimental results show that, with a hybrid setting, more than 50% of the motion discontinuity macroblocks are coded using the motion hints mode in low bit rate cases. The use of this mode leads to a gain of prediction PSNR of 1.11 dB, or equivalently 17.05% savings in bit rate, when compared to the H.264/AVC reference and considering both low and high bit rate applications.
Ashek Ahmmed, Md. Jahangir Alam 0005, Aous Thabit Naman, Mark R. Pickering, David S. Taubman
PCS1
2013 Motion segmentation initialization strategies for bi-directional inter-frame prediction
abstract
Experimental results and the latest standards have proved that segmentation based video coding systems can outperform the traditional block-based video coding systems. However, this approach requires the simultaneous estimation of both the shape and motion of moving objects in a video scene. In most of the cases neither the shape nor the motion are known initially. Another critical aspect of this tightly-coupled relationship is that inaccurate motion estimation may cause poor segmentation and erroneous segmentation may negatively impact motion estimation. While some of the existing approaches require user intervention and some use clues such as depth, colour or occlusion to separate the foreground from the background, we propose to use motion reliability information for this purpose. This is because the ingredients necessary for the calculation of motion reliability are the by-product of block-based motion estimation and compensation between the reference frames. Therefore, they require very little or no increase in the computational overhead. In this paper, we explore several motion segmentation initialization strategies based on motion reliability. The performances of these initialization approaches are investigated, in terms of the PSNR, for the predicted inter-frames.
Ashek Ahmmed, Rui Xu 0030, Aous Thabit Naman, Md. Jahangir Alam 0005, Mark R. Pickering, David S. Taubman
MMSP1
2013 Motion hints based inter-frame prediction for hybrid video coding
abstract
Experimental results and the latest standards have proved video coding systems with the ability to adapt the size and shape of the motion estimation area to the objects in the scene can outperform the traditional block-based video coding systems. In this paper, a segmentation-based coding strategy that employs bi-directional motion hints for interframe prediction is proposed. The appealing thing about motion hints is that they are continuous and invertible, even though the observed motion field for a frame will be discontinuous and non-invertible. The proposed scheme outperforms the rate-distortion performance of H.264/AVC reference by 1.1 dB and a bit rebate of 26.6% is achieved.
Ashek Ahmmed, Md. Jahangir Alam 0005, Mark R. Pickering, Rui Xu 0030, Aous Thabit Naman, David S. Taubman
PCS1