VLDB 2026 Research / reviewers in the wild / expert
Manoranjan Paul
dblp:99/774
· DBLP profile ↗
96ranked-venue papers
24as first author
28since 2021 · last 2026
0000-0001-6870-5056ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 22 first-author · 24 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning-Based CU Splitting in VVC: A ConvNeXt Approach for Intra-Mode Acceleration
Md. Zahirul Islam, Asifur Rahman Rahat, Noman Amin, S. M. Faisal, Manoranjan Paul, Zhenni Pan |
ISCAS | 5 |
| 2026 | Semantic First, Bits Later: Predictive Latent Transport for Smart-City Edge Media
Quazi Mamun, Manoranjan Paul |
WoWMoM | 2 |
| 2025 | Policy Gradient-Based Optimal Subset Selection for Few-Shot Vision-Language LearningabstractVision-Language models (VLMs) like Contrastive Language-Image Pre-Training (CLIP) have been extensively adapted for few-shot classification. Most few-shot methods rely on randomly selected samples from the dataset. However, since only a few samples are used, the sample selection process can significantly impact the performance of the downstream classification task. In this work, we propose a reinforcement learning-based policy gradient technique that employs a diversity and informativeness-based reward function to optimise the sample selection process. We evaluate various sample selection techniques based on downstream classification accuracy across three benchmark datasets, where the proposed method demonstrates promising results. Muhammad Khizer Ali, Manoranjan Paul, Anwaar Ulhaq, Muhammad Haris Khan, Quazi Mamun |
ICIP | 2 |
| 2024 | Estimating Soil Organic Carbon from Multispectral Images Using Physics-Informed Neural Networks
James Sargeant, Shyh Wei Teng, M. Manzur Murshed, Manoranjan Paul, David Brennan |
ACCV (7) | 4 |
| 2024 | Block-Wise Compression Of The Quantum Gray-Scale Image Using Lossy Preparation ApproachabstractQuantum computing draws huge attention due to the multidimensional information processing capability of the processor at the same time. The main idea of quantum imaging is to represent pixel intensities and their state using standard quantum circuits. For a medium or larger size image, the state label circuit takes more qubits as well as requires a huge connection leading to circuit complexity. In this work, a block-wise lossy SCMNEQR (state connection modification novel enhanced quantum representation) approach has been proposed to address more qubit connection issues. The average computation time decreased by 99.66% and 7.39% compared to the JPEG and DCT-EFRQI approaches respectively. The experimental results show that it outperforms the DCT-EFRQI (discrete cosine transform- efficient flexible representation of the quantum image) approach in terms of both representation and compression due to the use of block-wise division of transfer coefficient and state label alongside a novel reset gate. Md. Ershadul Haque, Manoranjan Paul |
ICME | 2 |
| 2024 | Impact analysis of recovery cases due to COVID-19 outbreak using deep learning model
Md. Ershadul Haque, Samiul Ul Hoque, Manoranjan Paul, Mahidur R. Sarker, Abdulla Al Suman, Tanvir Ul Huque |
Multim. Tools Appl. | 3 |
| 2024 | Efficient motion modelling with variable-sized blocks from hierarchical cuboidal partitioning
Priyabrata Karmakar, M. Manzur Murshed, Manoranjan Paul, David S. Taubman |
Multim. Tools Appl. | 3 |
| 2023 | A Novel State Connection Strategy for Quantum Computing to Represent and Compress Digital ImagesabstractQuantum image processing draws a lot of attention due to faster data computation and storage compared to classical data processing systems. Converting classical image data into the quantum domain and state label preparation complexity is still a challenging issue. The existing techniques normally connect the pixel values and the state position directly. Recently, the EFRQI (efficient flexible representation of the quantum image) approach uses an auxiliary qubit that connects the pixel-representing qubits to the state position qubits via Toffoli gates to reduce state connection. Due to the twice use of Toffoli gates for each pixel connection still it requires a significant number of bits to connect each pixel value. In this paper, we propose a new SCMFRQI (state connection modification FRQI) approach for further reducing the required bits by modifying the state connection using a reset gate rather than repeating the use of the same Toffoli gate connection as a reset gate. Moreover, unlike other existing methods, we compress images using block-level for further reduction of required qubits. The experimental results confirm that the proposed method outperforms the existing methods in terms of both image representation and compression points of view. Md. Ershadul Haque, Manoranjan Paul, Anwar Ulhaq, Tanmoy Debnath |
ICASSP | 2 |
| 2023 | Dynamic Point Cloud Compression Approach Using Hexahedron SegmentationabstractVideo-based point cloud compression (V-PCC) is the state-of-the-art standard for compressing dynamic point clouds, a newly developed media format. However, real-time V-PCC applications are challenging because of excessive encoding complexity. This approach relies on converting 3D data into 2D frames using patch generation and refinement processes, which is one of the most computationally demanding components of V-PCC. This paper proposes a suitable-sized hexahedron and its projection plane by investigating different sizes of regular hexahedron partitioning and projection planes to convert 3D data to 2D without needing the complex patch generation and refinement processes while preserving more data and accuracy. The experimental results demonstrate that the chosen hexahedron can reduce the patch generation time of the V-PCC significantly. The proposed technique also minimises the point cloud losses, thus improving the image quality, and requires a smaller size of 2D maps than the V-PCC to further reduce the required bits for compression. Faranak Tohidi, Manoranjan Paul |
ICIP | 2 |
| 2023 | You're Not the Boss of me, Algorithm: Increased User Control and Positive Implicit Attitudes Are Related to Greater Adherence to an Algorithmic AidabstractAbstract This study examined whether participants’ adherence to an algorithmic aid was related to the degree of control they were provided at decision point and their attitudes toward new technologies and algorithms. It also tested the influence of control on participants’ subjective reports of task demands whilst using the aid. A total of 159 participants completed an online experiment centred on a simulated forecasting task, which required participants to predict the performance of school students on a standardized mathematics test. For each student, participants also received an algorithm-generated forecast of their score. Participants were randomly assigned to either the ‘full control’ (adjust forecast as much as they wish), ‘moderate control’ (adjust forecast by 30%) or ‘restricted control’ (adjust forecast by 2%) group. Participants then completed an assessment of subjective task load, a measure of their explicit attitudes toward new technologies, demographic and experience items (age, gender and computer literacy) and a novel version of the Go/No-Go Association Task, which tested their implicit attitudes toward algorithms. The results revealed that participants who were provided with more control over the final forecast tended to deviate from it more greatly and reported lower levels of frustration. Furthermore, participants showing more positive implicit attitudes toward algorithms were found to deviate less from the algorithm’s forecasts, irrespective of the degree of control they were given. The findings allude to the importance of users’ control and preexisting attitudes in their acceptance of, and frustration in using a novel algorithmic aid, which may ultimately contribute to their intention to use them in the workplace. These findings can guide system developers and support workplaces implementing expert system technology. Ben W. Morrison, Joshua N. Kelson, Natalie M. V. Morrison, John Michael Innes, Gregory Zelic, Yeslam Al-Saggaf, Manoranjan Paul |
Interact. Comput. | 7 |
| 2023 | Rate-Distortion Modeling for Bit Rate Constrained Point Cloud CompressionabstractAs being one of the main representation formats of 3D real world and well-suited for virtual reality and augmented reality applications, point clouds have gained a lot of popularity. In order to reduce the huge amount of data, a considerable amount of research on point cloud compression has been done. However, given a target bit rate, how to properly choose the color and geometry quantization parameters for compressing point clouds is still an open issue. In this paper, we propose a rate-distortion model based quantization parameter selection scheme for bit rate constrained point cloud compression. Firstly, to overcome the measurement uncertainty in evaluating the distortion of the point clouds, we propose a unified model to combine the geometry distortion and color distortion. In this model, we take into account the correlation between geometry and color variables of point clouds and derive a dimensionless quantity to represent the overall quality degradation. Then, we derive the relationships of overall distortion and bit rate with the quantization parameters. Finally, we formulate the bit rate constrained point cloud compression as a constrained minimization problem using the derived polynomial models and deduce the solution via an iterative numerical method. Experimental results show that the proposed algorithm can achieve optimal decoded point cloud quality at various target bit rates, and substantially outperform the video-rate-distortion model based point cloud compression scheme. Pan Gao 0001, Shengzhou Luo, Manoranjan Paul |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | A Two-Step Discrete Cosine Basis Oriented Motion Modeling Approach for Enhanced Motion CompensationabstractVideo coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. Modern video coding systems are block-based wherein commonality modeling is carried out only from the perspective of the block that need be coded next. In this work, we argue for a commonality modeling approach that can provide a seamless blending between global and local homogeneity information in terms of motion. For this purpose, at first a prediction of the current frame, the frame that need be coded, is generated by performing a two-step discrete cosine basis oriented (DCO) motion modeling. The DCO motion model is employed rather than traditional translational or affine motion model since it has the ability to efficiently model complex motion fields by providing a smooth and sparse representation. Moreover, the proposed two-step motion modeling approach can yield better motion compensation at a reduced computational complexity since an informed guess is designed for initializing the motion search procedure. After that the current frame is partitioned into rectangular regions and the conformance of these regions to the learned motion model is investigated. Depending on the non-conformance to the estimated global motion model, an additional DCO motion model is introduced to increase the local motion homogeneity. In this way, the proposed approach generates a motion compensated prediction of the current frame through the minimization of both global and local motion commonality. Experimental results show an improved rate-distortion performance of a reference high efficiency video coding (HEVC) encoder, specifically up to around 9% savings in bit rate, that employs the DCO prediction frame as a reference frame for encoding the current frame. When compared to the more recent video coding standard, the versatile video coding (VVC) encoder, a bit rate savings of 2.37% is reported. Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering |
IEEE Trans. Image Process. | 2 |
| 2022 | An Edge Aware Motion Modeling Technique Leveraging on the Discrete Cosine Basis Oriented Motion Model and Frame Super ResolutionabstractTo capture motion homogeneity between successive frames, the edge position difference (EPD) measure based motion modeling (EPD-MM) has shown good motion compensation capabilities. The EPD-MM technique is underpinned by the fact that from one frame to next, edges map to edges and such mapping can be captured by an appropriate motion model. An example of such a motion model is the discrete cosine basis oriented (DCO) motion model, which can capture complex motion and has a smooth and sparse representation. However, for higher resolution video sequences, the baseline EPD-MM approach equipped with the DCO motion model, may fail to approximate the underlying motion field accurately. This is due to the difficulty in fitting motion model parameters by incorporating significantly large number of moving edge pixels. Observing the fact that in lower resolution version of the current frame$C$, the same scene structure is present although scaled down moving objects contain smaller number of edge pixels; in this paper we propose to carry out the EPD-MM technique, augmented by the DCO motion model, over lower resolution form of$C$. The resultant edge motion compensated prediction is then upsampled back to the original resolution of$C$, employing single image super resolution (SISR) technique. Experimental results show an improved prediction PSNR of 1.85 dB, on average, from the proposed approach compared to that of the baseline EPD-MM. Moreover, if this predicted frame is employed as an additional reference frame to encode$C$, bit rate savings of up to 7.90% is achievable over a HEVC reference. Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering, Andrew J. Lambert |
DCC | 2 |
| 2022 | Dilated Convolutional Neural Network-Based Deep Reference Picture Generation for Video CompressionabstractMotion estimation and motion compensation are indispensable parts of inter prediction in video coding. Since the motion vector of objects is mostly in fractional pixel units, original reference pictures may not accurately provide a suitable reference for motion compensation. In this paper, we propose a deep reference picture generator which can create a picture that is more relevant to the cur-rent encoding frame, thereby further reducing temporal redundancy and improving video compression efficiency. Inspired by the recent progress of Convolutional Neural Network(CNN), this paper pro-poses to use a dilated CNN to build the generator. Moreover, we insert the generated deep picture into Versatile Video Coding(VVC) as a reference picture and perform a comprehensive set of experiments to evaluate the effectiveness of our network on the latest VVC Test Model–VTM. The experimental results demonstrate that our pro-posed method achieves on average 9.7% bit saving compared with VVC under low-delay P configuration. Haoyue Tian, Pan Gao 0001, Manoranjan Paul |
ICASSP | 4 |
| 2022 | Efficient Scalable 360-degree Video Compression Scheme using 3D Cuboid PartitioningabstractVideo coding techniques minimize spatial and temporal redundancies inherent in video sequences based on non-overlapping block-based image partitioning. Due to depending on the information from already encoded neighboring blocks, these algorithms lack efficient techniques to exploit the overall global redundancies. Compared to the traditional block-based coding, the cuboid coding (2D) framework has been proven to be a more effective method of image compression that exploits global redundancy by considering homogeneous pixel correlation within a frame. In this paper, we improved the idea of 2D cuboid coding to exploit both local and global redundancy from a video sequence by adopting a three-dimensional (3D) cuboid partitioning scheme for SHVC compression improvement of 360-degree videos. The proposed method considers a group of successive frames as a 3D cuboid and recursively partitions it into sub-3D cuboids where static information over a selected GOP share the same cuboid and moving regions share new cuboids with better-defined objects. All the 3D cuboids are then encoded to create a coarse representation of the video stream. Experiments indicate that the proposed framework significantly outperforms its relevant benchmarks, notably by 17.18% (average) in BD-Rate reduction and 0.82 dB in BD-PSNR gain with respect to the standard SHVC codec. Fariha Afsana, Manoranjan Paul, M. Manzur Murshed, David S. Taubman |
ICIP | 2 |
| 2022 | Dynamic Point Cloud Compression with Cross-Sectional Approach
Faranak Tohidi, Manoranjan Paul, Anwaar Ulhaq |
PSIVT | 2 |
| 2022 | Dynamic Mesh Commonality Modeling Using the Cuboidal PartitioningabstractFor 3D object representation, volumetric contents like meshes and point clouds provide suitable formats. However, a dynamic mesh sequence may require significantly large amount of data because it consists of information that varies with time. Hence, for the facilitation of storage and transmission of such content, efficient compression technologies are required. MPEG has started standardization activities aiming to develop a mesh compression standard that would be able to handle dynamic meshes with time varying connectivity information and time varying attribute maps. The attribute maps are features associated with the mesh surface and stored as 2D images/videos. In this paper, we propose to capture the commonality information in the dynamic mesh attribute maps using the cuboidal partitioning algorithm. This algorithm is capable of modeling both the global and local commonality within an image in a compact and computationally efficient way. Experimental results show that the proposed approach can outperform the anchor HEVC codec, suggested by MPEG to encode such sequences, with a bit rate savings of up to 3.66%. Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, Mark R. Pickering |
VCIP | 2 |
| 2022 | Human pose based video compression via forward-referencing using deep learningabstractTo exploit high temporal correlations in video frames of the same scene, the current frame is predicted from the already-encoded reference frames using block-based motion estimation and compensation techniques. While this approach can efficiently exploit the translation motion of the moving objects, it is susceptible to other types of affine motion and object occlusion/deocclusion. Recently, deep learning has been used to model the high-level structure of human pose in specific actions from short videos and then generate virtual frames in future time by predicting the pose using a generative adversarial network (GAN). Therefore, modelling the high-level structure of human pose is able to exploit semantic correlation by predicting human actions and determining its trajectory. Video surveillance applications will benefit as stored “big” surveillance data can be compressed by estimating human pose trajectories and generating future frames through semantic correlation. This paper explores a new way of video coding by modelling human pose from the already-encoded frames and using the generated frame at the current time as an additional forward-referencing frame. It is expected that the proposed approach can overcome the limitations of the traditional backward-referencing frames by predicting the blocks containing the moving objects with lower residuals. Our experimental results show that the proposed approach can achieve on average up to 2.83 dB PSNR gain and 25.93% bitrate savings for high motion video sequences compared to standard video coding. S. M. A. K. Rajin, M. Manzur Murshed, Manoranjan Paul, Shyh Wei Teng, Jiangang Ma |
VCIP | 3 |
| 2022 | Efficient Scalable UHD/360-Video Coding by Exploiting Common Information With Cuboid-Based PartitioningabstractThe scalable extension of High Efficiency Video Coding, SHVC can code Ultra High-Definition (UHD) video, including 360-degree video for various devices to serve a single bitstream with different display resolutions and qualities. To improve the SHVC compression efficiency, this paper proposes a novel intra and inter-frame coding scheme by first separating the common/visually important information and then applying cuboid-based variable size block partitioning and coding process for the common/visually important information in the base layer. In cuboid-based partitioning a video frame is partitioned into arbitrary shaped rectangular regions, known as cuboids, based on the distribution of relatively homogeneous pixel values. As the cuboid adopts a variable block partitioning based on the homogeneity of the data value, the partitioned blocks have better alignment with the object boundary. Moreover, in the cuboid coding process, only the partitioning tree information and a single value for each block need to be coded which takes lower number of bits and computational time compared to the traditional SHVC base layer. To verify the performance of the proposed method we embedded the proposed scheme as a base layer into the standard SHVC reference software and used several popular UHD/360-degree videos. The experimental results indicate that the proposed scalable coding strategy achieves an average of 14.04% BD-Rate reduction and 0.61 dB BD-PSNR gain for UHD/360-video compared to the operation points provided by an SHVC conforming encoder. Fariha Afsana, Manoranjan Paul, M. Manzur Murshed, David S. Taubman |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | A Commonality Modeling Framework for Enhanced Video Coding Leveraging on the Cuboidal Partitioning Based Representation of FramesabstractVideo coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. Modern video coding systems are block-based wherein commonality modeling is carried out only from the perspective of the block that need be coded next. In this work, we argue for a commonality modeling approach that can provide a seamless blending between global and local homogeneity information. For this purpose, at first the frame that need be coded, is recursively partitioned into rectangular regions based on the homogeneity information of the entire frame. After that each obtained rectangular region’s feature descriptor is taken to be the average value of all the pixels’ intensities encompassing the region. In this way, the proposed approach generates a coarse representation of the current frame by minimizing both global and local commonality. This coarse frame is computationally simple and has a compact representation. It attempts to preserve important structural properties of the current frame which can be viewed subjectively as well as from improved rate-distortion performance of a reference scalable HEVC coder that employs the coarse frame as a reference frame for encoding the current frame. Ashek Ahmmed, M. Manzur Murshed, Manoranjan Paul, David S. Taubman |
IEEE Trans. Multim. | 3 |
| 2021 | Dynamic Point Cloud Texture Video Compression using the Edge Position Difference Oriented Motion ModelabstractImmersive media representation format based on point clouds has underpinned significant opportunities for extended reality applications. Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression (V-PCC) technique projects a dynamic point cloud into geometry and texture video sequences. The projected texture video is then coded using modern video coding standard like HEVC. Since the properties of projected texture video frames are different from traditional video frames, HEVC-based commonality modeling can be inefficient. An improved commonality modeling technique is proposed that employs edge position difference oriented motion model. Experimental results show that the proposed commonality modeling technique can yield savings in bit rate of up to 3.15% over the V-PCC HEVC reference encoder. Ashek Ahmmed, Manoranjan Paul, Mark R. Pickering |
DCC | 2 |
| 2021 | Dynamic Point Cloud Compression Using A Cuboid Oriented Discrete Cosine Based Motion ModelabstractImmersive media representation format based on point clouds has underpinned significant opportunities for extended reality applications. Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression technique projects a dynamic point cloud into geometry and texture video sequences. The projected texture video is then coded using modern video coding standard like HEVC. Since the properties of projected texture video frames are different from traditional video frames, HEVC-based commonality modeling can be inefficient. An improved commonality modeling technique is proposed that employs discrete cosine basis oriented motion models and the domains of such models are approximated by homogeneous regions called cuboids. Experimental results show that the proposed commonality modeling technique can yield savings in bit rate of up to 4.17%. Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman |
ICASSP | 2 |
| 2021 | Coding and Quality Evaluation of Affordable 6DoF Video Content
Pan Gao 0001, Manoranjan Paul |
ICIG (3) | 4 |
| 2021 | Human-Machine Collaborative Video Coding Through Cuboidal PartitioningabstractVideo coding algorithms encode and decode an entire video frame while feature coding techniques only preserve and communicate the most critical information needed for a given application. This is because video coding targets human perception, while feature coding aims for machine vision tasks. Recently, attempts are being made to bridge the gap between these two domains. In this work, we propose a video coding framework by leveraging on to the commonality that exists between human vision and machine vision applications using cuboids. This is because cuboids, estimated rectangular regions over a video frame, are computationally efficient, has a compact representation and object centric. Such properties are already shown to add value to traditional video coding systems. Herein cuboidal feature descriptors are extracted from the current frame and then employed for accomplishing a machine vision task in the form of object detection. Experimental results show that a trained classifier yields superior average precision when equipped with cuboidal features oriented representation of the current test frame. Additionally, this representation costs 7% less in bit rate if the captured frames are need be communicated to a receiver. Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman |
ICIP | 2 |
| 2021 | Dynamic Point Cloud Geometry Compression using Cuboid based Commonality Modeling FrameworkabstractPoint cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression (V-PCC) technique projects a dynamic point cloud into geometry and texture video sequences. The projected geometry and texture video frames are then encoded using modern video coding standard like HEVC. However, HEVC encoder is unable to exploit the global commonality that exists within a geometry frame and between successive geometry frames to a greater extent. This is because in HEVC, the current frame partitioning starts from a rigid $64 \times 64$ pixels level without considering the structure of the scene need be coded. In this paper, an improved commonality modeling framework is proposed, by leveraging on cuboid-based frame partitioning, to encode point cloud geometry frames. The associated frame-partitioning scheme is based on statistical properties of the current geometry frame and therefore yields a flexible block partitioning structure composed of cuboids. Additionally, the proposed commonality modeling approach is computationally efficient and has a compact representation. Experimental results show that if the V-PCC reference encoder is augmented by the proposed commonality modeling technique, a bit rate savings of 2.71% and 4.25% are achieved for full body and upper body of human point clouds’ geometry sequences respectively. Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman |
ICIP | 2 |
| 2021 | Features Of ICU Admission In X-Ray Images Of Covid-19 PatientsabstractThis paper presents an original methodology for extracting semantic features from X-rays images that correlate to severity from a data set with patient ICU admission labels through interpretable models. The validation is partially performed by a proposed method that correlates the extracted features with a separate larger data set that does not contain the ICU-outcome labels. The analysis points out that a few features explain most of the variance between patients admitted in ICUs or not. The methods herein can be viewed as a statistical approach highlighting the importance of features related to ICU admission that may have been only qualitatively reported. In between features shown to be over-represented in the external data set were ones like ‘Consolidation’ (1.67), ‘Alveolar’ (1.33), and ‘Effusion’ (1.3). A brief analysis on the locations also showed higher frequency in labels like ‘Bilateral’ (1.58) and Peripheral (1.28) in patients labelled with higher chances to be admitted in ICU. To properly handle the limited data sets, a state-of-the-art lung segmentation network was also trained and presented, together with the use of low-complexity and interpretable models to avoid overfitting. Douglas P. S. Gomes, Anwaar Ulhaq, Manoranjan Paul, Michael J. Horry, Subrata Chakraborty, Manash Saha, Tanmoy Debnath, D. M. Motiur Rahaman |
ICIP | 3 |
| 2021 | Disocclusion filling for depth-based view synthesis with adaptive utilization of temporal correlations
Pan Gao 0001, Manoranjan Paul |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Automatically Assessing Quality of Online Health ArticlesabstractToday Information in the world wide web is overwhelmed by unprecedented quantity of data on versatile topics with varied quality. However, the quality of information disseminated in the field of medicine has been questioned as the negative health consequences of health misinformation can be life-threatening. There is currently no generic automated tool for evaluating the quality of online health information spanned over broad range. To address this gap, in this paper, we applied data mining approach to automatically assess the quality of online health articles based on 10 quality criteria. We have prepared a labelled dataset with 53012 features and applied different feature selection methods to identify the best feature subset with which our trained classifier achieved an accuracy of [Formula: see text] varied over 10 criteria. Our semantic analysis of features shows the underpinning associations between the selected features & assessment criteria and further rationalize our assessment approach. Our findings will help in identifying high quality health articles and thus aiding users in shaping their opinion to make right choice while picking health related help from online. Fariha Afsana, Muhammad Ashad Kabir, Naeemul Hassan, Manoranjan Paul |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Leveraging Cuboids for Better Motion Modeling in High Efficiency Video CodingabstractIn conventional video compression systems, motion model is used to approximate the geometry of moving object boundaries. It is possible to relieve motion model from describing discontinuities in the underlying motion field, by incorporating motion hint that can predict the spatial structure of future frames using the structure of reference frames. However, formation of highly accurate motion hint is computationally demanding, in particular for high resolution video sequences. Cuboids, rectangular regions derived using statistical features, attempt to separate out different objects present in the scene; they are computationally efficient and have sparse representation. Leveraging on the advantages of cuboids, in this paper, we propose to discover homogeneous motion regions and their associated motion based on cuboids. Afterwards, the estimated motion models and their domains are applied to form a prediction of the current frame. Experimental results show that a savings in bit rate of 3.96% is achievable over standalone HEVC reference, if this predicted frame is used as an additional reference frame for the current frame. Ashek Ahmmed, M. Manzur Murshed, Manoranjan Paul |
ICASSP | 3 |
| 2020 | Efficient Low Bit-Rate Intra-Frame Coding using Common Information for 360-degree VideoabstractWith the growth of video technologies, super-resolution videos, including 360-degree immersive video has become a reality due to exciting applications such as augmented/virtual/mixed reality for better interaction and a wide-angle user-view experience of a scene compared to traditional video with narrow-focused viewing angle. The new generation video contents are bandwidth-intensive in nature due to high resolution and demand high bit rate as well as low latency delivery requirements that pose challenges in solving the bottleneck of transmission and storage burdens. There is limited optimisation space in traditional video coding schemes for improving video coding efficiency in intra-frame due to the fixed size of processing block. This paper presents a new approach for improving intra-frame coding especially at low bit rate video transmission for 360-degree video for lossy mode of HEVC. Prior to using traditional HEVC intra-prediction, this approach exploits the global redundancy of entire frame by extracting common important information using multi-level discrete wavelet transformation. This paper demonstrates that the proposed method considering only low frequency information of a frame and encoding this can outperform the HEVC standard at low bit rates. The experimental results indicate that the proposed intra-frame coding strategy achieves an average of 54.07% BD-rate reduction and 2.84 dB BD-PSNR gain for low bit rate scenario compared to the HEVC. It also achieves a significant improvement in encoding time reduction of about 66.84% on an average. Moreover, this finding also demonstrates that the existing HEVC block partitioning can be applied in the transform domain for better exploitation of information concentration as we applied HEVC on wavelet frequency domain. Fariha Afsana, Manoranjan Paul, M. Manzur Murshed, David S. Taubman |
MMSP | 2 |
| 2020 | A Coarse Representation of Frames Oriented Video Coding By Leveraging Cuboidal Partitioning of Image DataabstractVideo coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. In this work, we form a coarse representation of the current frame by minimizing commonality within that frame while preserving important structural properties of the frame. The building blocks of this coarse representation are rectangular regions called cuboids, which are computationally simple and has a compact description. Then we propose to employ the coarse frame as an additional source for predictive coding of the current frame. Experimental results show an improvement in bit rate savings over a reference codec for HEVC, with minor increase in the codec computational complexity. Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman |
MMSP | 2 |
| 2020 | Depth Sequence Coding With Hierarchical Partitioning and Spatial-Domain QuantizationabstractDepth coding in 3D-HEVC deforms object shapes due to block-level edge-approximation and lacks efficient techniques to exploit the statistical redundancy, due to the frame-level clustering tendency in depth data, for higher coding gain at near-lossless quality. This paper presents a standalone mono-view depth sequence coder, which preserves edges implicitly by limiting quantization to the spatial-domain and exploits the frame-level clustering tendency efficiently with a novel binary tree-based decomposition (BTBD) technique. The BTBD can exploit the statistical redundancy in frame-level syntax, motion components, and residuals efficiently with fewer block-level prediction/coding modes and simpler context modeling for context-adaptive arithmetic coding. Compared with the depth coder in 3D-HEVC, the proposed one has achieved significantly lower bitrate at lossless to near-lossless quality range for mono-view coding and rendered superior quality synthetic views from the depth maps, compressed at the same bitrate, and the corresponding texture frames. Shampa Shahriyar, M. Manzur Murshed, Mortuza Ali, Manoranjan Paul |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Rate-Distortion Optimal Joint Texture and Depth Map Coding for 3-D Video StreamingabstractFor high compression efficiency, 3-D video coding usually employs a multimode methodology to exploit the dependencies between multiple views as well as between texture and depth. However, different coding modes will posses differentiating error propagation behaviour when the compressed 3-D video bit stream is transmitted over packet-switched networks, and thus lead to different amount of visual distortions. Further, the texture and depth distortions are combined in a highly complex fashion to produce the overall view synthesis distortion. To minimize the expected view synthesis distortion, this paper proposes an efficient rate-distortion optimized algorithm for joint selection of texture and depth modes. Firstly, a statistical model is developed to estimate the overall view synthesis distortion, in which the channel distortions caused by error propagation under different coding modes are analyzed. Then, joint optimization of texture and depth modes is derived within an operational rate-distortion framework using the Lagrange multiplier method. The adjacent block dependency caused by warping operation is explicitly considered in optimization, for which we develop a dynamic programming method to find the optimal solution. Finally, we extend the Lagrange minimization method to the more general variable-block-size prediction case, where the optimal quadtree tree structure and the combined coding modes are jointly determined using a multi-level dual trellis. Experimental results are presented for a wide range of packet loss rates to illustrate the effectiveness of the proposed algorithm. Pan Gao 0001, Manoranjan Paul |
IEEE Trans. Multim. | 2 |
| 2019 | Discrete Cosine Basis Oriented Motion Modeling for Fisheye and 360 Degree Video CodingabstractMotion modeling plays a central role in video compression. This role is even more critical in fisheye video sequences since the wide-angle fisheye imagery has special characteristics as in exhibiting radial distortion. While the translational motion model employed by modern video coding standards, such as HEVC, is sufficient in most cases, using higher order models is beneficial; for this reason, the upcoming video coding standard, VVC, employs a 4-parameter affine model. Discrete cosine basis has the ability to efficiently model complex motion fields. In this work, we investigate the motion modeling behaviour of the discrete cosine basis equipped with higher frequency cosine vectors. In particular, the developed discrete cosine basis is used as a single high-order model to describe a fisheye frame's motion; we employ this motion to produce an extra prediction reference, which is added to the HEVC list of references. Experimental results show an increase in delta bit rate, over conventional HEVC, when higher frequency cosine vectors are added in the motion modeling process. Then leveraging on this modified discrete cosine basis, we propose to employ it for predicting the motion in 360 degree video frames because of their resemblance with fisheye images. In this case, a delta bit rate of 2% is achieved, over conventional HEVC. Ashek Ahmmed, Manoranjan Paul |
MMSP | 2 |
| 2019 | Discrete Cosine Basis Oriented Homogeneous Motion Discovery for 360-Degree Video Coding
Ashek Ahmmed, Manoranjan Paul |
PSIVT | 2 |
| 2019 | Detection of Age and Defect of Grapevine Leaves Using Hyper Spectral Imaging
Tanmoy Debnath, Sourabhi Debnath, Manoranjan Paul |
PSIVT | 3 |
| 2019 | Grapevine Nutritional Disorder Detection Using Image Processing
D. M. Motiur Rahaman, Tintu Baby, Alex Oczkowski, Manoranjan Paul, Lihong Zheng, Leigh M. Schmidtke, Bruno P. Holzapfel, Rob R. Walker, Suzy Y. Rogiers |
PSIVT | 4 |
| 2019 | Enhanced Transfer Learning with ImageNet Trained Classification Layer
Tasfia Shermin, Shyh Wei Teng, M. Manzur Murshed, Guojun Lu, Ferdous Sohel, Manoranjan Paul |
PSIVT | 6 |
| 2019 | Efficient Self-embedding Data Hiding for Image Integrity Verification with Pixel-Wise Recovery Capability
Faranak Tohidi, Manoranjan Paul, Mohammad Reza Hooshmandasl, Tanmoy Debnath, Hojjat Jamshidi |
PSIVT | 2 |
| 2019 | Strided fully convolutional neural network for boosting the sensitivity of retinal blood vessels segmentation
Toufique Ahmed Soomro, Ahmed J. Afifi, Junbin Gao, Olaf Hellwich, Lihong Zheng, Manoranjan Paul |
Expert Syst. Appl. | 6 |
| 2019 | A Residue Number System Hardware Design of Fast-Search Variable-Motion-Estimation Accelerator for HEVC/H.265abstractA residue number system (RNS) has an inherent parallel structure that can be utilized for improving computer hardware systems. An RNS represents large integer numbers as a smaller integer set, or residues of a modulo set, without carry propagation between them. Hence mathematical operations, such as addition or subtraction, can be performed on the residues independently. This paper proposes an RNS implementation of motion estimation for the latest video coding standard known as high-efficiency video coding (HEVC) or H.265. Since motion estimation is the most computationally intensive task in video coding, several simplified algorithms are proposed for mitigating the problem, but the majority of them result in a worsening peak signal-to-noise ratio (PSNR) or bit-rate performance, or sometimes both. This paper also proposes a modified algorithm based on a test-zone (TZ) search algorithm, a widely used fast-search algorithm with good rate-distortion performance, suitable for hardware implementation for encoding ultra-high-definition videos in real time. The results show that worst-case PSNR degradation and bit-rate increases compared with the TZ search in the HEVC reference software implementation are negligible, and the hardware gate count is less than for many other designs in the literature. Cheeckottu Vayalil Niras, Manoranjan Paul, Yinan Kong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Spatial and Motion Saliency Prediction Method Using Eye Tracker Data for Video SummarizationabstractVideo summarization is the process to extract the most significant contents of a video and to represent it in a concise form. The existing methods for video summarization could not achieve a satisfactory result for a video with camera movement and significant illumination changes. To solve these problems, in this paper, a new framework for video summarization is proposed based on eye tracker data, as human eyes can track moving object accurately in these cases. The smooth pursuit is the state of eye movement when a user follows a moving object in a video. This motivates us to implement a new method to distinguish smooth pursuit from other type of gaze points, such as fixation and saccade. The smooth pursuit provides only the location of moving objects in a video frame; however, it does not indicate whether the located moving objects are very attractive (i.e., salient regions) to viewers or not, as well as the amount of motion of the moving objects. The amount of salient regions and object motions are the two important features to measure the viewer's attention level for determining the key frames for video summarization. To find the most attractive objects, a new spatial saliency prediction method is also proposed by constructing a saliency map around each smooth pursuit gaze point based on human visual field, such as fovea, parafoveal, and perifovea regions. To identify the amount of object motions, the total distances between the current and the previous gaze points of viewers during smooth pursuit are measured as a motion saliency score. The motivation is that the movement of eye gaze is related to the motion of the objects during smooth pursuit. Finally, both spatial and motion saliency maps are combined to obtain an aggregated saliency score for each frame and a set of key frames are selected based on user selected or system default skimming ratio. The proposed method is implemented on Office video data set that contains videos with camera movements and illumination changes. Experimental results confirm the superior performance of the proposed spatial and motion saliency prediction method compared with the state-of-the-art methods. Manoranjan Paul, Md. Musfequs Salehin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Virtual View Quality Enhancement Using Side View Information for Free Viewpoint VideoabstractWith the advancement of displaying technologies, virtual viewpoint video needs to be synthesized from adjacent viewpoints to provide immersive perceptual viewing experience of a scene. View synthesized techniques suffer poor rendering quality due to holes created by occlusion in the warping process. Currently, spatial and temporal correlation techniques are used to improve the quality of the synthesized view. However, spatial correlation e. g. inpainting and inverse mapping (IM) techniques cannot fill holes efficiently due to low spatial correlation in the edge between foreground and background pixels. On the other hand, the temporal correlation among already synthesized frames through learning by Gaussian mixture modelling (GMM) may fill occluded areas efficiently. However, there are no frames for GMM learning when the user switches view instantly. To address the aforementioned issues, in the proposed view synthesis technique, we apply GMM on the adjacent viewpoint videos. Then, we utilize the number of GMM models to refine pixel intensities of the synthesized view by using a weighting factor between pixel intensities in GMM models and warped images. This technique provides a better pixel correspondence, which improves 0.47~0.58dB PSNR compared to the IM technique. D. M. Motiur Rahaman, Manoranjan Paul, Nusrat Jahan Shoumy |
VCIP | 2 |
| 2018 | Virtual View Synthesis for Free Viewpoint Video and Multiview Video Compression using Gaussian Mixture ModellingabstractHigh quality virtual views need to be synthesized from adjacent available views for free viewpoint video and multiview video coding (MVC) to provide users with a more realistic 3D viewing experience of a scene. View synthesis techniques suffer from poor rendering quality due to holes created by occlusion and rounding integer error through warping. To remove the holes in the virtual view, the existing techniques use spatial and temporal correlation in intra/inter-view images and depth maps. However, they still suffer quality degradation in the boundary region of foreground and background areas due to the low spatial correlation in texture images and low correspondence in inter-view depth maps. To overcome the above-mentioned limitations, we use a number of models in the Gaussian mixture modeling (GMM) to separate background and foreground pixels in our proposed technique. Here, the missing pixels introduced from the warping process are recovered by the adaptive weighted average of the pixel intensities from the corresponding GMM model(s) and warped image. The weights vary with time to accommodate the changes due to a dynamic background and the motions of the moving objects for view synthesis. We also introduce an adaptive strategy to reset the GMM modeling if the contributions of the pixel intensities drop significantly. Our experimental results indicate that the proposed approach provides 5.40-6.60-dB PSNR improvement compared with the relevant methods. To verify the effectiveness of the proposed view synthesis technique, we use it as an extra reference frame in the motion estimation for MVC. The experimental results confirm that the proposed view synthesis is able to improve PSNR by 3.15-5.13 dB compared with the conventional three reference frames. D. M. Motiur Rahaman, Manoranjan Paul |
IEEE Trans. Image Process. | 2 |
| 2017 | Joint texture and depth map coding for error-resilient 3-D video transmissionabstractThis paper addresses the problem of error-resilient source coding for 3-D video transmission over packet-loss networks. The proposed approach jointly optimizes the texture coding mode and the depth coding mode for each macroblock in the reference views. Firstly, a distortion model is developed to capture the effect of the texture distortion and depth distortion on the synthesized view. Then, joint optimization of texture and depth coding modes is derived based upon an operational rate-distortion framework using Lagrange multiplier method. In particular, a dual trellis-based algorithm is introduced in order to overcome the macroblock interdependencies of texture and depth map in the optimization procedure. Simulation results demonstrate that significant and consistent gains can be achieved over currently used techniques. Pan Gao 0001, Wei Xiang 0001, D. M. Motiur Rahaman, Manoranjan Paul |
ICIP | 4 |
| 2017 | Retinal blood vessel extraction method based on basic filtering schemesabstractThe eye disease such as Diabetic Retinopathy(DR) can be analysed through segmentation of retinal blood vessels. In the last five years, many methods for retinal blood vessels segmentation were proposed. These methods give arise to the improved accuracy, however the sensitivity of low contrast vessels is often ignored. The performance of diagnosis in terms of segmentation of vessels can be degraded due to missing tiny vessels. In this study, we propose a novel algorithm aiming at improving the performance of segmenting small vessels. The proposed approach adopts a morphological and filtering method to handle the background noise and uneven illumination and uses anisotropic diffusion filtering to coherent the vessels and give initial detection of vessels, followed by a double threshold based region growing method. Toufique Ahmed Soomro, Manoranjan Paul, Junbin Gao, Lihong Zheng |
ICIP | 2 |
| 2017 | A novel angle-restricted test zone search algorithm for performance improvement of HEVCabstractHigh Efficiency Video Coding (HEVC) is the latest video encoding standard and has approximately 50% bit-rate saving compared to its predecessor. However, the motion estimation (ME) is considerably complicated by the incorporation of varieties of partitioning modes and a quad-tree based coding structure, and also by increasing the basic coding unit size by a factor of 16. Motion estimation is the most complex task in the video encoding process, consuming 60-80% of overall encoding time. This paper proposes a new algorithm, angle-restricted test zone (ARTZ) for motion estimation which is based on a test zone (TZ) search, exploiting directional probabilities of motion vector search. In our experiments, this proposal achieves a time saving in motion estimation of about 20% to 50% compared to a TZ search in the HEVC test model (HM) implementation for UHD videos without significant degradation of PSNR. Cheeckottu Vayalil Niras, Manoranjan Paul, Yinan Kong |
ICIP | 2 |
| 2017 | A Novel No-reference Subjective Quality Metric for Free Viewpoint Video Using Human Eye Movement
Pallab Kanti Podder, Manoranjan Paul, M. Manzur Murshed |
PSIVT | 2 |
| 2017 | Hybrid Adaptive Prediction Mechanisms with Multilayer Propagation Neural Network for Hyperspectral Image Compression
Manoranjan Paul |
PSIVT | 2 |
| 2017 | Adaptive weighted non-parametric background model for efficient video coding
Subrata Chakraborty, Manoranjan Paul, M. Manzur Murshed, Mortuza Ali |
Neurocomputing | 2 |
| 2017 | Computerised approaches for the detection of diabetic retinopathy using retinal fundus images: a survey
Toufique Ahmed Soomro, Junbin Gao, Tariq Mahmood Khan, Ahmad Fadzil M. Hani, Mohammad A. U. Khan, Manoranjan Paul |
Pattern Anal. Appl. | 6 |
| 2017 | Improved depth coding for HEVC focusing on depth edge approximation
Pallab Kanti Podder, Manoranjan Paul, D. M. Motiur Rahaman, M. Manzur Murshed |
Signal Process. Image Commun. | 2 |
| 2016 | Lossless depth map coding using binary tree based decomposition and context-based arithmetic codingabstractDepth maps are becoming increasingly important in the context of emerging video coding and processing applications. Depth images represent the scene surface and are characterized by areas of smoothly varying grey levels separated by sharp edges at the position of object boundaries. To enable high quality view rendering at the receiver side, preservation of these characteristics is important. Lossless coding enables avoiding rendering artifacts in synthesized views due to depth compression artifacts. In this paper, we propose a binary tree based lossless depth coding scheme that arranges the residual frame into integer or binary residual bitmap. High spatial correlation in depth residual frame is exploited by creating large homogeneous blocks of adaptive size, which are then coded as a unit using context based arithmetic coding. On the standard 3D video sequences, the proposed lossless depth coding has achieved compression ratio in the range of 20 to 80. Shampa Shahriyar, M. Manzur Murshed, Mortuza Ali, Manoranjan Paul |
ICME | 4 |
| 2016 | Efficient multi-view video coding using 3D motion estimation and virtual frame
Manoranjan Paul |
Neurocomputing | 1 |
| 2016 | A novel motion classification based intermode selection strategy for HEVC performance improvement
Pallab Kanti Podder, Manoranjan Paul, M. Manzur Murshed |
Neurocomputing | 2 |
| 2015 | Cuboid Coding of Depth Motion Vectors Using Binary Tree Based DecompositionabstractMotion vectors of depth-maps in multiview and free-viewpoint videos exhibit strong spatial as well as inter-component clustering tendency. This paper presents a novel motion vector coding technique that first compresses the multidimensional bitmaps of macro block mode information and then encodes only the non-zero components of motion vectors. The bitmaps are partitioned into disjoint cuboids using binary tree based decomposition so that the 0's and 1's are either highly polarized or further sub-partitioning is unlikely to achieve any compression. Each cuboid is entropy-coded as a unit using binary arithmetic coding. This technique is capable of exploiting the spatial and inter-component correlations efficiently without the restriction of scanning the bitmap in any specific linear order as needed by run-length coding. As encoding of non-zero component values no longer requires denoting the zero value, further compression efficiency is achieved. Experimental results on standard multiview test video sequences have comprehensively demonstrated the superiority of the proposed technique, achieving overall coding gain against the state-of-the-art in the range [17%,51%] and on average 31%. Shampa Shahriyar, M. Manzur Murshed, Mortuza Ali, Manoranjan Paul |
DCC | 4 |
| 2015 | Efficient coding strategy for HEVC performance improvement by exploiting motion featuresabstractThe striking feature of High Efficiency Video Coding (HEVC) Standard is emphasized by 50% bit-rate reduction compared to its predecessor H.264/AVC while keeping the same perceptual image quality. The time complexity - a congenital issue of HEVC has also increased to intensify the compression ratio. However, it is really a demanding task for the researchers to reduce the encoding time while preserving expected quality of the video sequences. Our contribution is to trim down the computational time by efficient selection of appropriate block-partitioning modes in HEVC using motion features based on phase-correlation. In this paper, we use phase-correlation between current and reference blocks to extract three motion features and combine them to determine binary motion pattern of the current block. The motion pattern is then matched against a codebook of predefined pattern templates to determine a subset of the inter-modes. Only the selected modes are exhaustively motion estimated and compensated for a coding unit. The experimental outcomes demonstrate that the average computational time can be down scaled by 30% of the HEVC while providing improved rate-distortion performance. Pallab Kanti Podder, Manoranjan Paul, M. Manzur Murshed |
ICASSP | 2 |
| 2015 | Efficient Compression of Hyperspectral Images Using Optimal Compression Cube and Image Plane
Manoranjan Paul |
MMM (1) | 2 |
| 2015 | Lossless image coding using binary tree decomposition of prediction residualsabstractState-of-the-art lossless image compression schemes, such as, JPEG-LS and CALIC, have been proposed in the context adaptive predictive coding framework. These schemes involve a prediction step followed by context adaptive entropy coding of the residuals. It can be observed that there exist significant spatial correlation among the residuals after prediction. The efficient schemes proposed in the literature rely on context adaptive entropy coding to exploit this spatial correlation. In this paper, we propose an alternative approach to exploit this spatial correlation. The proposed scheme also involves a prediction stage. However, we resort to a binary tree based hierarchical decomposition technique to efficiently exploit the spatial correlation. On a set of standard test images, the proposed scheme, using the same predictor as JPEG-LS, achieved an overall compression gain of 2.1% against JPEG-LS. Mortuza Ali, M. Manzur Murshed, Shampa Shahriyar, Manoranjan Paul |
PCS | 4 |
| 2015 | Fast Coding Strategy for HEVC by Motion Features and Saliency Applied on Difference Between Successive Image Blocks
Pallab Kanti Podder, Manoranjan Paul, M. Manzur Murshed |
PSIVT | 2 |
| 2015 | A novel depth motion vector coding exploiting spatial and inter-component clustering tendencyabstractMotion vectors of depth-maps in multiview and free-viewpoint videos exhibit strong spatial as well as inter-component clustering tendency. This paper presents a novel coding technique that first compresses the multidimensional bitmaps of macroblock mode and then encodes only the non-zero components of motion vectors. The bitmaps are partitioned into disjoint cuboids using binary tree based decomposition so that the 0's and 1's are either highly polarized or further sub-partitioning is unlikely to achieve any compression. Each cuboid is entropy-coded as a unit using binary arithmetic coding. This technique is capable of exploiting the spatial and inter-component correlations efficiently without the restriction of scanning the bitmap in any specific linear order as needed by run-length coding. As encoding of non-zero component values no longer requires denoting the zero value, further compression efficiency is achieved. Experimental results on standard multiview test video sequences have comprehensively demonstrated the superiority of the proposed technique, achieving overall coding gain against the state-of-the-art in the range [22%, 54%] and on average 38%. Shampa Shahriyar, M. Manzur Murshed, Mortuza Ali, Manoranjan Paul |
VCIP | 4 |
| 2015 | Epileptic seizure detection by exploiting temporal correlation of electroencephalogram signalsabstractElectroencephalogram (EEG) has a great potential for diagnosis and treatment of brain disorders like epileptic seizure. Feature extraction and classification of EEG signals is the crucial task to detect the stages of ictal and interictal signals for treatment and precaution of epileptic patients. However, existing seizure and non‐seizure feature extraction techniques are not good enough for the classification of ictal and interictal EEG signals considering the non‐abruptness phenomena and inconsistency in different brain locations. In this study, the authors present a new approach for feature extraction and classification by exploiting temporal correlation within EEG signals for better seizure detection as any abruptness in the temporal correlation within a signal represents the transition of a phenomenon. In the proposed methods, they divide an EEG signal into a number of epochs and arrange them into two‐dimensional matrix and then apply different transformation/decomposition to extract a number of statistical features. These features are then used as an input into LS‐SVM to classify them. Experimental results show that the proposed methods outperform the existing state‐of‐the‐art method for better classification in terms of sensitivity, specificity and accuracy of ictal and interictal period of epilepsy for benchmark datasets and different brain locations. Mohammad Zavid Parvez, Manoranjan Paul |
IET Signal Process. | 2 |
| 2014 | A novel video coding scheme using a scene adaptive non-parametric background modelabstractVideo coding techniques utilising background frames, provide better rate distortion performance by exploiting coding efficiency in uncovered background areas compared to the latest video coding standard. Parametric approaches such as the mixture of Gaussian (MoG) based background modeling has been widely used however they require prior knowledge about the test videos for parameter estimation. Recently introduced non-parametric (NP) based background modeling techniques successfully improved video coding performance through a HEVC integrated coding scheme. The inherent nature of the NP technique naturally exhibits superior performance in dynamic background scenarios compared to the MoG based technique without a priori knowledge of video data distribution. Although NP based coding schemes showed promising coding performances, they suffer from a number of key challenges - (a) determination of the optimal subset of training frames for generating a suitable background that can be used as a reference frame during coding, (b) incorporating dynamic changes in the background effectively after the initial background frame is generated, (c) managing frequent scene change leading to performance degradation, and (d) optimizing coding quality ratio between an I-frame and other frames under bit rate constraints. In this study we develop a new scene adaptive coding scheme using the NP based technique, capable of solving the current challenges by incorporating a new continuously updating background generation process. Extensive experimental results are also provided to validate the effectiveness of the new scheme. Subrata Chakraborty, Manoranjan Paul, M. Manzur Murshed, Mortuza Ali |
MMSP | 2 |
| 2014 | Epileptic seizure detection by analyzing EEG signals using different transformation techniques
Mohammad Zavid Parvez, Manoranjan Paul |
Neurocomputing | 2 |
| 2014 | A Long-Term Reference Frame for Hierarchical B-Picture-Based Video CodingabstractGenerally, H.264/AVC video coding standard with hierarchical bipredictive picture (HBP) structure outperforms the classical prediction structures such as “IPPP...” and “IBBP...” through better exploitation of data correlation using reference frames and unequal quantization setting among frames. However, multiple reference frames (MRFs) techniques are not fully exploited in the HBP scheme because of the computational requirement for B-frames, unavailability of adjacent reference frames, and with no explicit sorting of the reference frames for foreground or background being used. To exploit MRFs fully and explicitly in background referencing, we observe that not a single frame of a video is appropriate to be the reference frame as no one covers adequate background of a video. To overcome the problems, we propose a new coding scheme with the HBP, which uses the most common frame in scene (McFIS), generated by background modeling, as a long-term reference (LTR) frame for the third unipredictive reference frame, so that foreground and background areas are expected to be referenced from the two frames in the HBP structure and the McFIS, respectively. There are two approaches to generate McFIS under the proposed methodology. In the first approach, we generate a McFIS using a number of original frames of a scene in a video and then encode it as an I-frame with a higher quality. For the rest of the scene, this generated I-frame is used as an LTR frame. In the second approach, we generate an McFIS from the decoded frames and then use it as an LTR frame, without the need to encode the McFIS. The first and the second approaches are suitable for a video with static background and dynamic background, respectively. In general, the second approach requires more computational time than that of the the first approach. The experiments confirm that the proposed scheme outperforms three state-of-the-art algorithms by improving the image quality significantly with reduced computational time. Manoranjan Paul, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Disparity-adjusted 3D multi-view video coding with dynamic background modellingabstractCapturing a scene using multiple cameras from different angles is expected to provide the necessary interactivity in the 3D space to satisfy end-users' demands for observing objects and actions from different angles and depths. Existing multiview video coding (MVC) technologies are not sufficiently agile to exploit the interactivity and inefficient in terms of image quality and computational time. In this paper a novel technique is proposed using disparity-adjusted 3D MVC (DA-3D-MVC) with 3D motion estimation (ME) and 3D coding to overcome the problems. In the proposed scheme, a 3D frame is formed using the same temporal frames of all disparity-adjusted views and ME is carried out for the current 3D macroblock using the immediate previous 3D frame as a reference frame. Then, 3D coding technique is used for better compression. As all the same temporal position frames of all views are encoded at the same time, the proposed scheme provides better interactivity and reduced computational time compared to the H.264/MVC. To improve the rate-distortion (RD) performance of the proposed technique, an additional reference frame comprising dynamic background is also used. Experimental results reveal that the proposed scheme outperforms the H.264/MVC in terms of RD performance, computational time, and interactivity. Manoranjan Paul, Christopher J. Evans, M. Manzur Murshed |
ICIP | 1 |
| 2012 | 3D motion estimation for 3D video codingabstractH.264/MVC multi-view video coding provides a better compression rate compared to the simulcast coding using hierarchical B-picture prediction structure exploiting inter- and intra-view redundancy. However, this technique imposes random access frame delay as well as requiring huge computational time. In this paper a novel technique is proposed using 3D motion estimation (3D-ME) to overcome the problems. In the 3D-ME technique, a 3D frame is formed using the same temporal frames of all views and ME is carried out for the current 3D frame using the immediate previous 3D frame as a reference frame. As the correlation among the intra-view images is higher compared to the correlation among the inter-view images, the proposed 3D-ME technique reduces the overall computational time and eliminates the frame delay with comparable rate-distortion (RD) performance compared to H.264/MVC. Another technique is also proposed in the paper where an extra reference 3D frame comprising dynamic background frames (the most common frame of a scene i.e., McFIS) of each view is used for 3D-ME. Experimental results reveal that the proposed 3D-ME-McFIS technique outperforms the H.264/MVC in terms of improved RD performance by reducing computational time and by eliminating the random access frame delay. Manoranjan Paul, Junbin Gao, Michael Antolovich |
ICASSP | 1 |
| 2012 | The Image Matting Method with Regularized MatteabstractImage matting refers to the problem of accurately extracting foreground objects in images and video. The most recent works in natural image matting relies on the local and manifold smoothness assumptions on foreground and background colors on which a cost function is established. In this paper, we present a framework of formulating new regularization for robust solutions and illustrate new algorithms using the standard benchmark images. Junbin Gao, Manoranjan Paul, Jun Liu 0003 |
ICME | 2 |
| 2011 | McFIS in hierarchical bipredictve pictures-based video coding for referencing the stable area in a sceneabstractH.264/AVC video coding standard with hierarchical bipredictive picture (HBP) generally outperforms the other prediction structures such as ‘IPPP…’ and ‘IBBP…’ through better exploitation of data correlation using the preceding and succeeding reference frames. However, due to the different coding order of frames, the HBP scheme could not fully exploit the data correlations using multiple reference frames for occluded background, repetitive motion, etc. In this paper, we propose a new HBP scheme which uses the most common reference frame in scene (McFIS) as a third reference frame with other two closest bipredictive reference frames assuming that foreground and background areas of the current frame are referenced from the two bipredicted frames and the McFIS respectively. The experimental results confirm that the proposed scheme outperforms two state-of-art algorithms by improving significant image quality with comparable computational time. Manoranjan Paul, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
ICIP | 1 |
| 2011 | Explore and Model Better I-Frames for Video CodingabstractIn video coding, an intra (I)-frame is used as an anchor frame for referencing the subsequence frames, as well as error propagation prevention, indexing, and so on. To get better rate-distortion performance, a frame should have the following quality to be an ideal I-frame: the best similarity with the frames in a group of picture (GOP), so that when it is used as a reference frame for a frame in the GOP we need the least bits to achieve the desired image quality, minimize the temporal fluctuation of quality, and also maintain a more consistent bit count per frame. In this paper we use a most common frame of a scene in a video sequence with dynamic background modeling and then encode it to replace the conventional I-frame. The extensive experimental results confirm the superiority of our proposed scheme in comparison with the existing state-of-art methods by significant image quality improvement and computational time reduction. Manoranjan Paul, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Optimal Compression Plane for Efficient Video CodingabstractAll existing video coding standards developed so far deem video as a sequence of natural frames (formed in the XY plane), and treat spatial redundancy (redundancy along X and Y directions) and temporal redundancy (redundancy along T direction) differently and separately. In this paper, we investigate into a new compression (redundancy reduction) method for video in which the frames are allowed to be formed in a non-XY plane. We are to exploit fuller extent of video redundancy, and propose an adaptive optimal compression plane determination process to be used as a preprocessing step prior to any standard video coding scheme. The essence of the scheme is to form the frames in the plane formed by two axes (among X, Y, and T) corresponding to signal correlation evaluation, which enables better prediction (therefore better compression). In spite of the simplicity of the proposed method, it can be used for both lossless and lossy compression, and with and without interframe prediction. Extensive experimental results show that the new coding method improves the performance of the video coding for a number of coding methods (inclusive of lossless and near-lossless Motion JPEG-LS, Motion JPEG, Motion JPG2K, H.264 intraonly profile, and H.264) and videos with different visual content. Anmin Liu, Weisi Lin, Manoranjan Paul, Fan Zhang 0093, Chenwei Deng |
IEEE Trans. Image Process. | 3 |
| 2011 | Direct Intermode Selection for H.264 Video Coding Using Phase CorrelationabstractThe H.264 video coding standard exhibits higher performance compared to the other existing standards such as H.263, MPEG-X. This improved performance is achieved mainly due to the multiple-mode motion estimation and compensation. Recent research tried to reduce the computational time using the predictive motion estimation, early zero motion vector detection, fast motion estimation, and fast mode decision, etc. These approaches reduce the computational time substantially, at the expense of degrading image quality and/or increase bitrates to a certain extent. In this paper, we use phase correlation to capture the motion information between the current and reference blocks and then devise an algorithm for direct motion estimation mode prediction, without excessive motion estimation. A bigger amount of computational time is reduced by the direct mode decision and exploitation of available motion vector information from phase correlation. The experimental results show that the proposed scheme outperforms the existing relevant fast algorithms, in terms of both operating efficiency and video coding quality. To be more specific, 82 ~92% of encoding time is saved compared to the exhaustive mode selection (against 58 ~74% in the relevant state-of-the-art), and this is achieved without jeopardizing image quality (in fact, there is some improvement over the exhaustive mode selection at mid to high bit rates) and for a wide range of videos and bitrates (another advantages over the relevant state-of-the-art). Manoranjan Paul, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
IEEE Trans. Image Process. | 1 |
| 2010 | Video coding using the most common frame in sceneabstractMotion estimation (ME) and motion compensation (MC) using variable block size, fractional search, and multiple reference frames (MRFs) help the recent video coding standard H.264 to improve the coding performance significantly over the other contemporary coding standards. The concept of MRF achieves better coding performance in the cases of repetitive motion, uncovered background, non-integer pixel displacement, lighting change, etc. The requirement of index codes of the reference frames, computational time in ME&MC, and memory buffer for pre-coded frames limits the number of reference frames used in practical applications. In typical video sequence, the previous frame is used as a reference frame with 68~92% of cases. In this paper, we propose a new video coding method using a reference frame (i.e., the most common frame in scene (McFIS)) generated by the Gaussian mixture based dynamic background modelling. The McFIS is not only more effective in terms of rate-distortion and computational time performance compared to the MRFs but also error resilient transmission channel. The experimental results show that the proposed coding scheme outperforms the H.264 standard video coding with five reference frames by at least 0.5 dB and reduced 60% of computation time. Manoranjan Paul, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
ICASSP | 1 |
| 2010 | Comparison between H.264/AVC and Motion jpeg2000 for super-high definition video codingabstractH.264/AVC FRExt (Fidelity Range Extensions) and Motion JPEG2000 are the latest inter-frame and intra-frame video coding standards, respectively. It is well known that an inter-frame method achieves higher coding efficiency compared with an intra-frame one, and the Motion JPEG2000 has been selected for digital cinema compression. In this paper, we attempt to compare these two different schemes with theoretical and experimental analysis for super-HD (high definition) visual signals. One additional contribution of the paper is that the impact of block partition, motion search range and skipped block size for inter-frame coding is discussed. Based on the analysis, we extend the standard H.264/AVC FRExt by using larger block size and search range. The experimental results show that this extension leads to higher coding efficiency and makes the H.264/AVC FRExt more suitable for super-HD video coding. Chenwei Deng, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Manoranjan Paul |
ICIP | 5 |
| 2010 | Two dimensional Singular Value Decomposition (2D-SVD) based video codingabstractIn this paper, we propose a low-complexity video codec based on two-dimensional Singular Value Decomposition (2D-SVD). We exploit the common temporal characteristics of video without resorting to motion estimation. It has been demonstrated that this codec has higher coding efficiency than the relevant existing low complexity codecs. Moreover, the proposed codec performs well to deal with packet loss that is unavoidable in error-prone transmission. Therefore it is with advantages and good potential for wireless video applications such as mobile video calls and wireless surveillance. Zhouye Gu, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Manoranjan Paul |
ICIP | 5 |
| 2010 | Enhanced Just Noticeable Difference (JND) estimation with image decompositionabstractContrast masking (CM) on edge and textured regions have to be distinguished since distortions on edge regions are easier to be noticed than that on textured regions. Therefore, how to efficiently estimate the CM on edge and textured regions of an image is a key issue for accurate JND (Just Noticeable Difference) estimation. An enhanced image domain JND estimator is devised in this paper with new model for CM. We use the total variation method to obtain a structural image (which contains edge information) and a textural image (which contains texture information) from the input image, and then evaluate the CM for the two images separately rather than the whole image, and hence edge and texture are better distinguished and the under-estimation of JND on textured regions can be effectively avoided. Experimental results of subjective viewing confirm that the proposed model is capable of determining more accurate visibility thresholds. Anmin Liu, Weisi Lin, Fan Zhang 0093, Manoranjan Paul |
ICIP | 4 |
| 2010 | Pattern based video coding with uncovered backgroundabstract1The pattern-based video coding (PVC) outperforms the H.264 through better exploitation of block partitioning and partial block skipping. In the PVC scheme the best pattern is determined against the moving regions (MRs) in a macroblock (MB) of the current frame against the co-located MB in the reference frame; motion estimation (ME) and motion compensation (MC) are carried out using the pattern covered MRs, and the rest of the regions are treated as skipped areas. The MRs can be due to the object areas and the uncovered background (UCB) areas. Thus, the ME & MC by the pattern for the MRs of the UCB would not be accurate if there is no similar region in the reference frame. As a result no coding gain can be achieved for the UCB. Recently a dynamic background frame termed as the McFIS (the most common frame of a scene) has been generated using Gaussian mixture models for object detection. In this paper we propose a new PVC technique which will use the McFIS as a reference frame to determine the MRs where only object areas will be captured as the MRs. Thus, the proposed technique overcomes the mismatch problem of the UCB for ME&MC. The experimental results confirm the superiority of the proposed scheme in comparison with the existing PVC and McFIS-based methods by achieving significant image quality gain. Manoranjan Paul, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
ICIP | 1 |
| 2010 | Optimal compression plane (OCP) - A new framework for H.264 video codingabstractThis paper presents a framework for video coding, which is compatible with the existing H.264/AVC standard (for preprocessing). Since the video sequences are of 3D data matrices and the traditional XY image frame plane is not always the best for video coding in terms of rate-distortion (RD) performance, optimal compression plane (OCP) determination models are developed to improve the coding efficiency. We design coding plane level RD cost function (in analog to the one in H.264 macroblock level) to measure the RD performance of each coding plane. We jointly consider the available computational resource and the RD performance in order to determine an adaptive OCP. The performance of the proposed coding framework is assessed with a number video sequences. Extensive experimental results show that the new coding framework can achieve better RD performance at approximately the same computational complexity, for videos with different visual content. Anmin Liu, Weisi Lin, Manoranjan Paul, Chenwei Deng |
ICME | 3 |
| 2010 | McFIS: Better I-frame for video codingabstractThe conventional Intra (I-) frame is used for error propagation prevention, backward/forward play, random access, indexing, etc. This frame is also used as an anchor frame for referencing the subsequence frames. To get better rate-distortion performance a frame should have the following quality to be an ideal I-frame: the best similarity with the frames in a GOP, so that (i) when it is used as a reference frame for a frame in the GOP we need less bits to achieve the desired image quality; (ii) if any frame is missing at the decoding end we can retrieve the missing frame from it. In this paper we will generate a most common frame of a scene (McFIS) in a video sequence using dynamic background modelling and then encode it to replace the conventional I-frame. By using McFIS as an I-frame, we not only gain the above mentioned two benefits but also ensure adaptive GOP for better rate-distortion performance compared to the existing coding schemes. The experimental results confirm the superiority of our proposed scheme in comparison with the existing state-of-art methods by significant image quality and computation time. Manoranjan Paul, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
ISCAS | 1 |
| 2010 | Just Noticeable Difference for Images With Decomposition Model for Separating Edge and Textured RegionsabstractIn just noticeable difference (JND) models, evaluation of contrast masking (CM) is a crucial step. More specifically, CM due to edge masking (EM) and texture masking (TM) needs to be distinguished due to the entropy masking property of the human visual system. However, TM is not estimated accurately in the existing JND models since they fail to distinguish TM from EM. In this letter, we propose an enhanced pixel domain JND model with a new algorithm for CM estimation. In our model, total-variation based image decomposition is used to decompose an image into structural image (i.e., cartoon like, piecewise smooth regions with sharp edges) and textural image for estimation of EM and TM, respectively. Compared with the existing models, the proposed one shows its advantages brought by the better EM and TM estimation. It has been also applied to noise shaping and visual distortion gauge, and favorable results are demonstrated by experiments on different images. Anmin Liu, Weisi Lin, Manoranjan Paul, Chenwei Deng, Fan Zhang 0093 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Video Coding Focusing on Block Partitioning and OcclusionabstractAmong the existing block partitioning schemes, the pattern-based video coding (PVC) has already established its superiority at low bit-rate. Its innovative segmentation process with regular-shaped pattern templates is very fast as it avoids handling the exact shape of the moving objects. It also judiciously encodes the pattern-uncovered background segments capturing high level of interblock temporal redundancy without any motion compensation, which is favoured by the rate-distortion optimizer at low bit-rates. The existing PVC technique, however, uses a number of content-sensitive thresholds and thus setting them to any predefined values risks ignoring some of the macroblocks that would otherwise be encoded with patterns. Furthermore, occluded background can potentially degrade the performance of this technique. In this paper, a robust PVC scheme is proposed by removing all the content-sensitive thresholds, introducing a new similarity metric, considering multiple top-ranked patterns by the rate-distortion optimizer, and refining the Lagrangian multiplier of the H.264 standard for efficient embedding. A novel pattern-based residual encoding approach is also integrated to address the occlusion issue. Once embedded into the H.264 Baseline profile, the proposed PVC scheme improves the image quality perceptually significantly by at least 0.5 dB in low bit-rate video coding applications. A similar trend is observed for moderate to high bit-rate applications when the proposed scheme replaces the bi-directional predictive mode in the H.264 High profile. Manoranjan Paul, M. Manzur Murshed |
IEEE Trans. Image Process. | 1 |
| 2009 | A novel pattern identification scheme using distributed video coding conceptsabstractPattern-based video coding focusing on moving region in a macroblock has already established its superiority over recent H.264 video coding standard at very low bit rate. Obviously, a large number of pattern templates approximate the moving regions better however, after a certain limit no coding gain is observed due to the increase number of pattern identification bits. Recently, distributed video coding schemes used syndrome coding to predict the original information in decoder using side information. In this paper a novel pattern identification scheme is proposed which predicts the pattern from the syndrome codes and side information in decoder so that actual pattern identification number is not needed in the bitstream. The experimental results confirm that this new scheme successfully improves the rate-distortion performance compared to the existing pattern-based video coding as well as H.264 standard. This new scheme will also open another window of syndrome coding application. Manoranjan Paul, M. Manzur Murshed |
ICASSP | 1 |
| 2009 | An Efficient Mode Selection Prior to the Actual Encoding for H.264/AVC EncoderabstractMany video compression algorithms require decisions to be made to select between different coding modes. In the case of H.264, this includes decisions about whether or not motion compensation is used, and the block size to be used for motion compensation. It has been proposed that constrained optimization techniques, such as the method of Lagrange multipliers, can be used to trade off between the quality of the compressed video and the bit rate generated. In this paper, we show that in many cases of practical interest, very similar results can be achieved with much simpler optimizations. Mode selection by simply minimizing the distortion with motion vectors and header information produces very similar performance to the full constrained optimization, while it reduces the mode selection and over all encoding time by 31% and 12%, respectively. The proposed approach can be applied together with fast motion search algorithms and the mode filtering algorithms for further speed up. Manoranjan Paul, Michael R. Frater, John F. Arnold |
IEEE Trans. Multim. | 1 |
| 2008 | On Stable Dynamic Background Generation Technique Using Gaussian Mixture Models for Robust Object DetectionabstractGaussian mixture models (GMM) is used to represent the dynamic background in a surveillance video to detect the moving objects automatically. All the existing GMM based techniques inherently use the proportion by which a pixel is going to observe the background in any operating environment. In this paper we first show that such a proportion not only varies widely across different scenarios but also forbids using very fast learning rate. We then propose a dynamic background generation technique in conjunction with basic background subtraction which detected moving objects with improved stability and superior detection quality on a wide range of operating environments in two sets of benchmark surveillance sequences. Mahfuzul Haque, M. Manzur Murshed, Manoranjan Paul |
AVSS | 3 |
| 2008 | Threshold-free pattern-based low bit rate video codingabstractPattern-based video coding (PVC) has already established its superiority over recent video coding standard H.264, at low bit rate because of an extra pattern-mode to segment out the arbitrary shape of the moving region within the macroblock (MB). To determine the pattern-mode, the PVC however uses three thresholds to reduce the number of MBs coded using the pattern- mode. By setting these content-sensitive thresholds to any predefined values, the technique risks ignoring some MBs that would otherwise be selected by the rate-distortion optimization function for this mode. Consequently, the ultimate achievable performance is sacrificed to save motion estimation times. In this paper, a novel PVC scheme is proposed by removing all thresholds to determine this mode and hence more efficient performance is achieved without knowing the content of the video sequences. To keep computational complexity in check, pattern motion is approximated from the motion vector of the MB. In addition, efficient pattern similarity metric and new Lagrangian multipliers are also developed. The experimental results confirm that this new scheme improves the image quality by at least 0.5 dB and 1.0 dB compared to the existing PVC and the H.264 respectively. Manoranjan Paul, M. Manzur Murshed |
ICIP | 1 |
| 2008 | Improved Gaussian mixtures for robust object detection by adaptive multi-background generationabstractAdaptive Gaussian mixtures are widely used to model the dynamic background for real-time object detection. Recently the convergence speed of this approach is improved and a relatively robust statistical framework is proposed by Lee (PAMI, 2005). However, object quality still remains unacceptable due to poor Gaussian mixture quality, susceptibility to background/foreground data proportion, and inability to handle intrinsic background motion. This paper proposes an effective technique to eliminate these drawbacks by modifying the new model induction logic and using intensity difference thresholding to detect objects from one or more believe-to-be backgrounds. Experimental results on two benchmark datasets confirm that the object quality of the proposed technique is superior to that of Leepsilas technique at any model learning rate. Mahfuzul Haque, M. Manzur Murshed, Manoranjan Paul |
ICPR | 3 |
| 2008 | A hybrid object detection technique from dynamic background using Gaussian mixture modelsabstractAdaptive background modelling based object detection techniques are widely used in machine vision applications for handling the challenges of real-world multimodal background. But they are constrained to specific environment due to relying on environment specific parameters, and their performances also fluctuate across different operating speeds. On the other side, basic background subtraction (BBS) is not suitable for real applications due to manual background initialization requirement and its inability to handle repetitive multimodal background. However, it shows better stability across different operating speeds and can better eliminate noise, shadow, and trailing effect than adaptive techniques as no model adaptability or environment related parameters are involved. In this paper, we propose a hybrid object detection technique for incorporating the strengths of both approaches. In our technique, Gaussian mixture models (GMM) is used for maintaining an adaptive background model and both probabilistic and basic subtraction decisions are utilized for calculating inexpensive neighbourhood statistics for guiding the final object detection decision. Experimental results with two benchmark datasets and comparative analysis with recent adaptive object detection technique show the strength of the proposed technique in eliminating noise, shadow, and trailing effect while maintaining better stability across variable operating speeds. Mahfuzul Haque, M. Manzur Murshed, Manoranjan Paul |
MMSP | 3 |
| 2008 | Optimal arbitrary shaped pattern-based video codingabstractVery low bit-rate video coding algorithms using content-based generated patterns to segment out moving regions at macroblock level have exhibited good potential for improved coding efficiency when embedded into the H.264 standard as extra mode. This content-based pattern generation (CPG) algorithm provides local optimal result as only one pattern can be optimally generated from a given set of moving regions. But, it failed to provide optimal results for multiple patterns from entire sets. Obviously, a global optimal solution for clustering the set and then generation of multiple patterns enhances the performance farther. But a global optimal solution is not achievable due to the non-polynomial nature of the clustering problem. In this paper, we proposed a near optimal content-based pattern generation (OCPG) algorithm which outperforms the existing approach. Coupling OCPG, generating a set of patterns after clustering the macroblocks into several disjoint sets, with direct pattern selection algorithm by allowing all the macroblocks in multiple pattern modes outperforms the existing pattern-based coding while both embedded into the H.264. Manoranjan Paul, M. Manzur Murshed |
MMSP | 1 |
| 2008 | An efficient video coding using phase-matched error from phase correlation informationabstractThe H.264 video coding standard exhibits high performance in terms of compression and image quality compared to the other existing standard such as H.263, MPEGX. This improved performance is achieved due to the mainly enormous computations in multiple mode motion estimation and compensation. Recent research tried to reduce the computational time using predictive motion estimation, early zero motion vector detection, fast motion estimation, fast mode decision etc. These approaches successfully reduce the computational time by degrading the image quality. Phase correlation technique is used to find the shift between two pictures. In this paper we used phase correlation technique to indicate the motion information between current and reference block and then we devise an algorithm to predict the motion estimation block size. Using phase correlation we are able to successfully predict the motion estimation mode directly instead of using exhaustive motion estimation by all possible modes, thus we save a huge amount of computational time. The experimental results show that we can save around 50% time in motion estimation without degrading the image quality. Manoranjan Paul, Golam Sorwar |
MMSP | 1 |
| 2007 | Efficient H.264/AVC Video Encoder Where Pattern Is Used as Extra Mode for Wide Range of Video Coding
Manoranjan Paul, M. Manzur Murshed |
MMM (2) | 1 |
| 2007 | An Optimal Content-Based Pattern Generation AlgorithmabstractVery low bit-rate video coding algorithms using predefined regular-shaped patterns to segment out moving objects at macroblock level have exhibited good potential for improved coding efficiency when embedded in the H.264 standard as an extra mode. Even the best-matched regular-shaped pattern from a predefined codebook cannot approximate the shape of the object well, and there is no guarantee that even a regular-shaped object will have a close match with one of the limited number of predefined patterns. Intuitively, improved coding performance can be achieved if patterns are dynamically extracted from the video content. This letter presents a content-based pattern generation (CPG) algorithm for a set of macro blocks, which is shown optimal when only one pattern is allowed to represent the entire set. Coupling CPG, generating a pattern codebook after clustering the macro blocks into several disjoint sets, with any pattern selection algorithm outperforms the existing regular-shaped pattern-based coding while both embedded in H.264. Manoranjan Paul, M. Manzur Murshed |
IEEE Signal Process. Lett. | 1 |
| 2005 | A real-time pattern selection algorithm for very low bit-rate video coding using relevance and similarity metricsabstractVery low bit-rate video coding using regularly shaped patterns to represent moving regions in macroblocks has good potential for improved coding efficiency. This paper presents a real-time pattern selection (RTPS) algorithm, which uses a pattern relevance and similarity metric to achieve faster pattern selection from a large codebook. For each applicable macroblock, the relevance metric is applied to create a customized pattern codebook (CPC) from which the best pattern is selected using the similarity metric. The CPC size is adapted to facilitate real-time selection. Results prove the quantitative and perceptual performance of RTPS is superior to both the Fixed-8 algorithm and H.263. Manoranjan Paul, M. Manzur Murshed, Laurence Dooley |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | A new efficient similarity metric and generic computation strategy for pattern-based very low bit-rate video codingabstractIn the context of very low bit-rate video coding, pattern representations of a moving region (MR) in block-based motion estimation and compensation has become increasingly attractive. Generally, all existing pattern-matching algorithms apply a similarity metric, involving elementary operations, to compute the mismatch between an MR and a particular fixed pattern in order to select the best-matching pattern from a fixed-size codebook of predefined patterns. An efficient similarity metric, together with a new generic computation strategy, is presented by considering only the mismatch areas of MRs. It is theoretically proven that for a specific MR in a macroblock, the new similarity metric selects exactly the same pattern as existing metrics, while the resulting computational coding efficiency is improved by between 21% and 58% compared with the H.263 low bit-rate coding standard. Manoranjan Paul, M. Manzur Murshed, Laurence Dooley |
ICASSP (3) | 1 |
| 2003 | A new real-time pattern selection algorithm for very low bit-rate video coding focusing on moving regionsabstractVery low bit-rate video coding, using regular shaped patterns to focus on moving regions in macroblocks, has gained significant attention recently. This paper presents a new real-time pattern selection (RTPS) algorithm using a large codebook of thirty two patterns. The algorithm uses a relevance measurement for all the patterns and a moving region, to eliminate a large number of irrelevant patterns prior to the actual best likelihood pattern selection procedure. Both theoretically and empirically it is proven that not only is the computational complexity of the new algorithm comparable to the contemporary algorithm that use a pattern codebook size of only eight patterns but also the new algorithm reduces the bit-rate significantly, while maintaining comparable subjective quality. Manoranjan Paul, M. Manzur Murshed, Laurence Dooley |
ICASSP (3) | 1 |
| 2003 | A real time generic variable pattern selection algorithm for very low bit-rate video codingabstractThe selection of an optimal regular-shaped pattern set for very low bit-rate video coding, focusing on moving regions has been the objective of much recent research in order to try and improve bit-rate efficiency. Selecting the optimal pattern set however, is an NP hard problem. This paper presents a generic variable pattern selection (GVPS) algorithm, which introduces a pattern selection parameter that is able to control the performance in terms of computational complexity as well as bit-rate and picture quality. While using a sub-optimal variable pattern set, GVPS obtains a coding performance comparable to near-optimal algorithms, such as the k-change neighbourhood solution, while being much less computationally intensive, so that it is able to process all types of video sequences in real-time, with minimal pre-processing overheads. Manoranjan Paul, M. Manzur Murshed, Laurence Dooley |
ICIP (3) | 1 |
| 2003 | A new real-time pattern selection algorithm for very low bit-rate video coding focusing on moving regionsabstractVery low bit-rate video coding, using regular shaped patterns to focus on moving regions in macroblocks, has gained significant attention. This paper presents a new real-time pattern selection (RTPS) algorithm using a large codebook of thirty two patterns. The algorithm uses a relevance measurement for all the patterns and a moving region, to eliminate a large number of irrelevant patterns prior to the actual best likelihood pattern selection procedure. Both theoretically and empirically it is proven that not only is the computational complexity of the new algorithm comparable to the contemporary algorithm that use a pattern codebook size of only eight patterns but also the new algorithm reduces the bit-rate significantly, while maintaining comparable subjective quality. Manoranjan Paul, M. Manzur Murshed, Laurence Dooley |
ICME | 1 |