Raouf Hamzaoui

dblp:57/5365 · DBLP profile ↗
← Back
88ranked-venue papers
10as first author
36since 2021 · last 2026
0000-0001-6699-7331ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 9 first-author · 34 since 2021Computer networks · 10Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Feature Compression for Cloud-Edge Multimodal 3D Object Detection
abstract
Machine vision systems, which can efficiently manage extensive visual perception tasks, are becoming increasingly popular in industrial production and daily life. Due to the challenge of simultaneously obtaining accurate depth and texture information with a single sensor, multimodal data captured by cameras and LiDAR is commonly used to enhance performance. Additionally, cloud-edge cooperation has emerged as a novel computing approach to improve user experience and ensure data security in machine vision systems. This paper proposes a pioneering solution to address the feature compression problem in multimodal 3D object detection. Given a sparse tensor-based object detection network at the edge device, we introduce two modes to accommodate different application requirements: Transmission-Friendly Feature Compression (T-FFC) and Accuracy-Friendly Feature Compression (A-FFC). In T-FFC mode, only the output of the last layer of the network's backbone is transmitted from the edge device. The received feature is processed at the cloud device through a channel expansion module and two spatial upsampling modules to generate multi-scale features. In A-FFC mode, we expand upon the T-FFC mode by transmitting two additional types of features. These added features enable the cloud device to generate more accurate multi-scale features. Experimental results on the KITTI dataset using the VirConv-L detection network showed that T-FFC was able to compress the features by a factor of 4933 with less than a 3% reduction in detection performance. On the other hand, A-FFC compressed the features by a factor of about 733 with almost no degradation in detection performance. We also designed optional residual extraction and 3D object reconstruction modules to facilitate the reconstruction of detected objects. The reconstructed objects effectively reflected the shape, occlusion, and details of the original objects.
Chongzhen Tian, Hui Yuan 0001, Raouf Hamzaoui, Liquan Shen, Sam Kwong
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 VP-JND: Visual Perception Assisted Deep Picture-Wise Just Noticeable Difference Prediction Model for Image Compression
abstract
The Picture-Wise Just Noticeable Difference (PW-JND) represents the visibility threshold of human vision when viewing distorted images. The PW-JND plays an important role in perceptual image processing and compression. However, predicting the PW-JND is challenging due to its dependence on image content, viewing conditions, and the viewer. In this paper, we propose a visual perception-assisted deep PW-JND (VP-JND) prediction model for image compression that combines data-driven methods with the perceptual mechanisms of human vision. First, we identify a correlation between PW-JND and conventional pixel-wise JND. Based on this observation, we design the VP-JND model, consisting of a pixel-wise JND model, a deep binary classifier (VP-JNDnet) and a binary block search algorithm for refining predictions. VP-JNDnet exploits the pixel-wise JND map of the original image to predict whether a compressed image is perceptually lossless. In addition, the model incorporates visual importance of content and regions by using a mixed attention module and calculating perceptual loss during training. Experimental results show that VP-JND achieved an average precision of 94.82% and a mean absolute difference of 3.92 in predicting the JPEG quality factor corresponding to the PW-JND on the MCL-JCI dataset, outperforming state-of-the-art JND models. When applied to perceptual lossless image coding, the predicted PW-JND enabled average bit rate savings of 89.35% for JPEG compression on MCL-JCI and 85.46%/41.13% for JPEG/BPG compression on KonJND-1k. These savings were relative to images compressed at the lowest distortion level. The source codes and trained models are publicly available at https://github.com/SYSU-Video/VP-JND.
Yun Zhang 0002, Shisheng Zhang, Na Li 0015, Chunling Fan, Raouf Hamzaoui
IEEE Trans. Circuits Syst. Video Technol.5
2026 FD-SCU: Frequency Decomposition-Based Spectrum Collaborative Upsampling for Point Cloud Color Attribute
abstract
Existing point cloud color upsampling methods typically treat color upsampling as an interpolation problem within a local color or implicit feature domain. This largely overlooks the ability of the frequency domain to capture color correlations in local point sets. To address this limitation, we propose a spectrum collaborative strategy that uses frequency decomposition on voxel blocks (VBs) to enhance point cloud color reconstruction. We first voxelize the low-resolution (LR) color point cloud to generate multiple VBs and introduce a virtual filling strategy that adaptively assigns colors to empty voxels in each VB, ensuring that the irregularly distributed color information fully occupies the VB. We then apply the discrete cosine transform, known for its strong frequency-domain representation of locally smooth signals, to each color-filled VB to obtain frequency coefficients. These frequency coefficients are separated into high-frequency (HF) and low-frequency (LF) components. The LF coefficients, together with the LR color point cloud, are fed into a multi-scale cross-domain feature extraction module to capture deep features. Next, a Gaussian perturbation-based feature expansion generates upsampled color features, which are used to regress a coarse upsampled color point cloud. Finally, a high-frequency-guided residual refinement module uses the HF coefficients to refine the coarse upsampled result and produce a high-fidelity color point cloud. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art methods. Our code will be publicly available at https://github.com/wangwenchaoxx/FD-SCU.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Weiqing Yan, Junhui Hou
IEEE Trans. Image Process.4
2026 LPCM: Learning-Based Predictive Coding for LiDAR Point Cloud Compression
abstract
In recent years, LiDAR point clouds have been widely used in many applications. Since the data volume of LiDAR point clouds is very huge, efficient compression is necessary to reduce their storage and transmission costs. However, existing learning-based compression methods do not exploit the inherent angular resolution of LiDAR and ignore the significant differences in the correlation of geometry information at different bitrates. The predictive geometry coding method in the geometry-based point cloud compression (G-PCC) standard uses the inherent angular resolution to predict the azimuth angles. However, it only models a simple linear relationship between the azimuth angles of neighboring points. Moreover, it does not optimize the quantization parameters for residuals on each coordinate axis in the spherical coordinate system. To address these issues, we propose a learning-based predictive coding method (LPCM) with both high-bitrate and low-bitrate coding modes. LPCM converts point clouds into predictive trees using the spherical coordinate system. In high-bitrate coding mode, we use a lightweight Long-Short-Term Memory-based predictive (LSTM-P) module that captures long-term geometry correlations between different coordinates to efficiently predict and compress the elevation angles. In low-bitrate coding mode, where geometry correlation degrades, we introduce a variational radius compression (VRC) module to directly compress the point radii. Then, we analyze why the quantization of spherical coordinates differs from that of Cartesian coordinates and propose a differential evolution (DE)-based quantization parameter selection method, which improves rate-distortion performance without increasing coding time. Experimental results show that LPCM achieved a D1-PSNR BD-rate reduction of 21.2% compared with the G-PCC lossless octree-based coding mode on SemanticKITTI, and 5.6% compared with the PredGeom on Ford, using the latest G-PCC test model TMC13 v31.0.
Hui Yuan 0001, Shiqi Jiang 0006, Da Ai, Wei Zhang 0072, Raouf Hamzaoui
IEEE Trans. Image Process.6
2026 Inter-LPCM: Learning-Based Inter-Frame Predictive Coding for LiDAR Point Cloud Compression
abstract
Because LiDAR sensors acquire point clouds with a fixed angular resolution, the resulting data can be systematically parameterized and efficiently compressed in the spherical coordinate system. Traditional spherical coordinate-based point cloud compression methods have shown strong rate-distortion (RD) performance, with the predictive geometry coding (PredGeom) method in the geometry-based point cloud compression (G-PCC) standard being a prominent example. While PredGeom includes an inter-frame prediction mode, it relies on a simple linear model, which limits its ability to capture complex motion patterns or structural dependencies. On the other hand, existing learning-based compression methods in the spherical domain do not exploit inter-frame correlations to reduce geometry redundancy. To address these limitations, we propose a learning-based inter-frame predictive coding method (Inter-LPCM). For azimuth prediction, we use a delta coding strategy based on the predefined angular resolution. To improve compression for radii, we introduce an inter-frame radius predictive (Inter-RP) model that estimates the current point's radius using neighboring points from both the current frame and the registered reference frame. In addition, we design a lightweight attention-based prediction (LAEP) model to predict elevation angles by capturing long-range geometric correlations across different coordinates. For quantization, we propose an RD-optimized method to select the quantization steps in the spherical coordinate system. For entropy coding, we design distinct models for each spherical coordinate component. These models are adapted to the statistical priors of each coordinate, which enables more accurate probability estimation. Experimental results show that Inter-LPCM, in its best RD configuration, achieved a D1-PSNR BD-rate reduction of 26.1% compared with the G-PCC lossless octree-based coding mode on SemanticKITTI, and 8.3% compared with the inter-frame prediction mode of PredGeom on Ford, using the latest G-PCC test model TMC13 v31.0. Our source code is publicly available at https://github.com/SDUChangSun/Inter-LPCM.
Hui Yuan 0001, Shiqi Jiang 0006, Chongzhen Tian, Raouf Hamzaoui
IEEE Trans. Image Process.6
2026 UGAE: Unified Geometry and Attribute Enhancement for G-PCC Compressed Point Clouds
abstract
Lossy compression of point clouds reduces storage and transmission costs; however, it inevitably leads to irreversible distortion in geometry structure and attribute information. To address these issues, we propose a unified geometry and attribute enhancement (UGAE) framework, which consists of three core components: post-geometry enhancement (PoGE), pre-attribute enhancement (PAE), and post-attribute enhancement (PoAE). In PoGE, a Transformer-based sparse convolutional U-Net is used to reconstruct the geometry structure with high precision by predicting voxel occupancy probabilities. Building on the refined geometry structure, PAE introduces an innovative enhanced geometry-guided recoloring strategy, which uses a detail-aware K-Nearest Neighbors (DA-KNN) method to achieve accurate recoloring and effectively preserve high-frequency details before attribute compression. Finally, at the decoder side, PoAE uses an attribute residual prediction network with a weighted mean squared error (W-MSE) loss to enhance the quality of high-frequency regions while maintaining the fidelity of low-frequency regions. UGAE significantly outperformed existing methods on three benchmark datasets: 8iVFB, Owlii, and MVUB. Compared to the latest G-PCC test model (TMC13v29), in terms of total bitrate setting, UGAE achieved an average BD-PSNR gain of 9.98 dB and -90.54% BD-bitrate for geometry under the D1 metric, as well as a 3.34 dB BD-PSNR improvement with -55.53% BD-bitrate for attributes. Additionally, it improved perceptual quality significantly. Our source code will be released on GitHub at: https://github.com/yuanhui0325/UGAE.
Hui Yuan 0001, Chongzhen Tian, Raouf Hamzaoui
IEEE Trans. Image Process.5
2025 PCAC-GAN: A Sparse-Tensor-Based Generative Adversarial Network for 3D Point Cloud Attribute Compression
abstract
Learning-based methods have proven successful in compressing geometric information for point clouds. For attribute compression, however, they still lag behind non-learning-based methods such as the MPEG G-PCC standard. To bridge this gap, we propose a novel deep learning-based point cloud attribute compression method that uses a generative adversarial network (GAN) with sparse convolution layers. Our method also includes a module that adaptively selects the resolution of the voxels used to voxelize the input point cloud. Sparse vectors are used to represent the voxelized point cloud, and sparse convolutions process the sparse tensors, ensuring computational efficiency. To the best of our knowledge, this is the first application of GANs to compress point cloud attributes. Our experimental results show that our method outperforms existing learning-based techniques and rivals the latest G-PCC test model (TMC13v23) in terms of visual quality.
Xiaolong Mao, Hui Yuan 0001, Xin Lu 0001, Raouf Hamzaoui, Wei Gao 0003
Comput. Vis. Media4
2025 PU-GSM: A Latent Geometry-Guided Self-Similarity Model for Point Cloud Upsampling
abstract
Existing point cloud upsampling methods typically treat upsampling as a local interpolation problem, neglecting the importance of global correlations within point sets, which can limit their performance. To address this limitation, we exploit the inherent self-similarity of point clouds from a global perspective and propose PU-GSM, a latent geometry-guided self-similarity model for upsampling. We first generate a lower-resolution sparse sub-point cloud (SPC) by downsampling the input point cloud (IPC). Then, we introduce a latent geometry-guided self-similarity model (LGSM) that learns a point distribution on the underlying surface of SPC by exploiting the inherent self-similarity of IPC. Next, we reuse the LGSM for the remaining points (i.e., the points left after removing SPC from IPC). Afterward, we introduce a gradient-aware dual domain refiner to generate and calibrate the upsampled point cloud from the learned point distribution. Finally, we propose an inference-free latent vector matching approach to regularize the upsampled point cloud by enhancing the feature similarity between the upsampled point cloud and the ground truth in latent space. Extensive experiments show that PU-GSM achieves better upsampling results compared to state-of-the-art methods. Our code will be available at: https://github.com/liuhaoyun/PU-GSM.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Weiqing Yan
IEEE Trans. Circuits Syst. Video Technol.3
2025 PCE-GAN: A Generative Adversarial Network for Point Cloud Attribute Quality Enhancement Based on Optimal Transport
abstract
Point cloud compression significantly reduces data volume but sacrifices reconstruction quality, highlighting the need for advanced quality enhancement techniques. Most existing approaches focus primarily on point-to-point fidelity, often neglecting the importance of perceptual quality as interpreted by the human visual system. To address this issue, we propose a generative adversarial network for point cloud quality enhancement (PCE-GAN), grounded in optimal transport theory, with the goal of simultaneously optimizing both data fidelity and perceptual quality. The generator consists of a local feature extraction (LFE) unit, a global spatial correlation (GSC) unit and a feature squeeze unit. The LFE unit uses dynamic graph construction and a graph attention mechanism to efficiently extract local features, placing greater emphasis on points with severe distortion. The GSC unit uses the geometry information of neighboring patches to construct an extended local neighborhood and introduces a transformer-style structure to capture long-range global correlations. The discriminator computes the deviation between the probability distributions of the enhanced point cloud and the original point cloud, guiding the generator to achieve high quality reconstruction. Experimental results show that the proposed method achieves state-of-the-art performance. Specifically, when applying PCE-GAN to the latest geometry-based point cloud compression (G-PCC) test model, it achieves an average BD-rate of -19.2% compared with the PredLift coding configuration and -18.3% compared with the RAHT coding configuration. Subjective comparisons show a significant improvement in texture clarity and color transitions, revealing finer details and more natural color gradients.
Hui Yuan 0001, Qi Liu 0029, Honglei Su, Raouf Hamzaoui, Sam Kwong
IEEE Trans. Image Process.5
2025 SPAC: Sampling-Based Progressive Attribute Compression for Dense Point Clouds
abstract
We propose an end-to-end attribute compression method for dense point clouds. The proposed method combines a frequency sampling module, an adaptive scale feature extraction module with geometry assistance, and a global hyperprior entropy model. The frequency sampling module uses a Hamming window and the Fast Fourier Transform to extract high-frequency components of the point cloud. The difference between the original point cloud and the sampled point cloud is divided into multiple sub-point clouds. These sub-point clouds are then partitioned using an octree, providing a structured input for feature extraction. The feature extraction module integrates adaptive convolutional layers and uses offset-attention to capture both local and global features. Then, a geometry-assisted attribute feature refinement module is used to refine the extracted attribute features. Finally, a global hyperprior model is introduced for entropy encoding. This model propagates hyperprior parameters from the deepest (base) layer to the other layers, further enhancing the encoding efficiency. At the decoder, a mirrored network is used to progressively restore features and reconstruct the color attribute through transposed convolutional layers. The proposed method encodes base layer information at a low bitrate and progressively adds enhancement layer information to improve reconstruction accuracy. Compared to the best anchor of the latest geometry-based point cloud compression (G-PCC) standard that was proposed by the Moving Picture Experts Group (MPEG), the proposed method can achieve an average Bjøntegaard delta bitrate of -24.58% for the Y component (resp. -21.23% for YUV components) on the MPEG Category Solid dataset and -22.48% for the Y component (resp. -17.19% for YUV components) on the MPEG Category Dense dataset. This is the first instance that a learning-based attribute codec outperforms the G-PCC standard on these datasets by following the common test conditions specified by MPEG. Our source code will be made publicly available on https://github.com/sduxlmao/SPAC.
Xiaolong Mao, Hui Yuan 0001, Shiqi Jiang 0006, Raouf Hamzaoui, Sam Kwong
IEEE Trans. Image Process.5
2025 Global Spatial-Temporal Information-Based Residual ConvLSTM for Video Space-Time Super-Resolution
abstract
By converting low-frame-rate, low-resolution videos into high-frame-rate, high-resolution ones, space-time video super-resolution techniques can enhance visual experiences and facilitate more efficient information dissemination. We propose a convolutional neural network (CNN) for space-time video super-resolution, namely GIRNet. Our method combines long-term global information and short-term local information from the video to better extract complete and accurate spatial-temporal information. To generate highly accurate features and thus improve performance, the proposed network integrates a feature-level temporal interpolation module with deformable convolutions and a global spatial-temporal information-based residual convolutional long short-term memory (convLSTM) module. In the feature-level temporal interpolation module, we leverage deformable convolution, which adapts to deformations and scale variations of objects across different scene locations. This provides a more efficient solution than conventional convolution for extracting features from moving objects. Our network effectively uses forward and backward feature information to determine inter-frame offsets, leading to the direct generation of interpolated frame features. In the global spatial-temporal information-based residual convLSTM module, the first convLSTM is used to derive global spatial-temporal information from the input features, and the second convLSTM uses the previously computed global spatial-temporal information feature as its initial cell state. This second convLSTM adopts residual connections to preserve spatial information, thereby enhancing the output features. Experiments on the Vimeo90 K dataset show that the proposed method outperforms open source state-of-the-art techniques in peak signal-to-noise-ratio (by 1.45 dB, 1.14 dB, and 0.2 dB over STARnet, TMNet, and 3DAttGAN, respectively), structural similarity index(by 0.027, 0.023, and 0.006 over STARnet, TMNet, and 3DAttGAN, respectively), and visual quality.
Congrui Fu, Hui Yuan 0001, Shiqi Jiang 0006, Liquan Shen, Raouf Hamzaoui
IEEE Trans. Multim.6
2025 EdgeRegNet: Edge Feature-Based Multimodal Registration Network Between Images and LiDAR Point Clouds
abstract
Cross-modal data registration has long been a critical task in computer vision, with extensive applications in autonomous driving and robotics. Accurate and robust registration methods are essential for aligning data from different modalities, forming the foundation for multimodal sensor data fusion and enhancing perception systems' accuracy and reliability. The registration task between 2D images captured by cameras and 3D point clouds captured by Light Detection and Ranging (LiDAR) sensors is usually treated as a visual pose estimation problem. High-dimensional feature similarities from different modalities are leveraged to identify pixel-point correspondences, followed by pose estimation techniques using least squares methods. However, existing approaches often resort to downsampling the original point cloud and image data due to computational constraints, inevitably leading to a loss in precision. Additionally, high-dimensional features extracted using different feature extractors from various modalities require specific techniques to mitigate cross-modal differences for effective matching. To address these challenges, we propose a method that uses edge information from the original point clouds and images for cross-modal registration. We retain crucial information from the original data by extracting edge points and pixels, enhancing registration accuracy while maintaining computational efficiency. The use of edge points and edge pixels allows us to introduce an attention-based feature exchange block to eliminate cross-modal disparities. Furthermore, we incorporate an optimal matching layer to improve correspondence identification. We validate the accuracy of our method on the KITTI and nuScenes datasets, demonstrating its state-of-the-art performance. Our code is publicly available on GitHub athttps://github.com/ESRSchao/EdgeRegNet.
Yuanchao Yue, Hui Yuan 0001, Qinglong Miao, Xiaolong Mao, Raouf Hamzaoui, Peter Eisert
IEEE Trans. Multim.5
2025 CS-Net: Contribution-Based Sampling Network for Point Cloud Simplification
abstract
Point cloud sampling plays a crucial role in reducing computation costs and storage requirements for various vision tasks. Traditional sampling methods, such as farthest point sampling, lack task-specific information and, as a result, cannot guarantee optimal performance in specific applications. Learning-based methods train a network to sample the point cloud for the targeted downstream task. However, they do not guarantee that the sampled points are the most relevant ones. Moreover, they may result in duplicate sampled points, which requires completion of the sampled point cloud through post-processing techniques. To address these limitations, we propose a contribution-based sampling network (CS-Net), where the sampling operation is formulated as a Top-$k$k operation. To ensure that the network can be trained in an end-to-end way using gradient descent algorithms, we use a differentiable approximation to the Top-$k$k operation via entropy regularization of an optimal transport problem. Our network consists of a feature embedding module, a cascade attention module, and a contribution scoring module. The feature embedding module includes a specifically designed spatial pooling layer to reduce parameters while preserving important features. The cascade attention module combines the outputs of three skip connected offset attention layers to emphasize the attractive features and suppress less important ones. The contribution scoring module generates a contribution score for each point and guides the sampling process to prioritize the most important ones. Experiments on the ModelNet40 and PU147 showed that CS-Net achieved state-of-the-art performance in two semantic-based downstream tasks (classification and registration) and two reconstruction-based tasks (compression and surface reconstruction). CS-Net also achieved high average precision for objection detection on the KITTI LiDAR point cloud dataset, demonstrating its effectiveness in three-dimensional object detection.
Chen Chen 0063, Hui Yuan 0001, Xiaolong Mao, Raouf Hamzaoui, Junhui Hou
IEEE Trans. Vis. Comput. Graph.5
2025 Progressive Knowledge Transfer Network Based on Human Visual Perception Mechanism for No-Reference Point Cloud Quality Assessment
abstract
Point cloud perceptual quality assessment plays a critical role in many applications, including compression and communication. We propose PKT-PCQA, a point-based no-reference point cloud quality assessment deep learning network that emulates the human visual system by using progressive knowledge transfer to convert coarse-grained quality classification knowledge into a fine-grained quality prediction task. PKT-PCQA exploits local and global features, as well as an attention mechanism based on spatial and channel attention modules. Experiments on three large and independent point cloud assessment datasets show that PKT-PCQA outperforms existing no-reference and reduced-reference point cloud quality assessment methods and achieves better or similar performance compared to several State-of-the-Art full-reference methods.
Honglei Su, Qi Liu 0029, Hui Yuan 0001, Raouf Hamzaoui
IEEE Trans. Vis. Comput. Graph.5
2024 Enhancing Octree-Based Context Models for Point Cloud Geometry Compression With Attention-Based Child Node Number Prediction
abstract
In point cloud geometry compression, most octree-based context models use the cross-entropy between the one-hot encoding of node occupancy and the probability distribution predicted by the context model as the loss. This approach converts the problem of predicting the number (a regression problem) and the position (a classification problem) of occupied child nodes into a 255-dimensional classification problem. As a result, it fails to accurately measure the difference between the one-hot encoding and the predicted probability distribution. We first analyze why the cross-entropy loss function fails to accurately measure the difference between the one-hot encoding and the predicted probability distribution. Then, we propose an attention-based child node number prediction (ACNP) module to enhance the context models. The proposed module can predict the number of occupied child nodes and map it into an 8-dimensional vector to assist the context model in predicting the probability distribution of the occupancy of the current node for efficient entropy coding. Experimental results demonstrate that the proposed module enhances the coding efficiency of octree-based context models.
Hui Yuan 0001, Xiaolong Mao, Xin Lu 0001, Raouf Hamzaoui
IEEE Signal Process. Lett.5
2024 PU-Mask: 3D Point Cloud Upsampling via an Implicit Virtual Mask
abstract
We present PU-Mask, a virtual mask-based network for 3D point cloud upsampling. Unlike existing upsampling methods, which treat point cloud upsampling as an “unconstrained generative” problem, we propose to address it from the perspective of “local filling”, i.e., we assume that the sparse input point cloud (i.e., the unmasked point set) is obtained by locally masking the original dense point cloud with virtual masks. Therefore, given the unmasked point set and virtual masks, our goal is to fill the point set hidden by the virtual masks. Specifically, because the masks do not actually exist, we first locate and form each virtual mask by a virtual mask generation module. Then, we propose a mask-guided transformer-style asymmetric auto-encoder (MTAA) to restore the upsampled features. Moreover, we introduce a second-order unfolding attention mechanism to enhance the interaction between the feature channels of MTAA. Next, we generate a coarse upsampled point cloud using a pooling technique that is specific to the virtual masks. Finally, we design a learnable pseudo Laplacian operator to calibrate the coarse upsampled point cloud and generate a refined upsampled point cloud. Extensive experiments demonstrate that PU-Mask is superior to the state-of-the-art methods. Our code will be made available at: https://github.com/liuhaoyun/PU-Mask.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Qi Liu 0029, Shuai Li 0005
IEEE Trans. Circuits Syst. Video Technol.3
2024 Crowdsourced Estimation of Collective Just Noticeable Difference for Compressed Video With the Flicker Test and QUEST+
abstract
The concept of videowise just noticeable difference (JND) was recently proposed for determining the lowest bitrate at which a source video can be compressed without perceptible quality loss with a given probability. This bitrate is usually obtained from estimates of the satisfied used ratio (SUR) at different encoding quality parameters. The SUR is the probability that the distortion corresponding to the quality parameter is not noticeable. Commonly, the SUR is computed experimentally by estimating the subjective JND threshold of each subject using a binary search, fitting a distribution model to the collected data, and creating the complementary cumulative distribution function of the distribution. The subjective tests consist of paired comparisons between the source video and compressed versions. However, as shown in this paper, this approach typically overestimates or underestimates the SUR. To address this shortcoming, we directly estimate the SUR function by considering the entire population as a collective observer. In our method, the subject for each paired comparison is randomly chosen, and a state-of-the-art Bayesian adaptive psychometric method (QUEST+) is used to select the compressed video in the paired comparison. Our simulations show that this collective method yields more accurate SUR results using fewer comparisons than traditional methods. We also perform a subjective experiment to assess the JND and SUR for compressed video. In the paired comparisons, we apply a flicker test that compares a video interleaving the source video and its compressed version with the source video. Analysis of the subjective data reveals that the flicker test provides, on average, greater sensitivity and precision in the assessment of the JND threshold than does the usual test, which compares compressed versions with the source video. Using crowdsourcing and the proposed approach, we build a JND dataset for 45 source video sequences that are encoded with both advanced video coding (AVC) and versatile video coding (VVC) at all available quantization parameters. Our dataset and the source code have been made publicly available at http://database.mmsp-kn.de/flickervidset-database.html.
Mohsen Jenadeleh, Raouf Hamzaoui, Ulf-Dietrich Reips, Dietmar Saupe
IEEE Trans. Circuits Syst. Video Technol.2
2024 Dependence-Based Coarse-to-Fine Approach for Reducing Distortion Accumulation in G-PCC Attribute Compression
abstract
Geometry-based point cloud compression (G-PCC) is a state-of-the-art point cloud compression standard. While G-PCC achieves excellent performance, its reliance on the predicting transform leads to a significant dependence problem, which can easily result in distortion accumulation. This not only increases bitrate consumption but also degrades reconstruction quality. To address these challenges, we propose a dependence-based coarse-to-fine approach for distortion accumulation in G-PCC attribute compression. Our method consists of three modules: level-based adaptive quantization, point-based adaptive quantization, and Wiener filter-based refinement level quality enhancement. The level-based adaptive quantization module addresses the interlevel-of-detail (LOD) dependence problem, while the point-based adaptive quantization module tackles the interpoint dependence problem. On the other hand, the Wiener filter-based refinement level quality enhancement module enhances the reconstruction quality of each point based on the dependence order among LODs. Extensive experimental results demonstrate the effectiveness of the proposed method. Notably, when the proposed method was implemented in the latest G-PCC test model (TMC13v23.0), a Bj$\phi$ntegaard delta rate of$-$4.9%,$-$12.7%, and$-$14.0% was achieved for the Luma, Chroma Cb, and Chroma Cr components, respectively.
Hui Yuan 0001, Raouf Hamzaoui
IEEE Trans. Ind. Informatics3
2024 Colored Point Cloud Quality Assessment Using Complementary Features in 3D and 2D Spaces
abstract
Point Cloud Quality Assessment (PCQA) plays an essential role in optimizing point cloud acquisition, encoding, transmission, and rendering for human-centric visual media applications. In this paper, we propose an objective PCQA model using Complementary Features from 3D and 2D spaces, called CF-PCQA, to measure the visual quality of colored point clouds. First, we develop four effective features in 3D space to represent the perceptual properties of colored point clouds, which include curvature, kurtosis, luminance distance and hue features of points in 3D space. Second, we project the 3D point cloud onto 2D planes using patch projection and extract a structural similarity feature of the projected 2D images in the spatial domain, as well as a sub-band similarity feature in the wavelet domain. Finally, we propose a feature selection and a learning model to fuse high dimensional features and predict the visual quality of the colored point clouds. Extensive experimental results show that the Pearson Linear Correlation Coefficients (PLCCs) of the proposed CF-PCQA were 0.9117, 0.9005, 0.9340 and 0.9826 on the SIAT-PCQD, SJTU-PCQA, WPC2.0 and ICIP2020 datasets, respectively. Moreover, statistical significance tests demonstrate that the CF-PCQA significantly outperforms the state-of-the-art PCQA benchmark schemes on the four datasets.
Mao Cui, Yun Zhang 0002, Chunling Fan, Raouf Hamzaoui, Qinglan Li
IEEE Trans. Multim.4
2024 Support Vector Regression-Based Reduced- Reference Perceptual Quality Model for Compressed Point Clouds
abstract
Video-based point cloud compression (V-PCC) is a state-of-the-art moving picture experts group (MPEG) standard for point cloud compression. V-PCC can be used to compress both static and dynamic point clouds in a lossless, near lossless, or lossy way. Many objective quality metrics have been proposed for distorted point clouds. Most of these metrics are full-reference metrics that require both the original point cloud and the distorted one. However, in some real-time applications, the original point cloud is not available, and no-reference or reduced-reference quality metrics are needed. Three main challenges in the design of a reduced-reference quality metric are how to build a set of features that characterize the visual quality of the distorted point cloud, how to select the most effective features from this set, and how to map the selected features to a perceptual quality score. We address the first challenge by proposing a comprehensive set of features consisting of compression, geometry, normal, curvature, and luminance features. To deal with the second challenge, we use the least absolute shrinkage and selection operator (LASSO) method, which is a variable selection method for regression problems. Finally, we map the selected features to the mean opinion score in a nonlinear space. Although we have used only 19 features in our current implementation, our metric is flexible enough to allow any number of features, including future more effective ones. Experimental results on the Waterloo point cloud dataset version 2 (WPC2.0) and the MPEG point cloud compression dataset (M-PCCD) show that our method, namely PCQAML, outperforms state-of-the-art full-reference and reduced-reference quality metrics in terms of Pearson linear correlation coefficient, Spearman rank order correlation coefficient, Kendall's rank-order correlation coefficient, and root mean squared error.
Honglei Su, Qi Liu 0029, Hui Yuan 0001, Qiang Shawn Cheng, Raouf Hamzaoui
IEEE Trans. Multim.5
2023 CAS-Net: Cascade Attention-Based Sampling Neural Network for Point Cloud Simplification
abstract
Point cloud sampling can reduce storage requirements and computation costs for various vision tasks. Traditional sampling methods, such as farthest point sampling, are not geared towards downstream tasks and may fail on such tasks. In this paper, we propose a cascade attention-based sampling network (CAS-Net), which is end-to-end trainable. Specifically, we propose an attention-based sampling module (ASM) to capture the semantic features and preserve the geometry of the original point cloud. Experimental results on the ModelNet40 dataset show that CAS-Net outperforms state-of-the-art methods in a sampling-based point cloud classification task, while preserving the geometric structure of the sampled point cloud.
Chen Chen 0063, Hui Yuan 0001, Hao Liu 0044, Junhui Hou, Raouf Hamzaoui
ICME5
2023 Relaxed forced choice improves performance of visual quality assessment methods
abstract
In image quality assessment, a collective visual quality score for an image or video is obtained from the individual ratings of many subjects. One commonly used format for these experiments is the two-alternative forced choice method. Two stimuli with the same content but differing visual quality are presented sequentially or side-by-side. Subjects are asked to select the one of better quality, and when uncertain, they are required to guess. The relaxed alternative forced choice format aims to reduce the cognitive load and the noise in the responses due to the guessing by providing a third response option, namely, “not sure”. This work presents a large and comprehensive crowdsourcing experiment to compare these two response formats: the one with the “not sure” option and the one without it. To provide unambiguous ground truth for quality evaluation, subjects were shown pairs of images with differing numbers of dots and asked each time to choose the one with more dots. Our crowdsourcing study involved 254 participants and was conducted using a within-subject design. Each participant was asked to respond to 40 pair comparisons with and without the “not sure” response option and completed a questionnaire to evaluate their cognitive load for each testing condition. The experimental results show that the inclusion of the “not sure” response option in the forced choice method reduced mental load and led to models with better data fit and correspondence to ground truth. We also tested for the equivalence of the models and found that they were different. The dataset is available at http://database.mmsp-kn.de/cogvqa-database.html.
Mohsen Jenadeleh, Johannes Zagermann, Harald Reiterer, Ulf-Dietrich Reips, Raouf Hamzaoui, Dietmar Saupe
QoMEX5
2023 GQE-Net: A Graph-Based Quality Enhancement Network for Point Cloud Color Attribute
abstract
In recent years, point clouds have become increasingly popular for representing three-dimensional (3D) visual objects and scenes. To efficiently store and transmit point clouds, compression methods have been developed, but they often result in a degradation of quality. To reduce color distortion in point clouds, we propose a graph-based quality enhancement network (GQE-Net) that uses geometry information as an auxiliary input and graph convolution blocks to extract local features efficiently. Specifically, we use a parallel-serial graph attention module with a multi-head graph attention mechanism to focus on important points or features and help them fuse together. Additionally, we design a feature refinement module that takes into account the normals and geometry distance between points. To work within the limitations of GPU memory capacity, the distorted point cloud is divided into overlap-allowed 3D patches, which are sent to GQE-Net for quality enhancement. To account for differences in data distribution among different color components, three models are trained for the three color components. Experimental results show that our method achieves state-of-the-art performance. For example, when implementing GQE-Net on a recent test model of the geometry-based point cloud compression (G-PCC) standard, 0.43 dB, 0.25 dB and 0.36 dB Bjφntegaard delta (BD)-peak-signal-to-noise ratio (PSNR), corresponding to 14.0%, 9.3% and 14.5% BD-rate savings were achieved on dense point clouds for the Y, Cb, and Cr components, respectively. The source code of our method is available at https://github.com/xjr998/GQE-Net.
Jinrui Xing, Hui Yuan 0001, Raouf Hamzaoui, Hao Liu 0044, Junhui Hou
IEEE Trans. Image Process.3
2023 No-Reference Bitstream-Layer Model for Perceptual Quality Assessment of V-PCC Encoded Point Clouds
abstract
No-reference bitstream-layer models for point cloud quality assessment (PCQA) use the information extracted from a bitstream for real-time and nonintrusive quality monitoring. We propose a no-reference bitstream-layer model for the perceptual quality assessment of video-based point cloud compression (V-PCC) encoded point clouds. First, we study the relationship between the perceptual coding distortion and the texture quantization parameter (TQP) when geometry encoding is lossless. The results indicate that the perceptual coding distortion depends on the texture complexity (TC). Next, we estimate TC using TQP and the texture bitrate per pixel (TBPP), both of which are extracted from the compressed bitstream without resorting to complete decoding. This allows us to build a texture distortion model as a function of TQP and TBPP. By combining this texture distortion model with a geometry distortion model that depends on the geometry quantization parameter (GQP), we obtain an overall no-reference bitstream-layer PCQA model that we call bitstreamPCQ. Experimental results show that the proposed model markedly outperforms existing models in terms of widely used performance criteria, including the Pearson linear correlation coefficient (PLCC), the Spearman rank order correlation coefficient (SRCC) and the root mean square error (RMSE).
Qi Liu 0029, Honglei Su, Tianxin Chen, Hui Yuan 0001, Raouf Hamzaoui
IEEE Trans. Multim.5
2022 PU-Refiner: A Geometry Refiner with Adversarial Learning for Point Cloud Upsampling
abstract
We present PU-Refiner, a generative adversarial network for point cloud upsampling. The generator of our network includes a coarse feature expansion module to create coarse upsampled features, a geometry generation module to regress a coarse point cloud from the coarse upsampled features, and a progressive geometry refinement module to restore the dense point cloud in a coarse-to-fine fashion based on the coarse upsampled point cloud. The discriminator of our network helps the generator produce point clouds closer to the target distribution. It makes full use of multi-level features to improve its classification performance. Extensive experimental results show that PU-Refiner is superior to five state-of-the-art point cloud upsampling methods. Code: https://github.com/liuhaoyun/PU-Refiner.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Wei Gao 0003, Shuai Li 0005
ICASSP3
2022 APCCPA '22: 1st International Workshop on Advances in Point Cloud Compression, Processing and Analysis
abstract
Point clouds are attracting much attention from academia, industry and standardization organizations such as MPEG, JPEG, and AVS. 3D Point clouds consisting of thousands or even millions of points with attributes can represent real-world objects and scenes in a way that enables an improved immersive visual experience and facilitates complex 3D vision tasks. In addition to various point cloud analysis and processing tasks (e.g., segmentation, classification, 3D object detection, registration), efficient compression for these large-scale 3D visual data is essential to make point cloud applications more effective. This workshop focuses on point cloud processing, analy sis, and compression in challenging situations to further improve visual experience and machine vision performance. Both learning-based and non-learning-based perception-oriented optimization algorithms for compression and processing are solicited. Contributions that advance the state-of-the-art in analysis tasks, are also welcomed.
Wei Gao 0003, Ge Li 0002, Hui Yuan 0001, Raouf Hamzaoui, Zhu Li 0001, Shan Liu 0001
ACM Multimedia4
2022 Large-Scale Crowdsourced Subjective Assessment of Picturewise Just Noticeable Difference
abstract
The picturewise just noticeable difference (PJND) for a given image, compression scheme, and subject is the smallest distortion level that the subject can perceive when the image is compressed with this compression scheme. The PJND can be used to determine the compression level at which a given proportion of the population does not notice any distortion in the compressed image. To obtain accurate and diverse results, the PJND must be determined for a large number of subjects and images. This is particularly important when experimental PJND data are used to train deep learning models that can predict a probability distribution model of the PJND for a new image. To date, such subjective studies have been carried out in laboratory environments. However, the number of participants and images in all existing PJND studies is very small because of the challenges involved in setting up laboratory experiments. To address this limitation, we develop a framework to conduct PJND assessments via crowdsourcing. We use a new technique based on slider adjustment and a flicker test to determine the PJND. A pilot study demonstrated that our technique could decrease the study duration by 50% and double the perceptual sensitivity compared to the standard binary search approach that successively compares a test image side by side with its reference image. Our framework includes a robust and systematic scheme to ensure the reliability of the crowdsourced results. Using 1,008 source images and distorted versions obtained with JPEG and BPG compression, we apply our crowdsourcing framework to build the largest PJND dataset, KonJND-1k (Konstanz just noticeable difference 1k dataset). A total of 503 workers participated in the study, yielding 61,030 PJND samples that resulted in an average of 42 samples per source image. The KonJND-1k dataset is available athttp://database.mmsp-kn.de/konjnd-1k-database.html
Hanhe Lin, Guangan Chen, Mohsen Jenadeleh, Vlad Hosu, Ulf-Dietrich Reips, Raouf Hamzaoui, Dietmar Saupe
IEEE Trans. Circuits Syst. Video Technol.6
2022 PUFA-GAN: A Frequency-Aware Generative Adversarial Network for 3D Point Cloud Upsampling
abstract
We propose a generative adversarial network for point cloud upsampling, which can not only make the upsampled points evenly distributed on the underlying surface but also efficiently generate clean high frequency regions. The generator of our network includes a dynamic graph hierarchical residual aggregation unit and a hierarchical residual aggregation unit for point feature extraction and upsampling, respectively. The former extracts multiscale point-wise descriptive features, while the latter captures rich feature details with hierarchical residuals. To generate neat edges, our discriminator uses a graph filter to extract and retain high frequency points. The generated high resolution point cloud and corresponding high frequency points help the discriminator learn the global and high frequency properties of the point cloud. We also propose an identity distribution loss function to make sure that the upsampled points remain on the underlying surface of the input low resolution point cloud. To assess the regularity of the upsampled points in high frequency regions, we introduce two evaluation metrics. Objective and subjective results demonstrate that the visual quality of the upsampled point clouds generated by our method is better than that of the state-of-the-art methods.
Hao Liu 0044, Hui Yuan 0001, Junhui Hou, Raouf Hamzaoui, Wei Gao 0003
IEEE Trans. Image Process.4
2021 Model-Based Rate-Distortion Optimized Video-Based Point Cloud Compression with Differential Evolution
Hui Yuan 0001, Raouf Hamzaoui, Ferrante Neri, Shengxiang Yang
ICIG (1)2
2021 Adaptive Quantization for Predicting Transform-Based Point Cloud Compression
Guoxia Sun, Hui Yuan 0001, Raouf Hamzaoui
ICIG (1)4
2021 Global Rate-distortion Optimization of Video-based Point Cloud Compression with Differential Evolution
abstract
In video-based point cloud compression (V-PCC), one geometry video and one color video are generated from a dynamic point cloud. Then, the two videos are compressed independently using a state-of-the-art video coder. In the Moving Picture Experts Group (MPEG) V-PCC test model, the quantization parameters for a given group of frames are constrained according to a fixed offset rule. For example, for the low-delay configuration, the difference between the quantization parameters of the first frame and the quantization parameters of the following frames in the same group is zero by default. We show that the rate-distortion performance of the V-PCC test model can be improved by lifting this constraint and considering the ratedistortion optimization problem as a multi-variable constrained combinatorial optimization problem where the variables are the quantization parameters of all frames. To solve the optimization problem, we use a variant of the differential evolution algorithm. Experimental results for the low-delay configuration show that our method can achieve a Bjøntegaard delta bitrate of up to -43.04% and more accurate rate control (average bitrate error to the target bitrate of 0.45% vs. 10.75%) compared to the state-of- the-art method, which optimizes the rate-distortion performance subject to the test model default offset rule. We also show that our optimization strategy can be used to improve the rate-distortion performance of two-dimensional video coders.
Hui Yuan 0001, Raouf Hamzaoui, Ferrante Neri, Shengxiang Yang
MMSP2
2021 Kalman filter-based prediction refinement and quality enhancement for geometry-based point cloud compression
abstract
A point cloud is a set of points representing a three-dimensional (3D) object or scene. To compress a point cloud, the Motion Picture Experts Group (MPEG) geometry-based point cloud compression (G-PCC) scheme may use three attribute coding methods: region adaptive hierarchical transform (RAHT), predicting transform (PT), and lifting transform (LT). To improve the coding efficiency of PT, we propose to use a Kalman filter to refine the predicted attribute values. We also apply a Kalman filter to improve the quality of the reconstructed attribute values at the decoder side. Experimental results show that the combination of the two proposed methods can achieve an average Bjøntegaard delta bitrate of −0.48%, −5.18%, and −6.27% for the Luma, Chroma Cb, and Chroma Cr components, respectively, compared with a recent G-PCC reference software.
Jian Sun 0013, Hui Yuan 0001, Raouf Hamzaoui
VCIP4
2021 Reduced Reference Perceptual Quality Model With Application to Rate Control for Video-Based Point Cloud Compression
abstract
In rate-distortion optimization, the encoder settings are determined by maximizing a reconstruction quality measure subject to a constraint on the bitrate. One of the main challenges of this approach is to define a quality measure that can be computed with low computational cost and which correlates well with the perceptual quality. While several quality measures that fulfil these two criteria have been developed for images and videos, no such one exists for point clouds. We address this limitation for the video-based point cloud compression (V-PCC) standard by proposing a linear perceptual quality model whose variables are the V-PCC geometry and color quantization step sizes and whose coefficients can easily be computed from two features extracted from the original point cloud. Subjective quality tests with 400 compressed point clouds show that the proposed model correlates well with the mean opinion score, outperforming state-of-the-art full reference objective measures in terms of Spearman rank-order and Pearson linear correlation coefficient. Moreover, we show that for the same target bitrate, rate-distortion optimization based on the proposed model offers higher perceptual quality than rate-distortion optimization based on exhaustive search with a point-to-point objective quality metric. Our datasets are publicly available at https://github.com/qdushl/Waterloo-Point-Cloud-Database-2.0.
Qi Liu 0029, Hui Yuan 0001, Raouf Hamzaoui, Honglei Su, Junhui Hou, Huan Yang 0001
IEEE Trans. Image Process.3
2021 Highly Efficient Multiview Depth Coding Based on Histogram Projection and Allowable Depth Distortion
abstract
Mismatches between the precisions of representing the disparity, depth value and rendering position in 3D video systems cause redundancies in depth map representations. In this paper, we propose a highly efficient multiview depth coding scheme based on Depth Histogram Projection (DHP) and Allowable Depth Distortion (ADD) in view synthesis. Firstly, DHP exploits the sparse representation of depth maps generated from stereo matching to reduce the residual error from INTER and INTRA predictions in depth coding. We provide a mathematical foundation for DHP-based lossless depth coding by theoretically analyzing its rate-distortion cost. Then, due to the mismatch between depth value and rendering position, there is a many-to-one mapping relationship between them in view synthesis, which induces the ADD model. Based on this ADD model and DHP, depth coding with lossless view synthesis quality is proposed to further improve the compression performance of depth coding while maintaining the same synthesized video quality. Experimental results reveal that the proposed DHP based depth coding can achieve an average bit rate saving of 20.66% to 19.52% for lossless coding on Multiview High Efficiency Video Coding (MV-HEVC) with different groups of pictures. In addition, our depth coding based on DHP and ADD achieves an average depth bit rate reduction of 46.69%, 34.12% and 28.68% for lossless view synthesis quality when the rendering precision varies from integer, half to quarter pixels, respectively. We obtain similar gains for lossless depth coding on the 3D-HEVC, HEVC Intra coding and JPEG2000 platforms.
Yun Zhang 0002, Linwei Zhu, Raouf Hamzaoui, Sam Kwong, Yo-Sung Ho
IEEE Trans. Image Process.3
2021 Guest Editorial Special Section on Hybrid Human-Artificial Intelligence for Multimedia Computing
abstract
The papers in this special section focus on hybrid human-artificial intelligene (AI) for multimedia computing. Multimedia computing has experienced a tremendous growth in the last decades, with applications ranging from multimedia information retrieval and analysis to multimedia compression and communication. However, the increasing volume and complexity of multimedia data driven by the large-scale spread of various new devices and sensors is posing a serious challenge to traditional multimedia computing algorithms. Artificial intelligence (AI), in particular deep learning techniques, has improved the performance of multimedia computing algorithms for many tasks, including computer vision and natural language processing. But unlike humans, AI is poor at solving tasks across multiple domains or in dealing with an uncontrolled dynamic environment. Hybrid Human-Artificial Intelligence (HH-AI) is an emerging field that aims at combining the benefits of human intelligence, such as semantic association, inference, and generalization with the computing power of AI.
Raouf Hamzaoui, Huansheng Ning, Chonggang Wang, Reza Malekian
IEEE Trans. Multim.1
2021 Model-Based Joint Bit Allocation Between Geometry and Color for Video-Based 3D Point Cloud Compression
abstract
In video-based 3D point cloud compression, the quality of the reconstructed 3D point cloud depends on both the geometry, and color distortions. Finding an optimal allocation of the total bitrate between the geometry coder, and the color coder is a challenging task due to the large number of possible solutions. To solve this bit allocation problem, we first propose analytical distortion, and rate models for the geometry, and color information. Using these models, we formulate the joint bit allocation problem as a constrained convex optimization problem, and solve it with an interior point method. Experimental results show that the rate-distortion performance of the proposed solution is close to that obtained with exhaustive search but at only 0.66$\%$of its time complexity.
Qi Liu 0029, Hui Yuan 0001, Junhui Hou, Raouf Hamzaoui, Honglei Su
IEEE Trans. Multim.4
2020 Feature learning for Human Activity Recognition using Convolutional Neural Networks
abstract
Abstract The use of Convolutional Neural Networks (CNNs) as a feature learning method for Human Activity Recognition (HAR) is becoming more and more common. Unlike conventional machine learning methods, which require domain-specific expertise, CNNs can extract features automatically. On the other hand, CNNs require a training phase, making them prone to the cold-start problem. In this work, a case study is presented where the use of a pre-trained CNN feature extractor is evaluated under realistic conditions. The case study consists of two main steps: (1) different topologies and parameters are assessed to identify the best candidate models for HAR, thus obtaining a pre-trained CNN model. The pre-trained model (2) is then employed as feature extractor evaluating its use with a large scale real-world dataset. Two CNN applications were considered: Inertial Measurement Unit (IMU) and audio based HAR. For the IMU data, balanced accuracy was 91.98% on the UCI-HAR dataset, and 67.51% on the real-world Extrasensory dataset. For the audio data, the balanced accuracy was 92.30% on the DCASE 2017 dataset, and 35.24% on the Extrasensory dataset.
Federico Cruciani, Anastasios Vafeiadis, Chris D. Nugent, Ian Cleland, Paul J. McCullagh, Konstantinos Votis, Dimitrios Giakoumis, Dimitrios Tzovaras, Liming Chen 0001, Raouf Hamzaoui
CCF Trans. Pervasive Comput. Interact.10
2020 Audio content analysis for unobtrusive event detection in smart homes
Anastasios Vafeiadis, Konstantinos Votis, Dimitrios Giakoumis, Dimitrios Tzovaras, Liming Chen 0001, Raouf Hamzaoui
Eng. Appl. Artif. Intell.6
2019 Interactive Subjective Study on Picture-level Just Noticeable Difference of Compressed Stereoscopic Images
abstract
The Just Noticeable Difference (JND) reveals the minimum distortion that the Human Visual System (HVS) can perceive. Traditional studies on JND mainly focus on background luminance adaptation and contrast masking. However, the HVS does not perceive visual content based on individual pixels or blocks, but on the entire image. In this work, we conduct an interactive subjective visual quality study on the Picture-level JND (PJND) of compressed stereo images. The study, which involves 48 subjects and 10 stereoscopic images compressed with H.265 intra coding and JPEG2000, includes two parts. In the first part, we determine the minimum distortion that the HVS can perceive against a pristine stereo image. In the second part, we explore the minimum distortion that each subject perceives against a distorted stereo image. Modeling the distribution of the PJND samples as Gaussian, we obtain their complementary cumulative distribution functions, which are known as Satisfied User Ratio (SUR) functions. Statistical analysis results demonstrate that the SUR is highly dependent on the image contents. The HVS is more sensitive to distortion in images with more texture details. The compressed stereoscopic images and the PJND samples are collected in a data set called SIAT-JSSI, which we release to the public.
Chunling Fan, Yun Zhang 0002, Raouf Hamzaoui, Qingshan Jiang
ICASSP3
2019 Two-Dimensional Convolutional Recurrent Neural Networks for Speech Activity Detection
abstract
Speech Activity Detection (SAD) plays an important role in mobile communications and automatic speech recognition (ASR). Developing efficient SAD systems for real-world applications is a challenging task due to the presence of noise. We propose a new approach to SAD where we treat it as a two-dimensional multilabel image classification problem. To classify the audio segments, we compute their Short-time Fourier Transform spectrograms and classify them with a Convolutional Recurrent Neural Network (CRNN), traditionally used in image recognition. Our CRNN uses a sigmoid activation function, max-pooling in the frequency domain, and a convolutional operation as a moving average filter to remove misclassified spikes. On the development set of Task 1 of the 2019 Fearless Steps Challenge, our system achieved a decision cost function (DCF) of 2.89%, a 66.4% improvement over the baseline. Moreover, it achieved a DCF score of 3.318% on the evaluation dataset of the challenge, ranking first among all submissions.
Anastasios Vafeiadis, Lefteris Fanioudakis, Ilyas Potamitis, Konstantinos Votis, Dimitrios Giakoumis, Dimitrios Tzovaras, Liming Chen 0001, Raouf Hamzaoui
INTERSPEECH8
2019 SUR-Net: Predicting the Satisfied User Ratio Curve for Image Compression with Deep Learning
abstract
The Satisfied User Ratio (SUR) curve for a lossy image compression scheme, e.g., JPEG, characterizes the probability distribution of the Just Noticeable Difference (JND) level, the smallest distortion level that can be perceived by a subject. We propose the first deep learning approach to predict such SUR curves. Instead of the direct approach of regressing the SUR curve itself for a given reference image, our model is trained on pairs of images, original and compressed. Relying on a Siamese Convolutional Neural Network (CNN), feature pooling, a fully connected regression-head, and transfer learning, we achieved a good prediction performance. Experiments on the MCL-JCI dataset showed a mean Bhattacharyya distance between the predicted and the original JND distributions of only 0.072.
Chunling Fan, Hanhe Lin, Vlad Hosu, Yun Zhang 0002, Qingshan Jiang, Raouf Hamzaoui, Dietmar Saupe
QoMEX6
2019 Picture-level just noticeable difference for symmetrically and asymmetrically compressed stereoscopic images: Subjective quality assessment study and datasets
Chunling Fan, Yun Zhang 0002, Huan Zhang 0008, Raouf Hamzaoui, Qingshan Jiang
J. Vis. Commun. Image Represent.4
2018 Peer-to-peer live video streaming with rateless codes for massively multiplayer online games
Christos Bouras, Eliya Buyukkaya, Muneeb Dawood, Raouf Hamzaoui, Vaggelis Kapoulas, Andreas Papazois, Gwendal Simon
Peer-to-Peer Netw. Appl.5
2016 Temporal and Inter-View Consistent Error Concealment Technique for Multiview Plus Depth Video
abstract
Multiview plus depth (MVD) is an emerging video format with many applications, including 3-D television and free viewpoint television. During the broadcast of a compressed MVD video, transmission errors may cause the loss of whole frames, resulting in significant degradation of video quality. Error concealment techniques have been widely used to deal with transmission errors in video communication. However, the existing solutions do not address the requirement that the reconstructed frames should be consistent with neighboring frames, i.e., corresponding pixels should have consistent color information. We propose a new consistency model for error concealment of MVD video that allows one to maintain a high level of consistency between frames of the same view (temporal consistency) and those of neighboring views (inter-view consistency). We then propose an algorithm that uses our model to implement concealment in a consistent way. Simulations with the reference software for the multiview video coding project of the joint video team of the ISO/IEC MPEG and ITU-T VCEG show that our method outperforms benchmark techniques, including a baseline approach based on the boundary matching algorithm, with respect to both reconstruction quality and view consistency.
Shadan Khan Khattak, Thomas Maugey, Raouf Hamzaoui, Pascal Frossard
IEEE Trans. Circuits Syst. Video Technol.3
2015 Resource allocation in underprovisioned multioverlay peer-to-peer live video sharing services
Jiayi Liu 0001, Eliya Buyukkaya, Raouf Hamzaoui, Gwendal Simon
Peer-to-Peer Netw. Appl.4
2013 Fast encoding techniques for Multiview Video Coding
Shadan Khan Khattak, Raouf Hamzaoui, Pascal Frossard
Signal Process. Image Commun.2
2013 Bayesian Early Mode Decision Technique for View Synthesis Prediction-Enhanced Multiview Video Coding
abstract
View synthesis prediction (VSP) is a coding mode that predicts video blocks from synthesised frames. It is particularly useful in a multi-camera setup with large inter-camera distances. Adding a VSP-based SKIP mode to a standard Multiview Video Coding (MVC) framework improves the rate-distortion (RD) performance but increases the time complexity of the encoder. This letter proposes an early mode decision technique for VSP SKIP-enhanced MVC. Our method uses the correlation between the RD costs of the VSP SKIP mode in neighbouring views and Bayesian decision theory to reduce the number of candidate coding modes for a given macroblock. Simulation results showed that our technique can save up to 36.20% of the encoding time without any significant loss in RD performance.
Shadan Khan Khattak, Raouf Hamzaoui, Thomas Maugey, Pascal Frossard
IEEE Signal Process. Lett.2
2012 Level-Based Peer-to-Peer Live Streaming with Rateless Codes
abstract
We propose a peer-to-peer system for streaming user-generated live video. Peers are arranged in levels so that video is delivered at about the same time to all peers in the same level, and peers in a higher level watch the video before those in a lower level. We encode the video bit stream with rate less codes and use trees to transmit the encoded symbols. Trees are constructed to minimize the transmission rate for the source while maximizing the number of served peers and guaranteeing on-time delivery and reliability at the peers. We formulate this objective as a height bounded spanning forest problem with nodal capacity constraint and compute a solution using a heuristic polynomial-time algorithm. We conduct ns-2 simulations to study the trade-off between used bandwidth and video quality for various packet loss rates and link latencies.
Eliya Buyukkaya, Muneeb Dawood, Jiayi Liu 0001, Fen Zhou 0001, Raouf Hamzaoui, Gwendal Simon
ISM6
2012 Minimizing server throughput for low-delay live streaming in content delivery networks
abstract
Large-scale live streaming systems can experience bottlenecks within the infrastructure of the underlying Content Delivery Network. In particular, the "equipment bottleneck" occurs when the fan-out of a machine does not enable the concurrent transmission of a stream to multiple other equipments. In this paper, we aim to deliver a live stream to a set of destination nodes with minimum throughput at the source and limited increase of the streaming delay. We leverage on rateless codes and cooperation among destination nodes. With rateless codes, a node is able to decode a video block of k information symbols after receiving slightly more than k encoded symbols. To deliver the encoded symbols, we use multiple trees where inner nodes forward all received symbols. Our goal is to build a diffusion forest that minimizes the transmission rate at the source while guaranteeing on-time delivery and reliability at the nodes. When the network is assumed to be lossless and the constraint on delivery delay is relaxed, we give an algorithm that computes a diffusion forest resulting in the minimum source transmission rate. We also propose an effective heuristic algorithm for the general case where packet loss occurs and the delivery delay is bounded. Simulation results for realistic settings show that with our solution the source requires only slightly more than the video bit rate to reliably feed all nodes.
Fen Zhou 0001, Eliya Buyukkaya, Raouf Hamzaoui, Gwendal Simon
NOSSDAV4
2012 Peer-to-peer live streaming for Massively Multiplayer Online Games
abstract
One of the most attractive features of Massively MuItiplayer Online Games (MMOGs) is the possibility for users to interact with a large number of other users in a variety of collaborative and competitive situations. Garners within an MMOG typically become members of active communities with mutual interests, shared adventures, and common objectives. This demonstration presents a peer-to-peer live video system that enables MMOG players to stream screen-captured video of their game. Players can use the system to show their skills, share experience with friends, or coordinate missions in strategy games.
Christos Bouras, Eliya Buyukkaya, Raouf Hamzaoui, Andreas Papazois, Alex Shani, Gwendal Simon, Fen Zhou 0001
P2P4
2012 Low-complexity multiview video coding
abstract
We consider the problem of complexity reduction in Multiview Video Coding (MVC). We provide a unique comprehensive study that integrates and compares the different low complexity encoding techniques that have been proposed at different levels of the MVC system. In addition, we propose a novel complexity reduction method that takes advantage of the relationship between disparity vectors along time. The relationship is exploited with respect to the motion activity in the frame, as well as with the position of the frame in the Group of Pictures. We integrate this technique into our unique comprehensive framework and evaluate the performance of the resulting system in different setups. We show that the effective combination of complexity reduction techniques results in saving up to 93% in encoding time at the cost of only 0.08 dB in peak signal-to-noise ratio (PSNR) and 1.64% increase in bitrate compared to the standard MVC implementation (JMVM 6.0).
Shadan Khan Khattak, Raouf Hamzaoui, Pascal Frossard
PCS2
2011 Unequal Error Protection Using Fountain Codes With Applications to Video Communication
abstract
Application-layer forward error correction (FEC) is used in many multimedia communication systems to address the problem of packet loss in lossy packet networks. One powerful form of application-layer FEC is unequal error protection which protects the information symbols according to their importance. We propose a method for unequal error protection with a Fountain code. When the information symbols were partitioned into two protection classes (most important and least important), our method required a smaller transmission bit budget to achieve low bit error rates compared to the two state-of-the-art techniques. We also compared our method to the two state-of-the-art techniques for video unicast and multicast over a lossy network. Simulations for the scalable video coding (SVC) extension of the H.264/AVC standard showed that our method required a smaller transmission bit budget to achieve high-quality video.
Raouf Hamzaoui, Marwan Al-Akaidi
IEEE Trans. Multim.2
2010 Adaptive Unicast Video Streaming With Rateless Codes and Feedback
abstract
Video streaming over the Internet and packet-based wireless networks is sensitive to packet loss, which can severely damage the quality of the received video. To protect the transmitted video data against packet loss, application-layer forward error correction (FEC) is commonly used. Typically, for a given source block, the channel code rate is fixed in advance according to an estimation of the packet loss rate. However, since network conditions are difficult to predict, determining the right amount of redundancy introduced by the channel encoder is not obvious. To address this problem, we consider a general framework where the sender applies rateless erasure coding to every source block and keeps on transmitting the encoded symbols until it receives an acknowledgment from the receiver indicating that the block was decoded successfully. Within this framework, we design transmission strategies that aim at minimizing the expected bandwidth usage while ensuring successful decoding subject to an upper bound on the packet loss rate. In real simulations over the Internet, our solution outperformed standard FEC and hybrid automatic repeat request approaches. For the quarter common intermediate formatForemansequence compressed with the H.264 video coder, the gain in average peak signal to noise ratio over the best previous scheme exceeded 3.5 dB at 90 kb/s.
Raouf Hamzaoui, Marwan Al-Akaidi
IEEE Trans. Circuits Syst. Video Technol.2
2010 Introduction of the TCSVT Associate Editors
Roberto Rinaldo, Alberto Signoroni, Raouf Hamzaoui, Wenjun Zhang 0001, Riccardo Bernardini, Francesco G. B. De Natale, Rastislav Lukac, Anthony Vetro, Houqiang Li
IEEE Trans. Circuits Syst. Video Technol.3
2009 Optimal Packet Loss Protection of Progressively Compressed 3-D Meshes
abstract
We consider a state-of-the-art system that uses layered source coding and forward error correction with Reed-Solomon codes to efficiently transmit 3-D meshes over lossy packet networks. Given a transmission bit budget, the performance of this system can be optimized by determining how many layers should be sent, how each layer should be packetized, and how many parity bits should be allocated to each layer such that the expected distortion at the receiver is minimum. The previous solution for this optimization problem uses exhaustive search, which is not feasible when the transmission bit budget is large. We propose instead an exact algorithm that solves this optimization problem in linear time and space. We illustrate the advantages of our approach by providing experimental results for the compressed progressive meshes (CPM) mesh compression technique.
Raouf Hamzaoui, Marwan Al-Akaidi
IEEE Trans. Multim.2
2007 Efficient Rate-Distortion Optimized Media Streaming for Tree-Structured Packet Dependencies
abstract
When streaming packetized media data over a lossy packet network, it is desirable to use transmission strategies that minimize the expected distortion subject to a constraint on the expected transmission rate. Because the computation of such optimal strategies is usually an intractable problem, fast heuristic techniques are often used. We first show that when the graph that gives the decoding dependencies between the data packets is reducible to a tree, optimal transmission strategies can be efficiently computed with dynamic programming algorithms. The proposed algorithms are much faster than other exact algorithms developed for arbitrary dependency graphs. They are slower than previous heuristic techniques but can provide much better solutions. We also show how to apply our algorithms to find high-quality approximate solutions when the dependency graph is not tree reducible. To validate our approach, we run simulations for MPEG1 and H.264 video data. We first consider a simulated packet erasure channel. Then we implement a real video streaming system and provide experimental results for an Internet connection.
Martin Röder, Jean Cardinal, Raouf Hamzaoui
IEEE Trans. Multim.3
2006 Constructing Dependency Trees for Rate-Distortion Optimized Media Streaming
abstract
Finding adequate packet transmission strategies for media streaming systems is a challenging algorithmic task. Recently, we proposed an efficient dynamic programming algorithm for streams in which the dependencies between packets, such as those prescribed between video frames by video codecs, can be modeled with a tree. In this contribution, we propose a heuristic algorithm for arbitrary dependency graphs. This algorithm consists of first transforming the dependency graph into a tree by adding dependencies, and then applying the dynamic programming algorithm on the tree thus obtained. The algorithm is both simple and efficient, as shown by experimental results on video sequences
Martin Röder, Jean Cardinal, Raouf Hamzaoui
ICASSP (5)3
2006 Optimal Error Protection of Progressively Compressed 3D Meshes
abstract
Given a number of available layers of source data and a transmission bit budget, we propose an algorithm that determines how many layers should be sent and how many protection bits should be allocated to each transmitted layer such that the expected distortion at the receiver is minimum. The algorithm is used for robust transmission of progressively compressed 3D models over a packet erasure channel. In contrast to the previous approach, which uses exhaustive search, the time complexity of our algorithm is linear in the transmission bit budget
Raouf Hamzaoui
ICME2
2006 Fast tree-trellis list Viterbi decoding
abstract
A list Viterbi algorithm (LVA) finds the n most likely paths in a trellis diagram of a convolutional code. One of the most efficient LVAs is the tree-trellis algorithm of Soong and Huang. We propose a new implementation of this algorithm. Instead of storing the candidate paths in a single list sorted according to the metrics of the paths, we show that it is computationally more efficient to use several unsorted lists, where all paths of the same list have the same metric. For an arbitrary integer bit metric, both the time and space complexity of our implementation are linear in n. Experimental results for a binary symmetric channel and an additive white Gaussian noise channel show that our implementation is much faster than all previous LVAs.
Martin Röder, Raouf Hamzaoui
IEEE Trans. Commun.2
2006 Branch and bound algorithms for rate-distortion optimized media streaming
abstract
We consider the problem of rate-distortion optimized streaming of packetized multimedia data over a single quality-of-service network using feedback and retransmissions. For a single data unit, we prove that the problem is NP-hard and provide efficient branch and bound algorithms that are much faster than the previously best solution based on dynamic programming. For a group of interdependent data units, we show how to compute optimal solutions with branch and bound algorithms. The branch and bound algorithms for a group of data units are much slower than the current state of the art, a heuristic technique known as sensitivity adaptation. However, in many real-world situations, they provide a significantly better rate-distortion performance.
Martin Röder, Jean Cardinal, Raouf Hamzaoui
IEEE Trans. Multim.3
2005 Dynamic programming algorithm for rate-distortion optimized media streaming
abstract
We propose a dynamic programming algorithm for finding optimal transmission policies for a single packet in rate-distortion optimized media streaming. The algorithm relies on an optimality assumption holding in particular when both the forward and round trip times have exponential distributions. In the other cases, we use the assumption as a heuristic principle. Simulations show that for realistic channel models, the algorithm provides optimal solutions and can be significantly faster than the previous fastest exact algorithm. The proposed algorithm can be used as a preprocessing step for streaming mutually dependent packets.
Martin Röder, Jean Cardinal, Raouf Hamzaoui
ICIP (2)3
2005 Robust layered multiple description coding of scalable media data for multicast
abstract
Layered multiple description codes allow robust transmission of scalable media data over packet erasure networks, while providing simple rate adaptation and bandwidth savings for shared bottleneck links. We show how to efficiently design layered multiple description codes for multicast and broadcast applications in memoryless packet erasure networks. Our approach offers a significantly better quality tradeoff among clients than the best previous solution.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
IEEE Signal Process. Lett.2
2005 Fast Algorithm for Distortion-Based Error Protection of Embedded Image Codes
abstract
We consider a joint source-channel coding system that protects an embedded bitstream using a finite family of channel codes with error detection and error correction capability. The performance of this system may be measured by the expected distortion or by the expected number of correctly decoded source bits. Whereas a rate-based optimal solution can be found in linear time, the computation of a distortion-based optimal solution is prohibitive. Under the assumption of the convexity of the operational distortion-rate function of the source coder, we give a lower bound on the expected distortion of a distortion-based optimal solution that depends only on a rate-based optimal solution. Then, we propose a local search (LS) algorithm that starts from a rate-based optimal solution and converges in linear time to a local minimum of the expected distortion. Experimental results for a binary symmetric channel show that our LS algorithm is near optimal, whereas its complexity is much lower than that of the previous best solution.
Raouf Hamzaoui, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Image Process.1
2004 On the Complexity of Rate-Distortion Optimal Streaming of Packetized Media
abstract
We consider the problem of rate-distortion optimal streaming of packetized media with sender-driven transmission over a single-QoS network using feedback and retransmissions. For a single data unit, we prove that the problem is NP-hard and provide efficient branch and bound algorithms that are in practice much faster than the best known solution. For a group of interdependent data units, we show how to compute optimal solutions with branch and bound algorithms. The branch and bound algorithms for a group of data units are slower than the current state of the art, the heuristic sensitivity adaptation algorithm, but provide a significantly better rate-distortion performance in many real-world situations.
Martin Röder, Jean Cardinal, Raouf Hamzaoui
Data Compression Conference3
2004 Real-time error protection of embedded codes for packet erasure and fading channels
abstract
Reliable real-time transmission of packetized embedded multimedia data over noisy channels requires the design of fast error control algorithms. For packet erasure channels, efficient forward error correction is obtained by using systematic Reed-Solomon (RS) codes across packets. For fading channels, state-of-the-art performance is given by a product channel code where each column code is an RS code and each row code is a concatenation of an outer cyclic redundancy check code and an inner rate-compatible punctured convolutional code. For each of these two systems, we propose a low-memory linear-time iterative improvement algorithm to compute an error protection solution. Experimental results for the two-dimensional and three-dimensional set partitioning in hierarchical trees coders showed that our algorithms provide close to optimal average peak signal-to-noise ratio performance, and that their running time is significantly lower than that of all previously proposed solutions.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
IEEE Trans. Circuits Syst. Video Technol.2
2004 Efficient channel code rate selection algorithms for forward error correction of packetized multimedia bitstreams in varying channels
abstract
We study joint source-channel coding systems for the transmission of images over varying channels without feedback. We consider the situation where the channel statistics are unknown to the transmitter and focus on systems that enable good performance over a wide range of channel conditions. We first propose a linear-time channel code rate selection algorithm for a hybrid transmission system that combines packetization of an embedded wavelet bitstream into independently decodable packets and forward error correction with a concatenated cyclic redundancy check/rate-compatible punctured convolutional (RCPC) channel coder. We then consider an extension of this hybrid system with additional Reed-Solomon (RS) coding across the packets and give a linear-time algorithm for the efficient selection of both the RS and RCPC code rates. Experimental results for a wireline/wireless link modeled as the combination of a packet erasure channel and a Rayleigh flat-fading channel showed that our schemes significantly outperformed the best previous forward error correction systems in many situations where the actual channel parameter values deviated from the ones used in the optimization of the source-channel rate allocation.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
IEEE Trans. Multim.2
2003 Fast forward error protection of packetized multimedia bitstreams for transmission over varying channels
abstract
We propose a real-time optimization algorithm that selects an appropriate channel code for hybrid systems that combine packetization of an embedded wavelet bitstream into independently decodable packets and forward error correction using a family of channel codes with error detection and error correction capability. Such systems are very powerful for the transmission of audio, images, and video over fading and erasure channels with varying statistics. We also give an implementation that uses an optimal packetization technique and a concatenated cyclic redundancy check/rate-compatible punctured convolutional coder. Experimental results show that the peak signal-to-noise ratio of the average mean square error of our system is up to 1.74 dB higher than that of the previous best hybrid system for a Rayleigh fading channel and a transmission rate of 0.25 bits per pixel. Finally, we compare the hybrid approach to a state-of-the-art approach that uses a product code to protect the information bitstream.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICC2
2003 Influence of channel fluctuations on optimal real-time scalable image transmission
abstract
Joint source-channel coding systems using scalable source codes and forward error correction allow reliable transmission of multimedia data over noisy channels. The performance of such systems highly depends on the source-channel bit allocation strategy. Rate-based error protection schemes, which maximize the expected source rate are very attractive for real-time applications because the optimization can be done very quickly and is independent of the source. In real-world communication, channel conditions are varying in time. Thus, it is important to frequently update the error protection. For two state-of-the-art joint source-channel coding systems, we show that a channel mismatch can lead to a poor performance. We study theoretically and experimentally the dependency of a rate-based optimal protection on the channel statistics and provide an efficient strategy for adjusting the error protection when a channel mismatch occurs.
Vladimir Stankovic 0001, Raouf Hamzaoui, Dietmar Saupe
ICIP (1)2
2003 Product code error protection of packetized multimedia bitstreams
abstract
Sherwood and Zeger (1997) proposed a source-channel coding system where the source code is an embedded bitstream and the channel code is a product code such that each row code is a concatenation of a cyclic redundancy check (CRC) and rate-compatible punctured convolutional codes (RCPC) and the column codes are Reed-Solomon (RS) codes. We improve this system for wireless applications by efficiently reorganizing the source code into a set of independently decodable packets, which makes it more robust in varying channels. We also give a linear-time algorithm for finding an optimal equal error protection for the resulting system. Experimental results show that the performance of our system significantly outperforms that of the current state-of-the-art in fading channels with varying statistics.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICIP (1)2
2003 Model-based real-time progressive transmission of images over noisy channels
abstract
Many unequal error protection algorithms used in image communication systems need the operational distortion-rate (D/R) curve of the source coder whose computation is time-consuming. We study the use of parametric models instead of the true D/R curves for wavelet-based embedded image and video coders. We propose a Weibull model and show its superiority to the previous models for real-time applications. For unequal error protection over binary symmetric and packet erasure channels, the Weibull model yielded performance similar to the one obtained with the true D/R curve while satisfying the real-time constraint.
Youssef Charfi, Raouf Hamzaoui, Dietmar Saupe
WCNC2
2003 Real-time unequal error protection for distortion-optimal progressive image transmission
abstract
For optimal progressive transmission of an embedded image code over a noisy channel, we consider an unequal error protection strategy that minimizes the average of the expected distortion over a set of intermediate rates. In contrast to previous work, we find a near-optimal solution in real-time. For a binary symmetric channel, two state-of-the-art source coders (SPIHT and JPEG200), and a rate-compatible punctured turbo coder as a channel coder, we compare our solution to the strategy that optimizes the end-to-end performance.
Vladimir Stankovic 0001, Youssef Charfi, Raouf Hamzaoui, Zixiang Xiong
WCNC3
2003 Real-time unequal error protection algorithms for progressive image transmission
abstract
We consider unequal error protection strategies for the efficient progressive transmission of embedded image codes over noisy channels. In progressive transmission, the reconstruction quality is important not only at the target transmission rate but also at the intermediate rates. An adequate error protection strategy may, thus, consist of optimizing the average performance over the set of intermediate rates. The performance can be the expected number of correctly decoded source bits or the expected distortion. For the rate-based performance, we prove some interesting properties of an optimal solution and give an optimal linear-time algorithm to compute it. For the distortion-based performance, we propose an efficient linear-time local search algorithm. For a binary symmetric channel, two state-of-the-art source coders (SPIHT and JPEG2000), we compare the progressive ability of our proposed solutions to that of the strategies that optimize the end-to-end performance of the system. Experimental results showed that the proposed solutions had a slightly worse performance at the target transmission rate and a better performance at most of the intermediate rates, especially at the lowest ones.
Vladimir Stankovic 0001, Raouf Hamzaoui, Youssef Charfi, Zixiang Xiong
IEEE J. Sel. Areas Commun.2
2003 Fast algorithm for rate-based optimal error protection of embedded codes
abstract
Embedded image codes are very sensitive to channel noise because a single bit error can lead to an irreversible loss of synchronization between the encoder and the decoder. P.G. Sherwood and K. Zeger (see IEEE Signal Processing Lett., vol.4, p.191-8, 1997) introduced a powerful system that protects an embedded wavelet image code with a concatenation of a cyclic redundancy check coder for error detection and a rate-compatible punctured convolutional coder for error correction. For such systems, V. Chande and N. Farvardin (see IEEE J. Select. Areas Commun., vol.18, p.850-60, 2000) proposed an unequal error protection strategy that maximizes the expected number of correctly received source bits subject to a target transmission rate. Noting that an optimal strategy protects successive source blocks with the same channel code, we give an algorithm that accelerates the computation of the optimal strategy of Chande and Farvardin by finding an explicit formula for the number of occurrences of the same channel code. Experimental results with two competitive channel coders and a binary symmetric channel showed that the speed-up factor over the approach of Chande and Farvardin ranged from 2.82 to 44.76 for transmission rates between 0.25 and 2 bits per pixel.
Vladimir Stankovic 0001, Raouf Hamzaoui, Dietmar Saupe
IEEE Trans. Commun.2
2002 Rate-Based versus Distortion-Based Optimal Joint Source-Channel Coding
abstract
We consider a joint source-channel coding system that protects an embedded wavelet bitstream against noise using a finite family of channel codes with error detection and error correction capability. The performance of this system may be measured by the expected distortion or by the expected number of correctly received source bits subject to a target total transmission rate. Whereas a rate-based optimal solution can be found in linear time, the computation of a distortion-based optimal solution is prohibitive. Under the assumption of the convexity of the operational distortion-rate function of the source coder, we give a lower bound on the expected distortion of a distortion-based optimal solution that depends only on a rate-based optimal solution. Then we show that a distortion-based optimal solution provides a stronger error protection than a rate-based optimal solution and exploit this result to reduce the time complexity of the distortion-based optimization. Finally, we propose a fast iterative improvement algorithm that starts from a rate-based optimal solution and converges to a local minimum of the expected distortion. Experimental results for a binary symmetric channel with the SPIHT coder and JPEG 2000 show that our lower bound is close to optimal. Moreover, the solution given by our local search algorithm has about the same quality as a distortion-based optimal solution, whereas its complexity is much lower than that of the previous best solution.
Raouf Hamzaoui, Vladimir Stankovic 0001
DCC1
2002 Packet loss protection of embedded data with fast local search
abstract
Unequal loss protection with systematic Reed-Solomon codes allows reliable transmission of embedded multimedia over packet erasure channels. The design of a fast algorithm with low memory requirements for the computation of an unequal loss protection solution is essential in real-time systems. Because the determination of an optimal solution is time-consuming, fast suboptimal solutions have been used. In this paper, we present a fast iterative improvement algorithm with negligible memory requirements. Experimental results for the JPEG2000, 2D, and 3D set partitioning in hierarchical trees (SPIHT) coders showed that our algorithm provided close to optimal peak signal-to-noise ratio (PSNR) performance, while its time complexity was significantly lower than that of all previously proposed algorithms.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICIP (2)2
2002 Fast list Viterbi decoding and application for source-channel coding of images
abstract
A list Viterbi algorithm (LVA) finds the n best paths in a trellis. We propose a new implementation of the tree-trellis LVA. Instead of storing all paths in a single sorted list, we show that it is more efficient to use several lists, where all paths of the same list have the same metric. For an integer metric, both the time and space complexity of our implementation are linear in n. Experimental results show that our implementation is much faster than all previous LVAs. This allows us to consider a large number of paths in acceptable time, which significantly improves the performance of a popular progressive source-channel coding system that protects embedded data with a concatenation of an outer error detecting code and an inner error correcting convolutional code.
Martin Röder, Raouf Hamzaoui
ICME (1)2
2002 Joint product code optimization for scalable multimedia transmission over wireless channels
abstract
State-of-the-art systems for the transmission of images over wireless channels generate an embedded bitstream and protect it with a product code where the row code is a concatenation of an outer cyclic redundancy check (CRC) code and an inner rate-compatible punctured convolutional (RCPC) code, and the column code is a Reed-Solomon (RS) code. In previous works, the product code was optimized by searching for the best RS protection for each RCPC code rate. We present a local search algorithm that jointly optimizes the RS and the RCPC codes. Experimental results show that our algorithm provides an approximately optimal solution, while its time complexity is much lower than that of the previous works.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICME (1)2
2001 Progressive fractal coding
abstract
Progressive coding is an important feature of compression schemes. Wavelet coders are well suited for this purpose because the wavelet coefficients can be naturally ordered according to decreasing importance. Progressive fractal coding is feasible, but it was proposed only for hybrid fractal-wavelet schemes. We introduce a progressive fractal image coder in the spatial domain. A Lagrange optimization based on rate-distortion performance estimates determines an optimal ordering of the code bits. The optimality is in the sense that the reconstruction error is monotonically decreasing and minimum at intermediate rates. The decoder recovers this ordering without side information. As a side effect, our work motivates improved bit allocation strategies for fractal coding.
Iván Kopilovic, Dietmar Saupe, Raouf Hamzaoui
ICIP (1)3
2001 Rate-distortion unequal error protection for fractal image codes
abstract
Fractal image codes are very sensitive to bit errors because the decoding of a block is dependent not only on the code information associated to this block but also on the code information associated to other blocks. We analyze the sensitivity of a fractal code to transmission errors in a binary symmetric channel. We provide two rate-distortion unequal error protection techniques that allocate the code bits to protection classes in a nearly optimal way. We give an implementation for BCH and RCPC channel codes and show that rate-compatible punctured convolutional (RCPC) codes are preferable. For a binary symmetric channel bit error probability of 0.1 and a total code rate of 0.5 bpp, the loss in reconstruction quality with our best implementation was about 3.38 dB for the 512 /spl times/ 512 Lenna image, yielding a PSNR of 27.12 dB.
Vladimir Stankovic 0001, Dietmar Saupe, Raouf Hamzaoui
ICIP (1)3
2001 Distortion Minimization with Fast Local Search for Fractal Image Compression
Raouf Hamzaoui, Dietmar Saupe, Michael Hiller
J. Vis. Commun. Image Represent.1
2000 Fast Code Enhancement with Local Search for Fractal Image Compression
abstract
Optimal fractal coding consists of finding, in a finite set of contractive affine mappings, one whose unique fixed point is closest to the original image. Optimal fractal coding is an NP-hard combinatorial optimization problem. Conventional coding is based on a greedy suboptimal algorithm known as collage coding. In a previous study, we proposed a local search algorithm that significantly improves on collage coding. However the algorithm, which requires the computation of many fixed points, is computationally expensive. In this paper we provide techniques that drastically reduce the time complexity of the algorithm.
Raouf Hamzaoui, Dietmar Saupe, Michael Hiller
ICIP1
2000 Local iterative improvement of fractal image codes
Raouf Hamzaoui, Hannes Hartenstein, Dietmar Saupe
Image Vis. Comput.1
2000 Combining fractal image compression and vector quantization
abstract
In fractal image compression, the code is an efficient binary representation of a contractive mapping whose unique fixed point approximates the original image. The mapping is typically composed of affine transformations, each approximating a block of the image by another block (called domain block) selected from the same image. The search for a suitable domain block is time-consuming. Moreover, the rate distortion performance of most fractal image coders is not satisfactory. We show how a few fixed vectors designed from a set of training images by a clustering algorithm accelerates the search for the domain blocks and improves both the rate-distortion performance and the decoding speed of a pure fractal coder, when they are used as a supplementary vector quantization codebook. We implemented two quadtree-based schemes: a fast top-down heuristic technique and one optimized with a Lagrange multiplier method. For the 8 bits per pixel (bpp) luminance part of the 512 x 512 Lena image, our best scheme achieved a peak-signal-to-noise ratio of 32.50 dB at 0.25 bpp.
Raouf Hamzaoui, Dietmar Saupe
IEEE Trans. Image Process.1
1999 A Video Codec Based on R/D-Optimized Adaptive Vector Quantization
abstract
Summary form only given. We present a new AVQ-based video coder for very low bitrates. To encode a block from a frame, the encoder offers three modes: (1) a block from the same position in the last frame can be taken; (2) the block can be represented with a vector from the codebook; or (3) a new vector, that sufficiently represents a block, can be inserted into the codebook. For mode 2 a mean-removed VQ scheme is used. The decision on how blocks are encoded and how the codebook is updated is done in an rate-distortion (R-D) optimized fashion. The codebook of shape blocks is updated once per frame. First results for an implementation of such a scheme have been reported previously. Here we extend the method to incorporate a wavelet image transform before coding in order to enhance the compression performance. In addition the rate-distortion optimization is comprehensively discussed. Our R-D optimization is based on an efficient convex-hull computation. This method is compared to common R-D optimizations that use a Lagrangian multiplier approach. In the discussion of our R-D method we show the similarities and differences between our scheme and the generalized threshold replenishment (GTR) method of Fowler et al. (1997). Furthermore, we demonstrate that the translation of our R-D optimized AVQ into the wavelet domain leads to an improved coding performance. We present coding results that show that one can achieve the same encoding quality as with comparable standard transform coding (H.263). In addition we offer an empirical analysis of the short- and long-term behavior of the adaptive codebook. This analysis indicates that the AVQ method uses the vectors in its codebook for some kind of long-term prediction.
Marcel Wagner, Ralf Herz, Hannes Hartenstein, Raouf Hamzaoui, Dietmar Saupe
Data Compression Conference4
1998 Rate-Distortion based Video Coding with Adaptive Mean-Removed Vector Quantization
Raouf Hamzaoui, Dietmar Saupe, Marcel Wagner
ICIP (3)1
1998 Optimal Hierarchical Partitions for Fractal Image Compression
abstract
In fractal image compression a partitioning of the image is required. In this paper we discuss the construction of rate-distortion optimal partitions. We begin with a fine scale partition which gives a fractal encoding with a high bit rate and a low distortion. The partition is hierarchical, thus, corresponds to a tree. We employ a pruning strategy based on the generalized BFOS algorithm. It extracts subtrees corresponding to partitions and fractal encodings which are optimal in the rate-distortion sense. First results are included for the case of fractal encodings based on rectangular (HV) partitions. We also provide a comparison with greedy partitions based on the traditional collage error criterion or just using block variance.
Dietmar Saupe, Matthias Ruhl, Raouf Hamzaoui, Luigi Grandi, Daniele Marini
ICIP (1)3
1997 Quadtree Based Variable Rate Oriented Mean Shape-Gain Vector Quantization
abstract
Mean shape-gain vector quantization (MSGVQ) is extended to include negative gains and square isometries. Square isometries together with a classification technique based on average block intensities enable us to enlarge the MSGVQ codebook size without any additional storage requirements while keeping the complexity of both the codebook generation and the encoding manageable. Variable rate codes are obtained with a quadtree segmentation based on a rate-distortion criterion. Experimental results show that our scheme performs favorably when compared to previous product code techniques or quadtree based VQ methods.
Raouf Hamzaoui, Bertram Ganz, Dietmar Saupe
Data Compression Conference1
1996 VQ-enhanced fractal image compression
abstract
A novel hybrid scheme combining fractal image compression with mean-removed shape-gain vector quantization is presented. The scheme uses a small set of VQ codebook blocks as a block classifier for the domain blocks and as an alternative means of coding when able to provide a satisfying distortion. Our scheme is shown to improve the performance of conventional fractal coding in all its aspects. The rate-distortion curve is ameliorated, and both the encoding and the decoding are faster.
Raouf Hamzaoui, Dietmar Saupe
ICIP (1)1