VLDB 2026 Research / reviewers in the wild / expert
Anique Akhtar
dblp:157/8456
· DBLP profile ↗
16ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0003-2701-6611ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 12 since 2021Computer networks · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InterGS-Lite: Light Weight Dynamic GS Coding with Vector Quantization of Prediction Residualsabstract3D Gaussian splatting enables real-time, photo-realistic scene rendering but is challenged by redundant, high-dimensional primitive data. To make immersive AR/VR and volumetric communication practical, compression must address both temporal and attribute redundancy. We introduce a streamlined inter-prediction coding pipeline that predicts each P-frame's color and geometry attributes from a nearby reference using a lightweight hybrid predictor, combining bilateral filtering and K-nearest-neighbor feature transfer, and encodes only residuals. Our approach targets the most rate-critical channels by applying codebook-based vector quantization to spherical harmonic components, while retaining efficient scalar quantization for other attributes. All quantized streams and codebooks are entropy coded. The proposed method achieves an average 89.1 % reduction in BD-rate and a 12.39 dB increase in BD-PSNR over GPCCv1, and outperforms previous SOTA interGS by 18.6 % in BD-rate and 1.81 dB in BD-PSNR. With practical decoding speeds (0.605 seconds per frame), interGS-Lite preserves real-time rendering and has potential for streaming and AR/VR applications, delivering consistent bitrate reductions with high visual quality. Zhu Li 0001, Anique Akhtar, Geert Van der Auwera |
DCC | 3 |
| 2026 | FDIP-PCAC: Frequency Domain Inter Prediction and Compensation for Dynamic Point Cloud Attributes CompressionabstractA point cloud is a 3D data representation that presents unique challenges due to its large volume, high dimensionality, and lack of structure. This work introduces FDIP-PCAC, a$F$requency$D$omain$I$nter$P$rediction and Compensation system for dynamic point cloud attribute compression. The method combines deep learning, geometry-driven modeling, and Graph Fourier Transform (GFT) within a five-phase encoder–decoder framework. Initially, the reference and target point cloud frames are partitioned into corresponding nodes using a binary tree. In Phase 1, each target node identifies its K-nearest reference nodes based on the geometry centroid. Phase 2 computes the GFT latent representation of these selected reference nodes. Phase 3 and 4 transfer the RGB colors from matched reference nodes to target nodes, generating Predicted Target Node and Second-Stage Predicted Target Node, respectively. In phase 5, the neural network-based FreqNeRF predicts the GFT latent representation of target nodes. The proposed system achieves superior compression and reconstruction performance compared to G-PCC across diverse datasets. It delivers average BD-rate reductions of 65.89% against RAHT-v30, 63.61% against PredLift-v30, 67.59% against RAHT-v23 and 66.82% against PredLift-v23. Sajid Umair, Zhu Li 0001, Anique Akhtar, Geert Van der Auwera |
DCC | 3 |
| 2025 | InterGS: Inter-Predictive Coding of Gaussian Splatting SequencesabstractDynamic multi-view video is increasingly important for applications in VR/AR, telepresence, and volumetric streaming. 3D Gaussian Splatting (3DGS) has emerged as a powerful learnable representation enabling real-time rendering of dynamic scenes, but each frame typically contains hundreds of thousands of Gaussian primitives, leading to prohibitive storage and transmission costs when compressed independently. To address this, we introduce InterGS, the inter-prediction framework designed specifically for Gaussian splatting sequences. InterGS exploits temporal coherence by predicting each P-frame’s high-dimensional color and geometry attributes from a preceding I-frame using a lightweight hybrid predictor that unifies bilateral filter prediction and nearest-neighbor feature prediction. A simple adaptive selector chooses the most accurate prediction for each Gaussian, and only the resulting prediction residuals are quantized and entropy-coded. Overall coding scheme is light weight in computational complexity, and highly parallelizeable for potential real time streaming applications. On three sequences from the MPEG-GSC benchmark, InterGS achieves an average BD-rate reduction of 88.31%, and on two AVS sequences it delivers a 64.47% reduction. These results demonstrate that InterGS offers a practical, high-efficiency solution for dynamic Gaussian splatting compression. Zhu Li 0001, Anique Akhtar, Geert Van der Auwera |
VCIP | 3 |
| 2025 | GeRF3D-PCAC: Dynamic Point Cloud Attributes Compression with GFT Domain Motion CompensationabstractA point cloud is a standard 3D data representation that poses challenges due to its large volume, high dimensionality, and unstructured nature. This paper introduces the Graph Fourier Transform Neural Radiance Field for Dynamic 3D Point Cloud Attribute Coding, termed GeRF3D-PCAC. We present a geometry-guided compression method for dynamic point cloud attributes using a four-step encoder-decoder framework. First, a binary tree partitions both reference(Ft) and target(Ft+1) point cloud frames, and each target partition locates its top-k nearest reference partitions using centroid-based KNN search. Second, the optimal reference node is chosen by comparing the Graph Fourier Transform (GFT) based latent representations of the target node and its candidates using mean squared error. Third, the target node is predicted by transferring RGB colors from geometry-matching points in the reference frame to the target geometry. Finally, GeRF3D predicts the GFT latent representation of target nodes, and three bitstreams, such as reference indices, quantized GeRF3D weights, and Y-channel residuals in the frequency domain, are generated and transmitted to the decoder. Our method outperforms G-PCC-PredLift-v23, G-PCC-RAHT-v23, and Unicorn in both compression and reconstruction across benchmark datasets. Sajid Umair, Zhu Li 0001, Anique Akhtar, Geert Van der Auwera |
VCIP | 3 |
| 2024 | ResNeRF-PCAC: Super Resolving Residual Learning NeRF for High Efficiency Point Cloud Attributes CodingabstractA point cloud (PC) is a popular 3D data representation that poses challenges due to its size, dimensionality, and unstructured nature. This paper introduces the Residual Neural Radiance Field for Point Cloud Attribute Coding (ResNeRFPCAC), a novel approach for point cloud attribute compression. ResNeRF-PCAC combines sparse convolutions with neural radiance fields, to create a highly efficient attribute coding solution. It initially downscales the point cloud to generate a coarse thumbnail point cloud and encodes it using the G-PCC attribute encoder. The thumbnail PC is upsampled using a super-resolution network to generate a recolored PC. Color attribute residuals are then computed between the original and the super-resolved recolored PC. A ResNeRF network is employed to predict these residuals. The trained ResNeRF weights are compressed into a bitstream. The thumbnail bitstream and the compressed model weights are then transmitted to the decoder. Sparse convolution-based super-resolving network weights are shared and common across all content and need not to be signaled. Experiments on the MPEG-8i dataset demonstrate superior performance in terms of reconstruction quality and compression ratio compared to G-PCCRAHT and G-PCC-Predlift for both v14 and v21. Sajid Umair, Birendra Kathariya, Zhu Li 0001, Anique Akhtar, Geert Van der Auwera |
ICIP | 4 |
| 2024 | Sparse Convolution Based Point Cloud Attributes Deblocking with Graph Fourier Latent RepresentationabstractThe rapid development of 3D modeling and computer vision has made point cloud data essential across various industries. Effective processing, transmission, and storage of these point clouds require innovative filtering and compression methods. Unlike traditional image and video media, point clouds are sparse and non-uniformly sampled, posing unique compression challenges. In this paper, we introduce a novel approach for deblocking and denoising point clouds using multi-scale sparse convolution-based attribute learning in the spectral domain. Our method leverages concurrent voxelized feature embedding for efficiency and utilizes sparsity to build a deeper learning structure. By employing the Graph Fourier Transform (GFT), we better understand spectral patterns and spatial correlations of attributes. Experiments show that our GFT latent representation significantly enhances reconstruction quality, achieving an 18.5% Y-BD rate reduction compared to GPCC TMC13v14 anchors. This extends our previous work, MUSCON, the first SparseConv-based multi-scale attribute upsampling solution for deblocking. Birendra Kathariya, Zhu Li 0001, Anique Akhtar, Geert Van der Auwera |
MMSP | 4 |
| 2024 | PointCU: Multiscale Sparse Convolutional Learning for Point Cloud Color UpsamplingabstractUpsampling a sparse point cloud is a common operation in various applications, such as high-quality reconstruction, rendering. Point-cloud geometry upsampling has been extensively studied for this purpose. However, when a sparse point cloud includes attributes, such as color, their upsampling must also be carried out in addition to geometry. Despite the apparent need of such work, not a lot of effective solutions are developed that can work with real world large scale point cloud. In this work, we propose a novel solution called PointCU, which is a sparse convolution learning-based Point cloud Color Upsampling method that enables high-fidelity dense point color reconstruction from sparse point color. The proposed method first prepares multiple representations of a dense point cloud through voxelization at different scales, then transfers the color from sparse to the newly created point clouds including the dense point cloud itself through devoxelization. Then by learning feature on these multiple point cloud representation through sparse convolution neural network (SparseCNN) while also expanding and fusing the feature to higher scale, PointCU achieves an excellent color super-resolving capability. Our experimental results on four times (4x) and eight times (8x) upsampling tasks demonstrate that the color upsampling performance of the proposed method is superior to the previous known color upsampling schemes by a large margin. Birendra Kathariya, Anique Akhtar, Zhu Li 0001, Geert Van der Auwera |
VCIP | 2 |
| 2024 | Inter-Frame Compression for Dynamic Point Cloud Geometry CodingabstractEfficient point cloud compression is essential for applications like virtual and mixed reality, autonomous driving, and cultural heritage. This paper proposes a deep learning-based inter-frame encoding scheme for dynamic point cloud geometry compression. We propose a lossy geometry compression scheme that predicts the latent representation of the current frame using the previous frame by employing a novel feature space inter-prediction network. The proposed network utilizes sparse convolutions with hierarchical multiscale 3D feature learning to encode the current frame using the previous frame. The proposed method introduces a novel predictor network for motion compensation in the feature domain to map the latent representation of the previous frame to the coordinates of the current frame to predict the current frame's feature embedding. The framework transmits the residual of the predicted features and the actual features by compressing them using a learned probabilistic factorized entropy model. At the receiver, the decoder hierarchically reconstructs the current frame by progressively rescaling the feature embedding. The proposed framework is compared to the state-of-the-art Video-based Point Cloud Compression (V-PCC) and Geometry-based Point Cloud Compression (G-PCC) schemes standardized by the Moving Picture Experts Group (MPEG). The proposed method achieves more than 88% BD-Rate (Bjøntegaard Delta Rate) reduction against G-PCCv20 Octree, more than 56% BD-Rate savings against G-PCCv20 Trisoup, more than 62% BD-Rate reduction against V-PCC intra-frame encoding mode, and more than 52% BD-Rate savings against V-PCC P-frame-based inter-frame encoding mode using HEVC. These significant performance gains are cross-checked and verified in the MPEG working group. Anique Akhtar, Zhu Li 0001, Geert Van der Auwera |
IEEE Trans. Image Process. | 1 |
| 2022 | Dynamic Point Cloud InterpolationabstractDense photorealistic point clouds can depict real-world dynamic objects in high resolution and with a high frame rate. Frame interpolation of such dynamic point clouds would enable the distribution, processing, and compression of such content. In this work, we propose a first point cloud interpolation framework for photorealistic dynamic point clouds. Given two consecutive dynamic point cloud frames, our framework aims to generate intermediate frame(s) between them. The proposed deep learning framework has three major components: the encoder module, the fusion network, and the multi-scale point cloud synthesis module. The encoder module extracts multi-scale features from two consecutive frames. The fusion network employs a novel 4D feature learning technique to merge the multi-scale features from consecutive frames. Finally, the multi-scale point cloud synthesis module hierarchically reconstructs the interpolated point cloud intermediate frame at different resolutions. We evaluate our framework on high-resolution point cloud datasets used in MPEG, JPEG Pleno, and AVS standards. The quantitative and qualitative results demonstrate the effectiveness of the proposed method. Anique Akhtar, Zhu Li 0001, Geert Van der Auwera, Jianle Chen |
ICASSP | 1 |
| 2022 | PU-Dense: Sparse Tensor-Based Point Cloud Geometry UpsamplingabstractDue to the increased popularity of augmented and virtual reality experiences, the interest in capturing high-resolution real-world point clouds has never been higher. Loss of details and irregularities in point cloud geometry can occur during the capturing, processing, and compression pipeline. It is essential to address these challenges by being able to upsample a low Level-of-Detail (LoD) point cloud into a high LoD point cloud. Current upsampling methods suffer from several weaknesses in handling point cloud upsampling, especially in dense real-world photo-realistic point clouds. In this paper, we present a novel geometry upsampling technique, PU-Dense, which can process a diverse set of point clouds including synthetic mesh-based point clouds, real-world high-resolution point clouds, real-world indoor LiDAR scanned objects, as well as outdoor dynamically acquired LiDAR-based point clouds. PU-Dense employs a 3D multiscale architecture using sparse convolutional networks that hierarchically reconstruct an upsampled point cloud geometry via progressive rescaling and multiscale feature extraction. The framework employs a UNet type architecture that downscales the point cloud to a bottleneck and then upscales it to a higher level-of-detail (LoD) point cloud. PU-Dense introduces a novel Feature Extraction Unit that incorporates multiscale spatial learning by employing filters at multiple sampling rates and receptive fields. The architecture is memory efficient and is driven by a binary voxel occupancy classification loss that allows it to process high-resolution dense point clouds with millions of points during inference time. Qualitative and quantitative experimental results show that our method significantly outperforms the state-of-the-art approaches by a large margin while having much lower inference time complexity. We further test our dataset on high-resolution photo-realistic datasets. In addition, our method can handle noisy data well. We further show that our approach is memory efficient compared to the state-of-the-art methods. Anique Akhtar, Zhu Li 0001, Geert Van der Auwera, Li Li 0040, Jianle Chen |
IEEE Trans. Image Process. | 1 |
| 2022 | Video-Based Point Cloud Compression Artifact RemovalabstractPhoto-realistic point cloud capture and transmission are the fundamental enablers for immersive visual communication. The coding process of dynamic point clouds, especially video-based point cloud compression (V-PCC) developed by the MPEG standardization group, is now delivering state-of-the-art performance in compression efficiency. V-PCC is based on the projection of the point cloud patches to 2D planes and encoding the sequence as 2D texture and geometry patch sequences. However, the resulting quantization errors from coding can introduce compression artifacts, which can be very unpleasant for the quality of experience (QoE). In this work, we developed a novel out-of-the-loop point cloud geometry artifact removal solution that can significantly improve reconstruction quality without additional bandwidth cost. Our novel framework consists of a point cloud sampling scheme, an artifact removal network, and an aggregation scheme. The point cloud sampling scheme employs a cube-based neighborhood patch extraction to divide the point cloud into patches. The geometry artifact removal network then processes these patches to obtain artifact-removed patches. The artifact-removed patches are then merged together using an aggregation scheme to obtain the final artifact-removed point cloud. We employ 3D deep convolutional feature learning for geometry artifact removal that jointly recovers both the quantization direction and the quantization noise level by exploiting projection and quantization prior. The simulation results demonstrate that the proposed method is highly effective and can considerably improve the quality of the reconstructed point cloud. Anique Akhtar, Wen Gao 0019, Li Li 0040, Zhu Li 0001, Shan Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Convolutional Neural Network-Based Occupancy Map Accuracy Improvement for Video-Based Point Cloud CompressionabstractIn video-based point cloud compression (V-PCC), a dynamic point cloud is projected onto geometry and attribute videos patch by patch for compression. In addition to the geometry and attribute videos, an occupancy map video is compressed into a V-PCC bitstream to indicate whether a two-dimensional (2D) point in the projected geometry video corresponds to any point in three-dimensional (3D) space. The occupancy map video is usually downsampled before compression to obtain a tradeoff between the bitrate and the reconstructed point cloud quality. Due to the accuracy loss in the downsampling process, some noisy points are generated, which leads to severe objective and subjective quality degradation of the reconstructed point cloud. To improve the quality of the reconstructed point cloud, we propose using a convolutional neural network (CNN) to improve the accuracy of the occupancy map video. We mainly make the following contributions. First, we improve the accuracy of the occupancy map video by formulating the problem as a binary segmentation problem since the pixel values of the occupancy map video are either 0 or 1. Second, in addition to the downsampled occupancy map video, we introduce a reconstructed geometry video as the other input of the CNN to provide more useful information in order to indicate the occupancy map video. To the best of our knowledge, this is the first learning-based work to improve the performance of V-PCC. Compared to state-of-the-art schemes, our proposed CNN-based approach achieves much more accurate occupancy map videos and significant bitrate savings. Li Li 0040, Anique Akhtar, Zhu Li 0001, Shan Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Point Cloud Geometry Prediction Across Spatial Scale using Deep LearningabstractA point cloud is a 3D data representation that is becoming increasingly popular. Due to the large size of a point cloud, the transmission of point cloud is not feasible without compression. However, the current point cloud lossy compression and processing techniques suffer from quantization loss which results in a coarser sub-sampled representation of point cloud. In this paper, we solve the problem of points lost during voxelization by performing geometry prediction across spatial scale using deep learning architecture. We perform an octree-type upsampling of point cloud geometry where each voxel point is divided into 8 sub-voxel points and their occupancy is predicted by our network. This way we obtain a denser representation of the point cloud while minimizing the losses with respect to the ground truth. We utilize sparse tensors with sparse convolutions by using Minkowski Engine with a UNet like network equipped with inception-residual network blocks. Our results show that our geometry prediction scheme can significantly improve the PSNR of a point cloud, therefore, making it an essential post-processing scheme for the compression-transmission pipeline. This solution can serve as a crucial prediction tool across scale for point cloud compression, as well as display adaptation. Anique Akhtar, Wen Gao 0019, Xiang Zhang 0004, Li Li 0040, Zhu Li 0001, Shan Liu 0001 |
VCIP | 1 |
| 2019 | Low Latency Scalable Point Cloud Communication in VANETs using V2I CommunicationabstractMobile edge and vehicle-based depth sending and real-time point cloud communication is an essential subtask enabling autonomous driving. In this paper, we propose a framework for point cloud multicast in VANETs using vehicle to infrastructure (V2I) communication. We employ a scalable Binary Tree embedded Quad Tree (BTQT) point cloud source encoder with bitrate elasticity to match with an adaptive random network coding (ARNC) to multicast different layers to the vehicles. The scalability of our BTQT encoded point cloud provides a trade-off in the received voxel size/quality vs channel condition whereas the ARNC helps maximize the throughput under a hard delay constraint. The solution is tested with the outdoor 3D point cloud dataset from MERL for autonomous driving. The users with good channel conditions receive a near lossless point cloud whereas users with bad channel conditions are still able to receive at least the base layer point cloud. Anique Akhtar, Rubayet Shafin Bradley Shafin, Jianan Bai 0001, Lianjun Li 0001, Lingjia Liu 0001 |
ICC | 1 |
| 2019 | Low Latency Scalable Point Cloud CommunicationabstractMobile edge and V2V Low latency streaming of 3D information is a crucial technology for smart city and autonomous driving. In this paper, we propose a joint source-channel coding framework for transmitting 3D point cloud data to different quality-of-service devices by creating a scalable representation of point cloud. We employ a scalable Binary Tree embedded Quad Tree point cloud encoder with adaptive modulation and coding schemes to guarantee the latency as well as the quality requirement of each user. We perform link level simulations using outdoor 3D point cloud dataset from LiDAR scans for auto-driving. The scalability of our encoded point cloud provides a trade-off in the received voxel size/quality vs channel condition under a hard latency constraint. The users with good channel conditions receive a near lossless point cloud whereas users with bad channel conditions are still able to receive at least the base layer point cloud. Anique Akhtar, Birendra Kathariya, Zhu Li 0001 |
ICIP | 1 |
| 2018 | Directional MAC protocol for IEEE 802.11ad based wireless local area networks
Anique Akhtar, Sinem Coleri Ergen |
Ad Hoc Networks | 1 |