Huiming Zheng

dblp:300/7290 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Ambiguity-Tolerant Cross-Modal Hashing with Partial Labels
abstract
Cross-modal hashing (CMH) has achieved remarkable success in large-scale cross-modal retrieval due to its low storage cost and high computational efficiency. However, most existing CMH methods rely on accurately annotated training data, which is often impractical in real-world applications due to the high cost and limited scalability of data annotation. In practice, annotators typically assign a candidate label set rather than a single precise label to each sample pair, resulting in partial labels with inherent ambiguity. Such ambiguous supervision poses significant challenges to conventional CMH methods that assume reliable and unambiguous labels. In this paper, we investigate a less-touched yet meaningful problem, i.e., cross-modal hashing with partial labels (PLCMH). PLCMH faces two major challenges: label ambiguity and modality-alignment barriers induced by misleading supervision. To address these issues, we propose a new approach named Ambiguity-Tolerant Cross-Modal Hashing (ATCH). Specifically, ATCH presents a Local Consensus Disambiguation (LCD) mechanism that resolves label ambiguity by effectively inferring stable and accurate label confidence based on local consensus within the Hamming space. Moreover, ATCH proposes a Confidence-Aware Contrastive Hashing (CACH) mechanism that derives both pseudo labels and trustworthiness scores from the label confidence vectors to learn discriminative hash codes, leading to effective modality alignment. Extensive experiments on three multimodal datasets demonstrate the superiority of ATCH.
Chao Su 0003, Xu Wang 0028, Yingke Chen, Huiming Zheng, Dezhong Peng, Yuan Sun 0016
AAAI5
2026 DiffPCGC: Efficient Point Cloud Geometry Compression via Diffusion Models
Huiming Zheng
ISCAS1
2026 Side-Information-Aided Deep Unfolding Network for Hybrid-Field Channel Estimation of Wideband Terahertz Ultra-Massive MIMO Systems
Sijia Wu, Hong Jiang 0003, Huiming Zheng, Liangdong Qu
IEEE Trans. Commun.3
2025 Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
abstract
Cross-modal hashing (CMH) has appeared as a popular technique for cross-modal retrieval due to its low storage cost and high computational efficiency in large-scale data. Most existing methods implicitly assume that multi-modal data is correctly labeled, which is expensive and even unattainable due to the inevitable imperfect annotations (i.e., noisy labels) in real-world scenarios. Inspired by human cognitive learning, a few methods introduce self-paced learning to gradually train the model from easy to hard samples, which is often used to mitigate the effects of feature noise or outliers. It is a less-touched problem that how to utilize SPL to alleviate the misleading of noisy labels on the hash model. To tackle this problem, we propose a new cognitive cross-modal retrieval method called Robust Self-paced Hashing with Noisy Labels (RSHNL), which can mimic the human cognitive process to identify the noise while embracing robustness against noisy labels. Specifically, we first propose a contrastive hashing learning (CHL) scheme to improve multi-modal consistency, thereby reducing the inherent semantic gap. Afterward, we propose center aggregation learning (CAL) to mitigate the intra-class variations. Finally, we propose Noise-tolerance Self-paced Hashing (NSH) that dynamically estimates the learning difficulty for each instance and distinguishes noisy labels through the difficulty level. For all estimated clean pairs, we further adopt a self-paced regularizer to gradually learn hash codes from easy to hard. Extensive experiments demonstrate that the proposed RSHNL performs remarkably well over the state-of-the-art CMH methods.
Ruitao Pu, Yuan Sun 0016, Zhenwen Ren, Xiaomin Song, Huiming Zheng, Dezhong Peng
AAAI6
2025 DiCA: Disambiguated Contrastive Alignment for Cross-Modal Retrieval with Partial Labels
abstract
Cross-modal retrieval aims to retrieve relevant data across different modalities. Driven by costly massive labeled data, existing cross-modal retrieval methods achieve encouraging results. To reduce annotation costs while maintaining performance, this paper focuses on an untouched but challenging problem, i.e., cross-modal retrieval with partial labels (PLCMR). PLCMR faces the dual challenges of annotation ambiguity and modality gap. To address these challenges, we propose a novel method termed disambiguated contrastive alignment (DiCA) for cross-modal retrieval with partial labels. Specifically, DiCA proposes a novel non-candidate boosted disambiguation learning mechanism (NBDL), which elaborately balances the trade-off between the losses on candidate and non-candidate labels that eliminate label ambiguity and narrow the modality gap. Moreover, DiCA presents an instance-prototype representation learning mechanism (IPRL) to enhance the model by further eliminating the modality gap at both the instance and prototype levels. Thanks to NBDL and IPRL, our DiCA effectively addresses the issues of annotation ambiguity and modality gap for cross-modal retrieval with partial labels. Experiments on four benchmarks validate the effectiveness of our proposed method, which demonstrates enhanced performance over existing state-of-the-art methods.
Chao Su 0003, Huiming Zheng, Dezhong Peng, Xu Wang 0028
AAAI2
2025 Octree-Based Learned Point Cloud Geometry Compression: A Lossy Perspective
abstract
In this paper, we mainly research lossy octree-based point cloud geometry compression. We analyze data characteristics of different point clouds and propose lossy approaches specifically (Fig. 1 (d-f)). For object point clouds that suffer from quantization step adjustment, we propose a new leaf nodes lossy compression method (Fig. 1 (a-b)), which achieves lossy compression by performing bit-wise coding and binary prediction on leaf nodes. For LiDAR point clouds, we discover the occupancy distribution similarity for octrees in the same depth. Therefore, we present variable rate approaches and propose a simple but effective rate control method. Experimental results demonstrate that the proposed leaf nodes lossy compression method significantly outperforms the previous octree-based method on object point clouds, and the proposed rate control method achieves about 1% bit error without finetuning on LiDAR point clouds.
Kaiyu Zheng, Wei Gao 0003, Huiming Zheng
DCC3
2025 PMR: Physical Model-Driven Multi-Stage Restoration of Turbulent Dynamic Videos
abstract
Geometric distortions and blurring caused by atmospheric turbulence degrade the quality of long-range dynamic scene videos. Existing methods struggle with restoring edge details and eliminating mixed distortions, especially under conditions of strong turbulence and complex dynamics. To address these challenges, we introduce a Dynamic Efficiency Index (DEI), which combines turbulence intensity, optical flow, and proportions of dynamic regions to accurately quantify video dynamic intensity under varying turbulence conditions and provide a high-dynamic turbulence training dataset. Additionally, we propose a Physical Model-Driven Multi-Stage Video Restoration (PMR) framework that consists of three stages: de-tilting for geometric stabilization, motion segmentation enhancement for dynamic region refinement, and de-blurring for quality restoration. PMR employs lightweight backbones and stage-wise joint training to ensure both efficiency and high restoration quality. Experimental results demonstrate that the proposed method effectively suppresses motion trailing artifacts, restores edge details and exhibits strong generalization capability, especially in real-world scenarios characterized by high-turbulence and complex dynamics. We will make the code and datasets openly available.
Jingyuan Ye, Huiming Zheng
ECAI6
2025 OpenMVC: An Open-Source Library for Learning-based Multi-view Compression
abstract
Amidst the swift advancement of 3D vision technology, Multi-view Compression (MVC) has become a crucial technique, widely applied in fields such as virtual reality, augmented reality, autonomous driving, telemedicine, and security surveillance. The technology effectively handles views from multiple cameras, utilizing the inter-view correlations to compress data efficiently. It substantially decreases the data transmission and storage requirements, enabling a richer and more realistic visual experience within the same bandwidth constraints. To further enhance compression performance, new methods continue to emerge. However, the absence of a unified benchmark testing library capable of effectively evaluating existing algorithms poses significant challenges to the further development of the field and the practical deployment of algorithms. To address this issue, we introduce OpenMVC, an Open-Source Library for Learning-based Multi-view Compression. We provide a comprehensive description and analysis of the performance advantages of existing algorithms. Furthermore, we conduct extensive and comprehensive benchmark testing of nine typical algorithms in the last five years, evaluating them in a consistent environment across various metrics. The open-source library for OpenMVC is available at https://openi.pcl.ac.cn/OpenAICoding/OpenMVC.
Huiming Zheng, Wei Gao 0003
ACM Multimedia1
2025 SCID-Compress900: A Multi-Scene Dataset of 4K and 1080P Screen Content Images for Image Compression Research
abstract
With the rapid growth of digital content and the increasing demand for high-resolution displays, the efficient compression of screen content images characterized by text, graphics, and UI elements has become an important research field. This paper introduces a new dataset named SCID-Compress900 specially designed for image compression research. The dataset consists of 900 high-quality screen content images, including 500 4K images and 400 1080P images. All these images are mainly composed of text/graphics, reflecting typical screen content scenarios such as office documents, software interfaces, and presentation slides. The dataset covers a diverse range of content, including various font sizes, graphic styles, and color modes, providing a comprehensive testbed for compression algorithms. To demonstrate the effectiveness of SCID-Compress900, we conduct benchmark tests using several deep learning-based image compression methods commonly employed by researchers. The experimental results show that SCID-Compress900 can well differentiate the performance of different compression algorithms. Compared with existing datasets, SCID-Compress900 offers higher resolution, larger scale, and more targeted content, making it an ideal resource for developing and evaluating advanced image compression algorithms for screen content. This dataset will not only promote the research and development of screen content compression technology but also contribute to the standardization and optimization of compression algorithms in practical applications. The project is available at https://openi.pcl.ac.cn/OpenDatasets/SCID-Compress900.
Huiming Zheng, Linjie Zhou, Wei Gao 0003
ACM Multimedia1
2025 Granular-ball fuzzy information-based outlier detector
Zhong Yuan, Dezhong Peng, Xiaomin Song, Huiming Zheng, Xinyu Su
Int. J. Approx. Reason.5
2025 Multi-granularity confidence learning for unsupervised text-to-image person re-identification with incomplete modality
Dezhong Peng, Haixiao Huang, Huiming Zheng
Knowl. Based Syst.5
2025 Label-Informed Outlier Detection Based on Granule Density
abstract
Outlier detection, crucial for identifying unusual patterns with significant implications across numerous applications, has drawn considerable research interest. Existing semisupervised methods typically treat data as purely numerical and in a deterministic manner, thereby neglecting the heterogeneity and uncertainty inherent in complex, real-world datasets. This article introduces a label-informed outlier detection method for heterogeneous data based on Granular Computing and Fuzzy Sets, namely Granule Density-based Outlier Factor (GDOF). Specifically, GDOF first employs label-informed fuzzy granulation to effectively represent various data types and develops granule density for precise density estimation. Subsequently, granule densities from individual attributes are integrated for outlier scoring by assessing attribute relevance with a limited number of labeled outliers. Experimental results on various real-world datasets show that GDOF stands out in detecting outliers in heterogeneous data with a minimal number of labeled outliers. The integration of Fuzzy Sets and Granular Computing in GDOF offers a practical framework for outlier detection in complex and diverse data types.
Baiyang Chen, Zhong Yuan, Dezhong Peng, Hongmei Chen 0001, Xiaomin Song, Huiming Zheng
IEEE Trans. Fuzzy Syst.6
2025 Deep Reversible Consistency Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval (CMR) typically involves learning common representations to directly measure similarities between multimodal samples. Most existing CMR methods commonly assume multimodal samples in pairs and employ joint training to learn common representations, limiting the flexibility of CMR. Although some methods adopt independent training strategies for each modality to improve flexibility in CMR, they utilize the randomly initialized orthogonal matrices to guide representation learning, which is suboptimal since they assume inter-class samples are independent of each other, limiting the potential of semantic alignments between sample representations and ground-truth labels. To address these issues, we propose a novel method termed Deep Reversible Consistency Learning (DRCL) for cross-modal retrieval. DRCL includes two core modules, i.e., Selective Prior Learning (SPL) and Reversible Semantic Consistency learning (RSC). More specifically, SPL first learns a transformation weight matrix on each modality and selects the best one based on the quality score as the Prior, which greatly avoids indiscriminateselection of priors learned from low-quality modalities. Then, RSC employs a Modality-invariant Representation Recasting mechanism (MRR) to recast the potential modality-invariant representations from sample semantic labels by the generalized inverse matrix of the prior. Since labels are devoid of modal-specific information, we utilize the recast features to guide the representation learning, thus maintaining semantic consistency to the fullest extent possible. In addition, a feature augmentation mechanism (FA) is introduced in RSC to encourage the model to learn over a wider data distribution for diversity. Finally, extensive experiments conducted on five widely used datasets and comparisons with 15 state-of-the-art baselines demonstrate the effectiveness and superiority of our DRCL.
Ruitao Pu, Dezhong Peng, Xiaomin Song, Huiming Zheng
IEEE Trans. Multim.5
2025 Identifying Outliers via Local Granular-Ball Density
abstract
Existing density-based outlier detection methods process data at the single-granularity level of individual samples, requiring pairwise distance calculations between all samples and exhibiting high sensitivity to noise. The single-granularity-based processing paradigm fails to mine the information at multiple levels of granularity in data, and most of these methods ignore the potential uncertainty information in data, such as fuzziness, resulting in an inability to effectively detect potential outliers in data. As a novel granular computing method, Granular-Ball Computing (GBC) is characterized by its multi-granularity and robustness, which makes it able to make up for the above drawbacks well. In this study, we propose local Granular-Ball Density-based Outlier (GBDO) detection to improve the performance of the density-based methods. In GBDO, we first identify the $k\text {-}$ similarity Granular-Ball (GB) neighborhoods of each GB via the fuzzy relations among them. Subsequently, the local reachability similarity density of the GBs is calculated through the reachability similarity we defined. Finally, the local GB outlier factors of the samples are calculated based on the local reachability similarity density of the GBs. We adopt a multi-granularity processing paradigm using GBs as the basic units, which reduces computational complexity and improves robustness to noisy data by leveraging the multi-granularity nature of GBs. The experimental results demonstrate the effectiveness of GBDO by comparing it with state-of-the-art methods. The source code and datasets are publicly available at https://github.com/Mxeron/GBDO.
Xinyu Su, Dezhong Peng, Xiaomin Song, Huiming Zheng, Zhong Yuan
IEEE Trans. Neural Networks Learn. Syst.5
2024 End-to-End RGB-D Image Compression via Exploiting Channel-Modality Redundancy
abstract
As a kind of 3D data, RGB-D images have been extensively used in object tracking, 3D reconstruction, remote sensing mapping, and other tasks. In the realm of computer vision, the significance of RGB-D images is progressively growing. However, the existing learning-based image compression methods usually process RGB images and depth images separately, which cannot entirely exploit the redundant information between the modalities, limiting the further improvement of the Rate-Distortion performance. With the goal of overcoming the defect, in this paper, we propose a learning-based dual-branch RGB-D image compression framework. Compared with traditional RGB domain compression scheme, a YUV domain compression scheme is presented for spatial redundancy removal. In addition, Intra-Modality Attention (IMA) and Cross-Modality Attention (CMA) are introduced for modal redundancy removal. For the sake of benefiting from cross-modal prior information, Context Prediction Module (CPM) and Context Fusion Module (CFM) are raised in the conditional entropy model which makes the context probability prediction more accurate. The experimental results demonstrate our method outperforms existing image compression methods in two RGB-D image datasets. Compared with BPG, our proposed framework can achieve up to 15% bit rate saving for RGB images.
Huiming Zheng, Wei Gao 0003
AAAI1
2024 Semantic-Aware Visual Decomposition for Point Cloud Geometry Compression
abstract
Focusing on encoding the Region of Interest (ROI) in point clouds and allocating more bitstream is a crucial area of research. In processing point cloud data, the foreground ROI region typically contains critical information, making it essential for applications like autonomous driving and robot navigation. However, previous point cloud compression methods often treat the entire point cloud uniformly and fail to fully harness the significance of the ROI. This study is dedicated to preserving vital information in point cloud by optimizing the Point Cloud Compression (PCC) process and allocating more bitstream to the foreground ROI region. To achieve this goal, we introduce a Semantic-Aware Visual Decomposition Point Cloud Geometry Compression (SAVD-PCGC) strategy. It involves the initial identification of foreground and background regions, followed by allocating of additional bitstream resources to machine vision critical areas by controlling compression model parameters. We also propose corresponding compensation methods to reduce distortion loss in compression. This separation of foreground and background coding strategy aims to maintain compression performance while ensuring high-quality of the ROI region, thereby improving the performance of downstream tasks. Experimental results demonstrate that our approach significantly enhances the performance of point cloud object detection compared to traditional PCC methods.
Liang Xie 0004, Wei Gao 0003, Huiming Zheng
DCC3
2024 SPCGC: Scalable Point Cloud Geometry Compression for Machine Vision
abstract
With the proliferation of sensor devices, the extensive utilization of three-dimensional data in multimedia continues to grow. Point clouds are widely adopted within this domain because they are one of the most intuitive representations of three-dimensional data. However, the substantial volume of point cloud data poses significant challenges for storage and transmission. Moreover, a considerable portion of the data loses its semantic information during transmission. Consequently, how can we ensure both the perceptual quality for the human and the performance of downstream tasks during the transmission? To address this issue, we propose a scalable point cloud geometry compression framework (SPCGC) for machine perception. This framework tackles the fidelity issues associated with point cloud compression and preserves more semantic information, enhancing the performance of machine vision tasks. Our solution consists of a base layer bitstream and an enhancement layer bitstream. The base layer bitstream contains geometry data, while the enhancement layer bitstream utilizes semantic-guided residual data. Additionally, we introduce two modules for extracting and coding residual features. And incorporate classification and segmentation losses from downstream tasks into the Rate-Distortion (RD) optimization. Our approach outperforms existing learning-based lossy point cloud coding methods through empirical validation in downstream tasks without sacrificing point cloud compression performance.
Liang Xie 0004, Wei Gao 0003, Huiming Zheng, Ge Li 0002
ICRA3
2024 OpenDIC: An Open-Source Library and Performance Evaluation for Deep-learning-based Image Compression
abstract
Deep learning technologies have been popular in the image compression field for some time. An increasing number of deep-learning-based models are proposed to improve Rate-Distortion (RD) performance. Previous algorithms are implemented in the specific platform and can not be applied in cross-platform environments. In this paper, we present an open-source algorithm library called OpenDIC, which integrates a variety of end-to-end image compression methods in cross-platform environments. The contribution and details of the algorithms used in the library are described. To evaluate the performance of these algorithms, we conduct a comprehensive performance test. We compare and analyze each algorithm according to RD performance, running time, and GPU memory occupancy. The algorithm library has been released at https://openi.pcl.ac.cn/OpenDIC/.
Wei Gao 0003, Huiming Zheng, Kaiyu Zheng, Zhuozhen Yu, Yuan Li 0076, Yongchi Zhang
ACM Multimedia2
2024 ROI-Guided Point Cloud Geometry Compression Towards Human and Machine Vision
abstract
Point cloud data is pivotal in applications like autonomous driving, virtual reality, and robotics. However, its substantial volume poses significant challenges in storage and transmission. In order to obtain a high compression ratio, crucial semantic details usually confront severe damage, leading to difficulties in guaranteeing the accuracy of downstream tasks. To tackle this problem, we are the first to introduce a novel Region of Interest (ROI)-guided Point Cloud Geometry Compression (RPCGC) method for human and machine vision. Our framework employs a dual-branch parallel structure, where the base layer encodes and decodes a simplified version of the point cloud, and the enhancement layer refines this by focusing on geometry details. Furthermore, the residual information of the enhancement layer undergoes refinement through an ROI prediction network. This network generates mask information, which is then incorporated into the residuals, serving as a strong supervision signal. Additionally, we intricately apply these mask details in the Rate-Distortion (RD) optimization process, with each point weighted in the distortion calculation. Our loss function includes RD loss and detection loss to better guide point cloud encoding for the machine. Experiment results demonstrate that RPCGC achieves exceptional compression performance and better detection accuracy (10% gain) than some learning-based compression methods at high bitrates in ScanNet and SUN RGB-D datasets.
Liang Xie 0013, Wei Gao 0003, Huiming Zheng, Ge Li 0002
ACM Multimedia3
2024 ViewPCGC: View-Guided Learned Point Cloud Geometry Compression
abstract
With the rise of immersive media applications such as digital museums, virtual reality, and interactive exhibitions, point clouds, as a three-dimensional data storage format, have gained increasingly widespread attention. The massive data volume of point clouds imposes extremely high requirements on transmission bandwidth in the above applications, gradually becoming a bottleneck for immersive media applications. Although existing learning-based point cloud compression methods have achieved specific successes in compression efficiency by mining the spatial redundancy of their local structural features, these methods often overlook the intrinsic connections between point cloud data and other modality data (such as image modality), thereby limiting further improvements in compression efficiency. To address the limitation, we innovatively propose a view-guided learned point cloud geometry compression scheme, namely ViewPCGC. We adopt a novel self-attention mechanism and cross-modality attention mechanism based on sparse convolution to align the modality features of the point cloud and the view image, removing view redundancy through Modality Redundancy Removal Module (MRRM). Simultaneously, side information of the view image is introduced into the Conditional Checkboard Entropy Model (CCEM), significantly enhancing the accuracy of the probability density function estimation for point cloud geometry. In addition, we design a View-Guided Quality Enhancement Module (VG-QEM) in the decoder, utilizing the contour information of the point cloud in the view image to supplement reconstruction details. The superior experimental performance demonstrates the effectiveness of our method. Compared to the state-of-the-art point cloud geometry compression methods, ViewPCGC exhibits an average performance gain exceeding 10% on D1-PSNR metric.
Huiming Zheng, Wei Gao 0003, Zhuozhen Yu, Tiesong Zhao, Ge Li 0002
ACM Multimedia1
2024 Deviation Control for Learned Image Compression
abstract
Most approaches in learned image compression follow the transform coding scheme. The characteristics of latent variables transformed from images significantly influence the performance of codecs. In this paper, we present visual analyses on latent features of learned image compression and find that the latent variables are spread over a wide range, which may lead to complex entropy coding processes. To address this, we introduce a Deviation Control (DC) method, which applies a constraint loss on latent features and entropy parameter μ. Training with DC loss, we obtain latent features with smaller values of coding symbols and σ, effectively reducing entropy coding complexity. Our experimental results show that the plug-and-play DC loss reduces entropy coding time by 30-40% and improves compression performance.
Haotian Zhang 0009, Xiaomin Song, Huiming Zheng, Li Li 0040, Dong Liu 0002
VCIP5
2024 Frame Level Content Adaptive λ for Neural Video Compression
abstract
Neural video compression (NVC) methods have made significant advances in recent years. In most NVC methods, all frames share the same Rate-Distortion trade-off parameter λ, which might be sub-optimal. Recently, inspired by traditional video codecs' hierarchical quality structure, DCVC-DC proposed allocating periodic weights to λ to equip NVC with the hierarchical quality structure. However, the inspiration from traditional video codecs is designed to complement their complex reference structure. Compared to traditional video codecs, NVC methods' reference structure is much simpler and may not require large fluctuations in their hierarchical quality structure. Moreover, DCVC-DC's fixed hierarchical quality structure ignored the influence of video content. We conduct an elaborate study on the hierarchical quality structure in DCVC-DC, shedding light on the potential for improving compression performance by proposing a content adaptive λ to achieve a more reasonable hierarchical quality structure based on the fixed hierarchical weights. Experimental results demonstrate that the proposed method achieves a better rate-distortion performance than allocating the fixed weights to the fixed λ. On DCVC-DC and DCVC-SDD, we achieved 4.9% and 8.8% bdrate reduction with our method.
Zhirui Zuo, Junqi Liao, Xiaomin Song, Huiming Zheng, Dong Liu 0002
VCIP5
2023 OpenDMC: An Open-Source Library and Performance Evaluation for Deep-learning-based Multi-frame Compression
abstract
Video streaming has become an essential component of our everyday routines. Nevertheless, video data imposes a significant strain on data usage, demanding substantial bandwidth and storage resources for effective transmission. To suit explosively increasing video transmission and storage requirements, deep-learning-based video compression has developed rapidly in the past few years. New methods have mushroomed in order to achieve better Rate-Distortion (RD) performance. However, the absence of an algorithm library that can effectively sort, classify, and conduct extensive benchmark testing on existing algorithms remains a challenge. In this paper, we present an open-source algorithm library called OpenDMC, which integrates a variety of end-to-end video compression methods in cross-platform environments. We provide comprehensive descriptions of the algorithms used in the library, including their contributions and implementation details. We perform a thorough benchmarking test to evaluate the performance of the algorithms. We meticulously compare and analyze each algorithm based on various metrics, including RD performance, running time, and GPU memory usage. The open-source library for OpenDMC is available at https://openi.pcl.ac.cn/OpenDMC/.
Wei Gao 0003, Shangkun Sun, Huiming Zheng, Yuyang Wu, Yongchi Zhang
ACM Multimedia3
2022 OpenPointCloud: An Open-Source Algorithm Library of Deep Learning Based Point Cloud Compression
abstract
This paper gives an overview of OpenPointCloud, the first open-source algorithm library containing outstanding deep learning methods on point cloud compression (PCC). We provide an introduction of our implementations, including 8 methods on lossless geometry PCC and lossy geometry PCC. Principles and contributions of these methods in our algorithm library are illustrated, which are also implemented with different deep learning programming frameworks, such as TensorFlow, Pytorch and TensorLayer. In order to systematically evaluate the performances of all these methods, we conduct a comprehensive benchmarking test. We provide analyses and comparisons of their performances according to their categories and draw constructive conclusions. This algorithm library has been released at https://git.openi.org.cn/OpenPointCloud.
Wei Gao 0003, Ge Li 0002, Huiming Zheng, Yuyang Wu, Liang Xie 0013
ACM Multimedia4