EDBT 2026 Demo / reviewers in the wild / expert
Tongda Xu
dblp:227/8096
· DBLP profile ↗
24ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0001-8712-0301ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian SplattingabstractImplicit neural representations (INRs) have achieved remarkable success in image representation and compression, but they require substantial training time and memory. Meanwhile, recent 2D Gaussian Splatting (GS) methods (\textit{e.g.}, GaussianImage) offer promising alternatives through efficient primitive-based rendering. However, these methods require excessive Gaussian primitives to maintain high visual fidelity. To exploit the potential of GS-based approaches, we present GaussianImage++, which utilizes limited Gaussian primitives to achieve impressive representation and compression performance. Firstly, we introduce a distortion-driven densification mechanism. It progressively allocates Gaussian primitives according to signal intensity. Secondly, we employ context-aware Gaussian filters for each primitive, which assist in the densification to optimize Gaussian primitives based on varying image content. Thirdly, we integrate attribute-separated learnable scalar quantizers and quantization-aware training, enabling efficient compression of primitive attributes. Experimental results demonstrate the effectiveness of our method. In particular, GaussianImage++ outperforms GaussianImage and INRs-based COIN in representation and compression performance while maintaining real-time decoding and low memory usage. Xingtong Ge, Tongda Xu, Dailan He, Jun Zhang 0004, Yan Wang 0105 |
AAAI | 4 |
| 2025 | CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image CompressionabstractExisting learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, which leads to suboptimal rate-distortion results. In this paper, we propose a stereo image compression framework, named CAMSIC. CAMSIC independently transforms each image to latent representation and employs a powerful decoder-free Transformer entropy model to capture both spatial and disparity dependencies, by introducing a novel content-aware masked image modeling (MIM) technique. Our content-aware MIM facilitates efficient bidirectional interaction between prior information and estimated tokens, which naturally obviates the need for an extra Transformer decoder. Experiments show that our stereo image codec achieves state-of-the-art rate-distortion performance on two stereo image datasets Cityscapes and InStereo2K with fast encoding and decoding speed. Shenyuan Gao, Zhening Liu 0001, Jiawei Shao, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang 0105, Jun Zhang 0004 |
AAAI | 7 |
| 2025 | PICD: Versatile Perceptual Image Compression with Diffusion RenderingabstractRecently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images. Tongda Xu, Jiahao Li 0001, Bin Li 0012, Yan Wang 0105, Ya-Qin Zhang, Yan Lu 0001 |
CVPR | 1 |
| 2025 | MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenesabstract4D Gaussian Splatting (4DGS) has recently emerged as a promising technique for capturing complex dynamic 3D scenes with high fidelity. It utilizes a 4D Gaussian representation and a GPU-friendly rasterizer, enabling rapid rendering speeds. Despite its advantages, 4DGS faces significant challenges, notably the requirement of millions of 4D Gaussians, each with extensive associated attributes, leading to substantial memory and storage cost. This paper introduces a memory-efficient framework for 4DGS. We streamline the color attribute by decomposing it into a per-Gaussian direct color component with only 3 parameters and a shared lightweight alternating current color predictor. This approach eliminates the need for spherical harmonics coefficients, which typically involve up to 144 parameters in classic 4DGS, thereby creating a memory-efficient 4D Gaussian representation. Furthermore, we introduce an entropy-constrained Gaussian deformation technique that uses a deformation field to expand the action range of each Gaussian and integrates an opacity-based entropy loss to limit the number of Gaussians, thus forcing our model to use as few Gaussians as possible to fit a dynamic scene well. With simple half-precision storage and zip compression, our framework achieves a storage reduction by approximately 190$\times$ and 125$\times$ on the Technicolor and Neural 3D Video datasets, respectively, compared to the original 4DGS. Meanwhile, it maintains comparable rendering speeds and scene representation quality, setting a new standard in the field. Code is available at https://github.com/Xinjie-Q/MEGA. Zhening Liu 0001, Yifan Zhang 0004, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang 0105, Zehong Lin, Shuicheng Yan, Jun Zhang 0004 |
ICCV | 6 |
| 2025 | Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a PosteriorabstractRecent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating the conditional score. While in this paper, we demonstrate that the conditional score approximation employed by DPS is not as effective as previously assumed, but rather aligns more closely with the principle of maximizing a posterior (MAP). This assertion is substantiated through an examination of DPS on 512$\times$512 ImageNet images, revealing that: 1) DPS’s conditional score estimation significantly diverges from the score of a well-trained conditional diffusion model and is even inferior to the unconditional score; 2) The mean of DPS’s conditional score estimation deviates significantly from zero, rendering it an invalid score estimation; 3) DPS generates high-quality samples with significantly lower diversity. In light of the above findings, we posit that DPS more closely resembles MAP than a conditional score estimator, and accordingly propose the following enhancements to DPS: 1) we explicitly maximize the posterior through multi-step gradient ascent and projection; 2) we utilize a light-weighted conditional score estimator trained with only 100 images and 8 GPU hours. Extensive experimental results indicate that these proposed improvements significantly enhance DPS's performance. The source code for these improvements is provided in https://github.com/tongdaxu/Rethinking-Diffusion-Posterior-Sampling-From-Conditional-Score-Estimator-to-Maximizing-a-Posterior. Tongda Xu, Xiyan Cai, Xingtong Ge, Dailan He, Ya-Qin Zhang, Yan Wang 0105 |
ICLR | 1 |
| 2025 | CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object QueryabstractCooperative perception enhances the individual perception capabilities of autonomous vehicles (AVs) by providing a comprehensive view of the environment. However, balancing perception performance and transmission costs remains a significant challenge. Current approaches that transmit regionlevel features across agents are limited in interpretability and demand substantial bandwidth, making them unsuitable for practical applications. In this work, we propose CoopDETR, a novel cooperative perception framework that introduces objectlevel feature cooperation via object query. Our framework consists of two key modules: single-agent query generation, which efficiently encodes raw sensor data into object queries, reducing transmission cost while preserving essential information for detection; and cross-agent query fusion, which includes Spatial Query Matching (SQM) and Object Query Aggregation (OQA) to enable effective interaction between queries. Our experiments on the OPV2V and V2XSet datasets demonstrate that CoopDETR achieves state-of-the-art performance and significantly reduces transmission costs to 1/782 of previous methods. Zhe Wang 0070, Shaocong Xu, Xucai Zhuang, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
ICRA | 4 |
| 2025 | PA-INR: Parallel adapter-based storage of edited implicit neural representation
Mingyi Ma, Xinzui Wang, Yan Wang 0105, Tongda Xu, Fucheng Cao, Shichen Su |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | Task-Aware Encoder Control for Deep Video CompressionabstractPrior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task, mandating a dedicated decoder per task. In contrast, traditional video codecs employ a flexible encoder controller, enabling the adaptation of a single codec to different tasks through mechanisms like mode prediction. Drawing inspiration from this, we introduce an innovative encoder controller for deep video compression for machines. This controller features a mode prediction and a Group of Pictures (GoP) selection module. Our approach centralizes control at the encoding stage, allowing for adaptable encoder adjustments across different tasks, such as detection and tracking, while maintaining compatibility with a standard pre-trained DvC decoder. Empirical evidence demonstrates that our method is applica-ble across multiple tasks with various existing pre-trained Dv'Cs. Moreover, extensive experiments demonstrate that our method outperforms previous DVC by about 25% bi-trate for different tasks, with only one pre-trained decoder. Xingtong Ge, Jixiang Luo, Tongda Xu, Guo Lu, Dailan He, Yan Wang 0105, Jun Zhang 0004, Hongwei Qin |
CVPR | 4 |
| 2024 | Boosting Neural Representations for Videos with a Conditional DecoderabstractImplicit neural representations (INRs) have emerged as a promising approach for video storage and processing, showing remarkable versatility across various video tasks. However, existing methods often fail to fully leverage their representation capabilities, primarily due to inadequate alignment of intermediate features during target frame decoding. This paper introduces a universal boosting framework for current implicit video representation approaches. Specifically, we utilize a conditional decoder with a temporal-aware affine transform module, which uses the frame index as a prior condition to effectively align intermediate features with target frames. Besides, we introduce a sinusoidal NeRV-like block to generate diverse intermediate features and achieve a more balanced parameter distribution, thereby enhancing the model's capacity. With a high-frequency information-preserving reconstruction loss, our approach successfully boosts multiple baseline INRs in the reconstruction quality and convergence speed for video regression, and exhibits superior inpainting and interpolation results. Further, we integrate a consistent entropy minimization technique and develop video codecs based on these boosted INRs. Experiments on the UVG dataset confirm that our enhanced codecs significantly outperform baseline INRs and offer competitive rate-distortion performance compared to traditional and learning-based codecs. Code is available at htt ps://github.com/Xin j ieQ/Boosting-NeRV. Dailan He, Xingtong Ge, Tongda Xu, Yan Wang 0105, Hongwei Qin, Jun Zhang 0004 |
CVPR | 5 |
| 2024 | GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
Xingtong Ge, Tongda Xu, Dailan He, Yan Wang 0080, Hongwei Qin, Guo Lu, Jun Zhang 0004 |
ECCV (9) | 3 |
| 2024 | GeneFormer: Learned Gene Compression using Transformer-Based Context ModelingabstractThe development of gene sequencing technology sparks an explosive growth of gene data. Thus, the storage of gene data has become an important issue. Recently, researchers begin to investigate deep learning-based gene data compression, which outperforms general traditional methods. In this paper, we propose a transformer-based gene compression method named GeneFormer. Specifically, we first introduce a modified transformer encoder with latent array to eliminate the dependency of the nucleotide sequence. Then, we design a multi-level-grouping method to accelerate and improve the compression process. Experimental results on real-world datasets show that our method achieves significantly better compression ratio compared with state-of-the-art method, and the decoding speed is significantly faster than all existing learning-based gene compression methods. We will release our code on github once the paper is accepted. Zhanbei Cui, Tongda Xu, Yu Liao, Yan Wang 0105 |
ICASSP | 2 |
| 2024 | ECM-OPCC: Efficient Context Model for Octree-Based Point Cloud CompressionabstractRecently, deep learning methods have shown promising results in point cloud compression. However, previous octree-based approaches either lack sufficient context or have high decoding complexity (e.g. > 900s). To address this problem, we propose a sufficient yet efficient context model and design an efficient deep learning codec for point clouds. Specifically, we first propose a segment-constrained multi-group coding strategy to exploit the autoregressive context while maintaining decoding efficiency. Then, we propose a dual transformer architecture to utilize the dependency of current node on its ancestors and siblings. We also propose a random-masking pre-train method to enhance our model. Experimental results show that our approach achieves state-of-the-art performance for both lossy and lossless point cloud compression, and saves a significant amount of decoding time compared with previous octree-based SOTA compression methods. Yiqi Jin, Tongda Xu, Yuhuan Lin, Yan Wang 0105 |
ICASSP | 3 |
| 2024 | Bandwidth-Efficient Inference for Nerual Image CompressionabstractWith neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottle-neck in implementing network inference on mobile and edge devices. In this paper, we propose an end-to-end differentiable bandwidth efficient neural inference method with the activation compressed by neural data compression method. Specifically, we propose a transform-quantization-entropy coding pipeline for activation compression with symmetric exponential Golomb coding and a data-dependent Gaussian entropy model for arithmetic coding. Optimized with existing model quantization methods, low-level task of image compression can achieve up to 19× bandwidth reduction with 6.21× energy saving. The code implementation is available at https://github.com/xyzysz/Bandwidth_efficient_nic. Shanzhi Yin, Tongda Xu, Yongsheng Liang 0001, Yanghao Li, Yan Wang 0002 |
ICASSP | 2 |
| 2024 | Idempotence and Perceptual Image CompressionabstractIdempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence constraint is equivalent to conditional generative codec. Based on this newfound equivalence, we propose a new paradigm of perceptual image codec by inverting unconditional generative model with idempotence constraints. Our codec is theoretically equivalent to conditional generative codec, and it does not require training new models. Instead, it only requires a pre-trained mean-square-error codec and unconditional generative model. Empirically, we show that our proposed approach outperforms state-of-the-art methods such as HiFiC and ILLM, in terms of Fréchet Inception Distance (FID). The source code is provided in https://github.com/tongdaxu/Idempotence-and-Perceptual-Image-Compression. Tongda Xu, Ziran Zhu, Dailan He, Yanghao Li, Zhe Wang 0070, Hongwei Qin, Yan Wang 0105, Ya-Qin Zhang |
ICLR | 1 |
| 2024 | Noise Dimension of GAN: An Image Compression PerspectiveabstractGenerative adversial network (GAN) is a type of generative model that maps a high-dimensional noise to samples in target distribution. However, the dimension of noise required in GAN is not well understood. Previous approaches view GAN as a mapping from a continuous distribution to another continous distribution. In this paper, we propose to view GAN as a discrete sampler instead. From this perspective, we build a connection between the minimum noise required and the bits to losslessly compress the images. Furthermore, to understand the behaviour of GAN when noise dimension is limited, we propose divergence-entropy trade-off. This trade-off depicts the best divergence we can achieve when noise is limited. And as rate distortion trade-off, it can be numerically solved when source distribution is known. Finally, we verifies our theory with experiments on image generation. Ziran Zhu, Tongda Xu, Yan Wang 0105 |
ICME | 2 |
| 2024 | EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object DetectionabstractIn autonomous driving, cooperative perception makes use of multi-view cameras from both vehicles and infrastructure, providing a global vantage point with rich semantic context of road conditions beyond a single vehicle viewpoint. Currently, two major challenges persist in vehicle-infrastructure cooperative 3D (VIC3D) object detection: 1) inherent pose errors when fusing multi-view images, caused by time asynchrony across cameras; 2) information loss in transmission process resulted from limited communication bandwidth. To address these issues, we propose a novel camera-based 3D detection framework for VIC3D task, Enhanced Multi-scale Image Feature Fusion (EMIFF). To fully exploit holistic perspectives from both vehicles and infrastructure, we propose Multi-scale Cross Attention (MCA) and Camera-aware Channel Masking (CCM) modules to enhance infrastructure and vehicle features at scale, spatial, and channel levels to correct the pose error introduced by camera asynchrony. We also introduce a Feature Compression (FC) module with channel and spatial compression blocks for transmission efficiency. Experiments show that EMIFF achieves SOTA on DAIR-V2X-C datasets, significantly outperforming previous early-fusion and late-fusion methods with comparable transmission costs. Zhe Wang 0070, Siqi Fan 0002, Xiaoliang Huo, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
ICRA | 4 |
| 2023 | Your Camera Improves Your Point Cloud CompressionabstractLiDAR point cloud compression is important for autonomous driving as it consumes a lot of storage and bandwidth. Although the fusion of camera and LiDAR for vision perception has been well studied, it remains unexplored that how we can improve the compression of LiDAR point cloud data using cross-modal information from cameras. In this paper’ we propose a multi-modality compression framework for LiDAR point cloud by exploiting the depth information predicted from its paired image. To the best of our knowledge’ our model is the first multi-modality compression framework for point cloud. Specifically’ we first represent point cloud based on octrees to reduce spatial redundancy. Then’ we propose a cross-modal fusion structure to improve the compression of these octrees’ with depth distribution extracted from the camera pixels and acts as side information. Compared to previous state-of-the-art (SOTA) method, our approach obtains up to 8.10% compression rate gain for LiDAR point cloud compression. Yuhuan Lin, Tongda Xu, Yanghao Li, Zhe Wang 0070, Yan Wang 0105 |
ICASSP | 2 |
| 2023 | Unified Learning-Based Lossy and Lossless Jpeg RecompressionabstractJPEG is still the most widely used image compression algorithm. Most image compression algorithms only consider uncompressed original image, while ignoring a large number of already existing JPEG images. Recently, JPEG recompression approaches have been proposed to further reduce the size of JPEG files. However, those methods only consider JPEG lossless recompression, which is just a special case of the rate-distortion theorem. In this paper, we propose a unified lossly and lossless JPEG recompression framework, which consists of learned quantization table and Markovian hierarchical variational autoencoders. Experiments show that our method can achieve arbitrarily low distortion when the bitrate is close to the upper bound, namely the bitrate of the lossless compression model. To the best of our knowledge, this is the first learned method that bridges the gap between lossy and lossless recompression of JPEG images. Jianghui Zhang, Jixiang Luo, Tongda Xu, Yan Wang 0105, Hongwei Qin |
ICIP | 5 |
| 2023 | MIXLIC: Mixing Global and Local Context Model for learned Image CompressionabstractLearned Image Compression (LIC) is considered as a future direction for image compression. However, existing LIC methods only marginally outperform the latest traditional codec VVC. This is because current LIC methods have not fully utilized the global and local information of images. In this paper, we propose a MIXing global and local context Module (MIXM), which combines global context extractor (GCE) with local context extractor (LCE) in a parallel design, capturing both global and local dependencies. Based on the MIXM, we further build the MIXLIC, a mixing global and local context model for learned image compression, which can fully utilize global and local context to improve performance. Experimental results show that our proposed method MIXLIC achieves state-of-the-art RD performance and gets more visually pleasant results compared with other learning-based methods and traditional codecs. Haihang Ruan, Tongda Xu, Zhiyong Tan, Yan Wang 0105 |
ICME | 3 |
| 2023 | Bit Allocation using OptimizationabstractIn this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equivalent to pixel-level bit allocation with precise rate & quality dependency model. Based on this equivalence, we establish a new paradigm of bit allocation using SAVI. Different from previous bit allocation methods, our approach requires no empirical model and is thus optimal. Moreover, as the original SAVI using gradient ascent only applies to single-level latent, we extend the SAVI to multi-level such as NVC by recursively applying back-propagating through gradient ascent. Finally, we propose a tractable approximation for practical implementation. Our method can be applied to scenarios where performance outweights encoding speed, and serves as an empirical bound on the R-D performance of bit allocation. Experimental results show that current state-of-the-art bit allocation algorithms still have a room of $\approx 0.5$ dB PSNR to improve compared with ours. Code is available at https://github.com/tongdaxu/Bit-Allocation-Using-Optimization. Tongda Xu, Han Gao 0012, Chenjian Gao, Dailan He, Jinyong Pi, Jixiang Luo, Mao Ye 0001, Hongwei Qin, Yan Wang 0080, Ya-Qin Zhang |
ICML | 1 |
| 2023 | Idempotent Learned Image Compression with Right-InverseabstractWe consider the problem of idempotent learned image compression (LIC).
The idempotence of codec refers to the stability of codec to re-compression.
To achieve idempotence, previous codecs adopt invertible transforms such as DCT and normalizing flow.
In this paper, we first identify that invertibility of transform is sufficient but not necessary for idempotence. Instead, it can be relaxed into right-invertibility. And such relaxation allows wider family of transforms.
Based on this identification, we implement an idempotent codec using our proposed blocked convolution and null-space enhancement.
Empirical results show that we achieve state-of-the-art rate-distortion performance among idempotent codecs. Furthermore, our codec can be extended into near-idempotent codec by relaxing the right-invertibility. And this near-idempotent codec has significantly less quality decay after $50$ rounds of re-compression compared with other near-idempotent codecs. Yanghao Li, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
NeurIPS | 2 |
| 2022 | Spatial Moment Pooling Improves Neural Image AssessmentabstractIn recent years, there has been widespread attention drawn to convolutional neural network (CNN) based blind image quality assessment (IQA). A large number of works start by extracting deep features from CNN. Then, those features are processed through spatial average pooling (SAP) and fully connected layers to predict quality. Inspired by full reference IQA and texture features, in this paper, we extend SAP (1stmoment) into spatial moment pooling (SMP) by incorporating higher order moments (such as variance, skewness). Moreover, we provide learning friendly normalization to circumvent numerical issue when computing gradients of higher moments. Experimental results suggest that simply upgrading SAP to SMP significantly enhances CNN-based blind IQA methods and achieves state of the art performance. Tongda Xu, Yifan Shao, Yan Wang 0080, Hongwei Qin |
ICIP | 1 |
| 2022 | Flexible Neural Image Compression via Code EditingabstractNeural image compression (NIC) has outperformed traditional image codecs in rate-distortion (R-D) performance. However, it usually requires a dedicated encoder-decoder pair for each point on R-D curve, which greatly hinders its practical deployment. While some recent works have enabled bitrate control via conditional coding, they impose strong prior during training and provide limited flexibility. In this paper we propose Code Editing, a highly flexible coding method for NIC based on semi-amortized inference and adaptive quantization. Our work is a new paradigm for variable bitrate NIC, and experimental results show that our method surpasses existing variable-rate methods. Furthermore, our approach is so flexible that it can also achieves ROI coding and multi-distortion trade-off with a single decoder. Our approach is compatible to all NIC methods with differentiable decoder NIC, and it can be even directly adopted on existing pre-trained models. Chenjian Gao, Tongda Xu, Dailan He, Yan Wang 0080, Hongwei Qin |
NeurIPS | 2 |
| 2022 | Multi-Sample Training for Neural Image CompressionabstractThis paper considers the problem of lossy neural image compression (NIC). Current state-of-the-art (SOTA) methods adopt uniform posterior to approximate quantization noise, and single-sample pathwise estimator to approximate the gradient of evidence lower bound (ELBO). In this paper, we propose to train NIC with multiple-sample importance weighted autoencoder (IWAE) target, which is tighter than ELBO and converges to log likelihood as sample size increases. First, we identify that the uniform posterior of NIC has special properties, which affect the variance and bias of pathwise and score function estimators of the IWAE target. Moreover, we provide insights on a commonly adopted trick in NIC from gradient variance perspective. Based on those analysis, we further propose multiple-sample NIC (MS-NIC), an enhanced IWAE target for NIC. Experimental results demonstrate that it improves SOTA NIC methods. Our MS-NIC is plug-and-play, and can be easily extended to neural video compression. Tongda Xu, Yan Wang 0080, Dailan He, Chenjian Gao, Han Gao 0012, Kunzan Liu, Hongwei Qin |
NeurIPS | 1 |