VLDB 2026 Research / reviewers in the wild / expert
Tong Chen 0004
dblp:22/1512-4
· DBLP profile ↗
25ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0001-5020-6099ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MG-VLQA: Multi-Granularity Quality Assessment for Image Compression via Visual Language ModelsabstractDespite significant advances in image compression, existing evaluation metrics remain poorly aligned with human visual perception-particularly under extremely low bitrates, where reconstructed images often suffer from abstract distortions or semantic degradation that are difficult for conventional metrics to capture. To address this limitation, we propose MG-VLQA, a novel multi-granularity quality assessment framework that leverages VisionLanguage Models (VLMs) to evaluate image reconstruction fidelity through the lens of semantic consistency with the original caption. Our method formulates a suite of captionderived questions spanning three complementary dimensions: (1) entity presence (semantic completeness), (2) detail fidelity (local appearance accuracy), and (3) inter-entity interactions (relational coherence). By simulating human-like perceptual judgment via VLMbased question answering and semantic similarity scoring, MG-VLQA provides a more interpretable, fine-grained, and perceptually relevant assessment of compression quality. Extensive experiments across multiple datasets and codecs demonstrate that our metric achieves higher correlation with human judgment and offers superior discriminative power. Hanfei Li, Anle Ke, Jiawen Gu, Tong Chen 0004, Zhan Ma 0001 |
DCC | 5 |
| 2026 | An Efficient Hardware Accelerator for JPEG-AI Image Compression on FPGA
Weize Ma, Siyuan Leng, Tong Chen 0004, Ming Lu 0003, Zhan Ma 0001 |
ISCAS | 4 |
| 2025 | RENO: Real-Time Neural Compression for 3D LiDAR Point CloudsabstractDespite the substantial advancements demonstrated by learning-based neural models in the LiDAR Point Cloud Compression (LPCC) task, realizing real-time compression—an indispensable criterion for numerous industrial applications—remains a formidable challenge. This paper proposes RENO, the first real-time neural codec for 3D LiDAR point clouds, achieving superior performance with a lightweight model. RENO skips the octree construction and directly builds upon the multiscale sparse tensor representation. Instead of the multi-stage inferring, RENO devises sparse occupancy codes, which exploit cross-scale correlation and derive voxels’ occupancy in a one-shot manner, greatly saving processing time. Experimental results demonstrate that the proposed RENO achieves real-time coding speed, 10 fps at 14-bit depth on a desktop platform (e.g., one RTX 3090 GPU) for both encoding and decoding processes, while providing 12.25% and 48.34% bit-rate savings compared to G-PCCv23 and Draco, respectively, at a similar quality. RENO model size is merely 1MB, making it attractive for practical applications. The source code is available at https://github.com/NJUVISION/RENO. Kang You, Tong Chen 0004, Dandan Ding, Muhammad Salman Asif, Zhan Ma 0001 |
CVPR | 2 |
| 2025 | On Quantizing Neural Representation for Variable-Rate Video CodingabstractThis work introduces NeuroQuant, a novel post-training quantization (PTQ) approach tailored to non-generalized Implicit Neural Representations for variable-rate Video Coding (INR-VC). Unlike existing methods that require extensive weight retraining for each target bitrate, we hypothesize that variable-rate coding can be achieved by adjusting quantization parameters (QPs) of pre-trained weights. Our study reveals that traditional quantization methods, which assume inter-layer independence, are ineffective for non-generalized INR-VC models due to significant dependencies across layers. To address this, we redefine variable-rate INR-VC as a mixed-precision quantization problem and establish a theoretical framework for sensitivity criteria aimed at simplified, fine-grained rate control. Additionally, we propose network-wise calibration and channel-wise quantization strategies to minimize quantization-induced errors, arriving at a unified formula for representation-oriented PTQ calibration. Our experimental evaluations demonstrate that NeuroQuant significantly outperforms existing techniques in varying bitwidth quantization and compression efficiency, accelerating encoding by up to eight times and enabling quantization down to INT2 with minimal reconstruction loss. This work introduces variable-rate INR-VC for the first time and lays a theoretical foundation for future research in rate-distortion optimization, advancing the field of video coding technology. The materials
will be available at https://github.com/Eric-qi/NeuroQuant. Junqi Shi, Zhujia Chen, Hanfei Li, Ming Lu 0003, Tong Chen 0004, Zhan Ma 0001 |
ICLR | 6 |
| 2025 | Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware DiffusionabstractExisting multimodal large model-based image compression frameworks often rely on a fragmented integration of semantic retrieval, latent compression, and generative models, resulting in suboptimal performance in both reconstruction fidelity and coding efficiency. To address these challenges, we propose a residual-guided ultra lowrate image compression named ResULIC, which incorporates residual signals into both semantic retrieval and the diffusion-based generation process. Specifically, we introduce Semantic Residual Coding (SRC) to capture the semantic disparity between the original image and its compressed latent representation. A perceptual fidelity optimizer is further applied for superior reconstruction quality. Additionally, we present the Compression-aware Diffusion Model (CDM), which establishes an optimal alignment between bitrates and diffusion time steps, improving compression-reconstruction synergy. Extensive experiments demonstrate the effectiveness of ResULIC, achieving superior objective and subjective performance compared to state-of-the-art diffusion-based methods with -80.7%, -66.3% BD-rate saving in terms of LPIPS and FID. Anle Ke, Xu Zhang 0027, Tong Chen 0004, Ming Lu 0003, Jiawen Gu, Zhan Ma 0001 |
ICML | 3 |
| 2025 | ConPCAC: Conditional Lossless Point Cloud Attribute Compression via Spatial DecompositionabstractA conditional lossless point cloud attribute compression method, dubbed ConPCAC, is proposed. The previous work typically codes point attributes in a point cloud in an autoregressive way, incurring unbearable coding time. By contrast, ConPCAC proposes a group-wise conditional entropy model for fast coding while preserving coding performance. Specifically, ConPCAC adopts a “Group Decomposition - Attribute Initialization - Latent Distribution Prediction” framework. First, it flexibly decomposes the original point cloud into multiple groups according to the geometry coordinate distribution. Then, the first group is coded using a base coder, e.g., the standardized G-PCC, and the following groups are progressively coded using a neural coder conditioned on their preceding groups. Two key units, Attribute Initialization (Init) and Latent Distribution Prediction (LDP), are devised in the neural coder. The Init unit employs the nearest neighbor to initialize the attributes of a group, and the LDP unit further predicts the attribute probability distribution for the group. In this way, ConPCAC enables full correlation exploration across groups and parallel processing among points in a group. Finally, the predicted probabilities are fed into the arithmetic engine to code the true attribute values of each group. Extensive experiments demonstrate the performance of ConPCAC. It achieves 14.59%, 10.32%, and 12.26% improvements over the latest G-PCC on the widely used 8iVFB, Owlii, and MVUB datasets, respectively, significantly outperforming state-of-the-art lossless PCAC methods. Moreover, its computational complexity is comparable to G-PCC and much lower than existing learning-based methods. Associated code and models will be released on the websitehttps://github.com/3dpcc/ConPCAC. Tong Chen 0004, Kang You, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Neural Compression System for Point Cloud Video StreamingabstractPoint cloud video streaming is promising for immersive media applications, which urges the development of efficient compression methods. However, existing approaches either suffer from poor performance or lack effective coder control mechanisms, making them impractical for networked point cloud services, where bandwidth is often constrained and fluctuates over time. Therefore, this paper proposes a system-level solution - a layered point cloud compressor, called Yak, to address these issues. Yak offers comprehensive support for both intra and inter-frame coding of geometry and attribute components in point cloud sequences. It consists of three layers: the Base Layer uses the standard G-PCC to encode a thumbnail counterpart downscaled from the input point cloud; the Enhancement Layer devises the end-to-end variational autoencoder to compress the original input conditioned on the base layer reconstruction, and the Dynamic Layer generates feature-space predictions as the temporal prior for conditional inter-frame coding. In addition, Yak devises the Content Analysis module to dynamically determine the optimal encoding parameters of each frame, by which bit budget is intelligently allocated for geometry and attribute components to maximize the overall rate-distortion (R-D) performance. Such accurate rate control relies on the parametric rate/distortion models whose parameters are initialized through one-pass template matching and frame-wise delta updating constrained by R-D optimization. Following standard evaluation guidelines, Yak has notably outperformed traditional rules-based methods such as MPEG G-PCC and V-PCC, as well as other learning-based approaches, while offering flexible networked adaption and affordable complexity. Junteng Zhang, Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Revisit Point Cloud Quality Assessment: Current Advances and a Multiscale-Inspired ApproachabstractThe demand for full-reference point cloud quality assessment (PCQA) has extended across various point cloud services. Unlike image quality assessment, where the reference and the distorted images are naturally aligned in coordinates and thus allow point-to-point (P2P) color assessment, the coordinates and attributes of a 3D point cloud may both suffer from distortion, making the P2P evaluation unsuitable. To address this, PCQA methods usually define a set of key points and construct a neighborhood around each key point for neighbor-to-neighbor (N2N) computation on geometry and attribute. However, state-of-the-art PCQA methods often exhibit limitations in certain scenarios due to insufficient consideration of key points and neighborhoods. To overcome these challenges, this paper proposes PQI, a simple yet efficient metric to index point cloud quality. PQI suggests using scale-wise key points to uniformly perceive distortions within a point cloud, along with a mild neighborhood size associated with each key point for compromised N2N computation. To achieve this, PQI employs a multiscale framework to obtain key points, ensuring comprehensive feature representation and distortion detection throughout the entire point cloud. Such a multiscale method merges every eight points into one in the downsampling processing, implicitly embedding neighborhood information into a single point and thereby eliminating the need for an explicitly large neighborhood. Further, within each neighborhood, simple features, such as geometry Euclidean distance difference and attribute value difference, are extracted. Feature similarity is then calculated between the reference and the distorted samples at each scale and linearly weighted to generate the final PQI score. Extensive experiments demonstrate the superiority of PQI, consistently achieving high performance across several widely recognized PCQA datasets. Moreover, PQI is highly appealing for practical applications due to its low complexity and flexible scale options. Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Variable-rate Neural Speech Compression with Multi-scale Feature Extraction and Improved Entropy ModelingabstractSpeech coding serves as a means of data compression, aiming to decrease the expenses related to data storage and transmission. The efficacy of compressing speech efficiently through neural networks has been demonstrated in methods using vector quantization (VQ). However, the complex procedure of VQ makes it challenging to fit into frameworks and limits compression at discrete bitrate points. This paper proposes a neural speech compression framework, which achieves flexible bitrate speech reconstruction through compact latent representation and better entropy estimation. Shaohan Sun, Yuzhuo Kong, Tong Chen 0004, Zhan Ma 0001 |
DCC | 3 |
| 2024 | NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image SequenceabstractThis paper proposes the NeRI, an implicit neural representation (INR) based LiDAR point cloud compressor. In NeRI, we first transform a sequence of 3D LiDAR frames into a 2D range image sequence through range image projection over time. Then, we employ a neural network conditioned on the temporal frame index and associated LiDAR sensor pose to fit input range images as closely as possible. The optimized network parameters, which implicitly represent the input LiDAR data, are later lossily compressed. NeRI decoder is then initialized using decoded parameters to generate range images for reconstructing the 3D LiDAR sequence accordingly. Extensive experimental results demonstrate the significant superiority of NeRI regarding the compression efficiency and decoding speed compared to state-of-the-art 2D and 3D compressors for LiDAR point cloud. Ruixiang Xue, Tong Chen 0004, Dandan Ding, Xun Cao, Zhan Ma 0001 |
ICASSP | 3 |
| 2024 | HINER: Neural Representation for Hyperspectral ImageabstractThis paper introduces HINER, a novel neural representation for compressing HSI and ensuring high-quality downstream tasks on compressed HSI. HINER fully exploits inter-spectral correlations by explicitly encoding of spectral wavelengths and achieves a compact representation of the input HSI sample through joint optimization with a learnable decoder. By additionally incorporating the Content Angle Mapper with the L1 loss, we can supervise the global and local information within each spectral band, thereby enhancing the overall reconstruction quality. For downstream classification on compressed HSI, we theoretically demonstrate the task accuracy is not only related to the classification loss but also to the reconstruction fidelity through a first-order expansion of the accuracy degradation, and accordingly adapt the reconstruction by introducing Adaptive Spectral Weighting. Owing to the monotonic mapping of HINER between wavelengths and spectral bands, we propose Implicit Spectral Interpolation for data augmentation by adding random variables to input wavelengths during classification model training. Experimental results on various HSI datasets demonstrate the superior compression performance of our HINER compared to the existing learned methods and also the traditional codecs. Our model is lightweight and computationally efficient, which maintains high accuracy for downstream classification task even on decoded HSIs at high compression ratios. Our materials will be released at https://github.com/Eric-qi/HINER. Junqi Shi, Mingyi Jiang, Ming Lu 0003, Tong Chen 0004, Xun Cao, Zhan Ma 0001 |
ACM Multimedia | 4 |
| 2024 | Compressing 3D Gaussian Splatting via a Generalizable Neural CoderabstractAs a promising technique for 3D representation, 3D Gaussian Splatting (3DGS) offers fast rendering speed and high fidelity while generating large data volumes. This challenges storage and transmission, so an efficient compression solution is required. Existing implicit methods require pre-scene optimization (online), leading to a long optimization time. By contrast, this paper regards the 3DGS as a point cloud and pre-trains a generalizable (offline) neural coder for compression. After obtaining the 3DGS representation, we focus on the data compression process, which is friendly to applications already equipped with a PCC codec. The neural coder employed is extended from a typical AIbased point cloud compression method, which uses a multiscale and multistage framework to exploit spatial correlations across scales and stages for conditional coding. Experimental results show that our method significantly outperforms existing 3DGS representations without compromising fidelity, achieving more than 39× and 6.8× compression ratio compared to the original 3DGS and SOTA Scaffold-GS, respectively. More importantly, our approach does not require additional time to optimize the compression model. Junteng Zhang, Tong Chen 0004, Hao Zhu 0004, Dandan Ding, Zhan Ma 0001 |
VCIP | 2 |
| 2023 | G-PCC++: Enhanced Geometry-based Point Cloud CompressionabstractMPEG Geometry-based Point Cloud Compression (G-PCC) standard is developed for lossy encoding of point clouds to enable immersive services over the Internet. However, lossy G-PCC introduces superimposed distortions from both geometry and attribute information, seriously deteriorating the Quality of Experience (QoE). This paper thus proposes the Enhanced G-PCC (GPCC++), to effectively address the compression distortion and restore the quality. G-PCC++ separates the enhancement into two stages: it first enhances the geometry and then maps the decoded attribute to the enhanced geometry for refinement. As for geometry restoration, a k Nearest Neighbors (kNN)-based Linear Interpolation is first used to generate a denser geometry representation, on top of which GeoNet further generates sufficient candidates to restore geometry through probability-sorted selection. For attribute enhancement, a kNN-based Gaussian Distance Weighted Mapping is devised to re-colorize all points in enhanced geometry tensor, which are then refined by AttNet for the final reconstruction. G-PCC++ is the first solution addressing the geometry and attribute artifacts together. Extensive experiments on several public datasets demonstrate the superiority of G-PCC++, e.g., on the solid point cloud dataset 8iVFB, G-PCC++ outperforms G-PCC by 88.24% (80.54%) BD-BR in D1 (D2) measurement of geometry and by 14.64% (13.09%) BD-BR in Y (YUV) attribute. Moreover, when considering both geometry and attribute, G-PCC++ also largely surpasses G-PCC by 25.58% BD-BR using PCQM assessment. Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
ACM Multimedia | 2 |
| 2023 | YOGA: Yet Another Geometry-based Point Cloud CompressorabstractA learning-based YOGA (Yet Another Geometry-based Point Cloud Compressor) is proposed. It is flexible, allowing for the separable lossy compression of geometry and color attributes, and variable-rate coding using a single neural model; it is high-efficiency, significantly outperforming the latest G-PCC standard quantitatively and qualitatively, e.g., 25% BD-BR gains using PCQM (Point Cloud Quality Metric) as the distortion assessment, and it is lightweight, e.g., similar runtime as the G-PCC codec, owing to the use of sparse convolution and parallel entropy coding. To this end, YOGA adopts a unified end-to-end learning-based backbone for separate geometry and attribute compression. The backbone uses a two-layer structure, where the downscaled thumbnail point cloud is encoded using G-PCC at the base layer, and upon G-PCC compressed priors, multiscale sparse convolutions are stacked at the enhancement layer to effectively characterize spatial correlations to compactly represent the full-resolution sample. In addition, YOGA integrates the adaptive quantization and entropy model group to enable variable-rate control, as well as adaptive filters for better quality restoration. Junteng Zhang, Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
ACM Multimedia | 2 |
| 2023 | 2C-Net: integrate image compression and classification via deep neural network
Tong Chen 0004, Shiliang Pu, Qiu Shen |
Multim. Syst. | 2 |
| 2023 | Toward Robust Neural Image Compression: Adversarial Attack and Model FinetuningabstractDeep neural network-based image compression has been extensively studied. However, the model robustness which is crucial to practical application is largely overlooked. We propose to examine the robustness of prevailing learned image compression models by injecting negligible adversarial perturbation into the original source image. Severe distortion in decoded reconstruction reveals the general vulnerability in existing methods regardless of their settings (e.g., network architecture, loss function, quality scale). A variety of defense strategies including geometric self-ensemble based pre-processing, and adversarial training, are investigated against the adversarial attack to improve the model’s robustness. Later the defense efficiency is further exemplified in real-life image recompression case studies. Overall, our methodology is simple, effective, and generalizable, making it attractive for developing robust learned image compression solutions. All materials are made publicly accessible athttps://njuvision.github.io/RobustNICfor reproducible research. Tong Chen 0004, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Decoder-Side Cross Resolution Synthesis for Video Compression EnhancementabstractThis paper proposes a decoder-side Cross Resolution Synthesis (CRS) module to pursue better compression efficiency beyond the latest Versatile Video Coding (VVC), where we encode intra frames at original high resolution (HR), compress inter frames at a lower resolution (LR), and then super-resolve decoded LR inter frames with the help from preceding HR intra and neighboring LR inter frames. For a LR inter frame, a motion alignment and aggregation network (MAN) is devised to produce temporally aggregated motion representation to best guarantee the temporal smoothness; Another texture compensation network (TCN) is utilized to generate texture representation from decoded HR intra frame for better augmenting spatial details; Finally, a similarity-driven fusion engine synthesizes motion and texture representations to upscale LR inter frames for the removal of compression and resolution re-sampling noises. We enhance the VVC using proposed CRS, showing averaged 8.76% and 11.93% Bjntegaard Delta Rate (BD-Rate) gains against the latest VVC anchor in Random Access (RA) and Low-delay P (LDP) settings respectively. In addition, experimental comparisons to the state-of-the-art super-resolution (SR) based VVC enhancement methods, and ablation studies are conducted to further report superior efficiency and generalization of the proposed algorithm. All materials will be made to public at https://njuvision.github.io/CRS for reproducible research. Ming Lu 0003, Tong Chen 0004, Zhenyu Dai, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Efficient Neural Image Decoding via Fixed-Point InferenceabstractRecent learned image coding has emerged with superior efficiency to conventional methods. It, however, is criticized for its complexity-exhaustive deep neural network (DNN) architectures, especially on resource-constrained mobile platforms. Thus we devise a two-stage approach: first, a range pre-processing is applied to constrain the dynamic range of feature map activation by leveraging its sparsity nature with densely clustered distribution, then a layer-wise range-adaptive quantization for convolutional parameter (e.g., weight, bias), and simple yet efficient linear scaling and range-dependent normalization for activation are executed, leading to a fully fixed-point inference architecture. All arithmetic operations and associated data tensors are processed using low-bit-width fixed-point numbers, yielding significant reductions of the computational complexity, memory space, and the elimination of platform-dependent inconsistency induced by floating-point operations. We first exemplify such fixed-point inference in a DNN-based image decoder, showing the comparable coding efficiency with its native floating-point model, against the same anchor using the High-Efficiency Video Coding (HEVC)-based intra image coder. We also extend proposed approach to the super-resolution network for learned resolution scaling-based video streaming, and VGG network-based classification tasks, both of which present negligible performance loss. These evidence the generalization of our approach for efficient DNN processing of various tasks. All materials are made publicly accessible at http://njuvision.github.io/fixed-point/. Weixin Hong, Tong Chen 0004, Ming Lu 0003, Shiliang Pu, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | End-to-End Learnt Image Compression via Non-Local Attention Optimization and Improved Context ModelingabstractThis article proposes an end-to-end learnt lossy image compression approach, which is built on top of the deep nerual network (DNN)-based variational auto-encoder (VAE) structure with Non-Local Attention optimization and Improved Context modeling (NLAIC). Our NLAIC 1) embeds non-local network operations as non-linear transforms in both main and hyper coders for deriving respective latent features and hyperpriors by exploiting both local and global correlations, 2) applies attention mechanism to generate implicit masks that are used to weigh the features for adaptive bit allocation, and 3) implements the improved conditional entropy modeling of latent features using joint 3D convolutional neural network (CNN)-based autoregressive contexts and hyperpriors. Towards the practical application, additional enhancements are also introduced to speed up the computational processing (e.g., parallel 3D CNN-based context prediction), decrease the memory consumption (e.g., sparse non-local processing) and reduce the implementation complexity (e.g., a unified model for variable rates without re-training). The proposed model outperforms existing learnt and conventional (e.g., BPG, JPEG2000, JPEG) image compression methods, on both Kodak and Tecnick datasets with the state-of-the-art compression efficiency, for both PSNR and MS-SSIM quality measurements. We have made all materials publicly accessible at https://njuvision.github.io/NIC for reproducible research. Tong Chen 0004, Zhan Ma 0001, Qiu Shen, Xun Cao, Yao Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Learned Video Compression via Joint Spatial-Temporal Correlation ExplorationabstractTraditional video compression technologies have been developed over decades in pursuit of higher coding efficiency. Efficient temporal information representation plays a key role in video coding. Thus, in this paper, we propose to exploit the temporal correlation using both first-order optical flow and second-order flow prediction. We suggest an one-stage learning approach to encapsulate flow as quantized features from consecutive frames which is then entropy coded with adaptive contexts conditioned on joint spatial-temporal priors to exploit second-order correlations. Joint priors are embedded in autoregressive spatial neighbors, co-located hyper elements and temporal neighbors using ConvLSTM recurrently. We evaluate our approach for the low-delay scenario with High-Efficiency Video Coding (H.265/HEVC), H.264/AVC and another learned video compression method, following the common test settings. Our work offers the state-of-the-art performance, with consistent gains across all popular test sequences. Lichao Huang, Ming Lu 0003, Tong Chen 0004, Zhan Ma 0001 |
AAAI | 5 |
| 2020 | Variable Bitrate Image Compression with Quality Scaling FactorsabstractRecently, learned image compression has emerged with significant coding efficiency improvement, and even shown noticeable gains over the state-of-the-art traditional codecs. In the mean time, most existing methods need to train separate models for different bitrate target. In this paper, we propose to embed a set of quality scaling factors (SFs) into learned image compression network, by which we can encode images across an entire bitrate range with a single model. This solution offers the comparable performance with those default approaches requiring multiple bitrate dependent models, and reduces the complexity significantly for practical implementation. Our work also demonstrates the generalization for various compression network structures, image contents, and training loss functions. Tong Chen 0004, Zhan Ma 0001 |
ICASSP | 1 |
| 2019 | Looking-Ahead: Neural Future Video Frame PredictionabstractWe have developed a Looking-Ahead system to facilitate the future video frame prediction via deep learning, which is of practical value in the domain like autonomous driving etc. The overall problem is decomposed into cascaded optical flow prediction and subsequent predictive frame post-processing for quality refinement. A pyramid flow calculation across existing frames is used to efficiently infer the motion of target frame; while a universal inpainting network is applied to restore those motion-induced occluded pixels. Compared with those published methods, our Looking-Ahead offers the state-of-the-art performance measured objectively with better Peak-Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM), and more appealing reconstructions. Changxu Zhang, Tong Chen 0004, Qiu Shen, Zhan Ma 0001 |
ICIP | 2 |
| 2019 | Extreme Image Coding via Multiscale Autoencoders with Generative Adversarial OptimizationabstractWe propose a MultiScale AutoEncoder (MSAE) based extreme image coding/compression framework to offer visually pleasing reconstruction at a very low bitrate. Our method leverages the "priors" at different resolution scale to improve the compression efficiency, and also employs the generative adversarial network (GAN) with multiscale discriminators to perform the end-to-end trainable rate-distortion optimization. We compare the perceptual quality of our reconstructions with traditional compression algorithms using High-Efficiency Video Coding (HEVC) based Intra Profile and JPEG2000 on the public Cityscapes, ADE20K and Kodak datasets, demonstrating the significant subjective quality improvement. However, objective measurements, such as PSNR, SSIM, etc, are often deteriorated by applying the generative adversarial optimization. Tong Chen 0004, Qiu Shen, Zhan Ma 0001 |
VCIP | 3 |
| 2019 | Codedretrieval: Joint Image Compression and Retrieval with Neural NetworksabstractWith the explosive increase of image data, the efficiency of both image compression and retrieval becomes unprecedentedly significant. However, these two tasks are usually isolated executed, which waste great computational resources in large-scale image applications. In this work, we propose a joint framework called CodedRetrieval, which can find a general feature expression for both compression and retrieval based on neural network. Additionally, a two stage training strategy is designed to achieve better balance between the two distinct tasks. Experimental results show that our method can achieve competitive perfomance on both compression and retrieval comparing to classic methods, while saving great amount of computation time. Tong Chen 0004, Qiu Shen, Zhan Ma 0001 |
VCIP | 3 |
| 2017 | DeepCoder: A deep neural network based video compressionabstractInspired by recent advances in deep learning, we present the DeepCoder - a Convolutional Neural Network (CNN) based video compression framework. We apply separate CNN nets for predictive and residual signals respectively. Scalar quantization and Huffman coding are employed to encode the quantized feature maps (fMaps) into binary stream. We use the fixed 32 × 32 block in this work to demonstrate our ideas, and performance comparison is conducted with the well-known H.264/AVC video coding standard with comparable rate-distortion performance. Here distortion is measured using Structural Similarity (SSIM) because it is more close to perceptual response. Tong Chen 0004, Qiu Shen, Tao Yue 0003, Xun Cao, Zhan Ma 0001 |
VCIP | 1 |