VLDB 2026 Research / reviewers in the wild / expert
Dandan Ding
dblp:71/7248
· DBLP profile ↗
64ranked-venue papers
11as first author
49since 2021 · last 2026
0000-0003-2911-1321ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 7 first-author · 40 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive AV2 In-loop Filtering via Guided Neural Model with Vectorized Quantization
Kequan Mao, Dandan Ding, Urvang Joshi, Debargha Mukherjee |
ISCAS | 2 |
| 2026 | Accelerating QTMT partitioning in versatile video coding using lightweight parameter-conditioned dynamic routing (recommended by ChinaMM 2025)
Xianlu Bian, Guosheng Yu, Dandan Ding |
Multim. Syst. | 4 |
| 2026 | Improving Occupancy Prediction for Multiscale Point Cloud Geometry CompressionabstractMultiscale sparse representation offers significant advantages in point cloud geometry compression, delivering state-of-the-art performance compared to both standardized solutions and other learned approaches. A crucial component of this framework is the cross-scale occupancy prediction, which employs the lower-scale reference representation either from the current frame alone or from both the current and temporal reference frames to establish conditional priors for either static or dynamic coding. However, existing works mainly use local computations,e.g., sparse convolutions andkNN attention, to exploit correlations in such a representation; these methods usually fail to adequately capture global coherence. In addition, the fixed configuration of lossless-lossy scales cannot adapt to temporal dynamics, which limits the reconstruction quality of temporal references in dynamic coding. These limitations constrain the generation of more effective priors used for conditional coding. To address these issues, we propose two new techniques. The first is KPA (Key Point-driven Attention), which integrates both local and global characteristics. The second is AdaScale (Adaptive Lossy/Lossless Scale), which decides whether the transitional scale should be in lossless or lossy mode based on temporal displacement, thereby enhancing the reconstruction quality of the temporal reference. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, including rules-based standard codecs like G-PCC and V-PCC, as well as learning-based approaches like Unicorn and TMAP, across both static/dynamic and lossy/lossless coding scenarios. Zehong Li, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | EDRIC: Embracing 1D Autoencoder for Real-Time Lossy LiDAR Reflectance CompressionabstractWhile recent advancements in LiDAR reflectance compression have improved rate-distortion performance, real-time processing remains an unresolved challenge. In this work, we introduce EDRIC, a highly effective neural compression framework that offers state-of-the-art compression efficiency while achieving real-time capability. To overcome the suboptimal downsampling scheme and inefficient feature extraction in conventional 3D frameworks, EDRIC serializes a 3D point cloud into a 1D sequence and introduces a lightweight 1D autoencoder to efficiently compress the serialized LiDAR reflectance signal. In addition, we explicitly incorporate geometric priors through a geometry-aware entropy model, effectively exploiting the interdependencies between reflectance attributes and underlying geometry. Extensive experiments on representative datasets (e.g., KITTI, Ford, nuScenes, and QNX) demonstrate that EDRIC achieves 8.38%~9.85% BD-BR reduction compared to the latest G-PCCv23 (RAHT) standard while operating $22\times $ faster (e.g., >50 frames per second on an RTX 4090 GPU). Furthermore, EDRIC comprises merely 2.3M parameters, making it a practical solution for real-world deployment. Kang You, Kequan Mao, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | RENO: Real-Time Neural Compression for 3D LiDAR Point CloudsabstractDespite the substantial advancements demonstrated by learning-based neural models in the LiDAR Point Cloud Compression (LPCC) task, realizing real-time compression—an indispensable criterion for numerous industrial applications—remains a formidable challenge. This paper proposes RENO, the first real-time neural codec for 3D LiDAR point clouds, achieving superior performance with a lightweight model. RENO skips the octree construction and directly builds upon the multiscale sparse tensor representation. Instead of the multi-stage inferring, RENO devises sparse occupancy codes, which exploit cross-scale correlation and derive voxels’ occupancy in a one-shot manner, greatly saving processing time. Experimental results demonstrate that the proposed RENO achieves real-time coding speed, 10 fps at 14-bit depth on a desktop platform (e.g., one RTX 3090 GPU) for both encoding and decoding processes, while providing 12.25% and 48.34% bit-rate savings compared to G-PCCv23 and Draco, respectively, at a similar quality. RENO model size is merely 1MB, making it attractive for practical applications. The source code is available at https://github.com/NJUVISION/RENO. Kang You, Tong Chen 0004, Dandan Ding, Muhammad Salman Asif, Zhan Ma 0001 |
CVPR | 3 |
| 2025 | Super Resolution-Based Video Coding via Lightweight Implicit Neural ModelingabstractThe super-resolution (SR)-based coding tool is widely employed in modern video coding standards. By encoding video frames at a reduced resolution and then restoring them to their original resolution during the in-loop filtering stage, this tool helps to further reduce the bitrate and improve the coding performance. Current video coding standards typically devise rule-based SR methods in their codecs, compromising the coding efficiency to maintain low computational complexity. As deep neural network (DNN)-based SR methods are proving more effective than rule-based approaches, this paper proposes integrating the neural SR into video codecs to enhance coding performance while minimizing the computational cost. To this end, we propose a Lightweight Implicit Neural Model (LIM). Specifically, our LIM, consisting of Lightweight Feature Aggregation Network (LFANet) and Coordinate Upsampling Network Based on B-spline Representation (CURNet), is developed to support SR-based coding at an arbitrary scale. We exemplify the proposed method on the ongoing AVM reference software and conduct extensive experiments to demonstrate its effectiveness. Compared with anchored AVM, our method improves the BD-Rate by 5.52%, which significantly outperforms state-of-the-art works. Meanwhile, its computational complexity is much lower than others, having only 22.5k parameters and 18.7k FLOPs/pixel complexity, which is attractive to real-world applications. Xianlu Bian, Dandan Ding, Urvang Joshi, Debargha Mukherjee |
DCC | 3 |
| 2025 | Optical Flow-Driven Fast CU Partition for Inter Prediction in Versatile Video Coding
Junhao Jiang, Shuangxing Tian, Dandan Ding |
ICIG (1) | 3 |
| 2025 | Efficient LiDAR Reflectance Compression via Scanning SerializationabstractReflectance attributes in LiDAR point clouds provide essential information for downstream tasks but remain underexplored in neural compression methods. To address this, we introduce SerLiC, a serialization-based neural compression framework to fully exploit the intrinsic characteristics of LiDAR reflectance. SerLiC first transforms 3D LiDAR point clouds into 1D sequences via scan-order serialization, offering a device-centric perspective for reflectance analysis. Each point is then tokenized into a contextual representation comprising its sensor scanning index, radial distance, and prior reflectance, for effective dependencies exploration. For efficient sequential modeling, Mamba is incorporated with a dual parallelization scheme, enabling simultaneous autoregressive dependency capture and fast processing. Extensive experiments demonstrate that SerLiC attains over 2$\times$ volume reduction against the original reflectance data, outperforming the state-of-the-art method by up to 22% reduction of compressed bits while using only 2% of its parameters. Moreover, a lightweight version of SerLiC achieves $\geq 10$ fps (frames per second) with just 111K parameters, which is attractive for real applications. Kang You, Dandan Ding, Zhan Ma 0001 |
ICML | 3 |
| 2025 | GeoQE: Enhancing Quality of Experience in Point Cloud Streaming
Chengfeng Han, Dandan Ding, Zhan Ma 0001 |
ACM Multimedia | 3 |
| 2025 | ARNet: Attribute Artifact Reduction for G-PCC Compressed Point CloudsabstractA learning-based adaptive loop filter is developed for the geometry-based point-cloud compression (G-PCC) standard to reduce attribute compression artifacts. The proposed method first generates multiple most probable sample offsets (MPSOs) as potential compression distortion approximations, and then linearly weights them for artifact mitigation. Therefore, we drive the filtered reconstruction as closely to the uncompressed PCA as possible. To this end, we devise an attribute artifact reduction network (ARNet) consisting of two consecutive processing phases: MPSOs derivation and MPSOs combination. The MPSOs derivation uses a two-stream network to model local neighborhood variations from direct spatial embedding and frequency-dependent embedding, where sparse convolutions are utilized to best aggregate information from sparsely and irregularly distributed points. The MPSOs combination is guided by the least-squares error metric to derive weighting coefficients on the fly to further capture the content dynamics of the input PCAs. ARNet is implemented as an in-loop filtering tool for G-PCC, where the linear weighting coefficients are encapsulated into the bitstream with negligible bitrate overhead. The experimental results demonstrate significant improvements over the latest G-PCC both subjectively and objectively. For example, our method offers a 22.12% YUV Bj⊘ntegaard delta rate (BD-Rate) reduction compared to G-PCC across various commonly used test point clouds. Compared with a recent study showing state-of-the-art performance, our work not only gains 13.23% YUV BD-Rate but also provides a 30 × processing speedup. Junteng Zhang, Dandan Ding, Zhan Ma 0001 |
Comput. Vis. Media | 3 |
| 2025 | Edge detection-driven LightGBM for fast intra partition of H.266/VVC
Guosheng Yu, Xianlu Bian, Dandan Ding |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | A Versatile Point Cloud Compressor Using Universal Multiscale Conditional Coding - Part II: AttributeabstractA universal multiscale conditional coding framework, Unicorn, is proposed to code the geometry and attribute of any given point cloud. Attribute compression is discussed in Part II of this paper, while geometry compression is given in Part I of this paper. We first construct the multiscale sparse tensors of each voxelized point cloud attribute frame. Since attribute components exhibit very different intrinsic characteristics from the geometry element, e.g., 8-bit RGB color versus 1-bit occupancy, we process the attribute residual between lower-scale reconstruction and current-scale data. Similarly, we leverage spatially lower-scale priors in the current frame and (previously processed) temporal reference frame to improve the probability estimation of attribute intensity through conditional residual prediction in lossless mode or enhance the attribute reconstruction through progressive residual refinement in lossy mode for better performance. The proposed Unicorn is a versatile, learning-based solution capable of compressing a great variety of static and dynamic point clouds in both lossy and lossless modes. Following the same evaluation criteria, Unicorn significantly outperforms standard-compliant approaches like MPEG G-PCC, V-PCC, and other learning-based solutions, yielding state-of-the-art compression efficiency with affordable encoding/decoding runtime. Jianqiang Wang 0006, Ruixiang Xue, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | A Versatile Point Cloud Compressor Using Universal Multiscale Conditional Coding - Part I: GeometryabstractA universal multiscale conditional coding framework, Unicorn, is proposed to compress the geometry and attribute of any given point cloud. Geometry compression is addressed in Part I of this paper, while attribute compression is discussed in Part II. We construct the multiscale sparse tensors of each voxelized point cloud frame and properly leverage lower-scale priors in the current and (previously processed) temporal reference frames to improve the conditional probability approximation or content-aware predictive reconstruction of geometry occupancy in compression. Unicorn is a versatile, learning-based solution capable of compressing static and dynamic point clouds with diverse source characteristics in both lossy and lossless modes. Following the same evaluation criteria, Unicorn significantly outperforms standard-compliant approaches like MPEG G-PCC, V-PCC, and other learning-based solutions, yielding state-of-the-art compression efficiency while presenting affordable complexity for practical implementations. Jianqiang Wang 0006, Ruixiang Xue, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | ConPCAC: Conditional Lossless Point Cloud Attribute Compression via Spatial DecompositionabstractA conditional lossless point cloud attribute compression method, dubbed ConPCAC, is proposed. The previous work typically codes point attributes in a point cloud in an autoregressive way, incurring unbearable coding time. By contrast, ConPCAC proposes a group-wise conditional entropy model for fast coding while preserving coding performance. Specifically, ConPCAC adopts a “Group Decomposition - Attribute Initialization - Latent Distribution Prediction” framework. First, it flexibly decomposes the original point cloud into multiple groups according to the geometry coordinate distribution. Then, the first group is coded using a base coder, e.g., the standardized G-PCC, and the following groups are progressively coded using a neural coder conditioned on their preceding groups. Two key units, Attribute Initialization (Init) and Latent Distribution Prediction (LDP), are devised in the neural coder. The Init unit employs the nearest neighbor to initialize the attributes of a group, and the LDP unit further predicts the attribute probability distribution for the group. In this way, ConPCAC enables full correlation exploration across groups and parallel processing among points in a group. Finally, the predicted probabilities are fed into the arithmetic engine to code the true attribute values of each group. Extensive experiments demonstrate the performance of ConPCAC. It achieves 14.59%, 10.32%, and 12.26% improvements over the latest G-PCC on the widely used 8iVFB, Owlii, and MVUB datasets, respectively, significantly outperforming state-of-the-art lossless PCAC methods. Moreover, its computational complexity is comparable to G-PCC and much lower than existing learning-based methods. Associated code and models will be released on the websitehttps://github.com/3dpcc/ConPCAC. Tong Chen 0004, Kang You, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Neural Compression System for Point Cloud Video StreamingabstractPoint cloud video streaming is promising for immersive media applications, which urges the development of efficient compression methods. However, existing approaches either suffer from poor performance or lack effective coder control mechanisms, making them impractical for networked point cloud services, where bandwidth is often constrained and fluctuates over time. Therefore, this paper proposes a system-level solution - a layered point cloud compressor, called Yak, to address these issues. Yak offers comprehensive support for both intra and inter-frame coding of geometry and attribute components in point cloud sequences. It consists of three layers: the Base Layer uses the standard G-PCC to encode a thumbnail counterpart downscaled from the input point cloud; the Enhancement Layer devises the end-to-end variational autoencoder to compress the original input conditioned on the base layer reconstruction, and the Dynamic Layer generates feature-space predictions as the temporal prior for conditional inter-frame coding. In addition, Yak devises the Content Analysis module to dynamically determine the optimal encoding parameters of each frame, by which bit budget is intelligently allocated for geometry and attribute components to maximize the overall rate-distortion (R-D) performance. Such accurate rate control relies on the parametric rate/distortion models whose parameters are initialized through one-pass template matching and frame-wise delta updating constrained by R-D optimization. Following standard evaluation guidelines, Yak has notably outperformed traditional rules-based methods such as MPEG G-PCC and V-PCC, as well as other learning-based approaches, while offering flexible networked adaption and affordable complexity. Junteng Zhang, Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Scalable Point Cloud Attribute CompressionabstractThis paper develops a Scalable Point Cloud Attribute Compression solution, termedScalablePCAC. In a two-layer example,ScalablePCACuses the standard G-PCC at the base layer to directly encode the thumbnail point cloud that is downscaled from the original input, and a learning-based model at the enhancement layer to compress and restore the full-resolution input point cloud conditioned on the base layer reconstruction. As such, the base layer provides a coarse reconstruction of the input point cloud and the enhancement layer further improves the quality. We then adopt a cross-layer rate allocation strategy that flexibly determines the resolution downscaling factor, the quantization parameter of the base layer, and the quality controlling factor of the enhancement layer to adapt the bitrate of the two layers for approximately optimal Rate-Distortion (R-D) performance. We conduct extensive experiments on popular point clouds following the MPEG common test conditions. Results demonstrate that the proposedScalablePCACachieves$>$10% BD-BR reduction against the latest G-PCC version 22 (TMC13v22) on the Y component; it also significantly outperforms existing learning-based solutions for point cloud attribute compression,e.g., compared with a recent work showing state-of-the-art performance, it achieves$>$20% BD-BR reduction. Junteng Zhang, Jianqiang Wang 0006, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | 4D Gaussian Videos with Motion LayeringabstractOnline free-view navigation in volumetric videos requires high-quality rendering and real-time streaming in order to provide immersive user experiences. However, existing methods ( e.g. , dynamic NeRF and 3DGS) may not handle dynamic scenes with complex motions, and their models may not be streamable due to storage and bandwidth constraints. In this paper, we propose a novel 4D Gaussian Video (4DGV) approach that enables the creation and streaming of photorealistic, volumetric videos for dynamic scenes over the Internet. The core of our 4DGV is a novel streamable group of Gaussians (GOG) representation based on motion layering. Each GOG consists of static and dynamic points obtained via lifting 2D segmentation into 3D in motion layering, where the deformation of each dynamic point is represented as the temporal offset of its attributes. We also adaptively convert static points back to dynamic points to handle the appearance change, (e.g. , moving shadows and reflections), of static objects through optimization. To support real-time streaming of 4DGVs, we show that by applying quantization on Gaussian attributes and H.265 encoding on deformation offsets, our GOG representation can be significantly compressed (to around 6% of the original model size) without sacrificing the accuracy (PSNR loss less than 0.01dB). Extensive experiments on standard benchmarks demonstrate that our method outperforms state-of-the-art volumetric video approaches, with superior rendering quality and minimum storage overheads. Pinxuan Dai, Peiquan Zhang, Ke Xu 0010, Yifan Peng 0001, Dandan Ding, Yujun Shen, Yin Yang 0002, Xinguo Liu, Rynson W. H. Lau, Weiwei Xu 0003 |
ACM Trans. Graph. | 6 |
| 2025 | Revisit Point Cloud Quality Assessment: Current Advances and a Multiscale-Inspired ApproachabstractThe demand for full-reference point cloud quality assessment (PCQA) has extended across various point cloud services. Unlike image quality assessment, where the reference and the distorted images are naturally aligned in coordinates and thus allow point-to-point (P2P) color assessment, the coordinates and attributes of a 3D point cloud may both suffer from distortion, making the P2P evaluation unsuitable. To address this, PCQA methods usually define a set of key points and construct a neighborhood around each key point for neighbor-to-neighbor (N2N) computation on geometry and attribute. However, state-of-the-art PCQA methods often exhibit limitations in certain scenarios due to insufficient consideration of key points and neighborhoods. To overcome these challenges, this paper proposes PQI, a simple yet efficient metric to index point cloud quality. PQI suggests using scale-wise key points to uniformly perceive distortions within a point cloud, along with a mild neighborhood size associated with each key point for compromised N2N computation. To achieve this, PQI employs a multiscale framework to obtain key points, ensuring comprehensive feature representation and distortion detection throughout the entire point cloud. Such a multiscale method merges every eight points into one in the downsampling processing, implicitly embedding neighborhood information into a single point and thereby eliminating the need for an explicitly large neighborhood. Further, within each neighborhood, simple features, such as geometry Euclidean distance difference and attribute value difference, are extracted. Feature similarity is then calculated between the reference and the distorted samples at each scale and linearly weighted to generate the final PQI score. Extensive experiments demonstrate the superiority of PQI, consistently achieving high performance across several widely recognized PCQA datasets. Moreover, PQI is highly appealing for practical applications due to its low complexity and flexible scale options. Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Learning to Restore Compressed Point Cloud Attribute: A Fully Data-Driven Approach and a Rules-Unrolling-Based OptimizationabstractThe emergence of holographic media drives the standardization of Geometry-based Point Cloud Compression (G-PCC) to sustain networked service provisioning. However, G-PCC inevitably introduces visually annoying artifacts, degrading the quality of experience (QoE). This work focuses on restoring G-PCC compressed point cloud attributes, e.g., RGB colors, to which fully data-driven and rules-unrolling-based post-processing filters are studied. At first, as compressed attributes exhibit nested blockiness, we develop a learning-based sample adaptive offset (NeuralSAO), which leverages a neural model using multiscale feature aggregation and embedding to characterize local correlations for quantization error compensation. Later, given statistically Gaussian distributed quantization noise, we suggest the utilization of a bilateral filter with Gaussian kernels to weigh neighbors by jointly considering their geometric and photometric contributions for restoration. Since local signals often present varying distributions, we propose estimating the smoothing parameters of the bilateral filter using an ultra-lightweight neural model. Such a bilateral filter with learnable parameters is called NeuralBF. The proposed NeuralSAO demonstrates the state-of-art restoration quality improvement, e.g., 20% BD-BR (Bjøntegaard delta rate) reduction over G-PCC on solid points clouds. However, NeuralSAO is computationally intensive and may suffer from poor generalization. On the other hand, although NeuralBF only achieves half of the gains of NeuralSAO, it is lightweight and exhibits impressive generalization across various samples. This comparative study between the data-driven large-scale NeuralSAO and the rules-unrolling-based small-scale NeuralBF helps to understand the capacity (i.e., performance, complexity, generalization) of underlying filters in terms of the quality restoration for compressed point cloud attribute. Junteng Zhang, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Another Way to the Top: Exploit Contextual Clustering in Learned Image CodingabstractWhile convolution and self-attention are extensively used in learned image compression (LIC) for transform coding, this paper proposes an alternative called Contextual Clustering based LIC (CLIC) which primarily relies on clustering operations and local attention for correlation characterization and compact representation of an image. As seen, CLIC expands the receptive field into the entire image for intra-cluster feature aggregation. Afterward, features are reordered to their original spatial positions to pass through the local attention units for inter-cluster embedding. Additionally, we introduce the Guided Post-Quantization Filtering (GuidedPQF) into CLIC, effectively mitigating the propagation and accumulation of quantization errors at the initial decoding stage. Extensive experiments demonstrate the superior performance of CLIC over state-of-the-art works: when optimized using MSE, it outperforms VVC by about 10% BD-Rate in three widely-used benchmark datasets; when optimized using MS-SSIM, it saves more than 50% BD-Rate over VVC. Our CLIC offers a new way to generate compact representations for image compression, which also provides a novel direction along the line of LIC development. Zhihao Duan, Ming Lu 0003, Dandan Ding, Fengqing Zhu 0001, Zhan Ma 0001 |
AAAI | 4 |
| 2024 | NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image SequenceabstractThis paper proposes the NeRI, an implicit neural representation (INR) based LiDAR point cloud compressor. In NeRI, we first transform a sequence of 3D LiDAR frames into a 2D range image sequence through range image projection over time. Then, we employ a neural network conditioned on the temporal frame index and associated LiDAR sensor pose to fit input range images as closely as possible. The optimized network parameters, which implicitly represent the input LiDAR data, are later lossily compressed. NeRI decoder is then initialized using decoded parameters to generate range images for reconstructing the 3D LiDAR sequence accordingly. Extensive experimental results demonstrate the significant superiority of NeRI regarding the compression efficiency and decoding speed compared to state-of-the-art 2D and 3D compressors for LiDAR point cloud. Ruixiang Xue, Tong Chen 0004, Dandan Ding, Xun Cao, Zhan Ma 0001 |
ICASSP | 4 |
| 2024 | Encoding Auxiliary Information to Restore Compressed Point Cloud Geometry
Gexin Liu, Dandan Ding, Zhan Ma 0001 |
IJCAI | 3 |
| 2024 | Pointsoup: High-Performance and Extremely Low-Decoding-Latency Learned Geometry Codec for Large-Scale Point Cloud Scenes
Kang You, Li Yu 0004, Pan Gao 0001, Dandan Ding |
IJCAI | 5 |
| 2024 | ELIM: Extremely Low-Complexity Implicit Neural Model for Super Resolution-Based CodingabstractThe super-resolution (SR)-based coding, which encodes a frame at a reduced resolution to achieve a lower bitrate, is a prevalent tool used in modern video coding standards. Accordingly, the low-resolution frame is restored to the full resolution at the reconstruction stage for subsequent reference. Therefore, the resolution restoration algorithm significantly affects the coding performance. This paper devises a highly efficient and extremely low complexity implicit neural model (ELIM) for SR-based encoding to support arbitrary scale factors. Specifically, ELIM consists of two stages: Feature Aggregation and Coordinate Upsampling. In Feature Aggregation, we embed a simplified attention block to the U-Net style framework to collect valuable information while reducing computational complexity through downsampling. In Coordinate Upsampling, in addition to the extracted content features, information including coordinate relative location and pixel cell size is fused to achieve better performance. We exemplify ELIM on the AV2 codec (the next generation of AV1). Extensive experiments demonstrate its superior performance: it achieves 4.17% BD-Rate gains over the anchor AV2 reference software with only 5,185 flops/pixel, significantly surpassing existing methods. The low complexity of ELIM is attractive to real applications. Dandan Ding, Urvang Joshi, Debargha Mukherjee |
PCS | 3 |
| 2024 | Compressing 3D Gaussian Splatting via a Generalizable Neural CoderabstractAs a promising technique for 3D representation, 3D Gaussian Splatting (3DGS) offers fast rendering speed and high fidelity while generating large data volumes. This challenges storage and transmission, so an efficient compression solution is required. Existing implicit methods require pre-scene optimization (online), leading to a long optimization time. By contrast, this paper regards the 3DGS as a point cloud and pre-trains a generalizable (offline) neural coder for compression. After obtaining the 3DGS representation, we focus on the data compression process, which is friendly to applications already equipped with a PCC codec. The neural coder employed is extended from a typical AIbased point cloud compression method, which uses a multiscale and multistage framework to exploit spatial correlations across scales and stages for conditional coding. Experimental results show that our method significantly outperforms existing 3DGS representations without compromising fidelity, achieving more than 39× and 6.8× compression ratio compared to the original 3DGS and SOTA Scaffold-GS, respectively. More importantly, our approach does not require additional time to optimize the compression model. Junteng Zhang, Tong Chen 0004, Hao Zhu 0004, Dandan Ding, Zhan Ma 0001 |
VCIP | 5 |
| 2024 | Leveraging occupancy map to accelerate video-based point cloud compressionabstractVideo-based Point Cloud Compression enables point cloud streaming over the internet by converting dynamic 3D point clouds to 2D geometry and attribute videos, which are then compressed using 2D video codecs like H.266/VVC. However, the complex encoding process of H.266/VVC, such as the quadtree with nested multi-type tree (QTMT) partition, greatly hinders the practical application of V-PCC. To address this issue, we propose a fast CU partition method dedicated to V-PCC to accelerate the coding process. Specifically, we classify coding units (CUs) of projected images into three categories based on the occupancy map of a point cloud: unoccupied, partially occupied, and fully occupied. Subsequently, we employ either statistic-based rules or machine-learning models to manage the partition of each category. For unoccupied CUs, we terminate the partition directly; for partially occupied CUs with explicit directions, we selectively skip certain partition candidates; for the remaining CUs (partially occupied CUs with complex directions and fully occupied CUs), we train an edge-driven LightGBM model to predict the partition probability of each partition candidate automatically. Only partitions with high probabilities are retained for further Rate–Distortion (R–D) decisions. Comprehensive experiments demonstrate the superior performance of our proposed method: under the V-PCC common test conditions , our method reduces encoding time by 52% and 44% in geometry and attribute, respectively, while incurring only 0.68% (0.66%) BD-Rate loss in D1 (D2) measurements and 0.79% (luma) BD-Rate loss in attribute, significantly surpassing state-of-the-art works. Gongchun Ding, Dandan Ding |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Content-Aware Rate Control for Geometry-Based Point Cloud CompressionabstractThe Geometry-based Point Cloud Compression (G-PCC) standard enables point cloud delivery over the internet through efficient compression. Limited by the transmission bandwidth, rate control is demanded in G-PCC for high-quality point cloud video streaming. This paper thus proposes a content-aware rate control solution for G-PCC. Given the target bitrate and distortion evaluation criteria, our method can predict the geometry and attribute quantizers for G-PCC while minimizing the overall distortion. Specifically, as the rate and distortion of both geometry and attribute are involved in G-PCC, we separately establish rate/distortion models for geometry and attribute. Moreover, recognizing the dependence of attribute compression on reconstructed geometry, we integrate the geometry quantizer into the attribute rate/distortion models to improve prediction accuracy. For dynamic coding scenarios, we leverage selective representative frames for efficient model parameter initialization. Additionally, we introduce a μ updating strategy that dynamically incorporates information from previous frames to update the existing models. Extensive experiments demonstrate the effectiveness of our proposed method. Under the G-PCC common test condition, our method achieves remarkable rate accuracy, with a 5.3% bitrate error for static coding and 0.3% for dynamic coding. Moreover, it achieves >15% BD-Rate gains over the G-PCC anchor. These results showcase its capabilities in delivering high-fidelity point cloud video streams within the bandwidth constraint. Junteng Zhang, Wenxi Ma, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | On Content-Aware Post-Processing: Adapting Statistically Learned Models to Dynamic ContentabstractLearning-based post-processing methods generally produce neural models that are statistically optimal on their training datasets. These models, however, neglect intrinsic variations of local video content and may fail to process unseen content. To address this issue, this article proposes a content-aware approach for the post-processing of compressed videos. We develop a backbone network, called BackboneFormer , where a Fast Transformer using Separable Self-Attention, Spatial Attention, and Channel Attention is devised to support underlying feature embedding and aggregation. Furthermore, we introduce Meta-learning to strengthen BackboneFormer for better performance. Specifically, we propose Meta Post-Processing (Meta-PP) which leverages the Meta-learning framework to drive BackboneFormer to capture and analyze input video variations for spontaneous updating. Since the original frame is unavailable to the decoder, we devise a Compression Degradation Estimation model where a low-complexity neural model and classic operators are used collaboratively to estimate the compression distortion. The estimated distortion is then utilized to guide the BackboneFormer model for dynamic updating of weighting parameters. Experimental results demonstrate that the proposed BackboneFormer itself gains about 3.61% Bjøntegaard delta bit-rate reduction over Versatile Video Coding in the post-processing task and “BackboneFormer + Meta-PP” attains 4.32%, costing only 50K and 61K parameters, respectively. The computational complexity of MACs is 49k/pixel and 50k/pixel, which represents only about 16% of state-of-the-art methods having similar coding gains. Gongchun Ding, Dandan Ding, Zhan Ma 0001, Zhu Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | A Reconfigurable Framework for Neural Network Based Video In-Loop FilteringabstractThis article proposes a reconfigurable framework for neural network based video in-loop filtering to guide large-scale models for content-aware processing. Specifically, the backbone neural model is decomposed into several convolutional groups and the encoder systematically traverses all candidate configurations combined by these groups to find the best one. The selected configuration index is then encapsulated as side information and passed to the decoder, enabling dynamic model reconfiguration during the decoding stage. The preceding reconfiguration process is only deployed in the inference stage on top of a pre-trained backbone model. Furthermore, we devise WMSPFormer , a wavelet multi-scale Poolformer, as the backbone network structure. WMSPFormer utilizes a wavelet-based multi-scale structure to losslessly decompose the input into multiple scales for spatial-spectral features aggregation. Moreover, it uses multi-scale pooling operations ( MSPoolformer ) instead of complicated matrix calculations to substitute the attention process. We also extend MSPoolformer to a large-scale version using more parameters, referred to as MSPoolformerExt . Extensive experiments demonstrate that the proposed WMSPFormer+Reconfig. and WMSPFormerExt+Reconfig. achieve a remarkable 7.13% and 7.92% BD-Rate reduction over the anchor H.266/VVC, outperforming most existing methods evaluated under the same training and testing conditions. In addition, the low-complexity nature of the WMSPFormer series makes it attractive for practical applications. Dandan Ding, Zhan Ma 0001, Zhu Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | GRNet:Geometry Restoration for G-PCC Compressed Point Clouds Using Auxiliary Density SignalingabstractThe lossy Geometry-based Point Cloud Compression (G-PCC) inevitably impairs the geometry information of point clouds, which deteriorates the quality of experience (QoE) in reconstruction and/or misleads decisions in tasks such as classification. To tackle it, this work proposes GRNet for the geometry restoration of G-PCC compressed large-scale point clouds. By analyzing the content characteristics of original and G-PCC compressed point clouds, we attribute the G-PCC distortion to two key factors: point vanishing and point displacement. Visible impairments on a point cloud are usually dominated by an individual factor or superimposed by both factors, which are determined by the density of the original point cloud. To this end, we employ two different models for coordinate reconstruction, termed Coordinate Expansion and Coordinate Refinement, to attack the point vanishing and displacement, respectively. In addition, 4-byte auxiliary density information is signaled in the bitstream to assist the selection of Coordinate Expansion, Coordinate Refinement, or their combination. Before being fed into the coordinate reconstruction module, the G-PCC compressed point cloud is first processed by a Feature Analysis Module for multiscale information fusion, in which kNN-based Transformer is leveraged at each scale to adaptively characterize neighborhood geometric dynamics for effective restoration. Following the common test conditions recommended in the MPEG standardization committee, GRNet significantly improves the G-PCC anchor and remarkably outperforms state-of-the-art methods on a great variety of point clouds (e.g., solid, dense, and sparse samples) both quantitatively and qualitatively. Meanwhile, GRNet runs fairly fast and uses a smaller-size model when compared with existing learning-based approaches, making it attractive to industry practitioners. Gexin Liu, Ruixiang Xue, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Channel Estimation and Detection for Intelligent Reflecting Surface-Assisted Orthogonal Time Frequency Space SystemsabstractOrthogonal time frequency space (OTFS) modulation is a promising technique for the next-generation communications in high-mobility scenarios. However, in the delay-Doppler (DD) domain, the received signals suffer from a decrease in power due to the non-coherent superposition of all symbols transmitted through wireless channels. To address this issue, this paper proposes the incorporation of an intelligent reflecting surface (IRS) to assist the transmission for the OTFS systems, and jointly designs the OTFS frame structure and IRS phase shifts to achieve a coherent combination of the received signals. To address the problem of outdated channel state information in high-mobility systems, we propose a location-aided channel estimation strategy at the IRS. Additionally, to mitigate the adverse effects of the fractional Doppler shifts, a delay and shifted-Doppler domain-based channel estimation method is designed at the base station. By utilizing the well-designed OTFS frame structure and IRS phase shifts, we propose a low-complexity iterative interference cancellation (IIC) detector, and analyze the lower bound for its symbol error probability. To provide a clear understanding for the process of the considered systems, we describe a two-stage transmission protocol. Finally, the numerical results are provided to evaluate the effectiveness and superiority of the proposed estimation methods and IIC detector. Qin Tao, Taoyu Xie, Xiaoling Hu 0001, Shuowen Zhang, Dandan Ding |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | Lossless Point Cloud Attribute Compression Using Cross-scale, Cross-group, and Cross-color PredictionabstractThis work extends the multiscale structure originally developed for point cloud geometry compression to point cloud attribute compression. To losslessly encode the attribute while maintaining a low bitrate, accurate probability prediction is critical. With this aim, we extensively exploit cross-scale, cross-group, and cross-color correlations of point cloud attribute to ensure accurate probability estimation and thus high coding efficiency. Specifically, we first generate multiscale attribute tensors through average pooling, by which, for any two consecutive scales, the decoded lower-scale attribute can be used to estimate the attribute probability in the current scale in one shot. Additionally, in each scale, we perform the probability estimation group-wisely following a predefined grouping pattern. In this way, both cross-scale and (same-scale) cross-group correlations are exploited jointly. Furthermore, cross-color redundancy is removed by allowing inter-color processing for YCoCg/RGB alike multi-channel attributes. The proposed method not only demonstrates state-of-the-art compression efficiency with significant performance gains over the latest G-PCC on various contents but also sustains low complexity with affordable encoding and decoding runtime. Jianqiang Wang 0006, Dandan Ding, Zhan Ma 0001 |
DCC | 2 |
| 2023 | G-PCC++: Enhanced Geometry-based Point Cloud CompressionabstractMPEG Geometry-based Point Cloud Compression (G-PCC) standard is developed for lossy encoding of point clouds to enable immersive services over the Internet. However, lossy G-PCC introduces superimposed distortions from both geometry and attribute information, seriously deteriorating the Quality of Experience (QoE). This paper thus proposes the Enhanced G-PCC (GPCC++), to effectively address the compression distortion and restore the quality. G-PCC++ separates the enhancement into two stages: it first enhances the geometry and then maps the decoded attribute to the enhanced geometry for refinement. As for geometry restoration, a k Nearest Neighbors (kNN)-based Linear Interpolation is first used to generate a denser geometry representation, on top of which GeoNet further generates sufficient candidates to restore geometry through probability-sorted selection. For attribute enhancement, a kNN-based Gaussian Distance Weighted Mapping is devised to re-colorize all points in enhanced geometry tensor, which are then refined by AttNet for the final reconstruction. G-PCC++ is the first solution addressing the geometry and attribute artifacts together. Extensive experiments on several public datasets demonstrate the superiority of G-PCC++, e.g., on the solid point cloud dataset 8iVFB, G-PCC++ outperforms G-PCC by 88.24% (80.54%) BD-BR in D1 (D2) measurement of geometry and by 14.64% (13.09%) BD-BR in Y (YUV) attribute. Moreover, when considering both geometry and attribute, G-PCC++ also largely surpasses G-PCC by 25.58% BD-BR using PCQM assessment. Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
ACM Multimedia | 3 |
| 2023 | YOGA: Yet Another Geometry-based Point Cloud CompressorabstractA learning-based YOGA (Yet Another Geometry-based Point Cloud Compressor) is proposed. It is flexible, allowing for the separable lossy compression of geometry and color attributes, and variable-rate coding using a single neural model; it is high-efficiency, significantly outperforming the latest G-PCC standard quantitatively and qualitatively, e.g., 25% BD-BR gains using PCQM (Point Cloud Quality Metric) as the distortion assessment, and it is lightweight, e.g., similar runtime as the G-PCC codec, owing to the use of sparse convolution and parallel entropy coding. To this end, YOGA adopts a unified end-to-end learning-based backbone for separate geometry and attribute compression. The backbone uses a two-layer structure, where the downscaled thumbnail point cloud is encoded using G-PCC at the base layer, and upon G-PCC compressed priors, multiscale sparse convolutions are stacked at the enhancement layer to effectively characterize spatial correlations to compactly represent the full-resolution sample. In addition, YOGA integrates the adaptive quantization and entropy model group to enable variable-rate control, as well as adaptive filters for better quality restoration. Junteng Zhang, Tong Chen 0004, Dandan Ding, Zhan Ma 0001 |
ACM Multimedia | 3 |
| 2023 | SPNE: sample-perturbed network entropy for revealing critical states of complex biological systemsabstractComplex biological systems do not always develop smoothly but occasionally undergo a sharp transition; i.e. there exists a critical transition or tipping point at which a drastic qualitative shift occurs. Hunting for such a critical transition is important to prevent or delay the occurrence of catastrophic consequences, such as disease deterioration. However, the identification of the critical state for complex biological systems is still a challenging problem when using high-dimensional small sample data, especially where only a certain sample is available, which often leads to the failure of most traditional statistical approaches. In this study, a novel quantitative method, sample-perturbed network entropy (SPNE), is developed based on the sample-perturbed directed network to reveal the critical state of complex biological systems at the single-sample level. Specifically, the SPNE approach effectively quantifies the perturbation effect caused by a specific sample on the directed network in terms of network entropy and thus captures the criticality of biological systems. This model-free method was applied to both bulk and single-cell expression data. Our approach was validated by successfully detecting the early warning signals of the critical states for six real datasets, including four tumor datasets from The Cancer Genome Atlas (TCGA) and two single-cell datasets of cell differentiation. In addition, the functional analyses of signaling biomarkers demonstrated the effectiveness of the analytical and computational results. Jiayuan Zhong, Dandan Ding, Juntan Liu, Rui Liu 0009, Pei Chen 0004 |
Briefings Bioinform. | 2 |
| 2023 | Accelerating QTMT-based CU partition and intra mode decision for versatile video coding
Gongchun Ding, Xiujun Lin, Dandan Ding |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Sparse Tensor-Based Multiscale Representation for Point Cloud Geometry CompressionabstractThis study develops a unified Point Cloud Geometry (PCG) compression method through the processing of multiscale sparse tensor-based voxelized PCG. We call this compression method SparsePCGC. The proposed SparsePCGC is a low complexity solution because it only performs the convolutions on sparsely-distributed Most-Probable Positively-Occupied Voxels (MP-POV). The multiscale representation also allows us to compress scale-wise MP-POVs by exploiting cross-scale and same-scale correlations extensively and flexibly. The overall compression efficiency highly depends on the accuracy of estimated occupancy probability for each MP-POV. Thus, we first design the Sparse Convolution-based Neural Network (SparseCNN) which stacks sparse convolutions and voxel sampling to best characterize and embed spatial correlations. We then develop the SparseCNN-based Occupancy Probability Approximation (SOPA) model to estimate the occupancy probability either in a single-stage manner only using the cross-scale correlation, or in a multi-stage manner by exploiting stage-wise correlation among same-scale neighbors. Besides, we also suggest the SparseCNN based Local Neighborhood Embedding (SLNE) to aggregate local variations as spatial priors in feature attribute to improve the SOPA. Our unified approach not only shows state-of-the-art performance in both lossless and lossy compression modes across a variety of datasets including the dense object PCGs (8iVFB, Owlii, MUVB) and sparse LiDAR PCGs (KITTI, Ford) when compared with standardized MPEG G-PCC and other prevalent learning-based schemes, but also has low complexity which is attractive to practical applications. Jianqiang Wang 0006, Dandan Ding, Zhu Li 0001, Xiaoxing Feng, Chuntong Cao, Zhan Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Neural Adaptive Loop Filtering for Video Coding: Exploring Multi-Hypothesis Sample RefinementabstractAdaptive loop filtering (ALF) is extensively investigated for lossy video coding to mitigate compression noise. Numerous learning-based ALFs have emerged recently and improved the coding efficiency significantly through the use of complexity-intensive, large-scale models trained on excessive samples, making it impractical for real-life applications. By contrast, lightweight, small-scale ALF models cannot promise convincing performance and model generalization. In principle, the ALF estimates the sample distortion for restoration. Instead of directly approximating the distortion as in existing solutions, we reformulate it as a Multi-hypothesis Sample Refinement (MSR) problem. To this end, we first generate multiple distortion hypotheses through a deep neural network (DNN) model. Then, these hypotheses are linearly superimposed to approximate the final distortion through the minimization of mean square error (MMSE) between the filtered reconstruction and its original, uncompressed input. Finally, the linear superimposition coefficients are explicitly signaled in the compressed bitstream. As seen, the superimposition coefficients inherently generalize the MSR to various content. And using DNNs to generate distortion hypotheses essentially models the spatial priors of a local block and its underlying compression error distribution. As a result, the MSR using a small-scale convolutional neural network (CNN) model with only 5k parameters and 5 KMACs/pixel achieves 4.35% (Intra) and 2.49% (Inter) BD-Rate (Bjøntegaard Delta Rate) gains over the AV1 anchor. Ablation studies further demonstrate that the MSR can be generalized to diverse small-scale network structures, different standards (e.g., H.265/HEVC and H.266/VVC), and diverse video content (e.g., screen videos). Dandan Ding, Guangkun Zhen, Debargha Mukherjee, Urvang Joshi, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Decoder-Side Cross Resolution Synthesis for Video Compression EnhancementabstractThis paper proposes a decoder-side Cross Resolution Synthesis (CRS) module to pursue better compression efficiency beyond the latest Versatile Video Coding (VVC), where we encode intra frames at original high resolution (HR), compress inter frames at a lower resolution (LR), and then super-resolve decoded LR inter frames with the help from preceding HR intra and neighboring LR inter frames. For a LR inter frame, a motion alignment and aggregation network (MAN) is devised to produce temporally aggregated motion representation to best guarantee the temporal smoothness; Another texture compensation network (TCN) is utilized to generate texture representation from decoded HR intra frame for better augmenting spatial details; Finally, a similarity-driven fusion engine synthesizes motion and texture representations to upscale LR inter frames for the removal of compression and resolution re-sampling noises. We enhance the VVC using proposed CRS, showing averaged 8.76% and 11.93% Bjntegaard Delta Rate (BD-Rate) gains against the latest VVC anchor in Random Access (RA) and Low-delay P (LDP) settings respectively. In addition, experimental comparisons to the state-of-the-art super-resolution (SR) based VVC enhancement methods, and ablation studies are conducted to further report superior efficiency and generalization of the proposed algorithm. All materials will be made to public at https://njuvision.github.io/CRS for reproducible research. Ming Lu 0003, Tong Chen 0004, Zhenyu Dai, Dandan Ding, Zhan Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Quadtree-based Guided CNN for AV1 In-loop FilteringabstractRecently, learning-based in-loop filtering has attracted lots of attention. State-of-the-art works generally deploy computationally expensive, large-scale neural networks, which is unfriendly to practical applications. Besides, since these models are generally pre-trained using a limited dataset and applied to various videos, they may fail in video contents excluded in the training dataset. To address these issues, this paper develops a Guide CNN in-loop filtering framework to obtain the restored signal. Our basic idea is to construct a subspace and use the projection of the original signal into this subspace to approximate the original signal itself. Specifically, we employ CNN to transform the degraded signal into M subsignals to construct the optimal subspace since the training of CNN is essentially an optimization procedure. Furthermore, unlike existing CNN models that process all blocks uniformly, our method leverages a quadtree structure to implement the Guided CNN through R-D optimization. As such, the best partition to Guided CNN can be determined. We exemplify the proposed method in AV1 codec. Experimental results show that the Guided CNN framework achieves 2.19% and 1.31% BD-Rate gains over the AV1 anchor in intra and inter coding mode, respectively, while the normal CNN achieves only 1.64% and 1.04%. Gongchun Ding, Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 3 |
| 2022 | PCGFormer: Lossy Point Cloud Geometry Compression via Local Self-AttentionabstractAlthough the multiscale sparse tensor using stacked convolutions has attained noticeable gains for lossy compression of point cloud geometry (PCG), its capability suffers because convolutions with fixed receptive field and fixed weights after training cannot aggregate sufficient information collection due to the extremely sparse and unevenly distributed nature of points. To best tackle the sparsity and adaptively exploit inter-point correlations, we apply local self-attention on$k$nearest neighbors (kNN) that are instantaneously formed for each point, with which attention-based mechanism can effectively characterize and embed spatial information conditioned on the dynamic neighborhood. This kNN self-attention is implemented using the prevalent Transformer architecture and stacked with sparse convolutions to capture neighborhood information in a progres-sively re-sampling framework, referred to as the PCGFormer. Compared with the MPEG standard Geometry-based PCC (G-PCC) using the latest octree codec, the proposed PCGFormer provides more than 90% and 87% BD-rate (Bjøntegaard Delta Rate) reduction in average across three different object point cloud datasets for point-to-point (D1) and point-to-plane (D2) distortion measures. Compared with the state-of-the-art learning-based approach, the PCGFormer achieves 17.39% and 15.75% BD-rate gains on D1 and D2, respectively. Gexin Liu, Jianqiang Wang 0006, Dandan Ding, Zhan Ma 0001 |
VCIP | 3 |
| 2022 | Low Light RAW Image Enhancement Using Paired Fast Fourier Convolution and TransformerabstractCompared to RGB images non-linearly mapped from RAW data through the Image Signal Processor (ISP), RAW data are linear to scene radiance and contain more native information, which is better to be modeled in many vision tasks. This work proposes to enhance low-light images in the RAW domain via a cross-scale framework using paired Fast Fourier Convolution (FFC) and Transformer, driving the network to characterize images effectively. The entire framework has three scales to abstract low-level, mid-level, and high-level representations of input images. We embed paired FFC and Transformer in each scale to attain spatial-spectral information extraction and aggregation. Specifically, by transforming features from the spatial domain into the spectral domain with FFC, pixel correlations can be effectively exploited locally and globally, generating representative features for the input image. Immediately, the Transformer using multi-head self-attention mechanism is applied to aggregate and embed important features. Experimental results demonstrate that our method significantly outperforms state-of-the-art low-light enhancement works in both full reference assessment metrics, including PSNR, MPSNR, and SSIM, and no-reference metrics, such as NIMA. Meanwhile, the perceptual quality of the proposed method is more visually pleasing than that of other methods. Hengyu Liu 0004, Dandan Ding, Zhan Ma 0001 |
VCIP | 3 |
| 2022 | Biprediction-Based Video Quality Enhancement via LearningabstractConvolutional neural networks (CNNs)-based video quality enhancement generally employs optical flow for pixelwise motion estimation and compensation, followed by utilizing motion-compensated frames and jointly exploring the spatiotemporal correlation across frames to facilitate the enhancement. This method, called the optical-flow-based method (OPT), usually achieves high accuracy at the expense of high computational complexity. In this article, we develop a new framework, referred to as biprediction-based multiframe video enhancement (PMVE), to achieve a one-pass enhancement procedure. PMVE designs two networks, that is, the prediction network (Pred-net) and the frame-fusion network (FF-net), to implement the two steps of synthesization and fusion, respectively. Specifically, the Pred-net leverages frame pairs to synthesize the so-called virtual frames (VFs) for those low-quality frames (LFs) through biprediction. Afterward, the slowly fused FF-net takes the VFs as the input to extract the correlation across the VFs and the related LFs, to obtain an enhanced version of those LFs. Such a framework allows PMVE to leverage the cross-correlation between successive frames for enhancement, hence capable of achieving high accuracy performance. Meanwhile, PMVE effectively avoids the explicit operations of motion estimation and compensation, hence greatly reducing the complexity compared to OPT. The experimental results demonstrate that the peak signal-to-noise ratio (PSNR) performance of PMVE is fully on par with that of OPT while its computational complexity is only 1% of OPT. Compared with other state-of-the-art methods in the literature, PMVE is also confirmed to achieve superior performance in both objective quality and visual quality at a reasonable complexity level. For instance, PMVE can surpass its best counterpart method by up to 0.42 dB in PSNR. Dandan Ding, Junchao Tong, Xinbo Gao 0001, Zoe Liu, Yong Fang 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Neural Reference Synthesis for Inter Frame CodingabstractThis work proposes the neural reference synthesis (NRS) to generate high-fidelity reference block for motion estimation and motion compensation (MEMC) in inter frame coding. The NRS is comprised of two submodules: one for reconstruction enhancement and the other for reference generation. Although numerous methods have been developed in the past for these two submodules using either handcrafted rules or deep convolutional neural network (CNN) models, they basically deal with them separately, resulting in limited coding gains. By contrast, the NRS proposes to optimize them collaboratively. It first develops two CNN-based models, namely EnhNet and GenNet. The EnhNet only uses spatial correlations within the current frame for reconstruction enhancement and the GenNet is then augmented by further aggregating temporal correlations across multiple frames for reference synthesis. However, a direct concatenation of EnhNet and GenNet without considering the complex temporal reference dependency across inter frames would implicitly induce iterative CNN processing and cause the data overfitting problem, leading to visually-disturbing artifacts and oversmoothed pixels. To tackle this problem, the NRS applies a new training strategy to coordinate the EnhNet and GenNet for more robust and generalizable models, and also devises a lightweight multi-level R-D (rate-distortion) selection policy for the encoder to adaptively choose reference blocks generated from the proposed NRS model or conventional coding process. Our NRS not only offers state-of-the-art coding gains, e.g., >10% BD-Rate (Bjøntegaard Delta Rate) reduction against the High Efficiency Video Coding (HEVC) anchor for a variety of common test video sequences encoded at a wide bit range in both low-delay and random access settings, but also greatly reduces the complexity relative to existing learning-based methods by utilizing more lightweight DNNs. All models are made publicly accessible at https://github.com/IVC-Projects/NRS for reproducible research. Dandan Ding, Chenran Tang, Zhan Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Multiscale Point Cloud Geometry CompressionabstractRecent years have witnessed the growth of point cloud based applications for both immersive media as well as 3D sensing for auto-driving, because of its realistic and fine-grained representation of 3D objects and scenes. However, it is a challenging problem to compress sparse, unstructured, and high-precision 3D points for efficient communication. In this paper, leveraging the sparsity nature of the point cloud, we propose a multiscale end-to-end learning framework that hierarchically reconstructs the 3D Point Cloud Geometry (PCG) via progressive re-sampling. The framework is developed on top of a sparse convolution based autoencoder for point cloud compression and reconstruction. For the input PCG which has only the binary occupancy attribute, our framework translates it to a down-scaled point cloud at the bottleneck layer which possesses both geometry and associated feature attributes. Then, the geometric occupancy is losslessly compressed using an octree codec and the feature attributes are lossy compressed using a learned probabilistic context model. Compared with the state-of-the-art Video-based Point Cloud Compression (V-PCC) and Geometry-based PCC (G-PCC) schemes standardized by the Moving Picture Experts Group (MPEG), our method achieves more than 40% and 70% BD-Rate (BjØntegaard Delta Rate) reduction, respectively. We would like to make all materials publicly accessible at https://njuvision.github.io/PCGCv2/ for reproducible research. Jianqiang Wang 0006, Dandan Ding, Zhu Li 0001, Zhan Ma 0001 |
DCC | 2 |
| 2021 | Advances in Video Compression System Using Deep Neural Network: A Review and Case StudiesabstractSignificant advances in video compression systems have been made in the past several decades to satisfy the near-exponential growth of Internet-scale video traffic. From the application perspective, we have identified three major functional blocks, including preprocessing, coding, and postprocessing, which have been continuously investigated to maximize the end-user quality of experience (QoE) under a limited bit rate budget. Recently, artificial intelligence (AI)-powered techniques have shown great potential to further increase the efficiency of the aforementioned functional blocks, both individually and jointly. In this article, we review recent technical advances in video compression systems extensively, with an emphasis on deep neural network (DNN)-based approaches, and then present three comprehensive case studies. On preprocessing, we show a switchable texture-based video coding example that leverages DNN-based scene understanding to extract semantic areas for the improvement of a subsequent video coder. On coding, we present an end-to-end neural video coding framework that takes advantage of the stacked DNNs to efficiently and compactly code input raw videos via fully data-driven learning. On postprocessing, we demonstrate two neural adaptive filters to, respectively, facilitate the in-loop and postfiltering for the enhancement of compressed frames. Finally, a companion website hosting the contents developed in this work can be accessed publicly at https://purdueviper.github.io/dnn-coding/. Dandan Ding, Zhan Ma 0001, Qingshuang Chen, Zoe Liu, Fengqing Zhu 0001 |
Proc. IEEE | 1 |
| 2021 | A progressive CNN in-loop filtering approach for inter frame coding
Dandan Ding, Lingyi Kong, Fengqing Zhu 0001 |
Signal Process. Image Commun. | 1 |
| 2021 | Point Cloud Upsampling via Perturbation LearningabstractGiven sparse point clouds, this paper develops a perturbation learning-based point cloud upsampling method to generate uniform, clean, and dense point clouds. We build a simple yet efficient neural network framework including feature extraction, perturbation learning, and coordinate reconstruction operations. In the feature extraction task, shallow-and-wide dense connections are applied to present the latent geometric information. Subsequently, the extracted features are expanded for perturbation learning. According to the theory of the differential geometry of surfaces, the position of an upsampled point can be approximated by its projection on a tangent plane in a sufficiently small neighborhood around the point. Inspired by this, we propose learning a 2D perturbation through multilayer perceptrons (MLPs) to estimate the coordinate shift from the input point to the upsampled point. Then, we concatenate the 2D perturbation with the extracted features for residual learning to fine-tune the coordinate shift. Finally, the coordinate reconstruction step transforms all the high-level features into an upsampled and consolidated point cloud. To enable the learning-based point cloud upsampling process above, we collect a large-scale point cloud dataset that contains 36000 pairs of training sets. The entire network size is only 5.01 MB, which is much smaller than the requirements of state-of-the-art point cloud upsampling models. Qualitative and quantitative evaluation results show that our proposed scheme outperforms most existing methods in terms of the Chamfer distance (CD), Hausdorff distance (HD), Jensen-Shannon divergence (JSD), and uniformity. In addition, our method is applied to real scan data, and its robustness is confirmed. Dandan Ding, Chi Qiu, Fuchang Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Lossy Point Cloud Geometry Compression via Region-Wise ProcessingabstractPoint cloud geometry (PCG) is used to precisely represent arbitrary-shaped 3D objects and scenes, is of great interest to vast applications which puts forward the pressing desire of high-efficiency PCG compression for transmission and storage. Existing PCG coding mostly relies on the octree model by which point-wise processing is applied without exploring nonlocal regional geometry similarity across the entire 3D surface. This work, instead, suggests the region-wise processing to leverage the region similarity to exploit inter-region redundancy for efficient lossy point cloud geometry compression. Towards this goal, a given PCG is first segmented into numerous local regions each of which comprises a portion of point cloud surface, and can be represented by a surface vector that describes the geometry shape numerically in a projected principal space. Subsequently, these regions are grouped into several discriminative clusters, assuring that inter-cluster similarity is minimized and intra-cluster similarity is maximized simultaneously, where the similarity is calculated using the regional surface vectors. In each cluster, we set a reference region having the largest similarity score to the others, which enables the non-reference region prediction from the reference one using alignment transform. In the end, we encode the reference regions directly using the lossless mode of the Geometry-based Point Cloud Compression (G-PCC), while corresponding non-reference regions are signaled using associated transform parameters. Compared with the state-of-the-art G-PCC using octree model, our region-wise approach can offer remarkable coding efficiency improvement, e.g., 32.4% and 22.0% Bjontegaard-delta rate (BD-Rate) gains for respective point-to-point ($D1$) and point-to-plane ($D2$) distortion evaluations, across a variety of common test sequences used in standard committee. Wenjie Zhu 0004, Yiling Xu, Dandan Ding, Zhan Ma 0001, Mike Nilsson |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Guided CNN Restoration with Explicitly Signaled Linear CombinationabstractState-of-the-art Convolutional Neural Network (CNN) based loop restoration generally involves a CNN structure with a large number of parameters and applies the CNN model to those degraded frames uniformly to generate their restored version, even though the contents within these frames are different. By contrast, in this paper, we propose a Guided CNN Restoration (GNR) scheme, where a CNN is used in conjunction with explicitly signaled guide parameters, with an aim to adapt the CNN model to different input contents. Specifically, the CNN architecture is designed such that the final restoration is constrained within the subspace generated by various output channels of the CNN, and meanwhile the weighting parameters for a linear combination of the output channels to obtain the final restoration are explicitly signaled by the encoders. The proposed GNR is incorporated into an AV1 encoder to replace the anchor in-loop filters and the weighting parameters are written into the encoded bitstream. Experimental results show that given a small CNN with 3,312 parameters, the proposed approach achieves a BD-rate reduction of 3.06% over the AV1 anchor, while the traditional CNN-based method only achieves 1.39%. Lingyi Kong, Dandan Ding, Fuchang Liu, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 2 |
| 2020 | A Switchable Deep Learning Approach for In-Loop Filtering in Video CodingabstractDeep learning provides a great potential for in-loop filtering to improve both coding efficiency and subjective quality in video coding. State-of-the-art work focuses on network structure design and employs a single powerful network to solve all problems. In contrast, this paper proposes a deep learning based systematic approach that includes an effective Convolutional Neural Network (CNN) structure, a hierarchical training strategy, and a video codec oriented switchable mechanism. First, we propose a novel CNN structure, i.e., Squeeze-and-Excitation Filtering CNN (SEFCNN), as an optional in-loop filter. To capture the non-linear interaction between channels, the SEFCNN is comprised of two subnets, i.e., Feature EXtracting (FEX) subnet and Feature ENhancing (FEN) subnet. Then, we develop a hierarchical model training strategy to adapt the two subnets to different coding scenarios. For high-rate videos with small artifacts, we train a single global model using the FEX for all types of frames, whereas for low-rate videos with large artifacts, different models are trained using both FEX and FEN for different types of frames. Finally, we propose an adaptive enhancing mechanism which is switchable between the CNN-based and the conventional methods. We selectively apply the CNN model to some frames or some regions in a frame. Experimental results show that the proposed scheme outperforms state-of-the-art work in coding efficiency, while the computational complexity is acceptable after GPU acceleration. Dandan Ding, Lingyi Kong, Zoe Liu, Yong Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | AV1 in-loop Filtering using a Wide-Activation Structured Residual NetworkabstractThe in-loop filter, which constitutes an important part in modern video coding, improves both subjective and objective quality of reconstructed frames. Lately, Convolutional Neural Network (CNN) has demonstrated its superiority over traditional methods in addressing in-loop filtering problem. In this paper, we develop a CNN-based in-loop filter, namely Wide Activation Residual Network (WARN), for AV1 encoder. On top of the plain Residual Network (ResNet), we introduce wide activation to each residual block, making a more reasonable allocation of network parameters. When incorporating WARN into video encoder, particular to inter coding, it is intricate to obtain the global optimum performance. After simplifying this as an end-to-end trainable problem, we propose a skipping method by taking advantage of the hierarchical reference structure in AV1. Experimental results show that our WARN achieves up to 14.42% and 9.64% BD-rate reduction in intra and inter coding, respectively. All the code and model of our approach are available at https://github.com/IVC-Projects/AV1_WARN. Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 2 |
| 2019 | Learning-Based Multi-Frame Video Quality EnhancementabstractConvolution neural network (CNN) has shown its great success in video quality enhancement. Existing methods mainly conduct enhancement tasks in the spatial domain, exploring the pixel correlations within one frame. Taking advantage of the similarity across successive frames, this paper develops a learning-based multi-frame approach, with an aim to explore the greatest potential for video quality enhancement leveraging the temporal correlation. First, we apply a learning-based optical flow to compensate the temporal motion across neighboring frames. Afterwards, a deep CNN network, which is structured in an early-fusion manner, is designed to discover the joint spatial-temporal correlations within a video. To ensure the generality of our CNN model, we further propose a robust training strategy. One high-quality frame and one moderate-quality frame are paired to enhance the remaining low-quality frames in between, which considers a trade-off between frame distances and various frame quality. Experimental results demonstrate that our method outperforms state-of-the-art work in objective quality. The code and model of our approach are published in Github (https://github.com/IVC-Projects/LMVE). Junchao Tong, Xilin Wu, Dandan Ding, Zoe Liu |
ICIP | 3 |
| 2019 | A CNN-based In-loop Filtering Approach for AV1 Video CodecabstractIn-loop filter using Convolutional Neural Network (CNN) has lately attracted lots of attention in video coding. CNN models may be trained to learn how to restore degradation introduced by compression in pictures, and hence effectively help improve the coding efficiency. State-of-the-art work in this field generally employs a single network to enhance reconstructed frames mainly in intra coding. In this paper, we develop a depth-variable network handling both intra and inter coding. The depth of our network is varied with the distortion levels of reconstructed frames. Moreover, we leverage a skip enhancing strategy for inter coding, which improves both the coding efficiency and the resulting visual quality, while maintaining low computational complexity. We apply our approach to AV1, a newly released video coding standard from AOM. Experimental results show that our approach achieves an average BD-rate reduction of 7.27% and 5.57% for intra and inter modes, respectively, compared to AV1 anchor. The code and model of our approach are published in our Github website [1]. Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
PCS | 1 |
| 2018 | Retrieving indoor objects: 2D-3D alignment using single image and interactive ROI-based refinement
Fuchang Liu, Shuangjian Wang, Dandan Ding, Qingshu Yuan, Zhengwei Yao, Hai-Sheng Li 0002 |
Comput. Graph. | 3 |
| 2014 | A reconfiguration system for video decoderabstractThis demonstration system shows a kind of video decoder's implementation in Reconfigurable Video Coding (RVC) framework on Open RVC-CAL Compiler (Orcc) platform. Differently from tradition video decoder, the reconfigurable video decoder is not a decoder conforming a special video coding standard, but dynamically built according to actual bitsteams, which may not conform any standard. The reconfigurable video decoder receives not only the compressed video bitstream but also the decoder description. As an example, in this demo, we reconfigure AVS and H.264/AVC decoders using Just-In-Time Adaptive Decoder Engine (Jade). Honggang Qi, Dandan Ding, Lu Yu 0003 |
VCIP | 3 |
| 2014 | A hardware-oriented IME algorithm and its implementation for HEVCabstractThe flexible coding structure in High Efficiency Video Coding (HEVC) introduces many challenges to real-time implementation of the integer-pel motion estimation (IME). In this paper, a hardware-oriented IME algorithm naming parallel clustering tree search (PCTS) is proposed, where various prediction units (PU) are processed simultaneously with a parallel scheme. The PCTS consists of four hierarchical search steps. After each search step, PUs with the same MV candidate are clustered to one group. And the next search step is shared by PUs in the same group. Owing to the top-down tree-structure search strategy of the PCTS, search processes are highly shared among different PUs and system throughput is thus significantly increased. As a result, the hardware implementation based on the proposed algorithm can support real-time video applications of QFHD (3840×2160) at 30fps. Dandan Ding, Lu Yu 0003 |
VCIP | 2 |
| 2014 | A cost-efficient hardware architecture of deblocking filter in HEVCabstractThis paper presents a hardware architecture of deblocking filter (DBF) for High Efficiency Video Coding (HEVC) by jointly considering system throughput and hardware cost. A hybrid pipeline with two processing levels is adopted to improve system performance. With the hybrid pipeline, only one 1-D filter and single-port on-chip SRAM are used. According to the data dependence between neighbouring edges, a shifted 16×16 basic processing unit as well as corresponding filtering order is proposed. It reduces memory cost and makes the DBF friendlier to work in a coding/decoding system. The proposed hardware architecture is synthesized under 0.13um standard CMOS technology and result shows that it consumes 17.6k gates at an operating frequency of 250MHz. Consequently, the design can support real-time processing of QFHD (3840×2160) video applications at 60 fps. Dandan Ding, Lu Yu 0003 |
VCIP | 2 |
| 2013 | A hardware CABAC encoder for HEVCabstractThis paper presents a hardware design of context-based adaptive binary arithmetic coding (CABAC) for the emerging High efficiency video coding (HEVC) standard. While aiming at higher compression efficiency, the CABAC in HEVC also invests a lot of effort in the pursuit of parallelism and reducing hardware cost. Simulation results show that our design processes 1.18 bins per cycle on average. It can work at 357 MHz with 48.940K gates targeting 0.13 μm CMOS process. This processing rate can support real-time encoding for all sequences under common test conditions of HEVC standard conforming to the main profile level 6.1 of main tier or main profile level 5.1 of high tier. Dandan Ding, Xingguo Zhu, Lu Yu 0003 |
ISCAS | 2 |
| 2013 | On hardware architecture and processing order of HEVC intra prediction moduleabstractThis article presents a parallel and memory optimized hardware architecture for intra prediction of the High Efficiency Video Coding (HEVC) standard. The architecture consists of 64 parallel reconfigurable Processing Elements as datapaths and supports all 35 intra prediction modes and all prediction sizes from 4×4 to 64×64. In order to avoid implementing large area memory-datapaths interconnections and save memory usage, the maximum number of reference registers is reduced from 129 to 72 by reclassifying 35 prediction modes into 3 general categories. In addition, a 3 stage hierarchical processing order including an S-shaped scan order of blocks and a Bidirectional Ring Register File is proposed to avoid bandwidth bottleneck and increase system throughput. This architecture is synthesized using TSMC 130nm technology and the working frequency is up to 400MHz with 324K gates area. Running at 300MHz, it supports real time 1080p@60fps full modes and full sizes HEVC intra prediction. Dandan Ding, Lu Yu 0003 |
PCS | 2 |
| 2011 | An efficient hardware design for HDTV H.264/AVC encoderabstractThis paper presents a hardware efficient high definition television (HDTV) encoder for H.264/AVC. We use a two-level mode decision (MD) mechanism to reduce the complexity and maintain the performance, and design a sharable architecture for normal mode fractional motion estimation (NFME), special mode fractional motion estimation (SFME), and luma motion compensation (LMC), to decrease the hardware cost. Based on these technologies, we adopt a four-stage macro-block pipeline scheme using an efficient memory management strategy for the system, which greatly reduces on-chip memory and bandwidth requirements. The proposed encoder uses about 1126k gates with an average Bjontegaard-Delta peak signal-to-noise ratio (BD-PSNR) decrease of 0.5 dB, compared with JM15.0. It can fully satisfy the real-time video encoding for 1080p@30 frames/s of H.264/AVC high profile. Dandan Ding, Bin-bin Yu, Lu Yu 0003 |
J. Zhejiang Univ. Sci. C | 2 |
| 2009 | A Proposed AVS Decoder Configuration in the Reconfigurable Video Coding FrameworkabstractThis demonstration shows an AVS intra decoder configuration in the RVC framework. It explains how to use the dataflow mechanism offered by the RVC framework to support AVS decoder configuration. It also shows the flexibility and convenience to reconfigure decoders in the RVC framework. In this work, the AVS VTL is established containing FUs from AVS. The proposed AVS decoder configuration is implemented by connecting some FUs from AVS VTL and reusing some FUs from MPEG VTL. The demonstration shows that the decoder can decode AVS conformance bitstreams correctly in RVC simulator. Dandan Ding, Honggang Qi, Lu Yu 0003, Tiejun Huang 0001, Wen Gao 0001 |
ISCAS | 1 |
| 2009 | A Hybrid Decoder Configuration of MPEG-4 and AVS in Reconfigurable Video Coding FrameworkabstractThis demonstration shows that the RVC framework offers a great flexibility in selecting coding tools for decoder reconfigurations to satisfy a wide variety of different applications by showing a hybrid decoder reconfiguration using coding tools from MPEG and AVS Video Tool Library (VTL). Compared with the original MPEG-4 Simple Profile (SP), complexity of the reconfigured decoder is reduced whereas the performance is improved gradually as the bitrate increases. Dandan Ding, Lu Yu 0003, Christophe Lucarz, Marco Mattavelli |
ISCAS | 1 |
| 2009 | Reconfigurable video coding framework and decoder reconfiguration instantiation of AVS
Dandan Ding, Honggang Qi, Lu Yu 0003, Tiejun Huang 0001, Wen Gao 0001 |
Signal Process. Image Commun. | 1 |