Fangdong Chen

dblp:125/2304 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
6since 2021 · last 2024
0000-0003-1422-7276ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2024 A Novel Visually-Lossless Compression Model for Low-Latency Interaction
abstract
Perceptual Lossless Compression (PLC) is a novel compression standard directed by the Audio Video coding Standard (AVS) work group. It defines a lightweight, low-latency, and visually lossless image compression framework, which offers an alternative mezzanine codec for most user-agnostic on-chip compression scenarios, alleviating the tension between growing transmission demands and expensive integration upgrades. In this paper, the technical designs in the development process of PLC will be fully introduced. The balance between feature modeling and ASIC implementation costs will be present throughout. A high throughput and low hardware complexity implementation will be detailed and evaluated namely HIM. Hopefully, the design of the PLC standard and HIM framework will bring new inspiration for the emerging low-latency interaction systems.
Huiwen Ren, Zetian Song, Danni Wang, Dongping Pan, Haitao Yang 0001, Fangdong Chen, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.9
2023 An Efficient Rate Control Scheme for Video Compression in Low-latency Interoperable Interfaces
abstract
Lightweight video compression has effectively alleviated the tension between growing transmission demands and expensive integration upgrades. Effective rate control algorithms are believed to be the crucial bottleneck for quality improvement during those ultra-high throughput coding processes. This paper proposes a novel rate control (RC) scheme that constructs a contextual adaptive bit estimation model through clustering historical compression information into block-gradient complexity categories. A buffer-aware tuning method and a flexible quantization parameter (QP) mapping algorithm are designed to determine the Luma/Chroma QP distribution where a simplified Lagrangian multiplier is further defined to preserve the stability of the overall compression process. As a result, the constant-bitrate compression towards low-latency interoperable ASICs is implemented with a promising RC performance.
Huiwen Ren, Zetian Song, Yan Wang 0011, Shanshe Wang, Fangdong Chen, Shiliang Pu, Siwei Ma 0001, Wen Gao 0001
DCC6
2023 Pixel-Wise Quantization for Image Compression
abstract
This paper proposes a pixel-wise quantization (PWQ) method, which allows to reduce the quantization parameters (QPs) of simple pixels adaptively for the purpose of enhancing the subjective quality, since the distortions on simple pixels are more noticeable than those on complex pixels. For the pixel-wise prediction in Fig. 1, the pixel-wise reconstruction is implemented and the transformation is disabled, where the symbol “=” (or $^{\prime \prime}\vee^{\prime \prime}/^{\prime \prime}\gt^{\prime \prime}$) means the current prediction is the average value of the left and right reconstructions (or the upper/left reconstruction). And the PWQ method is applied in the same prediction direction and reconstruction order, with adjusting the current pixel QP $(Q_{pixel})$ adaptively by (1), where Qcbdenotes the current block $\mathrm{Q}\mathrm{P}, T_{pred}$ denotes the predicted texture complexity based on the neighboring reconstruction pixels, and parameters $\delta, Q_{jnd}, Q_{thres}$ and Tthresare preseted on the encoder and decoder side. So no additional syntax need to be transmitted in the bitstream. Moreover, for the transformation-off non-pixel-wise prediction, the straightforward extension of the PWQ method is designed to divide the coding block into simple and complex areas based on the above reference pixels, and reduce the pixel QP in simple areas. Qualitative results in Fig. 1 show that, the PWQ method can significantly improve the subjective quality by reducing the distortions on simple pixels, especially in the flat areas near the object edge and between the words on the screen content, and realizes more fine-grained pixel-level quantization compared with the traditional block-level quantization.
Fangdong Chen, Xiaoyang Wu 0007, Shiliang Pu
DCC2
2023 A Spatio-Temporal Decomposition Network for Compressed Video Quality Enhancement
abstract
Compressed video quality enhancement has always been a widely concerned research. However, existing methods rarely build models from the consideration of object motion diversity and feature frequency distribution. In this paper, we propose a Spatio-Temporal Decomposition Network (STDN) to reduce the compressed distortion with motion classification and frequency separation. In the temporal domain, a novel deformable convolution is designed to estimate the various motion offsets of different categories objects, and then the adjacent frame features would be accurately fused with them. In the spatial domain, a frequency decomposition module is first proposed to decompose the features of different frequencies and process them with appropriate precision. Experiments show that our method can surpass all the existing methods in both of subjective and objective aspects, and achieve the performance of state-of-the-art.
Kai Wang 0036, Fangdong Chen, Zongmiao Ye, Xiaoyang Wu 0007, Shiliang Pu
ICASSP2
2022 Two-Stage Octave Residual Network for End-to-End Image Compression
abstract
Octave Convolution (OctConv) is a generic convolutional unit that has already achieved good performances in many computer vision tasks. Recent studies also have shown the potential of applying the OctConv in end-to-end image compression. However, considering the characteristic of image compression task, current works of OctConv may limit the performance of the image compression network due to the loss of spatial information caused by the sampling operations of inter-frequency communication. Besides, the correlation between multi-frequency latents produced by OctConv is not utilized in current architectures. In this paper, to address these problems, we propose a novel Two-stage Octave Residual (ToRes) block which strips the sampling operation from OctConv to strengthen the capability of preserving useful information. Moreover, to capture the redundancy between the multi-frequency latents, a context transfer module is designed. The results show that both ToRes block and the incorporation of context transfer module help to improve the Rate-Distortion performance, and the combination of these two strategies makes our model achieve the state-of-the-art performance and outperform the latest compression standard Versatile Video Coding (VVC) in terms of both PSNR and MS-SSIM.
Fangdong Chen, Yumeng Xu
AAAI1
2021 Angular Weighted Prediction for Next-Generation Video Coding Standard
abstract
Weighted prediction plays an imperative role in the video coding methods but the traditional weighted process with fixed weight values is not satisfactory for coding of oblique edge regions of two objects. The existing methods consume much computational complexity in pixel-wise weight derivation which need to calculate the distance between each pixel position and partition line. In this paper, angular prediction is utilized to derive pixel level weight values which reuses the logic in intra prediction to simplify the complexity. To further improve the accuracy of prediction, a refinement process is introduced on the motion vectors used for weighted prediction. Experimental results show that the proposed AWP mode outperforms the existing methods, and can bring 0.9% bitrate saving in random access test and 2.0% bitrate saving in low delay B test, respectively.
Fangdong Chen, Shiliang Pu
ICME2
2020 A Spatial RNN Codec for End-to-End Image Compression
abstract
Recently, deep learning has been explored as a promising direction for image compression. Removing the spatial redundancy of the image is crucial for image compression and most learning based methods focus on removing the redundancy between adjacent pixels. Intuitively, to explore larger pixel range beyond adjacent pixel is beneficial for removing the redundancy. In this paper, we propose a fast yet effective method for end-to-end image compression by incorporating a novel spatial recurrent neural network. Block based LSTM is utilized to remove the redundant information between adjacent pixels and blocks. Besides, the proposed method is a potential efficient system that parallel computation on individual blocks is possible. Experimental results demonstrate that the proposed model outperforms state-of-the-art traditional image compression standards and learning based image compression models in terms of both PSNR and MS-SSIM metrics. It provides a 26.73% bits-reduction than High Efficiency Video Coding (HEVC), which is the current official state-of-the-art video codec.
Chaoyi Lin, Jiabao Yao, Fangdong Chen
CVPR3
2019 An Attention Residual Neural Network with Recurrent Greedy Approach as Loop Filter for Inter Frames
abstract
Recently, the deep learning neural networks (DNNs) based filters have demonstrated their advantages to remove artifacts or improve the performance in the area of image/video coding. However, the existing DNNs-based filters can handle intra coding distortion or quality enhancement approaches, thus not suitable for inter frames without the Rate Distortion Optimization (RDO) strategy in encoder side, which is not desirable for the implementation on hardware since all the existing filters need to be executed to make the optimal choice for the current content in encoder side. Also the RDO algorithm can only make the local optimal choice, which does not represent the global optimum, because of the long-term dependency between the inter frames. In this paper, we propose an in-loop filter for inter frames to completely replace all the conventional filters in the codec, including de-blocking filter (DB), bilateral filter (BF), adaptive loop filter (ALF) and SAO (Sample Adaptive Offset) without the RDO strategy in encoder side. Moreover, the greedy heuristic is adopted to produce the optimal filter weights which approximates the global optimal solution in each stage during the off-line training. In the experiments, our method reduces the average BD-rate by 2.77%, 7.01%, 8.64% for luma and both chroma components with Random Access (RA) configuration by using only one set of parameters to handle multiple distortion and decrease the consumption of memory.
Jiabao Yao, Fangdong Chen, Chaoyi Lin, Shiliang Pu
ICME3
2017 Surveillance video coding with dynamic textural background detection
abstract
Texture scenes like flickering flames, swaying tree branches, flowing water exhibit a complex stochastic motion character. It presents a great challenge to compress these dynamic texture efficiently even with the state-of-the-art video encoder. Furthermore, these contents only contain a little helpful information in surveillance analysis. In this paper, we propose an approach for compressing the dynamic textures in the surveillance video. In the proposed scheme, the dynamic texture contents are detected by the histogram of motion direction (HMD) algorithm, and then removed at the encoder, these dynamic texture contents will be restored at the decoder directly. Objective and subjective results are presented, demonstrating that the proposed approach provides about 8.7% bitrate saving with visually plausible dynamic textures in comparison with High Efficiency Video Coding (HEVC).
Fangdong Chen, Dong Liu 0002, Zhibo Chen 0001, Weiping Li 0003
ICIP2
2017 Fast encoding of surveillance videos based on HEVC
abstract
An increasing number of deployed surveillance cameras raise a huge demand for higher efficiency video coding scheme, and the emerging background reference based video coding methods with High Efficiency Video Coding (HEVC) have achieved a large increase of coding efficiency on surveillance videos. However, the high encoding complexity of HEVC causes troubles for these methods to be adopted in practice, especially in real-time coding scenarios. Among all the factors resulting in the increase of encoding complexity, mode decision and motion estimation are both critical reasons. Therefore, a fast algorithm based on a block-level background generation method is proposed to address this problem for surveillance video. With the help of fast detection of background coding units, the Merge mode is early decided, some prediction modes at specific depths that have the least probabilities are skipped using early termination during rate-distortion optimization, and unnecessary motion estimation is also avoided. Thanks to the full utilization of the static camera premise, the proposed fast algorithm achieves 77.0% reduction of encoding time, while only incurs 1.0% performance loss that is negligible. More experimental results also verify that the proposed algorithm outperforms the state-of-the-art schemes.
Fangdong Chen, Dong Liu 0002, Houqiang Li, Feng Wu 0001
VCIP1
2017 Block-Composed Background Reference for High Efficiency Video Coding
abstract
A block-composed background reference method is proposed in this paper for High Efficiency Video Coding (HEVC). For a group of picture (GoP), the first reconstructed picture is served as an initial background reference, which probably includes foreground content. In the subsequent coding, some background coding tree units (CTUs) in every picture are selected to be compressed with high quality. These reconstructed CTUs are used to update the background reference as well as replace the foreground content. Finally, a high-quality background reference is generated to better exploit the long-term temporal correlation in the video. There are three key technical contributions in the proposed coding scheme. First, the background reference is generated gradually by block updating instead of picture updating, which makes the scheme free of bit-rate burst and more suitable for real-time applications and can generate high-quality background reference even with complicated foreground. Second, we propose an approach to select background CTUs by taking both temporal and spatial smoothness into account. Third, we propose a model to decide the coding parameters of the selected background CTUs based on the overall picture activity, which essentially pursues the GoP-level optimal performance when making CTU-level decision. The proposed background reference is implemented into HEVC, and the experimental results demonstrate a significant improvement in coding efficiency. Compared with HEVC, our method can averagely save 14% bits in surveillance and conferencing sequences with negligible increase of encoding and decoding complexity. In particular, it can still averagely save 7.3% bits in HEVC general test sequences. Obviously, the proposed scheme can be applied to more general video contents.
Fangdong Chen, Houqiang Li, Li Li 0040, Dong Liu 0002, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2016 Two-stage picture padding for high efficiency video coding
abstract
With the exponential growth of digital cameras and the aid of powerful video editing softwares, videos with various resolutions become ever more popular. Therefore, there is a great demand for video coding schemes supporting arbitrary resolutions with higher efficiency. In the latest standard, only a simple padding with direct copying is adopted to meet this requirement, and the coding efficiency is still unsatisfactory. To improve the performance, a novel two-stage picture padding is proposed in this paper. Based on the adopted transform, an efficient residual padding is employed in the first stage to minimize the burdened coding bits of padded parts. To obtain a reference picture with high-quality boundary, a template matching based approach is also proposed to rearrange the padded pixels in the second stage. Experimental results reveal that on top of HEVC, our method offers the performance with 2.1%, 1.8%, and 1.5% BD-rate reduction for RA-Main, LDB-Main and LDP-Main configurations, respectively. Compared with the state-of-the-art algorithm, it still outperforms for kinds of test sequences under all configurations.
Fangdong Chen, Xiaowei Qin, Houqiang Li
PCS1
2015 Improved Rate-Distortion Optimization Algorithms for HEVC Lossless Coding
Fangdong Chen, Houqiang Li
MMM (1)1
2015 Efficient background picture coding for videos obtained from static cameras
abstract
With the exponential growth of surveillance videos, conference videos and sports videos, videos with static cameras present an unprecedented challenge for high-efficiency video coding technology. The existing schemes developed for these videos mostly encode the background as the long-term reference (LTR) to further improve the coding efficiency. However, since the bit allocation of the long-term background reference is not intensively studied, the coding efficiency is still unsatisfactory. Based on the stability analysis of the video content, an efficient background picture coding algorithm for videos obtained from static cameras, which is embedded with the basic unit level bit allocation, is proposed in this paper. Experimental results reveal that on top of the default mode in HEVC, our method offers the performance with 10.8% BD-rate reduction on average. Compared with the state-of-the-art algorithm, it still outperforms for kinds of test sequences with negligible increases of computational complexity in both encoder and decoder.
Fangdong Chen, Li Li 0040, Dong Liu 0002, Houqiang Li, Zhuoyi Lv, Haitao Yang 0001
VCIP1
2015 Image semantic quality assessment for compression of car-plate images
abstract
We explore image semantic quality assessment (ISQA) for compression of images that are utilized for automatic image analyses, such as recognition and detection, rather than for human viewing. For such analyses purposes, we argue that the quality of compressed images should be evaluated from its preserved semantic-related features, instead of its pixel-wise fidelity (e.g. PSNR) or visual quality (e.g. SSIM). In this paper, we make an empirical study of an ISQA approach based on SIFT features extracted from both original and compressed car-plate images, and we formulate an optimization problem to find the operating point of an image compression system for car-plate recognition. Experimental results show that our proposed ISQA measure is significantly better than PSNR and SSIM in predicting the recognizability of compressed car-plate images. Accordingly, using our ISQA measure during compression leads to more than 50% bit-rate saving compared to using PSNR or SSIM.
Dong Liu 0002, Fangdong Chen
VCIP3
2014 Hybrid transform for HEVC-based lossless coding
abstract
The High Efficiency Video Coding (HEVC) with the transform bypass mode is simple but inefficient for lossless coding. For this reason, we propose a novel transform to further eliminate the redundancy between residues of different blocks in intra prediction. Dependent on intra prediction modes, the proposed transform is adaptable to exploit correlations of residues formed by different modes. In order to accurately obtain parameters of the transform matrix, an approach similar to the Wiener filtering method is adopted. Experimental results show that on top of the lossless coding mode in HEVC, our method offers the performance with a 7.4% bit-rate reduction on average for All Intra Main configuration. Compared with other representative algorithms, our proposal still shows an improvement in the compression ratio, without substantial increases of computational complexity in the encoder or decoder.
Fangdong Chen, Jinlei Zhang, Houqiang Li
ISCAS1
2013 A Video Communication System Based on Spatial Rewriting and ROI Rewriting
Fangdong Chen, Bin Li 0012, Houqiang Li
MMM (2)2