Jie Chen 0006

dblp:92/6289-6 · DBLP profile ↗
← Back
22ranked-venue papers in the field
2as first author
17since 2021 · last 2026
ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 21 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Sparse2Dense: A Keypoint-Driven Generative Framework for Human Video Compression and Vertex Prediction
abstract
For bandwidth-constrained multimedia applications, simultaneously achieving ultra-low bitrate human video compression and accurate vertex prediction remains a critical challenge, as it demands the harmonization of dynamic motion modeling, detailed appearance synthesis, and geometric consistency. To address this challenge, we propose Sparse2Dense, a keypoint-driven generative framework that leverages extremely sparse 3D keypoints as compact transmitted symbols to enable ultra-low bitrate human video compression and precise human vertex prediction. The key innovation is the multi-task learning-based and keypointaware deep generative model, which could encode complex human motion via compact 3D keypoints and leverage these sparse keypoints to estimate dense motion for video synthesis with temporal coherence and realistic textures. Additionally, a vertex predictor is integrated to learn human vertex geometry through joint optimization with video generation, ensuring alignment between visual content and geometric structure. Extensive experiments demonstrate that the proposed Sparse2Dense framework achieves competitive compression performance for human video over traditional/generative video codecs, whilst enabling precise human vertex prediction for downstream geometry applications. As such, Sparse2Dense is expected to facilitate bandwidth-efficient human-centric media transmission, such as realtime motion analysis, virtual human animation, and immersive entertainment.
Ru-Ling Liao, Yan Ye 0003, Jie Chen 0006, Shanzhi Yin, Xinrui Ju, Shiqi Wang 0001, Yibo Fan
DCC4
2026 A Dual Merge Mode for Video Coding
abstract
In the latest international video coding standard, Versatile Video Coding, merge mode utilizes spatial or temporal adjacent motion information as motion vector predictors to reduce the signaling overhead. However, only one merge candidate's motion information is selected from previously coded blocks for motion compensation and prediction. To improve the prediction efficiency, a dual merge mode which inherits two candidates' motion information from previously coded blocks is proposed. In the proposed method, motion compensation is executed utilizing the motion information from both merge candidates to generate two predictions, and the two predictions are ultimately combined with equal weight to derive the final predicted results. Both of the merge candidates can be uni-prediction or bi-prediction, requiring up to four motion compensations for the proposed dual merge mode. Experimental results show that the proposed dual merge mode achieves an average coding gain of 0.03 % over ECM-14.0 with negligible complexity increase, and the extended ablation studies verify the effectiveness of the proposed method.
Ru-Ling Liao, Jie Chen 0006, Yan Ye 0003
DCC3
2026 Affine with Symmetric Motion Vector Difference Mode
abstract
The Versatile Video Coding (VVC) standard incorporates affine motion compensation to represent complex motions including rotation and zoom. However, affine advanced motion vector prediction (AMVP) mode incurs significant signaling overhead, particularly in biprediction where motion vector differences (MVDs) must be signaled for both reference pictures. This paper extends the Symmetric Motion Vector Difference (SMVD) [1] method to affine motion fields. The proposed approach signals MVDs only from one reference picture, deriving the other according to POC distances.
Ru-Ling Liao, Jie Chen 0006, Yan Ye 0003
DCC3
2026 Improvements on Template Matching Merge Mode for the Next Generation of AVS
abstract
In search for more efficient video compression capability than versatile video coding (VVC), joint video experts team (JVET) launched exploration work on new video coding technologies beyond VVC in 2021. In 2023, audio video coding standard (AVS) workgroup started to explore the new video coding technologies beyond AVS3 and released a software platform named exploration video model (EVM). In EVM, the candidates of merge mode are reordered according to the template matching (TM) cost to reduce the candidate index signaling overhead. After reordering, the TM-based refinement is applied to the candidate to improve the accuracy of the motion derived in merge mode. However, it is found that the bi-prediction candidates are usually more effective than uni-prediction candidates despite the TM cost in non-low-delay pictures, while the current reordering method only relies on the TM cost. And when performing TM-based refinement, the motions of the two reference picture lists (RPLs) are processed independently without joint optimization. To solve these issues, an adaptive motion candidate reordering method and an improved TM motion refinement method which is based on the joint motion search are proposed in this paper. In the proposed candidate reordering method, the TM cost of uni-prediction candidate is weighted by a factor greater than 1 to prioritize the bi-prediction candidates. And in the proposed motion refinement method, the reference template of the bi-predicted block is obtained by averaging the templates of two predicted blocks, such that the motions of the two RPLs are jointly optimized. The proposed methods were implemented on top of EVM-0.9, it is reported that$\{-0.18 \%(\mathrm{Y}),-0.43 \%(\mathrm{U}),-0.08 \%(~\mathrm{V})\}$and$\{-0.03 \%(\mathrm{Y}),- 0.06 \%(\mathrm{U}), 0.18 \%(~\mathrm{V})\}$BD-rate reductions are achieved under random access and low-delay B configurations, respectively, with negligible runtime increase. Due to the good trade-off, the proposed methods have been adopted to EVM platform.
Yucheng Zhong, Jiabao Zhu, Wanglin Lai, LiCong Ma, Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003
DCC6
2026 Towards Efficient 3D Gaussian Human Avatar Compression: A Prior-Guided Framework
abstract
This paper proposes an efficient 3D avatar coding framework that leverages compact human priors and canonical-to-target transformation to enable high-quality 3D human avatar video compression at ultra-low bit rates. The framework begins by training a canonical Gaussian avatar using articulated splatting in a network-free manner, which serves as the foundation for avatar appearance modeling. Simultaneously, a human-prior template is employed to capture temporal body movements through compact parametric representations. This decomposition of appearance and temporal evolution minimizes redundancy, enabling efficient compression: the canonical avatar is shared across the sequence, requiring compression only once, while the temporal parameters, consisting of just 94 parameters per frame, are transmitted with minimal bit-rate. For each frame, the target human avatar is generated by deforming canonical avatar via Linear Blend Skinning transformation, facilitating temporalcoherent video reconstruction and novel view synthesis. Experimental results demonstrate that the proposed method significantly outperforms conventional 2D/3D codecs and existing learnable dynamic 3D Gaussian splatting compression method in terms of rate-distortion performance on mainstream multi-view human video datasets, paving the way for seamless immersive multimedia experiences in meta-verse applications.
Shanzhi Yin, Xinju Wu, Ru-Ling Liao, Jie Chen 0006, Shiqi Wang 0001, Yan Ye 0003
DCC5
2026 Adaptive Enhanced Affine Inter Mode for the Next Generation of AVS Standard
abstract
To meet the growing demand for advanced video coding, the Joint Video Experts Team (JVET) launched the Enhanced Compression Model (ECM) in April 2021 to explore technologies beyond Versatile Video Coding (VVC). Following this trend, the Audio Video Coding Standard (AVS) Working Group began exploring video coding technologies beyond AVS3 in 2023 and released the Exploration Video Model (EVM) as a development platform. In AVS3, the motion vector prediction (MVP) candidate list for the affine inter mode contains only one 4 -parameter spatial affine candidate, which severely limits its prediction accuracy. Furthermore, the MVD coding approach does not take advantage of the correlation between control points, which introduces redundancy in bitstream representation. To address these issues, this paper proposes an adaptive enhanced method. First, the number of MVP candidates for affine inter mode is extended from a single candidate to multiple ones, with an additional independent 6 -parameter candidate list. Inherited, constructed, and historical motion information are integrated to provide a richer set of prediction candidates. Second, a new adaptive coding method for MVD is designed to reduce bit overhead. Experimental results show that under the random access (RA) and low-delay B-picture (LDB) configurations, the proposed method achieves overall Bjøntegaard Delta Rate (BD-rate) reductions of$\{0.33 \%(\mathrm{Y}), 0.04 \%(\mathrm{U}), 0.34 \%(~\mathrm{V})\}$and$\{0.26 \%(\mathrm{Y}),-0.05 \% (\mathrm{U}), 0.40 \%(~\mathrm{V})\}$, respectively. Due to the excellent performance and acceptable computational complexity of specific components of the proposed scheme, these components have been integrated into the EVM platform as promising coding tools for the next generation AVS standard.
Yucheng Zhong, Jiabao Zhu, Wanglin Lai, LiCong Ma, Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003
DCC6
2026 Angular Weighted Prediction Mode Improvements Beyond AVS3 Standard
abstract
Audio video coding standard (AVS) workgroup started to explore the latest video coding technologies beyond AVS3 standard and released a software platform named exploration video model (EVM) in March 2023. In EVM-0.7, the construction of the motion vector (MV) candidate list for angle weighted prediction (AWP) only considers temporal and spatial neighboring motion information, ignoring non-adjacent candidates. And those motion candidates in the list are directly borrowed from previously coded blocks and thus may not match well with the current coding block. In this paper, we first propose an improved AWP MV candidate list construction method that incorporates more non-adjacent motion information. Second, we introduce an angle-adaptive refinement method to refine the motion candidate in the MV candidate list of AWP. The proposed method was implemented on top of EVM-0.7. And the experimental results show that it overall achieves$\{0.11 \%(\mathrm{Y}), 0.16 \%(\mathrm{U}),\ 0.14 \%(~\mathrm{V})$\} and$\{0.13 \%(\mathrm{Y}), 0.02 \%(\mathrm{U}), 0.13 \%(~\mathrm{V})\}$BD-rate gain on random access (RA) and low delay B (LDB) configurations, respectively, by applying the improved MV candidate list construction method, and$\{0.31 \%(\mathrm{Y}), 0.32 \%(\mathrm{U}),\ 0.43 \%(~\mathrm{V}))$and$\{0.27 \%(\mathrm{Y}), 0.05 \%(\mathrm{U}), 0.37 \%(~\mathrm{V})\}$BD-rate gain on RA and LDB configurations, respectively, by applying both the proposed MV candidate list construction method and the angle adaptive refinement method. Due to the attractive trade-off between performance and complexity, the improved MV candidate list construction was adopted into EVM software as the potential coding tool for the next generation of AVS standard.
Jiabao Zhu, Wanglin Lai, Yucheng Zhong, LiCong Ma, Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003
DCC6
2025 Regression-Based Geometric Partitioning Mode Coding
abstract
Geometric Partitioning Mode (GPM) is an effective coding tool for inter prediction that splits a block into two partitions and blends their predictions. This paper presents a new coding mode, Regression-based Geometric Partitioning Mode (RGPM), which derives a sample-based blending for bi-predictions using a reconstructed template. The RGPM can enhance flexibility in splitting and blending methods compared to GPM. Moreover, two extensions of RGPM scheme are investigated: 1) extending RGPM with template matching (TM) and merge with motion vector difference (MMVD) methods; 2) extending RGPM principle to Spatial Geometric Partitioning Mode (SGPM) for intra prediction. Experimental results show that RGPM with extensions provide 0.12%, 0.24% and 0.23% average luma BD-rate savings on top of enhanced compression model (ECM) in all intra, random access and low delay configurations, respectively. The proposed RGPM is currently adopted in ECM and its two extensions are under study in exploration experiments for future ECM developments.
Philippe Bordes, Kevin Reuze, Franck Galpin, Ke Jia, Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003
DCC6
2025 Template Matching Based Motion Refinement on Subblock Merge Mode
abstract
The subblock merge mode, in which the current coding block is split into multiple subblocks for motion compensation but still inherits the motion at the coding block level, improves the accuracy of the inter prediction and reduces the signaling overhead of motion information at the same time. And thus, it was adopted into versatile video coding (VVC) due to its high efficiency and continually improved in the enhanced compression model (ECM). However, the motion used in subblock merge mode was inherited from the previously coded blocks and may not match well with the current coding block. Thus, to improve the accuracy of the motion for the subblock merge mode, it is proposed to apply template matching (TM) based motion refinement. For subblock temporal motion vector predictor (SbTMVP) candidates, it is proposed to refine both the motion displacement and subblock motion vectors (MVs) based on TM; for affine motion candidates, it is proposed to refine the affine model, including base MV and non-translation parameters, based on TM. The proposed method was implemented on top of ECM, and the experiment results show that by applying the proposed method, it achieves {−0.23%(Y), −0.23%(U), −0.20%(V)} and {−0.09%(Y), −0.36%(U), −0.01%(V)} BD-rate reduction on random access (RA) and low delay B (LDB) configurations, respectively. Due to the good trade-off between performance and complexity, the proposed method was adopted into ECM.
Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003, Lei Zhao 0032, Kai Zhang 0007, Li Zhang 0136
DCC1
2025 Beyond GFVC: A Progressive Face Video Compression Framework with Adaptive Visual Tokens
abstract
Recently, deep generative models have greatly advanced the progress of face video coding towards promising rate-distortion performance and diverse application functionalities. Beyond traditional hybrid video coding paradigms, Generative Face Video Compression (GFVC) relying on the strong capabilities of deep generative models and the philosophy of early Model-Based Coding (MBC) can facilitate the compact representation and realistic reconstruction of visual face signal, thus achieving ultra-low bitrate face video communication. However, these GFVC algorithms are sometimes faced with unstable reconstruction quality and limited bitrate ranges. To address these problems, this paper proposes a novel Progressive Face Video Compression framework, namely PFVC, that utilizes adaptive visual tokens to realize exceptional trade-offs between reconstruction robustness and bandwidth intelligence. In particular, the encoder of the proposed PFVC projects the high-dimensional face signal into adaptive visual tokens in a progressive manner, whilst the decoder can further reconstruct these adaptive visual tokens for motion estimation and signal synthesis with different granularity levels. Experimental results demonstrate that the proposed PFVC framework can achieve better coding flexibility and superior rate-distortion performance in comparison with the latest Versatile Video Coding (VVC) codec and the state-of-the-art GFVC algorithms. The project page can be found at https://github.com/Berlin0610/PFVC.
Shanzhi Yin, Jie Chen 0006, Ru-Ling Liao, Lingyu Zhu 0006, Shiqi Wang 0001, Yan Ye 0003
DCC4
2024 Generative Face Video Coding Techniques and Standardization Efforts: A Review
abstract
Generative Face Video Coding (GFVC) techniques can exploit the compact representation of facial priors and the strong inference capability of deep generative models, achieving high-quality face video communication in ultra-low bandwidth scenarios. This paper conducts a comprehensive survey on the recent advances of the GFVC techniques and standardization efforts, which could be applicable to ultra low bitrate communication, user-specified animation/filtering and metaverse-related functionalities. In particular, we generalize GFVC systems within one coding framework and summarize different GFVC algorithms with their corresponding visual representations. Moreover, we review the GFVC standardization activities that are specified with supplemental enhancement information messages. Finally, we discuss fundamental challenges and broad applications on GFVC techniques and their standardization potentials, as well as envision their future trends. The project page can be found at https://github.com/Berlin0610/Awesome-Generative-Face-Video-Coding.
Jie Chen 0006, Shiqi Wang 0001, Yan Ye 0003
DCC2
2024 An Improvement to Subblock-based Temporal Motion Vector Prediction Beyond VVC
abstract
Temporal motion vector predictor (TMVP) is a well-known coding technology that has been included in recent video compression standards. The basic concept of TMVP is to utilize the motion continuity between temporal pictures. It obtains motion vector predictor from temporal collocated block by performing temporal motion vector scaling to align the reference pictures of the collocated block to that of to-be-coded block. To exploit the coding performance of TMVP, subblock-based temporal motion vector prediction (SbTMVP) is introduced in Versatile Video Coding (VVC) and is further improved in Enhanced Compression Model (ECM), the software platform for exploring compression capability beyond that of VVC. SbTMVP shares a similar concept to that of TMVP, but obtains motion vectors at the subblock level with a motion displacement which is derived from neighboring blocks. In both VVC and ECM design, the SbTMVP is treated as one of subblock-based merge candidates. In this paper, it is proposed to extend SbTMVP to advanced motion vector predictor (AMVP) for coding motion explicitly. To allow more flexibility, the motion displacement is directly signaled in the bitstream in the proposed mode. Experimental results show that on top of ECM, the proposed mode provides {0.11% (Y), 0.04% (U), 0.05% (V)} and {0.36% (Y), 0.25% (U), 0.33% (V)} BD-rate reduction in random access and low delay B configurations, respectively. The proposed mode is currently being studied in exploration experiments due to its promising compression efficiency gain.
Ru-Ling Liao, Jie Chen 0006, Yan Ye 0003
DCC2
2024 Intra Template Matching Prediction with Fusion Techniques
abstract
Intra template matching prediction (Intra TMP) is a promising intra prediction tool which generates the prediction block by copying from a reconstructed block of the current frame. The position of the reconstructed block is derived by template matching at both encoder and decoder. Intra TMP has been adopted in enhanced compression model (ECM) for both screen content and natural content due to its outstanding trade-off between coding efficiency and complexity. This paper describes an advanced intra TMP algorithm with fusion techniques to further improve the coding efficiency. The proposed intra TMP fusion scheme includes the following aspects: 1) extended template matching search range, 2) improved search procedure with multiple candidates, 3) adaptive fusion method with template updating, and 4) support fractional-pel precision in intra TMP. Experimental results show that the proposed intra TMP fusion scheme provides 0.76% average luma Bjøntegaard delta rate (BD-rate) reduction with negligible runtime increase over ECM-8.0 in all-intra configuration.
Fangjun Pu, Taoran Lu, Peng Yin 0002, Sean McCarthy, Jeeva Raj Arumugam, Ashwin Natesan, Vaibhav Valvaiker, Jay N. Shingala, Ru-Ling Liao, Jie Chen 0006, Yan Ye 0003, Lai Zhang, Haoping Yu
DCC11
2023 Decoder-side Affine Model Refinement for Video Coding beyond VVC
Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003
DCC1
2023 Gradient Linear Model for Chroma Intra Prediction
abstract
In Versatile Video Coding (VVC), Cross-component Linear Model (CCLM) predicts chroma samples by assuming a linear relationship between luma and chroma components. In performing CCLM for video in YUV 4:2:0 chroma format, collocated luma samples are firstly downsampled by a low-pass filter to match luma resolution with chroma, and one linear model of luma-chroma sample pairs is applied on the reconstructed luma samples to generate the predicted chroma samples. However, the low-pass downsampling procedure ignores relative spatial variations among luma samples in proximity, such as edge and gradient information. To solve this issue, a new coding technique, namely gradient linear model (GLM), is proposed for further compression efficiency exploration beyond VVC. Instead of using a low-pass filter in CCLM, the GLM utilizes high-pass gradient filters to generate the downsampled luma values. In this paper, two GLM schemes are provided with different trade-offs between coding gain and complexity, including: 1) a 2-parameter scheme that shares the CCLM module framework but replaces the downsampling filter with high-pass gradient filters; 2) a 3-parameter scheme that further combines the luma gradients with the low-pass downsampled luma values. Based on the enhanced compression model (ECM-5.0) software from the joint video experts team (JVET), simulation results show that the 2-parameter GLM achieves average Bjontegaard delta-rate (BD-rate) savings of {1.01%, 1.66%, 1.81%} and {0.69%, 0.95%, 1.12%} for {Y, U, V} components under the All Intra and Random Access configurations, respectively, and the 3-parameter GLM provides {1.28%, 3.23%, 3.28%} and {0.92%, 2.19%, 2.26%} BD-rate savings for {Y, U, V} components under the All Intra and Random Access configurations, respectively. Both of the proposed GLM schemes have been adopted to the ECM software platform.
Che-Wei Kuo, Xiaoyu Xiu, Hong-Jheng Jhu, Xianglin Wang, Yan Ye 0003, Jie Chen 0006, Ru-Ling Liao
DCC8
2023 Decoder-side Chroma Intra Mode Derivation in Video Coding
abstract
Decoder-side intra mode derivation (DIMD) is a promising coding tool in the enhanced compression model (ECM) developed by the joint video experts team (JVET). In DIMD, the intra prediction mode of a luma block is derived based on the gradient information of the adjacent luma samples at both encoder and decoder, rather than being explicitly signaled in the bitstream. Inspired by DIMD, a decoder-side chroma intra mode derivation (DCIMD) method is proposed in this paper to improve the coding efficiency of chroma intra prediction. In the proposed DCIMD, the gradient information of both adjacent luma samples and chroma samples is utilized to derive an angular intra mode to predict the current chroma block. There are two advantages to this proposed scheme: 1). the combination of luma information and chroma information can ensure that the derived angular intra mode is closely matched with the texture features of the current chroma block; 2). the number of chroma intra prediction modes can be effectively extended by introducing DCIMD as a new mode without incurring additional expensive bit overhead. Moreover, to further improve the coding efficiency of DCIMD, a chroma intra mode fusion (CIMF) method is proposed. In CIMF, predictions from the DCIMD chroma prediction mode and a cross-component correlation-based chroma prediction mode can be fused together with adaptive weights. Simulation results show that on top of ECM-4.0 the BD-rate savings of 0.07%, 1.17%, and 1.02% on average for Y, Cb, and Cr components, respectively, are achieved in all intra configuration with the proposed two methods. Both the proposed DCIMD and CIMF methods have been adopted into the ECM-5.0 by JVET.
Ru-Ling Liao, Jie Chen 0006, Yan Ye 0003
DCC3
2023 An Improvement to Merge Mode in ECM With Template Matching
abstract
In the development of video coding standard, decoder-side motion derivation technology has been proven to provide promising coding efficiency. With this type of technology, the motion information is derived at the decoder instead of being signaled in the bitstream by the encoder, and thus, the number of bits to be sent are reduced. A typical decoder-side derivation technology is template matching, which refines the motion by finding the closest match between neighboring reconstructed samples and corresponding reference samples in the reference pictures. In this paper, the template matching method is extended to temporal motion vector predictor, bi-prediction with CU-level weight and geometric partition modes in order to fully utilize its benefit. Specifically, template matching is used to determine the prediction direction and reference picture of temporal motion vector predictors, and to decide the bi-predicted weight of bi-prediction merge blocks. In addition, the motion of two geometric partitions are individually refined by the template matching mechanism. Simulation results show that on top of Enhanced Compression Model (ECM), which is the software platform for exploring activities beyond versatile video coding established by the joint video experts team (JVET), the three proposed methods achieve 0.20% luma BD-rate savings in random access configuration and 0.33% luma BD-rate saving in low delay B configuration with negligible encoding and decoding runtime impact. It is worth noting that all three proposed methods have been adopted to the ECM software platform.
Ru-Ling Liao, Yan Ye 0003, Jie Chen 0006
DCC3
2020 Advanced Geometric-Based Inter Prediction for Versatile Video Coding
abstract
Block-based partitioning is one of the fundamental techniques in video coding. Geometric-based block partitioning is a well-studied method to enable better spatial adaptation to the signal properties. This paper introduces the most recent proposal of advanced geometric-based inter prediction (GIP) made to the state-of-the-art are video coding standard - Versatile Video Coding (VVC). Implemented in the latest test model VTM-6.0 to generalize the existing triangle partition mode (TPM) and evaluated with the Joint Video Experts Team (JVET) Common Test Conditions (CTC) sequences, the proposed advanced GIP scheme provides luma BD-rate reduction of 0.56% for random access (RA) and 1.37% for low-delay (LB) test cases with 2% encoder runtime increase and negligible decoder runtime increase. Furthermore, BD-rate reductions up to 2.92% and 3.49% for RA and LB test cases can be achieved in the absence of multiple related VVC inter prediction tools.
Han Gao 0001, Ru-Ling Liao, Kevin Reuze, Semih Esenlik, Elena Alshina, Yan Ye 0003, Jie Chen 0006, Jiancong Luo, Chun-Chi Chen, Han Huang 0001, Wei-Jung Chien, Vadim Seregin, Marta Karczewicz
DCC7
2020 Luma Mapping with Chroma Scaling in Versatile Video Coding
abstract
This paper describes a new video coding tool in the Versatile Video Coding standard (VVC) named as luma mapping with chroma scaling (LMCS). Experimental compression performance results for LMCS and non-normative examples for deriving LMCS parameter values are also provided. LMCS has two main components: 1) a process for mapping input luma code values to a new set of code values for use inside the coding loop; and 2) a luma-dependent process for scaling chroma residue values. The first process, luma mapping, aims at improving the coding efficiency for standard and high dynamic range video signals by making better use of the range of luma code values allowed at a specified bit depth. The second process, chroma scaling, manages relative compression efficiency for the luma and chroma components of the video signal. The luma mapping process of LMCS is applied at the pixel sample level, and is implemented using a piecewise linear model. The chroma scaling process is applied at the chroma block level, and is implemented using a scaling factor derived from reconstructed neighboring luma samples of the chroma block.
Taoran Lu, Fangjun Pu, Peng Yin 0002, Sean McCarthy, Walt Husak, Tao Chen 0044, Edouard François, Christophe Chevance, Franck Hiron, Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003, Jiancong Luo
DCC10
2019 IDeRs: Iterative dehazing method for single remote sensing image
Long Xu 0001, Dong Zhao 0016, Yihua Yan, Sam Kwong, Jie Chen 0006, Ling-Yu Duan
Inf. Sci.5
2017 Compact Deep Invariant Descriptors for Video Retrieval
abstract
With emerging demand for large-scale video analysis, the Motion Picture Experts Group (MPEG) initiated the Compact Descriptor for Video Analysis (CDVA) standardization in 2014. In this work, we develop novel deep-learning features and incorporate them into the well-established CDVA evaluation framework to study its effectiveness in video analysis. In particular, we propose a Nested Invariance Pooling (NIP) method to obtain compact and robust Convolutional Neural Network (CNNs) descriptors. The CNNs descriptors are generated by applying three different pooling operations to the feature maps of CNNs in a nested way towards rotation and scale invariant feature representation. In particular, the rational, advantages and performance on the combination of CNNs and handcrafted descriptors are provided to better investigate the complementary effects of deep learnt and handcrafted features. Extensive experimental results show that the proposed CNNs descriptors outperform both state-of-the-art CNNs descriptors and canonical handcrafted descriptors adopted in CDVA Experimental Model (CXM) with significant mAP gains of 11.3% and 4.7%, respectively. Moreover, the combination of NIP derived deep invariant descriptors and handcrafted descriptors not only fulfills the lowest bitrate budget of CDVA, but also significantly advances the performance of CDVA core techniques.
Yihang Lou, Jie Lin 0001, Shiqi Wang 0001, Jie Chen 0006, Vijay Chandrasekhar 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
DCC5
2015 Optimizing Binary Fisher Codes for Visual Search
abstract
Fisher vectors (FV) aggregated from local invariant features (e.g., SIFT) is one of the state-of-the-art descriptors for visual search, due to high discriminability but small visual vocabulary. Nevertheless, a high-dimensional FV needs to be compressed into a compact descriptor for light storage and high matching eficiency. In this paper, we formulate the FV compression as a resource-constrained optimization problem. Our goal is to maximize search performance subject to the constraints of descriptor compactness, compression complexity in terms of memory usage and time cost. Accordingly, we present a selective binary Fisher codes (SBFC) to compress the raw FV. Firstly, to fulfill the constraint of compression complexity, we binarize the FV by a sign function, Secondly, we propose to select discriminative bits from the binarized FV (BFC) to maximize search performance, subject to the constraint of descriptor compactness. Extensive experiments over MPEG Compact Descriptor for Visual Search (CDVS) benchmark datasets have shown that S-BFC significantly improves search performance at a smaller descriptor size as well as much lower complexity, compared with the state-of-the-art FV compression algorithms like Hashing and Product Quantziation (PQ). A simplified version of SBFC, SBFC LS has been adopted by the MPEG CDVS standard. In the CDVS evaluation framework, SBFC LS has achieved promising performance mean Average Precision (mAP) 83% on average at much lower memory cost of 40KB.
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001
DCC4