Jielian Lin

dblp:304/1152 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-7957-2858ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Game-theory-based complexity allocation for 360-degree video coding
Jielian Lin, Kaiying Xing
Signal Process.1
2025 RDVC: Efficient Deep Video Compression With Regulable Rate and Complexity Optimization
abstract
Deep video coding has paved a way to break through the performance bottleneck of reigning hybrid video coding. However, unlike hybrid video codecs, existing deep video codecs cannot offer both flexible rates and regulable complexities within one single codec, which limits their applications. In this article, we propose a Regulable Deep Video Codec (RDVC) to address the above issue. First, we propose an Adaptive Feature Compression (AFC) network that generates variable rates while ensuring Rate-Distortion (RD) performance. The network introduces a two-stage coarse-to-fine rate adjustment that can be controlled by a user-specified rate level. Second, we propose a Spatio-Temporal Feature Propagation (STFP) mechanism to provide high-quality reference information for AFC process. Third, we also utilize slimmable convolutional components in our framework to adjust decoding complexity constrained by user configuration. Experimental results demonstrate that RDVC can adjust the codec structure flexibly according to different user configurations while maintaining advanced performance. On average, it reduces the bit-per-pixel (bpp) by 9.35%$/$58.12% while maintaining the same PSNR/MS-SSIM as the reference software VTM-13.2.
Xiaojie Wei 0001, Jielian Lin, Wei Gao 0003, Tiesong Zhao
IEEE Trans. Multim.2
2024 Efficient inter partitioning of versatile video coding based on supervised contrastive learning
Jielian Lin, Zhichen Zhang
Knowl. Based Syst.1
2024 Video Compression Artifacts Removal With Spatial-Temporal Attention-Guided Enhancement
abstract
Recently, many compression algorithms are applied to decrease the cost of video storage and transmission. This will introduce undesirable artifacts, which severely degrade visual quality. Therefore, Video Compression Artifacts Removal (VCAR) aims at reconstructing a high-quality video from its corrupted version of compression. Generally, this task is considered as a vision-related instead of media-related problem. In vision-related research, the visual quality has been significantly improved while the computational complexity and bitrate issues are less considered. In this work, we review the performance constraints of video coding and transfer to evaluate the VCAR outputs. Based on the analyses, we propose a Spatial-Temporal Attention-Guided Enhancement Network (STAGE-Net). First, we employ dynamic filter processing, instead of conventional optical flow method, to reduce the computational cost of VCAR. Second, we introduce self-attention mechanism to design Sequential Residual Attention Blocks (SRABs) to improve visual quality of enhanced video frames with bitrate constraints. Both quantitative and qualitative experimental results have demonstrated the superiority of our proposed method, which achieves high visual qualities and low computational costs.
Nanfeng Jiang, Jielian Lin, Tiesong Zhao, Chia-Wen Lin
IEEE Trans. Multim.3
2023 DeepSVC: Deep Scalable Video Coding for Both Machine and Human Vision
abstract
Nowadays, end-to-end video coding for both machine and human vision has become an emerging research topic. In complicated systems such as large-scale internet of video things (IoVT), feature streams and video streams can be separately encoded and delivered for machine judgement and human viewing. In this paper, we propose a deep scalable video codec (DeepSVC) to support three-layer scalability from machine to human vision. First, we design a semantic layer that encodes semantic features extracted from the captured video for machine analysis. This layer employs a conditional semantic compression (CSC) method to remove redundancies between semantic features. Second, we design a structure layer that can be combined with semantic layer to predict the captured video at a low quality. This layer effectively estimates video frames based on semantic layer with an interlayer frame prediction (IFP) network. Third, we design a texture layer that can be combined with the above two layers to reconstruct high-quality video signals. This layer also takes advantage of the IFP network to improve its coding efficiency. In large-scale IoVT systems, DeepSVC can deliver semantic layer for regular use and transmit the other layers on demand. Experimental results indicate that the proposed DeepSVC outperforms popular codecs for machine and human vision. Compared with scalable extension of H.265/HEVC (SHVC), the proposed DeepSVC reduces average bit-per-pixel (bpp) by 25.51%/27.63%/59.87% at the same mAP/PSNR/MS-SSIM. Sourcecode is available at: https://github.com/LHB116/DeepSVC.
Zhichen Zhang, Jielian Lin, Xu Wang 0006, Tiesong Zhao
ACM Multimedia4
2023 ELFIC: A Learning-based Flexible Image Codec with Rate-Distortion-Complexity Optimization
abstract
Learning-based image coding has attracted increasing attentions for its higher compression efficiency than reigning image codecs. However, most existing learning-based codecs do not support variable rates with a single encoder; their decoders are also of fixed, high computational complexity. In this paper, we propose an End-to-end, Learning-based and Flexible Image Codec (ELFIC) that supports variable rate and flexible decoding complexity. First, we propose a general image codec with Nonlinear Feature Fusion Transform (NFFT) as nonlinear transforms to improve its Rate-Distortion (RD) performance. Second, we propose an Instance-aware Decoding Complexity Allocation (IDCA) approach, which exploits image contents for a tradeoff between reconstruction quality and computational complexity in the decoding process. Third, we propose an RD-Complexity (RDC) optimization algorithm, which maximizes the image quality under given rate and complexity constraints for the whole framework. Experimental results show that ELFIC achie-ves variable rate, flexible decoding complexity with the state-of-the-art RD performance. It also supports a more efficient decoding process by focusing on image contents. Source codes are available at https://github.com/Zhichen-Zhang/ELFIC-Image-Compression.
Zhichen Zhang, Jielian Lin, Xu Wang 0006, Tiesong Zhao
ACM Multimedia4
2023 λ-Domain VVC Rate Control Based on Nash Equilibrium
abstract
With a significant Rate-Distortion (RD) improvement than H.265/HEVC, Versatile Video Coding (VVC) has set a new milestone in lossy video compression. It also incorporates the emerging$\lambda $-domain rate control technique, aiming at a higher visual quality under a fixed bit constraint. However, the challenge remains how to efficiently allocate bits to all frames and Coding Tree Units (CTUs). In this paper, we propose an effective solution by formulating the above task as a Nash equilibrium problem, where all CTUs are treated as players that bargains with each other. By introducing$\lambda $-domain RD models, a constrained optimization is derived with no closed-form solution. We then propose a two-step strategy to address this issue: a Newton method to iteratively calculate an intermediate variable, and a final solution of Nash equilibrium to obtain an approximately optimal$\lambda $. Finally, we utilize the derived$\lambda $to perform an effective CTU-level bit allocation, which is the very first attempt to introduce Nash equilibrium in$\lambda $-domain rate control. Experimental results with Common Test Conditions (CTC) demonstrate the effectiveness and superiority of our method, which outperforms the state-of-the-art CTU-level rate allocation algorithms for VVC.
Jielian Lin, Aiping Huang, Tiesong Zhao, Xu Wang 0006, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.1
2022 Learning-Based Multi-Stage Intra Partition for Versatile Video Coding
abstract
The latest standard, Versatile Video Coding (VVC), doubles the coding efficiency over the previous generation standard. However, better performance is at the cost of a sharp increase in coding complexity. In order to reduce the complexity of VVC intra coding, this paper proposes a multi-stage block partition decision framework based on deep learning. First, we propose a three-stage redundant modes removal framework that decreases the number of modes checked in the brute-force process. Then, we build a lightweight CNN to complete the classification task of each stage. To reduce the burden of CNN and adapt to different Coding Unit (CU) sizes, we pre-process the luminance component of CU and use the results as input of the network. Finally, the multi-threshold adjusting scheme is proposed for trading off complexity reduction with the bit-rate increase. The experimental results shows our method can reduce the encoding time ranging from 16.93% to 69.40% with the bit-rate increase ranging from 0.31% to 3.59%. Such results demonstrate that our method has superior performance with a wide range of adjustments compared with other state-of-the-art methods.
Hongji Zeng, Tiesong Zhao, Weize Feng, Jielian Lin, Xu Wang 0006
MMSP5
2022 DesnowFormer: an effective transformer-based image desnowing network
abstract
Single image desnowing is an important and challenge task for lots of computer vision applications, such as visual tracking and video surveillance. Although existing deep learning-based methods have achieved promising results, most of them rely on the local deep features and neglect global relationship information between the local regions. Therefore, inevitably leading to over-smooth or detail loss results. To solve this issue, we design a UNet-based end-to-end architecture for image desnowing. Specially, to better characterize global information and preserve image detail, we combine Window-based Self-Attention (WSA) transformer block with Residue Spatial Attention (RSA) to build basic unit of our network. Besides, to protect the structure of the image effectively, we also introduce a Residue Channel (RC) loss to guide high-quality image restoration. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed model achieves new state-of-the-art results.
Nanfeng Jiang, Junhong Lin 0001, Jielian Lin, Tiesong Zhao
VCIP4
2022 SSIM-Variation-Based Complexity Optimization for Versatile Video Coding
abstract
Hitherto, Versatile Video Coding (VVC) has a more magnificent overall performance than High Efficiency Video Coding (HEVC). The Quadtree with Nested Multi-Type Tree (QTMT) coding block structure can substantially enhance video coding quality in VVC. However, the coding gain also leads to a greater coding complexity. Therefore, this letter proposes a Fast Decision Scheme Based on Structural Similarity Index Metric Variation (FDS-SSIMV) to solve this problem. Firstly, the Structural Similarity Index Metric Variation (SSIMV) characteristic among the sub coding units of the spit mode is illustrated. Next, to evaluate the SSIMV value, SSIMV measure strategies are designed for different split modes in this letter. Then, the desired split modes are selected by the SSIMV values. Experimental results show that the proposed method achieves an average encoding Time Saving (TS) and Bjøntegaard Delta Bit Rate (BDBR) with 64.74% and 2.79%, respectively, outperforming the benchmarks.
Jielian Lin, Zhichen Zhang, Tiesong Zhao
IEEE Signal Process. Lett.1
2021 Game Theory-driven Rate Control for 360-Degree Video Coding
abstract
The 360-degree video (omnidirectional video) has become popular recently due to its capability of providing immersive experience, which is generally achieved via spherical moving pictures with freedom of viewpoint changing. Nevertheless, the support of full-view visual contents has inevitably reshaped its perceptual quality metric and dramatically increased its bitrate output after video coding. Therefore in 360-degree video coding, the Rate Control (RC) problem, which aims to maximize the resulted perceptual quality under bitrate constraint, has become a challenging task yet to be addressed. In this paper, we observe a latitude-based bitrate discrepancy in equirectangular-projected 360-degree video coding and further utilize this feature in bitrate allocation under panoramic vision. We introduce game theory to find optimal inter/intra-frame bit allocations that maximize the overall RC performance in terms of utility function. Finally, an overall framework is proposed that is capable of providing both an improved bitrate accuracy and an enhanced perceptual quality. Experimental results demonstrate the efficiency of proposed method, with promising RC performances for 4K and 8K 360-degree videos.
Tiesong Zhao, Jielian Lin, Xu Wang 0006, Yuzhen Niu
ACM Multimedia2