Dongjian Yang

dblp:383/3997 · DBLP profile ↗
← Back
2ranked-venue papers in the field
2as first author
2since 2021 · last 2026
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (2 first)
YearPublicationVenuePosition
2026 Towards B-Frame Neural Video Compression with Hybrid Implicit Motion Modeling
abstract
This paper proposes a novel neural B-frame video compression framework with hybrid implicit motion modeling. In our approach, implicit motion modeling replaces the rate-consuming yet less effective flow-based explicit motion modeling to improve overall RD performance. Specifically, an interpolated frame is first generated from the forward and backward reference frames to enrich the temporal priors. A Hybrid Temporal Prior Extractor (HTPE) is then introduced to exploit these priors, where a hybrid feature extractor combining Content-Aware Depthwise Separable Convolution (CADSC) and Linear Attention Duality (LAD) adaptively captures local and global temporal features, respectively. Finally, the enriched temporal prior features are leveraged in the main encoder/decoder to enable implicit motion modeling, and are further integrated into the entropy model to improve the accuracy of entropy estimation for the discrete latent representation.
Dongjian Yang, Xiaopeng Fan 0001, Hengyu Man, Debin Zhao
DCC1
2025 Neural Image Compression with Multi-Scale Depthwise Separable Dilated Convolution and Multi-Distribution Mixture Entropy Model
abstract
Recently, neural image compression (NIC) has made remarkable progress. Two key parts of NIC are the encoder-decoder and the entropy model. For the encoder-decoder, a larger effective receptive field (ERF) means a stronger transformation ability. Existing methods usually enlarge the ERF at the expense of complexity, which is intolerable. To address this issue, we propose a multi-scale depthwise separable dilated convolution (MSDSDC) to build the encoder-decoder. Specifically, we first construct a depthwise separable dilated convolution (DSDC) by using the depthwise separable strategy in dilated convolution to reduce its complexity. Subsequently, multi-scale features extracted by three DSDCs with varying dilation rates are fused to expand the ERF of the encoder-decoder, consequently enhancing its transformation capability. Besides, we design a multi-distribution mixture entropy model (MDMEM) to further enhance the flexibility of latent representation probability modeling. The experimental results demonstrate that our proposed method achieves the best balance between rate-distortion performance and complexity.
Dongjian Yang, Xiaopeng Fan 0001, Xiandong Meng, Debin Zhao
DCC1