EDBT 2026 Demo / reviewers in the wild / expert
Donghui Feng 0003
dblp:19/6847-3
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0001-8984-0425ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distilling Complexity-Scalable Learned Image Compression Models via Neural Architecture Search
Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Cheems Wang, Qunshan Gu, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Linear Attention Modeling for Learned Image CompressionabstractRecent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low complexity design. In this paper, we propose LALIC, a linear attention modeling for learned image compression. Specially, we propose to use Bi-RWKV blocks, by utilizing the Spatial Mix and Channel Mix modules to achieve more compact feature extraction, and apply the Conv based Omni-Shift module to adapt to two-dimensional latent representation. Furthermore, we propose a RWKV-based Spatial-Channel ConTeXt model (RWKV-SCCTX), that leverages the Bi-RWKV to modeling the correlation between neighboring features effectively. To our knowledge, our work is the first work to utilize efficient Bi-RWKV models with linear attention for learned image compression. Experimental results demonstrate that our method achieves competitive RD performances by outperforming VTM-9.1 by -15.26%, -15.41%, -17.63% in BD-rate on Kodak, CLIC and Tecnick datasets. The code is available at https://github.com/sjtu-medialab/RwkvCompress. Donghui Feng 0003, Zhengxue Cheng, Shen Wang 0013, Ronghua Wu, Hongwei Hu, Guo Lu, Li Song 0001 |
CVPR | 1 |
| 2025 | Towards a New Paradigm of Visual Signal CompressionabstractUltra-low bitrate image compression is a challenging and demand- ing topic. With the development of Large Multimodal Models (LMMs), a Cross Modality Compression (CMC) paradigm of Image-Text- Image has emerged. Compared with traditional codecs, this semantic- level compression can reduce image data size to 0.1% or even lower, which has strong potential applications. However, CMC has cer- tain defects in consistency with the original image and perceptual quality. To inspire insights into such a problem, we introduce CMC- Bench, a benchmark of the cooperative performance of Image-to- Text (I2T) and Text-to-Image (T2I) models for image compression. This benchmark covers 18,000 and 40,000 images respectively to verify 6 mainstream I2T and 12 T2I models, including 160,000 sub- jective preference scores annotated by human experts. At ultra-low bitrates, it proves that the combination of some I2T and T2I models has surpassed the most advanced visual signal codecs; meanwhile, it highlights where LMMs can be further optimized toward the compression task. We encourage LMM developers to participate in this test to promote the evolution of visual signal codec protocols. Chunyi Li 0001, Xiele Wu, Haoning Wu 0001, Donghui Feng 0003, Guo Lu, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ACM Multimedia | 4 |
| 2025 | A Multi-Grid Implicit Neural Representation for Multi-View Videos
Qingyue Ling, Zhengxue Cheng, Donghui Feng 0003, Shen Wang 0013, Guo Lu, Heming Sun, Jiro Katto, Li Song 0001 |
PCS | 3 |
| 2025 | MISC: Ultra-Low Bitrate Image Semantic Compression Driven by Large Multimodal ModelabstractWith the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, all existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. During recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC. Chunyi Li 0001, Guo Lu, Donghui Feng 0003, Haoning Wu 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Instance-Adaptive Spatial-Temporal Enhancement for Efficient Video CompressionabstractEfficiently compressing HD/UHD content has long been challenging due to high bitrate costs. Instance-adaptive enhancement methods try to tackle this issue by compressing a video at reduced resolution and enhancing it using a neural model specifically overfitted for this video. However, existing methods focus solely on spatial super-resolution (SR) and under-utilize the videos' temporal redundancy. Their limited management of the model's updated parameters also causes excessive overfitting overheads. Therefore, this paper introduces IASTE, the first instance-adaptive enhancement method based on spatial-temporal enhancement (STE), and incorporates low-rank adaptation (LoRA) for efficient model overfitting. Specifically, we downscale videos spatially and temporally to reduce the data volume and achieve efficient video compression. Then, we overfit a specific STE model for each video and use it to enhance the decoded video's spatiotemporal resolution. Leveraging the video swin transformer's strong capability in capturing spatiotemporal correlations, we design a lightweight and efficient model to implement video STE. The model is overfitted for each video using LoRA. By freezing the pre-trained model and selectively updating a few low-rank matrices, the bitrate overhead for model storage can be mitigated. Experiments prove that compared to directly compressing high-frame-rate (HFR), high-resolution (HR) videos, our method achieves around 30% BD-Rate gains on the CTC and UVG datasets, about 15% gains on the YoutubeUGC dataset, and about 10% gains on the ultra-long videos in the Xiph dataset. Yan Zhao 0041, Zhengxue Cheng, Jiangchuan Li, Donghui Feng 0003, Qunshan Gu, Cheems Wang, Guo Lu, Li Song 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | AsymLLIC: Asymmetric Lightweight Learned Image CompressionabstractLearned image compression (LIC) methods often employ symmetrical encoder and decoder architectures, evitably increasing decoding time. However, practical scenarios demand an asymmetric design, where the decoder requires low complexity to cater to diverse low-end devices, while the encoder can accommodate higher complexity to improve coding performance. In this paper, we propose an asymmetric lightweight learned image compression (AsymLLIC) architecture with a novel training scheme, enabling the gradual substitution of complex decoding modules with simpler ones. Building upon this approach, we conduct a comprehensive comparison of different decoder network structures to strike a better trade-off between complexity and compression performance. Experiment results validate the efficiency of our proposed method, which not only achieves comparable performance to VVC but also offers a lightweight decoder with only 51.47 GMACs computation and 19.65M parameters. Furthermore, this design methodology can be easily applied to any LIC models, enabling the practical deployment of LIC techniques. Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Guo Lu, Li Song 0001, Wenjun Zhang 0001 |
VCIP | 3 |
| 2024 | Coarse-to-fine Transformer For Lossless 3D Medical Image CompressionabstractThe rapid advancements in medical imaging have led to a growing demand for high-performance lossless compression of large 3D medical image datasets. Unlike natural images, medical images typically feature three-dimensional structures, and high bit-depth, necessitating specialized compression techniques. Based on a decoder-only transformer, we propose a learnable dual-decoder model for lossless compression of 3D medical images. Our approach packs voxels into patches, which are processed by a patch-level decoder to extract the patch feature. The voxels, along with the patch feature, are subsequently fed into a voxel-level decoder to model each voxel. This coarse-to-fine modeling strategy reduces the computational time for each voxel and enables long-range modeling dependencies. Experimental results demonstrate that our proposed model achieves state-of-the-art compression performance, with an approximately 15% improvement in compression performance over the traditional JP3D benchmark on various datasets. Guo Lu, Donghui Feng 0003, Zhengxue Cheng, Guosheng Yu, Li Song 0001 |
VCIP | 3 |
| 2024 | A Character Position-Aware Compression Framework for Screen Text ImageabstractText patterns typically exhibit distinct boundaries and sparse color histograms. However, in current hybrid codec frameworks, the positions of coding units are often misaligned with the text patterns, resulting in prediction and color mapping tools consuming a large number of bits to indicate these patterns. Nowadays, some text detection and recognition methods have been proposed to accurately locate and analyze the text regions in screen images. Combined with these techniques, we propose a character position-aware compression framework for screen text image. On the encoder side, a low-complexity detection method is adopted to locate the text characters. Then it copies the detected characters to the position aligned with the coding unit (CU) grid to form a text layer. This text-layer representation can further increase the efficiency of existing screen content coding tools such as Intra Block Copy (IBC). Moreover, we design several compression tools based on this representation. We extend the two Motion Vector (MV) prediction modes: Adaptive Motion Vector Prediction (AMVP) and Merge. We modify the MV encoding syntax according to the layout characteristics of the text layer. We present a Gradient-guided In-loop Filter (GIF) to sharpen the text lines using a convolutional network. Experiments conducted on VVC reference software VTM all_intra configuration show that the proposed framework can achieve an average bitrate savings of 4.6% and 3.6% under the w/ GIF and w/o GIF versions, with a corresponding increase in CPU encoding complexity of 72% and 10%. Guo Lu, Huanbang Chen, Donghui Feng 0003, Shen Wang 0013, Yan Zhao 0041, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Content Adaptive Checkerboard Context Model for Learned Image CompressionabstractLearned image compression methods are becoming popular and have achieved excellent performance, of which joint context and hyperprior architectures are the mainstream. In order to avoid the time-consuming serial decoding pipeline introduced by the autoregressive context model, the checkerboard context model (CCM) is proposed to implement fast two-pass coding. However, CCM sets half of the latents as anchors to extract spatial context for the other non-anchors, which is rough and redundant. We propose a more precise and flexible content adaptive checkerboard context model to decrease the numbers and bit consumption of anchors. By introducing pseudo-anchors for simple regions in latents, our method can preserve the capability of fast two-pass coding and outperform CCM in Rate-Distortion performance on several baseline models with negligible computational overhead. Guo Lu, Donghui Feng 0003, Li Song 0001 |
ISCAS | 3 |
| 2022 | Complexity-Oriented Per-Shot Video Coding OptimizationabstractCurrent per-shot encoding schemes aim to improve the compression efficiency by shot-level optimization. It splits a source video sequence into shots and imposes optimal sets of encoding parameters on each shot. Per-shot encoding achieved approximately 20% bitrate savings over baseline fixed QP encoding at the expense of pre-processing complexity. However, the adjustable parameter space of the current per-shot encoding schemes only has spatial resolution and QP/CRF, resulting in a lack of encoding flexibility. In this paper, we extend the per-shot encoding framework in the complexity dimension. We believe that per-shot encoding with flexible complexity will help in deploying user-generated content. We propose a rate-distortion-complexity optimization process for encoders and a methodology to determine the coding parameters under the constraints of complexities and bitrate ladders. Experimental results show that our proposed method achieves complexity constraints ranging from 100% to 3% in a dense form compared to the slowest per-shot anchor. With similar complexities of the per-shot scheme fixed in specific presets, our proposed method achieves BDrate gain up to −19.17%. Hongcheng Zhong, Jun Xu 0040, Donghui Feng 0003, Li Song 0001 |
ICME | 4 |
| 2022 | Position-based Motion Vector Prediction for Textual Image CodingabstractTextual content is becoming increasingly important in video conferencing, while existing screen content encoding tools still produce a high bitrate in text regions. The main coding tool Intra Block Copy (IBC) inherits the MV prediction mechanism in inter-frame coding, but the adjacent text characters typically have irrelevant MVs, making it inefficient to predict MV using only neighbor MVs. To solve the problem, we propose the Position-based Motion Vector Prediction, to cache IBC AMVP PU positions as predictors. One character can find the previously encoded position to construct a good MV prediction. Experiment results show the effectiveness of the proposed prediction scheme. Donghui Feng 0003, Guo Lu, Li Song 0001 |
PCS | 1 |
| 2022 | Edge-Based Video Compression Texture Synthesis Using Generative Adversarial NetworkabstractIt has been recognized that texture patterns with abundant high-frequency components, such as grass and water, produce visual masking effects, and the distortion in textures is hard to be perceived by human eyes than structure regions. However, modern video codecs in a rate-distortion optimized manner usually consume a lot of bits to encode textures, leading to the insufficiency in perceptual coding performance. Nowadays, with the rapid development of deep learning, learning based texture synthesis methods have been proposed to replace the coding process of prediction residuals to reduce the rate cost. In this paper, we present a deep texture synthesizer named edge-based texture synthesis framework (ETSF). At encoder side, the framework detects texture regions by semantic and fidelity classification criteria, and the detected regions are quantized coarsely by the hybrid coding framework. In texture characterization, ETSF extracts low-level edge features representing pixel intensity variation. Feature processing tools are developed to remove the spatiotemporal redundancy of edges. The processed edge information is compressed and transmitted. To effectively recover textures, we design an edge-based texture synthesis generative adversarial network (ETSGAN) at the decoder of ETSF, which can incorporate edge information into convolutional layers and generate realistic textures. Experimental results on a collected texture dataset show that the proposed ETSF can achieve an average of -12.8%, -14.2% and -9.6% MOS BD-rate under lowdelay_B, lowdelay_P and random_access configurations of VVC coding, respectively. Jun Xu 0040, Donghui Feng 0003, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | DVRCNN: Dark Video Post-processing Method for VVC
Donghui Feng 0003, Han Zhang 0030, Li Song 0001 |
MMM (1) | 1 |