VLDB 2026 Research / reviewers in the wild / expert
Shen Wang 0013
dblp:80/920-13
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0004-8882-8837ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distilling Complexity-Scalable Learned Image Compression Models via Neural Architecture Search
Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Cheems Wang, Qunshan Gu, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Linear Attention Modeling for Learned Image CompressionabstractRecent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low complexity design. In this paper, we propose LALIC, a linear attention modeling for learned image compression. Specially, we propose to use Bi-RWKV blocks, by utilizing the Spatial Mix and Channel Mix modules to achieve more compact feature extraction, and apply the Conv based Omni-Shift module to adapt to two-dimensional latent representation. Furthermore, we propose a RWKV-based Spatial-Channel ConTeXt model (RWKV-SCCTX), that leverages the Bi-RWKV to modeling the correlation between neighboring features effectively. To our knowledge, our work is the first work to utilize efficient Bi-RWKV models with linear attention for learned image compression. Experimental results demonstrate that our method achieves competitive RD performances by outperforming VTM-9.1 by -15.26%, -15.41%, -17.63% in BD-rate on Kodak, CLIC and Tecnick datasets. The code is available at https://github.com/sjtu-medialab/RwkvCompress. Donghui Feng 0003, Zhengxue Cheng, Shen Wang 0013, Ronghua Wu, Hongwei Hu, Guo Lu, Li Song 0001 |
CVPR | 3 |
| 2025 | A Multi-Grid Implicit Neural Representation for Multi-View Videos
Qingyue Ling, Zhengxue Cheng, Donghui Feng 0003, Shen Wang 0013, Guo Lu, Heming Sun, Jiro Katto, Li Song 0001 |
PCS | 4 |
| 2024 | Neural Rate Control for Learned Video CompressionabstractThe learning-based video compression method has made significant progress in recent years, exhibiting promising compression performance compared with traditional video codecs. However, prior works have primarily focused on advanced compression architectures while neglecting the rate control technique. Rate control can precisely control the coding bitrate with optimal compression performance, which is a critical technique in practical deployment. To address this issue, we present a fully neural network-based rate control system for learned video compression methods. Our system accurately encodes videos at a given bitrate while enhancing the rate-distortion performance. Specifically, we first design a rate allocation model to assign optimal bitrates to each frame based on their varying spatial and temporal characteristics. Then, we propose a deep learning-based rate implementation network to perform the rate-parameter mapping, precisely predicting coding parameters for a given rate. Our proposed rate control system can be easily integrated into existing learning-based video compression methods. The extensive experimental results show that the proposed method achieves accurate rate control on several baseline methods while also improving overall rate-distortion performance. Guo Lu, Yunuo Chen 0002, Shen Wang 0013, Yibo Shi, Jing Wang 0194, Li Song 0001 |
ICLR | 4 |
| 2024 | AsymLLIC: Asymmetric Lightweight Learned Image CompressionabstractLearned image compression (LIC) methods often employ symmetrical encoder and decoder architectures, evitably increasing decoding time. However, practical scenarios demand an asymmetric design, where the decoder requires low complexity to cater to diverse low-end devices, while the encoder can accommodate higher complexity to improve coding performance. In this paper, we propose an asymmetric lightweight learned image compression (AsymLLIC) architecture with a novel training scheme, enabling the gradual substitution of complex decoding modules with simpler ones. Building upon this approach, we conduct a comprehensive comparison of different decoder network structures to strike a better trade-off between complexity and compression performance. Experiment results validate the efficiency of our proposed method, which not only achieves comparable performance to VVC but also offers a lightweight decoder with only 51.47 GMACs computation and 19.65M parameters. Furthermore, this design methodology can be easily applied to any LIC models, enabling the practical deployment of LIC techniques. Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Guo Lu, Li Song 0001, Wenjun Zhang 0001 |
VCIP | 1 |
| 2024 | A Character Position-Aware Compression Framework for Screen Text ImageabstractText patterns typically exhibit distinct boundaries and sparse color histograms. However, in current hybrid codec frameworks, the positions of coding units are often misaligned with the text patterns, resulting in prediction and color mapping tools consuming a large number of bits to indicate these patterns. Nowadays, some text detection and recognition methods have been proposed to accurately locate and analyze the text regions in screen images. Combined with these techniques, we propose a character position-aware compression framework for screen text image. On the encoder side, a low-complexity detection method is adopted to locate the text characters. Then it copies the detected characters to the position aligned with the coding unit (CU) grid to form a text layer. This text-layer representation can further increase the efficiency of existing screen content coding tools such as Intra Block Copy (IBC). Moreover, we design several compression tools based on this representation. We extend the two Motion Vector (MV) prediction modes: Adaptive Motion Vector Prediction (AMVP) and Merge. We modify the MV encoding syntax according to the layout characteristics of the text layer. We present a Gradient-guided In-loop Filter (GIF) to sharpen the text lines using a convolutional network. Experiments conducted on VVC reference software VTM all_intra configuration show that the proposed framework can achieve an average bitrate savings of 4.6% and 3.6% under the w/ GIF and w/o GIF versions, with a corresponding increase in CPU encoding complexity of 72% and 10%. Guo Lu, Huanbang Chen, Donghui Feng 0003, Shen Wang 0013, Yan Zhao 0041, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Low-Complexity Multi-Model CNN in-Loop Filter for AVS3abstractConvolutional neural network (CNN) has demonstrated powerful capabilities in many image/video processing tasks. In this paper, a low-complexity multi-model CNN in-loop filtering scheme is proposed for AVS3. Firstly, we carefully choose simplified ResNet as the lightweight single model of our proposed network. Subsequently, based on the selected single model, the multi-model iterative training framework is proposed to train a multi-model filter, where the network depth and the number of multi-models are customized for different ranges of bit rate to achieve the trade-off between model performance and computational complexity. Experimental results show that our method achieves on average 6.06% BD-rate reduction on Y component under all intra configuration. Compared to other CNN filters with comparable performance, our proposed multi-model filter can significantly reduce the decoder complexity, and the experimental results indicate that the decoding time can be saved by 26.6% on average. Shen Wang 0013, Yibing Fu, Li Song 0001, Wenjun Zhang 0001 |
ICASSP | 1 |
| 2022 | An Attention Based CNN with Temporal Hierarchical Deployment for AVS3 Inter In-loop FilteringabstractConvolutional Neural Network (CNN) based in-loop filter in video coding has demonstrated its superiority in benefiting coding efficiency and enhancing visual quality. In this paper, we develop a lightweight CNN-based in-loop filter for AVS3 encoder. The proposed network consists of several residual blocks with two attention branches, namely Dual Attention Network (DAN). The added channel attention branch and spatial attention branch can take advantage of the correlation between channels and pixels, improving the quality of reconstructed frames. In addition, by analyzing the inter prediction reference structure, we propose a temporal hierarchical deployment strategy to incorporate DAN into AVS3 video encoder. Therefore reconstructed frames with different distortions and referenced levels can be enhanced according to their temporal layer. Experiments prove the effectiveness of our strategy and results show our method achieves up to 6.57% and on average 3.64% BD-rate reduction on Y component under Random Access configuration. Yibing Fu, Shen Wang 0013, Li Song 0001, Wenjun Zhang 0001 |
ISCAS | 2 |