EDBT 2026 Demo / reviewers in the wild / expert
Jian Zhang 0018
dblp:07/314-18
· DBLP profile ↗
12ranked-venue papers in the field
3as first author
3since 2021 · last 2026
0000-0001-5486-3125ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 12 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio-Visual Cross-Modal Compression for Generative Face Video CodingabstractGenerative face video coding (GFVC) is vital for modern applications like video conferencing, yet existing methods primarily focus on video motion while neglecting the significant bitrate contribution of audio. Despite the well-established correlation between audio and lip movements, this cross-modal coherence has not been systematically exploited for compression. To address this, we propose an Audio-Visual Cross-Modal Compression (AVCC) framework that jointly compresses audio and video streams. Our framework extracts motion information from video and tokenizes audio features, then aligns them through a unified audio-video diffusion process. This allows synchronized reconstruction of both modalities from a shared representation. In extremely low-rate scenarios, AVCC can even reconstruct one modality from the other. Experiments show that AVCC significantly outperforms the Versatile Video Coding (VVC) standard and state-of-the-art GFVC schemes in rate-distortion performance, paving the way for more efficient multimodal communication systems. Youmin Xu, Mengxi Guo, Shijie Zhao 0001, Li Zhang 0006, Jian Zhang 0018 |
DCC | 7 |
| 2022 | Semantic Neural Rendering-based Video Coding: Towards Ultra-Low Bitrate Video ConferencingabstractProviding high video quality under the lowest possible bitrate constraint is one of the critical challenges in video coding technology. Inspired by the continuous development of motion imitation [1], the model-based video coding method is derived from extracting a series of features or parameters representing the person's motion and reconstructing each frame by motion imitation model at the decoder. Thus, we propose a Semantic Neural Rendering-based Video Coding framework (SNRVC) to transmit video at ultra-low bitrate while maintaining high subjective quality. At the encoder, we first extract the motion parameters with specific semantic meanings from each frame and then compress the first frame and the parameters of the subsequent frames by truncating to different decimals and differential pulse code modulation coding. Finally, the decoded image and parameters are fed into the motion imitator [2] to obtain each reconstructed frame consistent with the movements of the original frame. Our SNRVC can achieve better visual quality than traditional and model-based methods [3] at the ultra-low bitrate below 0.01 bpp. Youmin Xu, Jianhui Chang, Jian Zhang 0018 |
DCC | 4 |
| 2021 | Invertible Resampling-Based Layered Image CompressionabstractFlow-based generative models are successfully applied in image generation tasks, where an invertible neural network (INN) is built up based on flow steps. Learning-based compression commonly transforms the input into a compact space and then implements a reconstruction network in the decoder accordingly. By utilizing low-resolution images, traditional or adaptive downsamplers with their corresponding traditional or learned upsamplers usually achieve better coding quality at low bit-rate. This paper proposes a novel image compression framework named Invertible Resampling-based Layered Image Compression (IRLIC). A rescaling network is built by splitting the input image into a downsampled image and a high-frequency part which is transformed into pre-defined distribution by INN, with symmetrical upsampling. Thus, reliable rescaling is applied in the total lossy compression framework, where only the downsampled image and the reconstruction residual are needed to recover a compressed image. Our IRLIC achieves superior performance to the current methods like BPG and other learning-based image compressions at the bit-rate below 1.8bpp. Youmin Xu, Jian Zhang 0018 |
DCC | 2 |
| 2017 | Effective Quadtree Plus Binary Tree Block Partition Decision for Future Video CodingabstractBlock partition structure has been recognized as a crucial module in video coding scheme. Recently, a quadtree plus binary tree (QTBT) block partition structure has been proposed in the Joint Video Exploration Team (JVET) development. Compared to the quadtree structure in HEVC, QTBT can achieve better coding performance with hugely increased encoding complexity. Here, we propose an effective QTBT partition decision algorithm to achieve a good trade-off between computational complexity and coding performance. In particular, at the Coding Tree Unit level, the partition parameters of QTBT are dynamically derived to adapt to the local characteristics without transmitting any overhead. Subsequently, at the Coding Unit level, a joint-classifier decision tree structure is designed to eliminate unnecessary iterations and meanwhile control the risk of false prediction. Experimental results show that the proposed algorithm can achieve 64% encoding time reduction on average with only 1.26% increase in terms of bit rate. This greatly benefits the practical implementations of QTBT in real application scenarios. Zhao Wang 0004, Shiqi Wang 0001, Jian Zhang 0018, Shanshe Wang, Siwei Ma 0001 |
DCC | 3 |
| 2017 | Globally Variance-Constrained Sparse Representation for Rate-Distortion Optimized Image RepresentationabstractSparse representation is efficient to approximately recover signals by a linear composition of a few bases from an over-complete dictionary. However, in the scenario of data compression, its efficiency and popularity are hindered due to the extra overhead for encoding the sparse coefficients. Therefore, how to establish an accurate rate model in sparse coding and dictionary learning becomes meaningful, which has been not fully exploited in the context of sparse representation. According to the Shannon entropy inequality, the variance of data source can bound its entropy, thus can reflect the actual coding bits. Therefore, a Globally Variance-Constrained Sparse Representation (GVCSR) model is proposed, where a variance-constrained rate term is introduced to the conventional sparse representation. To solve the non-convex optimization problem, we employ the Alternating Direction Method of Multipliers (ADMM) for sparse coding and dictionary learning, both of which have shown state-of-the-art rate-distortion performance in image representation. Xiang Zhang 0004, Siwei Ma 0001, Zhouchen Lin, Jian Zhang 0018, Shiqi Wang 0001, Wen Gao 0001 |
DCC | 4 |
| 2016 | Adaptive Motion Vector Resolution Scheme for Enhanced Video CodingabstractIn the state-of-the-art H.265/HEVC video coding standard, the motion vector is always fixed to be 1/4-pixel resolution for the entire video sequence regardless of the different video contents, which is not efficient for prediction coding. In this paper, we propose a frame level adaptive motion vector resolution selection scheme based on a rate-distortion model in terms of motion vector resolution. In the proposed rate-distortion model, the relationship between the distortion and the motion vector resolution is approximated with a linear model. And a rate model of motion vector is built, which reflects the relationship between the coding bits of motion vector and its value. With the proposed rate-distortion model, an optimal motion vector resolution minimizing the total rate-distortion cost will be selected for each frame. Experimental results show that the proposed scheme can achieve 1.5%, 1.3% and 2.5% BD-rate gain on average for Random Access, Lowdelay-B and Lowdelay-P configurations without complexity increment. Zhao Wang 0004, Jian Zhang 0018, Nan Zhang 0015, Siwei Ma 0001 |
DCC | 2 |
| 2016 | Structure-driven Adaptive Non-local Filter for High Efficiency Video Coding (HEVC)abstractDeblocking filter (DF) Is High Efficiency Video Coding (HEVC) is Only Applied to all Samples Adjacent to prediction units (PU), or transform units (TU), which actually exists two issues. The first one is that DF in HEVC does not fully exploit nonlocal similarity structure information in video. The second one is that DF is HEVC does not consider the inside pixels, which often suffer from quantization distrotion. To alleviate these issues, in this paper, a structure-driven adaptive non-local filter (SANF) Is Proposed By Simultaneously Enforcing The Intrinsic Local Sparsity And The Non-Local Self-Similarity Of Each Frame. Not only SANF deals with the boundary pixels, but also the inside area, which is able to effectively reduce block artifacts while enhancing the quality of the deblocked frames. Applying SANF to luma and chroma components after DF, simulation results demonstrate that the proposed SANF can save BD-rate reduction up to 10.3% with ALF off. For luma component, SANF achieves 4.1%. 3.3%, 4.4% BD-rate saving for all intra, low delay B and random access configurations, respectively with ALF off. furthermore, the performance with ALF on is also discussed. Jian Zhang 0018, Chuanmin Jia, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001 |
DCC | 1 |
| 2016 | Nonconvex Lp Nuclear Norm based ADMM Framework for Compressed SensingabstractCompressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression methodology. Recent studies further show that image prior models play an important role in image CS recovery. By exploiting the non-local self-similarity of natural images and clustering similar patches, low-rank prior model is adopted in this paper. Different from traditional nuclear norm, we extend thelp(0plpnuclear norm prior model for image CS recovery, which is able to more accurately enforce image structural sparsity and self-similarity at the same time. The proposed optimization problem is efficiently solved within the alternative direction multiplier method (ADMM) framework. Experimental results demonstrate that the proposedlpnuclear norm based ADMM framework for image CS recovery framework exhibits good convergence and achieves significant performance improvements over the current state-of-the-art methods. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001 |
DCC | 2 |
| 2016 | Compressive-Sensed Image Coding via Stripe-based DPCMabstractThese years have seen the advances of compressive sensing (CS), but efficient coding of sensed measurements is still an issue. In this paper, we propose an image coding system based on the compressive sensing paradigm via stripe-based differential pulse-code modulation (DPCM). In the system, we sample and encode an image in a unit of multiple rows, which we call a stripe. Through extensive experiments, we observe that the correlation between measurements of adjacent stripes are much higher than that of the neighboring blocks. Based on this, we combine the stripe-based CS acquisition with the DPCM framework and design a mechanism that predicts a stripe of measurements from its preceding stripe of measurements. The produced measurement residuals are then quantized and entropy-encoded into binary coding bits, which are tremendously reduced compared to the traditional block-based framework. Furthermore, we provide an image CS reconstruction algorithm corresponding to the stripe-based acquisition. Experiments verify that the reconstruction quality is no worse or even better than the block-based case when much lower bitrate is consumed. In a rate-distortion point of view, the proposed system also outperforms the methods using block-based sampling and achieves the state-of-the-art performance for compressive-sensed image coding. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001 |
DCC | 2 |
| 2015 | Block-Based Compressive Sensing Coding of Natural Images by Local Structural Measurement MatrixabstractGaussian random matrix (GRM) has been widely used to generate linear measurements in compressive sensing (CS) of natural images. However, in practice, there actually exist two problems with GRM. One is that GRM is non-sparse and complicated, leading to high computational complexity and high difficulty in hardware implementation. The other is that regardless of the characteristics of signal the measurements generated by GRM are also random, which results in low efficiency of compression coding. In this paper, we design a novel local structural measurement matrix (LSMM) for block-based CS coding of natural images by utilizing the local smooth property of images. The proposed LSMM has two main advantages. First, LSMM is a highly sparse matrix, which can be easily implemented in hardware, and its reconstruction performance is even superior to GRM at low CS sampling sub rate. Second, the adjacent measurement elements generated by LSMM have high correlation, which can be exploited to greatly improve the coding efficiency. Furthermore, this paper presents a new framework with LSMM for block-based CS coding of natural images, including measurement generating, measurement coding and CS reconstruction. Experimental results show that the proposed framework with LSMM for block-based CS coding of natural images greatly enhances the existing CS coding performance when compared with other state-of-the-art image CS coding schemes. Xinwei Gao, Jian Zhang 0018, Wenbin Che, Xiaopeng Fan 0001, Debin Zhao |
DCC | 2 |
| 2013 | Structural Group Sparse Representation for Image Compressive Sensing RecoveryabstractCompressive Sensing (CS) theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet, contour let and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for image compressive sensing recovery via structural group sparse representation (SGSR) modeling, which enforces image sparsity and self-similarity simultaneously under a unified framework in an adaptive group domain, thus greatly confining the CS solution space. In addition, an efficient iterative shrinkage/thresholding algorithm based technique is developed to solve the above optimization problem. Experimental results demonstrate that the novel CS recovery strategy achieves significant performance improvements over the current state-of-the-art schemes and exhibits nice convergence. Jian Zhang 0018, Debin Zhao, Feng Jiang 0001, Wen Gao 0001 |
DCC | 1 |
| 2012 | Compressed Sensing Recovery via Collaborative SparsityabstractCompressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression approach. Its theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. So one of the most significant challenges in CS is to seek a domain where a signal can exhibit a high degree of sparsity and hence be recovered faithfully. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for compressed sensing recovery via collaborative sparsity (RCoS), which enforces local two-dimensional sparsity and nonlocal three-dimensional sparsity simultaneously in an adaptive hybrid space-transform domain, thus substantially utilizing intrinsic sparsities of natural images and greatly confining the CS solution space. In addition, an efficient augmented Lagrangian based technique is developed to solve the above optimization problem. Experimental results on a wide range of natural images are presented to demonstrate the efficacy of the new CS recovery strategy. Jian Zhang 0018, Debin Zhao, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
DCC | 1 |