VLDB 2026 Research / reviewers in the wild / expert
Debin Zhao
dblp:16/3958
· DBLP profile ↗
24ranked-venue papers in the field
3as first author
3since 2021 · last 2026
0000-0003-3434-9967ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 24 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flow-Guided ConvLSTM with Quality-Aware Reconstruction for Learned Video Compression
Xiandong Meng, Hengyu Man, Xiaopeng Fan 0001, Debin Zhao |
DCC | 5 |
| 2026 | Towards B-Frame Neural Video Compression with Hybrid Implicit Motion ModelingabstractThis paper proposes a novel neural B-frame video compression framework with hybrid implicit motion modeling. In our approach, implicit motion modeling replaces the rate-consuming yet less effective flow-based explicit motion modeling to improve overall RD performance. Specifically, an interpolated frame is first generated from the forward and backward reference frames to enrich the temporal priors. A Hybrid Temporal Prior Extractor (HTPE) is then introduced to exploit these priors, where a hybrid feature extractor combining Content-Aware Depthwise Separable Convolution (CADSC) and Linear Attention Duality (LAD) adaptively captures local and global temporal features, respectively. Finally, the enriched temporal prior features are leveraged in the main encoder/decoder to enable implicit motion modeling, and are further integrated into the entropy model to improve the accuracy of entropy estimation for the discrete latent representation. Dongjian Yang, Xiaopeng Fan 0001, Hengyu Man, Debin Zhao |
DCC | 4 |
| 2025 | Neural Image Compression with Multi-Scale Depthwise Separable Dilated Convolution and Multi-Distribution Mixture Entropy ModelabstractRecently, neural image compression (NIC) has made remarkable progress. Two key parts of NIC are the encoder-decoder and the entropy model. For the encoder-decoder, a larger effective receptive field (ERF) means a stronger transformation ability. Existing methods usually enlarge the ERF at the expense of complexity, which is intolerable. To address this issue, we propose a multi-scale depthwise separable dilated convolution (MSDSDC) to build the encoder-decoder. Specifically, we first construct a depthwise separable dilated convolution (DSDC) by using the depthwise separable strategy in dilated convolution to reduce its complexity. Subsequently, multi-scale features extracted by three DSDCs with varying dilation rates are fused to expand the ERF of the encoder-decoder, consequently enhancing its transformation capability. Besides, we design a multi-distribution mixture entropy model (MDMEM) to further enhance the flexibility of latent representation probability modeling. The experimental results demonstrate that our proposed method achieves the best balance between rate-distortion performance and complexity. Dongjian Yang, Xiaopeng Fan 0001, Xiandong Meng, Debin Zhao |
DCC | 4 |
| 2017 | Convolutional Neural Networks Based Intra Prediction for HEVCabstractSummary form only given. Traditional intra prediction methods for HEVC rely on using the nearest reference lines for predicting a block, which ignore much richer context between the current block and its neighboring blocks and therefore cause inaccurate prediction especially when weak spatial correlation exists between the current block and the reference lines. To overcome this problem, in this paper, an intra-prediction convolutional neural network (IPCNN) is proposed for intra prediction, which exploits the rich context of the current block and therefore is capable of improving the accuracy of predicting the current block. Meanwhile, the reconstruction of the three nearest blocks can also be refined. To the best of our knowledge, this is the first paper that directly applies CNNs to intra prediction for HEVC. Experimental results validate the effectiveness of applying CNNs to intra prediction and the proposed method can achieve 0.70% bitrate reduction compared to HEVC reference software HM-14.0. Wenxue Cui, Tao Zhang 0013, Shengping Zhang, Feng Jiang 0001, Wangmeng Zuo, Zhaolin Wan, Debin Zhao |
DCC | 7 |
| 2017 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractSummary form only given. Traditional image coding standards (such as JPEG and JPEG2000) make the decoded image suffer from many blocking artifacts or noises since the use of big quantization steps. To overcome this problem, we proposed an end-to-end compression framework based on two CNNs, as shown in Figure 1, which produce a compact representation for encoding using a third party coding standard and reconstruct the decoded image, respectively. To make two CNNs effectively collaborate, we develop a unified end-to-end learning framework to simultaneously learn CrCNN and ReCNN such that the compact representation obtained by CrCNN preserves the structural information of the image, which facilitates to accurately reconstruct the decoded image using ReCNN and also makes the proposed compression framework compatible with existing image coding standards. Wen Tao, Feng Jiang 0001, Shengping Zhang, Jie Ren 0016, Wuzhen Shi, Wangmeng Zuo, Xun Guo 0002, Debin Zhao |
DCC | 8 |
| 2017 | Reduced Reference Image Quality Assessment Based on Entropy of Classified PrimitivesabstractThe human visual perception is a layered progressive process that brain assimilates visual information gradually, from primary information, structural information to detailed information. Recently, the visual primitives (atoms in the dictionary) extracted by sparse representation have been shown to be highly related to the layered progressive process of human visual perception. In this paper, the visual primitives are first classified into three categories: DCprimary, sketch and texture in terms of their inherent properties regarding tothe perceptual information. Then, we propose a novel reduced reference (RR) image quality assessment (IQA) metric using perceptual information represented by entropy of classified primitives (EoCP). Specifically, EoCP is a measurement of the distribution statistics of the visual primitives, which can represent the visual information. The differences of EoCPs between the reference image and its distorted version are calculated as features to characterize perceptual loss. The extracted features (only three scalars) are used to compute the quality score by a prediction function which is trained using support vector regression(SVR). Experimental results on LIVE, CSIQ and TID2013 image databases demonstrate that the proposed metric achieves high consistency with the human perception and show competitive performance with state-of-the-art IQA metrics. Zhaolin Wan, Yutao Liu 0002, Debin Zhao |
DCC | 4 |
| 2015 | Block-Based Compressive Sensing Coding of Natural Images by Local Structural Measurement MatrixabstractGaussian random matrix (GRM) has been widely used to generate linear measurements in compressive sensing (CS) of natural images. However, in practice, there actually exist two problems with GRM. One is that GRM is non-sparse and complicated, leading to high computational complexity and high difficulty in hardware implementation. The other is that regardless of the characteristics of signal the measurements generated by GRM are also random, which results in low efficiency of compression coding. In this paper, we design a novel local structural measurement matrix (LSMM) for block-based CS coding of natural images by utilizing the local smooth property of images. The proposed LSMM has two main advantages. First, LSMM is a highly sparse matrix, which can be easily implemented in hardware, and its reconstruction performance is even superior to GRM at low CS sampling sub rate. Second, the adjacent measurement elements generated by LSMM have high correlation, which can be exploited to greatly improve the coding efficiency. Furthermore, this paper presents a new framework with LSMM for block-based CS coding of natural images, including measurement generating, measurement coding and CS reconstruction. Experimental results show that the proposed framework with LSMM for block-based CS coding of natural images greatly enhances the existing CS coding performance when compared with other state-of-the-art image CS coding schemes. Xinwei Gao, Jian Zhang 0018, Wenbin Che, Xiaopeng Fan 0001, Debin Zhao |
DCC | 5 |
| 2014 | Multiple Description Image Coding with Local Random MeasurementsabstractIn this paper, an effective multiple description image coding technique is developed to achieve competitive coding efficiency at low encoder complexity, while being standard compliant. The new technique is particularly suitable for visual communication over packet-switched networks and with resource-deficient wireless devices. To keep the encoder simple and standard compliant, multiple descriptions are produced by quincunx spatial multiplexing. Each side description is a polyphase down sampled version of the input image, but the conventional low-pass filter prior to downsampling is replaced by a local random binary convolution kernel. The pixels of each resulting side description are local random measurements and placed in the original spatial configuration. The advantages of local random measurements are two folds: 1) preservation of high-frequency image features that are otherwise discarded by low-pass filtering, 2) each side description remains a conventional image and can therefore be coded by any standardized codec to remove statistical redundancy of larger scales. The decoder performs joint upsampling of received description(s) and recovers the image from local random measurements in a framework of compressive sensing. Experimental results demonstrate that the proposed multiple description image codec is competitive in rate-distortion performance compared with existing methods, with a unique strength of recovering fine details and sharp edges at low bit rates. Xianming Liu 0005, Xiaolin Wu 0001, Debin Zhao |
DCC | 3 |
| 2013 | Image Super-Resolution via Hierarchical and Collaborative Sparse RepresentationabstractIn this paper, we propose an efficient image super-resolution algorithm based on hierarchical and collaborative sparse representation (HCSR). Motivated by the observation that natural images typically exhibit multi-modal statistics, we propose a hierarchical sparse coding model which includes two layers: the first layer encodes individual patches, and the second layer jointly encodes the set of patches that belong to the same homogeneous subset of image space. We further present a simple alternative to achieve such target by identifying optimal sparse representation that is adaptive to specific statistics of images. Specially, we cluster images from the offline training set into regions of similar geometric structure, and model each region (cluster) by learning adaptive bases describing the patches within that cluster using principal component analysis (PCA). This cluster-specific dictionary is then exploited to optimally estimate the underlying HR pixel values using the idea of collaborative sparse coding, in which the similarity between patches in the same cluster is further considered. It conceptually and computationally remedies the limitation of many existing algorithms based on standard sparse coding, in which patches are independently encoded. Experimental results demonstrate the proposed method appears to be competitive with state-of-the-art algorithms. Xianming Liu 0005, Deming Zhai, Debin Zhao, Wen Gao 0001 |
DCC | 3 |
| 2013 | Progressive Image Restoration through Hybrid Graph Laplacian RegularizationabstractIn this paper, we propose a unified framework to perform progressive image restoration based on hybrid graph Laplacian regularized regression. We first construct a multi-scale representation of the target image by Laplacian pyramid, then progressively recover the degraded image in the scale space from coarse to fine so that the sharp edges and texture can be eventually recovered. On one hand, within each scale, a graph Laplacian regularization model represented by implicit kernel is learned which simultaneously minimizes the least square error on the measured samples and preserves the geometrical structure of the image data space by exploring non-local self-similarity. In this procedure, the intrinsic manifold structure is considered by using both measured and unmeasured samples. On the other hand, between two scales, the proposed model is extended to the parametric manner through explicit kernel mapping to model the inter-scale correlation, in which the local structure regularity is learned and propagated from coarser to finer scales. Experimental results on benchmark test images demonstrate that the proposed method achieves better performance than state-of-the-art image restoration algorithms. Deming Zhai, Xianming Liu 0005, Debin Zhao, Hong Chang 0001, Wen Gao 0001 |
DCC | 3 |
| 2013 | Structural Group Sparse Representation for Image Compressive Sensing RecoveryabstractCompressive Sensing (CS) theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet, contour let and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for image compressive sensing recovery via structural group sparse representation (SGSR) modeling, which enforces image sparsity and self-similarity simultaneously under a unified framework in an adaptive group domain, thus greatly confining the CS solution space. In addition, an efficient iterative shrinkage/thresholding algorithm based technique is developed to solve the above optimization problem. Experimental results demonstrate that the novel CS recovery strategy achieves significant performance improvements over the current state-of-the-art schemes and exhibits nice convergence. Jian Zhang 0018, Debin Zhao, Feng Jiang 0001, Wen Gao 0001 |
DCC | 2 |
| 2012 | Distributed Soft Video Broadcast (DCAST) with Explicit MotionabstractVideo broadcasting is a popular application of wireless network. However, the existing layered approaches can hardly accommodate users with diverse channel conditions as analog communication can do. The newly emerged `soft cast' approach, utilizing soft broadcast, provides smooth multicast performance but is not very efficient in inter frame compression. In this work, we propose a motion-aligned wireless video multicast scheme DCAST. Instead of using conventional close loop prediction (CLP), DCAST is based on distributed source coding (DSC) theory. This helps DCAST to avoid error propagation but still achieve high compression efficiency in inter frame coding. DCAST outperforms soft cast 5dB in video PSNR while maintaining the similar graceful degradation feature as soft cast. Xiaopeng Fan 0001, Feng Wu 0001, Debin Zhao, Oscar C. Au, Wen Gao 0001 |
DCC | 3 |
| 2012 | Multi-scale Spatial Error Concealment via Hybrid Bayesian RegressionabstractIn this paper, we propose a novel multi-scale spatial error concealment algorithm to combine the modeling strengthes of the parametric and nonparametric Bayesian regression. We progressively recover missing blocks in the scale space from coarse to fine so that the sharp edges and texture in the finest scale can be eventually recovered. On one hand, in each scale, the nonparametric part of the methodology is used to exploit the intra-scale correlation, which relies on the data itself to dictate the structure of the model. In this procedure, the non-local self-similarity property is utilized as a fruitful resource for abstracting a priori knowledge of images. On the other hand, the parametric part is used to explicitly model the inter-scale correlation, in which the local structure regularity is thoroughly explored to recover the sharp edges and major texture features of images. It is not respected if only the nonparametric modeling is considering. We achieve the best of both worlds within a multi-scale framework. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art error concealment algorithms. Xianming Liu 0005, Deming Zhai, Guangtao Zhai, Debin Zhao, Ruiqin Xiong, Wen Gao 0001 |
DCC | 4 |
| 2012 | Compressed Sensing Recovery via Collaborative SparsityabstractCompressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression approach. Its theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. So one of the most significant challenges in CS is to seek a domain where a signal can exhibit a high degree of sparsity and hence be recovered faithfully. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for compressed sensing recovery via collaborative sparsity (RCoS), which enforces local two-dimensional sparsity and nonlocal three-dimensional sparsity simultaneously in an adaptive hybrid space-transform domain, thus substantially utilizing intrinsic sparsities of natural images and greatly confining the CS solution space. In addition, an efficient augmented Lagrangian based technique is developed to solve the above optimization problem. Experimental results on a wide range of natural images are presented to demonstrate the efficacy of the new CS recovery strategy. Jian Zhang 0018, Debin Zhao, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
DCC | 2 |
| 2011 | Transductive Regression with Local and Global Consistency for Image Super-ResolutionabstractIn this paper, we propose a novel image super-resolution algorithm, referred to as interpolation based on transductive regression with local and global consistency (TRLGC). Our algorithm first constructs a set of local interpolation models which can predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Furthermore, a graph-Laplacian based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the accumulated loss of the locally linear regression, square error of prediction bias on the available LR samples and the manifold regularization term, which could be solved with a closed-form solution as a convex optimization problem. In this way, a transductive regression algorithm with local and global consistency is developed. Experimental results on benchmark test images demonstrate that the proposed image super-resolution method achieves very competitive performance with the state-of-the-art algorithms. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
DCC | 2 |
| 2011 | Up-sampling Dependent Frame Rate Reduction for Low Bit-Rate Video CodingabstractSummary form only given. In low bit rate video coding, the frame rate of input sequence can be reduced to the half or even smaller portion by skipping or deleting frames before compression, and then the temporal resolution is restored via up-sampling at the decoder side. Numerous algorithms have been developed to address the problem of temporal resolution improvement. Actually, the quality of up-sampled frames depends on not only the performance of up-sampling method but also the information maintained in the down-sampled video sequence. To improve the quality of up-sampled frames and smooth the quality between the up-sampled and decompressed frames, this paper proposes an up-sampling dependent frame rate reduction, which is shown in Fig. 1. The proposed low bit rate video coding scheme is composed of up-sampling dependent frame rate reduction, compression, decompression and up-sampling components. The proposed frame rate reduction method is hinged to the temporal up-sampling. It is noted that there is a feedback between frame rate reduction and up-sampling in the proposed up-sampling dependent frame rate reduction, of which the goal is to obtain a down-sampled sequence maintaining more information about the frames to be up-sampled at the decoder side. Yongbing Zhang 0002, Haoqian Wang, Debin Zhao |
DCC | 3 |
| 2010 | Error Resilient Dual Frame Motion Compensation with Uneven Quality ProtectionabstractSummary form only given. In this paper, an error resilient JU-DFMC is proposed for video transmission over error-prone channels. In the proposed error resilient JU-DFMC, a new error resilient prediction structure of DFMC is firstly presented. The LQF can adaptively select reference frame according to different packet loss rate. Then the MB information is divided into two partition header information (A) and texture coefficients (B). Based on the partition, an end-to-end distortion model is applied for macroblock (MB) level mode decision. Finally a frame level rate distortion cost scheme is proposed to determine how many times the header information will be transmitted in a high quality frame (HQF). The HQF (LTR) is given more protection. The experimental results show that the proposed method can achieve better performance than the previous DFMC schemes. In the future, how to determine LQF header transmission times will be further exploited. Debin Zhao, Siwei Ma 0001 |
DCC | 2 |
| 2010 | Auto Regressive Model and Weighted Least Squares Based Packet Video Error ConcealmentabstractIn this paper, auto regressive (AR) model is applied to error concealment for block-based packet video encoding. Each pixel within the corrupted block is restored as the weighted summation of corresponding pixels within the previous frame in a linear regression manner. Two novel algorithms using weighted least squares method are proposed to derive the AR coefficients. First, we present a coefficient derivation algorithm under the spatial continuity constraint, in which the summation of the weighted square errors within the available neighboring blocks is minimized. The confident weight of each sample is inversely proportional to the distance between the sample and the corrupted block. Second, we provide a coefficient derivation algorithm under the temporal continuity constraint, where the summation of the weighted square errors around the target pixel within the previous frame is minimized. The confident weight of each sample is proportional to the similarity of geometric proximity as well as the intensity gray level. The regression results generated by the two algorithms are then merged to form the ultimate restorations. Various experimental results demonstrate that the proposed error concealment strategy is able to increase the peak signal-to-noise ratio (PSNR) compared to other methods. Yongbing Zhang 0002, Xinguang Xiang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
DCC | 4 |
| 2009 | Compression-Induced Rendering Distortion Analysis for Texture/Depth Rate Allocation in 3D Video CompressionabstractIn 3D video applications, the virtual view is generally rendered by the compressed texture and depth. The texture and depth compression with different bit-rate overheads can lead to different virtual view rendering qualities. In this paper, we analyze the compression-induced rendering distortion for the virtual view. Based on the 3D warping principle, we first address how the texture and depth compression affects the virtual view quality, and then derive an upper bound for the compression-induced rendering distortion. The derived distortion bound depends on the compression-induced depth error and texture intensity error. Simulation results demonstrate that the theoretical upper bound is an approximate indication of the rendering quality and can be used to guide sequence-level texture/depth rate allocation for 3D video compression. Yanwei Liu 0001, Siwei Ma 0001, Qingming Huang, Debin Zhao, Wen Gao 0001, Nan Zhang 0015 |
DCC | 4 |
| 2008 | Performance Analysis of Dual Frame Motion CompensationabstractIn dual frame video coding, one short-term reference frame (STR) and one long-term reference frame (LTR) are available for motion compensation. The STR is the previous frame of current frame. The LTR remains static for a few frames, and then jump forward. In this paper, for different GOP length and bits allocation of the LTR, the coding performance of dual frame motion compensation is analyzed. The rate-distortion modeling of multi-hypothesis motion compensated prediction is employed to analyze the performance of dual frame motion compensation. Xiangyang Ji, Debin Zhao, Zhi Bian, Wen Gao 0001 |
DCC | 3 |
| 2007 | An Enhanced Robust Entropy Coder for Video Codecs Based on Context-Adaptive Reversible VLCabstractThis paper proposes an enhanced RVLC coder, context-adaptive reversible variable length coder (CRVLC), for DCT coefficients by using the techniques of data sub-partitioning and context modeling. The data sub-partitioning means that the data part of DCT coefficients is split into several small sub-partitions. As each sub-partition can be reversibly decoded by RVLC, more data as well as higher error resilience can be obtained. The context modeling exploits the correlation of DCT coefficients for further compression. This modeling defines the contexts by hierarchical-dependent information. The information is also available in the backward decoding, so that it supports the reversible decoding. And with it the data outputted by CRVLC can be naturally placed into multiple sub-partitions. Qiang Wang 0011, Debin Zhao, Siwei Ma 0001, Wen Gao 0001 |
DCC | 2 |
| 2001 | Morphological Representation of DCT Data for Image Coding
Debin Zhao, Wen Gao 0001 |
Data Compression Conference | 1 |
| 1999 | Extending DACLIC for Near Lossless Compression with Postprocessing of Greyscale ImagesabstractSummary form only given. A lossless/near lossless coding scheme, DACLIC, is presented. The proposed scheme attempts to remove redundancy in a given image in the spatial domain. The redundancy removal is achieved by block direction prediction and context-based error modeling. The block direction operation in DACLIC first partitions an image into blocks. Pixels within each incoming block are analyzed resulting in a best directional prediction for that block. The best direction is chosen from a given set that results in the minimum prediction error. Removal of redundancy by the block direction technique is not possible for removing all possible redundancy in a given image. Another decorrelation part of the DACLIC scheme is the context-based error modeling which exploits context-dependent DPCM error structures. The DACLIC scheme is primarily used as a lossless image compression technique. However, the scheme can be easily extended to near-lossless compression applications by introducing a small quantization loss. This small quantization loss is restricted to an absolute error not exceeding a prescribed value n for all pixels in a given image. Application of block direction and context modeling reduces a given image into residuals. This residual typically has a lower entropy than the given image. A quadtree Rice coder (QRC) is proposed as an entropy coder of DACLIC. An arithmetic coder is also given as an option. The QRC operates on residual blocks with low computing complexity that compares favorably with the residual coding method used by LOCO-I as the proposed QRC is two-dimensional in nature. For near-lossless compression with a larger value of n, banding artifacts are visible in the decoded image. In the DACLIC system, a postprocessing technique is proposed to remove the banding artifacts. Debin Zhao, Wen Gao 0001 |
Data Compression Conference | 1 |
| 1998 | FACOLA - Face Coder Based on Location and AttentionabstractSummary form only given. FACOLA (face coder based on location and attention) is proposed for potential applications such as compression of face pictures used in IC and ID cards. The face locator locates a face in an image using template matching and eigenface techniques. The attention detector detects high attention and low attention in the located face according to their different frequency characteristics. For high attention, usually high quality or lossless coding is required. The DCT is not suitable for such a case because it will not lead to an efficient compaction of the image energy and variable length coding (VLC) cannot be done efficiently. So a DPCM coder with three directional predictors is presented instead of the DCT. The prediction difference is quantized and the entropy coded DPCM coder supports lossless and lossy compression. The DCT coder is adopted for low attention and non-face area compression using different quantization factors. Debin Zhao, Wen Gao 0001 |
Data Compression Conference | 1 |