VLDB 2026 Research / reviewers in the wild / expert
Wenxue Cui
dblp:199/8380
· DBLP profile ↗
24ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0001-8656-0954ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reconstructing Temporal Heterogeneity: A Multidomain Collaborative Analysis Framework for Robust Time-Series Forecasting
Hengrui Li, Wenxue Cui, Yifeng Wang 0001, Chunshan Dong, Wenju Li, Jiangpeng Shi, Yongbing Zhang 0002, Shaohui Liu |
IEEE Internet Things J. | 2 |
| 2026 | Saliency guided deep unfolding network for compressive sensing
Lang Yuan, Wenxue Cui, Xiaopeng Fan 0001 |
Multim. Syst. | 2 |
| 2026 | Geometry-Aware Cholesky Projection for Indoor Radio Map SamplingabstractMulti-frequency radio maps are vital for integrated sensing and communication, offering potential applications in indoor localization and smart homes. However, it is challenging to sample at sparse measurement locations and estimate the indoor radio map at unmeasured locations. Previous sampling algorithms do not consider the rough geometry of the furniture in the indoor environment. In order to utilize the room geometry to reduce the cost of precise on-site measurements, this work proposes a Geometry-Aware Cholesky Projection algorithm that effectively utilizes inaccurate indoor geometry information to suggest better measurement locations. Additionally, this study statistically analyzes the data distribution characteristics of 3D radio environments and radio map datasets, revealing a correlation between the room geometry and worst-case error variance. These insights justify the use of geometry information to enhance sampling efficiency in radio map reconstruction. With the proposed sampling algorithm and an autoencoder pretrained on uniform randomly masked radio maps, we find that the proposed algorithm outperforms state-of-the-art sampling algorithms. Shitong Chai, Mengyao Ma, Jiahui Li 0001, Shitong Wu, Wenxue Cui, Xiaopeng Fan 0001 |
IEEE Signal Process. Lett. | 6 |
| 2024 | SC-HVPPNet: Spatial and Channel Hybrid-Attention Video Post-Processing Network with CNN and TransformerabstractConvolutional Neural Network (CNN) and Transformer have attracted much attention recently for video post-processing (VPP). However, the interaction between CNN and Transformer in existing VPP methods is not fully explored, leading to inefficient communication between the local and global extracted features. In this paper, we explore the interaction between CNN and Transformer in the task of VPP, and propose a novel Spatial and Channel Hybrid-Attention Video Post-Processing Network (SC-HVPPNet), which can cooperatively exploit the image priors in both spatial and channel domains. Specifically, in the spatial domain, a novel spatial attention fusion module is designed, in which two attention weights are generated to fuse the local and global representations collaboratively. In the channel domain, a novel channel attention fusion module is developed, which can blend the deep representations at the channel dimension dynamically. Extensive experiments show that SC-HVPPNet notably boosts video restoration quality, with average bitrate savings of 5.29%, 12.42%, and 13.09% for Y, U, and V components in the VTM-11.0-NNVC RA configuration. Wenxue Cui, Shaohui Liu, Feng Jiang 0001 |
ICME | 2 |
| 2024 | Mesh Denoising Using Filtering Coefficients Jointly Aware of Noise and GeometryabstractMesh denoising is a fundamental task in geometry processing, and recent studies have demonstrated the remarkable superiority of deep learning-based methods in this field. However, existing works commonly rely on neural networks without explicit designs for noise and geometry which are actually fundamental factors in mesh denoising. In this paper, by jointly considering noise intensity and geometric characteristics, a novel Filtering Coefficient Learner (FCL for short) for mesh denoising is developed, which delicately generates coefficients to filter face normals. Specifically, FCL produces filtering coefficients consisting of a noise-aware component and a geometry-aware component. The first component is inversely proportional to the noise intensity of each face, resulting in smaller coefficients for faces with stronger noise. For the effective assessment of the noise intensity, a noise intensity estimation module is designed, which predicts the angle between paired noisy-clean normals based on a mean filtering angle. The second component is derived based on two types of geometric features, namely the category feature and face-wise features. The category feature provides a global description of the input patch, while the face-wise features complement the perception of local textures. Extensive experiments have validated the superior performance of FCL over SOTA works in both noise removal and feature preservation. Xianqi Zhang, Wenxue Cui, Ruiqin Xiong, Xiaopeng Fan 0001, Debin Zhao |
ACM Multimedia | 3 |
| 2024 | GeneWorker: An end-to-end robotic reinforcement learning approach with collaborative generator and worker networks
Hao Wang 0212, Hengyu Man, Wenxue Cui, Riyu Lu, Chenxin Cai, Xiaopeng Fan 0001 |
Neural Networks | 3 |
| 2024 | Deep Unfolding Network for Image Compressed Sensing by Content-Adaptive Gradient Updating and Deformation-Invariant Non-Local ModelingabstractInspired by certain optimization solvers, the deep unfolding network (DUN) has attracted much attention in recent years for image compressed sensing (CS). However, there still exist the following two issues: 1) In existing DUNs, most hyperparameters are usually content independent, which greatly limits their adaptability for different input contents. 2) In each iteration, a plain convolutional neural network is usually adopted, which weakens the perception of wider context prior and therefore depresses the expressive ability. In this article, inspired by the traditional Proximal Gradient Descent (PGD) algorithm, a novel DUN for image compressed sensing (dubbed DUN-CSNet) is proposed to solve the above two issues. Specifically, for the first issue, a novel content adaptive gradient descent network is proposed, in which a well-designed step size generation sub-network is developed to dynamically allocate the corresponding step sizes for different textures of input image by generating a content-aware step size map, realizing a content-adaptive gradient updating. For the second issue, considering the fact that many similar patches exist in an image but have undergone a deformation, a novel deformation-invariant non-local proximal mapping network is developed, which can adaptively build the long-range dependencies between the nonlocal patches by deformation-invariant non-local modeling, leading to a wider perception on context priors. Extensive experiments manifest that the proposed DUN-CSNet outperforms existing state-of-the-art CS methods by large margins. Wenxue Cui, Xiaopeng Fan 0001, Jian Zhang 0018, Debin Zhao |
IEEE Trans. Multim. | 1 |
| 2024 | Rate-Adaptive Neural Network for Image Compressive SensingabstractDeep learning-based image compressive sensing (CS) methods have achieved great success in the past few years. However, most of them are content-independent, with a spatially uniform sampling rate allocation for the entire image. Such practises may potentially degrade the performance of image CS with block-based sampling, since the content of different blocks in an image is different. In this article, we propose a novel rate-adaptive image CS neural network (dubbed RACSNet) to achieve adaptive sampling rate allocation based on the content characteristics of the image with a single model. Specifically, a measurement domain-based reconstruction distortion is first used to guide the sampling rate allocation for different blocks in an image without access to the ground truth image. Then, a step-wise training strategy is designed to train a reusable sampling matrix, which is capable of sampling image blocks to generate the compressed measurements under arbitrary sampling rates. Subsequently, a pyramid-shaped initial reconstruction sub-network and a hierarchical deep reconstruction sub-network that fuse the measurement information of different scales are put forward to reconstruct image blocks from the compressed measurements. Finally, a reconstruction distortion map and an improved loss function are developed to eliminate the blocking artifacts and further enhance the CS reconstruction. Experimental results on both objective metrics and subjective visual qualities show that the proposed RACSNet achieves significant improvements over the state-of-the-art methods. Shengping Zhang, Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Multim. | 3 |
| 2024 | Deep Network for Image Compressed Sensing Coding Using Local Structural SamplingabstractExisting image compressed sensing (CS) coding frameworks usually solve an inverse problem based on measurement coding and optimization-based image reconstruction, which still exist the following two challenges: (1) the widely used random sampling matrix, such as the Gaussian Random Matrix (GRM), usually leads to low measurement coding efficiency, and (2) the optimization-based reconstruction methods generally maintain a much higher computational complexity. In this article, we propose a new convolutional neural network based image CS coding framework using local structural sampling (dubbed CSCNet) that includes three functional modules: local structural sampling, measurement coding, and Laplacian pyramid reconstruction. In the proposed framework, instead of GRM, a new local structural sampling matrix is first developed, which is able to enhance the correlation between the measurements through a local perceptual sampling strategy. Besides, the designed local structural sampling matrix can be jointly optimized with the other functional modules during the training process. After sampling, the measurements with high correlations are produced, which are then coded into final bitstreams by the third-party image codec. Last, a Laplacian pyramid reconstruction network is proposed to efficiently recover the target image from the measurement domain to the image domain. Extensive experimental results demonstrate that the proposed scheme outperforms the existing state-of-the-art CS coding methods while maintaining fast computational speed. Wenxue Cui, Xiaopeng Fan 0001, Shaohui Liu, Xinwei Gao, Debin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Hierarchical Interactive Reconstruction Network for Video Compressive SensingabstractDeep network-based image and video Compressive Sensing (CS) has attracted increasing attentions in recent years. However, in the existing deep network-based CS methods, a simple stacked convolutional network is usually adopted, which not only weakens the perception of rich contextual prior knowledge, but also limits the exploration of the correlations between temporal video frames. In this paper, we propose a novel Hierarchical InTeractive Video CS Reconstruction Network(HIT-VCSNet), which can cooperatively exploit the deep priors in both spatial and temporal domains to improve the reconstruction quality. Specifically, in the spatial domain, a novel hierarchical structure is designed, which can hierarchically extract deep features from keyframes and non-keyframes. In the temporal domain, a novel hierarchical interaction mechanism is proposed, which can cooperatively learn the correlations among different frames in the multi-scale space. Extensive experiments manifest that the proposed HIT-VCSNet outperforms the existing state-of-the-art video and image CS methods in a large margin. Wenxue Cui, Feng Jiang 0001 |
ICASSP | 2 |
| 2023 | G2-DUN: Gradient Guided Deep Unfolding Network for Image Compressive SensingabstractInspired by certain optimization solvers, the deep unfolding network (DUN) usually inherits a multi-phase structure for image compressive sensing (CS). However, in existing DUNs, the message transmission within and between phases still faces two issues: 1) the roughness of transmitted information, e.g., the low-dimensional representations. 2) the inefficiency of transmitted policy, e.g., simply concatenating deep features. In this paper, by unfolding the Proximal Gradient Descent (PGD) algorithm, a novel gradient guided DUN (G2 -DUN) for image CS is proposed, in which a gradient map is delicately introduced within each phase for providing richer informational guidance at both intra-phase and inter-phase levels. Specifically, corresponding to the gradient descent (GD) of PGD, a gradient guided GD module is designed, in which the gradient map can adaptively guide step size allocation for different textures of input image, realizing a content-aware gradient updating. On the other hand, corresponding to the proximal mapping (PM) of PGD, a gradient guided PM module is developed, in which the gradient map can dynamically guide the exploring of deep textural priors in multi-scale space, achieving the dynamic perception of the proposed deep model. By introducing the gradient map, the proposed message transmission system not only facilitates the informational communication between different functional modules within each phase, but also strengthens the inferential cooperation among cascaded phases. Extensive experiments manifest that the proposed G2 -DUN outperforms existing state-of-the-art CS methods. Wenxue Cui, Xiaopeng Fan 0001, Shaohui Liu, Debin Zhao |
ACM Multimedia | 1 |
| 2023 | FCNet: Learning Noise-Free Features for Point Cloud DenoisingabstractThe acquisition of point clouds is usually accompanied by noise due to imperfect laser scanning or image-based reconstruction techniques. Deep learning-based methods have achieved impressive performance in point cloud denoising. However, the features captured by a denoising network from noisy point clouds are usually contaminated by noise during training. The feature noise will lead to the oscillation of back-propagated gradients, which interferes with parameter optimization and reduces the denoising performance. In this paper, we propose to explicitly clean up feature noise for point cloud denoising from two aspects: feature noise cleaning and network training. From the first aspect, we propose the feature clean network (FCNet for short) to explicitly clean up the feature noise. From the second aspect, we train FCNet by a teacher-student learning model to learn the noise-free features under the guidance of feature domain losses. Specifically, FCNet is designed with emphasis on two modules: non-local self-similarity (NSS) and weighted average pooling (WAP). NSS module smooths features through a non-local filter based on the inherent non-local self-similarity of point clouds. WAP module applies original weights calculated by the statistical outlier removal algorithm to suppress the feature noise induced by outliers. In the teacher-student learning model, we introduce a clean input using the noisy point and its clean neighbors. The teacher network accepts the clean input to capture noise-free features. The student network is trained to imitate the teacher network to learn noise-free features by minimizing the feature loss. The experiments on synthetic and real scanned point clouds show that FCNet outperforms state-of-the-art point cloud denoising methods. Wenxue Cui, Ruiqin Xiong, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Progressive Token Reduction and Compensation for Hyperspectral Image RepresentationabstractHyperspectral images (HSIs) have been widely used in Earth observation because they contain continuous and detailed spectral information which is beneficial for the fine-grained diagnosis of the land cover. In the past few years, convolutional neural network (CNN)-based methods show limitations in modeling spectral-wise long-range dependences. Recently, transformer-based deep learning methods are proposed and have shown superiority in modeling the continuous representation of the spectral signatures because the self-attention (SA) mechanism has a global receptive field. Due to the special tokenization of the transformer-based methods, the redundant tokens contained in spectral embeddings are always involved in SA operation. Redundant tokens do not positively contribute to classification. Specifically, the overlapped group-wise tokenization approach may aggravate the Hughes phenomenon and impose additional computations. To address this issue, a lightweight spatial–spectral pyramid transformer (SSPT) framework is proposed to efficiently extract the spatial–spectral features of HSI by progressively reducing redundant tokens in an end-to-end manner. In particular, a token reduction (TR) method is proposed to decide which tokens will be involved by computing and comparing token attentiveness between spectral embeddings and the class token. In addition, for those tokens that are defined as redundant information, a token compensation mechanism is proposed to automatically extract supplementary information for classification. Extensive experiments on three standard datasets quantitatively show the superiority of our methods, and the ablation experiments qualitatively prove our hypothesis about the feature distribution in transformer architecture. Junjun Jiang, Huayi Li, Wenxue Cui, Guoyuan Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Image Compressed Sensing Using Non-Local Neural NetworkabstractDeep network-based image Compressed Sensing (CS) has attracted much attention in recent years. However, the existing deep network-based CS schemes either reconstruct the target image in a block-by-block manner that leads to serious block artifacts or train the deep network as a black box that brings about limited insights of image prior knowledge. In this paper, a novel image CS framework using non-local neural network (NL-CSNet) is proposed, which utilizes the non-local self-similarity priors with deep network to improve the reconstruction quality. In the proposed NL-CSNet, two non-local subnetworks are constructed for utilizing the non-local self-similarity priors in the measurement domain and the multi-scale feature domain respectively. Specifically, in the subnetwork of measurement domain, the long-distance dependencies between the measurements of different image blocks are established for better initial reconstruction. Analogically, in the subnetwork of multi-scale feature domain, the affinities between the dense feature representations are explored in the multi-scale space for deep reconstruction. Furthermore, a novel loss function is developed to enhance the coupling between the non-local representations, which also enables an end-to-end training of NL-CSNet. Extensive experiments manifest that NL-CSNet outperforms existing state-of-the-art CS methods, while maintaining fast computational speed. Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Multim. | 1 |
| 2022 | Fast Hierarchical Deep Unfolding Network for Image Compressed SensingabstractBy integrating certain optimization solvers with deep neural network, deep unfolding network (DUN) has attracted much attention in recent years for image compressed sensing (CS). However, there still exist several issues in existing DUNs: 1) For each iteration, a simple stacked convolutional network is usually adopted, which apparently limits the expressiveness of these models. 2) Once the training is completed, most hyperparameters of existing DUNs are fixed for any input content, which significantly weakens their adaptability. In this paper, by unfolding the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA), a novel fast hierarchical DUN, dubbed FHDUN, is proposed for image compressed sensing, in which a well-designed hierarchical unfolding architecture is developed to cooperatively explore richer contextual prior information in multi-scale spaces. To further enhance the adaptability, series of hyperparametric generation networks are developed in our framework to dynamically produce the corresponding optimal hyperparameters according to the input content. Furthermore, due to the accelerated policy in FISTA, the newly embedded acceleration module makes the proposed FHDUN save more than 50% of the iterative loops against recent DUNs. Extensive CS experiments manifest that the proposed FHDUN outperforms existing state-of-the-art CS methods, while maintaining fewer iterations. Wenxue Cui, Shaohui Liu, Debin Zhao |
ACM Multimedia | 1 |
| 2021 | Adaptive Flexible 3D Histogram WatermarkingabstractWatermarking technology has attracted increasing attentions in the past few years, and a great deal traditional and deep learning-based methods have been proposed. However, these methods usually suffer from the following three challenges: First, the current algorithms are designed separately for images or videos, and there is no universal solution. Second, most algorithms cannot resist screen recording, which limits its application. Third, some algorithms can only embed fixed-length watermarks and cannot handle the embedding capacity flexibly. In this paper, a novel watermarking scheme is proposed based on spatial and temporal histograms, in which two types of histogram watermarking are designed: One is constructed in the spatial domain, using the low-frequency characteristics of the image to change the shape of the histogram and embed the watermark. The other is established in the time domain, which uses the similarity of adjacent frames and combines texture features to modify the shape of the temporal histogram to embed the watermark. Experimental results compared with the state-of-the-art demonstrate that the proposed scheme achieves superior performance. Shaohui Liu, Wenxue Cui, Jinghua Zeng, Feng Jiang 0001, Debin Zhao |
ICME | 3 |
| 2021 | Protecting the Ownership of Deep Learning Models with An End-to-End Watermarking FrameworkabstractDeep neural network (DNN), as a key component of deep learning technology, plays a vital role in its development. Most major technology companies use deep neural network as a key component to build their artificial intelligence products and service. Building a deep neural network model requires us to pay a huge price: large-scale labeled data sets, a large number of computing resources, and highly specialized domain knowledge. Therefore, we believe that the model owner owns the intellectual property rights of the model, and it is very important to design a technology that protects the intellectual property rights of the deep neural network model and allows the owner to externally verify its copyright. Through statistical analysis of a large number of pre-trained network parameters, we propose an end-to-end network model protection framework-Deep Water based on the distribution of network model parameters. First, we propose a new research problem: embedding watermarks into deep neural networks. We also define the requirements for watermarking in deep neural networks, the embedding situation, and the types of attacks. Secondly, we propose a general framework for embedding the watermark into the parameter distribution function of each layer of the convolutional network. Our method does not harm the performance of the network where the watermark is placed, because the watermark is embedded when the host network is trained. Finally, we conducted a comprehensive experiment to reveal the potential of watermarking deep neural networks as the basis for this new research work. We proved that our framework can embed watermarks in the process of training deep neural networks from scratch and in the process of fine-tuning and distillation without compromising its performance. Even after migration learning and watermark overlay operations, the embedded watermark will not disappear. Even if 65% of the parameters are trimmed, the watermark remains intact. Wei Zhang 0192, Wenxue Cui, Feng Jiang 0001, Chifu Yang |
TrustCom | 2 |
| 2020 | Multi-Stage Residual Hiding for Image-Into-Audio SteganographyabstractThe widespread application of audio communication technologies has speeded up audio data flowing across the Internet, which made it a popular carrier for covert communication. In this paper, we present a cross-modal steganography method for hiding image content into audio carriers while preserving the perceptual fidelity of the cover audio. In our framework, two multi-stage networks are designed: the first network encodes the decreasing multilevel residual errors inside different audio subsequences with the corresponding stage sub-networks, while the second network decodes the residual errors from the modified carrier with the corresponding stage sub-networks to produce the final revealed results. The multi-stage design of proposed framework not only make the controlling of payload capacity more flexible, but also make hiding easier because of the gradual sparse characteristic of residual errors. Qualitative experiments suggest that modifications to the carrier are unnoticeable by human listeners and that the decoded images are highly intelligible. Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Yongliang Liu, Debin Zhao |
ICASSP | 1 |
| 2019 | Continuous Bidirectional Optical Flow for Video Frame Sequence InterpolationabstractExisting optical flow-based frame interpolation frameworks usually suffer from two problems. First, it is difficult to accurately estimate both large motion and fine motion in the optical flow estimation stage. Second, the hole problem and occlusion problem cannot be efficiently solved in the pixel synthesis step. In this paper, we propose a novel optical flowbased frame interpolation framework, which consists of two submodules: optical flow network and pixel synthesis network. In the optical flow network, we estimate bidirectional optical flow sequences iteratively, which makes full use of the continuity of motion and therefore improves the accuracy of the optical flow estimation. Besides, a novel multi-scale architecture is developed to capture finer motions. In the pixel synthesis network, we fuse the statistical information generated during forward warping to solve the hole problem and the occlusion problem. Experimental results demonstrate that the proposed method achieves superior performance compared to state-of-the-art methods. Donghao Gu, Zhaojing Wen, Wenxue Cui, Rui Wang 0093, Feng Jiang 0001, Shaohui Liu |
ICME | 3 |
| 2018 | An Efficient Deep Convolutional Laplacian Pyramid Architecture for Cs Reconstruction At Low Sampling RatiosabstractThe compressed sensing (CS) has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been proposed and obtained superior performance. However, these methods suffer from blocking artifacts or ringing effects at low sampling ratios in most cases. To address this problem, we propose a deep convolutional Laplacian Pyramid Compressed Sensing Network (LapC-SNet) for CS, which consists of a sampling sub-network and a reconstruction sub-network. In the sampling sub-network, we utilize a convolutional layer to mimic the sampling operator. In contrast to the fixed sampling matrices used in traditional CS methods, the filters used in our convolutional layer are jointly optimized with the reconstruction sub-network. In the reconstruction sub-network, two branches are designed to reconstruct multi -scale residual images and muti -scale target images progressively using a Laplacian pyramid architecture. The proposed LapCSNet not only integrates multi-scale information to achieve better performance but also reduces computational cost dramatically. Experimental results on benchmark datasets demonstrate that the proposed method is capable of reconstructing more details and sharper edges against the state-of-the-arts methods. Wenxue Cui, Heyao Xu, Xinwei Gao, Shengping Zhang, Feng Jiang 0001, Debin Zhao |
ICASSP | 1 |
| 2018 | Deep Neural Network Based Sparse Measurement Matrix for Image Compressed SensingabstractGaussian random matrix (GRM) has been widely used to generate linear measurements in compressed sensing (CS) of natural images. However, there actually exist two disadvantages with GRM in practice. One is that GRM has large memory requirement and high computational complexity, which restrict the applications of CS. Another is that the CS measurements randomly obtained by GRM cannot provide sufficient reconstruction performances. In this paper, a Deep neural network based Sparse Measurement Matrix (DSMM) is learned by the proposed convolutional network to reduce the sampling computational complexity and improve the CS reconstruction performance. Two sub-networks are included in the proposed network, which are the sampling sub-network and the reconstruction sub-network. In the sampling sub-network, the sparsity and the normalization are both considered by the limitation of the storage and the computational complexity. In order to improve the CS reconstruction performance, a reconstruction sub-network are introduced to help enhance the sampling sub-network. So by the offline iterative training of the proposed end-to-end network, the DSMM is generated for accurate measurement and excellent reconstruction. Experimental results demonstrate that the proposed DSMM outperforms GRM greatly on representative CS reconstruction methods. Wenxue Cui, Feng Jiang 0001, Xinwei Gao, Wen Tao, Debin Zhao |
ICIP | 1 |
| 2018 | Classification Guided Deep Convolutional Network for Compressed SensingabstractCompressed Sensing (CS) has been successfully applied to image compression in the past few years. However, there are still several challenges that restrict its applications in practice including large memory requirement and unsatisfactory reconstruction performance. To address these challenges, in this paper, we propose a classification guided deep convolutional network for image compressed sensing (CCSNet), which includes a sampling sub-network and a reconstruction sub-network. In the sampling sub-network, multiple convolutional layers are used to sample the original image, which significantly reduces the parameters of the sampling matrix while causes performance degradation moderately compared against existing convolution based sampling methods. In the reconstruction sub-network, a novel two-branch architecture is proposed to improve the adaptability of the model to various textures in natural images. The first branch, named the classification branch, is to classify the sampled measurements of the original image to one of the predefined textural classes. The second branch, named the reconstruction branch, consists of multiple sub-branches, which are responsible for reconstructing the original images belonging to the corresponding textural classes. By jointly utilizing two sub-networks, the entire network can be trained in the form of end-to-end metric with a joint loss function. Experimental results demonstrate that the proposed method provides a significant quality improvement in terms of PSNR compared against state-of-the-art methods. Wenxue Cui, Shaohui Liu, Shengping Zhang, Yashu Liu 0003, Heyao Xu, Xinwei Gao, Feng Jiang 0001, Debin Zhao |
ICPR | 1 |
| 2018 | An Efficient Deep Quantized Compressed Sensing Coding Framework of Natural ImagesabstractTraditional image compressed sensing (CS) coding frameworks solve an inverse problem that is based on the measurement coding tools (prediction, quantization, entropy coding, etc.) and the optimization based image reconstruction method. These CS coding frameworks face the challenges of improving the coding efficiency at the encoder, while also suffering from high computational complexity at the decoder. In this paper, we move forward a step and propose a novel deep network based CS coding framework of natural images, which consists of three sub-networks: sampling sub-network, offset sub-network and reconstruction sub-network that responsible for sampling, quantization and reconstruction, respectively. By cooperatively utilizing these sub-networks, it can be trained in the form of an end-to-end metric with a proposed rate-distortion optimization loss function. The proposed framework not only improves the coding performance, but also reduces the computational cost of the image reconstruction dramatically. Experimental results on benchmark datasets demonstrate that the proposed method is capable of achieving superior rate-distortion performance against state-of-the-art methods. Wenxue Cui, Feng Jiang 0001, Xinwei Gao, Shengping Zhang, Debin Zhao |
ACM Multimedia | 1 |
| 2017 | Convolutional Neural Networks Based Intra Prediction for HEVCabstractSummary form only given. Traditional intra prediction methods for HEVC rely on using the nearest reference lines for predicting a block, which ignore much richer context between the current block and its neighboring blocks and therefore cause inaccurate prediction especially when weak spatial correlation exists between the current block and the reference lines. To overcome this problem, in this paper, an intra-prediction convolutional neural network (IPCNN) is proposed for intra prediction, which exploits the rich context of the current block and therefore is capable of improving the accuracy of predicting the current block. Meanwhile, the reconstruction of the three nearest blocks can also be refined. To the best of our knowledge, this is the first paper that directly applies CNNs to intra prediction for HEVC. Experimental results validate the effectiveness of applying CNNs to intra prediction and the proposed method can achieve 0.70% bitrate reduction compared to HEVC reference software HM-14.0. Wenxue Cui, Tao Zhang 0013, Shengping Zhang, Feng Jiang 0001, Wangmeng Zuo, Zhaolin Wan, Debin Zhao |
DCC | 1 |