Chenqiu Zhao

dblp:187/1639 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0003-4574-4815ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Local Autoregression with Finite-Support Random Variables for Image Generation
Chenqiu Zhao, Anup Basu
ICPR (3)1
2024 LMoW: A Latent Random Variable Model for Unconditional Human Motion Generation
Justin Rozeboom, Hanran Song, Chenqiu Zhao, Anup Basu
MMAsia4
2024 Accelerating Inference of Networks in the Frequency Domain
Chenqiu Zhao, Guanfang Dong, Anup Basu
MMAsia1
2024 Learning Temporal Distribution and Spatial Correlation Toward Universal Moving Object Segmentation
abstract
The goal of moving object segmentation is separating moving objects from stationary backgrounds in videos. One major challenge in this problem is how to develop a universal model for videos from various natural scenes since previous methods are often effective only in specific scenes. In this paper, we propose a method called Learning Temporal Distribution and Spatial Correlation (LTS) that has the potential to be a general solution for universal moving object segmentation. In the proposed approach, the distribution from temporal pixels is first learned by our Defect Iterative Distribution Learning (DIDL) network for a scene-independent segmentation. Notably, the DIDL network incorporates the use of an improved product distribution layer that we have newly derived. Then, the Stochastic Bayesian Refinement (SBR) Network, which learns the spatial correlation, is proposed to improve the binary mask generated by the DIDL network. Benefiting from the scene independence of the temporal distribution and the accuracy improvement resulting from the spatial correlation, the proposed approach performs well for almost all videos from diverse and complex natural scenes with fixed parameters. Comprehensive experiments on standard datasets including LASIESTA, CDNet2014, BMC, SBMI2015 and 128 real world videos demonstrate the superiority of proposed approach compared to state-of-the-art methods with or without the use of deep learning networks. To the best of our knowledge, this work has high potential to be a general solution for moving object segmentation in real world environments. The code and real-world videos can be found on GitHub https://github.com/guanfangdong/LTS-UniverisalMOS.
Guanfang Dong, Chenqiu Zhao, Xichen Pan, Anup Basu
IEEE Trans. Image Process.2
2024 RAST: Restorable Arbitrary Style Transfer
abstract
The objective of arbitrary style transfer is to apply a given artistic or photo-realistic style to a target image. Although current methods have shown some success in transferring style, arbitrary style transfer still has several issues, including content leakage. Embedding an artistic style can result in unintended changes to the image content. This article proposes an iterative framework called Restorable Arbitrary Style Transfer (RAST) to effectively ensure content preservation and mitigate potential alterations to the content information. RAST can transmit both content and style information through multi-restorations and balance the content-style tradeoff in stylized images using the image restoration accuracy. To ensure RAST’s effectiveness, we introduce two novel loss functions: multi-restoration loss and style difference loss. We also propose a new quantitative evaluation method to assess content preservation and style embedding performance. Experimental results show that RAST outperforms state-of-the-art methods in generating stylized images that preserve content and embed style accurately.
Yingnan Ma, Chenqiu Zhao, Bingran Huang, Anup Basu
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Principal Component Approximation Network for Image Compression
abstract
In this work, we propose a novel principal component approximation network (PCANet) for image compression. The proposed network is based on the assumption that a set of images can be decomposed into several shared feature matrices, and an image can be reconstructed by the weighted sum of these matrices. The proposed PCANet is specifically devised to learn and approximate these feature matrices and weight vectors, which are used to encode images for compression. Unlike previous deep learning-based methods, a distinctive aspect of our approach is its consideration of network size in the bit-rate computation. Despite this inclusion, our proposed method yields promising results. Through extensive experiments conducted on standard datasets, we demonstrate the effectiveness of our approach in comparison to state-of-the-art techniques. To the best of our knowledge, this is the first machine learning approach that includes the size of networks during bitrate computation in image compression.
Chenqiu Zhao, Anup Basu
ACM Trans. Multim. Comput. Commun. Appl.2
2023 RAST: Restorable Arbitrary Style Transfer via Multi-restoration
abstract
Arbitrary style transfer aims to reproduce the target image with the artistic or photo-realistic styles provided. Even though existing approaches can successfully transfer style information, arbitrary style transfer still faces many challenges, such as the content leak issue. Specifically, the embedding of artistic style can lead to content changes. In this paper, we solve the content leak problem from the perspective of image restoration. In particular, an iterative architecture is proposed to achieve the Restorable Arbitrary Style Transfer (RAST), which can realize transmission of both content and style information through multi-restorations. We control the content-style balance in stylized images by the accuracy of image restoration. In order to ensure effectiveness of the proposed RAST architecture, we design two novel loss functions: multi-restoration loss and style difference loss. In addition, we propose a new quantitative evaluation method to measure content preservation performance and style embedding performance. Comprehensive experiments comparing with state-of-the-art methods demonstrate that our proposed architecture can produce stylized images with superior performance on content preservation and style embedding.
Yingnan Ma, Chenqiu Zhao, Anup Basu
WACV2
2022 Universal Background Subtraction Based on Arithmetic Distribution Neural Network
abstract
We propose a universal background subtraction framework based on the Arithmetic Distribution Neural Network (ADNN) for learning the distributions of temporal pixels. In our ADNN model, the arithmetic distribution operations are utilized to introduce the arithmetic distribution layers, including the product distribution layer and the sum distribution layer. Furthermore, in order to improve the accuracy of the proposed approach, an improved Bayesian refinement model based on neighboring information, with a GPU implementation, is incorporated. In the forward pass and backpropagation of the proposed arithmetic distribution layers, histograms are considered as probability density functions rather than matrices. Thus, the proposed approach is able to utilize the probability information of the histogram and achieve promising results with a very simple architecture compared to traditional convolutional neural networks. Evaluations using standard benchmarks demonstrate the superiority of the proposed approach compared to state-of-the-art traditional and deep learning methods. To the best of our knowledge, this is the first method to propose network layers based on arithmetic distribution operations for learning distributions during background subtraction.
Chenqiu Zhao, Kangkang Hu, Anup Basu
IEEE Trans. Image Process.1
2021 Deep Variation Transformation Network for Foreground Detection
abstract
In existing literature, the distribution of pixel observations is analyzed with models designed for the video foreground detection task. However, it is possible that the background and foreground share similar observations, causing false detections. We propose a novel foreground detection method called Deep Variation Transformation Network (DVTN), focusing on analyzing the pixel variations instead of distributions. In particular, pixel variations are represented by a sequence of pixel observations, and DVTN is trained to transform the pixel variations into a new space, where the observations can be classified easily. Following this, the output of DVTN is utilized by a linear classifier to label pixels as foreground or background. As a result of the global analysis and the strong learning ability of DVTN, the proposed approach adaptively learns a good transformation from pixel variations to probabilities of labels to improve performance. Comprehensive experiments on several benchmark datasets demonstrate the superiority of our DVTN approach compared to both state-of-the-art deep learning and traditional methods, especially in scenes lacking texture and color information. Code is available at https://github.com/Zhangjunyin/DVTN.
Yongxin Ge, Junyin Zhang, Xinyu Ren, Chenqiu Zhao, Anup Basu
IEEE Trans. Circuits Syst. Video Technol.4
2020 Multi-Scale Deep Pixel Distribution Learning for Concrete Crack Detection
abstract
A number of methods including image processing technologies (IPTs) and deep learning methods, have been used to detect defects in civilian infrastructure. These methods have been introduced to extract features representing cracks in concrete surfaces. Inspired by recent advances of a pixel distribution learning method in background subtraction, we propose a novel multi-scale deep learning method (MS-DPDL) for concrete crack detection. The designed CNN network is trained on the dataset CRACK500 [1], [2] and tested on it for concrete segmentation. To show good transferability of our proposed model, it is later tested on the dataset Concrete Crack Images for Classification [3]. Several existing deep learning methods are used to compare the performance of the proposed MS-DPDL method. Results show that our method has good performance and can effectively find concrete cracks in practical situations.
Xuanyi Wu, Jianfei Ma, Chenqiu Zhao, Anup Basu
ICPR4
2020 Dynamic Deep Pixel Distribution Learning for Background Subtraction
abstract
Previous approaches to background subtraction usually approximate the distribution of pixels with artificial models. In this paper, we focus on automatically learning the distribution, using a novel background subtraction model named Dynamic Deep Pixel Distribution Learning (D-DPDL). In our D-DPDL model, a distribution descriptor named Random Permutation of Temporal Pixels (RPoTP) is dynamically generated as the input to a convolutional neural network for learning the statistical distribution, and a Bayesian refinement model is tailored to handle the random noise introduced by the random permutation. Because the temporal pixels are randomly permutated to guarantee that only statistical information is retained in RPoTP features, the network is forced to learn the pixel distribution. Moreover, since the noise is random, the Bayesian theorem is naturally selected to propose an empirical model as a compensation based on the similarity between pixels. Evaluations using standard benchmark demonstrates the superiority of the proposed approach compared with the state-of-the-art, including traditional methods as well as deep learning methods.
Chenqiu Zhao, Anup Basu
IEEE Trans. Circuits Syst. Video Technol.1
2019 Background Subtraction Based on Integration of Alternative Cues in Freely Moving Camera
abstract
Previous approaches to background subtraction in freely moving camera typically focus on improving the accuracy of motion estimation. In this paper, we propose that the accurate background subtraction is possible with the integration of alternative cues about foreground and background. We also put forward a novel background subtraction framework called the integration of foreground and background cues. Here, the foreground cues are extracted by the Gaussian mixture model compensated with image alignment, while the background cues are obtained from the spatiotemporal features filtered by the homography transformation. Subsequently, the integration is devised as a hierarchical competition procedure based on super-pixels under multiple levels with the underlying motivation to utilize the exclusiveness between these cues for the compensation of their corresponding defects. The result of competition between foreground and background cues in a particular super-pixel is used as the proximity, and the foreground is segmented by combining super-pixels with proximity under multiple levels. Comprehensive evaluations using standard benchmarks demonstrate the superiority of our work compared with the state-of-the-art.
Chenqiu Zhao, Aneeshan Sain, Ying Qu 0007, Yongxin Ge, Haibo Hu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2018 Background Subtraction Based on Deep Pixel Distribution Learning
abstract
Previous approaches to background subtraction typically address the problem by formulating a representation of the background, and comparing the background to new frames. In this work, we focus on the essence of background subtraction, which is the classification of a pixel's current observation in comparison to historical observations, and propose a Deep Pixel Distribution Learning (DPDL) model for background subtraction. In the DPDL model, a novel pixel-based feature, called the Random Permutation of Temporal Pixels (RPoTP), is used to represent the distribution of past observations for a particular pixel, in which the temporal correlation between observations is deliberately obfuscated. Subsequently a convolutional neural network (CNN) is used to learn the distribution for determining whether the current observation is foreground or background, with the random permutation enabling the framework to focus primarily on the distribution of observations, rather than be misled by learning spurious temporal correlations. In addition, the pixel-wise representation allows for a large number of RPoTP features to be captured even with a limited number of groundtruth frames, with the DPDL model being effective even with only a single groundtruth frame. The proposed framework is able to achieve promising results in diverse natural scenes, and a comprehensive evaluation on standard benchmarks demonstrates the superiority of our work to state-of-the-art methods. The source code ispublicly available at https://github.com/zhaochenqiu/DPDL
Chenqiu Zhao, Tat-Jen Cham, Xinyu Ren, Jianfei Cai 0001, Haichen Zhu
ICME1
2018 Background Modeling by Stability of Adaptive Features in Complex Scenes
abstract
The single-feature-based background model often fails in complex scenes, since a pixel is better described by several features, which highlight different characteristics of it. Therefore, the multi-feature-based background model has drawn much attention recently. In this paper, we propose a novel multi-feature-based background model, named stability of adaptive feature (SoAF) model, which utilizes the stabilities of different features in a pixel to adaptively weigh the contributions of these features for foreground detection. We do this mainly due to the fact that the features of pixels in the background are often more stable. In SoAF, a pixel is described by several features and each of these features is depicted by a unimodal model that offers an initial label of the target pixel. Then, we measure the stability of each feature by its histogram statistics over a time sequence and use them as weights to assemble the aforementioned unimodal models to yield the final label. The experiments on some standard benchmarks, which contain the complex scenes, demonstrate that the proposed approach achieves promising performance in comparison with some state-of-the-art approaches.
Dan Yang 0001, Chenqiu Zhao, Xiaohong Zhang 0002, Sheng Huang 0001
IEEE Trans. Image Process.2
2016 Collaborative Sparse Preserving Projections for Feature Extraction
abstract
Sparsity Preserving Projections (SPP) is a well known approach for feature extraction and dimensionality reduction. Its success is mainly attributed to its high quality graph which is constructed by sparse representation. As an instance of graph embedding, SPP can be formulated as regression model. Thus we apply the idea of collaborative graph embedding, which reformulates SPP as a collaborative representation model via imposing a L2-norm constraint to projections from the perspective of linear regression, to further enhance SPP. We call this novel SPP method Collaborative Sparsity Preserving Projections (CSPP). Experiment results on four popular face datasets, namely Yale, ORL, FERET and AR, show the effectiveness in feature extraction and the improvement of CSPP over SPP.
Yunsong Wu, Qianying Huang, Xiaohong Zhang 0002, Chenqiu Zhao
ICSS4