Chuangbai Xiao

dblp:14/914 · also Chuang-Bai Xiao · DBLP profile ↗
← Back
32ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-4676-2479ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Computer networks · 3 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Self-supervised multi-scale uniform motion deblurring via alternating optimization
Lening Guo, Jing Yu 0005, Chuangbai Xiao
Pattern Recognit.4
2025 ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
abstract
Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference speed to exchange with high-quality sample generation limits its application scenarios. In order to obtain high quality synthetic singing voice more efficiently, we propose a singing voice synthesis method based on the consistency model, ConSinger, to achieve high-fidelity singing voice synthesis with minimal steps. The model is trained by applying consistency constraint and the generation quality is greatly improved at the expense of a small amount of inference speed. Our experiments show that ConSinger is highly competitive with the baseline model in terms of generation speed and quality. Audio samples are available at https://keylxiao.github.io/consinger.
Yulin Song, Guorui Sang, Jing Yu 0016, Chuangbai Xiao
ICASSP4
2025 Adversarial Dilated Large-Kernel Attention Networks for Cross-Domain Few-Shot Hyperspectral Image Classification
abstract
Labeling hyperspectral image (HSI) data is time-consuming and laborious, resulting in limited labeled samples for training deep learning-based classifiers. To address this challenge, we propose ADLKAN, a cross-domain few-shot learning framework based on an adversarial dilated large kernel attention network. First, the ADLKAN consists of an adversarial dilated large kernel attention feature extractor, which can enhance the ability to capture sparse features. The feature extractor includes two cooperative spectral-spatial attention blocks with large kernel convolution, each of which extracts spectral and spatial features separately. These features are then fused to capture local contextual information and long-range dependencies in the spatial and spectral domains with attention networks. Second, ADLKAN also designs a multi-scale conditional domain discriminator to alleviate the weakness of few-shot learning under domain shift. The detailed multi-scale features are extracted by three sub-discriminators on different input scales. Finally, Adversarial training is then adopted to optimize the embedding feature extractor and strengthen its consistency in uncertain regions of samples from different domains. Experimental results on five public HSI datasets show that the proposed ADLKAN outperforms other state-of-the-art methods.
Jun Zhou 0001, Chuangbai Xiao
IEEE Trans. Geosci. Remote. Sens.4
2024 Prototypical Network With Residual Capsule for Few-Shot Hyperspectral Image Classification
abstract
While deep learning has been widely used in the hyperspectral image (HSI) classification, lacking labeled HSI poses significant challenges to effective and sufficient learning. To address this issue, this letter introduces a prototypical network with residual capsule (PN-ResCapsNet) for few-shot HSI classification. Compared with the convolutional neural networks (CNNs), the capsule networks can better capture spatial relationships. To better extract HSI features, residual structures and self-attention (SE) mechanisms are incorporated, which can overcome the limitation of shallow feature extraction in capsule networks. Moreover, a bias-reduction (BR) method, consisting of an interclass BR module and an intraclass BR module, is designed to rectify the prototypes, which can mitigate the bias between the support and the query set and alleviate the bias between the calculated prototype and the expected prototype. Experiments on widely used HSI datasets illustrate that the proposed method outperforms several state-of-the-art methods and achieves the overall accuracies of 92.07%, 81.02%, and 85.35% on Kennedy Space Center (KSC), Houston (HT), and Pavia University (PU) datasets, respectively.
Ruihan Fan, Jun Zhou 0001, Baoqing Guo, Chuangbai Xiao
IEEE Geosci. Remote. Sens. Lett.5
2024 Efficient scheme to perform semantic segmentation on 3-D brain tumor using 3-D u-net architecture
Zeeshan Shaukat, Qurratul Ain Farooq, Chuangbai Xiao, Faheem Akhtar Rajpoot, Muhammad Azeem 0001, Abdul Ahad Zulfiqar
Multim. Tools Appl.3
2024 Few-Shot Hyperspectral Image Classification Using Relational Generative Adversarial Network
abstract
Hyperspectral image (HSI) classification is an essential task in remote sensing, but its performance is greatly affected by limited labeled samples. Currently, generative adversarial network (GAN)-based methods can generate the virtue samples to augment the training set. However, with limited labeled data, GANs perform poorly in capturing features during sample generation. Very few relation networks (RNs) and few-shot learning (FSL) methods considered data augmentation to enhance performance. To address this challenge, we propose FSHyperRGAN, a few-shot HSI classification method based on relational GAN, which uses GANs to augment the training samples for RNs, while leveraging relation feature extraction to guide the generation of specific class samples. FSHyperRGAN comprises four modules: a data processing module converting the HSI data to 1-D and 3-D features, an adversarial generation (AG) module synthesizing virtual samples conditioned on labels, a data embedding and reconstruction (DER) module encoding latent spaces for accurate sample reconstruction while preserving category characteristics, and a relation computation (RC) module computing relation scores across generated, reconstructed, and original samples. In addition, a relational feature matching scheme is also applied, which can use virtual samples to guide classification. Two FSHyperRGAN frameworks are designed, 1D-FSHyperRGAN and 3D-FSHyperRGAN, which can be utilized for 1-D spectral or 3-D spatial-spectral classification. Experiments on widely used HSI datasets illustrate that the proposed method outperforms several state-of-the-art methods and achieves overall accuracies of 92.64%, 86.82%, 83.64%, and 84.57% on KSC, PaviaU, Houston, and WHU-Hi-HongHu datasets, respectively.
Baoqing Guo, Jun Zhou 0001, Chuangbai Xiao
IEEE Trans. Geosci. Remote. Sens.5
2024 Spectral-Spatial Large Kernel Attention Network for Hyperspectral Image Classification
abstract
Due to its ability to capture long-range dependencies, self-attention mechanism based transformer models are introduced for hyperspectral image classification. However, the self-attention mechanism has only spatial adaptability but ignores channel adaptability, thus cannot well extract complex spectral-spatial information in hyperspectral images. To tackle this problem, in this paper, we propose a novel spectral-spatial large kernel attention network (SSLKA) for hyperspectral image classification. SSLKA consists of two consecutive cooperative spectral-spatial attention blocks with large convolution kernels, which can efficiently extract features in spectral and spatial domains simultaneously. In each cooperative spectral-spatial attention block, we employ the spectral attention branch and the spatial attention branch to generate the attention maps, respectively, and then fuse the extracted spatial features with the spectral features. With large kernel attention, we can enhance the classification performance by fully exploiting local contextual information, capturing long-range dependencies, as well as be adaptive in the channel dimension. Experimental results on widely used benchmark datasets show that our method achieves higher classification accuracy in terms of overall accuracy, average accuracy, and Kappa than several state-of-the-art methods.
Chunran Wu, Jun Zhou 0001, Chuangbai Xiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Deep reinforcement learning for fault-tolerant workflow scheduling in cloud environment
Tingting Dong, Hengliang Tang, Chuangbai Xiao
Appl. Intell.4
2023 Self-Supervised Spectral-Spatial Graph Prototypical Network for Few-Shot Hyperspectral Image Classification
abstract
In recent years, deep learning has been widely applied to hyperspectral image (HSI) classification with great success. However, since labeling hyperspectral images is time-consuming and labor-intensive, a limited number of labeled hyperspectral images are available, making it difficult to train feature extractors and classifiers. To address this challenge, this paper introduces a self-supervised spectral-spatial graph prototypical network for few-shot HSI classification (S4GPN). In addition, we combine self-supervised learning with few-shot learning to provide additional semantic information to improve classification accuracy. Our method consists of three stages, including the prototype network (PN) stage, the self-supervised learning (SSL) stage, and the fusion stage (Fusion stage), with each stage progressively improving the classification performance. In the PN stage, we perform supervised learning using a prototype network structure and leverage the supervised information to guide self-supervised learning. In the SSL stage, we use the SimSiam structure in a novel way for model training after the augmentation of spectral and spatial data separately. Finally, the features learned in the first two stages are fused to improve the quality of the feature representation. These three stages use the same structured feature extractors. To extract more diverse and discriminative feature representations for the HSI classification task, our method uses the graph convolution network (GCN) and the dense network (Densenet) to extract spectral information and spatial information, respectively. Experiments on four data sets and comparisons with the state-of-the-art methods demonstrate that our proposed S4GPN outperforms other methods for HSI classification.
Shan Ma, Jun Zhou 0001, Jing Yu 0005, Chuangbai Xiao
IEEE Trans. Geosci. Remote. Sens.5
2022 COSPLAY: Concept Set Guided Personalized Dialogue Generation Across Both Party Personas
abstract
Maintaining a consistent persona is essential for building a human-like conversational model. However, the lack of attention to the partner makes the model more egocentric: they tend to show their persona by all means such as twisting the topic stiffly, pulling the conversation to their own interests regardless, and rambling their persona with little curiosity to the partner. In this work, we propose COSPLAY(COncept Set guided PersonaLized dialogue generation Across both partY personas) that considers both parties as a "team": expressing self-persona while keeping curiosity toward the partner, leading responses around mutual personas, and finding the common ground. Specifically, we first represent self-persona, partner persona and mutual dialogue all in the concept sets. Then, we propose the Concept Set framework with a suite of knowledge-enhanced operations to process them such as set algebras, set expansion, and set distance. Based on these operations as medium, we train the model by utilizing 1) concepts of both party personas, 2) concept relationship between them, and 3) their relationship to the future dialogue. Extensive experiments on a large public dataset, Persona-Chat, demonstrate that our model outperforms state-of-the-art baselines for generating less egocentric, more human-like, and higher quality responses in both automatic and human evaluations.
Piji Li, Wei Wang 0138, Chuangbai Xiao
SIGIR6
2022 A state-of-the-art technique to perform cloud-based semantic segmentation using deep learning 3D U-Net architecture
abstract
Glioma is the most aggressive and dangerous primary brain tumor with a survival time of less than 14 months. Segmentation of tumors is a necessary task in the image processing of the gliomas and is important for its timely diagnosis and starting a treatment. Using 3D U-net architecture to perform semantic segmentation on brain tumor dataset is at the core of deep learning. In this paper, we present a unique cloud-based 3D U-Net method to perform brain tumor segmentation using BRATS dataset. The system was effectively trained by using Adam optimization solver by utilizing multiple hyper parameters. We got an average dice score of 95% which makes our method the first cloud-based method to achieve maximum accuracy. The dice score is calculated by using Sørensen-Dice similarity coefficient. We also performed an extensive literature review of the brain tumor segmentation methods implemented in the last five years to get a state-of-the-art picture of well-known methodologies with a higher dice score. In comparison to the already implemented architectures, our method ranks on top in terms of accuracy in using a cloud-based 3D U-Net framework for glioma segmentation.
Zeeshan Shaukat, Qurat ul Ain Farooq, Shanshan Tu, Chuangbai Xiao
BMC Bioinform.4
2022 Route recommendation for evacuation networks using MMPP/M/1/N queueing models
Ting Sun 0006, Kaiqi Xiong, Chuangbai Xiao
Comput. Commun.4
2022 Stochastic Depth Residual Network for Hyperspectral Image Classification
abstract
The convolutional neural network (CNN) is a feed-forward neural network with deep structure and convolution operation. In the hyperspectral image (HSI) classification, CNN has demonstrated excellent performance in extracting spectral and spatial information. However, the inherent complexity and high dimension of HSIs still limit the performance of most neural network models. The powerful feature extraction ability of CNN is normally achieved by dozens or more layers, which brings a series of problems such as gradient vanishing, overfitting, and slow training speed. In order to address these problems, this article presents a CNN architecture-based stochastic depth residual network (SDRN), which is specially designed for HSI data. This model takes the original 3-D cube as the input and 3-D convolution is used to extract abundant spectral and spatial features through corresponding residual blocks. In order to reduce the training time, we adopt a stochastic depth strategy. For each small batch, a sublayer is randomly discarded by an identity function. During the testing stage, the residual network with complete depth is used. Experiments on three datasets and a comparison of the state-of-art methods show that SDRN has great advantages in accuracy and training time compared with state-of-the-art HSI classification methods.
Jun Zhou 0001, Bin Qian 0006, Jing Yu 0005, Chuangbai Xiao
IEEE Trans. Geosci. Remote. Sens.6
2022 Superpixel Spectral-Spatial Feature Fusion Graph Convolution Network for Hyperspectral Image Classification
abstract
Recently, convolutional neural networks (CNNs) have demonstrated impressive capabilities in the representation and classification of hyperspectral remote sensing images. Traditional CNNs require massive data to sufficiently train the network. To tackle this problem, graph convolutional network (GCN) has been introduced for hyperspectral image classification. GCN methods usually construct the graph from either spectral or spatial domain, which has not adequately explored the information in the joint spectral–spatial domain. In this article, we propose a superpixel spectral–spatial feature fusion graph convolution network for hyperspectral image classification (S3FGCN). S3FGCN can comprehensively use information in spectral, spatial, and spectral–spatial domains with limited data. Moreover, to enhance the performance, we explore a shared weights’ GCN in the spectral–spatial domain. To further improve the efficiency, superpixels are used to construct the adjacency matrix. Finally, dynamic sampling is adopted to make the model focus more on difficult samples. In the experiments on four datasets, S3FGCN demonstrates better accuracy compared with the state-of-the-art hyperspectral image classification methods.
Jun Zhou 0001, Bin Qian 0006, Lijuan Duan, Chuangbai Xiao
IEEE Trans. Geosci. Remote. Sens.6
2021 Change or Not: A Simple Approach for Plug and Play Language Models on Sentiment Control
abstract
Text generation with sentiment control is difficult without fine-tuning or modifying the model architecture. Plug and Play Language Model (PPLM) utilizes an external sentiment classifier to update the hidden states of GPT-2 at each time step. It does not change the parameters but achieves competitive performance. However, fluency is impaired due to the instability of the hidden states. Moreover, the classifier is not strong because of the way it is trained with partial texts, hence it is difficult to guide the generation in the process. To solve the above problems, in this paper, we first propose a fixed threshold method based on the Valence-Arousal-Dominance (VAD) lexicon to decide whether to change a word, which keeps the fluency of the original LM to the greatest extent. Furthermore, for the improvement of sentiment alignment, we propose a dynamic threshold method that utilizes VAD-based loss to make the threshold dynamic. Experiments demonstrate that our methods outperform the baseline with a great margin significantly both on fluency and sentiment accuracy.
Rang Li, Changjian Hu, Chuangbai Xiao
AAAI5
2021 Edge-Based Blur Kernel Estimation Using Sparse Representation and Self-similarity
Jing Yu 0005, Lening Guo, Chuangbai Xiao, Zhenchun Chang
ICIG (2)3
2021 Effective *-flow schedule for optical circuit switching based data center networks: A comprehensive survey
Yinan Tang, Tongtong Yuan, Bo Liu 0011, Chuangbai Xiao
Comput. Networks4
2020 Tracking area list allocation scheme based on overlapping community algorithm
Shanshan Tu, Muhammad Waqas 0001, Qiangqiang Lin, Sadaqat ur Rehman, Muhammad Hanif 0001, Chuangbai Xiao, M. Majid Butt, Chin-Chen Chang 0001
Comput. Networks6
2020 Task scheduling based on deep reinforcement learning in a cloud manufacturing environment
abstract
Summary Cloud manufacturing promotes the transformation of intelligence for the traditional manufacturing mode. In a cloud manufacturing environment, the task scheduling plays an important role. However, as the number of problem instances increases, the solution quality and computation time always go against. Existing task scheduling algorithms can get local optimal solutions with the high computational cost, especially for large problem instances. To tackle this problem, a task scheduling algorithm based on a deep reinforcement learning architecture (RLTS) is proposed to dynamically schedule tasks with precedence relationship to cloud servers to minimize the task execution time. Meanwhile, the Deep‐Q‐Network, as a kind of deep reinforcement learning algorithms, is employed to consider the problem of complexity and high dimension. In the simulation, the performance of the proposed algorithm is compared with other four heuristic algorithms. The experimental results show that RLTS can be effective to solve the task scheduling in a cloud manufacturing environment.
Tingting Dong, Chuangbai Xiao
Concurr. Comput. Pract. Exp.3
2020 Cloud-based efficient scheme for handwritten digit recognition
Zeeshan Shaukat, Qurat ul Ain Farooq, Chuangbai Xiao, Sana Sahiba, Allah Ditta
Multim. Tools Appl.4
2019 A Convolutional Neural Network for Aspect-Level Sentiment Classification
abstract
Sentiment analysis, including aspect-level sentiment classification, is an important basic natural language processing (NLP) task. Aspect-level sentiment can provide complete and in-depth results. Words with different contexts variably influence the aspect-level sentiment polarity of sentences, and polarity varies based on different aspects of a sentence. Recurrent neural networks (RNNs) are regarded as effective models for handling NLP and have performed well in aspect-level sentiment classification. Extensive literature exists on sentiment classification that utilizes convolutional neural networks (CNNs); however, no literature on aspect-level sentiment classification that uses CNNs is available. In the present study, we develop a CNN model for handling aspect-level sentiment classification. In our model, attention-based input layers are incorporated into CNN to introduce aspect information. In our experiment, in which a benchmark dataset from Twitter is compared with other models, incorporating aspect information into CNN improves aspect-level sentiment classification performance without using syntactic parser or other language features.
Yongping Xing, Chuangbai Xiao, Ziming Ding
Int. J. Pattern Recognit. Artif. Intell.2
2019 Blur kernel estimation using sparse representation and cross-scale self-similarity
abstract
Blind image deconvolution, i.e., estimating both the latent image and the blur kernel from the only observed blurry image, is a severely ill-posed inverse problem. In this paper, we propose a blur kernel estimation method for blind motion deblurring using sparse representation and cross-scale self-similarity of image patches as priors to recover the latent sharp image from a single blurry image. Sparse representation indicates that image patches can always be represented well as a sparse linear combination of atoms in an appropriate dictionary. Cross-scale self-similarity results in that any image patch can in some way be well approximated by a number of other similar patches across different image scales. Our method is based on the observations that almost any image patch in a natural image has multiple similar patches in down-sampled versions of the image, and down-sampling produces image patches that are sharper than those in the blurry image itself. In our method, the dictionary for sparse representation is trained adaptively from sharper patches sampled from the down-sampled latent image estimate to make the similar patches of the latent sharp image well represented sparsely, and meanwhile, all patches from the latent image estimate are optimized to be as close to the sharper similar patches searched from the down-sampled version to enforce the sharp recovery of the latent image by constructing a non-local regularization. Experimental results on both simulated and real blurry images demonstrate that our method outperforms state-of-the-art blind deblurring methods.
Jing Yu 0005, Zhenchun Chang, Chuangbai Xiao
Multim. Tools Appl.3
2018 Fault Detection Based on AP Clustering and PCA
abstract
To improve the accuracy, reduce the time consumption and obtain the number of faults, a fault detection method based on AP (affinity propagation) clustering and PCA (principal component analysis) was proposed. Firstly, discontinuous points in seismic horizons were searched out by the connected component labeling method. Secondly, the AP clustering algorithm was used to cluster the discontinuous points and the points of the same cluster were used to determine a fault, meanwhile, the faults existing in a seismic section were quantified. Finally, the PCA was adopted to calculate the principal direction of the discontinuous points contained in the same cluster. As a result, the corresponding cluster center and the principal direction determined a straight line, and the part that intercepted by the clustered edge was the fault we wanted. In the proposed method, the time consumption of correlation calculation of the traditional method was reduced; the computing work was simplified and the number of the faults in the seismic section was obtained. To confirm the feasibility and advancement of the proposed method, comparative experiments were done on the seismic model data and the real seismic section. The results show that the accuracy of the proposed method was better and the time cost was greatly reduced.
Chuangbai Xiao, Jing Yu 0016, Zhenli Wang
Int. J. Pattern Recognit. Artif. Intell.2
2017 An Improved Budget-Deadline Constrained Workflow Scheduling Algorithm on Heterogeneous Resources
abstract
In recent years, there are many scheduling algorithms for execution of workflow applications using Quality of Service (QoS) parameters. In this paper, we improve a scheduling workflow algorithm considering the time and cost constraints on heterogeneous resources, which is called Budget-Deadline constrained using Sub-Deadline scheduling (BDSD). With the deadline and budget constraints required by the user, we use the BDSD algorithm to find a scheduling which satisfy with the both constraints. We use the planning successful rate (PSR) to show the effectiveness of our algorithm. In the simulation experiment, we use the random workflow applications and real workflow applications to experiment. The simulation results show that compared with other algorithms, our BDSD algorithm has a high PSR and low-time complexity of O(n2m) for n tasks and m processors.
Ting Sun 0006, Chuangbai Xiao, Xiujie Xu, Guozhong Tian
CSCloud2
2017 Blind image deblurring based on sparse representation and structural self-similarity
abstract
In this paper, we propose a blind motion deblurring method based on sparse representation and structural self-similarity from a single image. The priors for sparse representation and structural self-similarity are explicitly added into the recovery of the latent image by means of sparse and multi-scale nonlocal regularizations, and the down-sampled version of the observed blurry image is used as training samples in the dictionary learning for sparse representation so that the sparsity of the latent image over this dictionary can be guaranteed, which implicitly makes use of multi-scale similar structures. Experimental results on both simulated and real blurry images demonstrate that our method outperforms existing state-of-the-art blind deblurring methods.
Jing Yu 0005, Zhenchun Chang, Chuangbai Xiao
ICASSP3
2017 Real-Time Multi-camera Video Stitching Based on Improved Optimal Stitch Line and Multi-resolution Fusion
Dongbin Xu, He-Meng Tao, Jing Yu 0016, Chuangbai Xiao
ICIG (3)4
2016 Ship wake detection for SAR images with complex backgrounds based on morphological dictionary learning
abstract
The ship wake detection of SAR images is useful not only in estimating the speed and the direction of moving ships, but also in finding small ships which are hard to be detected. The traditional ship wake detection methods of SAR images can achieve satisfactory results in simple backgrounds, but hardly work in complex backgrounds. In this paper, we propose a novel method based on the morphological component analysis and the dictionary learning to detect ship wakes in complex backgrounds. In our method, the SAR image is decomposed into a cartoon component containing ship wakes and a sea-background texture component by adaptive-ly learning the ship wake dictionary and the sea-background texture dictionary; and then the shearlet transform is used to enhance ship wakes in the cartoon component. Experimental results show our method outperforms the traditional methods for SAR images in complex backgrounds.
Guozheng Yang, Jing Yu 0005, Chuangbai Xiao
ICASSP3
2011 A Classification Algorithm to Distinguish Image as Haze or Non-haze
abstract
The technology of image dehazing can only work for haze images, but in batch and real-time processing, only relying on human visual system judge whether the image is haze or non-haze image, is unrealistic, so how to determine whether there are haze or non-haze images is needed to be solved. In this paper, we proposed a method to judge whether a given image is haze. According to the difference between the haze and non-haze images, we extract three eigen values, including image visibility, intensity of dark channel and image contrast, then combine with support vector machine to make judgment of image state which is haze or non-haze, obtaining high recognition rate. Experimental results show that our method is feasible and effective. Our method for bath and real-time processing provide the basis for judging image state, promoting the wide application of image dehazing.
Xiaoliang Yu, Chuangbai Xiao, Mike Deng
ICIG2
2009 A new strategy to predict the search range in H.264/AVC
abstract
Motion estimation is a very time consuming part in H.264 codec. In order to reduce motion estimation time, many strategies have been used during this process. Dynamic search range is one of them. In this paper, based on the analysis of the problems existed in current algorithm, we propose a new strategy to predict motion estimation search range by using the information of image size, block mode and QP. Experiment results show that the proposed algorithm can reduce up to 16.94% of the motion estimation time and save 9.15% of the total encoding time only with 0.001 dB PSNR loss and 0.51% BitRate increment comparing with existing DSR algorithm in UMHexagon Search algorithm. In addition, this algorithm is very easy to combine with other Motion Estimation algorithms in JM.
Xiaobing Lu, Chuangbai Xiao
ICME2
2009 An Efficient Method for Early Detecting All-Zero Quantized DCT Coefficients for H.264/AVC
abstract
The residual signal gained after motion estimation needs to be transformed and quantized prior to entropy coding in H.264/AVC. In many cases, especially in low bitrate situation, the probability that obtain all zero quantized coefficients block (AZB) is high. In this paper, we study the characteristics of 4×4 integer transform first, and then we use the divide and rule strategy to inverse the formulae of integer transform and quantization. By this way, we get a new sufficient condition for AZB detection. The experiment results show that our algorithm can find out more AZBs compared with existing methods without any loss and save more computations at the same time.
Chuangbai Xiao, Shoudao Wang, Mu Ling
SMC2
2008 An adaptive fast multiple reference frames selection algorithm for H.264/AVC
abstract
To make full use of the temporal correlation of video sequences, H.264/AVC adopts multiple reference frames to enhance the coding quality. While the performance being improved, the complexity of computation has been increased linearly, too. In this paper, we study motion characteristic of video sequences and spatial correlation in video frames first, and then we propose an adaptive fast multiple reference frames selection algorithm. It can decrease the number of reference frames for motion compensation, and reduce the complexity of coding adaptively according to the features of video sequences. The results show that our algorithm achieves 45% coding time saving on average with unnoticeable quality loss.
Chuangbai Xiao
ICASSP2
2008 Gibbs artifact reduction for POCS super-resolution image reconstruction
Chuangbai Xiao, Jing Yu 0016, Kaina Su
Frontiers Comput. Sci. China1