Weiping Li 0003

dblp:77/4748-3 · DBLP profile ↗
← Back
52ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0003-3750-1653ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 since 2021Artificial intelligence and machine learning · 8 · 2 since 2021Systems, architecture and hardware · 7 · 1 since 2021Theory of computation · 5Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Versatile learned video compression
Runsen Feng, Zongyu Guo, Zhizheng Zhang 0004, Weiping Li 0003, Zhibo Chen 0001
J. Vis. Commun. Image Represent.4
2025 Towards Defining an Efficient and Expandable File Format for AI-Generated Contents
abstract
Recently, AI-generated content (AIGC) has gained significant traction due to its powerful creation capability. However, the storage and transmission of large amounts of high-quality AIGC images inevitably pose new challenges for recent file formats. To overcome this, we define a new file format for AIGC images, named AIGIF, enabling ultra-low bitrate coding of AIGC images. Unlike compressing AIGC images intuitively with pixel-wise space as existing file formats, AIGIF instead compresses the generation syntax. This raises a crucial question: Which generation syntax elements, e.g., text prompt, device configuration, etc, are necessary for compression/transmission? To answer this question, we systematically investigate the effects of three essential factors: platform, generative model, and data configuration. We experimentally find that a well-designed composable bitstream structure incorporating the above three factors can achieve an impressive compression ratio of even up to 1/10,000 while still ensuring high fidelity. We also introduce an expandable syntax in AIGIF to support the extension of the most advanced generation models to be developed in the future.
Runsen Feng, Xin Li 0082, Weiping Li 0003, Zhibo Chen 0001
ISCAS4
2023 NVTC: Nonlinear Vector Transform Coding
abstract
In theory, vector quantization (VQ) is always better than scalar quantization (SQ) in terms of rate-distortion (RD) performance [33]. Recent state-of-the-art methods for neural image compression are mainly based on nonlinear transform coding (NTC) with uniform scalar quantization, overlooking the benefits of VQ due to its exponentially increased complexity. In this paper, we first investigate on some toy sources, demonstrating that even if modern neural networks considerably enhance the compression performance of SQ with nonlinear transform, there is still an insurmountable chasm between SQ and VQ. Therefore, revolving around VQ, we propose a novel framework for neural image compression named Nonlinear Vector Transform Coding (NVTC). NVTC solves the critical complexity issue of VQ through (1) a multi-stage quantization strategy and (2) nonlinear vector transforms. In addition, we apply entropy-constrained VQ in latent space to adaptively determine the quantization boundaries for joint rate-distortion optimization, which improves the performance both theoretically and experimentally. Compared to previous NTC approaches, NVTC demonstrates superior rate-distortion performance, faster decoding speed, and smaller model size. Our code is available at https://github.com/USTC-IMCL/NVTC.
Runsen Feng, Zongyu Guo, Weiping Li 0003, Zhibo Chen 0001
CVPR3
2022 Two-Step Fast Mode Decision for Intra Coding of Screen Content
abstract
With the rapid development of screen content video applications, screen content coding (SCC) is urgently needed to be used in commercial codecs. However, the extra encoding complexity introduced by the new SCC tools has posed a great challenge for its practical deployment. In this paper, motivated by our observations that there should be a fine-grained mapping between image content and candidate modes, we propose a two-step fast mode decision method to reduce the encoding complexity. First, we propose to use a convolution neural network (CNN) to automatically extract useful features for fine-grained content classification. Second, we build a precise and concise mapping from CUs to candidate modes by simultaneously considering CU content type, CU size, and mode complexity. Note that the spatial correlations between neighboring CUs and current CU are also utilized in candidate modes derivation. In addition to the two-step fast mode decision method, a content-aware early termination algorithm is further proposed to reduce the encoding complexity. Extensive experiments demonstrate that our method achieves better performance compared with state-of-the-art ones, with 50.13% total encoding complexity reduction and only 0.92% BD-rate increase.
Changsheng Gao, Li Li 0040, Dong Liu 0002, Zhibo Chen 0001, Weiping Li 0003, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Spatiotemporal Generative Adversarial Network-Based Dynamic Texture Synthesis for Surveillance Video Coding
abstract
Dynamic texture refers to the content in video sequences that is characterized by spatial repetition and temporal variation, such as swaying foliage and flowing water. It is a great challenge to compress the dynamic textures efficiently in the current prediction/transform hybrid video coding framework. However, these textures have little information for machine vision, and human visual perception is less sensitive to the textures than to the structures. Thus, we propose a spatiotemporal generative adversarial network (GAN) based dynamic texture synthesis method for surveillance video coding. We detect and remove the dynamic texture content at encoder side, which is irrelevant to machine vision. We generate the dynamic texture content using the proposed GAN at decoder side, so that the reconstructed videos can be observed by human without deteriorating perceptual quality. Specifically, we design a GAN network to synthesize dynamic textures by exploiting the correlation between spatial and temporal neighbors; we present a surveillance video coding scheme with the dynamic texture detection/synthesis method; we build a high-quality dynamic texture dataset, and we collect a dynamic texture testing dataset that goes beyond the existing video coding test datasets by focusing on surveillance scenes. The proposed video coding scheme has been implemented on top of the High Efficiency Video Coding (HEVC) reference software. Experiments have been conducted to evaluate the quantitative and qualitative performance of the proposed coding scheme. Our method achieves 7.4% and 7.6% bit-rate savings in low-delay-B and low-delay-P settings, respectively, at similar visual quality levels in comparison with HEVC.
Dong Liu 0002, Zhibo Chen 0001, Feng Wu 0001, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.5
2021 Corrections to "Blind quality assessment for image superresolution using deep two-stream convolutional networks"
Wei Zhou 0021, Qiuping Jiang, Yuwang Wang, Zhibo Chen 0001, Weiping Li 0003
Inf. Sci.5
2021 Hierarchical visual comfort assessment for stereoscopic image retargeting
Zhibo Chen 0001, Weiping Li 0003
Signal Process. Image Commun.3
2021 Weakly Supervised Reinforced Multi-Operator Image Retargeting
abstract
Image retargeting aims to adjust the resolution and aspect ratio to an arbitrary size while preserving important content of the image. Usually multi-operator image retargeting demonstrates better generalization than single operator scheme due to heterogeneous characteristics of different regions in the image. Most existing multi-operator retargeting methods search the optimal operator at each step with exponential complexity and with the possibility of falling into local optimum. Therefore, in order to produce better results with lower computational costs, we formulate the multi-operator retargeting as a Markov decision-making process and apply Reinforcement Learning (RL) to achieve global optimum. Instead of using traditional image-level measures, we design a high-level semantic and aesthetic reward function to better match human visual perception. With the priori in reward, we further propose a weakly supervised Semantics and Aesthetics aware Multi-operator Image Retargeting (SAMIR) framework. Particularly, the semantic part of the reward helps to constrain the severe deformations that may occur during retargeting process, while the aesthetic part guarantees the sensory quality, which can effectively measure the perceptual effects of different operators on various image content. The operator of each step is learned in an end-to-end manner. In addition, retargeting can be performed in arbitrary target size, step size, and direction. The experiment results on both representative aesthetic datasets and retargeting datasets consistently show that our model outperforms the state-of-the-art methods.
Zhibo Chen 0001, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.3
2021 Asynchronous Episodic Deep Deterministic Policy Gradient: Toward Continuous Control in Computationally Complex Environments
abstract
Deep deterministic policy gradient (DDPG) has been proved to be a successful reinforcement learning (RL) algorithm for continuous control tasks. However, DDPG still suffers from data insufficiency and training inefficiency, especially, in computationally complex environments. In this article, we propose asynchronous episodic DDPG (AE-DDPG), as an expansion of DDPG, which can achieve more effective learning with less training time required. First, we design a modified scheme for data collection in an asynchronous fashion. Generally, for asynchronous RL algorithms, sample efficiency or/and training stability diminish as the degree of parallelism increases. We consider this problem from the perspectives of both data generation and data utilization. In detail, we redesign experience replay by introducing the idea of episodic control so that the agent can latch on good trajectories rapidly. In addition, we also inject a new type of noise in action space to enrich the exploration behaviors. Experiments demonstrate that our AE-DDPG achieves higher rewards and requires less time consumption than most popular RL algorithms in learning to run task which has a computationally complex environment. Not limited to the control tasks in the computationally complex environments, AE-DDPG also achieves higher rewards and two-fold to four-fold improvement in sample efficiency on average compared with other variants of DDPG in MuJoCo environments. Furthermore, we verify the effectiveness of each proposed technique component through abundant ablation study.
Zhizheng Zhang 0004, Jiale Chen 0001, Zhibo Chen 0001, Weiping Li 0003
IEEE Trans. Cybern.4
2020 Region Normalization for Image Inpainting
abstract
Feature Normalization (FN) is an important technique to help neural network training, which typically normalizes features across spatial dimensions. Most previous image inpainting methods apply FN in their networks without considering the impact of the corrupted regions of the input image on normalization, e.g. mean and variance shifts. In this work, we show that the mean and variance shifts caused by full-spatial FN limit the image inpainting network training and we propose a spatial region-wise normalization named Region Normalization (RN) to overcome the limitation. RN divides spatial pixels into different regions according to the input mask, and computes the mean and variance in each region for normalization. We develop two kinds of RN for our image inpainting network: (1) Basic RN (RN-B), which normalizes pixels from the corrupted and uncorrupted regions separately based on the original inpainting mask to solve the mean and variance shift problem; (2) Learnable RN (RN-L), which automatically detects potentially corrupted and uncorrupted regions for separate normalization, and performs global affine transformation to enhance their fusion. We apply RN-B in the early layers and RN-L in the latter layers of the network respectively. Experiments show that our method outperforms current state-of-the-art methods quantitatively and qualitatively. We further generalize RN to other inpainting networks and achieve consistent performance improvements.
Tao Yu 0012, Zongyu Guo, Xin Jin 0014, Shilin Wu, Zhibo Chen 0001, Weiping Li 0003, Zhizheng Zhang 0004, Sen Liu 0001
AAAI6
2020 Blind quality assessment for image superresolution using deep two-stream convolutional networks
Wei Zhou 0021, Qiuping Jiang, Yuwang Wang, Zhibo Chen 0001, Weiping Li 0003
Inf. Sci.5
2020 AI-GAN: Asynchronous interactive generative adversarial network for single image rain removal
Xin Jin 0014, Zhibo Chen 0001, Weiping Li 0003
Pattern Recognit.3
2020 Learned Fast HEVC Intra Coding
abstract
In High Efficiency Video Coding (HEVC), excellent rate-distortion (RD) performance is achieved in part by having a flexible quadtree coding unit (CU) partition and a large number of intra-prediction modes. Such an excellent RD performance is achieved at the expense of much higher computational complexity. In this paper, we propose a learned fast HEVC intra coding (LFHI) framework taking into account the comprehensive factors of fast intra coding to reach an improved configurable tradeoff between coding performance and computational complexity. First, we design a low-complex shallow asymmetric-kernel CNN (AK-CNN) to efficiently extract the local directional texture features of each block for both fast CU partition and fast intra-mode decision. Second, we introduce the concept of the minimum number of RDO candidates (MNRC) into fast mode decision, which utilizes AK-CNN to predict the minimum number of best candidates for RDO calculation to further reduce the computation of intra-mode selection. Third, an evolution optimized threshold decision (EOTD) scheme is designed to achieve configurable complexity-efficiency tradeoffs. Finally, we propose an interpolation-based prediction scheme that allows for our framework to be generalized to all quantization parameters (QPs) without the need for training the network on each QP. The experimental results demonstrate that the LFHI framework has a high degree of parallelism and achieves a much better complexity-efficiency tradeoff, achieving up to 75.2% intra-mode encoding complexity reduction with negligible rate-distortion performance degradation, superior to the existing fast intra-coding schemes.
Zhibo Chen 0001, Jun Shi 0004, Weiping Li 0003
IEEE Trans. Image Process.3
2020 On the Energy-Delay Tradeoff in Streaming Data: Finite Blocklength Analysis
abstract
This paper investigates basic trade-offs between energy and delay in wireless communication systems using finite blocklength theory. We first assume that data arrive in constant stream of bits, which are put into packets and transmitted over a communications link. Our results show that depending on exactly how energy is measured, in general energy depends √d-1or √Vd-1log d, where d is the delay. This means on that the energy decreases quite slowly with increasing delay. Furthermore, to approach the absolute minimum of -1.59 dB on energy, bandwidth has to increase very rapidly, much more than what is predicted by infinite blocklength theory. We then consider the scenario when data arrive stochastically in packets and can be queued. We devise a scheduling algorithm based on finite blocklength theory and develop bounds for the energy-delay performance. Our results again show that the energy decreases quite slowly with increasing delay.
Mirza Uzair Baig, Lei Yu 0003, Zixiang Xiong, Anders Høst-Madsen, Houqiang Li, Weiping Li 0003
IEEE Trans. Inf. Theory6
2019 Learned Scalable Image Compression with Bidirectional Context Disentanglement Network
abstract
In this paper, we propose a learned scalable/progressive image compression scheme based on deep neural networks (DNN), named Bidirectional Context Disentanglement Network (BCD-Net). For learning hierarchical representations, we first adopt bit-plane decomposition to decompose the information coarsely before the deep-learning-based transformation. However, the information carried by different bit-planes is not only unequal in entropy but also of different importance for reconstruction. We thus take the hidden features corresponding to different bit-planes as the context and design a network topology with bidirectional flows to disentangle the contextual information for more effective compressed representations. Our proposed scheme enables us to obtain the compressed codes with scalable rates via a one-pass encoding-decoding. Experiment results demonstrate that our proposed model outperforms the state-of-the-art DNN-based scalable image compression methods in both PSNR and MS-SSIM metrics. In addition, our proposed model achieves better performance in MS-SSIM metric than conventional scalable image codecs. Effectiveness of our technical components is also verified through sufficient ablation experiments.
Zhizheng Zhang 0004, Zhibo Chen 0001, Weiping Li 0003
ICME4
2019 Exploiting Weight-Level Sparsity in Channel Pruning with Low-Rank Approximation
abstract
Acceleration and compression on Deep Neural Networks (DNNs) have become a critical problem to develop intelligence on resource-constrained hardware, especially on Internet of Things (IoT) devices. Previous works based on channel pruning can be easily deployed and accelerated without specialized hardware and software. However, weight-level sparsity is not well explored in channel pruning, which results in relatively low compression rate. In this work, we propose a framework that combines channel pruning with low-rank decomposition to tackle this problem. First, the low-rank decomposition is utilized to eliminate redundancy within filter, and achieves acceleration in shallow layers. Then, we apply channel pruning on the decomposed network in a global way, and obtains further acceleration in deep layers. In addition, a spectral norm-based indicator is proposed to balance low-rank approximation and channel pruning. We conduct a series of ablation experiments and prove that low-rank decomposition can effectively improve channel pruning by generating small and compact filters. To further demonstrate the hardware compatibility, we deploy the pruned networks on the FPGA, and the networks produced by our method have obviously low latency.
Zhen Chen 0013, Sen Liu 0001, Zhibo Chen 0001, Weiping Li 0003
ISCAS5
2019 Importance-Aware Filter Selection for Convolutional Neural Network Acceleration
abstract
Convolutional Neural Networks(CNNs) are widely used in many fields, including artificial intelligence, computer vision and video coding. However, CNNs are typically over-parameterized and contain significant redundancy. Traditional model acceleration methods mainly rely on specific manual rules. This usually leads to sub-optimal results with relatively limited compression ratio. Recent works have deployed the self-learning agent on the layer-level acceleration but still combined with human-designed criterias. In this paper, we proposed a filter-based model acceleration method to directly and automatically decide which filters should be pruned with the reinforcement learning method DDPG. We designed a novel reward function with the reward shaping technique for the training process. Our method is utilized on the models trained on MNIST and CIFAR-10 datasets and achieves both higher acceleration ratio and less accuracy loss than the conventional methods simultaneously.
Zikun Liu 0002, Zhen Chen 0013, Weiping Li 0003
VCIP3
2019 Multi-tracker fusion via adaptive outlier detection
Ning Wang 0020, Wengang Zhou 0001, Weiping Li 0003, Houqiang Li
Multim. Tools Appl.4
2019 Multi-View Vehicle Type Recognition With Feedback-Enhancement Multi-Branch CNNs
abstract
Vehicle type recognition (VTR) is a quite common requirement and one of the key challenges in real surveillance scenarios, such as intelligent traffic and unmanned driving. Usually coarse-grained and fine-grained VTRs are applied in different applications, and the challenge from multiple viewpoints is critical for both cases. In this paper, we propose a feedback-enhancement multi-branch CNN (FM-CNN) to solve the challenge in these two cases. The proposed FM-CNN takes three derivatives of an image as input and leverages the advantages of hierarchical details, feedback enhancement, model average, and stronger robustness to translation and mirroring. A single global cross-entropy loss is insufficient to train such a complex CNN and so we add extra branch losses to enhance feedbacks to each branch. Though reusing pre-trained parameters, we propose a novel parameter update method to adapt FM-CNN to task-specific local visual patterns and global information in new datasets. To test the effectiveness of FM-CNN, we create our own multi-view VTR (MVVTR) data set since there are no such data sets available. And, for fine-grained VTR, we use the CompCars data set. Compared with state-of-the-art classification solutions without special preprocessing, the proposed FM-CNN demonstrates better performance in both coarse-grained and fine-grained scenarios. For coarse-grained VTR, it achieves 94.9% Top-1 accuracy on the MVVTR data set. For fine-grained VTR, it achieves 91.0% Top-1 and 97.8% Top-5 accuracies on the CompCars data set.
Zhibo Chen 0001, Chenlu Ying, Chaoyi Lin, Sen Liu 0001, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.5
2019 Attention-Based 3D-CNNs for Large-Vocabulary Sign Language Recognition
abstract
Sign language recognition (SLR) is an important and challenging research topic in the multimedia field. Conventional techniques for SLR rely on hand-crafted features, which achieve limited success. In this paper, we present attention-based 3D-convolutional neural networks (3D-CNNs) for SLR. The framework has two advantages: 3D-CNNs learn spatio-temporal features from raw video without prior knowledge and the attention mechanism helps to select the clue. When training 3D-CNN for capturing spatio-temporal features, spatial attention is incorporated into the network to focus on the areas of interest. After feature extraction, temporal attention is utilized to select the significant motions for classification. The proposed method is evaluated on two large scale sign language data sets. The first one, collected by ourselves, is a Chinese sign language data set that consists of 500 categories. The other is the ChaLearn14 benchmark. The experiment results demonstrate the effectiveness of our approach compared with state-of-the-art algorithms.
Jie Huang 0011, Wengang Zhou 0001, Houqiang Li, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.4
2019 Dual-Stream Interactive Networks for No-Reference Stereoscopic Image Quality Assessment
abstract
The goal of objective stereoscopic image quality assessment (SIQA) is to predict the human perceptual quality of stereoscopic/3D images automatically and accurately. Compared with traditional 2D image quality assessment, the quality assessment of stereoscopic images is more challenging because of complex binocular vision mechanisms and multiple quality dimensions. In this paper, inspired by the hierarchical dual-stream interactive nature of the human visual system, we propose a stereoscopic image quality assessment network (StereoQA-Net) for no-reference stereoscopic image quality assessment. The proposed StereoQA-Net is an end-to-end dual-stream interactive network containing left and right view sub-networks, where the interaction of the two sub-networks exists in multiple layers. We evaluate our method on the LIVE stereoscopic image quality databases. The experimental results show that our proposed StereoQA-Net outperforms state-of-the-art algorithms on both symmetrically and asymmetrically distorted stereoscopic image pairs of various distortion types. In a more general case, the proposed StereoQA-Net can effectively predict the perceptual quality of local regions. In addition, cross-dataset experiments also demonstrate the generalization ability of our algorithm.
Wei Zhou 0021, Zhibo Chen 0001, Weiping Li 0003
IEEE Trans. Image Process.3
2018 Video-Based Sign Language Recognition Without Temporal Segmentation
abstract
Millions of hearing impaired people around the world routinely use some variants of sign languages to communicate, thus the automatic translation of a sign language is meaningful and important. Currently, there are two sub-problems in Sign Language Recognition (SLR), i.e., isolated SLR that recognizes word by word and continuous SLR that translates entire sentences. Existing continuous SLR methods typically utilize isolated SLRs as building blocks, with an extra layer of preprocessing (temporal segmentation) and another layer of post-processing (sentence synthesis). Unfortunately, temporal segmentation itself is non-trivial and inevitably propagates errors into subsequent steps. Worse still, isolated SLR methods typically require strenuous labeling of each word separately in a sentence, severely limiting the amount of attainable training data. To address these challenges, we propose a novel continuous sign recognition framework, the Hierarchical Attention Network with Latent Space (LS-HAN), which eliminates the preprocessing of temporal segmentation. The proposed LS-HAN consists of three components: a two-stream Convolutional Neural Network (CNN) for video feature representation generation, a Latent Space (LS) for semantic gap bridging, and a Hierarchical Attention Network (HAN) for latent space based recognition. Experiments are carried out on two large scale datasets. Experimental results demonstrate the effectiveness of the proposed framework.
Jie Huang 0011, Wengang Zhou 0001, Qilin Zhang 0004, Houqiang Li, Weiping Li 0003
AAAI5
2018 SDM: Semantic Distortion Measurement for Video Encryption
abstract
Semantic information is important in video encryption. However, existing image quality assessment (IQA) methods, such as the peak signal to noise ratio (PSNR), are still widely applied to measure the encryption security. Generally, these traditional IQA methods aim to evaluate the image quality from the perspective of visual signal rather than semantic information. In this paper, we propose a novel semantic-level full-reference image quality assessment (FR-IQA) method named Semantic Distortion Measurement (SDM) to measure the degree of semantic distortion for video encryption. Then, based on a semantic saliency dataset, we verify that the proposed SDM method outperforms state-of-the-art algorithms. Furthermore, we construct a Region Of Semantic Saliency (ROSS) video encryption system to demonstrate the effectiveness of our proposed SDM method in the practical application.
Yongquan Hu, Wei Zhou 0021, Shuxin Zhao, Zhibo Chen 0001, Weiping Li 0003
FG5
2018 Decouple and Stretch: A Boost to Channel Pruning
abstract
Deep Neural Networks (DNNs) have shown superior performance on a variety of artificial intelligence problems. Reducing the resource usage of DNN is critical to adding intelligence on Internet of Things (IoT) devices. Channel pruning based network compression shows effective reduction simultaneously on storage, memory and computation without specialized software on general platforms. But limited by pruning flexibility, channel pruning methods have relatively low compression rate for a given target performance. In this paper, we demonstrate that channel pruning becomes more robust to decision errors by reducing the granularity of filters. Then we propose a Decouple and Stretch (DS) scheme to enhance channel pruning. Under this scheme, each filter in a specific layer is decoupled into two small spatial-wise filters, and the spatial-wise filters are stretched into two successive convolutional layers. Our scheme obtains up to 49% improvement on compression and 35% improvement on acceleration. To further demonstrate hardware compatibility, we deploy pruned networks on the FPGA, and the network produced by Decouple and Stretch scheme is more hardware-friendly with latency reduced by 42%.
Zhen Chen 0013, Sen Liu 0001, Jun Xia 0003, Weiping Li 0003
IPCCC5
2018 Automating Robotic Furniture with A Collaborative Vision-based Sensing Scheme
abstract
Automating teleoperated robots is an essential task for transforming human-robot interaction design into practical applications. In this paper, we present a Collaborative Vision-based Sensing Scheme (CVSS) for automating mobile robotic furniture in the household environment. Using multiple cameras to perceive users' spatial information and capture their body postures respectively, our sensing scheme can provide formerly teleoperated robots with sufficient situational awareness of the users' surroundings and enable them to understand the interactive willingness of humans. As an application instance, we introduce the design of a furniture-type robot, called Automan, whose prototype is a teleoperated mechanical ottoman reported in [1]. Utilizing CVSS, we enable Automan the similar functions as teleoperated ottoman to offer and withdraw services for humans. To evaluate our automation method, we conducted the subjective experiments with 20 participants to verify the effectiveness of it in comparison with the teleoperated ottomans in terms of interactive experience. And the result of paired samples t test indicates that the participants can't significantly distinguish whether the interactive ottoman is autonomous or teleoperated (p≫0.05). Furthermore, we also explored and analyzed users' satisfaction for different behavior styles between autonomous and teleoperated ottomans.
Zhizheng Zhang 0004, Zhibo Chen 0001, Weiping Li 0003
RO-MAN3
2018 Blind Stereoscopic Video Quality Assessment: From Depth Perception to Overall Experience
abstract
Stereoscopic video quality assessment (SVQA) is a challenging problem. It has not been well investigated on how to measure depth perception quality independently under different distortion categories and degrees, especially exploit the depth perception to assist the overall quality assessment of 3D videos. In this paper, we propose a new depth perception quality metric (DPQM) and verify that it outperforms existing metrics on our published 3D video extension of High Efficiency Video Coding (3D-HEVC) video database. Furthermore, we validate its effectiveness by applying the crucial part of the DPQM to a novel blind stereoscopic video quality evaluator (BSVQE) for overall 3D video quality assessment. In the DPQM, we introduce the feature of auto-regressive prediction-based disparity entropy (ARDE) measurement and the feature of energy weighted video content measurement, which are inspired by the free-energy principle and the binocular vision mechanism. In the BSVQE, the binocular summation and difference operations are integrated together with the fusion natural scene statistic measurement and the ARDE measurement to reveal the key influence from texture and disparity. Experimental results on three stereoscopic video databases demonstrate that our method outperforms state-of-the-art SVQA algorithms for both symmetrically and asymmetrically distorted stereoscopic video pairs of various distortion types.
Zhibo Chen 0001, Wei Zhou 0021, Weiping Li 0003
IEEE Trans. Image Process.3
2018 Distortion Bounds for Source Broadcast Problems
abstract
This paper investigates the joint source-channel coding problem of sending a memoryless source over a memoryless broadcast channel. An inner bound and several outer bounds on the admissible distortion region are derived, which, respectively, generalize and unify several existing bounds. As a consequence, we also obtain an inner bound and an outer bound for the degraded broadcast channel case. When specialized to the Gaussian or binary source broadcast, the inner bound and outer bound not only recover the best known inner bound and outer bound in the literature but also generate some new results. Besides, we also extend the inner bound and outer bounds to the Wyner-Ziv source broadcast problem, i.e., source broadcast with side information available at decoders. Some new bounds are obtained when specialized to the Wyner-Ziv Gaussian and Wyner-Ziv binary cases.
Lei Yu 0003, Houqiang Li, Weiping Li 0003
IEEE Trans. Inf. Theory3
2017 Surveillance video coding with dynamic textural background detection
abstract
Texture scenes like flickering flames, swaying tree branches, flowing water exhibit a complex stochastic motion character. It presents a great challenge to compress these dynamic texture efficiently even with the state-of-the-art video encoder. Furthermore, these contents only contain a little helpful information in surveillance analysis. In this paper, we propose an approach for compressing the dynamic textures in the surveillance video. In the proposed scheme, the dynamic texture contents are detected by the histogram of motion direction (HMD) algorithm, and then removed at the encoder, these dynamic texture contents will be restored at the decoder directly. Objective and subjective results are presented, demonstrating that the proposed approach provides about 8.7% bitrate saving with visually plausible dynamic textures in comparison with High Efficiency Video Coding (HEVC).
Fangdong Chen, Dong Liu 0002, Zhibo Chen 0001, Weiping Li 0003
ICIP5
2017 Video restoration based on a novel second order nonlocal total variation model
Zhenbo Lu, Qing Ling 0001, Houqiang Li, Weiping Li 0003
Signal Process.4
2017 Source-Channel Secrecy for Shannon Cipher System
abstract
Recently, a secrecy measure based on list-reconstruction has been proposed, in which a wiretapper is allowed to produce a list of 2mRLreconstruction sequences and the secrecy is measured by the minimum distortion over the entire list. In this paper, we show that this list secrecy problem is equivalent to the one with secrecy measured by a new quantity lossy equivocation, which is proved to be the minimum optimistic one-achievable source coding rate (the minimum coding rate needed to reconstruct the source within target distortion with positive probability for infinitely many blocklengths) of the source with the wiretapped signal as two-sided information, and also can be seen as a lossy extension of conventional equivocation. Upon this (or list) secrecy measure, we study source-channel secrecy problem in the discrete memoryless Shannon cipher system with noisy wiretap channel. Two inner bounds and an outer bound on the achievable region of secret key rate, list rate, wiretapper distortion, and distortion of legitimate user are given. The inner bounds are derived by using uncoded scheme and (operationally) separate scheme, respectively. Thanks to the equivalence between lossy-equivocation secrecy and list secrecy, information spectrum method is leveraged to prove the outer bound. As special cases, the admissible region for the case of degraded wiretap channel or lossless communication for legitimate user has been characterized completely. For both these two cases, separate scheme is proved to be optimal. Interestingly, however, separation indeed suffers performance loss for other certain cases. Besides, we also extend our results to characterize the achievable region for Gaussian communication case. As a side product, optimistic lossy source coding has also been addressed.
Lei Yu 0003, Houqiang Li, Weiping Li 0003
IEEE Trans. Inf. Theory3
2017 Joint Source-Channel Secrecy Using Uncoded Schemes: Towards Secure Source Broadcast
abstract
This paper investigates a joint source-channel secrecy problem for the Shannon cipher broadcast system. We suppose list secrecy is applied, i.e., a wiretapper is allowed to produce a list of reconstruction sequences and the secrecy is measured by the minimum distortion over the entire list. For discrete communication cases, we propose a permutation-based uncoded scheme, which cascades a random permutation with a symbol-by-symbol mapping. Using this scheme, we derive an inner bound for the admissible region of secret key rate, list rate, wiretapper distortion, and distortions of legitimate users. For the converse part, we easily obtain an outer bound for the admissible region from an existing result. Comparing the outer bound with the inner bound shows that the proposed scheme is optimal under certain conditions. Besides, we extend the proposed scheme to the scalar and vector Gaussian communication scenarios, and characterize the corresponding performance as well. For these two cases, we also propose another uncoded scheme, orthogonal-transform-based scheme, which achieves the same performance as the permutation-based scheme. Interestingly, by introducing the random permutation or the random orthogonal transform into the traditional uncoded scheme, the proposed uncoded schemes, on one hand, provide a certain level of secrecy, and on the other hand, do not lose any performance in terms of the distortions for legitimate users.
Lei Yu 0003, Houqiang Li, Weiping Li 0003
IEEE Trans. Inf. Theory3
2016 Comparative Deep Learning of Hybrid Representations for Image Recommendations
abstract
In many image-related tasks, learning expressive and discriminative representations of images is essential, and deep learning has been studied for automating the learning of such representations. Some user-centric tasks, such as image recommendations, call for effective representations of not only images but also preferences and intents of users over images. Such representations are termed hybrid and addressed via a deep learning approach in this paper. We design a dual-net deep network, in which the two sub-networks map input images and preferences of users into a same latent semantic space, and then the distances between images and users in the latent space are calculated to make decisions. We further propose a comparative deep learning (CDL) method to train the deep network, using a pair of images compared against one user to learn the pattern of their relative distances. The CDL embraces much more training data than naive deep learning, and thus achieves superior performance than the latter, with no cost of increasing network complexity. Experimental results with real-world data sets for image recommendations have shown the proposed dual-net network and CDL greatly outperform other state-of-the-art image recommendation solutions.
Chenyi Lei, Dong Liu 0002, Weiping Li 0003, Zhengjun Zha, Houqiang Li
CVPR3
2016 Popularity-driven content caching
abstract
This paper presents a novel cache replacement method — Popularity-Driven Content Caching (PopCaching). PopCaching learns the popularity of content and uses it to determine which content it should store and which it should evict from the cache. Popularity is learned in an online fashion, requires no training phase and hence, it is more responsive to continuously changing trends of content popularity. We prove that the learning regret of PopCaching (i.e., the gap between the hit rate achieved by PopCaching and that by the optimal caching policy with hindsight) is sublinear in the number of content requests. Therefore, PopCaching converges fast and asymptotically achieves the optimal cache hit rate. We further demonstrate the effectiveness of PopCaching by applying it to a movie.douban.com dataset that contains over 38 million requests. Our results show significant cache hit rate lift compared to existing algorithms, and the improvements can exceed 40% when the cache capacity is limited. In addition, PopCaching has low complexity.
Suoheng Li, Jie Xu 0001, Mihaela van der Schaar, Weiping Li 0003
INFOCOM4
2016 Distortion bounds for source broadcast over degraded channel
abstract
This paper investigates the joint source-channel coding problem of sending a memoryless source over a memoryless degraded broadcast channel. An inner bound and an outer bound on the achievable distortion region are derived, which respectively generalize and unify several existing bounds. Moreover, when specialized to Gaussian source broadcast or binary source broadcast, the inner bound and outer bound could recover the best known inner bound and outer bound in the literature. Besides, the inner bound and outer bound are also extended to Wyner-Ziv source broadcast problem, i.e., source broadcast with degraded side information available at decoders. Some new bounds are obtained when specialized to Wyner-Ziv Gaussian case and Wyner-Ziv binary case.
Lei Yu 0003, Houqiang Li, Weiping Li 0003
ISIT3
2016 Respiration Motion State Estimation on 4D CT Rib Cage Images
Wengang Zhou 0001, Weiping Ding 0002, Houqiang Li, Weiping Li 0003
MMM (1)5
2016 3D-HEVC visual quality assessment: Database and bitstream model
abstract
Visual Quality Assessment of 3D/stereoscopic video (3D VQA) is significant for both quality monitoring and optimization of the existing 3D video services. In this paper, we build a 3D video database based on the latest 3D-HEVC video coding standard, to investigate the relationship among video quality, depth quality, and overall quality of experience (QoE) of 3D/stereoscopic video. We also analyze the pivotal factors to the video and depth qualities. Moreover, we develop a No-Reference 3D-HEVC bitstream-level objective video quality assessment model, which utilizes the key features extracted from the 3D video bitstreams to assess the perceived quality of the stereoscopic video. The model is verified to be effective on our database as compared with widely used 2D Full-Reference quality metrics as well as a state-of-the-art 3D FR pixel-level video quality metric.
Wei Zhou 0021, Ning Liao, Zhibo Chen 0001, Weiping Li 0003
QoMEX4
2016 No-reference image quality assessment based on global and local content perception
abstract
Existing no-reference image quality assessment (NR-IQA) methods mainly focus on designing the low-level features related to image degradation. However, the evaluation of image quality is the human visual perception of image content, involving the integrated analysis of global high-level semantics and local low-level characteristics. From this perspective, we propose a NR-IQA framework based on global and local content perception. We adopt the deep convolutional neural network (DCNN) to extract the semantic feature implied in global image content. The perception of visual quality associated with local content utilizes the visual attention and filtering mechanisms of human visual system. The overall image quality is estimated by combining the semantic and local characteristic features generated from the perceptions. Experimental results on the LIVE IQA database demonstrate that our method is superior to the state-of-the-art NR-IQA algorithms and competitive to the popular full-reference IQA methods. Further experiments on the TID2008 dataset show that the proposed approach is robust for various kinds of distortion types.
Cuirong Sun, Houqiang Li, Weiping Li 0003
VCIP3
2016 Comments on "Approximate Characterizations for the Gaussian Source Broadcast Distortion Region"
abstract
Recently, Tianet al.[1]considered joint source-channel coding of transmitting a Gaussian source over$K$-user Gaussian broadcast channel, and derived an outer bound on the admissible distortion region. In[1], they stated “due to its nonlinear form, it appears difficult to determine whether it is always looser than the trivial outer bound in all distortion regimes with bandwidth compression”. However, in this correspondence we solve this problem and prove that for the bandwidth expansion case (with$K\geq 2$), this outer bound is strictly tighter than the trivial outer bound with each user being optimal in the point-to-point setting; while for the bandwidth compression or bandwidth match case, this outer bound actually degenerates to the trivial outer bound. Therefore, our results imply that on one hand, the outer bound given in[1]is nontrivial only for Gaussian broadcast communication ($K\geq 2$) with bandwidth expansion; on the other hand, unfortunately, no nontrivial outer bound exists so far for Gaussian broadcast communication ($K\geq 2$) with bandwidth compression.
Lei Yu 0003, Houqiang Li, Weiping Li 0003
IEEE Trans. Inf. Theory3
2016 Social Diffusion Analysis With Common-Interest Model for Image Annotation
abstract
Automatic image annotation has been extensively studied, mostly from a content-based approach, whose effectiveness is restricted by the “semantic gap” between low-level image features and semantic annotations, and by the irrelevance of annotations to image content. We propose a social diffusion analysis approach to image annotation, which exploits abundant social diffusion records about how images are disseminated within online social networks. Specifically, we propose a common-interest model to analyze social diffusion records, with the assumption that the diffusion pattern of an image in social networks is highly related to the relevance between image annotations and user preferences. In our proposed model, user preferences are represented as common interests of pairwise users instead of individual user interests. We find the notion of common interests not only facilitates the analysis of social diffusion patterns, but also leads to more accurate profiling of user preferences compared to individual interests. Based on the common-interest model, we design an image annotation framework via social diffusion analysis, which consists of the mining of common interests from social diffusion records, the feature extraction from diffusion graphs and common interests, and the automatic annotation by the learning-to-rank method. Experimental results on real-world data sets show that our proposed common-interest based approach outperforms individual-interest based methods, and also achieves superior performance than state-of-the-art content-based image annotation methods.
Chenyi Lei, Dong Liu 0002, Weiping Li 0003
IEEE Trans. Multim.3
2016 Trend-Aware Video Caching Through Online Learning
abstract
This paper presents Trend-Caching, a novel cache replacement method that optimizes cache performance according to the trends of video content. Trend-Caching explicitly learns the popularity trend of video content and uses it to determine which video it should store and which it should evict from the cache. Popularity is learned in an online fashion and requires no training phase, hence it is more responsive to continuously changing trends of videos. We prove that the learning regret of Trend-Caching (i.e., the gap between the hit rate achieved by Trend-Caching and that by the optimal caching policy with hindsight) is sublinear in the number of video requests, thereby guaranteeing both fast convergence and asymptotically optimal cache hit rate. We further validate the effectiveness of Trend-Caching by applying it to a movie.douban.com dataset that contains over 38 million requests. Our results show significant cache hit rate lift compared to existing algorithms, and the improvements can exceed 40% when the cache capacity is limited. Furthermore, Trend-Caching has low complexity.
Suoheng Li, Jie Xu 0001, Mihaela van der Schaar, Weiping Li 0003
IEEE Trans. Multim.4
2015 A Bayesian adaptive weighted total generalized variation model for image restoration
abstract
In recent years, the Total Generalized Variation (TGV) model has received lots of attention in image processing community. Though this model can restore image with natural intensity transitions, its spatial identical parameter setting limits its performance. In this paper, we propose a novel Adaptive Weighted Total Generalized Variation model for image restoration. We analyze the TGV model from Bayesian Probability view and derive a novel adaptive parameter calculation scheme for it, exploiting the image's self-similarity. Experiment results on image deblurring and reconstruction show that by adapting the parameters in TGV model to image contents, the proposed model can restore image's edges and details well and achieve significant improvement over state of the art variational based models.
Zhenbo Lu, Houqiang Li, Weiping Li 0003
ICIP3
2015 Sign Language Recognition using 3D convolutional neural networks
abstract
Sign Language Recognition (SLR) targets on interpreting the sign language into text or speech, so as to facilitate the communication between deaf-mute people and ordinary people. This task has broad social impact, but is still very challenging due to the complexity and large variations in hand actions. Existing methods for SLR use hand-crafted features to describe sign language motion and build classification models based on those features. However, it is difficult to design reliable features to adapt to the large variations of hand gestures. To approach this problem, we propose a novel 3D convolutional neural network (CNN) which extracts discriminative spatial-temporal features from raw video stream automatically without any prior knowledge, avoiding designing features. To boost the performance, multi-channels of video streams, including color information, depth clue, and body joint positions, are used as input to the 3D CNN in order to integrate color, depth and trajectory information. We validate the proposed model on a real dataset collected with Microsoft Kinect and demonstrate its effectiveness over the traditional approaches based on hand-crafted features.
Jie Huang 0011, Wengang Zhou 0001, Houqiang Li, Weiping Li 0003
ICME4
2015 Image deblocking via group sparsity optimization
abstract
Block-wise compressed image often suffers from the blocking artifacts. In this paper, we propose a novel deblocking scheme for compressed image, by combining image's sparse property and its self-similarity together, called group sparsity optimization. Instead of processing each image patch individually, in the proposed scheme, similar patches in one group are required to be well-represented on learned dictionary collaboratively, using group sparsity regularization. The group sparsity not only imposes every patch's representation to be sparse, bus also requires patches' coefficients in the group share the similar pattern. The experiment results on standard test images demonstrate that our scheme can improve the PSNR of the compressed images by an average of 1.25 dB, and outperform state of the art deblocking approaches.
Zhenbo Lu, Houqiang Li, Weiping Li 0003
ISCAS3
2015 Wireless Cooperative Video Coding Using a Hybrid Digital-Analog Scheme
abstract
Wireless video broadcast/multicast and mobile video communication pose a challenge to the conventional video transmission strategy (which applies separate digital source coding and digital channel coding in a point-to-point communication system) and, therefore, cooperative communication has been proposed to improve the received quality of the receivers with bad channels and the robustness to fading (mobility). In this paper we propose a novel wireless cooperative video coding (WCVC) framework. Specifically, we present a cooperative joint source-channel coding scheme that is based on hybrid digital-analog coding and integrates the advantages of digital coding and analog coding. Compared with most state-of-the-art cooperative video delivery methods, no matter for cooperation scenario or noncooperation scenario, the proposed scheme can avoid the staircase effect and realize continuous quality scalability on condition that the channel quality is within the expected range, and it has strong adaptability to channel variation with higher coding efficiency and better fairness among all receivers. Therefore, it is very suitable for wireless cooperative/noncooperative video broadcast/multicast transmission and cooperative/noncooperative mobile video applications. The experimental results show that the proposed WCVC outperforms SoftCast (which is a new analog scheme) and SVC + hierarchical modulation (HM) (which combines H.264/SVC codec and HM technique) no matter the noncooperative scenario or for cooperative scenario, which verify the effectiveness of our proposed WCVC framework.
Lei Yu 0003, Houqiang Li, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.3
2014 Wireless Scalable Video Coding Using a Hybrid Digital-Analog Scheme
abstract
Wireless video broadcast/multicast and mobile video communication pose a challenge to the conventional video transmission scheme (which consists of separate digital source coding and digital channel coding). The reason is that the separate coding scheme is based on nonscalable coding design, and this unavoidably leads to cliff effect as well as limits the ability to support multiple users with diverse channel conditions. In this paper, we propose a novel wireless scalable video coding (WSVC) framework. Specifically, we present a hybrid digital-analog (HDA) joint source-channel coding (JSCC) scheme that integrates the advantages of digital coding and analog coding. Moreover, the proposed JSCC is able to broadcast one video with different resolutions to fit various devices with different display resolutions. Compared to most state-of-the-art video delivery methods, it avoids the staircase effect and realizes continuous quality scalability (CQS) on condition that the channel quality is within the expected range, and it has strong adaptability to channel variation with higher coding efficiency and better fairness among all receivers. Therefore, it is very suitable for wireless video broadcast/multicast transmission and mobile video applications. The experimental results show that for broadcasting/multicasting the videos with CIF and QCIF resolutions the proposed WSVC outperforms SoftCast (which is a new analog scheme) average 0.60-5.90 dB and 3.39-9.97 dB respectively, outperforms SVC+HM (which combines H.264/SVC codec and hierarchical modulation technique) average 3.87-9.13 dB and 0.27-10.47 dB respectively, and outperforms DCast (which is an up-to-date video delivery scheme) about 0.2-3.3 dB for the video with CIF resolution. The experimental results verify the effectiveness of our proposed WSVC framework.
Lei Yu 0003, Houqiang Li, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.3
2013 Noise reduction for hyperspectral images based on structural sparse and low-rank matrix decomposition
abstract
In this paper, a noise reduction approach for hyperspectral images (HSIs) is presented. Due to the assorted noise sources of HSIs, it seems difficult to describe the noise in a concise manner. Commonly, noise reduction algorithms are dedicated to a certain kind of noise, such as random or striping noise. Most of them in addition have somewhat idealized hypotheses. For example, the random noise is white or signal-independent, or the observed scene is spatially homogeneous or quasi-homogeneous. Thus a practically efficient and universal denoising method is preferred. Thanks to the low-rank characteristic of HSI signal, and the structural sparsity of HSI noise, we draw inspiration from low-rank matrix decomposition and the emerging mixed norm, to propose a method dealing with various patterns of noise simultaneously. Both simulated and real data experiments show the effectiveness of the proposed approach.
Zhenbo Lu, Qingbo Lu, Houqiang Li, Weiping Li 0003
IGARSS5
2013 Hybrid digital-analog scheme for video transmission over wireless
abstract
In this paper, we propose a novel wireless video transmission scheme named HDA-Cast, which is a hybrid digital-analog (HDA) coding scheme that integrates the advantages of digital coding and analog coding. Relative to most state-of-the-art video transmission methods, it avoids the “cliff effect” provided that the channel quality is within the expected range, gives better fairness among all receivers for multicast, and has strong adaption to channel variation. The evaluation results show that our HDA-Cast is 3.5-9.6 dB better than the SoftCast which is an up-to-date analog scheme. Owing to its strong adaption to channel variation, it can be regarded as a kind of wireless scalable video coding (WSVC).
Lei Yu 0003, Houqiang Li, Weiping Li 0003
ISCAS3
2013 Detection of Blotch and Scratch in Video Based on Video Decomposition
abstract
In old video restoration, automatic detection of common defects, e.g., scratches and blotches, has always been emphasized. While prior thoughts mainly focus on detecting blotches and linear, vertical scratches separately, this paper contributes to a more generalized and challenging issue: simultaneous detection of blotches and complex scratches in video, with much less knowledge of them. We investigate the characteristics of blotches and scratches in space and time domain, and propose a novel detection method based on two main steps: cartoon-texture decomposition in the space domain and content-defect separation in the time domain. We then formulate it into convex optimization problems and develop corresponding algorithms. The experiment results demonstrate that the proposed method is of high detection accuracy, verifying the effectiveness of our detection via a video decomposition method.
Houqiang Li, Zhenbo Lu, Zhangyang Wang, Qing Ling 0001, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.5
2013 Robust Temporal-Spatial Decomposition and Its Applications in Video Processing
abstract
In this paper, we propose a robust temporal-spatial decomposition (RTSD) model and discuss its applications in video processing. A video sequence usually possesses high correlations among and within its frames. Fully exploiting the temporal and spatial correlations enables efficient processing and better understanding of the video sequence. Considering that the video sequence typically contains slowly changing background and rapidly changing foreground as well as noise, we propose to decompose the video frames into three parts: the temporal-spatially correlated part, the feature compensation part, and the sparse noise part. Accordingly, the decomposition problem can be formulated as the minimization of a convex function, which consists of a nuclear norm, a total variation (TV)-like norm, and anl1norm. Since the minimization is nontrivial to handle, we develop a two-stage strategy to solve this decomposition problem, and discuss different alternatives to fulfil each stage of decomposition. The RTSD model treats video frames as a unity from both the temporal and spatial point of view, and demonstrates robustness to noise and certain background variations. Experiments on video denoising and scratch detection applications verify the effectiveness of the proposed RTSD model and the developed algorithms.
Zhangyang Wang, Houqiang Li, Qing Ling 0001, Weiping Li 0003
IEEE Trans. Circuits Syst. Video Technol.4
2012 An adaptive down-sampling based video coding with hybrid super-resolution method
abstract
It has been proven that performance of video coding at low bit rates can be improved by down-sampling a video before compression and then using super-resolution to up-sample it after decompression. Such techniques are especially important for limited bandwidth communications. In this paper we propose an adaptively down-sampling based coding (DBC) method which performs rate distortion (RD) optimization to determine the coding structure between regular coding and down-sampling coding. In order to restore the original resolution of down-sampling coded video signals, a hybrid super-resolution (SR) algorithm which combines motion compensation (MC) based SR and wiener filter based SR is used. Experimental results show that our method has improvement both in rate-distortion performance and perceived visual quality at low bit rate.
Zeng Hu, Houqiang Li, Weiping Li 0003
ISCAS3
2012 Mixed Gaussian-impulse video noise removal via temporal-spatial decomposition
abstract
This paper presents a novel denoising scheme for video sequences corrupted by mixed Gaussian-impulse noise. From a global viewpoint, such a video sequence contains three parts: temporal-spatially correlated video content, uncorrelated dense Gaussian noise, and uncorrelated sparse impulse noise. This fact motivates us to formulate the mixed Gaussian-impulse noise removal task as a temporal-spatial decomposition problem, which amounts to a convex program. A two-stage algorithm is developed to solve this problem efficiently. Effectiveness of the proposed algorithm on mixed Gaussian-impulse noise removal is validated through experiments. The results are satisfactory in both visual quality and PSNR values, while very few prior knowledge of noise statistic is required compared to most state-of-the-art methods.
Zhangyang Wang, Houqiang Li, Qing Ling 0001, Weiping Li 0003
ISCAS4
2011 Smoothing rate control for multiple video streams using game theory
abstract
To fairly allocate network bandwidth for different streams and to achieve a smooth quality between different video sequences are two key issues in multi-stream applications. In this paper, a new joint rate control scheme for multiple video streams based on game theory is proposed. We achieve a fair rate allocation in video quality between different sequences by Nash Bargaining theory under the constraints of buffer and bandwidth. Also, by grouping all the different sequences' frames within the same timeslot into a coalition, we can avoid the fluctuation in a sequence by bargaining between current coalition and remaining coalitions. Some experiments have been conducted to verify the proposed method.
Meng Liu 0006, Houqiang Li, Weiping Li 0003
ISCAS3