Zhaoyang Zhang 0002

dblp:81/6236-2 · DBLP profile ↗
← Back
41ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0009-3611-7060ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2Computer networks · 1
YearPublicationVenuePosition
2025 FGMIA: Feature-Guided Model Inversion Attacks Against Face Recognition Models
abstract
Model Inversion Attacks (MIAs) against face recognition systems aim to reconstruct facial images of specific individuals from the recognition models. Existing MIA approaches commonly optimize the latent variables of Generative Adversarial Networks (GANs) iteratively, which can result in non-smooth optimizations due to the complexity and entanglement of latent space. Furthermore, the optimization guided by the target model’s gradients may generate high-confidence images with poor perceptual similarity to the target class. This paper introduces a novel perspective by reformulating the inversion attack as a conditional data distribution learning task. Based on this, we propose a Feature-Guided Model Inversion Attack (FGMIA), which learns the facial data distribution and integrates feature guidance as a conditional signal. Specifically, we treat the deconstructed target model as a feature encoder, which provides guidance during the training of a specialized feature-guided diffusion model. During the attack, feature encodings implicit in the target model are extracted and utilized to guide the reconstruction of private data. Extensive experiments demonstrate that FGMIA accurately reconstructs private data from face recognition models and significantly improves evaluation accuracy and perceptual similarity compared to state-of-the-art methods while maintaining comparable target confidence scores. Our code is available at https://github.com/MMCTTT/FGMIA_codes.
Shen Wang 0004, Guopu Zhu, Zhaoyang Zhang 0002, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2024 Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy
abstract
The widespread collection and analysis of private speech signals have become increasingly prevalent, raising significant privacy concerns. To protect speech signals from unauthorized analysis, adversarial attack methods for deceiving speaker recognition models have been proposed. While a few of these methods are specifically designed for real-time protection of speech signals, they introduce significant delays that can severely impact speech communication when applied to streaming speech data. In this paper, we present a novel approach that aims to offer real-time protection for speech signals without delays. By utilizing observed data only, we generate initial adversarial seed perturbations and refine them to obtain the necessary adversarial perturbations predicted for adjacent unobserved signals. This refinement process is conducted via a proposed model called PAPG. On the basis of perturbation prediction, we develop a streaming audio processing framework that generates perturbations in synchronization with the playback of the original signal, effectively eliminating delays. The experimental results demonstrate that under the proposed attack, the average Top-1 accuracy of various advanced speaker recognition methods is reduced by 89%, and the average equal error rate (EER) increases to 36%. Remarkably, these results are achieved without delays while maintaining superior perceptual quality.
Zhaoyang Zhang 0002, Shen Wang 0004, Guopu Zhu, Dechen Zhan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.1
2023 Query-Efficient Adversarial Attack With Low Perturbation Against End-to-End Speech Recognition Systems
abstract
With the widespread use of automated speech recognition (ASR) systems in modern consumer devices, attack against ASR systems have become an attractive topic in recent years. Although related white-box attack methods have achieved remarkable success in fooling neural networks, they rely heavily on obtaining full access to the details of the target models. Due to the lack of prior knowledge of the victim model and the inefficiency in utilizing query results, most of the existing black-box attack methods for ASR systems are query-intensive. In this paper, we propose a new black-box attack called the Monte Carlo gradient sign attack (MGSA) to generate adversarial audio samples with substantially fewer queries. It updates an original sample based on the elements obtained by a Monte Carlo tree search. We attribute its high query efficiency to the effective utilization of the dominant gradient phenomenon, which refers to the fact that only a few elements of each origin sample have significant effect on the output of ASR systems. Extensive experiments are performed to evaluate the efficiency of MGSA and the stealthiness of the generated adversarial examples on the DeepSpeech system. The experimental results show that MGSA achieves 98% and 99% attack success rates on the LibriSpeech and Mozilla Common Voice datasets, respectively. Compared with the state-of-the-art methods, the average number of queries is reduced by 27% and the signal-to-noise ratio is increased by 31%.
Shen Wang 0004, Zhaoyang Zhang 0002, Guopu Zhu, Xinpeng Zhang 0001, Yicong Zhou, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.2
2015 Depth upsampling method via Markov random fields without edge-misaligned artifacts
abstract
Recently, the widely use of time-of-flight sensors captures depth information for dynamic scenes in real time, which promotes the developing of many 3D image or video processing applications. However, such depth maps are noisy and have low resolutions. In this paper, we propose an edge-based depth map super-resolution method via solving a labeling optimization problem in MRF. The inputs are low quality depth map and the according high-resolution color image. The proposed method not only avoids the texture-copy artifacts, but also preserves the edges of depth which do not exist in the color image. We compare our algorithm with the state of the art on the benchmark dataset. The experimental results prove the validity and robustness of our approach.
Yifan Zuo 0001, Ping An 0001, Zhaoyang Zhang 0002
ICIP4
2015 Fast TU size decision algorithm for HEVC encoders using Bayesian theorem detection
Liquan Shen, Zhaoyang Zhang 0002, Xinpeng Zhang 0001, Ping An 0001, Zhi Liu 0003
Signal Process. Image Commun.2
2015 A 3D-HEVC Fast Mode Decision Algorithm for Real-Time Applications
abstract
3D High Efficiency Video Coding (3D-HEVC) is an extension of the HEVC standard for coding of multiview videos and depth maps. It inherits the same quadtree coding structure as HEVC for both components, which allows recursively splitting into four equal-sized coding units (CU). One of 11 different prediction modes is chosen to code a CU in inter-frames. Similar to the joint model of H.264/AVC, the mode decision process in HM (reference software of HEVC) is performed using all the possible depth levels and prediction modes to find the one with the least rate distortion cost using a Lagrange multiplier. Furthermore, both motion estimation and disparity estimation need to be performed in the encoding process of 3D-HEVC. Those tools achieve high coding efficiency, but lead to a significant computational complexity. In this article, we propose a fast mode decision algorithm for 3D-HEVC. Since multiview videos and their associated depth maps represent the same scene, at the same time instant, their prediction modes are closely linked. Furthermore, the prediction information of a CU at the depth level X is strongly related to that of its parent CU at the depth level X-1 in the quadtree coding structure of HEVC since two corresponding CUs from two neighboring depth levels share similar video characteristics. The proposed algorithm jointly exploits the inter-view coding mode correlation, the inter-component (texture-depth) correlation and the inter-level correlation in the quadtree structure of 3D-HEVC. Experimental results show that our algorithm saves 66% encoder runtime on average with only a 0.2% BD-Rate increase on coded views and 1.3% BD-Rate increase on synthesized views.
Liquan Shen, Ping An 0001, Zhaoyang Zhang 0002, Qianqian Hu, Zhengchuan Chen
ACM Trans. Multim. Comput. Commun. Appl.3
2014 Efficient depth coding in 3D video to minimize coding bitrate and complexity
Liquan Shen, Zhaoyang Zhang 0002
Multim. Tools Appl.2
2014 Adaptive Inter-Mode Decision for HEVC Jointly Utilizing Inter-Level and Spatiotemporal Correlations
abstract
High Efficiency Video Coding (HEVC) adopts the quadtree structured coding unit (CU), which allows recursive splitting into four equally sized blocks. At each depth level, it enables SKIP mode, merge mode, inter 2N × 2N, inter 2N × N, inter N × 2N, inter 2N × nU, inter 2N × nD, inter nL x 2N, inter nR × 2N, inter N × N (only available for the smallest CU), intra 2N × 2N, and intra N × N (only available for the smallest CU) in inter-frames. Similar to H.264/AVC, the mode decision process in HEVC is performed using all the possible depth levels (or CU sizes) and prediction modes to find the one with the least rate distortion (RD) cost using Lagrange multiplier. This achieves the highest coding efficiency, but leads to a very high computational complexity. Since the optimal prediction mode is highly content dependent, it is not efficient to use all the modes. In this paper, we propose a fast inter-mode decision algorithm for HEVC by jointly using the inter-level correlation of quadtree structure and the spatiotemporal correlation. There exist strong correlations of the prediction mode, the motion vector and RD cost between different depth levels and between spatially temporally adjacent CUs. We statistically analyze the prediction mode distribution at each depth level and the coding information correlation among the adjacent CUs. Based on the analysis results, three adaptive inter-mode decision strategies are proposed including early SKIP mode decision, prediction size correlation-based mode decision and RD cost correlation-based mode decision. Experimental results show that the proposed overall algorithm can save 49%-52% computational complexity on average with negligible loss of coding efficiency, exhibiting applicability to various types of video sequences.
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003
IEEE Trans. Circuits Syst. Video Technol.2
2014 Effective CU Size Decision for HEVC Intracoding
abstract
In high efficiency video coding (HEVC), the tree structured coding unit (CU) is adopted to allow recursive splitting into four equally sized blocks. At each depth level (or CU size), it enables up to 35 intraprediction modes, including a planar mode, a dc mode, and 33 directional modes. The intraprediction via exhaustive mode search exploited in the test model of HEVC (HM) effectively improves coding efficiency, but results in a very high computational complexity. In this paper, a fast CU size decision algorithm for HEVC intracoding is proposed to speed up the process by reducing the number of candidate CU sizes required to be checked for each treeblock. The novelty of the proposed algorithm lies in the following two aspects: 1) an early determination of CU size decision with adaptive thresholds is developed based on the texture homogeneity and 2) a novel bypass strategy for intraprediction on large CU size is proposed based on the combination of texture property and coding information from neighboring coded CUs. Experimental results show that the proposed effective CU size decision algorithm achieves a computational complexity reduction up to 67%, while incurring only 0.06-dB loss on peak signal-to-noise ratio or 1.08% increase on bit rate compared with that of the original coding in HM.
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003
IEEE Trans. Image Process.2
2013 Perceptual multiview video coding based on foveated just noticeable distortion profile in DCT domain
abstract
Recently just noticeable distortion (JND) has been highly successful in improving the video coding efficiency. Foveated JND (FJND) is an extension of the JND by further exploiting the human vision characteristic. However, there is a challenge to quickly estimate the foveation point and accurately combine the foveation factor with the spatio-temporal JND model of DCT domain for FJND. In this paper, a new FJND model in DCT domain is proposed, which adaptively searches the foveation point by exploiting the property of image signature and builds a foveation model based on the contrast threshold. Experimental results demonstrate that the proposal model remarkably reduces the complexity of FJND in searching foveation. In a number of video coding experiments, we find that, in terms of coding efficiency, the proposed perceptual coding based on FJND for multiview video coding (MVC) significantly outperforms the existing algorithm.
Xiwu Shang, Lidong Luo, Zhaoyang Zhang 0002
ICIP4
2013 A novel H.264 rate control algorithm with consideration of visual attention
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
Multim. Tools Appl.3
2013 An Effective CU Size Decision Method for HEVC Encoders
abstract
The emerging high efficiency video coding standard (HEVC) adopts the quadtree-structured coding unit (CU). Each CU allows recursive splitting into four equal sub-CUs. At each depth level (CU size), the test model of HEVC (HM) performs motion estimation (ME) with different sizes including 2N × 2N, 2N × N, N × 2N and N × N. ME process in HM is performed using all the possible depth levels and prediction modes to find the one with the least rate distortion (RD) cost using Lagrange multiplier. This achieves the highest coding efficiency but requires a very high computational complexity. In this paper, we propose a fast CU size decision algorithm for HM. Since the optimal depth level is highly content-dependent, it is not efficient to use all levels. We can determine CU depth range (including the minimum depth level and the maximum depth level) and skip some specific depth levels rarely used in the previous frame and neighboring CUs. Besides, the proposed algorithm also introduces early termination methods based on motion homogeneity checking, RD cost checking and SKIP mode checking to skip ME on unnecessary CU sizes. Experimental results demonstrate that the proposed algorithm can significantly reduce computational complexity while maintaining almost the same RD performance as the original HEVC encoder.
Liquan Shen, Zhi Liu 0003, Xinpeng Zhang 0001, Zhaoyang Zhang 0002
IEEE Trans. Multim.5
2012 Fast Segment-Based Algorithm for Multi-view Depth Map Generation
Yifan Zuo 0001, Ping An 0001, Qiuwen Zhang, Zhaoyang Zhang 0002
ICIC (2)4
2012 Content-Adaptive Motion Estimation Algorithm for Coarse-Grain SVC
abstract
A joint model of scalable video coding (SVC) uses exhaustive mode and motion searches to select the best prediction mode and motion vector for each macroblock (MB) with high coding efficiency at the cost of computational complexity. If major characteristics of a coding MB such as the complexity of the prediction mode and the motion property can be identified and used in adjusting motion estimation (ME), one can design an algorithm that can adapt coding parameters to the video content. This way, unnecessary mode and motion searches can be avoided. In this paper, we propose a content-adaptive ME for SVC, including analyses of mode complexity and motion property to assist mode and motion searches. An experimental analysis is performed to study interlayer and spatial correlations in the coding information. Based on the correlations, the motion and mode characteristics of the current MB are identified and utilized to adjust each step of ME at the enhancement layer including mode decision, search-range selection, and prediction direction selection. Experimental results show that the proposed algorithm can significantly reduce the computational complexity of SVC while maintaining nearly the same rate distortion performance as the original encoder.
Liquan Shen, Zhaoyang Zhang 0002
IEEE Trans. Image Process.2
2012 Unsupervised Salient Object Segmentation Based on Kernel Density Estimation and Two-Phase Graph Cut
abstract
In this paper, we propose an unsupervised salient object segmentation approach based on kernel density estimation (KDE) and two-phase graph cut. A set of KDE models are first constructed based on the pre-segmentation result of the input image, and then for each pixel, a set of likelihoods to fit all KDE models are calculated accordingly. The color saliency and spatial saliency of each KDE model are then evaluated based on its color distinctiveness and spatial distribution, and the pixel-wise saliency map is generated by integrating likelihood measures of pixels and saliency measures of KDE models. In the first phase of salient object segmentation, the saliency map based graph cut is exploited to obtain an initial segmentation result. In the second phase, the segmentation is further refined based on an iterative seed adjustment method, which efficiently utilizes the information of minimum cut generated using the KDE model based graph cut, and exploits a balancing weight update scheme for convergence of segmentation refinement. Experimental results on a dataset containing 1000 test images with ground truths demonstrate the better segmentation performance of our approach.
Zhi Liu 0003, Liquan Shen, Yinzhu Xue, King Ngi Ngan, Zhaoyang Zhang 0002
IEEE Trans. Multim.6
2011 An improved depth map estimation for coding and view synthesis
abstract
Inaccuracy depth estimation may influence on depth coding and virtual view rendering in the free-viewpoint television (FTV) system, an improved depth map estimation is proposed to solve the problem for coding and view synthesis. Firstly, check the consistency of initial depth, and the influence of initial miss-matches is minimized by introduction of an additional adaptive matching error selection that penalizes the unreliable matches. Then according to certain criteria, the multi-reference depth maps are merged into one disparity map to improve the quality of disparity map. Finally, a multilateral filtering is used to preserve details in the depth map and simultaneously smooth the depths in occluded areas at object boundary, less texture and discontinuity regions. Experimental results show a significant improvement of the initial input depth maps and coding efficiency, as well as a reduction of view synthesis artifacts.
Qiuwen Zhang, Ping An 0001, Liquan Shen, Zhaoyang Zhang 0002
ICIP5
2011 Efficient rendering distortion estimation for depth map compression
abstract
A depth map represents three-dimensional (3D) scene information and is used to synthesize virtual views in 3D video. Since the quality of synthesized virtual views highly depends on the quality of depth map, efficient depth compression is crucial to realize the 3D video system. However compressing depth map using existing video coding techniques yields unacceptable distortions while rendering virtual views. To solve this problem, we propose an efficient depth map compression method for the view rendering, a novel distortion metric base on view rendering distortions instead of distortion of depth map itself. First, we derive relationships between distortions in coded depth map and rendered view. Then, a region based video characteristics distortion model is proposed for precisely estimation distortion in view synthesis. Finally, experimental results have shown that 1.8 dB coding gain in terms of PSNR and subjective quality improvement of synthesized views are achieved by the proposed method.
Qiuwen Zhang, Ping An 0001, Zhaoyang Zhang 0002
ICIP4
2011 Robust real-time multi-user pupil detection and tracking under various illumination and large-scale head motion
Yuangqing Wang, Zhaoyang Zhang 0002
Comput. Vis. Image Underst.3
2011 Unsupervised image segmentation based on analysis of binary partition tree for salient object extraction
Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002
Signal Process.3
2011 Low-Complexity Mode Decision for MVC
abstract
The finalized international standard for multiview video coding (MVC) is an extension of H.264. In the joint model of MVC, variable size motion estimation (ME) and disparity estimation (DE) are introduced to achieve the highest coding efficiency with the cost of very high computational complexity. A low complexity mode decision algorithm is proposed to reduce complexity of ME and DE. An experimental analysis is performed to study inter-view correlation in the coding information such as the prediction mode and rate-distortion (RD) cost. Based on the correlation, we propose four efficient mode decision techniques, including early SKIP mode decision, adaptive early termination, fast mode size decision, and selective intra coding in inter frame. Experimental results show that the proposed algorithm can significantly reduce computational complexity of MVC while maintaining almost the same RD performance.
Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2010 Nonparametric saliency detection using kernel density estimation
abstract
This paper proposes a nonparametric saliency model based on kernel density estimation (KDE) mainly aiming at content-based applications such as salient object segmentation. A set of KDE models are constructed on the basis of regions segmented using the mean shift algorithm. For each pixel, a set of color likelihood measures to all KDE models are calculated, and then the color saliency and spatial saliency of each KDE model are evaluated based on its color distinctiveness and spatial distribution. The final saliency map is generated by combining saliency measures of KDE models and color likelihood measures of pixels. Experimental results demonstrate the better saliency detection performance of our saliency model.
Zhi Liu 0003, Yinzhu Xue, Liquan Shen, Zhaoyang Zhang 0002
ICIP4
2010 An adaptive early termination of mode decision using inter-layer correlation in scalable video coding
abstract
The scalable video coding (SVC) standard adopts the variable size motion estimation (ME) to select the best coding mode for each macroblock (MB). Although this technique achieves the highest possible coding efficiency, it results in extremely large computation complexity which obstructs SVC from the practical application. In this paper, we propose an adaptive early termination of fast mode decision algorithm in SVC. It makes use of the coding information of spatial neighbor MBs and the corresponding MBs in base layer to early terminate the mode decision procedure. Experimental results show that the proposed fast mode decision algorithm can achieve computational saving up to 67% with no significant loss of rate distortion (RD) performance.
Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002
ICIP5
2010 Unsupervised salient object segmentation from color images
abstract
This paper proposes an efficient approach for unsupervised segmentation of salient objects from color images. A set of Gaussian models are first estimated based on a pre-segmentation result of the input image, and then for each pixel, a set of normalized color likelihood measures to each Gaussian model are calculated. The color saliency and spatial saliency of Gaussian models are exploited to generate the pixel-wise saliency map. By thresholding the saliency map, the pixels are classified into object seed pixels, background seed pixels and uncertain pixels to obtain the trimap. For each pixel, the probability belonging to salient object/background is evaluated using kernel density estimation, and the geodesic distances to salient object and background are calculated based on the object likelihood map. By comparing the two geodesic distances, uncertain pixels are finally classified into salient object or background. Experimental results demonstrate the better segmentation performance of the proposed approach.
Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002
VCIP4
2010 Rate control algorithm based on frame complexity estimation for MVC
abstract
Rate control has not been well studied for multi-view video coding (MVC). In this paper, we propose an efficient rate control algorithm for MVC by improving the quadratic rate-distortion (R-D) model, which reasonably allocate bit-rate among views based on correlation analysis. The proposed algorithm consists of four levels for rate bits control more accurately, of which the frame layer allocates bits according to frame complexity and temporal activity. Extensive experiments show that the proposed algorithm can efficiently implement bit allocation and rate control according to coding parameters.
Tao Yan 0003, Ping An 0001, Liquan Shen, Zhaoyang Zhang 0002
VCIP4
2010 Automatic segmentation of focused objects from images with low depth of field
Zhi Liu 0003, Liquan Shen, Zhongmin Han, Zhaoyang Zhang 0002
Pattern Recognit. Lett.5
2010 Early SKIP mode decision for MVC using inter-view correlation
Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001
Signal Process. Image Commun.4
2010 Efficient SKIP Mode Detection for Coarse Grain Quality Scalable Video Coding
abstract
Scalable video coding (SVC) was recently standardized by the Joint Video Team as an extension of H.264. In SVC, a computationally expensive exhaustive mode decision is employed to select the best coding mode for each macroblock (MB), which achieves a high coding efficiency. In order to reduce computational complexity, we propose an efficient SKIP mode detection approach for coarse grain quality SVC. It makes use of the coding information of spatial neighboring MBs and the co-located MB in base layer to predict the SKIP mode MB and early terminate its mode decision procedure. Experimental results show that the proposed early SKIP mode decision approach can achieve the average computational saving about 54% with almost no loss of rate distortion (RD) performance in the enhancement layer.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
IEEE Signal Process. Lett.4
2010 View-Adaptive Motion Estimation and Disparity Estimation for Low Complexity Multiview Video Coding
abstract
The emerging international standard for multiview video coding (MVC) is an extension of H.264/advanced video coding. In the joint mode of MVC, both motion estimation (ME) and disparity estimation (DE) are included in the encoding process. This achieves the highest coding efficiency but requires a very high computational complexity. In this letter, we propose a fast ME and DE algorithm that adaptively utilizes the inter-view correlation. The coding mode complexity and the motion homogeneity of a macroblock (MB) are first analyzed according to the coding modes and motion vectors from the corresponding MBs in the neighbor views, which are located by means of global disparity vector. According to the coding mode complexity and the motion homogeneity, the proposed algorithm adjusts the search strategies for different types of MBs in order to perform a precise search according to video content. Experimental results demonstrate that the proposed algorithm can save 85% computational complexity on average, with negligible loss of coding efficiency.
Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001
IEEE Trans. Circuits Syst. Video Technol.4
2009 Fast mode decision for multiview video coding
abstract
In the draft of multi-view coding (MVC), variable size motion estimation and disparity estimation are employed to select the best coding mode for each macroblock. These techniques achieve the highest possible coding efficiency, but they result in extremely large computation complexity which obstructs MVC from practical application. This paper proposes a fast mode size decision algorithm for MVC in inter-frame coding. It makes use of the mode distribution correlation between neighbor views to deduct the executions of unnecessary modes. Experimental results show that the proposed fast mode decision algorithm reduces the computational complexity significantly with negligible coding efficiency.
Liquan Shen, Tao Yan 0003, Zhi Liu 0003, Zhaoyang Zhang 0002, Ping An 0001
ICIP4
2009 Frame-level bit allocation based on incremental PID algorithm and frame complexity estimation
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
J. Vis. Commun. Image Represent.3
2009 Selective VS-MRF-ME and intra coding in H.264 based on spatiotemporal continuity of motion field
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
Signal Process. Image Commun.3
2009 An Efficient Intermode Decision Algorithm Based on Motion Homogeneity for H.264/AVC
abstract
The latest video coding standard H.264/AVC significantly outperforms previous standards in terms of coding efficiency. H.264/AVC adopts variable block sizes ranging from 4 times 4 to 16 times 16 in inter frame coding, and achieves significant gain in coding efficiency compared to coding a macroblock (MB) using regular block size. However, this new feature causes extremely high computation complexity when rate-distortion optimization (RDO) is performed using the scheme of full mode decision. This paper presents an efficient intermode decision algorithm based on motion homogeneity evaluated on a normalized motion vector (MV) field, which is generated using MVs from motion estimation on the block size of 4 times 4. Three directional motion homogeneity measures derived from the normalized MV field are exploited to determine a subset of candidate intermodes for each MB, and unnecessary RDO calculations on other intermodes can be skipped. Experimental results demonstrate that our algorithm can reduce the entire encoding time about 40% on average, without any noticeable loss of coding efficiency.
Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2008 Fast Inter Mode Decision Using Spatial Property of Motion Field
abstract
Variable size motion estimation with multiple reference frames has been adopted by the new video coding standard H.264. It can achieve significant coding efficiency compared to coding a macroblock (MB) in regular size with single reference frame. On the other hand, it causes high computational complexity of motion estimation at the encoder. Rate distortion optimized (RDO) decision is one powerful method to choose the best coding mode among all combinations of block sizes and reference frames, but it requires extremely high computation. In this paper, a fast inter mode decision is proposed to decide best prediction mode utilizing the spatial continuity of motion field, which is generated by motion vectors from 4times4 motion estimation. Motion continuity of each MB is decided based on the motion edge map detected by the Sobel operator. Based on the motion continuity of a MB, only a small number of block sizes are selected in motion estimation and RDO computation process. Simulation results show that our algorithm can save more than 50% computational complexity, with negligible loss of coding efficiency.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
IEEE Trans. Multim.3
2007 A novel algorithm to fast mode decision with consideration about video texture in H.264
abstract
H.264 employs 7 different size block types for motion estimation that can significantly improve the coding performance compared with the previous video coding standards. However, H.264 requires extremely high computation with the R-D optimized decision since so many prediction modes are used. In this paper, a novel inter mode decision algorithm (NIMDA) is proposed that utilizes SADs of each 4X4 block and texture characteristic to reduce the candidate mode set after the 16 X16 prediction mode is tested. The simulation results show that the proposed algorithm reduces the entire encoding time by 64.52% with only negligible coding loss.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
AICCSA3
2007 Video nature considerations for multi-frame selection algorithm in H.264
abstract
H.264 allows motion estimation performing on multiple reference frames. This new feature improves the prediction accuracy of inter-coding blocks significantly. However, the coding gain comes at the cost of a much higher computational complexity. The reference software JM adopts full search scheme, and the computational complexity of motion estimation increases linearly with the number of allowed reference frames. In fact, the reduction of prediction residues is highly dependent on the nature of sequences, not on the number of searched frames. In this paper, with consideration of video nature and the available information from previous searched reference frames, an adaptive multi-frame selection algorithm (AMFSA) is proposed to speed up the matching process for multiple reference frames in the H.264 video coding system. The proposed algorithm can effectively reduce 63.6% on average.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
AICCSA3
2007 A Novel Video Object Tracking Approach Based on Kernel Density Estimation and Markov Random Field
abstract
In this paper, we propose a novel video object tracking approach based on kernel density estimation and Markov random field (MRF). The interested video objects are first segmented by the user, and a nonparametric model based on kernel density estimation is initialized for each video object and the remaining background, respectively. A temporal saliency map is also initialized for each object to memorize the temporal trajectory. Based on the probabilities evaluated on the non-parametric models, each pixel in the current frame is first classified into the corresponding video object or background using the maximum likelihood criterion. Starting from the initial classification result, a MRF model that combines spatial smoothness and temporal coherency is selectively exploited to generate more reliable video objects. The nonparametric model and the temporal saliency map for each video object are updated and propagated for the future tracking. Experimental results on several MPEG-4 test sequences demonstrate the good segmentation performance of our approach.
Zhi Liu 0003, Liquan Shen, Zhongmin Han, Zhaoyang Zhang 0002
ICIP (3)4
2007 An Adaptive and Fast H.264 Multi-Frame Selection Algorithm Based on Information from Previous Searches
abstract
The H.264 video coding standard adopts multiple reference frames for motion estimation. This new feature improves the prediction accuracy of inter-coding blocks significantly, but it results in a considerable increase in encoder complexity, mainly regarding to multi-frame selection and motion estimation. The reference software JM adopts the full search scheme, and the increased computation is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the nature of sequences, not on the number of searched frames. In this paper, we propose an adaptive and fast multi-frame selection algorithm (AFMFS) based on motion vectors and SAD information coming from previous searches to adaptively terminate the procedure of multiple reference frames selection. Compared with the full search algorithm and the flexible multi-reference frame search criterion (FMRFSC), simulation results show that the proposed algorithm can save 55.28% and 40.23% computation cost on average, respectively, while it still maintains similar coding efficiency.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
ICME3
2007 Real-time spatiotemporal segmentation of video objects in the H.264 compressed domain
Zhi Liu 0003, Zhaoyang Zhang 0002
J. Vis. Commun. Image Represent.3
2007 An adaptive and fast fractional pixel search algorithm in H.264
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003, Wenjun Zhang 0001
Signal Process.2
2007 An Adaptive and Fast Multiframe Selection Algorithm for H.264 Video Coding
abstract
H.264 allows motion estimation performing on multiple reference frames. This new feature improves the prediction accuracy of inter-coding blocks significantly, but it is extremely computational intensive. The reference software JM adopts full search scheme, and the increased computation load is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the content of sequences, not on the number of searched frames. In this letter, we propose an adaptive and fast multiframe selection algorithm (AFMFSA) to speed up the searching procedure for multiple reference frames. Simulation results show that the proposed algorithm can deduct 56.0%–74.2% computation load of motion estimation on average.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
IEEE Signal Process. Lett.3
2004 Intermediate View Synthesis from Stereoscopic Videoconference Images
Chaohui Lu, Ping An 0001, Zhaoyang Zhang 0002
ICCSA (4)3