Yue Wang 0032

dblp:33/4822-32 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
10since 2021 · last 2023
0009-0004-3451-5704ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2023 QUTY: Towards Better Understanding and Optimization of Short Video Quality
abstract
Short video applications such as TikTok and Instagram have attracted tremendous attention recently. However, it is very limited for industry and academia to understand the user's Quality of Experience (QoE) on short video, let alone how to improve the QoE in short video streaming.
Haodan Zhang, Yixuan Ban, Zongming Guo, Zhimin Xu 0001, Yue Wang 0032, Xinggong Zhang
MMSys6
2023 User-Generated Video Quality Assessment: A Subjective and Objective Study
abstract
Recently, we have observed an exponential increase of user-generated content (UGC) videos. The distinguished characteristic of UGC videos originates from the video production and delivery chain, as they are usually acquired and processed by non-professional users before uploading to the hosting platforms for sharing. As such, these videos usually undergo multiple distortion stages that may affect visual quality before ultimately being viewed. Inspired by the increasing consensus that the optimization of the video coding and processing shall be fully driven by the perceptual quality, in this paper, we propose to study the quality of the UGC videos from both objective and subjective perspectives. We first construct a UGC video quality assessment (VQA) database, aiming to provide useful guidance for the UGC video coding and processing in the hosting platform. The database contains source UGC videos uploaded to the platform and their transcoded versions that are ultimately enjoyed by end-users, along with their subjective scores. Furthermore, we develop an objective quality assessment algorithm that automatically evaluates the quality of the transcoded videos based on the corrupted reference, which is in accordance with the application scenarios of UGC video sharing in the hosting platforms. The information from the corrupted reference is well leveraged and the quality is predicted based on the inferred quality maps with deep neural networks (DNN). Experimental results show that the proposed method yields superior performance. Both subjective and objective evaluations of the UGC videos also shed lights on the design of perceptual UGC video coding.
Yang Li 0153, Shengbin Meng, Xinfeng Zhang 0001, Meng Wang 0017, Shiqi Wang 0001, Yue Wang 0032, Siwei Ma 0001
IEEE Trans. Multim.6
2022 Invertible Single Image Rescaling via Steganography
abstract
High-resolution (HR) images are typically downscaled by the bicubic method to fit the different resolutions of target devices and save the transmission bandwidth. However, the bicubic downsampling may cause the loss of high-frequency information (HFI), raising great challenges for recovering the original high-resolution images. Inspired by image steganography, wherein certain information needs to be written into an image as steganographic information and recovered later, in this work, we propose an invertible steganography rescaling network (ISRN) for single image rescaling. The proposed ISRN can preserve HFI that would otherwise be lost as stegano-graphic information in low-resolution (LR) images. Moreover, we develop an attentive steganography separate block (ASSB) to decompose the steganographic information into a case-aware part that can be embedded into the downscaled image and a case-agnostic part that can be captured using a prior distribution. Extensive experiments demonstrate that the proposed ISRN outperforms prior arts in terms of both quantitative and qualitative evaluations.
Mengxi Guo, Shijie Zhao 0001, Yue Li 0015, Li Zhang 0006, Yue Wang 0032
ICME6
2022 DIG: A Data-Driven Impact-Based Grouping Method for Video Rebuffering Optimization
Shengbin Meng, Chunyu Qiao, Yue Wang 0032, Zongming Guo
MMM (2)4
2022 End-to-end Image Compression with Swin-Transformer
abstract
In this paper, we propose an end-to-end image compression framework, which cooperates with the swin-transformer modules to capture the localized and non-localized similarities in image compression. In particular, the swin-transformer modules are deployed in the analysis and synthesis stages, interleaving with convolution layers. The transformer layers are expected to perceive more flexible receptive fields, such that the spatially localized and non-localized redundancies could be more effectively eliminated. The proposed method reveals the excellent capability of signal conjunction and prediction, leading to the improvement of the rate and distortion performance. Experimental results show that the proposed method is superior to the existing methods on both natural scene and screen content images, where 22.46% BD-Rate savings are achieved when compared with the BPG. Over 30% BD-Rate gains could be observed with screen content images when compared with the classical hyper-prior end-to-end coding method.
Meng Wang 0017, Kai Zhang 0007, Li Zhang 0006, Yue Li 0015, Yue Wang 0032, Shiqi Wang 0001
VCIP6
2021 Adaptive Dual Tree Structure For Screen Content Coding
abstract
The quad-tree plus binary-tree (QTBT) partition structures was adopted into the versatile video coding (VVC) standard. In the QTBT partition structure, a dual tree structure can be applied on intra slices where the partitioning structures for luminance and chrominance components are separate. The dual tree structure is efficient for camera captured videos. However, it may be not the case for screen content videos. Therefore, adaptive dual tree structure is proposed in this paper wherein the coding structure of each coding tree unit is switched between separate and joint coding structure to adapt the textures adaptively. The proposed scheme is implemented into the reference software of the VVC standard-VTM. Simulation results report that up to 4.5% BD bitrate saving can achieve on the typical screen content videos when compared to VTM5.
Weijia Zhu, Jizheng Xu, Li Zhang 0006, Yue Wang 0032
ICASSP4
2021 Implicit Seleted Transform Skip Method For Avs3
abstract
AVS3 is an emerging video coding standard, and screen content coding is a very important feature of AVS3. This paper presents a method of Implicit-Selected Transform Skip (ISTS) to further improve the screen content coding performance. With ISTS, transform skip mode is introduced as an optional substitution to transform-coding on residual signals of blocks with intra-prediction. The indication of whether to apply transform skip is hidden into the Parity of the Number of Non-zero Coefficients (PNNC) of a residual block, instead of being signaled to the decoder. Moreover, the coefficients of an intra-coded block are reordered to make the coefficients more compact. Experimental results show that the proposed method can achieve 12.04%, 8.15% and 10.19% BD-rate savings on average under All Intra (AI), Low Delay (LD) and Random Access (RA) configurations, respectively, with the encoding time reduced by 5% to 8%. ISTS has been adopted into AVS3.
Yuhuai Zhang, Kai Zhang 0007, Li Zhang 0006, Hongbin Liu 0004, Yue Wang 0032, Siwei Ma 0001, Wen Gao 0001
ICIP5
2021 Quality Assessment of End-to-End Learned Image Compression: The Benchmark and Objective Measure
abstract
Recently, learning-based lossy image compression has achieved notable breakthroughs with their excellent modeling and representation learning capabilities. Comparing to traditional image codecs based on block partitioning and transform, these data-driven approaches with artificial-neural-network (ANN) structures bring significantly different distortion patterns. Efficient objective image quality assessment (IQA) measures play the key role in quantitative evaluation and optimization of image compression algorithms. In this paper, we construct a large-scale image database for quality assessment of compressed images. In the proposed database, 100 reference images are compressed to different quality levels by 10 codecs, involving both traditional and learning-based codecs. Based on this database, we present a benchmark for existing IQA methods and reveal the challenges of IQA on learning-based compression distortions. Furthermore, we develop an objective quality assessment framework in which a self-attention module is adopted to leverage multi-level features from reference and compressed images. Extensive experiments demonstrate the superiority of our method in terms of prediction accuracy. The subjective and objective study of various compressed images also shed lights on the optimization of image compression methods.
Yang Li 0153, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Yue Wang 0032
ACM Multimedia6
2021 Low Complexity Trellis-Coded Quantization in Versatile Video Coding
abstract
The forthcoming Versatile Video Coding (VVC) standard adopts the trellis-coded quantization, which leverages the delicate trellis graph to map the quantization candidates within one block into the optimal path. Despite the high compression efficiency, the complex trellis search with soft-decision quantization may hinder the applications due to high complexity and low throughput capacity. To reduce the complexity, in this paper, we propose a low complexity trellis-coded quantization scheme in a scientifically sound way with theoretical modeling of the rate and distortion. As such, the trellis departure point can be adaptively adjusted, and unnecessarily visited branches are accordingly pruned, leading to the shrink of total trellis stages and simplification of transition branches. Extensive experimental results on the VVC test model show that the proposed scheme is effective in reducing the encoding complexity by 11% and 5% with all intra and random access configurations, respectively, at the cost of only 0.11% and 0.05% BD-Rate increase. Meanwhile, on average 24% and 27% quantization time savings can be achieved under all intra and random access configurations. Due to the excellent performance, the VVC test model has adopted one implementation of the proposed scheme.
Meng Wang 0017, Shiqi Wang 0001, Li Zhang 0006, Yue Wang 0032, Siwei Ma 0001, Sam Kwong
IEEE Trans. Image Process.5
2021 Band Representation-Based Semi-Supervised Low-Light Image Enhancement: Bridging the Gap Between Signal Fidelity and Perceptual Quality
abstract
It has been widely acknowledged that under-exposure causes a variety of visual quality degradation because of intensive noise, decreased visibility, biased color, etc. To alleviate these issues, a novel semi-supervised learning approach is proposed in this paper for low-light image enhancement. More specifically, we propose a deep recursive band network (DRBN) to recover a linear band representation of an enhanced normal-light image based on the guidance of the paired low/normal-light images. Such design philosophy enables the principled network to generate a quality improved one by reconstructing the given bands based upon another learnable linear transformation which is perceptually driven by an image quality assessment neural network. On one hand, the proposed network is delicately developed to obtain a variety of coarse-to-fine band representations, of which the estimations benefit each other in a recursive process mutually. On the other hand, the extracted band representation of the enhanced image in the recursive band learning stage of DRBN is capable of bridging the gap between the restoration knowledge of paired data and the perceptual quality preference to high-quality images. Subsequently, the band recomposition learns to recompose the band representation towards fitting perceptual regularization of high-quality images with the perceptual guidance. The proposed architecture can be flexibly trained with both paired and unpaired data. Extensive experiments demonstrate that our method produces better enhanced results with visually pleasing contrast and color distributions, as well as well-restored structural details.
Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001
IEEE Trans. Image Process.4
2020 Consistent Video Style Transfer via Compound Regularization
abstract
Recently, neural style transfer has drawn many attentions and significant progresses have been made, especially for image style transfer. However, flexible and consistent style transfer for videos remains a challenging problem. Existing training strategies, either using a significant amount of video data with optical flows or introducing single-frame regularizers, have limited performance on real videos. In this paper, we propose a novel interpretation of temporal consistency, based on which we analyze the drawbacks of existing training strategies; and then derive a new compound regularization. Experimental results show that the proposed regularization can better balance the spatial and temporal performance, which supports our modeling. Combining with the new cost formula, we design a zero-shot video style transfer framework. Moreover, for better feature migration, we introduce a new module to dynamically adjust inter-channel distributions. Quantitative and qualitative results demonstrate the superiority of our method over other state-of-the-art style transfer methods. Our project is publicly available at: https://daooshee.github.io/CompoundVST/.
Wenjing Wang 0001, Jizheng Xu, Li Zhang 0006, Yue Wang 0032, Jiaying Liu 0001
AAAI4
2020 From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement
abstract
Under-exposure introduces a series of visual degradation, i.e. decreased visibility, intensive noise, and biased color, etc. To address these problems, we propose a novel semi-supervised learning approach for low-light image enhancement. A deep recursive band network (DRBN) is proposed to recover a linear band representation of an enhanced normal-light image with paired low/normal-light images, and then obtain an improved one by recomposing the given bands via another learnable linear transformation based on a perceptual quality-driven adversarial learning with unpaired data. The architecture is powerful and flexible to have the merit of training with both paired and unpaired data. On one hand, the proposed network is well designed to extract a series of coarse-to-fine band representations, whose estimations are mutually beneficial in a recursive process. On the other hand, the extracted band representation of the enhanced image in the first stage of DRBN (recursive band learning) bridges the gap between the restoration knowledge of paired data and the perceptual quality preference to real high-quality images. Its second stage (band recomposition) learns to recompose the band representation towards fitting perceptual properties of high-quality images via adversarial learning. With the help of this two-stage design, our approach generates enhanced results with well-reconstructed details and visually promising contrast and color distributions. Qualitative and quantitative evaluations demonstrate the superiority of our DRBN.
Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001
CVPR4
2020 Fixed-Length Coding for Escape Samples in Palette Mode
abstract
Palette mode is a powerful tool for screen content coding in the upcoming versatile video coding (VVC) standard. In the palette mode, escape samples are employed to handle an outlier case. In this paper, a quantization parameter (QP) dependent fixed-length binarization is proposed for escape samples coding to 1) simplify the design and 2) improve coding efficiency. The length is calculated according to the QP for the current block and correspondingly, the dequantization process can be also implemented by only left shifting. The proposed method is evaluated with VVC reference software VTM-6.0 on typical sequences containing "text and graphics with motion". Experimental results report that the proposed scheme can achieve up to 2.0% BD-rate savings compared to VTM-6.0 with a much simpler design.
Weijia Zhu, Jizheng Xu, Li Zhang 0006, Yue Wang 0032
DCC4
2020 Learning to Fool the Speaker Recognition
abstract
Due to the widespread deployment of fingerprint/face/speaker recognition systems, attacking deep learning based biometric systems has drawn more and more attention. Previous research mainly studied the attack to the vision-based system, such as fingerprint and face recognition. While the attack for speaker recognition has not been investigated yet, although it has been widely used in our daily life. In this paper, we attempt to fool the state-of-the-art speaker recognition model and present speaker recognition attacker, a lightweight model to fool the deep speaker recognition model by adding imperceptible perturbations onto the raw speech waveform. We find that the speaker recognition system is also vulnerable to the attack, and we achieve a high success rate on the non-targeted attack. Besides, we also present an effective method to optimize the speaker recognition attacker to obtain a trade-off between the attack success rate with the perceptual quality. Experiments on the TIMIT dataset show that we can achieve a sentence error rate of 99.2% with an average SNR 57.2dB and PESQ 4.2 with speed rather faster than real-time.
Jiguo Li 0002, Xinfeng Zhang 0001, Jizheng Xu, Li Zhang 0006, Yue Wang 0032, Siwei Ma 0001, Wen Gao 0001
ICASSP5
2020 Universal Adversarial Perturbations Generative Network For Speaker Recognition
abstract
Attacking deep learning based biometric systems has drawn more and more attention with the wide deployment of fingerprint/face/speaker recognition systems, given the fact that the neural networks are vulnerable to the adversarial examples, which have been intentionally perturbed to remain almost imperceptible for human. In this paper, we demonstrated the existence of the universal adversarial perturbations (UAPs) for the speaker recognition systems. We proposed a generative network to learn the mapping from the low-dimensional normal distribution to the UAPs subspace, then synthesize the UAPs to perturbe any input signals to spoof the well-trained speaker recognition model with high probability. Experimental results on TIMIT and LibriSpeech datasets demonstrate the effectiveness of our model.
Jiguo Li 0002, Xinfeng Zhang 0001, Chuanmin Jia, Jizheng Xu, Li Zhang 0006, Yue Wang 0032, Siwei Ma 0001, Wen Gao 0001
ICME6
2020 APL: Adaptive Preloading of Short Video with Lyapunov Optimization
abstract
Short video applications, like TikTok, have attracted many users across the world. It can feed short videos based on users' preferences and allow users to slide the boring content anywhere and anytime. To reduce the loading time and keep playback smoothness, most of the short video apps will preload the recommended short videos in advance. However, these apps preload short videos in fixed size and fixed order, which can lead to huge playback stall and huge bandwidth waste. To deal with these problems, we present an Adaptive Preloading mechanism for short videos based on Lyapunov Optimization, also called APL, to achieve near-optimal playback experience, i.e., maximizing playback smoothness and minimizing bandwidth waste considering users' sliding behaviors. Specifically, we make three technical contributions: (1) We design a novel short video streaming framework which can dynamically preload the recommended short videos before the current video is downloaded completely. (2) We formulate the preloading problem into a playback experience optimization problem to maximize the playback smoothness and minimize the bandwidth waste. (3) We transform the playback experience optimization problem during the whole viewing process into a single-step greedy algorithm based on the Lyapunov optimization theory to make the online decisions during playback. Through extensive experiments based on the real datasets that generously provided by TikTok, we demonstrate that APL can reduce the stall ratio by 81%/12% and bandwidth waste by 11%/31% compared with no-preloading/fixed-preloading mechanism.
Haodan Zhang, Yixuan Ban, Xinggong Zhang, Zongming Guo, Zhimin Xu 0001, Shengbin Meng, Yue Wang 0032
VCIP8
2020 Interweaved Prediction for Video Coding
abstract
In the emerging next generation video coding standard Versatile Video Coding (VVC) developed by the Joint Video Exploration Team (JVET), sub-block-based inter-prediction plays a key role in promising coding tools such as Affine Motion Compensation (AMC) and sub-block-based Temporal Motion Vector Prediction (sbTMVP). With sub-block-based inter-prediction, a coding block is divided into sub-blocks, and the motion information of each sub-block is derived individually. Although sub-block-based inter-prediction can provide a higher quality prediction benefiting from a finer motion granularity, it still suffers two problems: uneven prediction quality and boundary discontinuity. In this paper, we present a method of interweaved prediction to further improve sub-block-based inter-prediction. With interweaved prediction, a coding block with AMC or sbTMVP mode is divided into sub-blocks with two different dividing patterns, so that a corner position of a sub-block in one dividing pattern coincides with the central position of a sub-block in the other dividing pattern. Then two auxiliary predictions are generated by AMC or sbTMVP with the two dividing patterns, independently. The final prediction is calculated as a weighted-sum of the two auxiliary predictions. Theoretical analysis and statistical data prove that interweaved prediction can significantly mitigate the two problems in sub-block-based inter-prediction. Simulation results show that the proposed methods can achieve 0.64% BD-rate saving on average with the random access configurations. On sequences with rich affine motions, the average BD-rate saving can be up to 2.54%.
Kai Zhang 0007, Li Zhang 0006, Hongbin Liu 0004, Jizheng Xu, Zhipin Deng, Yue Wang 0032
IEEE Trans. Image Process.6
2019 History-Based Motion Vector Prediction in Versatile Video Coding
abstract
In this paper, History-based Motion Vector Prediction (HMVP) is presented for video coding. With the proposed method, a table of HMVP candidates is maintained and updated on-the-fly. After decoding one inter-coded block, the table is updated by appending the associated motion information to the table as a new HMVP candidate. A First-In-First-Out (FIFO) rule is applied to manage the table. The HMVP candidates could be added to the Advanced Motion Vector Prediction (AMVP) candidate list as additional motion vector predictors. And they could also be added to the merge candidate list as additional merge candidates. With the proposed method, the motion information of previously coded blocks even not adjacent to the current block can be utilized for more efficient motion vector prediction. Simulation results have validated the efficiency of HMVP, wherein up to 4% BD rate saving could be achieved. The proposed method has been adopted by the next generation video coding standard, named Versatile Video Coding (VVC) developed by Joint Video Exploration Team (JVET).
Li Zhang 0006, Kai Zhang 0007, Hongbin Liu 0004, Hsiao-Chiang Chuang, Yue Wang 0032, Ji-Zheng Xu, Pengwei Zhao, Dingkun Hong
DCC5
2019 Interweaved Prediction for Affine Motion Compensation
abstract
With affine motion compensation (AMC) in the emerging next generation video coding standard Versatile Video Coding (VVC), a coding-block is divided into sub-blocks, and each sub-block is assigned with an individual motion vector derived by the affine model. The sub-block-based design for AMC faces a dilemma. With smaller sub-blocks, AMC can achieve a better coding performance but suffers a higher complexity burden. In this contribution, an interweaved prediction approach is proposed for AMC to address the dilemma. With the interweaved prediction, a coding block is divided into sub-blocks with two different dividing patterns. Then two auxiliary predictions are generated by AMC with the two dividing patterns respectively. The final prediction is calculated as a weighted-sum of the two auxiliary predictions. The interweaved prediction is only applied on the luma-component for affine-coded blocks with uni-prediction. Simulation results show 0.53% Bjøntegaard-Delta rate savings in average can be achieved compared to VTM-3.0 under random access configurations. The coding gain on sequences with affine motions is up to 3.3%.
Kai Zhang 0007, Li Zhang 0006, Hongbin Liu 0004, Ji-Zheng Xu, Yue Wang 0032
ICIP5
2019 Coding Prior Based High Efficiency Restoration for Compressed Video
abstract
Lossy compression introduces complex compression artifacts such as the blocking, ringing and blurring artifacts, making decoded videos unpleasant for human visual system. In this paper, we propose a coding prior based high efficiency restoration algorithm to remove these compression artifacts. To improve the quality of the restored compressed videos, we take full advantage of side information from coding streams as coding prior which is ignored or not fully exploited by most existing post-processing methods. In particular, the unfiltered frames and the prediction frames are derived from coding streams which are utilized as coding priors and a high efficiency neural network is designed for these information to improve the overall quality of restored videos. Extensive experimental results on HEVC coding streams demonstrate that our proposed method can significantly improve both the objective and subjective quality of compressed videos.
Longtao Feng, Xinfeng Zhang 0001, Shanshe Wang, Yue Wang 0032, Siwei Ma 0001
ICIP4
2019 Two-Pass Bi-Directional Optical Flow Via Motion Vector Refinement
abstract
Bi-directional optical flow (BDOF) is an efficient coding tool that has been recently adopted into Versatile Video Coding (VVC) standard. With BDOF, bi-predictive prediction samples of one coding block are enhanced via higher-precision motion vectors (MVs) derived from its two reference blocks. In this way, the energy of prediction error could be reduced, resulting in better coding performance. In VVC, the derived motion information is only used to enhance prediction samples. In this paper, it is proposed to use the derived motion information to also refine decoded MVs. The refined MVs may be used as spatial motion vector prediction (MVP) for the following coding units (CUs), as the temporal MVP for the subsequent pictures, and in the deblocking filtering process. Furthermore, the refined MVs can be used to perform motion compensation (MC) again to further improve the quality of the prediction samples. Simulation results show that the proposed methods can achieve -1.18% BD-rate saving in average under the random access configuration on top of the existing BDOF design in VVC.
Hongbin Liu 0004, Li Zhang 0006, Kai Zhang 0007, Hsiao-Chiang Chuang, Yue Wang 0032, Jizheng Xu
ICIP5
2019 Adaptive Motion Vector Resolution for Affine-Inter Mode Coding
abstract
Affine Motion Model (AMM) based inter prediction, which can represent complex motions such as zooming, rotation or shearing, has been adopted into the Versatile Video Coding (VVC) standard. AMM is defined by Control Point Motion Vectors (CPMVs) in VVC. On the other hand, Adaptive Motion Vector Resolution (AMVR) has also been adopted into VVC standard due to a favorable trade-off between the Motion Vector (MV) precision and the bit consumption on MV Differences (MVDs). However, AMVR is only applied to the Translational Motion Model (TMM), and AMM cannot benefit from it. In this paper, it is proposed to extend AMVR to AMM. Specifically, 1-pixel, 1/4-pixel and 1/16-pixel MV precisions are allowed and can be selected adaptively by each affine-inter mode coded Coding Unit (CU). Simulation results reportedly show that the proposed method can achieve 0.32% BD-rate saving on average under the Random Access configuration.
Hongbin Liu 0004, Li Zhang 0006, Kai Zhang 0007, Jizheng Xu, Yue Wang 0032, Jiancong Luo, Yuwen He
PCS5
2019 Compound Palette Mode for Screen Content Coding
abstract
The Joint Video Exploration Team (JVET) has been developing an emerging standard Versatile Video Coding (VVC), which includes screen contents as one of its requirements. Intra block copy (IBC) and palette coding are the two powerful coding tools for screen content coding. In this paper, a compound palette mode is proposed to exploit the advantages of both IBC and palette coding, which allows samples to be reconstructed by either IBC predictions or palette entries. The proposed method is evaluated with VVC reference software VTM4 on typical sequences containing "text and graphics with motion". Experimental results report significant coding gain that the proposed scheme can achieve 7.80% and 1.03% BD-rate savings under AI conditions on average when compared with VTM4 and the existing palette scheme.
Weijia Zhu, Jizheng Xu, Li Zhang 0006, Kai Zhang 0007, Hongbin Liu 0004, Yue Wang 0032
PCS6
2018 CUB360: Exploiting Cross-Users Behaviors for Viewport Prediction in 360 Video Adaptive Streaming
abstract
To ensure 360-degree video's continuous playback and reduce the bandwidth waste, predicting user's future fixation is indispensable. However, existing methods concentrate either on user's motion information or content information. None of them consider users watching behaviors' inconsistency which embodies user's attention distribution more explicitly. So in this paper, we exploit Cross-Users Behaviors for viewport prediction in 360-degree video adaptive streaming, namely CUB360, trying to concurrently consider user's personalized information and cross-users behaviors information to predict future viewport. Besides, we use a QoE-driven framework to optimize existing video streaming approaches and propose a general algorithm aiming at solving the NP problem at a low complexity. Extensive experimental results over real datasets demonstrate that compared with traditional adaptive streaming method, our proposal can significantly boost the prediction accuracy by 20.2% absolutely and 48.1 % relatively. Besides, the mean quality can get 30.28% gain while quality variance can be reduced by 29.89%.
Yixuan Ban, Lan Xie, Zhimin Xu 0001, Xinggong Zhang, Zongming Guo, Yue Wang 0032
ICME6
2012 Spatio-temporal ssim index for video quality assessment
abstract
An ideal objective metric for video quality assessment (VQA) should achieve consistency between video distortion prediction and psychological perception of human visual system (HVS), and is important in many video processing applications. In general, both spatial distortion and temporal distortion should be carefully considered in the designing of VQA metrics. In this paper, we propose a novel spatio-temporal structural information based video quality metric. Motivated by the fact that pixels in natural videos are highly structured in both spatial domain and temporal domain, we propose to perform structural similarity evaluation in x-y, x-t and y-t dimensions respectively and pooled them adaptively based on local spatio-temporal activities. Experimental results on LIVE database show that such a conceptually simple and computationally efficient algorithm is competitive with state-of-the-art VQA metrics, and is very robust to various types of video distortions.
Yue Wang 0032, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001
VCIP1
2012 Novel Spatio-Temporal Structural Information Based Video Quality Metric
abstract
Video quality assessment (VQA) is very important for many video processing applications, e.g., compression, archiving, restoration, and enhancement. An ideal video quality metric should achieve consistency between video distortion prediction and psychological perception of human visual system. Different from the quality assessment of single images, motion information and temporal distortion should be carefully considered for VQA. Most of previous VQA algorithms deal with the motion information through two ways: either incorporating motion characteristics into a temporal weighting scheme to account for their affects on the spatial distortion, or modeling the temporal distortion and spatial distortion independently. Optical flows need to be estimated in the two ways. In this paper, we propose a different methodology to deal with the motion information. Instead of explicitly calculating the optical flow and independently modeling the temporal distortion, both the spatial edge features and temporal motion characteristics are accounted for by some structural features in the localized spacetime regions. We propose to represent the structural information by two descriptors extracted from the 3-D structure tensors, which are the largest eigenvalue as well as its corresponding eigenvector. Experimental results on LIVE database and VQEG FR-TV Phase-I database show that the proposed VQA metric is competitive with state-of-the-art VQA metrics, while keeping relatively low computing complexity.
Yue Wang 0032, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2011 Joint just noticeable difference model based on depth perception for stereoscopic images
abstract
Just noticeable difference (JND) model can reflect the least perceptible distortion from images, including 2D images and stereoscopic images. As we know, for the perception of human visual system (HVS), stereoscopic images have quite different characteristics from 2D images, since stereoscopic images contain not only planar information, but also depth information. This paper proposes a joint JND (JJND) model based on depth perception for stereoscopic images. Firstly, disparity estimation is performed in order to decompose the image into the occlusion region and the non-overlapped region. Then, different JND thresholds are applied on different regions, according to the depth information of the region, which can be derived from the disparity of the region. Experimental results verified our model's validity for stereoscopic images.
Xiaoming Li 0002, Yue Wang 0032, Debin Zhao, Tingting Jiang 0001, Nan Zhang 0015
VCIP2
2011 Advanced spatial and Temporal Direct Mode for B picture coding
abstract
The direct mode in H.264/AVC can efficiently improve the coding performance of B pictures, since it exploits the spatial or temporal correlation by deriving its motion vector from previously encoded information. Therefore, it does not require any additional motion information and could save many bits. Considering the spatial and temporal correlation has not been fully exploited in the current direct mode, in this paper, we propose an advanced Spatial and Temporal Direct Mode (STDM) for B picture coding. The motion vector is selected from a set of spatial-temporal neighboring motion vectors, and the selection criterion is to minimize a spatial-temporal cost function. The framework of Decoder-side Motion Vector Derivation (DMVD) is utilized, where encoder and decoder use the same derivation process to obtain motion vectors, thus no index for the chosen motion vectors need to be coded and transmitted. Simulation results show that the proposed method significantly outperforms the current SDM and TDM in H.264/AVC.
Yue Wang 0032, Li Zhang 0006, Siwei Ma 0001, Wen Gao 0001
VCIP1
2010 Image quality assessment based on local orientation distributions
abstract
Image quality assessment (IQA) is very important for many image and video processing applications, e.g. compression, archiving, restoration and enhancement. An ideal image quality metric should achieve consistency between image distortion prediction and psychological perception of human visual system (HVS). Inspired by that HVS is quite sensitive to image local orientation features, in this paper, we propose a new structural information based image quality metric, which evaluates image distortion by computing the distance of Histograms of Oriented Gradients (HOG) descriptors. Experimental results on LIVE database show that the proposed IQA metric is competitive with state-of-the-art IQA metrics, while keeping relatively low computing complexity.
Yue Wang 0032, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001
PCS1
2010 Frame rate up conversion via Bayesian motion estimation
abstract
In this paper, a novel block-based motion compensated frame interpolation (MCI) algorithm is proposed to enhance the temporal resolution of video sequences. We formulated motion estimation into MAP framework, and solved it via Bayesian belief propagation. By effectively incorporating a priori knowledge of the motion field and optimizing the whole motion field synchronously, it could derive more accurate motion vectors than traditional methods. Finally, adaptive overlapped block motion compensation (OBMC) is used to reduce blocking artifacts. Experimental results show that the proposed method outperforms other methods in both objective and subjective quality.
Yue Wang 0032, Siwei Ma 0001, Wen Gao 0001
VCIP1